Graphic content generation method, device, equipment and storage medium
By obtaining key information of text information and using AI text and picture generation model to generate matching pictures, the high cost and quality limitations caused by relying on paper books in the prior art are solved, and high-quality picture and text content generation is achieved.
Patent Information
- Application Number
- CN202311030903.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-14
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2043-08-14
AI Technical Summary
The existing process of generating graphic and text content relies on paper books, resulting in high generation costs and limited quality due to the quality of paper books.
By obtaining key information in text information, calling the preconfigured text and picture generation model to generate matching pictures, and integrating text information to generate picture content, using the ability of AI text and picture generation model to automatically generate pictures that match text information.
It realizes high-quality graphic and text content generation, reduces generation costs, and ensures the semantic matching and diversity of generated pictures and text information.
Smart Images

Figure CN117032869B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of text and image generation, and more specifically, to a method, device, equipment and storage medium for generating text and image content. Background Art
[0002] Graphic content refers to electronic content that combines text and illustrations. It can be reading materials like e-books or electronic content like interview transcripts. Common examples of graphic content include children's picture books and illustrated novels. The combination of text and images can enhance readers' interest in graphic content. For example, children's picture books, which often contain more illustrations and a small amount of text, can effectively promote children's language development and cultivate their interest in reading.
[0003] Currently, the production of graphic content is usually based on the production process of converting paper books into digital data through collection, processing and manipulation. This method of generating graphic content is costly and limited by the quality of paper books. When the image quality in paper books is low, the quality of the generated electronic graphic content is also not high. Summary of the Invention
[0004] In view of the above problems, this application is proposed to provide a method, device, equipment and storage medium for generating graphic content, so as to solve the problems that the existing graphic content generation process relies on paper books, resulting in high generation costs and quality limited by paper books. The specific solution is as follows:
[0005] In a first aspect, a method for generating graphic content is provided, comprising:
[0006] Get the text information of the image to be generated;
[0007] Obtaining key information from the text information;
[0008] Calling a preconfigured text-image generation model to generate an image based on the key information;
[0009] The text information and the image are integrated to obtain generated graphic content.
[0010] Preferably, the obtaining of text information for the image to be generated includes:
[0011] Obtain key element information input by the user;
[0012] A preconfigured artificial intelligence model is called to generate the text information based on the key element information.
[0013] Preferably, obtaining key information from the text information includes:
[0014] Dividing the text information according to the set division granularity to obtain each part of the divided text information;
[0015] Get the key information of each part of text information separately.
[0016] Preferably, the step of separately obtaining key information of each portion of text information includes:
[0017] For each part of text information, a preconfigured artificial intelligence model is called to generate key information corresponding to each part of text information.
[0018] Preferably, calling a preconfigured text-image generation model to generate an image based on the key information includes:
[0019] For each part of text information, analyze the target role contained therein;
[0020] The configured personalized text-image generation model corresponding to the target character is called to generate an image based on key information corresponding to the partial text information, wherein the personalized text-image generation model corresponding to the target character supports the ability to generate different images so that the identity of the corresponding image of the target character in different images is consistent.
[0021] Preferably, before calling the preconfigured text-image generation model to generate the image based on the key information, the method further includes:
[0022] Analyzing each character information contained in the text information;
[0023] For each character, a personalized text-image generation model corresponding to the character is configured. The personalized text-image generation model corresponding to the character supports the ability to generate different images so that the identity of the corresponding image of the character in different images is consistent.
[0024] Preferably, if there is only one target role, the process of calling the configured personalized text-image generation model corresponding to the target role and generating an image based on the key information corresponding to the partial text information includes:
[0025] A personalized text-image generation model corresponding to the target character is called to generate an image based on key information corresponding to the partial text information.
[0026] Preferably, if there are more than two target roles, the process of calling the configured personalized text-image generation model corresponding to the target role and generating an image based on the key information corresponding to the partial text information includes:
[0027] Selecting any one of the target characters, calling a personalized text-image generation model corresponding to the selected target character, and generating a first image based on key information corresponding to the partial text information;
[0028] traversing and selecting the target characters from the target characters that have not been selected, each time after selecting the target character, performing an occlusion process on the image corresponding to the currently selected target character in the first picture to obtain a second picture, calling the personalized text-image generation model corresponding to the currently selected target character, and generating a new first picture again based on the key information corresponding to the partial text information and the second picture;
[0029] After traversing the target character, the first picture generated last time is used as the final picture.
[0030] Preferably, for each role, the process of configuring a personalized text-image generation model corresponding to the role includes:
[0031] Add a text feature extractor to the general text-image generation model to obtain the edited text-image generation model;
[0032] Obtaining a training data set corresponding to each character, wherein the training data set includes a plurality of training data pairs consisting of image samples and their description texts, wherein the identity of the image corresponding to the same character in different image samples is consistent;
[0033] For each role: the descriptive text of the image sample in the corresponding training data set is input into the edited text-image generation model, and the general text-image generation model in the edited text-image generation model is used to process the descriptive text, and the learnable text vector of the text label corresponding to the role is input into the text feature extractor to extract text features, and the text features are used to offset the features extracted by the hidden layer of the general text-image generation model, and the decoding module of the general text-image generation model decodes the offset-processed features to obtain a generated image, and the model parameters are updated with the goal of making the generated image close to the image sample to obtain a personalized text-image generation model corresponding to the role.
[0034] Preferably, before fusing the text information and the picture, the method further includes:
[0035] The generated image is converted into a specified style, where the specified style is a preset unified style, or the specified style is a style that matches the text information.
[0036] In a second aspect, a device for generating graphic content is provided, comprising:
[0037] A text acquisition unit, used to acquire text information for generating a picture;
[0038] A key information acquisition unit, configured to acquire key information from the text information;
[0039] An image generation unit, configured to call a preconfigured text-image generation model to generate an image based on the key information;
[0040] The text-image fusion unit is used to fuse the text information and the image to obtain generated image-text content.
[0041] In a third aspect, a graphic content generation device is provided, comprising: a memory and a processor;
[0042] The memory is used to store programs;
[0043] The processor is used to execute the program to implement each step of the aforementioned graphic content generation method.
[0044] In a fourth aspect, a storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the graphic content generation method as described above are implemented.
[0045] By means of the above technical solution, the present application first obtains the text information that needs to generate the illustrations. In order to generate the illustrations, key information is obtained from the text information as reference text information when generating the illustrations. The pre-configured text-image generation model is further called, and with the help of the powerful text-image generation model's ability to generate pictures with matching semantics based on text, pictures that match the key information can be generated based on the key information obtained above, and finally the text information and the generated pictures are fused to obtain the graphic content. Obviously, the present application solution can automatically generate pictures that match the text information with the help of the ability of the AI text-image generation model, and in view of the diversity and text consistency of the pictures generated by the text-image generation model, the quality of the generated pictures can be guaranteed and they match the semantics of the text information, and finally high-quality graphic content is obtained. The whole process does not rely on paper books, and the generation cost is greatly reduced. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present application. The same reference symbols are used throughout the drawings to represent the same components. In the drawings:
[0047] Figure 1 A schematic diagram of a flow chart of a method for generating graphic content provided in an embodiment of the present application;
[0048] Figure 2 An example of an electronic picture book illustration is provided;
[0049] Figure 3 This example shows another type of illustration for an electronic picture book;
[0050] Figure 4 This example illustrates another type of illustration for an electronic picture book;
[0051] Figure 5 This diagram illustrates the training process of a personalized text-image generation model.
[0052] Figure 6 A schematic diagram of the structure of a graphic content generation device provided in an embodiment of the present application;
[0053] Figure 7 A schematic diagram of the structure of a graphic content generation device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0054] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0055] Before introducing this application plan, let me first explain the English terms involved in this article:
[0056] Prompt: Instructions. When interacting with AI (such as large artificial intelligence models), you need to send instructions to the AI. This can be a text description, such as "Please recommend me a pop song" when interacting with the AI, or it can be a parameter description in a certain format, such as asking the AI to draw a picture according to a certain format, which requires describing the relevant drawing parameters.
[0057] Large-scale AI models, also known as large-scale deep learning models, are AI models based on deep learning technology. They consist of hundreds of millions of parameters and can achieve complex tasks such as natural speech processing, image recognition, image generation, and speech recognition by learning and training on large amounts of data. Large-scale AI models can include large models and large language models. Both large models and large language models refer to machine learning models with very large parameters, but their application scenarios and focus differ slightly. The definitions of the two are provided below.
[0058] Text-image generation models: These are large models that generate images based on auxiliary information (which can be text descriptions or both text descriptions and images). Common text-image generation models include the stable diffusion model and other large models.
[0059] This application provides a graphic content generation solution that can be applied to generate various types of electronic graphic content. Graphic content in this application refers to electronic content containing text and illustrations, which can be reading materials such as e-books or electronic content such as interview records. Common graphic content includes children's picture books, novel story picture books, etc. For ease of understanding, the following embodiments are exemplified by using electronic picture books as graphic content.
[0060] The present application solution can be implemented based on a terminal with data processing capabilities, which can be a mobile phone, computer, server, cloud, etc.
[0061] Next, combine Figure 1 The method for generating graphic content of the present application may include the following steps:
[0062] Step S100: Obtain text information for generating a picture.
[0063] The graphic content consists of text and accompanying pictures. In this step, the text information can be obtained first, and then the accompanying pictures corresponding to the text information can be generated to form the graphic content. When the graphic content is an electronic picture book, the obtained text information can be the picture book text information.
[0064] The text information obtained in this step may be text information input by a user, or may be text information downloaded from the Internet, imported locally, or obtained from other terminals.
[0065] In addition, the text information obtained in this step may also be generated text information, for example, text information automatically generated with the help of an artificial intelligence model.
[0066] In some scenarios, users can input key element information and then call the artificial intelligence model to generate text information based on the key element information of the picture book. The artificial intelligence model can be a large artificial intelligence model, which can then use its super semantic understanding and text generation capabilities to generate rich text information. Taking the graphic content of an electronic picture book as an example, the key element information input by the user can be the key element information of the picture book. For example, the key element information of the picture book input by the user can be "Please use animals such as the big bad wolf, the little white rabbit, and the giant panda as the main characters to generate novel content." An example of text information generated by calling a large artificial intelligence model is as follows:
[0067] One day, a big bad wolf met a little white rabbit in the forest. The little white rabbit was very scared, but bravely asked the big bad wolf, "Why are you so fierce? Don't you want to live in peace with other animals?" The big bad wolf was surprised to hear the little white rabbit's words. It had never thought that it could live in peace with other animals. So, the big bad wolf decided to change its behavior and stop preying on other animals..."
[0068] Step S110: Acquire key information from the text information.
[0069] Specifically, reference text information needs to be provided when calling the text-image generation model. Considering that the content of the obtained text information may be long and may contain information irrelevant to the generation of illustrations, key information is extracted from the text information in this step. The key information can represent the information in the text information that is related to the generation of illustrations.
[0070] It should be noted that the key information obtained from the text information in this step can be part of the intercepted text information, or it can be new text information generated based on the text information, such as text information after refining and summarizing the text information.
[0071] Step S120: calling a pre-configured text-image generation model to generate an image based on the key information.
[0072] In this embodiment, a preconfigured text-to-image generation model can be invoked, leveraging its ability to generate images from text to generate matching images based on the key information acquired in the previous step. The text-to-image generation model invoked in this step can be any of the various existing large models capable of generating images from text, such as the stable diffusion model. Furthermore, existing text-to-image generation models can be appropriately improved, as detailed in subsequent embodiments.
[0073] Step S130: Fusing the text information and the image to obtain generated graphic content.
[0074] Specifically, the text information and the picture can be combined according to the set page layout structure of the text and the picture to obtain graphic content, such as an electronic picture book.
[0075] The method for generating graphic content provided by the embodiment of the present application first obtains the text information for which the illustrations need to be generated. In order to generate the illustrations, key information is obtained from the text information as reference text information when generating the illustrations. The pre-configured text-image generation model is further called, and with the help of the powerful ability of the text-image generation model to generate pictures with matching semantics based on text, pictures matching the key information can be generated based on the key information obtained above, and finally the text information and the generated pictures are fused to obtain graphic content. Obviously, the present application solution can automatically generate pictures matching the text information with the help of the ability of the AI text-image generation model, and in view of the diversity and text consistency of the pictures generated by the text-image generation model, the quality of the generated pictures can be guaranteed and they match the semantics of the text information, and finally graphic content with higher content quality is obtained. The whole process does not rely on paper books, and the generation cost is greatly reduced.
[0076] Further, referring to the description of step S100 above, the acquired text information can be automatically generated based on the artificial intelligence model. Therefore, the method provided in this embodiment can batch generate graphic and text content created by AI, and can provide creative reference materials to graphic and text content producers to stimulate creative inspiration.
[0077] In some embodiments of the present application, considering that common graphic content includes pictures at different granularities, for example, pictures at sentence level, paragraph level, chapter level, etc. To this end, in this embodiment, the text information can be divided according to the set division granularity to obtain the divided text information.
[0078] Here, the division granularity can be set by the user or automatically determined based on the text information. Taking the example of assigning pictures to chapters, the text information can be divided according to chapters in this embodiment to obtain the text information of each chapter after division.
[0079] In this embodiment, when dividing text information, a rule-based division method can be adopted, or the division can be performed by calling a preconfigured model, for example, calling an artificial intelligence model to instruct the model to divide the text information, etc.
[0080] After dividing the acquired text into several parts, the key information corresponding to each part of the text is obtained. The text-image generation model is then invoked to generate an image that matches each part of the text based on the key information corresponding to each part of the text. When fusing text and images, the text and image can be combined based on the matching relationship between each part of the text and the image. This can include embedding a matching image within each part of the text, or combining matching images around each part of the text, or other image-text combinations.
[0081] In this embodiment, for each part of text information, part of the information can be intercepted as key information, or an artificial intelligence model can be called to generate key information corresponding to each part of text information.
[0082] This embodiment provides an example of the result of extracting key information from a portion of text information:
[0083] The text after segmentation reads: "One day, a big bad wolf met a little white rabbit in the forest. The little white rabbit was very scared, but still bravely asked the big bad wolf, 'Why are you so fierce? Don't you want to live in peace with other animals?'" The big bad wolf was surprised by the little white rabbit's words. It had never thought that it could live in peace with other animals. So, the big bad wolf decided to change its behavior and stop preying on other animals.
[0084] For some of the above text information, examples of key information extracted are: "The big bad wolf changed from ferocious to friendly, and decided to change its behavior after meeting the little white rabbit and stop preying on other animals. He also met the giant panda and learned the importance of peace and friendship."
[0085] In some embodiments of the present application, a process of calling a text-image generation model to generate a matching image based on key information corresponding to each part of text information is introduced.
[0086] In this embodiment, considering the corresponding text Figure 1 Generally, the main characters contained in the text need to be included. For example, in an electronic picture book, the illustrations need to include the main characters contained in the corresponding text, such as the Big Bad Wolf and the Little White Rabbit. In addition, a piece of graphic content may contain multiple illustrations, and different illustrations may contain the same character. To ensure the consistency of the illustrations of the graphic content, in this embodiment, the identity of the corresponding image of the same character contained in different images within the graphic content can be set to be consistent.
[0087] Here, the identity consistency of the corresponding images of the same character in different images can be understood as the overall performance of the corresponding images of the same character in different images is consistent and cannot be destroyed, that is, readers can intuitively judge that the character images in different images belong to the same identity.
[0088] Reference Figure 2 and Figure 3 , Figure 2 Contains a turtle character. Figure 3 It contains a little white rabbit and a turtle character. However, obviously Figure 2 and Figure 3 The identity of the turtle has changed, and readers can easily see that Figure 2 and Figure 3The turtle in the picture is not a turtle with the same identity, which will affect the coherence of the reader's reading of the text and picture content.
[0089] In this embodiment, in order to avoid the above problems, when generating different images of graphic content, the identity of the corresponding image of the same character in different images is kept consistent. Figure 4 As shown, the turtle character is included, and Figure 2 The identity of the turtle's role is consistent.
[0090] Specifically, in this embodiment, the information of each character contained in the text information can be pre-analyzed, and then a personalized text-image generation model corresponding to each character can be configured. The personalized text-image generation model corresponding to each character supports the ability to generate different images, while maintaining the identity of the corresponding character in different images.
[0091] In this embodiment, the personalized text and image generation model configured for each character can not only maintain the identity of different characters in the text and image content, but also maintain the generalization ability of the model itself, that is, it can generate different types of objects such as trees, portraits, animals, science fiction, etc.
[0092] On this basis, for each part of text information:
[0093] Analyze the target role contained therein, and then call the configured personalized text-image generation model corresponding to the target role to generate a matching image based on the key information corresponding to the partial text information.
[0094] In the solution provided by this embodiment, by configuring a personalized text-image generation model corresponding to each character, the model supports the ability to maintain the identity of the corresponding character in different images when generating different images. Furthermore, the model's own generalization capabilities can be maintained. When generating an image for each portion of text information, the personalized text-image generation model corresponding to the target character contained in the text information can be called based on the target character, ensuring the consistency of the target character's image identity in the multiple generated images, thereby ensuring the coherence of the illustrations of the text and image content.
[0095] Furthermore, considering that the same image may involve multiple main characters, for example, a certain part of text information contains multiple characters at the same time, such as a certain text information "At noon, the little white rabbit was still lying under the big tree to sleep, at this time the turtle was still crawling forward with all its strength, and had already surpassed the position of the little white rabbit". This text information contains both the "little white rabbit" and the "turtle" characters.
[0096] As described in the previous embodiments, this application can configure a corresponding personalized text-image generation model for each character, such as configuring a personalized text-image generation model for "Little White Rabbit" and "Turtle" respectively. If the personalized text-image generation model corresponding to only one character is used to generate the image corresponding to the text information, the identity of the other character in the generated image will be changed.
[0097] For example, Figure 2 This is an image generated by the personalized text-image generation model configured for "turtle". When generating an image for a text message: "At noon, the little white rabbit was still sleeping under the tree, while the tortoise was still crawling forward and had already surpassed the little white rabbit", if the personalized text-image generation model corresponding to "little white rabbit" is selected to generate the image, the generated image may be Figure 3 As shown. Obviously, Figure 3 The identity of the "turtle" in Figure 2 The identity of the "turtle" in the novel has changed.
[0098] To address the issue of role identity changes when a picture contains multiple characters, this embodiment provides a solution. Specifically:
[0099] For each portion of text information, if there is one target role included therein, the personalized text-image generation model corresponding to the one target role can be directly called to generate an image based on the key information corresponding to the portion of text information.
[0100] For each portion of text information, if there are more than two target characters included therein, the process of calling the configured personalized text-image generation model corresponding to the target character and generating an image based on the key information corresponding to the portion of text information may include:
[0101] S1. Select any one of the target characters, call a personalized text-image generation model corresponding to the selected target character, and generate a first image based on key information corresponding to the partial text information.
[0102] S2. Traverse and select the target characters from the target characters that have not been selected. After each target character is selected, mask the image corresponding to the currently selected target character in the first picture to obtain a second picture, call the personalized text-image generation model corresponding to the currently selected target character, and generate a new first picture again based on the key information corresponding to the partial text information and the second picture.
[0103] S3. After traversing the target character, the first picture generated last time is used as the final picture.
[0104] Combine Figure 2-Figure 4 The process of the above steps S1-S3 is exemplified.
[0105] Regarding a picture book about the tortoise and the hare, part of the content is as follows:
[0106] Paragraph 1: "An hour after setting out, the turtle was the only one on the road, slowly moving forward. The little turtle was neither anxious nor impatient, and was not discouraged at all."
[0107] …
[0108] Paragraph 5: "At noon, the little white rabbit was still sleeping under the tree, while the tortoise was still crawling forward with all his might, and had already surpassed the little white rabbit."
[0109] …
[0110] For the first paragraph of text information, since it only contains the target character "turtle", after extracting the key information, the personalized text-image generation model corresponding to "turtle" is called, and the generated image is as follows: Figure 2 shown.
[0111] For the fifth paragraph of text information, which contains both "little white rabbit" and "turtle", the two target characters are included. Therefore, after extracting the key information, the personalized text-image generation model corresponding to "little white rabbit" is first called. The generated image is as follows: Figure 3 As shown. Figure 3 The "turtle" identity in is wrong, so Figure 3 The “turtle” in the image is blocked, and then the blocked image and the extracted key information are sent to the personalized text-image generation model corresponding to the “turtle” to reshape the image. Figure 3 The image of the "turtle" in the image is generated as follows Figure 4 As shown. Obviously, Figure 4 The identities of the "tortoise" and the "little white rabbit" are both correct.
[0112] The solution provided by this embodiment can solve the problem of identity errors of multiple subject characters in the same image, ensure the consistency of the image identity of the target character in the multiple generated images, and ensure the coherence of the illustrations of the graphic content.
[0113] In some embodiments of the present application, a process of configuring a corresponding personalized text-image generation model for each character is introduced.
[0114] Reference Figure 5 As shown, a text feature extractor can be added to the general text-image generation model to obtain an edited text-image generation model.
[0115] A training data set corresponding to each character is obtained, wherein the training data set includes a plurality of training data pairs consisting of image samples and their description texts, and the identity of the image corresponding to the same character in different image samples is consistent.
[0116] For each role: the descriptive text of the image sample in the corresponding training data set is input into the edited text-image generation model, and the general text-image generation model in the edited text-image generation model is used to process the descriptive text, and the learnable text vector of the text label corresponding to the role is input into the text feature extractor to extract text features, and the text features are used to offset the features extracted by the hidden layer of the general text-image generation model, and the decoding module of the general text-image generation model decodes the offset-processed features to obtain a generated image, and the model parameters are updated with the goal of making the generated image close to the image sample to obtain a personalized text-image generation model corresponding to the role.
[0117] During training, the model parameters of the general text-image generation model can remain fixed. The text feature extractor can use the pre-trained textEncoder from the CLIP (Contrastive Language–Image Pre-training) model. Before training, a text label can be assigned to each character. For example, the text label assigned to the character "Turtle" could be "Little Turtle Cancan." Before training, the text vectors for the character's text labels can be initialized. For example, a learnable text vector for the text label can be set using one-hot encoding. The text vectors for the character's text labels are learnable. During training, the text vectors for the text labels are continuously updated until the final text vectors for the character's text labels are obtained after training.
[0118] The general text-image generation model can adopt stable diffusion or other large model structures. Taking stable diffusion as an example, it includes the attention module cross-attention.
[0119] The text vector of the text label corresponding to the role is extracted with text features by the text feature extractor. The dimension of the text features is consistent with the dimension of the features extracted by the attention module. The extracted text features can be fed into the attention module of the general text-image generation model to perform offset processing on the features extracted by the attention module. Specifically, the two matrices W and V of the attention module can be updated. W and V affect the similarity between the generated image and the input text, thereby achieving the purpose of training a personalized text-image generation model. In addition, in order not to affect the generalization ability of the general text-image generation model, the training dataset corresponding to each role can also include the sample image-description text data pairs used in the training of the general text-image generation model.
[0120] Based on the solution provided by this embodiment, a corresponding personalized text-image generation model can be created for each character. This personalized text-image generation model adds a text feature extractor to the general text-image generation model to extract features from the text vector of the text label corresponding to the character. The extracted text features are used to offset the features extracted from the hidden layer of the general text-image generation model, and an image is generated after decoding. While ensuring the generalization capability of the general text-image generation model, this application ensures that the identity of the corresponding character in the image generated by the created personalized text-image generation model remains consistent.
[0121] In addition, the personalized text and image generation model created in this embodiment only adds a text feature extractor compared to the general text and image generation model. The training process can fix the parameters of the general text and image generation model, and only needs to update the learnable text vector of the text label corresponding to the role, so that the amount of trainable parameters is small, reducing the training difficulty, and the personalized text and image generation model corresponding to each role only needs to save a small number of parameters, which is easy to deploy.
[0122] In some embodiments of the present application, considering that graphic content generally has a unified style, in order to ensure the consistency of the style of the illustrations in the final generated graphic content, in this embodiment, before fusing the text information and the image in the aforementioned step S130, the following steps may be further added:
[0123] The generated image is converted into a specified style, where the specified style is a preset unified style, or the specified style is a style that matches the text information.
[0124] This application can pre-define some style types, such as cartoon style, realistic style, two-dimensional style, science fiction style, etc. When generating graphic content, the user can specify the style of the generated graphic content. In addition, this application can also perform style recognition on text information, for example, through the use of a pre-trained style recognition model to identify the style corresponding to the text information and use it as the specified style.
[0125] For the generated images, they can be converted into a specified style through a style conversion model, and then fused with text information to generate graphic content, ensuring the consistency of the overall style of the graphic content.
[0126] The following describes a graphic content generation device provided in an embodiment of the present application. The graphic content generation device described below and the graphic content generation method described above can refer to each other.
[0127] See also Figure 6 , Figure 6 This is a schematic diagram of the structure of a graphic content generation device disclosed in an embodiment of the present application.
[0128] like Figure 6 As shown, the device may include:
[0129] A text acquisition unit 11 is used to acquire text information for generating a picture;
[0130] A key information acquisition unit 12, configured to acquire key information from the text information;
[0131] The image generation unit 13 is used to call a preconfigured text-image generation model to generate an image based on the key information;
[0132] The text-image fusion unit 14 is configured to fuse the text information and the image to obtain generated image-text content.
[0133] Optionally, the process of the text acquisition unit acquiring the text information for generating the accompanying image includes:
[0134] Obtain key element information input by the user;
[0135] A preconfigured artificial intelligence model is called to generate text information based on the key element information.
[0136] Optionally, the process of the key information obtaining unit obtaining the key information from the text information includes:
[0137] Dividing the text information according to the set division granularity to obtain each part of the divided text information;
[0138] Get the key information of each part of text information separately.
[0139] Optionally, the process of the key information obtaining unit obtaining the key information of each part of the text information separately includes:
[0140] For each part of text information, a preconfigured artificial intelligence model is called to generate key information corresponding to each part of text information.
[0141] Optionally, the image generation unit calls a preconfigured text-image generation model to generate an image based on the key information, including:
[0142] For each part of text information, analyze the target role contained therein;
[0143] The configured personalized text-image generation model corresponding to the target character is called to generate an image based on key information corresponding to the partial text information, wherein the personalized text-image generation model corresponding to the target character supports the ability to generate different images so that the identity of the corresponding image of the target character in different images is consistent.
[0144] Optionally, the device of the present application may further include:
[0145] A text analysis unit, configured to analyze the information of each character contained in the text information before the image generation unit calls a preconfigured text-image generation model to generate an image based on the key information;
[0146] The personalized text and image generation model configuration unit is used to configure a personalized text and image generation model corresponding to each role, and the personalized text and image generation model corresponding to the role supports the ability to generate different images with the identity of the corresponding image of the role in different images being consistent.
[0147] If there is only one target character, the image generation unit calls a personalized text-image generation model configured corresponding to the target character to generate an image based on key information corresponding to the partial text information, including:
[0148] A personalized text-image generation model corresponding to the target character is called to generate an image based on key information corresponding to the partial text information.
[0149] If there are more than two target characters, the image generation unit calls the configured personalized text-image generation model corresponding to the target character to generate an image based on the key information corresponding to the partial text information, including:
[0150] Selecting any one of the target characters, calling a personalized text-image generation model corresponding to the selected target character, and generating a first image based on key information corresponding to the partial text information;
[0151] traversing and selecting the target characters from the target characters that have not been selected, each time after selecting the target character, performing an occlusion process on the image corresponding to the currently selected target character in the first picture to obtain a second picture, calling the personalized text-image generation model corresponding to the currently selected target character, and generating a new first picture again based on the key information corresponding to the partial text information and the second picture;
[0152] After traversing the target character, the first picture generated last time is used as the final picture.
[0153] Optionally, the process of configuring the personalized text-image generation model corresponding to each role by the personalized text-image generation model configuration unit includes:
[0154] Add a text feature extractor to the general text-image generation model to obtain the edited text-image generation model;
[0155] Obtaining a training data set corresponding to each character, wherein the training data set includes a plurality of training data pairs consisting of image samples and their description texts, wherein the identity of the image corresponding to the same character in different image samples is consistent;
[0156] For each role: the descriptive text of the image sample in the corresponding training data set is input into the edited text-image generation model, and the general text-image generation model in the edited text-image generation model is used to process the descriptive text, and the learnable text vector of the text label corresponding to the role is input into the text feature extractor to extract text features, and the text features are used to offset the features extracted by the hidden layer of the general text-image generation model, and the decoding module of the general text-image generation model decodes the offset-processed features to obtain a generated image, and the model parameters are updated with the goal of making the generated image close to the image sample to obtain a personalized text-image generation model corresponding to the role.
[0157] Optionally, the device of the present application may further include:
[0158] The style conversion unit is used to convert the generated image into a specified style before the text-image fusion unit fuses the text information and the image, where the specified style is a preset unified style, or the specified style is a style that matches the text information.
[0159] The graphic content generation device provided in the embodiment of the present application can be applied to graphic content generation devices, such as mobile phones, computers, servers, and cloud computing. Figure 7 The hardware structure diagram of the graphic content generation device is shown. Figure 7 ,The hardware structure of the graphic content generating device may include: at least one processor 1, at least one communication interface 2, at least one memory 3 and at least one communication bus 4;
[0160] In the embodiment of the present application, the number of the processor 1, the communication interface 2, the memory 3, and the communication bus 4 is at least one, and the processor 1, the communication interface 2, and the memory 3 communicate with each other through the communication bus 4;
[0161] The processor 1 may be a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention;
[0162] The memory 3 may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory;
[0163] The memory stores a program, and the processor can call the program stored in the memory, wherein the program is used to:
[0164] Get the text information of the image to be generated;
[0165] Obtaining key information from the text information;
[0166] Calling a preconfigured text-image generation model to generate an image based on the key information;
[0167] The text information and the image are integrated to obtain generated graphic content.
[0168] Optionally, the detailed functions and extended functions of the program may refer to the above description.
[0169] An embodiment of the present application further provides a storage medium, which may store a program suitable for execution by a processor, wherein the program is used to:
[0170] Get the text information of the image to be generated;
[0171] Obtaining key information from the text information;
[0172] Calling a preconfigured text-image generation model to generate an image based on the key information;
[0173] The text information and the image are integrated to obtain generated graphic content.
[0174] Optionally, the detailed functions and extended functions of the program may refer to the above description.
[0175] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0176] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can be referenced to each other.
[0177] The above description of the disclosed embodiments will enable those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is to be construed in the widest manner consistent with the principles and novel features disclosed herein.
Claims
1. A method for generating graphic content, characterized in that: include: Get the text information of the image to be generated; Obtaining key information from the text information; Calling a preconfigured text-image generation model to generate an image based on the key information; Fusion of the text information and the image to obtain generated graphic content; The calling of a preconfigured text-image generation model to generate an image based on the key information includes: For each part of text information, analyze the target role contained therein; If there are more than two target characters, any one of the target characters is selected, and a personalized text-image generation model corresponding to the selected target character is called to generate a first image based on key information corresponding to the partial text information; wherein the personalized text-image generation model corresponding to the target character supports the ability to generate different images while maintaining the identity of the corresponding image of the target character in different images and the generalization ability of the model itself; traversing and selecting the target characters from the target characters that have not been selected, each time after selecting the target character, performing an occlusion process on the image corresponding to the currently selected target character in the first picture to obtain a second picture, calling the personalized text-image generation model corresponding to the currently selected target character, and generating a new first picture again based on the key information corresponding to the partial text information and the second picture; After traversing the target character, the first picture generated last time is used as the final picture.
2. The method according to claim 1, characterized in that The step of obtaining text information for generating a picture includes: Obtain key element information input by the user; A preconfigured artificial intelligence model is called to generate the text information based on the key element information.
3. The method according to claim 1, characterized in that Obtain key information from the text information, including: Dividing the text information according to the set division granularity to obtain each part of the divided text information; Get the key information of each part of text information separately.
4. The method according to claim 3, characterized in that The key information of each part of the text information is obtained separately, including: For each part of text information, a preconfigured artificial intelligence model is called to generate key information corresponding to each part of text information.
5. The method according to claim 1, wherein Before calling the preconfigured text-image generation model and generating the image based on the key information, the method further includes: Analyzing each character information contained in the text information; For each character, a personalized text-image generation model corresponding to the character is configured. The personalized text-image generation model corresponding to the character supports the ability to generate different images so that the identity of the corresponding image of the character in different images is consistent.
6. The method according to claim 1, characterized in that If there is only one target character, the process of calling the personalized text-image generation model configured corresponding to the target character and generating an image based on the key information corresponding to the partial text information includes: A personalized text-image generation model corresponding to the target character is called to generate an image based on key information corresponding to the partial text information.
7. The method according to claim 5, characterized in that For each role, a process of configuring a personalized text-image generation model corresponding to the role includes: Add a text feature extractor to the general text-image generation model to obtain the edited text-image generation model; Obtaining a training data set corresponding to each character, wherein the training data set includes a plurality of training data pairs consisting of image samples and their description texts, wherein the identity of the image corresponding to the same character in different image samples is consistent; For each role: the descriptive text of the image sample in the corresponding training data set is input into the edited text-image generation model, and the general text-image generation model in the edited text-image generation model is used to process the descriptive text, and the learnable text vector of the text label corresponding to the role is input into the text feature extractor to extract text features, and the text features are used to offset the features extracted by the hidden layer of the general text-image generation model, and the decoding module of the general text-image generation model decodes the offset-processed features to obtain a generated image, and the model parameters are updated with the goal of making the generated image close to the image sample to obtain a personalized text-image generation model corresponding to the role.
8. The method according to any one of claims 1 to 7, characterized in that Before fusing the text information and the picture, the method further includes: The generated image is converted into a specified style, where the specified style is a preset unified style, or the specified style is a style that matches the text information.
9. A graphic content generating device, characterized in that: include: A text acquisition unit, used to acquire text information for generating a picture; A key information acquisition unit, configured to acquire key information from the text information; An image generation unit, configured to call a preconfigured text-image generation model to generate an image based on the key information; A text-image fusion unit, configured to fuse the text information and the image to obtain generated image-text content; The image generation unit is specifically configured to: For each part of text information, analyze the target role contained therein; If there are more than two target characters, any one of the target characters is selected, and a personalized text-image generation model corresponding to the selected target character is called to generate a first image based on key information corresponding to the partial text information; wherein the personalized text-image generation model corresponding to the target character supports the ability to generate different images while maintaining the identity of the corresponding image of the target character in different images and the generalization ability of the model itself; traversing and selecting the target characters from the target characters that have not been selected, each time after selecting the target character, performing an occlusion process on the image corresponding to the currently selected target character in the first picture to obtain a second picture, calling the personalized text-image generation model corresponding to the currently selected target character, and generating a new first picture again based on the key information corresponding to the partial text information and the second picture; After traversing the target character, the first picture generated last time is used as the final picture.
10. A graphic content generating device, characterized in that: include: memory and processor; The memory is used to store programs; The processor is configured to execute the program to implement each step of the method for generating graphic and text content according to any one of claims 1 to 8.
11. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, each step of the method for generating graphic content according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
A illustration generating method, device and system
CN113449139A