Insertion generation method and device, electronic equipment and storage medium

By obtaining plot information and character image information of literary creation, generating text prompt words and using diffusion models to generate illustrations, the problems of low efficiency in generation of traditional illustrations and difficult to meet personalized needs are solved, high-quality and personalized illustration generation is achieved, and user reading experience is improved.

CN120032019APending Publication Date: 2025-05-23BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510120967.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The generation of traditional illustrations relies on manual labor, is inefficient and difficult to adapt to personalized needs, and cannot effectively enhance users' reading interest and immersion.

Method used

By obtaining plot information of different chapters of literary creation, determining role and storyboard information, obtaining character image information, and generating text prompt words based on this information, and using the target diffusion model to generate illustrations corresponding to plot information.

Benefits of technology

It realizes personalized and diverse illustration generation, improves user satisfaction with illustrations, and enhances user stickiness and retention rate for text reading.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032019A_ABST
    Figure CN120032019A_ABST
Patent Text Reader

Abstract

The invention provides an illustration generation method and device, electronic equipment and a storage medium, and relates to the technical field of artificial intelligence such as large models and text graphs. According to the specific implementation scheme, plot information of different chapters of literature creation is obtained; on the basis of the plot information, determining roles and split mirror information; obtaining image information of the role based on the at least one chapter original text; generating a text cue word according to the split mirror information and the image information of the role; and generating an illustration corresponding to the plot information based on the text cue word through the target diffusion model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the fields of artificial intelligence technology such as large models and text graphics, and in particular to an illustration generation method, device, electronic device and storage medium. Background Art

[0002] As users' requirements for reading experience continue to increase, traditional pure text novel readers can meet basic reading needs, but they are relatively simple in terms of visual stimulation and emotional resonance.

[0003] Illustrations in novels are an important means to enhance readers' immersion, which can improve users' reading interest and immersion. However, the generation of traditional illustrations relies on manual labor, which is inefficient and difficult to adapt to personalized needs. Summary of the invention

[0004] The present disclosure provides a method, device, electronic device and storage medium for generating illustrations.

[0005] According to one aspect of the present disclosure, a method for generating illustrations is provided, comprising: obtaining plot information of different chapters of a literary work; determining character and storyboard information based on the plot information; obtaining image information of the character based on at least one original chapter text; generating text prompt words based on the storyboard information and the image information of the character; and generating illustrations corresponding to the plot information based on the text prompt words through a target diffusion model.

[0006] According to another aspect of the present disclosure, there is provided an illustration generation device, comprising: a first acquisition module for acquiring plot information of different chapters of a literary work; a determination module for determining character and storyboard information based on the plot information; a second acquisition module for acquiring image information of the character based on at least one original chapter text; a first generation module for generating text prompt words based on the storyboard information and the image information of the character; and a second generation module for generating illustrations corresponding to the plot information based on the text prompt words through a target diffusion model.

[0007] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the illustration generation method described in the above-mentioned one aspect embodiment.

[0008] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, on which a computer program / instructions are stored, and the computer instructions are used to enable the computer to execute the illustration generation method described in the above-mentioned embodiment.

[0009] According to another aspect of the present disclosure, a computer program product is provided, including a computer program / instruction, wherein when the computer program / instruction is executed by a processor, the illustration generating method described in the above-mentioned first embodiment is implemented.

[0010] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.

[0012] Figure 1 A flowchart of a method for generating an illustration provided in an embodiment of the present disclosure;

[0013] Figure 2 A schematic diagram of a process for generating text prompt words provided in an embodiment of the present disclosure;

[0014] Figure 3 A flowchart of another illustration generation method provided by an embodiment of the present disclosure;

[0015] Figure 4 A flowchart of another illustration generation method provided by an embodiment of the present disclosure;

[0016] Figure 5 A schematic diagram of a process for generating a target character image provided by an embodiment of the present disclosure;

[0017] Figure 6 A schematic diagram of a process for fine-tuning a diffusion model provided in an embodiment of the present disclosure;

[0018] Figure 7 A schematic diagram of a flowchart of an illustration updating process in an illustration generating method provided in an embodiment of the present disclosure;

[0019] Figure 8 A schematic diagram of a flow chart for identifying and processing hand deformities provided in an embodiment of the present disclosure;

[0020] Fig. 9 A schematic diagram of the structure of an illustration generating device provided in an embodiment of the present disclosure;

[0021] Fig.10 The present invention is a block diagram of an electronic device for implementing the illustration generation method according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0022] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0023] The following describes an illustration generation method, an apparatus, an electronic device, and a storage medium according to embodiments of the present disclosure with reference to the accompanying drawings.

[0024] Artificial Intelligence (AI) is a discipline that studies how computers can simulate certain thought processes and intelligent behaviors of human beings (such as learning, reasoning, thinking, planning, etc.). It includes both hardware-level technologies and software-level technologies. AI hardware technologies generally include computer vision technology, speech recognition technology, natural language processing technology, as well as learning / deep learning, big data processing technology, knowledge graph technology, and other aspects.

[0025] Figure 1 A flowchart of a method for generating an illustration provided in an embodiment of the present disclosure.

[0026] like Figure 1 As shown, the illustration generation method may include:

[0027] S101, obtaining plot information of different chapters of a literary creation.

[0028] It should be noted that the execution subject of the illustration generation method in the embodiment of the present disclosure may be a hardware device with data processing capabilities and / or the necessary software required to drive the hardware device to work. Optionally, the execution subject may include a server, a user terminal and other intelligent devices. Optionally, the user terminal includes but is not limited to a mobile phone, a computer, an intelligent voice interaction device, etc. Optionally, the server includes but is not limited to a network server, an application server, and may also be a server of a distributed system, or a server combined with a blockchain, etc. The embodiment of the present disclosure is not specifically limited.

[0029] In some embodiments, plot information may be extracted from different chapters of a literary work based on a pre-trained Large Language Model (LLM).

[0030] Optionally, the plot information includes, but is not limited to, plot description, original plot text, character information, scene location and other information.

[0031] In some embodiments, the literary creation is input into the LLM, and the LLM parses the content of each chapter of the literary creation to obtain plot information of different chapters.

[0032] In some embodiments, for each chapter, LLM may divide each chapter into a set number of plots and extract plot information from each plot, that is, each chapter may contain multiple plot information.

[0033] For example, Figure 2 For example, chapter 1 is divided into three plots, namely plot 1, plot 2 and plot 3. For plot 1, the plot description, original plot text, character information, scene location and other information of plot 1 are obtained as plot information 1 of plot 1. Furthermore, plot information 2 and plot information 3 of plot 2 and plot 3 can be obtained, and plot information 1, plot information 2 and plot information 3 are used as the plot information of chapter 1.

[0034] S102, determining role and storyboard information based on the plot information.

[0035] In some embodiments, characters may be extracted from plot information, and a global character list may be generated based on the extracted characters to integrate the characters in the literary creation.

[0036] In some embodiments, splittable elements, such as scene size, lighting effects, scene layout, etc., can be identified from the plot information, and LLM can be used to generate description information of the splittable elements as storyboard information.

[0037] S103, acquiring character image information based on at least one chapter original text.

[0038] In some embodiments, the original texts of chapters associated with multiple characters may be input into the LLM, and the LLM extracts the original description information of each character from the original texts of the chapters and performs standardization processing on the original description information to obtain the image information of the character.

[0039] In some embodiments, the original description information of the character includes information such as gender, age, appearance characteristics and clothing.

[0040] Optionally, the format of the character's image information may be unified to improve the consistency and accuracy of generating prompt words.

[0041] S104, generating text prompt words according to the storyboard information and the character's image information.

[0042] In some embodiments, the storyboard information and the character's image information may be combined to obtain text prompts, wherein the text prompts at least include the character's appearance, scene features, and scene information.

[0043] For example, Figure 2 What is shown is a schematic diagram of a flow chart for generating text prompt words.

[0044] 1. Plot extraction: Use LLM to analyze the content of each chapter of literary creation and extract core plot information, including plot description, original plot text, character information, scene location and other information.

[0045] 2. Character information integration and description generation: Extract characters from plot information and generate a global character list. Based on the character list and the original chapter text, use LLM to obtain the original description information of each character, including gender, age, appearance characteristics, and clothing.

[0046] 3. Character description standardization: Further standardize the original description information to obtain the character's image information to ensure that the character's image information remains consistent and adopts a unified format to facilitate the subsequent generation of prompt words.

[0047] 4. Storyboard information generation: Based on the storyboard-capable elements in the plot information, use LLM to generate storyboard information for each plot, including scene size, lighting effects, scene layout, etc.

[0048] 5. Text prompt word construction: Combine the character's image information and the storyboard information to generate the text prompt words required by the target diffusion model. The text prompt words include the character's appearance, scene characteristics, and scene type, ensuring that the generated illustrations are highly consistent with the plot of the literary creation.

[0049] S105, generating illustrations corresponding to the plot information based on the text prompt words through the target diffusion model.

[0050] In some embodiments, the target character image corresponding to the character in the plot information can be determined based on the image information of the character in the text prompt words, and the target diffusion model can be used to generate the character image and scene image contained in the plot information according to the storyboard information and the target character image in the text prompt words as illustrations corresponding to the plot information.

[0051] In some embodiments, the target character image can be obtained from a pre-established character image library. The character image library is composed of character fine-tuning parameters corresponding to different character images. That is, the character fine-tuning parameters can be matched in the character image library according to the character image information, and in response to the existence of matching character fine-tuning parameters, the target diffusion model generates the target character image using the character fine-tuning parameters.

[0052] Optionally, the similarity between the character's image information and the character fine-tuning parameters may be calculated. If there is a character fine-tuning parameter whose similarity is greater than a set threshold, it is determined to use the character fine-tuning parameter to generate the target character image.

[0053] In some embodiments, if there are no matching image fine-tuning parameters in the character image library, a character base map can be generated based on the character's image information, and the character base map can be input into the face optimization model to optimize the character's facial information, and the optimized character base map can be input into the clothing optimization model to optimize the character's clothing information, thereby obtaining the target character image.

[0054] In some embodiments, after obtaining the target character image, it is also possible to determine whether there is any deformity in the hand area of ​​the target character image, and if there is any deformity, the target character image is redrawn according to the degree of the deformity.

[0055] Optionally, if the deformity is mild, the hand area is redrawn, and if the deformity is severe, the target character image is redrawn.

[0056] According to the illustration generation method provided by the embodiment of the present disclosure, the plot information of the literary creation is obtained, and the character and storyboard information are determined from the plot information. Then, the image information of the character is determined according to the original text of the chapter, and text prompt words are generated according to the image information and the storyboard information, so as to use the target diffusion model to generate illustrations corresponding to the plot information based on the text prompt words. The generation of illustrations based on the large language model and the diffusion model provides personalized and diversified visual content support. Using the image information and storyboard information as text prompt words enables the diffusion model to accurately restore the character details, thereby improving the user's satisfaction with the illustrations and further improving the user's stickiness and retention rate for text reading.

[0057] Figure 3 A flowchart of a method for generating an illustration provided in an embodiment of the present disclosure.

[0058] like Figure 3 As shown, the illustration generation method may include:

[0059] S301, obtaining plot information of different chapters of a literary creation.

[0060] The relevant contents of step S301 can be found in the above embodiment and will not be described again here.

[0061] S302, determining role and storyboard information based on the plot information.

[0062] For details on determining the role based on the plot information in step S302, please refer to the above embodiments, which will not be described in detail here.

[0063] In some embodiments, the plot information may be identified to obtain elements that can be divided into scenes, and storyboard information may be generated based on the elements that can be divided into scenes. The elements that can be divided into scenes include, but are not limited to, scene size, lighting effects, scene layout, etc. The storyboard information is obtained from the plot information, and the illustrations are generated using the storyboard information, which improves the quality of the illustrations and enhances the coherence of the illustrations.

[0064] Optionally, LLM may identify the separable elements from the plot information, and generate description information of the separable elements according to the separable elements to obtain the storyboard information.

[0065] S303, determining the original text of the chapter associated with the role, and extracting the original description information of the role from the original text of the associated chapter.

[0066] In some embodiments, a role identifier may be assigned to each role, and based on the identifier of any role, a search may be conducted in the chapters of the literary work to determine whether any role is described. If so, the context of any role in the chapter may be used as the original text of the chapter associated with the role.

[0067] Taking role A as an example, its corresponding identifier is A. Use LLM to search in chapter 1. If role A is found in chapter 1, the context of role A is used as the original text of the chapter associated with role A.

[0068] Furthermore, the original description information of the character can be extracted from the original text of the chapter, including: the character's basic information, personality traits, interpersonal relationships, character image and other information.

[0069] Optionally, the original description information may be extracted from the original text of the chapter based on keywords of the original description information.

[0070] S304, performing standard conversion on the original description information to obtain the character's image information.

[0071] In some embodiments, in order to ensure the consistency of the image characteristics of the character, the original description information of the character can be standardized. Optionally, the basic information of the character can be standardized, and the image of the character can also be standardized to obtain the image information of the character.

[0072] In some embodiments, the format of the character's image information can also be unified to unify the format of the text prompt words and enhance the consistency of the character image. For example, a unified document format, such as font, font size, paragraph format, etc., can be formulated.

[0073] S305, generating text prompt words according to the storyboard information and the character's image information.

[0074] S306, generating illustrations corresponding to the plot information based on the text prompt words through the target diffusion model.

[0075] The relevant contents of steps S305-S306 can be found in the above embodiment and will not be repeated here.

[0076] According to the illustration generation method provided by the embodiment of the present disclosure, the original description information of the character is extracted from the original text associated with the character, and the original description information is standardized to obtain the image information of the character. By standardizing the original description information, the consistency of the character image can be ensured.

[0077] Figure 4 A flowchart of a method for generating an illustration provided in an embodiment of the present disclosure.

[0078] like Figure 4 As shown, the illustration generation method may include:

[0079] S401, obtaining plot information of different chapters of a literary creation.

[0080] S402, determining role and storyboard information based on the plot information.

[0081] S403, acquiring character image information based on at least one chapter original text.

[0082] S404, generating text prompt words according to the storyboard information and the character's image information.

[0083] The relevant contents of steps S401 - S404 can be found in the above embodiment and will not be described again here.

[0084] S405, obtaining a target character image according to the character image information in the text prompt word.

[0085] In some embodiments, a character image is matched from a character image library. If the character image library contains character fine-tuning parameters that match the character's image information, a character image is generated by using the character fine-tuning parameters, and the character image is used as a target character image. If the character image library does not contain a character image that matches the character's image information, a character base map is generated and optimized to obtain a target character image. Since there are different ways of obtaining the target character image, different needs for generating illustrations can be flexibly responded to, which not only improves the quality of the generated illustrations, but also greatly enhances the ability to adapt to complex character generation scenarios.

[0086] In some embodiments, a character model can be used to construct a character image library to improve the efficiency of illustration generation and ensure that the character images are consistent in style, color and details. A global character atlas of multiple sample characters in different states is generated, wherein the character images in the global character atlas are complete character images including the front and side views of the sample characters, as well as images from different angles and with different expressions.

[0087] Furthermore, for the global character atlas of each sample character, the character images in the global character atlas are preprocessed and labeled to obtain a first training sample. The first training sample is used for training to improve the accuracy of model training.

[0088] Further, the character model is trained with Lora based on the first training sample, and after the training is completed, the character fine-tuning parameters corresponding to the sample character are obtained. The character fine-tuning parameters refer to the key information of the character, including age, gender, hairstyle, body characteristics and clothing details.

[0089] In some embodiments, after the character image library is constructed, character matching can be performed in the pre-constructed character image library according to the character image information, fine-tuning parameters of successfully matched characters can be obtained, and a target character image can be generated based on the character fine-tuning parameters.

[0090] Optionally, the degree of match between the character's image information and the character's fine-tuning parameters can be calculated, and in response to the degree of match being greater than a set threshold, it is determined that the character's fine-tuning parameters are matched successfully, and then the character's fine-tuning parameters are called to generate a target character image, thereby ensuring the consistency of the character in the target character image. For example, the similarity between the image information and the character's fine-tuning parameters can be calculated as the degree of match.

[0091] In some embodiments, if there is no successfully matched character fine-tuning parameter in the character image library, a first character base image may be generated according to the character image information in the text prompt word, and the facial features of the first character base image may be further optimized to obtain a second character base image, so as to ensure that the face of the character in the second character base image remains highly unified in multiple scenes.

[0092] Furthermore, the clothing features of the second character base image are optimized to obtain the target character image, thereby ensuring that the clothing style matches the character description and shows a natural transition in different scenes. For example, the Pulid model can be used to optimize facial features, and the Redux model can be used to optimize clothing features.

[0093] That is, the first character base image is input into the Pulid model to optimize the facial features to obtain the second character base image, and the second character base image is input into the Redux model to optimize the clothing features to obtain the target character image.

[0094] Figure 5 The figure shows a flow chart of generating a target character image. Before generating the target character image, a character image library is constructed, and a first training sample is obtained by obtaining a global character atlas of sample characters in different states and performing operations such as image cropping and labeling on the global character atlas. Figure 5 The first training sample in the example includes N training samples of the same person in different formats. For example, there are frontal image, side image and smiling image in bitmap image format (Portable Network Graphics, PNG); for another example, there are frontal image, side image and smiling image in text format (txt).

[0095] Furthermore, the character model is trained with Lora using the first training sample to obtain key information such as character age, gender, hairstyle, body characteristics, and clothing details corresponding to the sample character, as character fine-tuning parameters corresponding to the sample character.

[0096] After the character image library is established, the character's image information is used to match it in the character image library. If there are successfully matched character fine-tuning parameters in the character image library, the character fine-tuning parameters are loaded and used to generate the target character image and output it.

[0097] If there are no successfully matched character fine-tuning parameters in the character image library, the character image information is used to generate the first character base map, and the Pulid model and Redux model are used to optimize the facial features and clothing features to obtain the target character image and output it.

[0098] S406, generating illustrations corresponding to the plot information according to the storyboard information in the text prompt words and the target character image through the target diffusion model.

[0099] In some embodiments, by inputting the storyboard information and the target character image into the target diffusion model, the target diffusion model can generate images of different characters and different storyboards as illustrations corresponding to the plot information. For example, if the character in the target character image is character A, and the storyboard information indicates that the scene is a park, the diffusion model can generate an image of character A in a park as an illustration corresponding to the plot information.

[0100] In some embodiments, before generating illustrations corresponding to the plot information, the diffusion model is fine-tuned according to the painting style information to obtain a target diffusion model, so that the target diffusion model can be used to generate illustrations corresponding to the plot information, so that the illustrations have a high level of visual effects.

[0101] In some embodiments, the illustration style information may be determined according to the subject matter of the literary creation. For example, if the subject matter of the literary creation is ancient romance, the illustration style information is ancient style. Then, the target style fine-tuning parameters that match the illustration style information may be determined from the style fine-tuning parameters corresponding to different styles.

[0102] That is to say, the diffusion model can be pre-trained with Lora based on sample images of different painting styles to obtain the style fine-tuning parameters corresponding to the different painting styles. Optionally, the sample images are pre-processed and annotated to obtain second training samples of different painting styles, and the diffusion model is fine-tuned with Lora based on the second training samples, and the style fine-tuning parameters corresponding to the different painting styles are obtained after the fine-tuning is completed, so that the diffusion model can meet the needs of different painting styles, while improving the diversity of illustration generation and ensuring the consistency of the same painting style.

[0103] Furthermore, the diffusion model can be updated by fine-tuning parameters based on the target style to obtain a target diffusion model, so that the target diffusion model can generate illustrations that match the subject matter of the literary creation.

[0104] In some embodiments, after obtaining the illustration corresponding to the plot information, the illustration can be inserted into the literary creation to improve the readability of the literary creation and enhance the user's sense of immersion when reading. The location information of the plot information is determined, and the insertion position of the illustration corresponding to the plot information is determined based on the location information, and the illustration is inserted at the insertion position.

[0105] Optionally, the insertion position of the illustration can be determined based on the typesetting and layout of the literary text and the position information of the plot information to ensure the coherence of the illustration with the plot and the matching of the literary text.

[0106] Figure 6 The flowchart of fine-tuning the diffusion model is shown. Sample images of different painting styles are obtained and pre-processed to obtain sample images of uniform size. Further, the sample images of uniform size are annotated to obtain second training samples of different painting styles. Figure 6 In the example, the second training sample includes n sample images in different formats, including n sample images in png format and n sample images in txt format. By using the second sample image to perform Lora fine-tuning on the diffusion model, the style fine-tuning parameters corresponding to different styles can be obtained after the fine-tuning. During the Lora fine-tuning process, random noise can be added to the diffusion model to optimize the denoising performance of the diffusion model.

[0107] According to the illustration generation method provided by the embodiment of the present disclosure, the target character image is obtained according to the character's image information, and the target diffusion model is used to generate illustrations corresponding to the plot information according to the storyboard information and the target character image. By using different methods of obtaining the target character image, different requirements for generating illustrations can be flexibly met, which not only improves the quality of the generated illustrations, but also greatly enhances the adaptability to complex character generation scenarios.

[0108] Based on the above embodiments, the present disclosure can explain the process of updating illustrations, such as Figure 7 As shown, the illustration updating process may include:

[0109] S701, extracting a hand region image of the character from the illustration.

[0110] It is understandable that in the image generation based on the diffusion model, the detailed depiction of the hand has always been a recognized difficulty. The generated hand images often have problems such as deformity, incorrect number of phalanges, unnatural interlacing of fingers, etc., which seriously affect the quality and practicality of the generated images. Therefore, the present disclosure can identify the deformity of the hand area of ​​the illustrations generated using the target diffusion model, and update the illustrations with deformities, thereby ensuring the natural effect of the hand area and improving the quality of the illustrations.

[0111] In some embodiments, target detection and segmentation may be performed on the illustration to detect the character's hands from the illustration, and the illustration may be segmented to obtain an image of the character's hand region.

[0112] In some embodiments, the hand region image of the character may be segmented from the illustration based on the detection frame of the target detection.

[0113] S702, performing deformity recognition on the hand area image to determine deformity recognition information of the hand.

[0114] In some embodiments, deformity identification can be performed based on image features of the hand region image, by extracting image features of the hand region image and identifying whether there is a deformity based on the image features. Optionally, the image features can be compared with the deformity features, and if the similarity is greater than a set threshold, it is determined that there is a deformity in the hand region image, and the hand deformity value and deformity location are determined as deformity identification information.

[0115] In some embodiments, a classification model for detecting hand deformity information is pre-trained, and a hand deformity image is input into the classification model, and the classification model identifies the hand deformity, thereby outputting the hand deformity value and the deformity location as deformity identification information.

[0116] Optionally, whether the hand has deformity and the deformity level of the hand deformity can be determined according to the hand deformity value in the deformity identification information, wherein the deformity level includes: no deformity, slight deformity and severe deformity.

[0117] Optionally, in response to the hand deformity value being less than a first set value, it is determined that the character's hand does not have any deformity; in response to the hand deformity value being greater than or equal to the first set value and less than or equal to the second set value, it is determined that the character's hand has a deformity and the deformity level is mild deformity; in response to the hand deformity value being greater than the second set value, it is determined that the character's hand has a deformity and the deformity level is severe deformity.

[0118] S703: In response to the deformity identification information indicating that the character's hand is deformed, updating the illustration corresponding to the plot information.

[0119] In some embodiments, when the hand deformity value in the deformity identification information indicates that the character's hand is deformed, the method for updating the illustration corresponding to the plot information can be determined according to the size of the hand deformity value, so that the illustration can be updated using the corresponding update method, so that refined hand repair can be achieved and the coordination of the hand with the overall scene can be ensured.

[0120] In some embodiments, the hand deformity value and deformity part of the character can be determined according to the deformity identification information, and the first set value and the second set value for determining the deformity level can be obtained. In response to the hand deformity value being greater than or equal to the first set value, and the hand deformity value being less than or equal to the second set value, the deformity part is redrawn to obtain the redrawing data of the deformity part, that is, in response to the deformity level being slightly deformed, the deformity part is redrawn to obtain the redrawing data of the deformity part.

[0121] Optionally, the deformed part may be redrawn using an image redrawing model to obtain redrawing data of the deformed part, and then the deformed part may be replaced according to the redrawing data of the deformed part to update the illustration corresponding to the plot information.

[0122] Optionally, based on the position of the detection frame of the deformed part, the deformed part may be occluded, and the occluded part may be redrawn to obtain redrawing data.

[0123] In some embodiments, in order to achieve refined hand redrawing, hand prompt words and redrawing constraints are obtained, and based on the hand prompt words and redrawing constraints, details of the occluded parts are repaired to obtain redrawing data of the deformed parts.

[0124] In some embodiments, in response to the hand deformity value being greater than a second set value, the illustration corresponding to the plot information is globally redrawn. In other words, in response to the deformity level being severe, the illustration corresponding to the plot information is globally redrawn to ensure the coordination of the hand with the overall scene.

[0125] Optionally, when performing global redrawing, illustrations corresponding to the plot information may be regenerated as updated illustrations based on text prompts generated from the storyboard information and the character's image information.

[0126] Figure 8 The figure shows a schematic diagram of the process of identifying and treating hand deformities. Figure 8 As shown, by identifying the hand deformity value and the deformity part of the hand region image, the deformity level of the hand deformity is determined according to the hand deformity value. If the deformity level indicates no deformity, the illustration corresponding to the hand region image is output.

[0127] If the deformity level indicates a mild deformity, the redrawing model is used to redraw the deformed part. The hand prompt words and redrawing constraints are obtained to redraw the deformed part based on the hand prompt words and redrawing constraints to obtain redrawing data of the deformed part. The deformed part is then replaced with the redrawn data to obtain an updated illustration, which is then output.

[0128] If the deformity level indicates severe deformity, an illustration corresponding to the plot information is regenerated according to the text prompt word and output as an updated illustration.

[0129] According to the illustration generation method provided by the embodiment of the present disclosure, by performing deformity recognition on the hand area image in the illustration to determine whether the character's hand has any deformity, and updating the illustration when any deformity exists, an illustration without deformity in the hand area can be obtained, thereby enhancing the natural effect of the hand area and improving the quality of the illustration.

[0130] Corresponding to the illustration generation methods provided in the above-mentioned embodiments, an embodiment of the present disclosure further provides an illustration generation device. Since the illustration generation device provided in the embodiment of the present disclosure corresponds to the illustration generation methods provided in the above-mentioned embodiments, the implementation methods of the above-mentioned illustration generation methods are also applicable to the illustration generation device provided in the embodiment of the present disclosure and will not be described in detail in the following embodiments.

[0131] Fig. 9 A schematic diagram of the structure of an illustration generating device provided in an embodiment of the present disclosure.

[0132] like Fig. 9As shown, the illustration generating device 900 of the embodiment of the present disclosure includes a first acquiring module 901 , a determining module 902 , a second acquiring module 903 , a first generating module 904 and a second generating module 905 .

[0133] The first acquisition module 901 is used to acquire plot information of different chapters of a literary creation;

[0134] A determination module 902, for determining role and storyboard information based on the plot information;

[0135] A second acquisition module 903, used to acquire character image information based on at least one chapter original text;

[0136] A first generating module 904, for generating text prompt words according to the storyboard information and the image information of the character;

[0137] The second generation module 905 is used to generate illustrations corresponding to the plot information based on the text prompt words through the target diffusion model.

[0138] In one embodiment of the present disclosure, the second acquisition module 903 is further used to: determine the original text of the chapter associated with the role, and extract the original description information of the role from the associated original text of the chapter; perform standardized conversion on the original description information to obtain the image information of the role.

[0139] In one embodiment of the present disclosure, the determination module 902 is further used to: identify the plot information, obtain the splittable elements, and generate the splittable information according to the splittable elements.

[0140] In one embodiment of the present disclosure, the second generation module 905 is further used to: obtain a target character image based on the character's image information in the text prompt word; and generate illustrations corresponding to the plot information based on the storyboard information in the text prompt word and the target character image through a target diffusion model.

[0141] In one embodiment of the present disclosure, the second generation module 905 is further used to: perform character matching in a pre-built character image library according to the character's image information, obtain fine-tuning parameters of successfully matched characters, and generate a target character image based on the character fine-tuning parameters.

[0142] In one embodiment of the present disclosure, the second generation module 905 is also used to: generate a first character base map based on the image information of the character in the text prompt word; optimize the facial features of the first character base map to obtain a second character base map; optimize the clothing features of the second character base map to obtain a target character image.

[0143] In one embodiment of the present disclosure, the second generation module 905 is further used to: generate global character atlases for multiple sample characters in different states; for the global character atlas of each sample character, preprocess and label the character images in the global character atlas to obtain a first training sample; perform Lora training on the character model based on the first training sample, and obtain the character fine-tuning parameters corresponding to the sample character after the training is completed.

[0144] In one embodiment of the present disclosure, the second generation module 905 is further used to: extract a hand area image of the character from the illustration; perform deformity recognition on the hand area image to determine deformity recognition information of the hand; and update the illustration corresponding to the plot information in response to the deformity recognition information indicating that the character's hand is deformed.

[0145] In one embodiment of the present disclosure, the second generation module 905 is further used to: determine the hand deformity value and deformity part of the character according to the deformity identification information; in response to the hand deformity value being greater than or equal to the first set value, and the hand deformity value being less than or equal to the second set value, redraw the deformity part to obtain redrawing data of the deformity part; according to the redrawing data of the deformity part, replace the deformity part to update the illustration corresponding to the plot information; or, in response to the hand deformity value being greater than the second set value, globally redraw the illustration corresponding to the plot information.

[0146] In one embodiment of the present disclosure, the second generation module 905 is further used to: perform occlusion processing on the deformed part based on the detection frame position of the deformed part; obtain hand prompt words and redrawing constraints, and based on the hand prompt words and redrawing constraints, perform detail repair on the occluded part to obtain redrawing data of the deformed part.

[0147] In one embodiment of the present disclosure, the second generation module 905 is also used to: determine the style information of the illustration according to the subject matter of the literary creation; determine the target style fine-tuning parameters that match the illustration style information from the style fine-tuning parameters corresponding to different styles; and update the diffusion model based on the target style fine-tuning parameters to obtain the target diffusion model.

[0148] In one embodiment of the present disclosure, the second generation module 905 is further used to: obtain second training samples under different painting styles, perform Lora fine-tuning on the diffusion model based on the second training samples, and obtain painting style fine-tuning parameters corresponding to different painting styles after the fine-tuning is completed.

[0149] In one embodiment of the present disclosure, the second generating module 905 is further used to: determine the position information of the plot information, determine the insertion position of the illustration corresponding to the plot information based on the position information, and insert the illustration at the insertion position.

[0150] According to the illustration generation device provided by the embodiment of the present disclosure, the plot information of the literary creation is obtained, and the character and storyboard information are determined from the plot information. Then, the image information of the character is determined according to the original text of the chapter, and text prompt words are generated according to the image information and the storyboard information, so as to use the target diffusion model to generate illustrations corresponding to the plot information based on the text prompt words. The generation of illustrations based on the large language model and the diffusion model provides personalized and diversified visual content support. Using the image information and storyboard information as text prompt words enables the diffusion model to accurately restore the character details, thereby improving the user's satisfaction with the illustrations and further improving the user's stickiness and retention rate for text reading.

[0151] In the technical solution disclosed herein, the acquisition, storage and application of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0152] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.

[0153] Fig.10 A schematic block diagram of an example electronic device 1000 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0154] like Fig.10 As shown, the device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program / instruction stored in a read-only memory (ROM) 1002 or a computer program / instruction loaded from a storage unit 1006 to a random access memory (RAM) 1003. In the RAM 1003, various programs and data required for the operation of the device 1000 can also be stored. The computing unit 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.

[0155] A number of components in the device 1000 are connected to the I / O interface 1005, including: an input unit 1006 such as a keyboard, a mouse, etc.; an output unit 1007 such as various types of displays, speakers, etc.; a storage unit 1008 such as a disk, an optical disk, etc.; and a communication unit 1009 such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1009 allows the device 1000 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0156] The computing unit 1001 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 1001 performs the various methods and processes described above, such as the illustration generation method. For example, in some embodiments, the illustration generation method may be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as a storage unit 1006. In some embodiments, part or all of the computer program / instructions may be loaded and / or installed on the device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program / instructions are loaded into the RAM 1003 and executed by the computing unit 1001, one or more steps of the illustration generation method described above may be performed. Alternatively, in other embodiments, the computing unit 1001 may be configured to execute the illustration generating method in any other appropriate manner (eg, by means of firmware).

[0157] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs / instructions that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0158] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0159] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0160] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0161] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.

[0162] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs / instructions running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0163] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in the disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in the disclosure can be achieved, and this document does not limit them here.

[0164] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A method for generating illustrations, wherein: The method comprises: Get plot information for different chapters of a literary creation; Based on the plot information, determine role and storyboard information; Acquire the image information of the character based on at least one chapter original text; Generate text prompt words according to the storyboard information and the image information of the character; An illustration corresponding to the plot information is generated based on the text prompt word through a target diffusion model.

2. The method according to claim 1, wherein: The step of obtaining the character's image information based on at least one chapter original text includes: Determine the original text of the chapter associated with the role, and extract the original description information of the role from the original text of the associated chapter; The original description information is subjected to standardized conversion to obtain the image information of the character.

3. The method according to claim 1, wherein: The process of determining the storyboard information includes: The plot information is identified to obtain elements that can be divided into scenes, and the storyboard information is generated according to the elements that can be divided into scenes.

4. The method according to claim 1, wherein: The step of generating an illustration corresponding to the plot information based on the text prompt word through a target diffusion model includes: Acquire a target character image according to the character image information in the text prompt word; The target diffusion model is used to generate illustrations corresponding to the plot information according to the storyboard information in the text prompt word and the target character image.

5. The method according to claim 4, wherein: The step of obtaining a target character image according to the character image information in the text prompt word includes: According to the image information of the character, character matching is performed in a pre-built character image library, fine-tuning parameters of successfully matched characters are obtained, and a target character image is generated based on the character fine-tuning parameters.

6. The method according to claim 4, wherein: The step of obtaining a target character image according to the character image information in the text prompt word includes: Generate a first character base map according to the image information of the character in the text prompt word; Optimizing the facial features of the first character base image to obtain a second character base image; The clothing features of the second character base image are optimized to obtain the target character image.

7. The method according to claim 5, wherein: The pre-construction process of the character image library includes: Generate global character atlases of multiple sample characters in different states; For each sample character in the global character atlas, preprocessing and labeling the character images in the global character atlas to obtain a first training sample; Lora training is performed on the character model based on the first training sample, and after the training is completed, the character fine-tuning parameters corresponding to the sample character are obtained.

8. The method according to any one of claims 1 to 7, wherein: After generating the illustration corresponding to the plot information based on the text prompt word by using the target diffusion model, the method further includes: Extracting a hand region image of the character from the illustration; Performing deformity recognition on the hand region image to determine deformity recognition information of the hand; In response to the deformity identification information indicating that the character's hand is deformed, the illustration corresponding to the plot information is updated.

9. The method according to claim 8, wherein: The updating of the illustration corresponding to the plot information includes: Determine the hand deformity value and deformity position of the character according to the deformity identification information; In response to the hand deformity value being greater than or equal to a first set value, and the hand deformity value being less than or equal to a second set value, redrawing the deformed part to obtain redrawing data of the deformed part; According to the redrawing data of the deformed part, the deformed part is replaced to update the illustration corresponding to the plot information; or, In response to the hand deformity value being greater than the second set value, the illustration corresponding to the plot information is globally redrawn.

10. The method according to claim 9, wherein: The step of redrawing the deformed part to obtain redrawn data of the deformed part includes: Based on the position of the detection frame of the deformed part, performing occlusion processing on the deformed part; A hand prompt word and a redrawing constraint condition are obtained, and based on the hand prompt word and the redrawing constraint condition, details of the blocked part are repaired to obtain redrawing data of the deformed part.

11. The method according to any one of claims 1 to 7, wherein: Before generating the illustration corresponding to the plot information based on the text prompt word by using the target diffusion model, the method further includes: Determine the style of illustrations based on the subject matter of the literary creation; Determining a target style fine-tuning parameter that matches the illustration style information from the style fine-tuning parameters corresponding to different styles; The diffusion model is updated based on the target painting style fine-tuning parameters to obtain the target diffusion model.

12. The method according to claim 11, wherein: The fine-tuning process of the candidate diffusion model includes: Second training samples under different painting styles are obtained, and Lora fine-tuning is performed on the diffusion model based on the second training samples. After the fine-tuning is completed, painting style fine-tuning parameters corresponding to the different painting styles are obtained.

13. The method according to any one of claims 1 to 7, wherein: After generating the illustration corresponding to the plot information based on the text prompt word by using the target diffusion model, the method further includes: The position information of the plot information is determined, and based on the position information, an insertion position of an illustration corresponding to the plot information is determined, and the illustration is inserted at the insertion position.

14. An illustration generating device, wherein: The device comprises: The first acquisition module is used to obtain plot information of different chapters of literary creation; A determination module, used to determine role and storyboard information based on the plot information; A second acquisition module, used for acquiring the image information of the character based on at least one original chapter text; A first generating module, used for generating text prompt words according to the storyboard information and the image information of the character; The second generation module is used to generate illustrations corresponding to the plot information based on the text prompt words through a target diffusion model.

15. The device according to claim 14, wherein: The second acquisition module is further used for: Determine the original text of the chapter associated with the role, and extract the original description information of the role from the original text of the associated chapter; The original description information is subjected to standardized conversion to obtain the image information of the character.

16. The device according to claim 14, wherein: The determining module is further used for: The plot information is identified to obtain elements that can be divided into scenes, and the storyboard information is generated according to the elements that can be divided into scenes.

17. The device according to claim 14, wherein: The second generating module is further used for: Acquire a target character image according to the character image information in the text prompt word; The target diffusion model is used to generate illustrations corresponding to the plot information according to the storyboard information in the text prompt word and the target character image.

18. The device according to claim 17, wherein: The second generating module is further used for: According to the image information of the character, character matching is performed in a pre-built character image library, fine-tuning parameters of successfully matched characters are obtained, and a target character image is generated based on the character fine-tuning parameters.

19. The device according to claim 17, wherein: The second generating module is further used for: Generate a first character base map according to the image information of the character in the text prompt word; Optimizing the facial features of the first character base image to obtain a second character base image; The clothing features of the second character base image are optimized to obtain the target character image.

20. The device according to claim 18, wherein The second generating module is further used for: Generate global character atlases of multiple sample characters in different states; For each sample character in the global character atlas, preprocessing and labeling the character images in the global character atlas to obtain a first training sample; Lora training is performed on the character model based on the first training sample, and after the training is completed, the character fine-tuning parameters corresponding to the sample character are obtained.

21. The device according to any one of claims 14 to 20, wherein: The second generating module is further used for: Extracting a hand region image of the character from the illustration; Performing deformity recognition on the hand region image to determine deformity recognition information of the hand; In response to the deformity identification information indicating that the character's hand is deformed, the illustration corresponding to the plot information is updated.

22. The device according to claim 21, wherein The second generating module is further used for: Determine the hand deformity value and deformity position of the character according to the deformity identification information; In response to the hand deformity value being greater than or equal to a first set value, and the hand deformity value being less than or equal to a second set value, redrawing the deformed part to obtain redrawing data of the deformed part; According to the redrawing data of the deformed part, the deformed part is replaced to update the illustration corresponding to the plot information; or, In response to the hand deformity value being greater than the second set value, the illustration corresponding to the plot information is globally redrawn.

23. The device according to claim 22, wherein: The second generating module is further used for: Based on the position of the detection frame of the deformed part, performing occlusion processing on the deformed part; A hand prompt word and a redrawing constraint condition are obtained, and based on the hand prompt word and the redrawing constraint condition, details of the blocked part are repaired to obtain redrawing data of the deformed part.

24. The device according to any one of claims 14 to 20, wherein: The second generating module is further used for: Determine the style of illustrations based on the subject matter of the literary creation; Determining a target style fine-tuning parameter that matches the illustration style information from the style fine-tuning parameters corresponding to different styles; The diffusion model is updated based on the target painting style fine-tuning parameters to obtain the target diffusion model.

25. The device according to claim 24, wherein: The second generating module is further used for: Second training samples under different painting styles are obtained, and Lora fine-tuning is performed on the diffusion model based on the second training samples. After the fine-tuning is completed, painting style fine-tuning parameters corresponding to the different painting styles are obtained.

26. The device according to any one of claims 14 to 20, wherein: The second generating module is further used for: The position information of the plot information is determined, and based on the position information, an insertion position of an illustration corresponding to the plot information is determined, and the illustration is inserted at the insertion position.

27. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 13.

28. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-13.

29. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the method according to any one of claims 1 to 13 is implemented.