Method and apparatus for generating illustration, and device and storage medium

The automated generation of e-book illustrations through machine learning models solves the problem of high cost and low efficiency of manually generated illustrations, achieves efficient generation of high-quality illustrations, and increases users' interest in reading.

WO2025199829A1PCT designated stage Publication Date: 2025-10-02BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/084222
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-27
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

In the existing technology, generating e-book illustrations requires a lot of manpower and is inefficient, making it difficult to quickly and conveniently increase users' reading interest.

Method used

Through machine learning models, illustrations are generated based on the descriptive information and character information of text works, including descriptive information extraction of text fragments, character information acquisition, character image generation and illustration generation, and illustrations are automatically generated using trained neural network models.

Benefits of technology

While ensuring the quality of illustrations, the efficiency of generating illustrations is improved, the interest of e-books is enhanced, and the user reading experience is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024084222_02102025_PF_FP_ABST
    Figure CN2024084222_02102025_PF_FP_ABST
Patent Text Reader

Abstract

A method and apparatus for generating an illustration, and a device and a storage medium. The method comprises: on the basis of a text work, determining description information of a text segment and character information of at least one character, wherein the text work comprises a plurality of text segments, and the character information comprises attribute information of the character in at least one dimension (510); for a character among the at least one character, generating a character image of the character on the basis of the character information of the character (520); and on the basis of the description information of the text segment and the character image of the character, generating an illustration of the text segment (530). By inserting illustrations into a text work on the basis of the content of the text work, image data can be automatically generated for the text work, thereby increasing the interest for readers when reading the text work.
Need to check novelty before this filing date? Find Prior Art

Description

Method, device, apparatus and storage medium for generating illustrations Technical Field

[0001] Example embodiments of the present disclosure generally relate to the field of computers, and more particularly, to methods, devices, apparatuses, and computer-readable storage media for generating illustrations for textual works. Background Art

[0002] With the development of digital technology, more and more applications and websites are capable of presenting electronic publications, also known as e-books. To enhance user interest in e-books (particularly textual works such as novels, essays, prose, poetry, and scripts), illustrations can be inserted into e-books to increase their appeal and, in turn, enhance user interest. It is desirable to be able to quickly and easily access illustrations associated with e-books.

[0003] Summary of the Invention

[0004] In a first aspect of the present disclosure, a method for generating illustrations is provided. The method comprises: determining, from a text work, descriptive information of a text segment and character information of at least one character; wherein the text work includes multiple text segments, and the character information includes attribute information of the character in at least one dimension; generating, for each of the at least one characters, a character image based on the character information; and generating an illustration for the text segment based on the descriptive information of the text segment and the character image.

[0005] In a second aspect of the present disclosure, a device for generating illustrations is provided. The device includes: a determination module configured to determine, based on a text work, descriptive information of a text segment and character information of at least one character; wherein the text work includes multiple text segments, and the character information includes attribute information of the character in at least one dimension; an image generation module configured to generate, for a character in the at least one character, a character image based on the character information; and an image generation module configured to generate an illustration for the text segment based on the descriptive information of the text segment and the character image.

[0006] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the method of the first aspect.

[0007] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein a computer program is stored on the computer-readable storage medium, and the computer program can be executed by a processor to implement the method of the first aspect.

[0008] It should be understood that the content described in this summary section is not intended to limit the key features or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:

[0010] FIG1 shows a schematic diagram of an example environment in which embodiments of the present disclosure can be implemented;

[0011] 2A and 2B are schematic diagrams of example architectures for generating illustrations for textual works, respectively, according to some embodiments of the present disclosure;

[0012] FIG3 shows a schematic diagram of an example character image according to some embodiments of the present disclosure;

[0013] FIG4 shows a schematic diagram of an example illustration according to some embodiments of the present disclosure;

[0014] FIG5 illustrates a flowchart of a process for generating illustrations for a textual work according to some embodiments of the present disclosure;

[0015] FIG6 shows a schematic structural block diagram of an apparatus for generating illustrations for a text work according to some embodiments of the present disclosure; and

[0016] FIG7 illustrates a block diagram of an electronic device in which one or more embodiments of the present disclosure may be implemented. DETAILED DESCRIPTION

[0017] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0018] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, i.e., "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may be included below.

[0019] Herein, unless explicitly stated otherwise, executing a step “in response to A” does not mean executing the step immediately after “A” but may include one or more intermediate steps.

[0020] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.

[0021] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0022] For example, in response to receiving a user's active request, a prompt message is sent to the user to clearly remind the user that the operation requested to be performed will require obtaining and using the user's personal information, so that the user can independently choose whether to provide personal information to the electronic device, application, server or storage medium and other software or hardware that performs the operation of the technical solution of the present disclosure based on the prompt message.

[0023] As an optional but non-limiting implementation, in response to receiving a user's active request, a prompt message may be sent to the user, for example, in the form of a pop-up window, in which the prompt message may be presented in text form. Furthermore, the pop-up window may also include a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0024] It is understandable that the above notification and the process of obtaining user authorization are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0025] As used herein, the term "model" can learn the association between corresponding inputs and outputs from training data, so that after training is completed, corresponding outputs can be generated for given inputs. The generation of the model can be based on machine learning technology. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multiple layers of processing units. A neural network model is an example of a model based on deep learning. In this article, "model" may also be referred to as "machine learning model", "learning model", "machine learning network" or "learning network", and these terms are used interchangeably in this article.

[0026] A "neural network" is a machine learning network based on deep learning. A neural network is capable of processing inputs and providing corresponding outputs. It typically includes an input layer, an output layer, and one or more hidden layers between the input and output layers. Neural networks used in deep learning applications typically include many hidden layers, thereby increasing the depth of the network. The layers of a neural network are connected in sequence so that the output of the previous layer is provided as input to the next layer, where the input layer receives the input of the neural network and the output of the output layer serves as the final output of the neural network. Each layer of a neural network includes one or more nodes (also called processing nodes or neurons), each of which processes the input from the previous layer.

[0027] Generally speaking, machine learning can be roughly divided into three stages, namely the training stage, the testing stage, and the application stage (also called the inference stage). In the training stage, a given model can be trained using a large amount of training data, and the parameter values ​​are continuously updated iteratively until the model can obtain consistent inferences that meet the expected goals from the training data. Through training, the model can be considered to be able to learn the association between input and output (also called input-to-output mapping) from the training data. The parameter values ​​of the trained model are determined. In the testing stage, the test input is applied to the trained model to test whether the model can provide the correct output, thereby determining the performance of the model. The testing stage can sometimes be integrated into the training stage. In the application or inference stage, the trained model can be used to process the actual model input based on the parameter values ​​obtained through training to determine the corresponding model output.

[0028] As briefly mentioned above, to enhance user interest when browsing e-books, illustrations can be inserted into e-books to increase their appeal, thereby increasing user interest. Traditionally, staff members manually draw illustrations for novels and insert them into the novels. This requires significant labor costs and is inefficient in generating illustrations. For ease of description, the following textual works will only use novels as examples. Alternatively and / or additionally, textual works may include, but are not limited to, novels, short essays, prose, poetry, and screenplays.

[0029] According to an embodiment of the present disclosure, a method for generating illustrations is proposed. According to the solution of the embodiment of the present disclosure, descriptive information of a text segment and character information of at least one character are determined based on a text work; the text work includes multiple text segments, and the character information includes attribute information of the character in at least one dimension. For each of the at least one character, a character image is generated based on the character information. An illustration for the text segment is generated based on the descriptive information of the text segment and the character image.

[0030] In this way, multiple illustrations can be generated quickly and easily based on the text content of the novel while ensuring the quality of the illustrations, which can improve the efficiency of generating illustrations. In addition, inserting illustrations into the novel based on the novel content can increase the interest of readers when reading the novel.

[0031] FIG1 shows a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. As shown in FIG1 , the environment 100 may include an electronic device 110 .

[0032] The electronic device 110 may obtain the target novel 102 and generate at least one illustration 112 that matches the target novel 102 (for example, it may include illustrations 112-1, 112-2, ..., 112-N, where N is a positive integer. For ease of description, the one or more illustrations may be collectively referred to as illustrations 112 below). In some embodiments, the electronic device 110 may obtain the text content of the target novel 102 and generate the at least one illustration 112 based on the obtained text content of the target novel 102. In some embodiments, if the at least one illustration 112 includes multiple illustrations, different illustrations may correspond to different text segments of the target novel 102.

[0033] In some embodiments, the electronic device 110 can generate at least one illustration 112 that matches the target novel 102 with the help of a trained machine learning model 120. The machine learning model can be, for example, an image generation model. The machine learning model 120 can include, for example, but is not limited to, any appropriate model such as a Transformer model, a convolutional neural network (CNN), a recurrent neural network (RNN), a deep neural network (DNN), etc. The machine learning model 120 can be a model local to the electronic device 110, or a model installed in another electronic device 110 (for example, installed in a remote device). It should be noted that the machine learning model 120 can include multiple models, and the present disclosure does not limit the number and type of models specifically included in the machine learning model 120.

[0034] The electronic device 110 may include any computing system with computing capabilities, such as various computing devices / systems, terminal devices, server devices, etc. The terminal device may be any type of mobile terminal, fixed terminal or portable terminal, including a mobile phone, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a handheld computer, a portable game terminal, a VR / AR device, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a gaming device or any combination thereof, including accessories and peripherals of these devices or any combination thereof. The server device may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks, and big data and artificial intelligence platforms. The server-side device may include, for example, a computing system / server, such as a mainframe, an edge computing node, a computing device in a cloud environment, and the like.

[0035] It should be understood that the structure and function of each element in the environment 100 are described for exemplary purposes only and do not imply any limitation on the scope of the present disclosure. Some exemplary embodiments of the present disclosure will be described below with reference to the accompanying drawings.

[0036] Figure 2A shows a schematic diagram of an example architecture 200A for generating illustrations for a text work according to some embodiments of the present disclosure. The architecture 200A can be implemented at the electronic device 110. For ease of discussion, the architecture 200A will be described with reference to the environment 100 of Figure 1. Figure 2A shows an overview of the illustration generation process, where the text work 201 can include multiple text segments 202-1, ..., 202-N (individually and / or collectively referred to as text segments 202). Descriptive information 204 of a text segment (e.g., text segment 202-1) can be determined based on the text work 201. Role information 203 of at least one character of the text work 201 can be determined, and the role information 203 includes attribute information of the character in at least one dimension. Here, the dimensions can include, for example, the name, gender, age, occupation, appearance, expression, and clothing of the character, etc. For a character in at least one character, a character image diagram 205 of the character can be generated based on the character information 203 of the character.

[0037] Furthermore, illustrations of the text segment 202 can be generated based on the description information 204 of the text segment 202 and the character image 205. In this way, corresponding illustrations can be generated for each text segment in the text work, thereby increasing the interest of readers when reading the text work.

[0038] In some embodiments, the text work 201 may include multiple text fragments, and each text fragment may be determined based on the following method: obtaining structural information of the text work, and dividing the text work into the multiple text fragments based on the structural information. Here, the structural information may include the directory structure of the text work, for example, the multiple text fragments may be determined according to the hierarchy of multi-level titles defined in the directory structure. Rules for dividing text fragments may be pre-specified, for example, a text fragment may include a "chapter", a "section", or one or more paragraphs, etc. In this way, text fragments may be divided according to different precisions, and illustrations that better match the content of the text fragment may be generated.

[0039] For more details on illustration generation, see FIG2B , which illustrates a schematic diagram of an example architecture 200B for generating illustrations for a text work according to some embodiments of the present disclosure. The architecture 200B may be implemented at the electronic device 110 . For ease of discussion, the architecture 200B will be described with reference to the environment 100 of FIG1 .

[0040] As shown in FIG2B , architecture 200B includes a description information extraction unit 210 and a character information acquisition unit 220. The description information extraction unit 210 can be used, for example, to extract description information 215 (e.g., a summary) from a text segment of a novel. In some embodiments, the electronic device 110 can obtain various text segments from the novel and provide them to the description information extraction unit 210. The electronic device 110 can then process the various text segments based on predetermined rules to obtain description information for the multiple text segments.

[0041] The predetermined rules here may, for example, indicate the maximum number of text units (e.g., chapters, paragraphs, sentences, or words, etc.) included in each text segment. For example, the predetermined rules may indicate that each text segment may include at most one chapter, one section, one paragraph, or 50 text units, etc. Based on such predetermined rules, the full text of the novel may be segmented (e.g., with one chapter, one section, one paragraph, or 50 text units as a text segment) to obtain multiple text segments. The description information extraction unit 210 may then extract description information corresponding to each of the multiple text segments.

[0042] In some embodiments, the text segment here may be, for example, a text segment within a predetermined range of the current page presented in the e-book reader. The electronic device 110 may, for example, obtain a text segment within a predetermined range of the current page presented in the e-book reader in response to a page turning operation detected in the e-book reader, and provide the text segment to the description information extraction unit 210. It can be understood that the current page presents at least one text segment. The text segment within the predetermined range of the current page obtained by the electronic device 110 is at least a portion of at least one text segment. For example, if the current page includes two text segments, the electronic device 110 may obtain the text segment presented in the upper half of the current page. In this case, the description information extraction unit 210 may, for example, extract description information 215 for the text segment only from the obtained text segment. Thus, the electronic device 110 can obtain text segments and generate illustrations as the reader turns pages.

[0043] In some embodiments, the description information may include environmental information about the character's environment and the character's action information. The description information extraction unit 210 may, for example, determine description information 215 for a text segment by summarizing at least one character in the text segment and the environmental information and action information associated with the at least one character. For example, for text segment A, "Character A woke up early, made breakfast, put away the mess of toys in the living room, mopped the floor, and then took two steamed buns and went out."

[0044] The description information extraction unit 210 may determine that text segment A only includes character A. The description information generated by the description information extraction unit 210 may include, for example, the environment information "home," and the action information, for example, "Character A gets up early to do housework." It should be noted that not every text segment includes characters. For example, text segment B, "From now on, they will split the bill, regardless of expenses, half each," does not include either a character or an action associated with the character. Therefore, the description information extraction unit 210 may not generate description information corresponding to text segment B.

[0045] The character information acquisition unit 220 can acquire character information 225 for at least one character. This at least one character can be determined by the character information acquisition unit 220 based on the entire text or the current text segment, or it can be determined by the electronic device 110 and provided to the character information acquisition unit 220. Specifically, the electronic device 110 / character information acquisition unit 220 can determine at least one character in the novel. For example, if the description information for the three text segments includes "Character A gets up early to do housework," "Character A hears Mom and Dad arguing," and "Character A decides to go out for a walk," the electronic device 110 / character information acquisition unit 220 can determine that the three characters are "Character A," "Dad," and "Mom."

[0046] In some embodiments, the role information acquisition unit 220 can also determine the number of occurrences of the target role in the text segment for the target role among at least one role, and then determine the main role. For example, the role information of the target role can be acquired in response to determining that the number of occurrences meets a predetermined condition. The predetermined condition here can, for example, indicate a predetermined number of times (for example, 3 times, 5 times, or any other number), and the role information acquisition unit 220 can, for example, acquire the role information of the target role in response to determining that the number of occurrences reaches a predetermined number. In this way, only the role information of the characters with a large number of occurrences can be acquired, which can reduce the final illustration generation cost. Alternatively and / or additionally, assuming that the text segment only includes one character and the number of occurrences of the character is lower than the predetermined number, the role information of the character can still be acquired.

[0047] The role information 225 of at least one role may include multiple attributes of the at least one role. The multiple attributes here may include, for example, any one or more of the role's name, gender, age, occupation, appearance, expression, and clothing.

[0048] In some embodiments, to determine the character information of at least one character, the character information of the target character can be determined based on the portion of the text work associated with the target character. Furthermore, the character information of the target character can be updated based on the portion of the text segment associated with the target character. For example, the basic attributes of the character, such as name, gender, age, and occupation, can be determined from the entire text work. Furthermore, the special attributes of the character in the current segment of the text being processed, such as the current appearance, expression, and clothing, etc., can be determined.

[0049] For example, if text segment 1 depicts a winter scene, then based on this text segment 1, the character's attire can be determined to be a "coat." If text segment 2 depicts a summer scene, then based on this text segment 2, the character's attire can be determined to be a "dress." A mapping relationship can exist between character information and each text segment in the text work. This ensures that the character information matches the character's basic characteristics and reflects the character's current state as the story progresses within the text work.

[0050] In some embodiments, based on a specific determination method, multiple attributes can be divided into two parts, wherein the first part can be directly determined from the novel, and the second part can be indirectly determined from the novel, or can be manually set. For example, the name, gender, etc. of the character in the multiple attributes can be the attributes of the first part. The character information acquisition unit 220 can, for example, determine the attributes of the first part from the novel. It should be noted that the character information acquisition unit 220 can obtain the first part of the attributes of character A from the full text of the novel, and is not limited to the text fragment.

[0051] For example, if the novel explicitly describes a character's appearance, demeanor, clothing, and other attributes, these attributes can be used as the attributes of the first portion. If the novel does not include text associated with the attributes of the second portion, the electronic device 110 can receive user input from a user (e.g., a relevant staff member) and determine the attributes of the second portion based on the user input. For example, the electronic device 110 can provide a setting control in an e-book reader for setting the second portion of the character's multiple attributes, and in response to receiving a setting operation on the setting control, set the second portion of the multiple attributes based on the setting operation. Such a setting control can be, for example, an input box. The electronic device 110 can, for example, receive user input via the input box and determine the second portion of the multiple attributes based on the user input. The electronic device 110 can, for example, provide the determined attributes of the second portion to the character information acquisition unit 220 so that the character information acquisition unit 220 acquires the attributes of the second portion. In this way, the user is allowed to specify the attributes of the character (e.g., clothing style and color, etc.) according to their needs during the reading process, thereby generating illustrations that meet their needs.

[0052] The electronic device 110 can generate illustrations 112 of the novel based on the description information 215 of the text segment and the role information 225 of at least one character. The electronic device 110 can generate the illustrations 112 based on any appropriate method and using the description information 215 of the text segment and the role information 225 of at least one character. The present disclosure does not limit the specific method of generating illustrations. For example, the electronic device 110 can generate the illustrations 112 based on pre-acquired rules or algorithms. In some embodiments, the electronic device 110 can generate the illustrations with the help of a trained machine learning model. In this case, the architecture 200B can also include a prompt word determination unit 230 and a machine learning model 120.

[0053] The prompt word determination unit 230 can, for example, be configured to generate prompt words 235 for the machine learning model 120 based on the description information and the role information. The prompt word determination unit 230 can, for example, obtain a predetermined prompt word template and fill the prompt word template with the description information and the role information to generate the prompt word 235. For example, the prompt word template can include: environmental information, role information, and action information. The obtained various information can be filled into the corresponding positions of the template to generate the prompt word.

[0054] In some embodiments, in order to ensure the uniformity of the characters in subsequently generated illustrations, the prompt word determination unit 230 may also call the character image determination unit 250 and the character model generation unit 260. Here, the character image may, for example, represent the character image of the character from multiple angles, and the character image determination unit 250 may, for example, determine the character image 255 for at least one character based on the character information of at least one character in the novel. The character image determination unit 250 may determine the character image 255 in any appropriate manner. For example, the character image determination unit 250 may generate the character image 255 for at least one character based on the character information of at least one character using a trained image generation model. Alternatively or additionally, in some embodiments, the character image determination unit 250 may also directly obtain the character image input by the user (e.g., the character image drawn by the illustrator for at least one character).

[0055] FIG3 illustrates a schematic diagram of an example character image 300 according to some embodiments of the present disclosure. Character image determination unit 250 may, for example, generate character image 300 for character A based on character information for character A, such as "character A, 25 years old, female, curly hair, wearing a dress." Character image 300 may include multiple images of character A at various angles (e.g., image 301 tilted 45 degrees to the side, a side view image 302, a front view image 303, and a back view image 304).

[0056] For each character, the character model generation unit 260 can generate a character model 265 describing the character based on the character image 255 corresponding to the character. Character model 265 can be, for example, a LoRA model. The LoRA model can be understood as a plug-in to the Stable Diffusion (SD) model (a generative model), which can be used to meet a specific style or specified character attributes.

[0057] The process of generating a character model based on a character image can be understood as storing the character image in the form of a character model. The prompt word determination unit 230 can subsequently flexibly call different character models to call different character images. The prompt word determination unit 230 can obtain a character model 265 for at least one character and determine a prompt word 235 based on the character model 265 and the description information 215 of the text segment. For example, the prompt word determination unit 230 can generate the prompt word "Character A <Model A> makes breakfast at home" based on the description information "Character A gets up early to do housework" and the character model A corresponding to character A.

[0058] In some embodiments, the prompt word determination unit 230 can also update the prompt word 235 based on the weight index of the character model. The weight index can be used to indicate the similarity between the character in the illustration and the character image corresponding to the character. For example, if the prompt word is "Character A <Model A, 0.5> makes breakfast at home", then the prompt word indicates that the similarity between the character A in the subsequently generated illustration and the character image corresponding to character A is 50%. It can be understood that the higher the weight index, the higher the similarity between the character in the illustration and the character image corresponding to the character, and the more similar the two are. Using the embodiments of the present disclosure, the character details in each illustration can be adjusted while ensuring the consistency of the appearance of the novel characters.

[0059] In some embodiments, the prompt word determination unit 230 can also determine the style of the illustration based on the background environment of the novel, and update the prompt word 235 based on the style. For example, if the background of the novel is a modern urban background, the style of the illustration can be determined to be "comic style", and the prompt word can be, for example, "Character A <Model A, 0.5> makes breakfast at home in comic style." If the background of the novel is an ancient martial arts background, the style of the illustration can be determined to be "ink style", and the prompt word can be, for example, "Character A <Model A, 0.5> makes breakfast at home in ink style." Using the embodiments of the present disclosure, illustrations with richer visual effects can be generated in a more flexible manner.

[0060] The prompt word determination unit 230 may provide the determined prompt word 235 to the machine learning model 120. The machine learning model 120 may then generate an illustration 112 based on the acquired prompt word 235. For example, FIG4 shows a schematic diagram of an example illustration 400 according to some embodiments of the present disclosure. If the prompt word is "Character A <Model A, 0.5> making breakfast at home," the machine learning model 120 may call Model A and generate an illustration 400 showing Character A making breakfast at home based on a weight index of 0.5.

[0061] In some embodiments, the electronic device 110 may obtain at least one illustration corresponding to at least one character (each character may correspond to multiple illustrations), and insert the illustration into a position in the novel associated with the text segment. For example, if the electronic device 110 generates illustration A based on text segment A, the electronic device 110 may insert illustration A into text segment A (for example, inside text segment A, before / after text segment A, etc.). The electronic device 110 may also, for example, adjust the illustration based on the adjustment operation in response to receiving an adjustment operation for adjusting the illustration. The adjustment operation here may, for example, include an update operation, a delete operation, and / or a move operation for the illustration. For example, the electronic device 110 may delete illustration A inserted into the novel in response to receiving a delete operation from the user for illustration A. The electronic device 110 may move illustration A in response to receiving a move operation from the user to move illustration A from position A to position B, and the moved illustration A is located at position B.

[0062] In summary, according to the embodiments of the present disclosure, multiple illustrations can be generated based on the text content of a novel conveniently and quickly while ensuring the quality of the illustrations, which can improve the efficiency of generating illustrations. In addition, inserting illustrations into a novel based on the content of the novel can increase the interest of readers when reading the novel.

[0063] The above describes the specific details of each step of generating an illustration for a text work, providing a method for generating an illustration for a text work. FIG5 shows a flowchart of a process 500 for generating an illustration for a text work according to some embodiments of the present disclosure. Process 500 can be implemented at electronic device 110. Process 500 is described below with reference to FIG1.

[0064] At block 510 , description information of a text segment and role information of at least one role are determined based on the text work; wherein the text work includes a plurality of text segments, and the role information includes attribute information of the role in at least one dimension.

[0065] At block 520 , for a role in at least one of the roles, a character avatar diagram of the role is generated based on the role information of the role.

[0066] At block 530 , an illustration of the text segment is generated based on the description information of the text segment and the character image of the character.

[0067] In some embodiments, the plurality of text segments is determined based on: obtaining structural information of the text work; and dividing the text work into the plurality of text segments based on the structural information.

[0068] In some embodiments, at least one role is determined based on: determining the number of occurrences of a target role among multiple roles in a text work in a text segment; and in response to determining that the number of occurrences meets a predetermined condition, using the target role as a role among at least one role.

[0069] In some embodiments, determining role information of at least one role includes: determining the role information of the target role based on a portion of the text work associated with the target role; and updating the role information of the target role based on a portion of the text fragment associated with the target role.

[0070] In some embodiments, the descriptive information includes environmental information of the character's environment and action information of the character, and generating illustrations of text fragments includes: generating a character model for describing the character based on the character image; generating prompt words for the machine learning model using the environmental information, action information and character model; and generating illustrations based on the prompt words.

[0071] In some embodiments, generating the prompt word further includes: updating the prompt word based on the weight index of the role model.

[0072] In some embodiments, generating the prompt word further includes: determining a style of the illustration based on the context of the text work; and updating the prompt word based on the style.

[0073] In some embodiments, process 500 further includes: inserting an illustration into a location in the text work associated with the text fragment; and in response to receiving an adjustment operation for adjusting the illustration, adjusting the illustration based on the adjustment operation, the adjustment operation including at least any one of: an update operation, a delete operation, and a move operation for the illustration.

[0074] In some embodiments, the attribute information of at least one dimension includes at least any one of the following: the character's name, gender, age, occupation, appearance, expression, and clothing, and the first part of the attribute information of at least one dimension is determined from the text work.

[0075] In some embodiments, the process 500 is implemented in an electronic book reader for reading a text work, and the process 500 further includes: providing a setting control in the electronic book reader for setting a second part of the attribute information of at least one dimension; and in response to receiving a setting operation for the setting control, setting the second part of the attribute information of at least one dimension based on the setting operation.

[0076] According to some embodiments of the present disclosure, the text segment is a text segment within a predetermined range of a current page presented in an electronic book reader, and the method is performed in response to a page turning operation detected in the electronic book reader.

[0077] According to some embodiments of the present disclosure, a device for generating illustrations for a text work is also provided. Figure 6 shows a schematic block diagram of a device 600 for generating illustrations for a text work according to some embodiments of the present disclosure. Device 600 can be implemented as or included in electronic device 110. The various modules / components in device 600 can be implemented using hardware, software, firmware, or any combination thereof.

[0078] As shown in Figure 6, the device 600 includes: a determination module 610, which is configured to determine the description information of the text segment and the role information of at least one character based on the text work; wherein the text work includes multiple text segments, and the role information includes attribute information of the character in at least one dimension; an image generation module 620, which is configured to generate a character image diagram of the character based on the role information of the character in at least one character; and an image generation module 630, which is configured to generate an illustration of the text segment based on the description information of the text segment and the role image diagram of the character.

[0079] In some embodiments, the plurality of text segments is determined based on: obtaining structural information of the text work; and dividing the text work into the plurality of text segments based on the structural information.

[0080] In some embodiments, at least one role is determined based on: determining the number of occurrences of a target role among multiple roles in a text work in a text segment; and in response to determining that the number of occurrences meets a predetermined condition, using the target role as a role among at least one role.

[0081] In some embodiments, the determination module 610 is further configured to: determine the role information of the target role based on the portion of the text work associated with the target role; and update the role information of the target role based on the portion of the text segment associated with the target role.

[0082] In some embodiments, the descriptive information includes environmental information of the character's environment and action information of the character, and the image generation module is further configured to: generate a character model for describing the character based on the character image; generate prompt words for the machine learning model using the environmental information, action information and character model; and generate illustrations based on the prompt words.

[0083] In some embodiments, the image generation module 630 is further configured to update the prompt word based on the weight index of the role model.

[0084] In some embodiments, the image generation module 630 is further configured to: determine the style of the illustration based on the context of the textual work; and update the prompt word based on the style.

[0085] In some embodiments, the device 600 further includes: an insertion module configured to insert an illustration into a position associated with a text fragment in a text work; and an adjustment module configured to adjust the illustration based on the adjustment operation in response to receiving an adjustment operation for adjusting the illustration, the adjustment operation including at least any one of the following: an update operation, a deletion operation, and a move operation for the illustration.

[0086] In some embodiments, the attribute information of at least one dimension includes at least any one of the following: the character's name, gender, age, occupation, appearance, expression, and clothing, and the first part of the attribute information of at least one dimension is determined from the text work.

[0087] In some embodiments, the device 600 is implemented in an electronic book reader for reading text works, and the device 600 further includes: a providing module configured to provide a setting control in the electronic book reader for setting the second part of the attribute information of at least one dimension; and a setting module configured to set the second part of the attribute information of at least one dimension based on the setting operation in response to receiving a setting operation for the setting control.

[0088] In some embodiments, the text snippet is a text snippet within a predetermined range of a current page presented in an electronic book reader, and the apparatus is invoked in response to a page turning operation detected in the electronic book reader.

[0089] The units and / or modules included in the device 600 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units and / or modules can be implemented using software and / or firmware, such as machine executable instructions stored on a storage medium. In addition to or as an alternative to machine executable instructions, some or all of the units and / or modules in the device 600 can be implemented at least in part by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that can be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0090] FIG7 shows a block diagram of an electronic device 700 in which one or more embodiments of the present disclosure may be implemented. It should be understood that the electronic device 700 shown in FIG7 is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. The electronic device 700 shown in FIG7 may be used to implement the electronic device 110 of FIG1 and / or the apparatus 600 of FIG6 .

[0091] As shown in FIG7 , electronic device 700 is in the form of a general-purpose computing device. Components of electronic device 700 may include, but are not limited to, one or more processors or processing units 710, memory 720, storage device 730, one or more communication units 740, one or more input devices 750, and one or more output devices 760. Processing unit 710 may be a real or virtual processor and is capable of performing various processes according to programs stored in memory 720. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to increase the parallel processing capabilities of electronic device 700.

[0092] The electronic device 700 typically includes a plurality of computer storage media. Such media can be any available media accessible to the electronic device 700, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 720 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory) or some combination thereof. The storage device 730 can be a removable or non-removable medium and can include a machine-readable medium, such as a flash drive, a disk or any other medium, which can be used to store information and / or data and can be accessed within the electronic device 700.

[0093] The electronic device 700 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 7 , a disk drive for reading from or writing to a removable, non-volatile disk (e.g., a “floppy disk”) and an optical drive for reading from or writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. The memory 720 may include a computer program product 725 having one or more program modules configured to perform various methods or actions of various implementations of the present disclosure.

[0094] The communication unit 740 enables communication with other computing devices via a communication medium. Additionally, the functions of the components of the electronic device 700 can be implemented as a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the electronic device 700 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.

[0095] Input device 750 may be one or more input devices, such as a mouse, keyboard, or trackball. Output device 760 may be one or more output devices, such as a display, a speaker, or a printer. Electronic device 700 may also communicate with one or more external devices (not shown) via communication unit 740, as needed, such as storage devices, display devices, or the like, with one or more devices that allow a user to interact with electronic device 700, or with any device that allows electronic device 700 to communicate with one or more other computing devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).

[0096] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above.

[0097] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0098] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0099] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0100] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple implementations of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part for a module, program segment or instruction, and a part for a module, program segment or instruction comprises one or more executable instructions for realizing the logical function of the specification. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.

[0101] While various implementations of the present disclosure have been described above, the foregoing description is intended to be illustrative, not exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is selected to best explain the principles of the implementations, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A method for generating an illustration, comprising: Determining, based on a text work, descriptive information of a text segment and role information of at least one character; wherein the text work includes a plurality of the text segments, and the role information includes attribute information of the character in at least one dimension; For a role among the at least one role, generating a role avatar image of the role based on the role information of the role; and An illustration of the text segment is generated according to the description information of the text segment and the character image of the character.

2. The method of claim 1 , wherein the plurality of text segments are determined based on: Obtaining structural information of the textual work; and The text work is divided into the plurality of text segments based on the structural information.

3. The method of claim 1 , wherein the at least one role is determined based on: For a target character among the multiple characters in the text work, determining the number of occurrences of the target character in the text segment; and In response to determining that the number of occurrences satisfies a predetermined condition, the target character is selected as a character in the at least one character.

4. The method according to claim 3, wherein determining the role information of the at least one role comprises: For the target character, determining the character information of the target character based on a portion of the text work associated with the target character; as well as Based on the portion of the text segment associated with the target role, the role information of the target role is updated.

5. The method according to claim 1, wherein the description information includes environmental information of the environment in which the character is located and action information of the character, and generating the illustration of the text segment comprises: generating a role model for describing the role based on the role image; generating prompt words for a machine learning model using the environmental information, the action information, and the role model; as well as The illustration is generated based on the prompt word.

6. The method according to claim 5, wherein generating the prompt word further comprises: The prompt word is updated based on the weight index of the role model.

7. The method according to claim 5, wherein generating the prompt word further comprises: determining the style of the illustration based on the context of the textual work; as well as The prompt word is updated based on the style.

8. The method according to claim 1, further comprising: inserting the illustration into the text work at a location associated with the text fragment; as well as In response to receiving an adjustment operation for adjusting the illustration, the illustration is adjusted based on the adjustment operation, where the adjustment operation includes at least any one of the following: an update operation, a delete operation, and a move operation for the illustration.

9. The method according to claim 1, wherein the attribute information of at least one dimension includes at least any one of the following: the name, gender, age, occupation, appearance, expression, and clothing of the character, and the first part of the attribute information of at least one dimension is determined from the text work.

10. The method according to claim 9, wherein the method is implemented in an electronic book reader for reading the text work, and the method further comprises: Providing a setting control in the electronic book reader for setting a second part of the attribute information of the at least one dimension; as well as In response to receiving a setting operation for the setting control, a second portion of the attribute information of the at least one dimension is set based on the setting operation.

11. The method of claim 10, wherein the text segment is a text segment within a predetermined range of a current page presented in the electronic book reader, and the method is performed in response to a page turning operation detected in the electronic book reader.

12. A device for generating illustrations, comprising: a determination module configured to determine, based on a text work, description information of a text segment and role information of at least one character; wherein the text work includes a plurality of the text segments, and the role information includes attribute information of the character in at least one dimension; an image generation module configured to generate, for a character among the at least one character, a character image of the character based on the character information of the character; and The image generation module is configured to generate an illustration of the text segment according to the description information of the text segment and the character image of the character.

13. An electronic device comprising: at least one processing unit; as well as At least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to any one of claims 1 to 11 when executed by the at least one processing unit.

14. A computer-readable storage medium having a computer program stored thereon, wherein the computer program can be executed by a processor to implement the method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Method and device for generating video based on story text

    CN108470036A

  • Image generation method and device, electronic equipment and storage medium

    CN116894881A

  • Information interaction processing method, device and equipment and computer storage medium

    CN116954437A

  • Content generation method and device, computer equipment and storage medium

    CN117171369A

  • Story video generation method and device, storage medium and equipment

    CN117332118A