Image generation method, apparatus, and related product

By acquiring the first and second texts related to the e-book, and using a large language model and an image generation model to generate images that match the e-book content, the problem of generating matching images efficiently and accurately in existing technologies is solved, thereby improving the user's reading experience and image generation efficiency.

CN119991882BActive Publication Date: 2026-04-17DOUYIN VISION CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
DOUYIN VISION CO LTD
Filing Date
2025-03-03
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently and accurately generate illustrations for e-book content.

Method used

By acquiring the first text related to the e-book, the target object is identified, and the second text related to the target object is acquired from the e-book. An image matching the first text is then generated using a large language model and an image generation model.

Benefits of technology

It enables efficient and accurate generation of illustrations for e-book content, improving the user's reading experience and image generation efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991882B_ABST
    Figure CN119991882B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide an image generation method, device and related product, wherein the method comprises: obtaining an image generation request for a first text; the first text is a text associated with an electronic book; in response to the image generation request, determining a target object described by the first text, and obtaining a second text related to the target object in the electronic book; generating image generation prompt information according to the first text and the second text, and generating an image matched with the first text according to the image generation prompt information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to an image generation method, apparatus and related products. Background Technology

[0002] In related technologies, AI (Artificial Intelligence) technology can be used to assist users in generating the images they need. For example, users input descriptive information about the desired image into a raw image model trained on AI technology, and the model generates the image based on that descriptive information. Considering that users need to generate illustrations for e-book content while reading, how to efficiently and accurately generate illustrations for e-book content has become one of the problems that needs to be solved. Summary of the Invention

[0003] This disclosure provides an image generation method, apparatus, and related products that can efficiently and accurately generate images for e-book content.

[0004] In a first aspect, embodiments of this disclosure provide an image generation method, including:

[0005] Obtain an image generation request for the first text; the first text is text associated with the e-book;

[0006] In response to the image generation request, the target object described by the first text is determined, and the second text related to the target object is obtained in the e-book;

[0007] Based on the first text and the second text, an image-generating prompt message is generated, and based on the image-generating prompt message, an image matching the first text is generated.

[0008] Secondly, embodiments of this disclosure provide an image generation apparatus, comprising:

[0009] The request acquisition unit is used to acquire an image generation request for a first text; the first text is text associated with an e-book;

[0010] A text acquisition unit is configured to, in response to the image generation request, determine the target object described by the first text, and acquire second text related to the target object in the e-book;

[0011] An image generation unit is configured to generate image generation prompt information based on the first text and the second text, and generate an image matching the first text based on the image generation prompt information.

[0012] Thirdly, embodiments of this disclosure provide an electronic device, including: a processor; and a memory configured to store computer-executable instructions, which, when executed, cause the processor to perform the method described in the first aspect above.

[0013] Fourthly, embodiments of this disclosure provide a computer-readable storage medium for storing computer-executable instructions that, when executed by a processor, implement the method described in the first aspect.

[0014] Fifthly, embodiments of this disclosure provide a computer program product, the computer program product including a computer program, which, when executed by a processor, implements the method described in the first aspect above.

[0015] In one or more embodiments of this disclosure, firstly, an image generation request for first text is obtained, where the first text is text associated with an e-book. Then, in response to the image generation request, a target object described by the first text is determined. In the e-book, second text related to the target object is obtained. Finally, based on the first and second texts, image generation prompt information is generated. Based on the image generation prompt information, an image matching the first text is generated. Therefore, through this embodiment, it is possible to obtain first text related to an e-book, determine the target object described by the first text, obtain second text related to the target object in the e-book, and generate an image matching the first text based on the first and second texts, achieving the effect of efficiently and accurately generating illustrations for relevant content in an e-book. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in one or more embodiments or related technologies of this disclosure, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments recorded in this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a schematic flowchart of an image generation method provided in an embodiment of the present disclosure;

[0018] Figure 2 A schematic diagram of a first text provided for an embodiment of this disclosure;

[0019] Figure 3 A schematic diagram of a first text provided for another embodiment of this disclosure;

[0020] Figure 4This is a schematic diagram of a scene generated by an embodiment of the present disclosure;

[0021] Figure 5 This is a schematic diagram illustrating the result of image generation according to an embodiment of the present disclosure;

[0022] Figure 6 A scene illustration of a custom raw image provided in an embodiment of this disclosure;

[0023] Figure 7 This is a schematic diagram illustrating the generation of an image matching a first text from an existing image in an e-book, as provided in an embodiment of this disclosure.

[0024] Figure 8 This is a schematic diagram of the structure of an image generation apparatus provided in an embodiment of the present disclosure;

[0025] Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure. Detailed Implementation

[0026] To enable those skilled in the art to better understand the technical solutions in one or more embodiments of this disclosure, the technical solutions in one or more embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of the embodiments. Based on one or more embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of this disclosure.

[0027] It is understood that before using the technical solutions disclosed in the embodiments of this disclosure, relevant parties should be informed of the type, scope of use, and usage scenarios of the information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and authorization from the relevant parties should be obtained.

[0028] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0029] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0030] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0031] This disclosure provides an image generation method, apparatus, and related products, capable of efficiently and accurately generating images for e-book content. The image generation method can be applied to a terminal device or a server, and is implemented by the terminal device or server. The terminal device includes, but is not limited to, various types of user terminals such as laptops, tablets, desktop computers, set-top boxes, mobile devices (e.g., mobile phones, portable music players, personal digital assistants, dedicated messaging devices, portable gaming devices), smartphones, smart speakers, smartwatches, smart TVs, and in-vehicle terminals. The server can be a single server or a server cluster.

[0032] Figure 1 This is a schematic flowchart of an image generation method provided in an embodiment of the present disclosure, as shown below. Figure 1 As shown, the process includes:

[0033] Step S102: Obtain an image generation request for the first text; the first text is text associated with the e-book;

[0034] Step S104: In response to the image generation request, determine the target object described by the first text, and obtain the second text related to the target object in the e-book;

[0035] Step S106: Generate image prompt information based on the first text and the second text, and generate an image that matches the first text based on the image prompt information.

[0036] In this embodiment, firstly, an image generation request for a first text, which is text associated with an e-book, is obtained. Then, in response to the image generation request, the target object described by the first text is determined. Within the e-book, second text related to the target object is obtained. Finally, based on the first and second texts, image generation prompts are generated. Based on these prompts, an image matching the first text is generated. Therefore, this embodiment can efficiently and accurately generate illustrations for the relevant content of an e-book by obtaining the first text related to the e-book, determining the target object described by the first text, obtaining the second text related to the target object within the e-book, and generating an image matching the first text based on both texts.

[0037] In step S102 above, the first text is text associated with the e-book, such as the original text in the e-book or user comments on the e-book. In this embodiment, if a user's image generation operation for the first text is detected, then based on the image generation operation, it is determined that a user's image generation request for the first text has been received. The image generation request is used to request the generation of a matching image for the first text. The e-books involved in the various embodiments of this disclosure can be any e-book read through a reading application, and are not limited here.

[0038] In some embodiments, obtaining an image generation request for the first text includes:

[0039] In response to a selection command for the main body of the e-book, the content to be selected is determined within the main body of the e-book.

[0040] If a first request to generate an image based on selected content is received, the selected content is determined to be the first text, and the first request is determined to be an image generation request for the first text.

[0041] In this embodiment, the user can select the main text of the e-book. If the user's selection operation on the main text of the e-book is detected, a selection instruction for the main text of the e-book is generated based on the selection operation. In response to the selection instruction, the selection content corresponding to the selection instruction is determined in the main text of the e-book, that is, the user's selection content is determined in the main text of the e-book.

[0042] Next, if a first request to generate an image based on the selected content is received, the user's selected content is determined to be first text, and the first request is determined to be an image generation request for the first text. For example, if a user trigger operation on an image generation control corresponding to the selected content is detected, based on the trigger operation, it is determined that a first request to generate an image based on the user's selected content has been received, the user's selected content is determined to be first text, and the first request is determined to be an image generation request for the first text.

[0043] Figure 2 A schematic diagram of a first text provided for an embodiment of this disclosure, as shown below. Figure 2 As shown, users can select content in the main text of the e-book by long-pressing. In response to the user's selection, the selected content is determined and the "intelligent image generation" control corresponding to the selected content is displayed. If the user triggers the control, it is determined that the user's first request to generate an image based on the selected content has been received, the selected content is determined as the first text, and the first request is determined as an image generation request for the first text. Figure 2 In the image, the original text of the ebook selected by the user is indicated by shading.

[0044] As can be seen, through this embodiment, users can select content of interest in the main text of an e-book and request the generation of matching images for that content, thereby improving the user's e-book reading experience and image generation efficiency.

[0045] In some embodiments, obtaining an image generation request for the first text includes:

[0046] In response to a command to input comments for an e-book, retrieve the comments input for the e-book.

[0047] If a second request to generate an image based on comment information is received, the comment information is identified as the first text, and the second request is identified as an image generation request for the first text.

[0048] In this embodiment, users can comment on e-books, such as commenting on chapters or paragraphs. If a user's comment input is detected, a comment input command is generated based on this input. In response to the comment input command, the user's comment input is acquired and displayed. The user's comment input can be for a chapter or a paragraph of the e-book.

[0049] Next, if a second request to generate an image based on comment information is received, the comment information entered by the user is determined as the first text, and the second request is determined as an image generation request for the first text. For example, if a user trigger operation on an image generation control corresponding to the input comment information is detected, based on the trigger operation, it is determined that a second request to generate an image based on comment information has been received, the comment information entered by the user is determined as the first text, and the second request is determined as an image generation request for the first text.

[0050] Figure 3 A schematic diagram of the first text provided for another embodiment of this disclosure, such as Figure 3 As shown, users can target paragraphs in the ebook ( Figure 3 (Shaded area) Input comment information. In response to the user's comment information input operation, display the user's input comment information and the "Smart Image Generation" control corresponding to the comment information. If the user triggers the control, it is determined that a second request to generate an image based on the comment information has been received. The user's input comment information is determined as the first text, and the second request is determined as an image generation request for the first text. Figure 3 The first text in the article includes the phrase "This plane is really cool".

[0051] In some embodiments, reference Figure 3, after the user triggers the paragraph comment entry of the e - book and enters the paragraph comment page, if the user triggers the "Intelligent Image Generation" control on the paragraph comment page without entering comment information, the paragraph ( Figure 3 shaded in the figure) for which the user wants to post comment information can be obtained, and this paragraph is used as the first text to generate a matching image for the first text.

[0052] In some embodiments, referring to Figure 3 , after the user triggers the paragraph comment entry of the e - book and enters the paragraph comment page, if the user triggers the "Intelligent Image Generation" control on the paragraph comment page while entering comment information, the paragraph ( Figure 3 shaded in the figure) for which the user wants to post comment information and the comment information entered by the user are jointly used as the first text to generate a matching image for the first text.

[0053] It can be seen that through this embodiment, the user can comment on the chapters or paragraphs in the e - book and request to generate a matching image for the entered comment information, improving the user's e - book reading experience and image generation efficiency.

[0054] In step S104 above, in response to the image generation request, determine the target object described by the first text. The first text has the described target object. When the first text is the original text of the e - book, determine the object described by the original text of the e - book as the target object. For example, the first text is the original text of the e - book selected by the user: "Her big eyes, long black straight hair, and ethereal temperament amazed everyone". According to the content of the e - book, it can be determined that the original text of the e - book describes the character A in the e - book, so the target object can be determined as the character A.

[0055] When the first text is the comment information for the e - book, if the first text has information indicating the described target object, determine the target object according to this information. If the first text does not have information indicating the described target object, determine the original text of the e - book associated with the first text in the e - book, and determine the target object according to the content of the first text and the associated original text of the e - book. For example, the first text is the comment information of the user for the paragraph in the e - book, and this comment information is related to character B, then the character B in the e - book is determined as the target object. Another example is that the first text is the comment information of the user for the paragraph in the e - book: "This description is amazing, full of a sense of picture". Then the paragraph associated with this comment information can be determined in the e - book, and the object in this paragraph is determined as the target object.

[0056] Next, the system retrieves a second text related to the target object from the ebook. When the first text is the original ebook text selected by the user, the second text may include text related to that original ebook text; when the first text is a comment entered by the user, the second text may include the paragraph or chapter to which the comment is directed.

[0057] In some embodiments, in an ebook, obtaining second text related to the target object includes:

[0058] Determine the scene scene corresponding to the first text; the target object is located within the scene scene;

[0059] In the e-book, obtain the first sub-text related to the above scene and the second sub-text related to the object form of the target object in the above scene, and use the first sub-text and the second sub-text as the second text.

[0060] As described above, the first text is either the original ebook text or commentary on the ebook. Regardless of whether the first text is the original ebook text or commentary, the second text can be obtained through the following methods.

[0061] First, determine the scene scene corresponding to the first text. When the first text describes a scene scene, the scene scene corresponding to the first text includes the scene scene described by the first text. For example, if the first text is a comment on a paragraph, then the scene scene corresponding to the first text includes the scene scene described by the comment. The target object, as the object described by the first text, resides in the scene scene corresponding to the first text.

[0062] When the first text is not used to describe a scene, the scene corresponding to the first text includes the scene described in the original ebook text associated with the first text. For example, if the first text is a comment, "This description is wonderful," the paragraph associated with that comment in the ebook is identified, and the scene described in that paragraph is taken as the scene corresponding to the first text. This scene could be, for example, multiple characters drinking together. The target object, as the object described in the first text, resides in the scene corresponding to the first text. In the example above, the target object is each object in the scene described in the paragraph.

[0063] Next, in the ebook, a first sub-text related to the aforementioned scene and a second sub-text related to the object form of the target object within the aforementioned scene are retrieved, and the first and second sub-texts are used as the second text. The first sub-text can represent the specific content of the aforementioned scene, or the specific content of a scene related to the aforementioned scene. The second sub-text can represent the object form of the target object within the aforementioned scene, or the object form of the target object in a scene related to the aforementioned scene. The object form includes the target object's appearance, clothing, hairstyle, posture, etc.

[0064] Taking the first text as a comment on a specific paragraph as an example, the comment is related to Little B, and the scene corresponding to the first text includes the scene described in the comment. Therefore, in the ebook, the first sub-text related to the scene described in the comment is retrieved. Additionally, the second sub-text related to Little B's object form in the aforementioned scene is retrieved. The second sub-text can be text describing Little B's appearance. The first and second sub-texts are then used as the second text.

[0065] Taking the comment information "This description is really wonderful" as the first text as the comment information of the paragraph, since the scene corresponding to the first text includes the scene described in the paragraph, the paragraph can be obtained as the first sub-text. If the paragraph describes the object form of each target object, the paragraph can also be obtained as the second sub-text. The first sub-text and the second sub-text are then used as the second text.

[0066] Taking the first text as the user's selected e-book text, "Her big eyes, long straight black hair, and otherworldly temperament amazed everyone," as an example, the scene corresponding to the first text is the scene where character A appears. In the e-book, the first sub-text related to the appearance of character A is obtained, and the second sub-text related to the appearance of character A when she appears is obtained. The first sub-text and the second sub-text are used as the second text.

[0067] As can be seen, through this embodiment, when the first text has a corresponding scene, the scene scene corresponding to the first text can be determined, the target object is in the scene scene, the first sub-text related to the above scene scene and the second sub-text related to the object form of the target object in the above scene scene are obtained in the e-book, and the first sub-text and the second sub-text are used as the second text, thereby deeply understanding the meaning represented by the first text from both the scene scene where the target object is located and the object form of the target object, and preparing for generating an image that matches the first text.

[0068] In some embodiments, in an ebook, obtaining second text related to the target object includes:

[0069] Determine the storyline corresponding to the first text; the target object is located within the storyline.

[0070] In the ebook, a third subtext related to the storyline and a fourth subtext related to the target object's behavior information in the storyline are obtained, and the third and fourth subtexts are used as the second text.

[0071] As described above, the first text is either the original ebook text or a comment on the ebook. When the first text is the original ebook text and that original ebook text is used to describe the storyline, or when the first text is a comment on the original ebook text describing the storyline, the second text can be determined in the following way.

[0072] First, determine the storyline corresponding to the first text. When the first text is used to describe a storyline, the storyline corresponding to the first text includes the storyline described in the first text. For example, if the first text is the original ebook, then the storyline corresponding to the first text includes the storyline described in the original ebook. The target object, as the object described in the first text, resides within the storyline corresponding to the first text.

[0073] When the first text is not used to describe the storyline, the storyline corresponding to the first text includes the scenes described in the original ebook text associated with the first text. For example, if the first text is a comment, "This description is wonderful," the paragraph associated with that comment in the ebook is identified, and the storyline described in that paragraph is considered the storyline corresponding to the first text. The target object, as the object described in the first text, resides within the storyline corresponding to the first text.

[0074] Next, within the ebook, a third sub-text related to the aforementioned storyline and a fourth sub-text related to the target object's behavioral information within the aforementioned storyline are retrieved, and these three sub-texts are used as the second text. The third sub-text represents the specific content of the aforementioned storyline. The fourth sub-text represents the target object's behavioral information within the aforementioned storyline. Behavioral information includes dialogue, actions, etc.

[0075] As can be seen, through this embodiment, when the first text has a corresponding storyline, the storyline corresponding to the first text can be determined, the target object is in the storyline, the third sub-text related to the storyline and the fourth sub-text related to the target object's behavior information in the storyline are obtained in the e-book, and the third sub-text and the fourth sub-text are used as the second text, thereby deeply understanding the meaning represented by the first text from both the storyline in which the target object is located and the behavior information of the target object, and preparing for generating an image that matches the first text.

[0076] It's worth noting that in practice, in one scenario, for any first text, the second text can be determined by identifying the corresponding scene. For a first text with a storyline, the second text can be determined by identifying the corresponding storyline. In another scenario, one can first analyze whether the first text has a corresponding storyline. A storyline needs to include at least two elements: the characters involved in the story and the story's progression. If the first text has a corresponding storyline, the second text can be determined by identifying the corresponding storyline. If the first text does not have a corresponding storyline, the second text can be determined by identifying the corresponding scene.

[0077] After determining the second text, in step S106 above, image generation prompts are generated based on the first and second texts. These image generation prompts are information input into the image generation model to generate the image; they can also be called image generation prompt words. The image generation model can be a painting model trained using AI technology. This model learns from a large amount of image data and artistic style information to automatically generate artworks or perform image style transfer. Furthermore, based on the image generation prompts, an image matching the first text is generated.

[0078] In some embodiments, generating an image-based prompt message based on the first text and the second text includes:

[0079] Using a Large Language Model (LLM), the feature information of image elements is determined based on the text content of the first text and the text content of the second text. Image elements include the object shape of the objects in the image, the image scene environment, the image composition method, the image style type, and the image color tone.

[0080] Using a large language model, image generation prompts are generated based on the feature information of image elements.

[0081] In this embodiment, firstly, using a large language model, based on the text content of the first text and the text content of the second text, the element information of the image elements is determined. Image elements include the object form of the object in the image, the image scene environment, the image composition method, the image style type, and the image color tone. The object in the image refers to the main object in the image to be generated, including the target object determined above. The object form of the object in the image includes the target object's actions, expression, appearance, clothing, etc. The image scene environment refers to the background environment where the target object is located, such as being in a fairyland or indoors. The image composition method includes, but is not limited to, the aspect ratio of the image and the position of the target object in the image. Image style types can be exemplified as ancient style, modern style, technological style, comic style, landscape style, architectural style, etc. The image color tone includes, but is not limited to, the main color tone of the image to be generated, such as a brownish-yellow color scheme.

[0082] In one example, the object shape of the target object can be determined based on the first text and the second sub-text, the image scene environment can be determined based on the first sub-text, and the image composition, image style type and image color tone that match the image scene environment can be determined.

[0083] Next, using a large language model, image generation prompts are generated based on the feature information of the image elements. For example, the large language model uses the feature information of each image element as the image generation prompt. The input to the large language model can be a first text and a second text, and the output can be the feature information of each image element, thus facilitating the image generation model to generate the image based on the feature information of each image element.

[0084] As can be seen, this embodiment can utilize the semantic understanding capabilities of a large language model to generate image generation prompts that are suitable for the image generation model to understand, based on the first and second texts, thereby improving image generation efficiency.

[0085] In some embodiments, feature information of image features is determined based on the text content of a first text and the text content of a second text using a large language model, including:

[0086] Using a large language model, the text content used to describe image elements in the text content of the first text and the text content of the second text is extracted to obtain the element information of at least one first image element in each image element.

[0087] Using a large language model, for the second image element whose element information has not been extracted from each image element, the element information of the second image element is obtained by expanding based on the element information of the first image element and the book type of the e-book.

[0088] In this embodiment, firstly, the text content describing image elements in the text content of the first text and the text content of the second text is extracted using a large language model to obtain element information of at least one first image element among all image elements. This step may extract element information of all image elements or only some image elements. If all image element information is extracted, the large language model directly outputs the element information of all image elements. If only some image element information is extracted, the large language model expands the element information of the second image elements for which element information has not been extracted, based on the element information of the first image elements and the book type of the e-book.

[0089] For example, if the large language model extracts a specific description of the image scene environment from the first sub-text and extracts the object form of the target object from the first and second sub-texts, then the large language model can also determine the image composition, image style type, and image color tone based on the type of e-book (such as ancient romance, modern romance, etc.), the object form of the target object, and the specific description of the image scene environment.

[0090] In a specific example, the first text is the original text of the e-book selected by the user: "Her big eyes, long, straight black hair, and ethereal temperament amazed everyone." The first subtext is: "Little A always walks with an ethereal air, always accompanied by four maids in white." The second subtext is: "Although Little A is young, the light gray whisk she carries and her blue dress make her appear experienced and capable." In this example, the specific description of the image scene environment, "ethereal air, four maids in white," can be extracted from the first subtext. The object form of the target object, "big eyes, long, straight black hair, ethereal temperament, young age, light gray whisk, blue dress, experienced and capable," can be extracted from the first and second subtexts. The large language model can then determine the image composition as a character-centric composition, the image style as a fantasy / xia (fantasy / martial arts) image, and the image color scheme as predominantly white and light colors, based on the e-book type (ancient romance), the object form of the target object, and the specific description of the image scene environment.

[0091] As can be seen, through this embodiment, the text content used to describe image elements can be extracted from the text content of the first text and the text content of the second text, and the element information of at least one first image element among each image element can be obtained. For the second image elements among each image element for which element information has not been extracted, the element information of the second image element can be expanded based on the element information of the first image element and the book type of the e-book. Thus, the image generation prompt information contains the element information of all image elements, without being constrained by the user's first text, thereby improving the matching degree between the generated image and the first text.

[0092] In step S106 above, an image matching the first text is generated based on the image generation prompt information. In some embodiments, the image generation prompt information can be input into an image generation model, which then generates the image matching the first text based on the prompt information. In some embodiments, generating an image matching the first text based on the image generation prompt information includes:

[0093] Obtain existing images of the e-book; the existing images include at least one of the following: e-book cover image, e-book illustrations, e-book character images, e-book comic images, and images from related videos of the e-book;

[0094] The target image is determined from existing images based on the image generation prompts using an image generation model; the content of the target image is related to the content of the image generated by the prompts.

[0095] Using an image generation model, an image matching the first text is generated based on the image generation prompt and the target image.

[0096] In this embodiment, existing images of the e-book can be acquired in advance. These existing images can be images generated for the e-book by its readers or users, including at least one of the following: e-book cover image, e-book illustrations, e-book character images, e-book comic images, and images from related videos of the e-book. Specifically, e-book character images refer to images generated for characters in the e-book, e-book comic images refer to comic works derived from the e-book, and images from related videos of the e-book can be video frames from short videos derived from the e-book. These existing images of the e-book can be stored in the image generation model in advance.

[0097] Then, the image generation prompt information is input into the image generation model. Based on the prompt information, the model determines the target image from the existing images. The content of the target image is related to the content suggested by the image generation prompt information. For example, if the prompt information suggests generating an entrance image for character A, then the target image can be either a character image of A or a video frame from a short video showing A's entrance.

[0098] Finally, using an image generation model, an image matching the first text is generated based on the image generation prompt and the target image.

[0099] As can be seen, through this embodiment, since the target image can be determined from the existing images in the e-book and used as an auxiliary parameter for the generated image, the generated image can not only match the first text, but also be similar in style to the existing images in the e-book, thereby improving the accuracy of image generation.

[0100] In some embodiments, an image matching the first text is generated based on an image generation model, according to an image generation prompt and a target image, including:

[0101] Using an image generation model, at least one of the following is determined as image reference information: object shape, scene environment, composition, style, and color tone of the target image.

[0102] Using an image generation model, an image matching the first text is generated based on image generation prompts and image reference information.

[0103] In this embodiment, firstly, the image generation model identifies at least one of the following in the target image: object shape, image scene environment, image composition method, image style type, and image color tone, as image reference information. The more relevant the image content of the target image is to the image content suggested by the image generation prompt, the more image reference information can be determined.

[0104] For example, the image generation prompt information is used to prompt the generation of an entrance image for character A. If the cover image of an e-book is obtained as the target image, but A is not in the cover image, the image composition, image style type, and image color tone of the target image can be obtained as image reference information. If the video frame of A's appearance in a short video derived from the e-book is obtained as the target image, the object shape of A, the image scene environment, the image composition, the image style type, and the image color tone of the target image can be obtained as image reference information.

[0105] Next, an image generation model is used to generate an image that matches the first text, based on image generation prompts and image reference information. When generating the image based on the image generation prompts, the image generation model refers to the style, color tone, and other characteristics of the target image, so that the generated image not only matches the first text but also matches the style of existing images in the e-book.

[0106] In one example, prompt words can be input into the image generation model. These prompt words inform the model to use the generated image information as the primary prompt word and the image reference information as the secondary prompt word. When the primary and secondary prompt words conflict, the image generation model evaluates the credibility of both the primary and secondary prompt words. Based on the credibility of these two prompt words, the more credible prompt word is selected. The credibility of the primary prompt word is related to whether it is directly extracted from the first and second texts or derived from book type extensions. The credibility of the secondary prompt words is determined based on user interaction with the target image, such as the number of likes a user gives to the target image.

[0107] As can be seen, through this embodiment, at least one of the following can be referenced in the target image: object shape, image scene environment, image composition method, image style type, and image color tone, combined with image generation prompts to generate an image, thereby improving the accuracy of image generation.

[0108] Figure 4 This is a schematic diagram of a scene generated by an embodiment of the present disclosure, such as... Figure 4 As shown, assuming the user triggers... Figure 2 or Figure 3 The "Smart Image Generator" control in the middle will then redirect to... Figure 4 The page shown displays a prompt message informing the user that the desired image is being generated by AI. Figure 5 This is a schematic diagram of the image generation result provided in an embodiment of the present disclosure, such as... Figure 5 As shown, once the AI-generated image is finished, the image generated by the AI ​​can be displayed. Here, we take the generation of an image of an airplane as an example for illustration. This final image is generated by AI.

[0109] like Figure 5 As shown, users can zoom in, download, and perform other operations on the generated image. Users can also trigger the "Post a comment with this image" control to copy the image to the comment section and post it as a comment. In one example, the generated image could be an emoji, in which case users can copy the emoji to the comment section and send it. Figure 5As shown, users can also trigger the "Regenerate" control, which, in response to the trigger, will re-execute the command. Figure 1 The process involves regenerating images for the user. To improve image generation efficiency, two images can be generated for the first text at a time, and the user can switch between images by swiping horizontally.

[0110] Furthermore, the above method may also include:

[0111] In response to a custom image command for the first text, the first text and image style type are displayed; the image style type is determined based on the first text and the ebook's book type.

[0112] In response to a modification instruction for a first text and / or image style type, an image matching the first text is generated based on the modified first text and / or modified image style type.

[0113] refer to Figure 4 and Figure 5 As shown, if the user triggers Figure 4 The "Custom Raw Image" control in the middle, or, triggering Figure 5 The "Edit Image Effects" control generates a custom image command for the first text based on the triggered operation. In response to this command, it displays the first text and the image style type, which is determined based on the first text and the ebook's book type. Examples of book types include ancient Chinese style and fantasy / cultivation genres.

[0114] Figure 6 This is a scene illustration of a custom raw image provided in an embodiment of the present disclosure, such as... Figure 6 As shown, if the user triggers Figure 4 The "Custom Raw Image" control in the middle, or, triggering Figure 5 The "Edit Image Effects" control in the middle will then jump to Figure 6 The page shown displays the first text (the original text of the ebook selected by the user or the comment information entered by the user), and the image style type determined based on the text content of the first text and the book type of the ebook. This image style type may be different from the image style type determined based on the first and second texts mentioned above. The image style type displayed here can be understood as the style type initially determined by the large language model.

[0115] Next, if a user's modification operation on the first text and / or image style type is received, a modification instruction for the first text and / or image style type is generated. Based on this modification instruction, the modified first text and / or modified image style type is obtained. Based on the modified first text and / or modified image style type, an image matching the first text is generated. If the user triggers... Figure 6 The "Generate Image" control in the middle can return to Figure 4 The scene shown is rendered as a raw image.

[0116] In one scenario, if a user modifies the initial text, it can be done through... Figure 1 The process involves identifying the target object described in the modified first text, retrieving updated second text related to that target object from the ebook, and generating updated image generation prompts based on the modified first text and the updated second text. The image style type in the updated image generation prompts is then used to guide the user... Figure 6 The image style type confirmed in the scene is then used to generate a prompt message based on the updated image, and an image matching the modified first text and the image style type confirmed by the user is generated.

[0117] In another scenario, if the user modifies the image style type, it can be done through... Figure 1 The process involves identifying the target object described in the first text, retrieving second text related to the target object from the ebook, and generating updated image generation prompts based on the first and second texts. The updated image generation prompts include an image style type for the user. Figure 6 The image style type obtained by modifying the scene shown is then used to generate prompt information based on the updated image, and an image matching the first text and the image style type obtained by the user is generated.

[0118] In another scenario, if the user modifies the initial text and image style type, it can be done through... Figure 1 The process involves identifying the target object described in the modified first text, retrieving updated second text related to that target object from the ebook, and generating updated image generation prompts based on the modified first text and the updated second text. The image style type in the updated image generation prompts is then used to guide the user... Figure 6 The image style type obtained by modifying the scene shown is then used to generate prompt information based on the updated image, and an image matching the modified first text and the image style type obtained by the user is generated.

[0119] certainly, Figure 6 The scene shown can also be used as an entry point for raw images. Figure 6 In the scenario shown, users can also delete the first text, re-enter the desired image prompts, and select the desired image style type to generate the desired image.

[0120] As can be seen, this embodiment also provides a strategy for user-defined raw images, thereby generating images that match the user's expectations.

[0121] In some embodiments, when a user triggers the paragraph comment entry of an e-book and enters the paragraph comment page, if the user triggers the "intelligent image generation" control on the paragraph comment page without entering any comment information, the paragraph in which the user wants to post a comment can be obtained, the paragraph can be used as the first text, and a matching image can be generated for the first text.

[0122] As described above, in some embodiments, after generating image generation prompts, a target image can be determined from existing images in the ebook using an image generation model. The content of the target image is related to the content prompted by the image generation prompts. The image generation model determines at least one of the following: the object shape of the object in the target image, the image scene environment, the image composition method, the image style type, and the image color tone, as image reference information. Based on the image generation prompts and the image reference information, the image generation model then generates an image that matches the first text. A specific accompanying drawing illustrates this process.

[0123] Figure 7 This is a schematic diagram illustrating the generation of an image matching a first text from an existing image in an e-book, as provided in one embodiment of this disclosure. Figure 7 As shown, the first text is the original text of the ebook selected by the user. Of course, the first text can also be a comment posted by the user about the ebook, or it can include both the original text of the ebook selected by the user and the comment posted by the user. Here, we use the original text of the ebook as the first text for illustration. The image generation model generates image generation prompts based on the original text of the ebook, "In a corner of a lush bamboo forest, there is a cute panda with big eyes." The image generation prompts include:

[0124] The object in the picture is described as: "A cute panda with big eyes";

[0125] Image scene environment: Bamboo forest;

[0126] Image composition: Composition with animals as the main subject;

[0127] Image style type: Cute;

[0128] Image color scheme: predominantly black and white.

[0129] like Figure 7As shown, the image generation model, based on the image generation prompts, identifies the target image as a panda character image among existing images in the ebook (including the ebook cover image, character images, corresponding comic images, and video frames from the corresponding short video). This target image shows the panda eating bamboo. Next, the image generation model generates an image of a panda eating bamboo in a bamboo forest, matching the target image with the first text. This image incorporates the object form of the target image (panda eating bamboo), improving the accuracy of image generation. Figure 7 In the image, panda-related illustrations were generated using AI.

[0130] In summary, through the above embodiments, it is possible to obtain first text related to the e-book, determine the target object described by the first text, obtain second text related to the target object in the e-book, and generate an image matching the first text based on the first and second texts, thereby achieving the effect of efficiently and accurately generating illustrations for the relevant content of the e-book.

[0131] Figure 8 This is a schematic diagram of the structure of an image generation apparatus provided in an embodiment of the present disclosure, as shown below. Figure 8 As shown, the device includes:

[0132] The request acquisition unit 81 is used to acquire an image generation request for a first text; the first text is text associated with an e-book;

[0133] The text acquisition unit 82 is configured to, in response to the image generation request, determine the target object described by the first text, and acquire second text related to the target object in the e-book;

[0134] The image generation unit 83 is configured to generate image generation prompt information based on the first text and the second text, and generate an image matching the first text based on the image generation prompt information.

[0135] Optionally, the request acquisition unit is specifically configured to: in response to a selection instruction for the main text portion of the e-book, determine the selection content corresponding to the selection instruction in the main text portion; if a first request to generate an image based on the selection content is received, determine the selection content as the first text, and determine the first request as the image generation request for the first text.

[0136] Optionally, the request acquisition unit is specifically configured to: in response to an input instruction for comment information on the e-book, acquire the comment information input for the e-book; if a second request to generate an image based on the comment information is received, determine the comment information as the first text, and determine the second request as the image generation request for the first text.

[0137] Optionally, the text acquisition unit is specifically used to: determine the scene scene corresponding to the first text; the target object is in the scene scene; in the e-book, acquire a first sub-text related to the scene scene and a second sub-text related to the object form of the target object in the scene scene, and use the first sub-text and the second sub-text as the second text.

[0138] Optionally, the text acquisition unit is specifically used to: determine the storyline corresponding to the first text; the target object is in the storyline; in the e-book, acquire a third sub-text related to the storyline and a fourth sub-text related to the target object's behavior information in the storyline, and use the third sub-text and the fourth sub-text as the second text.

[0139] Optionally, the image generation unit is specifically used to: determine the element information of image elements based on the text content of the first text and the text content of the second text using a large language model; the image elements include the object shape of the object in the image, the image scene environment, the image composition method, the image style type, and the image color tone; and generate image generation prompt information based on the element information of the image elements using the large language model.

[0140] Optionally, the image generation unit is further configured to: extract the text content used to describe the image elements from the text content of the first text and the text content of the second text using the large language model to obtain the element information of at least one first image element among the various image elements; and, using the large language model, expand the element information of the second image element based on the element information of the first image element and the book type of the e-book for the second image element among the various image elements where the element information has not been extracted.

[0141] Optionally, the image generation unit is specifically configured to: acquire existing images of the e-book; the existing images include at least one of the following: the cover image of the e-book, illustrations of the e-book, character images of the e-book, comic images of the e-book, and images from related videos of the e-book; determine a target image from the existing images using an image generation model based on the image generation prompt information; the image content of the target image is related to the image content prompted by the image generation prompt information; and generate an image matching the first text using the image generation model based on the image generation prompt information and the target image.

[0142] Optionally, the image generation unit is further configured to: determine, through the image generation model, at least one of the following: object shape of the object in the target image, image scene environment, image composition method, image style type, and image color tone, as image reference information; and generate an image matching the first text based on the image generation prompt information and the image reference information, through the image generation model.

[0143] Optionally, the above apparatus further includes a custom generation unit, configured to: display the first text and an image style type in response to a custom image generation instruction for the first text; the image style type is determined based on the first text and the book type of the e-book; and generate an image matching the first text in response to a modification instruction for the first text and / or the image style type, according to the modified first text and / or the modified image style type.

[0144] The image generation apparatus in this embodiment can implement the various processes of the above-described image generation method embodiments and achieve the same effects and functions, which will not be repeated here.

[0145] One embodiment of this disclosure also provides an electronic device. Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure, as shown below. Figure 9As shown, electronic devices can vary significantly due to differences in configuration or performance. They may include one or more processors 901 and memories 902, with the memory 902 storing one or more application programs or data. The memory 902 can be temporary or persistent storage. The application programs stored in the memory 902 may include one or more modules (not shown), each module including a series of computer-executable instructions within the electronic device. Furthermore, the processor 901 may be configured to communicate with the memory 902, executing the series of computer-executable instructions stored in the memory 902 on the electronic device. The electronic device may also include one or more power supplies 903, one or more wired or wireless network interfaces 904, one or more input or output interfaces 905, one or more keyboards 906, etc.

[0146] In one specific embodiment, the electronic device includes a processor; and a memory configured to store computer-executable instructions, which, when executed, cause the processor to perform the following process:

[0147] Obtain an image generation request for the first text; the first text is text associated with the e-book;

[0148] In response to the image generation request, the target object described by the first text is determined, and the second text related to the target object is obtained in the e-book;

[0149] Based on the first text and the second text, an image-generating prompt message is generated, and based on the image-generating prompt message, an image matching the first text is generated.

[0150] The electronic device in this embodiment can implement the various processes of the above-described image generation method embodiments and achieve the same effects and functions, which will not be repeated here.

[0151] Another embodiment of this disclosure also provides a computer-readable storage medium for storing computer-executable instructions that, when executed by a processor, implement the following process:

[0152] Obtain an image generation request for the first text; the first text is text associated with the e-book;

[0153] In response to the image generation request, the target object described by the first text is determined, and the second text related to the target object is obtained in the e-book;

[0154] Based on the first text and the second text, an image-generating prompt message is generated, and based on the image-generating prompt message, an image matching the first text is generated.

[0155] The computer-readable storage medium in this disclosure embodiment can implement the various processes of the above-described image generation method embodiment and achieve the same effects and functions, which will not be repeated here.

[0156] Another embodiment of this disclosure also provides a computer program product, the computer program product including a computer program, which, when executed by a processor, implements the following process:

[0157] Obtain an image generation request for the first text; the first text is text associated with the e-book;

[0158] In response to the image generation request, the target object described by the first text is determined, and the second text related to the target object is obtained in the e-book;

[0159] Based on the first text and the second text, an image-generating prompt message is generated, and based on the image-generating prompt message, an image matching the first text is generated.

[0160] The computer program product in this disclosure embodiment can implement the various processes of the above-described image generation method embodiment and achieve the same effects and functions, which will not be repeated here.

[0161] In various embodiments of this disclosure, the computer-readable storage medium includes read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk, etc.

[0162] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many methodological improvements today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that a methodological improvement cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program and "integrate" a digital system onto a PLD themselves, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must also be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also understand that by simply performing some logic programming on the method flow using one of these hardware description languages ​​and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.

[0163] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.

[0164] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.

[0165] For ease of description, the above apparatus is described by dividing it into various functional units. Of course, in implementing the embodiments of this disclosure, the functions of each unit can be implemented in one or more software and / or hardware.

[0166] Those skilled in the art will understand that one or more embodiments of this disclosure can be provided as a method, system, or computer program product. Therefore, one or more embodiments of this disclosure can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, one or more embodiments of this disclosure can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0167] This disclosure is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0168] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0169] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0170] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0171] One or more embodiments of this disclosure can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a particular task or implement a particular abstract data type. One or more embodiments of this disclosure can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can reside in local and remote computer storage media, including storage devices.

[0172] The various embodiments in this disclosure are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0173] The above description is merely an embodiment of this disclosure and is not intended to limit the scope of this disclosure. Various modifications and variations can be made to this disclosure by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the scope of the claims of this disclosure.

Claims

1. An image generation method, characterized in that, include: Obtain an image generation request for the first text; the first text is text associated with the e-book; The first text includes one or more of the following: the main text of the e-book, and comment information about the e-book; In response to the image generation request, the target object described by the first text is determined, and in the e-book, second text related to the target object is obtained; the second text is used to represent at least one or more of the object form of the target object and the behavioral information of the target object; Wherein, the second text includes the first and second sub-texts in the e-book, or includes the third and fourth sub-texts in the e-book, or includes the first, second, third, and fourth sub-texts in the e-book; the first sub-text is related to the scene corresponding to the first text; the second sub-text is related to the object form of the target object in the scene; the third sub-text is related to the storyline corresponding to the first text; and the fourth sub-text is related to the behavioral information of the target object in the storyline. Based on the first text and the second text, an image-generating prompt message is generated, and based on the image-generating prompt message, an image matching the first text is generated.

2. The method according to claim 1, characterized in that, The step of obtaining the image generation request for the first text includes: In response to a selection instruction for the main text portion of the e-book, the selection content corresponding to the selection instruction is determined in the main text portion; If a first request to generate an image based on the selected content is received, then the selected content is determined to be the first text, and the first request is determined to be the image generation request for the first text.

3. The method according to claim 1, characterized in that, The step of obtaining the image generation request for the first text includes: In response to a comment input command for the e-book, obtain the comment input for the e-book; If a second request to generate an image based on the comment information is received, the comment information is identified as the first text, and the second request is identified as the image generation request for the first text.

4. The method according to claim 1, characterized in that, The step of generating an image and prompting information based on the first text and the second text includes: Using a large language model, based on the text content of the first text and the text content of the second text, the element information of the image elements is determined; the image elements include the object shape of the object in the image, the image scene environment, the image composition method, the image style type, and the image color tone; Based on the feature information of the image elements, the large language model generates image generation prompts.

5. The method according to claim 4, characterized in that, The step of determining the feature information of image elements based on the text content of the first text and the text content of the second text using a large language model includes: Using the large language model, the text content used to describe the image elements in the text content of the first text and the text content of the second text is extracted to obtain the element information of at least one first image element among the various image elements; Using the large language model, for the second image element whose element information has not been extracted from each of the image elements, the element information of the second image element is obtained by expanding based on the element information of the first image element and the book type of the e-book.

6. The method according to claim 1, characterized in that, The step of generating prompt information based on the image and generating an image matching the first text includes: Obtain existing images of the e-book; the existing images include at least one of the following: the cover image of the e-book, the illustrations of the e-book, the character images of the e-book, the comic images of the e-book, and images from related videos of the e-book; Using an image generation model, a target image is determined from the existing images based on the image generation prompt information; the image content of the target image is related to the image content prompted by the image generation prompt information. Using the image generation model, an image matching the first text is generated based on the image generation prompt information and the target image.

7. The method according to claim 6, characterized in that, The step of generating an image matching the first text based on the image generation prompt information and the target image using the image generation model includes: The image generation model determines at least one of the following in the target image: object shape, image scene environment, image composition method, image style type, and image color tone, as image reference information. Using the image generation model, an image matching the first text is generated based on the image generation prompt information and the image reference information.

8. The method according to claim 1, characterized in that, The method further includes: In response to a custom image command for the first text, the first text and image style type are displayed; the image style type is determined based on the first text and the book type of the e-book. In response to a modification instruction for the first text and / or the image style type, an image matching the first text is generated based on the modified first text and / or the modified image style type.

9. An image generation apparatus, characterized in that, include: The request acquisition unit is used to acquire an image generation request for a first text; the first text is text associated with an e-book; The first text includes one or more of the following: the main text of the e-book, and comment information about the e-book; A text acquisition unit is configured to, in response to the image generation request, determine the target object described by the first text, and acquire second text related to the target object in the e-book; the second text is used to represent at least one or more of the object form of the target object and the behavioral information of the target object; Wherein, the second text includes the first and second sub-texts in the e-book, or includes the third and fourth sub-texts in the e-book, or includes the first, second, third, and fourth sub-texts in the e-book; the first sub-text is related to the scene corresponding to the first text; the second sub-text is related to the object form of the target object in the scene; the third sub-text is related to the storyline corresponding to the first text; and the fourth sub-text is related to the behavioral information of the target object in the storyline. An image generation unit is configured to generate image generation prompt information based on the first text and the second text, and generate an image matching the first text based on the image generation prompt information.

10. An electronic device, characterized in that, include: processor; as well as, A memory configured to store computer-executable instructions, which, when executed, cause the processor to perform the method described in any one of claims 1-8.

11. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store computer-executable instructions that, when executed by a processor, implement the method described in any one of claims 1-8.

12. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the method described in any one of claims 1-8.

Citation Information

Patent Citations

  • Text illustrating method and device

    CN116385597A

  • Image generation method and device, electronic equipment and storage medium

    CN116894881A