Image generation method and device and related product
By acquiring and analyzing text related to e-books, generating image generation prompt information and generating matching images, the problem of users having difficulty generating pictures for e-book content is solved, and efficient and accurate image generation is achieved.
Patent Information
- Application Number
- CN202510246246.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-03-03
AI Technical Summary
In the e-book reading scenario, it is difficult for users to generate pictures of the relevant content of the e-book efficiently and accurately.
By obtaining text related to the e-book, determining the target object described by the text, and obtaining the second text related to the target object in the e-book, generating image generation prompt information, and finally generating an image matching the first text based on this information.
It realizes efficient and accurate generation and pictures of the relevant content of e-books, improving the user's e-book reading experience and image generation efficiency.
Smart Images

Figure CN119991882A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to an image generation method, device and related products. Background Art
[0002] In the related technology, AI (Artificial Intelligence) technology can be used to assist users in generating the required images. For example, the user inputs the description information of the image to be generated into the raw image model trained based on AI technology, and the raw image model generates the image based on the description information. Considering that when users read e-books, there is a need to generate illustrations for the relevant content of the e-books. Therefore, how to efficiently and accurately generate illustrations for the relevant content of the e-books has become one of the problems that need to be solved. Summary of the invention
[0003] The embodiments of the present disclosure provide an image generation method, device and related products, which can efficiently and accurately generate illustrations for the relevant content of an e-book.
[0004] In a first aspect, an embodiment of the present disclosure provides an image generation method, comprising: Obtaining an image generation request for a first text; the first text is text associated with an electronic book; In response to the image generation request, determining a target object described by the first text, and acquiring a second text related to the target object in the electronic book; Image generation prompt information is generated according to the first text and the second text, and an image matching the first text is generated according to the image generation prompt information.
[0005] In a second aspect, an embodiment of the present disclosure provides an image generating device, including: A request acquisition unit, configured to acquire an image generation request for a first text; the first text is a text associated with an electronic book; a text acquisition unit, configured to determine a target object described by the first text in response to the image generation request, and acquire a second text related to the target object in the electronic book; An image generating unit is used to generate image generation prompt information according to the first text and the second text, and generate an image matching the first text according to the image generation prompt information.
[0006] In a third aspect, an embodiment of the present disclosure provides an electronic device, comprising: a processor; and a memory configured to store computer-executable instructions, wherein the computer-executable instructions, when executed, enable the processor to implement the method described in the first aspect above.
[0007] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, wherein the computer-readable storage medium is used to store computer-executable instructions, and the computer-executable instructions implement the method described in the first aspect when executed by a processor.
[0008] In a fifth aspect, an embodiment of the present disclosure provides a computer program product, wherein the computer program product includes a computer program, and when the computer program is executed by a processor, the method described in the first aspect is implemented.
[0009] In one or more embodiments of the present disclosure, first, an image generation request for a first text is obtained, where the first text is text associated with an e-book. Then, in response to the image generation request, a target object described by the first text is determined. In the e-book, a second text related to the target object is obtained. Finally, image generation prompt information is generated based on the first text and the second text. Based on the image generation prompt information, an image matching the first text is generated. It can be seen that through this embodiment, a first text related to an e-book can be obtained, and the target object described by the first text can be determined. In the e-book, a second text related to the target object is obtained. Based on the first text and the second text, an image matching the first text is generated, thereby achieving the effect of efficiently and accurately generating illustrations for the relevant content of the e-book. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the technical solutions in one or more embodiments of the present disclosure or related technologies, the drawings required for use in the description of the embodiments or related technologies are briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative labor. Figure 1 A schematic diagram of a flow chart of an image generation method provided by an embodiment of the present disclosure; Figure 2 A schematic diagram of a first text provided for an embodiment of the present disclosure; Figure 3 A schematic diagram of a first text provided for another embodiment of the present disclosure; Figure 4 A schematic diagram of a scene for image generation provided by an embodiment of the present disclosure; Figure 5 A schematic diagram of image generation results provided by an embodiment of the present disclosure; Figure 6 A schematic diagram of a scene of a custom raw image provided by an embodiment of the present disclosure; Figure 7A schematic diagram of generating an image matching a first text based on an existing image in an electronic book provided by an embodiment of the present disclosure; Figure 8 A schematic diagram of the structure of an image generating device provided by an embodiment of the present disclosure; Fig. 9 A schematic diagram of the structure of an electronic device provided in one embodiment of the present disclosure. DETAILED DESCRIPTION
[0011] In order to enable those skilled in the art to better understand the technical solutions in one or more embodiments of the present disclosure, the technical solutions in one or more embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in one or more embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. Based on one or more embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the protection scope of the present disclosure.
[0012] It is understandable that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, scope of use, usage scenarios, etc. of the information involved in the present disclosure should be informed to the relevant parties and their authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0013] For example, in response to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested to be performed will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, application, server, or storage medium that performs the operation of the technical solution of the present disclosure according to the prompt message.
[0014] As an optional but non-limiting implementation, in response to receiving an active request from the user, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0015] It is understandable that the above notification and the process of obtaining user authorization are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that meet the relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0016] The disclosed embodiments provide an image generation method, device and related products, which can efficiently and accurately generate illustrations for the relevant content of an e-book. The image generation method can be applied to a terminal device or a server, and implemented by the terminal device or the server. The terminal device includes but is not limited to a laptop, a tablet computer, a desktop computer, a set-top box, a mobile device (e.g., a mobile phone, a portable music player, a personal digital assistant, a dedicated messaging device, a portable gaming device), a smart phone, a smart speaker, a smart watch, a smart TV, a vehicle-mounted terminal and other types of user terminals. The server includes a single server form and a server cluster form.
[0017] Figure 1 A flowchart of an image generation method provided by an embodiment of the present disclosure is shown in FIG. Figure 1 As shown, the process includes: Step S102, obtaining an image generation request for a first text; the first text is a text associated with an electronic book; Step S104, in response to the image generation request, determining a target object described by the first text, and acquiring a second text related to the target object in the electronic book; Step S106, generating image generation prompt information according to the first text and the second text, and generating an image matching the first text according to the image generation prompt information.
[0018] In this embodiment, first, an image generation request for a first text is obtained, and the first text is text associated with an electronic book. Then, in response to the image generation request, the target object described by the first text is determined, and in the above electronic book, a second text related to the target object is obtained. Finally, image generation prompt information is generated based on the first text and the second text, and an image matching the first text is generated based on the image generation prompt information. It can be seen that through this embodiment, the first text related to the electronic book can be obtained, and the target object described by the first text can be determined. In the above electronic book, the second text related to the target object is obtained, and an image matching the first text is generated based on the first text and the second text, so as to achieve the effect of efficiently and accurately generating illustrations for the relevant content of the electronic book.
[0019] In the above step S102, the first text is a text associated with the electronic book, for example, the first text is the original text in the electronic book or the user's comment information on the electronic book. In this embodiment, if an image generation operation of the user for the first text is detected, then based on the image generation operation, it is determined that an image generation request of the user for the first text is received. The image generation request is used to request the generation of a matching image for the first text. The electronic book involved in each embodiment of the present disclosure can be any electronic book that is read through a reading application, and is not limited here.
[0020] In some embodiments, obtaining an image generation request for a first text includes: In response to a selection instruction for a body part of the electronic book, determining in the body part a selection content corresponding to the selection instruction; If a first request for generating an image based on selected content is received, the selected content is determined as a first text, and the first request is determined as a request for generating an image for the first text.
[0021] In this embodiment, the user can select the main text of the e-book. If the user's selection operation on the main text of the e-book is detected, a selection instruction for the main text of the e-book is generated based on the selection operation. In response to the selection instruction, the selection content corresponding to the selection instruction is determined in the main text of the e-book, that is, the user's selection content is determined in the main text of the e-book.
[0022] Next, if a first request for generating an image based on the selected content is received, the user's selected content is determined as the first text, and the first request is determined as a request for generating an image for the first text. For example, if a trigger operation of a user for generating an image control corresponding to the selected content is detected, based on the trigger operation, it is determined that a first request for generating an image based on the selected content is received, the user's selected content is determined as the first text, and the first request is determined as a request for generating an image for the first text.
[0023] Figure 2 A schematic diagram of a first text provided by an embodiment of the present disclosure, such as Figure 2 As shown, the user can make a selection in the main text of the e-book by long pressing. In response to the user's selection operation, the user's selected content is determined, and the "Smart Image Generation" control corresponding to the selected content is displayed. If the user triggers the control, it is determined that the user's first request to generate an image based on the selected content is received, the selected content is determined as the first text, and the first request is determined as an image generation request for the first text. Figure 2 In the figure, the original text of the e-book selected by the user is indicated by a shadow.
[0024] It can be seen that through this embodiment, the user can select the content of interest in the main text of the e-book and request to generate a matching image for the content of interest, thereby improving the user's e-book reading experience and image generation efficiency.
[0025] In some embodiments, obtaining an image generation request for a first text includes: In response to an instruction to input review information for an electronic book, obtaining review information input for the electronic book; If a second request to generate an image based on the comment information is received, the comment information is determined as the first text, and the second request is determined as a request to generate an image for the first text.
[0026] In this embodiment, the user can comment on the e-book, for example, comment on a chapter or paragraph in the e-book. If the user inputs comment information on the e-book, a comment information input instruction for the e-book is generated based on the input operation, and in response to the comment information input instruction, the comment information input by the user on the e-book is obtained and displayed. The comment information input by the user on the e-book can be comment information on a chapter of the e-book or comment information on a paragraph of the e-book.
[0027] Next, if a second request to generate an image based on the comment information is received, the comment information input by the user is determined as the first text, and the second request is determined as a request to generate an image for the first text. For example, if a trigger operation of a user to generate an image control corresponding to the input comment information is detected, based on the trigger operation, it is determined that a second request to generate an image based on the comment information is received, the comment information input by the user is determined as the first text, and the second request is determined as a request to generate an image for the first text.
[0028] Figure 3 A schematic diagram of a first text provided for another embodiment of the present disclosure, such as Figure 3 As shown, users can select a paragraph in the e-book ( Figure 3 The control unit (shown by the shadow in the middle) inputs comment information, displays the comment information input by the user in response to the comment information input operation, and displays the "Smart Image Generation" control corresponding to the comment information; if the user triggers the control, it is determined that a second request for generating an image based on the comment information is received, the comment information input by the user is determined as the first text, and the second request is determined as an image generation request for the first text. Figure 3 , the first text includes "This plane is really cool."
[0029] In some embodiments, reference Figure 3 When the user triggers the paragraph comment entry of the e-book and enters the paragraph comment page, if the user triggers the "Smart Image" control in the paragraph comment page without entering any comment information, the paragraph that the user wants to comment on can be obtained ( Figure 3 (shown by the shaded area in the middle), the paragraph is taken as the first text, and a matching image is generated for the first text.
[0030] In some embodiments, reference Figure 3When the user triggers the paragraph comment entry of the e-book and enters the paragraph comment page, if the user triggers the "Smart Image Generation" control in the paragraph comment page while entering the comment information, the paragraph ( Figure 3 The image is shown in the figure with shadow in the middle) and the comment information input by the user are taken as the first text, and a matching image is generated for the first text.
[0031] It can be seen that through this embodiment, the user can comment on the chapters or paragraphs in the e-book and request to generate a matching image for the input comment information, thereby improving the user's e-book reading experience and image generation efficiency.
[0032] In step S104, in response to the image generation request, the target object described by the first text is determined. The first text has a target object described. When the first text is an original e-book text, the object described by the original e-book text is determined as the target object. For example, if the first text is an original e-book text selected by the user, "She has big eyes, long black straight hair, and an unworldly temperament, which amazes everyone", it can be determined based on the content of the e-book that the original e-book text describes the character Xiao A in the e-book, and the target object can be determined to be the character Xiao A.
[0033] When the first text is comment information for an e-book, if the first text contains information indicating the target object being described, the target object is determined based on the information; if the first text does not contain information indicating the target object being described, the original e-book text associated with the first text in the e-book is determined, and the target object is determined based on the content of the first text and the associated original e-book text. For example, if the first text is a user's comment information for a paragraph in an e-book, "I really want to beat up Little B," the character Little B in the e-book is determined as the target object. For another example, if the first text is a user's comment information for a paragraph in an e-book, "This description is so amazing, the picture is full," the paragraph associated with the comment information can be determined in the e-book, and the object in the paragraph can be determined as the target object.
[0034] Next, a second text related to the target object is obtained in the electronic book. When the first text is the original text of the electronic book selected by the user, the second text may include text related to the original text of the electronic book, and when the first text is comment information input by the user, the second text may include the paragraph or chapter targeted by the comment information.
[0035] In some embodiments, in the electronic book, obtaining a second text related to the target object includes: Determining a picture scene corresponding to the first text; the target object is in the picture scene; In the electronic book, a first subtext related to the picture scene and a second subtext related to the object form of the target object in the picture scene are obtained, and the first subtext and the second subtext are used as the second text.
[0036] According to the above description, the first text is the original text of the e-book or the comment information on the e-book. No matter whether the first text is the original text of the e-book or the comment information on the e-book, the second text can be determined in the following manner.
[0037] First, determine the picture scene corresponding to the first text. When the first text is used to describe the picture scene, the picture scene corresponding to the first text includes the picture scene described by the first text. For example, if the first text is the comment information for the paragraph "I really want to beat up Xiao B", then the picture scene corresponding to the first text includes the picture scene of Xiao B being beaten up. The target object, as the object described by the first text, is located in the picture scene corresponding to the first text. In the above example, the target object is Xiao B.
[0038] When the first text is not used to describe a scene, the scene corresponding to the first text includes the scene described in the original text of the e-book associated with the first text. For example, the first text is the comment information "This paragraph of description is really wonderful", and the paragraph associated with the comment information is determined in the e-book, and the scene described in the paragraph is used as the scene corresponding to the first text. The scene can be exemplified as multiple characters drinking together. The target object, as the object described by the first text, is located in the scene corresponding to the first text. In the above example, the target object is each object in the scene described in the above paragraph.
[0039] Next, in the electronic book, a first subtext related to the above-mentioned picture scene and a second subtext related to the object form of the target object in the above-mentioned picture scene are obtained, and the first subtext and the second subtext are used as the second text. The first subtext can represent the specific content of the above-mentioned picture scene, or the specific content of the scene related to the above-mentioned picture scene. The second subtext can represent the object form of the target object in the above-mentioned picture scene, or the object form of the target object in the scene related to the above-mentioned picture scene. The object form includes the appearance, clothing, hairstyle, posture, etc. of the target object.
[0040] Take the first text as the comment information for the paragraph "I really want to beat up Xiao B" as an example. The scene corresponding to the first text includes the scene of Xiao B being beaten up. Then, in the e-book, the first subtext related to the scene of Xiao B being beaten up is obtained. The first subtext can be a text describing the dispute between Xiao B and other characters. In addition, the second subtext related to the object form of Xiao B in the above scene is obtained. The second subtext can be a text describing the appearance of Xiao B, or a text used to describe the appearance of Xiao B when he or she is arguing with other characters. The first subtext and the second subtext are used as the second text.
[0041] Taking the first text as the comment information on the paragraph "This paragraph is really wonderful" as an example, since the picture scene corresponding to the first text includes the picture scene described in the paragraph, the paragraph can be obtained as the first sub-text. If the paragraph describes the object form of each target object, the paragraph can also be obtained as the second sub-text. The first sub-text and the second sub-text are regarded as the second text.
[0042] Taking the first text as the original text of the e-book selected by the user, "Her big eyes, long black straight hair, and otherworldly temperament have amazed everyone" as an example, the scene corresponding to the first text is the scene where the character Xiao A appears. In the e-book, a first sub-text related to the appearance of the character Xiao A is obtained, and a second sub-text related to the appearance of the character Xiao A when he appears is obtained, and the first sub-text and the second sub-text are used as the second text.
[0043] It can be seen that through this embodiment, when the first text has a corresponding picture scene, the picture scene corresponding to the first text can be determined, and the target object is in the picture scene. In the e-book, a first sub-text related to the above-mentioned picture scene and a second sub-text related to the object form of the target object in the above-mentioned picture scene are obtained, and the first sub-text and the second sub-text are used as the second text, so as to deeply understand the meaning represented by the first text from two aspects of the picture scene in which the target object is located and the object form of the target object, and prepare for generating an image matching the first text.
[0044] In some embodiments, in the electronic book, obtaining a second text related to the target object includes: Determining a storyline corresponding to the first text; the target object is in the storyline; In the electronic book, a third subtext related to the storyline and a fourth subtext related to the behavior information of the target object in the storyline are obtained, and the third subtext and the fourth subtext are used as the second text.
[0045] According to the above description, the first text is the original text of the e-book or the comment information on the e-book. When the first text is the original text of the e-book and the original text of the e-book is used to express the plot, or when the first text is the comment information published on the original text of the e-book expressing the plot in the e-book, the second text can be determined in the following way.
[0046] First, determine the storyline corresponding to the first text. When the first text is used to describe the storyline, the storyline corresponding to the first text includes the storyline described by the first text. For example, if the first text is an original e-book that describes a story about multiple warriors seeking revenge on their enemies, then the storyline corresponding to the first text includes the storyline described in the original e-book. The target object, as the object described by the first text, is located in the storyline corresponding to the first text. In the above example, the target object includes multiple warriors and enemies.
[0047] When the first text is not used to describe the plot, the plot corresponding to the first text includes the scenes described in the original e-book text associated with the first text. For example, the first text is the comment information "This paragraph is really wonderful", and the paragraph associated with the comment information is determined in the e-book, and the plot described in the paragraph is used as the plot corresponding to the first text. The plot can be exemplified as a story of multiple warriors seeking revenge on their enemies. The target object, as the object described by the first text, is located in the plot corresponding to the first text. In the above example, the target objects include multiple warriors and enemies.
[0048] Next, in the electronic book, a third subtext related to the plot and a fourth subtext related to the behavior information of the target object in the plot are obtained, and the third subtext and the fourth subtext are used as the second text. The third subtext can represent the specific content of the plot. The fourth subtext can represent the behavior information of the target object in the plot. The behavior information includes dialogue, action, etc.
[0049] Taking the first text as the original e-book text, which is used to describe the story of multiple warriors seeking revenge on their enemies, as an example, the third sub-text may include the original e-book text, and may also include other text content in the e-book that describes the revenge story. The fourth sub-text may include the original e-book text, and may also include other text content in the e-book that describes the fighting process and dialogues of the target object in this revenge story.
[0050] It can be seen that through this embodiment, when the first text has a corresponding storyline, the storyline corresponding to the first text can be determined, and the target object is in the storyline. In the e-book, the third sub-text related to the storyline and the fourth sub-text related to the behavior information of the target object in the storyline are obtained, and the third sub-text and the fourth sub-text are used as the second text, so as to deeply understand the meaning represented by the first text from the two aspects of the storyline in which the target object is located and the behavior information of the target object, and prepare for generating an image matching the first text.
[0051] It is worth mentioning that in actual implementation, in one case, for any first text, the second text can be determined by determining the corresponding picture scene. For a first text with a storyline, the second text can be determined by determining the corresponding storyline. In another case, it is possible to first analyze whether the first text has a corresponding storyline. The storyline needs to include at least two factors: the characters involved in the story and the story process. If the first text has a corresponding storyline, the second text can be determined by determining the corresponding storyline. If the first text does not have a corresponding storyline, the second text can be determined by determining the corresponding picture scene.
[0052] After determining the second text, in the above step S106, image generation prompt information is generated according to the first text and the second text. Image generation prompt information is information used to be input into the image generation model to generate an image, and the image generation prompt information can also be called image generation prompt words. The image generation model can be a painting model trained based on AI technology, which can automatically generate artworks or perform image style transfer and other operations by learning a large amount of image data and artistic style information. Then, according to the image generation prompt information, an image matching the first text is generated.
[0053] In some embodiments, generating image generation prompt information according to the first text and the second text includes: Determine, by a Large Language Model (LLM), element information of image elements according to the text content of the first text and the text content of the second text; the image elements include the object form of the object in the image, the image scene environment, the image composition method, the image style type, and the image color tone; Through the large language model, image generation prompt information is generated based on the element information of the image elements.
[0054] In this embodiment, first, through the large language model, according to the text content of the first text and the text content of the second text, the element information of the image elements is determined, and the image elements include the object form of the object in the image, the image scene environment, the image composition method, the image style type, and the image color tone. The object in the image refers to the main object in the image to be generated, and the object in the image includes the target object determined above. The object form of the object in the image includes the action, expression, appearance, clothing, etc. of the target object. The image scene environment refers to the background environment in which the target object is located, such as being in the fairyland or indoors. The image composition method includes but is not limited to the aspect ratio of the image, the position of the target object in the image, etc. The image style type can be exemplified as ancient style, modern, technological, comic, landscape, architecture, etc. The image color tone includes but is not limited to the main color tone of the image to be generated, such as the main color tone is brown-yellow.
[0055] In one example, the object form of the target object can be determined based on the first text and the second subtext, the image scene environment can be determined based on the first subtext, and the image composition mode, image style type and image color tone matching the image scene environment can be determined.
[0056] Next, the image generation prompt information is generated according to the element information of the image elements through the large language model. For example, the element information of each image element is used as the image generation prompt information through the large language model. The input of the large language model can be the first text and the second text, and the output is the element information of each image element, so that the image generation model can generate an image according to the element information of each image element.
[0057] It can be seen that through this embodiment, the semantic understanding ability of the large language model can be utilized to generate image generation prompt information suitable for understanding by the image generation model based on the first text and the second text, thereby improving the image generation efficiency.
[0058] In some embodiments, determining the element information of the image element according to the text content of the first text and the text content of the second text by using the large language model includes: Extracting text content used to describe image elements from the text content of the first text and the text content of the second text by using a large language model to obtain element information of at least one first image element among the image elements; Through the large language model, for the second image element whose element information is not extracted from each image element, the element information of the second image element is expanded based on the element information of the first image element and the book type of the electronic book.
[0059] In this embodiment, first, the text content used to describe the image elements in the text content of the first text and the text content of the second text is extracted through the large language model to obtain the element information of at least one first image element in each image element. Through this step, the element information of all image elements may be extracted, or the element information of some image elements may be extracted. If the element information of all image elements is extracted, the large language model directly outputs the element information of all image elements. If the element information of some image elements is extracted, the large language model expands the element information of the second image element for the second image element whose element information is not extracted in each image element based on the element information of the first image element and the book type of the electronic book.
[0060] For example, the large language model extracts a specific description of the image scene environment from the first sub-text, and extracts the object form of the target object from the first text and the second sub-text. The large language model can also determine the image composition method, image style type and image color tone based on the type of e-book (such as ancient romance, modern romance, etc.), the object form of the target object, and the specific description of the image scene environment.
[0061] In a specific example, the first text is the original text of the e-book selected by the user, "She has big eyes, long straight black hair, and a refined temperament that amazes everyone", the first sub-text is "Little A always walks with a fairy-like air, and is always accompanied by four maids in white", and the second sub-text is "Although Little A is young, she looks sophisticated and capable with a light gray whisk and a green dress." In this example, the specific description of the image scene environment can be extracted from the first sub-text, "fairy-like air, four little maids in white", and the object form of the target object can be extracted from the first text and the second sub-text, "big eyes, long straight black hair, a refined temperament, young, a light gray whisk, a green dress, sophisticated and capable". The large language model can also determine the image composition method as a composition method with characters as the main body, determine the image style type as a fairy-like image, and determine the image color tone as mainly white and light colors according to the type of the e-book, ancient romance, the object form of the target object, and the specific description of the image scene environment.
[0062] In another specific example, the first text is the comment information on the paragraph "I really want to beat up Xiao B", the first subtext is "When Xiao B gets angry, his face turns red, and he throws things around", and the second subtext is "Every time Xiao B is beaten, tears flow from his left eye because he has no tears in his right eye". In this example, the specific description of the image scene environment "throwing things around" can be extracted from the first subtext, and the object form of the target object "being beaten, tears flow from the left eye, and red face" can be extracted from the first text, the first subtext, and the second subtext. Then, the large language model can also determine the image composition method as a composition method with characters as the main body, determine the image style type as an indoor image, and determine that the image color tone is mainly black, white and gray according to the type of the e-book (funny rebirth text), the object form of the target object, and the specific description of the image scene environment. The large language model can also be extended to obtain the object form of the target object including "being hit on the head with a hammer".
[0063] It can be seen that through this embodiment, the text content used to describe the image elements can be first extracted from the text content of the first text and the text content of the second text to obtain the element information of at least one first image element in each image element, and for the second image element whose element information is not extracted in each image element, based on the element information of the first image element and the book type of the e-book, the element information of the second image element is expanded to obtain, so that the image generation prompt information contains the element information of all image elements without being constrained by the user's first text, thereby improving the matching degree between the generated image and the first text.
[0064] In the above step S106, an image matching the first text is generated according to the image generation prompt information. In some embodiments, the image generation prompt information can be input into an image generation model, and the image generation model generates an image matching the first text according to the image generation prompt information. In some embodiments, generating an image matching the first text according to the image generation prompt information includes: Acquire an existing image of the e-book; the existing image includes at least one of the following images: an e-book cover image, an e-book illustration, an e-book character image, an e-book comic image, and an image in a video related to the e-book; By using the image generation model, according to the image generation prompt information, a target image is determined in the existing images; the image content of the target image is related to the image content prompted by the image generation prompt information; The image generation model is used to generate prompt information and a target image according to the image, and an image matching the first text is generated.
[0065] In this embodiment, existing images of the e-book may be acquired in advance. The existing images may be images generated for the e-book by the readers or users of the e-book, including at least one of the cover image of the e-book, illustrations of the e-book, character images of the e-book, comic images of the e-book, and images in the related videos of the e-book. Among them, the character images of the e-book refer to images generated for the characters in the e-book, the comic images of the e-book refer to comic works derived from the e-book, and the images in the related videos of the e-book may be video frames in short videos derived from the e-book. The existing images of the e-book may be stored in the image generation model in advance.
[0066] Then, the image generation prompt information is input into the image generation model, and the target image is determined from the existing images according to the image generation prompt information through the image generation model, and the image content of the target image is related to the image content prompted by the image generation prompt information. For example, if the image generation prompt information is used to prompt the generation of an appearance image for the character Xiao A, then the character image of Xiao A or the video frame when Xiao A appears in the short video can be obtained as the target image.
[0067] Finally, an image matching the first text is generated by generating prompt information and a target image according to the image through an image generation model.
[0068] It can be seen that through this embodiment, since the target image can be determined from the existing images in the e-book and used as an auxiliary parameter for generating the image, the generated image can not only match the first text but also be similar in style to the existing images in the e-book, thereby improving the accuracy of image generation.
[0069] In some embodiments, generating an image matching the first text by using an image generation model and generating prompt information and a target image according to the image includes: Determine at least one of the object form, image scene environment, image composition, image style type, and image color tone of an object in a target image as image reference information through an image generation model; An image matching the first text is generated by using an image generation model and based on the image generation prompt information and the image reference information.
[0070] In this embodiment, the image generation model is first used to identify at least one of the object form, image scene environment, image composition method, image style type, and image color tone of the object in the target image as image reference information. The more relevant the image content of the target image is to the image content suggested by the image generation prompt information, the more image reference information can be determined.
[0071] For example, image generation prompt information is used to prompt the generation of an appearance image for the character Xiao A. If the cover image of an e-book is obtained as the target image, and Xiao A is not in the cover image, the image composition method, image style type, and image color tone of the target image can be obtained as image reference information. If the video frame when Xiao A appears in a short video derived from the e-book is obtained as the target image, the object form, image scene environment, image composition method, image style type, and image color tone of Xiao A in the target image can be obtained as image reference information.
[0072] Next, the image generation model generates an image matching the first text according to the image generation prompt information and the image reference information. When generating an image according to the image generation prompt information, the image generation model refers to the style type, color tone, etc. of the target image, so that the generated image not only matches the first text, but also matches the style of the existing images in the electronic book.
[0073] In one example, a prompt word can be input into the image generation model, and the prompt word is used to inform the image generation model to use the image generation prompt information as the main raw image prompt word, and use the image reference information as the auxiliary raw image prompt word for image generation. When the main raw image prompt word conflicts with the auxiliary raw image prompt word, the image generation model is set to evaluate the credibility of the main raw image prompt word and the credibility of the auxiliary raw image prompt word, and based on the credibility of the main raw image prompt word and the credibility of the auxiliary raw image prompt word, a more credible prompt word is selected from the main raw image prompt word and the auxiliary raw image prompt word. Among them, the credibility of the main raw image prompt word is related to whether the main generation prompt word is directly extracted from the first text and the second text, or is expanded based on the book type. The credibility of the auxiliary raw image prompt word is determined based on the user's interactive behavior with respect to the target image, such as determined based on the number of likes given by the user to the target image.
[0074] It can be seen that through this embodiment, the image can be generated by referring to at least one of the object form, image scene environment, image composition method, image style type, and image color tone of the object in the target image, combined with the image generation prompt information, to improve the accuracy of image generation.
[0075] Figure 4 A schematic diagram of a scene for image generation provided by an embodiment of the present disclosure, such as Figure 4 As shown, assuming the user triggers Figure 2 or Figure 3 In the "Smart Image Generator" control, jump to Figure 4 The page shown in the figure displays a prompt message, which prompts the user that the image required is being generated by AI. Figure 5 A schematic diagram of the image generation result provided by an embodiment of the present disclosure is shown in FIG. Figure 5As shown, when the AI image generation is completed, the image generated by AI can be displayed. Here, the generation of an image of an airplane is used as an example for illustration, and the appearance image is generated by AI.
[0076] like Figure 5 As shown, the user can zoom in, download, and other operations on the generated image. The user can trigger the "Use this image to post a comment" control to copy the image to the comment area and post it as a comment. In one example, the generated image can be an emoticon package, and the user can copy the emoticon package to the comment area and send it. Figure 5 As shown, the user can also trigger the "regenerate" control, and in response to the trigger operation, re-execute Figure 1 In order to improve the efficiency of image generation, two images can be generated for the first text each time, and the user can switch between the images by sliding the screen horizontally.
[0077] Furthermore, the above method may also include: In response to a custom image generation instruction for a first text, the first text and an image style type are displayed; the image style type is determined based on the first text and the book type of the electronic book; In response to a modification instruction for the first text and / or the image style type, an image matching the first text is generated according to the modified first text and / or the modified image style type.
[0078] refer to Figure 4 and Figure 5 As shown, if the user triggers Figure 4 In the "Custom Image" control, or trigger Figure 5 , then according to the trigger operation, a custom image generation instruction for the first text is generated, and in response to the custom image generation instruction for the first text, the first text and the image style type are displayed, and the image style type is determined based on the first text and the book type of the e-book. Examples of book types include ancient style, cultivation of immortals, etc.
[0079] Figure 6 A schematic diagram of a scene of a custom raw image provided by an embodiment of the present disclosure, such as Figure 6 As shown, if the user triggers Figure 4 In the "Custom Image" control, or trigger Figure 5 , jump to the Edit Image Effects control in Figure 6The page shown in FIG. 1 shows a first text (the original text of the e-book selected by the user or the comment information entered by the user) and an image style type determined based on the text content of the first text and the book type of the e-book. The image style type may be different from the image style type determined based on the first text and the second text. The image style type displayed here can be understood as the style type preliminarily determined by the large language model.
[0080] Next, if a modification operation of the user on the first text and / or image style type is received, a modification instruction for the first text and / or image style type is generated, and according to the modification instruction, the modified first text and / or the modified image style type is obtained, and according to the modified first text and / or the modified image style type, an image matching the first text is generated. Figure 6 Generate Image control in , you can return Figure 4 The scene shown is drawn.
[0081] In one case, the user modifies the first text, then Figure 1 The process of determining a target object described by the modified first text, obtaining an updated second text related to the target object in the electronic book, generating updated image generation prompt information according to the modified first text and the updated second text, wherein the image style type in the updated image generation prompt information is provided to the user in Figure 6 The image style type confirmed in the scene shown, then, based on the updated image generation prompt information, an image matching the modified first text and the image style type confirmed by the user is generated.
[0082] In another case, if the user modifies the image style type, he can use Figure 1 The process of determining a target object described by a first text, obtaining a second text related to the target object in the e-book, generating updated image generation prompt information based on the first text and the second text, wherein the image style type in the updated image generation prompt information is provided to the user in Figure 6 The image style type modified in the scene shown is then generated according to the updated image prompt information to generate an image that matches the first text and the image style type modified by the user.
[0083] In another case, the user modifies the first text and image style type, and can modify the first text and image style type by Figure 1The process of determining a target object described by the modified first text, obtaining an updated second text related to the target object in the electronic book, generating updated image generation prompt information according to the modified first text and the updated second text, wherein the image style type in the updated image generation prompt information is provided to the user in Figure 6 The image style type modified in the scene shown is then generated according to the updated image prompt information to generate an image that matches the modified first text and the image style type modified by the user.
[0084] certainly, Figure 6 The scene shown can also be used as a raw image entry. Figure 6 In the scenario shown, the user can also delete the first text, re-enter the required image prompt words, and select the required image style type to generate a desired image.
[0085] It can be seen that, through this embodiment, a strategy for users to customize raw images is also provided, so as to generate images that match the user's expectations.
[0086] In some embodiments, when a user triggers the paragraph comment entry of an e-book and enters the paragraph comment page, if the user triggers the "Smart Image Generation" control in the paragraph comment page without entering any comment information, the paragraph on which the user wants to comment can be obtained, and the paragraph can be used as the first text, and a matching image can be generated for the first text.
[0087] According to the previous description, in some embodiments, after the image generation prompt information is generated, the target image can be determined from the existing images in the electronic book through the image generation model, and the image content of the target image is related to the image content prompted by the image generation prompt information. Through the image generation model, at least one of the object form, image scene environment, image composition method, image style type, and image color tone of the object in the image of the target image is determined as image reference information. Through the image generation model, an image matching the first text is generated according to the image generation prompt information and the image reference information. Here, a specific figure is used to illustrate the above process.
[0088] Figure 7 A schematic diagram of generating an image matching a first text based on an existing image in an electronic book provided by an embodiment of the present disclosure, such as Figure 7As shown, the first text is the original text of the e-book selected by the user in the e-book. Of course, the first text can also be the comment information published by the user on the e-book, or include the original text of the e-book selected by the user in the e-book and the comment information published by the user on the e-book. Here, the first text is taken as an example for illustration. The image generation model generates image generation prompt information based on the original text of the e-book "In a corner of the green bamboo forest, there is a cute panda with big eyes". The image generation prompt information includes: The object form of the object in the picture: "a cute panda with big eyes"; Image scene environment: bamboo forest; Image composition: animal-based composition; Image style type: cute; Image color tone: mainly black and white.
[0089] like Figure 7 As shown, the image generation model determines that the target image is the character image of the same panda in the existing images of the e-book (including the cover image of the e-book, the character image of the e-book, the comic image corresponding to the e-book, and the video frame of the short video corresponding to the e-book) according to the image content generated by the image generation prompt information. The panda is eating bamboo in the target image. Then, the image generation model generates a panda eating bamboo in a bamboo forest based on the target image as the image matching the first text. The image incorporates the object form of the target image (panda eating bamboo), which improves the accuracy of image generation. Figure 7 In the video, panda-related diagrams were generated using AI.
[0090] In summary, through the above embodiments, it is possible to obtain a first text related to an e-book and determine the target object described by the first text. In the above e-book, a second text related to the target object is obtained, and based on the first text and the second text, an image matching the first text is generated, thereby achieving the effect of efficiently and accurately generating illustrations for the relevant content of the e-book.
[0091] Figure 8 A schematic diagram of the structure of an image generating device provided by an embodiment of the present disclosure is shown in FIG. Figure 8 As shown, the device comprises: The request acquisition unit 81 is used to acquire an image generation request for a first text; the first text is a text associated with an electronic book; A text acquisition unit 82 is used to determine a target object described by the first text in response to the image generation request, and acquire a second text related to the target object in the electronic book; The image generating unit 83 is configured to generate image generating prompt information according to the first text and the second text, and generate an image matching the first text according to the image generating prompt information.
[0092] Optionally, the request acquisition unit is specifically used to: in response to a selection instruction for the main text portion of the electronic book, determine the selected content corresponding to the selection instruction in the main text portion; if a first request to generate an image based on the selected content is received, determine the selected content as the first text, and determine the first request as the image generation request for the first text.
[0093] Optionally, the request acquisition unit is specifically used to: in response to a comment information input instruction for the electronic book, acquire comment information input for the electronic book; if a second request to generate an image based on the comment information is received, determine the comment information as the first text, and determine the second request as the image generation request for the first text.
[0094] Optionally, the text acquisition unit is specifically used to: determine a picture scene corresponding to the first text; the target object is in the picture scene; in the e-book, obtain a first sub-text related to the picture scene and a second sub-text related to the object form of the target object in the picture scene, and use the first sub-text and the second sub-text as the second text.
[0095] Optionally, the text acquisition unit is specifically used to: determine a storyline corresponding to the first text; the target object is in the storyline; in the e-book, obtain a third sub-text related to the storyline and a fourth sub-text related to the behavior information of the target object in the storyline, and use the third sub-text and the fourth sub-text as the second text.
[0096] Optionally, the image generation unit is specifically used to: determine element information of image elements according to the text content of the first text and the text content of the second text through a large language model; the image elements include the object form of the object in the image, the image scene environment, the image composition method, the image style type, and the image color tone; and generate image generation prompt information according to the element information of the image elements through the large language model.
[0097] Optionally, the image generation unit is further specifically used to: extract, through the large language model, text content used to describe the image elements from the text content of the first text and the text content of the second text to obtain the element information of at least one first image element among the image elements; and, through the large language model, expand the element information of the second image element from which the element information is not extracted among the image elements based on the element information of the first image element and the book type of the electronic book.
[0098] Optionally, the image generation unit is specifically used to: obtain an existing image of the e-book; the existing image includes at least one of the following images: a cover image of the e-book, illustrations of the e-book, character images of the e-book, comic images of the e-book, and images in related videos of the e-book; through an image generation model, a target image is determined from the existing images according to the image generation prompt information; the image content of the target image is related to the image content prompted by the image generation prompt information; through the image generation model, an image matching the first text is generated according to the image generation prompt information and the target image.
[0099] Optionally, the image generation unit is also specifically used to: determine at least one of the object form, image scene environment, image composition method, image style type, and image color tone of the object in the target image through the image generation model as image reference information; and generate an image matching the first text according to the image generation prompt information and the image reference information through the image generation model.
[0100] Optionally, the above-mentioned device also includes a customized generation unit, which is used to: display the first text and image style type in response to a customized image generation instruction for the first text; the image style type is determined based on the first text and the book type of the e-book; in response to a modification instruction for the first text and / or the image style type, generate an image matching the first text according to the modified first text and / or the modified image style type.
[0101] The image generating device in the embodiment of the present disclosure can implement each process of the above-mentioned image generating method embodiment and achieve the same effects and functions, which will not be repeated here.
[0102] An embodiment of the present disclosure further provides an electronic device, Fig. 9 A schematic diagram of the structure of an electronic device provided by an embodiment of the present disclosure, such as Fig. 9As shown, the electronic device may have relatively large differences due to different configurations or performances, and may include one or more processors 901 and memory 902, and one or more applications or data may be stored in the memory 902. Among them, the memory 902 may be a temporary storage or a permanent storage. The application stored in the memory 902 may include one or more modules (not shown in the figure), and each module may include a series of computer executable instructions in the electronic device. Furthermore, the processor 901 may be configured to communicate with the memory 902 to execute a series of computer executable instructions in the memory 902 on the electronic device. The electronic device may also include one or more power supplies 903, one or more wired or wireless network interfaces 904, one or more input or output interfaces 905, one or more keyboards 906, etc.
[0103] In a specific embodiment, the electronic device includes a processor; and a memory configured to store computer executable instructions, wherein when the computer executable instructions are executed, the processor implements the following process: Obtaining an image generation request for a first text; the first text is text associated with an electronic book; In response to the image generation request, determining a target object described by the first text, and acquiring a second text related to the target object in the electronic book; Image generation prompt information is generated according to the first text and the second text, and an image matching the first text is generated according to the image generation prompt information.
[0104] The electronic device in the embodiment of the present disclosure can implement each process of the above-mentioned image generation method embodiment and achieve the same effects and functions, which will not be repeated here.
[0105] Another embodiment of the present disclosure further provides a computer-readable storage medium, wherein the computer-readable storage medium is used to store computer-executable instructions, and when the computer-executable instructions are executed by a processor, the following process is implemented: Obtaining an image generation request for a first text; the first text is text associated with an electronic book; In response to the image generation request, determining a target object described by the first text, and acquiring a second text related to the target object in the electronic book; Image generation prompt information is generated according to the first text and the second text, and an image matching the first text is generated according to the image generation prompt information.
[0106] The computer-readable storage medium in the embodiment of the present disclosure can implement each process of the above-mentioned image generation method embodiment and achieve the same effects and functions, which will not be repeated here.
[0107] Another embodiment of the present disclosure further provides a computer program product, the computer program product comprising a computer program, and when the computer program is executed by a processor, the following process is implemented: Obtaining an image generation request for a first text; the first text is text associated with an electronic book; In response to the image generation request, determining a target object described by the first text, and acquiring a second text related to the target object in the electronic book; Image generation prompt information is generated according to the first text and the second text, and an image matching the first text is generated according to the image generation prompt information.
[0108] The computer program product in the embodiment of the present disclosure can implement each process of the above-mentioned image generation method embodiment and achieve the same effects and functions, which will not be repeated here.
[0109] In various embodiments of the present disclosure, the computer-readable storage medium includes a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0110] In the 1990s, it was very clear whether the improvement of a technology was hardware improvement (for example, improvement of the circuit structure of diodes, transistors, switches, etc.) or software improvement (improvement of the method flow). However, with the development of technology, many improvements of the method flow today can be regarded as direct improvements of the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that the improvement of a method flow cannot be implemented with a hardware entity module. For example, a programmable logic device (PLD) (such as a field programmable gate array (FPGA)) is such an integrated circuit whose logical function is determined by the user's programming of the device. Designers can "integrate" a digital system on a PLD by programming themselves, without having to ask chip manufacturers to design and make dedicated integrated circuit chips. Moreover, nowadays, instead of manually making integrated circuit chips, this kind of programming is mostly implemented by "logic compiler" software, which is similar to the software compiler used when developing and writing programs, and the original code before compilation must also be written in a specific programming language, which is called hardware description language (HDL). There is not only one kind of HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also know that it is only necessary to program the method flow slightly in the above-mentioned hardware description languages and program it into the integrated circuit, and then it is easy to obtain the hardware circuit that implements the logic method flow.
[0111] The controller may be implemented in any suitable manner, for example, the controller may take the form of a microprocessor or processor and a computer-readable medium storing a computer-readable program code (e.g., software or firmware) executable by the (micro)processor, a logic gate, a switch, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller, examples of which include but are not limited to the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320, and the memory controller may also be implemented as part of the control logic of the memory. It is also known to those skilled in the art that, in addition to implementing the controller in a purely computer-readable program code manner, the controller may be implemented in the form of a logic gate, a switch, an application-specific integrated circuit, a programmable logic controller, and an embedded microcontroller by logically programming the method steps. Therefore, such a controller may be considered as a hardware component, and the devices for implementing various functions included therein may also be considered as structures within the hardware component. Or even, the devices for implementing various functions may be considered as both software modules for implementing the method and structures within the hardware component.
[0112] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0113] For the convenience of description, the above devices are described in terms of functions and are divided into various units. Of course, when implementing the embodiments of the present disclosure, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0114] It should be understood by those skilled in the art that one or more embodiments of the present disclosure may be provided as a method, system or computer program product. Therefore, one or more embodiments of the present disclosure may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, one or more embodiments of the present disclosure may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0115] The present disclosure is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present disclosure. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0116] These computer program instructions may also be stored in a computer readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0117] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0118] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0119] One or more embodiments of the present disclosure may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. One or more embodiments of the present disclosure may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communications network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0120] Each embodiment in the present disclosure is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0121] The above description is only an embodiment of the present disclosure and is not intended to limit the present disclosure. For those skilled in the art, the present disclosure may have various changes and variations. Any modification, equivalent substitution, improvement, etc. made within the spirit and principle of the present disclosure shall be included in the scope of the claims of the present disclosure.
Claims
1. An image generation method, characterized in that: include: Obtain an image generation request for a first text; The first text is a text associated with the electronic book; In response to the image generation request, determining a target object described by the first text, and acquiring a second text related to the target object in the electronic book; Image generation prompt information is generated according to the first text and the second text, and an image matching the first text is generated according to the image generation prompt information.
2. The method according to claim 1, characterized in that The obtaining of an image generation request for the first text includes: In response to a selection instruction for a body part of the electronic book, determining, in the body part, a selection content corresponding to the selection instruction; If a first request for generating an image based on the selected content is received, the selected content is determined as the first text, and the first request is determined as the image generation request for the first text.
3. The method according to claim 1, characterized in that The obtaining of an image generation request for the first text includes: In response to an instruction for inputting comment information for the electronic book, obtaining comment information input for the electronic book; If a second request for generating an image based on the comment information is received, the comment information is determined as the first text, and the second request is determined as the image generation request for the first text.
4. The method according to claim 1, characterized in that The step of acquiring a second text related to the target object in the electronic book includes: Determining a picture scene corresponding to the first text; wherein the target object is in the picture scene; In the electronic book, a first subtext related to the picture scene and a second subtext related to the object form of the target object in the picture scene are obtained, and the first subtext and the second subtext are used as the second text.
5. The method according to claim 1, characterized in that The step of acquiring a second text related to the target object in the electronic book includes: Determining a storyline corresponding to the first text; wherein the target object is in the storyline; In the electronic book, a third subtext related to the storyline and a fourth subtext related to the behavior information of the target object in the storyline are obtained, and the third subtext and the fourth subtext are used as the second text.
6. The method according to claim 1, characterized in that The step of generating image generation prompt information according to the first text and the second text includes: Determining, by means of a large language model, element information of an image element according to the text content of the first text and the text content of the second text; the image element includes an object form of an object in the image, an image scene environment, an image composition method, an image style type, and an image color tone; Image generation prompt information is generated according to the element information of the image elements through the large language model.
7. The method according to claim 6, characterized in that The determining, by using the large language model, element information of the image element according to the text content of the first text and the text content of the second text includes: Extracting text content used to describe the image elements from the text content of the first text and the text content of the second text by using the large language model to obtain the element information of at least one first image element among the image elements; By using the large language model, for the second image element whose element information is not extracted from each of the image elements, the element information of the second image element is expanded based on the element information of the first image element and the book type of the electronic book.
8. The method according to claim 1, characterized in that The step of generating prompt information according to the image to generate an image matching the first text includes: Acquire an existing image of the electronic book; the existing image includes at least one of the following images: a cover image of the electronic book, an illustration of the electronic book, a character image of the electronic book, a comic image of the electronic book, and an image in a video related to the electronic book; Determining a target image from the existing images according to the image generation prompt information through an image generation model; the image content of the target image is related to the image content prompted by the image generation prompt information; An image matching the first text is generated by using the image generation model and based on the image generation prompt information and the target image.
9. The method according to claim 8, characterized in that The step of generating an image matching the first text by using the image generation model according to the image generation prompt information and the target image includes: Determine at least one of the object form, image scene environment, image composition method, image style type, and image color tone of the object in the target image as image reference information through the image generation model; An image matching the first text is generated by the image generation model according to the image generation prompt information and the image reference information.
10. The method according to claim 1, characterized in that The method further comprises: In response to a custom image generation instruction for the first text, displaying the first text and an image style type; the image style type is determined based on the first text and the book type of the electronic book; In response to a modification instruction for the first text and / or the image style type, an image matching the first text is generated according to the modified first text and / or the modified image style type.
11. An image generating device, characterized in that: include: A request obtaining unit, configured to obtain an image generation request for a first text; The first text is a text associated with the electronic book; a text acquisition unit, configured to determine a target object described by the first text in response to the image generation request, and acquire a second text related to the target object in the electronic book; An image generating unit is used to generate image generation prompt information according to the first text and the second text, and to generate an image matching the first text according to the image generation prompt information.
12. An electronic device, characterized in that: include: processor; as well as, A memory configured to store computer executable instructions, which, when executed, cause the processor to implement the method of any one of claims 1 to 10.
13. A computer-readable storage medium, characterized in that: The computer-readable storage medium is used to store computer-executable instructions, and the computer-executable instructions implement the method according to any one of claims 1 to 10 when executed by a processor.
14. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 10 is implemented.
Citation Information
Patent Citations
Method and device for generating image based on text, electronic equipment and medium
CN115631251A
Text illustrating method and device
CN116385597A
Image generation method and device, electronic equipment and storage medium
CN116894881A
Material image processing method and device, storage medium and electronic equipment
CN116999859A
Text illustration generation method and device, equipment and storage medium
CN117830451A
Cited By
Video generation method and device based on electronic book and related product
CN120529144A
Image generation method and device and related product
CN120672909A
Image generation method, apparatus, and related product
CN120672909B
Multimedia content generation method and device, electronic equipment, medium and product
CN121505080A