Image generation method and device and related product

By using a generative model to generate matching emoticon packages in e-books, the problem that emoticon packages in existing technologies are difficult to meet user needs is solved, and the efficiency and accuracy of user comment information release are improved.

CN120672909AActive Publication Date: 2025-09-19DOUYIN VISION CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510786934.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-09-19
Estimated Expiration
2045-06-12

AI Technical Summary

Technical Problem

Existing emoticon packages are difficult to meet the needs of users when they post comments, which makes it more difficult to obtain matching emoticon packages and reduces the user's posting efficiency.

Method used

A generative model is used to generate a matching first image based on the first text and related text in the e-book, and the image text is obtained in combination with the image content. A second image is generated to be carried in the comment information, thereby realizing personalized generation of emoticons.

Benefits of technology

It reduces the difficulty for users to obtain matching emoticon packages and improves the efficiency of users posting comment information. The generated emoticon packages are closely related to the e-book content to meet user needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120672909A_ABST
    Figure CN120672909A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an image generation method and device and a related product, and the method comprises the steps: obtaining a first image matched with a first text in response to a comment request for the first text in an electronic book; obtaining an image copywriting matched with the first image according to the image content of the first image; generating a second image corresponding to the first text based on the first image and the image copywriting; issuing comment information corresponding to the first text; the comment information carries the second image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular to an image generation method, device, and related products. Background Art

[0002] In related technologies, while reading e-books, users can post comments on passages or chapters of interest. When posting comments, users often want to use emoticons. However, current emoticons are primarily built-in emoticon packages provided by input methods or reading applications. When users post comments, these built-in emoticon packages may not match the emoticon packages they need to reference, making it difficult for users to obtain matching emoticon packages and reducing the efficiency of posting comments. Summary of the Invention

[0003] The embodiments of the present disclosure provide an image generation method, device and related products, which can generate a matching second image for the first text in a scenario where a user posts comment information on the first text in an e-book. The second image can be carried in the comment information corresponding to the first text, which is equivalent to an emoticon package in the comment information corresponding to the first text, thereby achieving the effect of generating an emoticon package according to user needs, reducing the difficulty for users to obtain matching emoticon packages, and improving the efficiency of users posting comment information.

[0004] In a first aspect, an embodiment of the present disclosure provides an image generation method, comprising: In response to a comment request for a first text in an electronic book, obtaining a first image matching the first text; Acquiring an image text matching the first image according to the image content of the first image; generating a second image corresponding to the first text based on the first image and the image text; Posting comment information corresponding to the first text; the comment information carries the second image.

[0005] In a second aspect, an embodiment of the present disclosure provides an image generating device, including: An image acquisition unit, configured to acquire a first image matching the first text in response to a comment request for the first text in the electronic book; a text acquisition unit, configured to acquire an image text matching the first image according to the image content of the first image; an image generating unit, configured to generate a second image corresponding to the first text based on the first image and the image text; A comment publishing unit is used to publish comment information corresponding to the first text; the comment information carries the second image.

[0006] In a third aspect, an embodiment of the present disclosure provides an electronic device, comprising: a processor; and a memory configured to store computer-executable instructions, wherein the computer-executable instructions, when executed, enable the processor to implement the method described in the first aspect above.

[0007] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, which is used to store computer-executable instructions. When the computer-executable instructions are executed by a processor, they implement the method described in the first aspect above.

[0008] In a fifth aspect, an embodiment of the present disclosure provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the method described in the first aspect above.

[0009] In one or more embodiments of the present disclosure, first, in response to a comment request for a first text in an e-book, a first image matching the first text is obtained, then, based on the image content of the first image, an image text matching the first image is obtained, then, based on the first image and the image text, a second image corresponding to the first text is generated, and finally, a comment message corresponding to the first text is published, and the comment message carries the second image. It can be seen that through this embodiment, in a scenario where a user publishes a comment message for the first text in an e-book, a matching second image can be generated for the first text, and the second image can be carried in the comment message corresponding to the first text, which is equivalent to an emoticon package in the comment message corresponding to the first text, thereby achieving the effect of generating an emoticon package according to user needs, reducing the difficulty for users to obtain matching emoticon packages, and improving the efficiency of users publishing comment information. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the technical solutions in one or more embodiments of the present disclosure or related technologies, the following briefly introduces the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings described below are only some embodiments described in the present disclosure. Those skilled in the art can also derive other drawings based on these drawings without inventive effort. Figure 1 A flow chart of an image generation method provided in one embodiment of the present disclosure Figure 2a A schematic diagram of publishing comment information for a first text provided by an embodiment of the present disclosure; Figure 2b A schematic diagram of a comment page provided in accordance with an embodiment of the present disclosure; Figure 2c A schematic diagram of a candidate image provided by an embodiment of the present disclosure; Figure 2d A schematic diagram of a first image provided in an embodiment of the present disclosure; Figure 3a A schematic diagram of candidate texts provided in one embodiment of the present disclosure; Figure 3b A schematic diagram of text selection provided in an embodiment of the present disclosure; Figure 4a A schematic diagram of a second image provided in an embodiment of the present disclosure; Figure 4b A schematic diagram of comment publishing provided in one embodiment of the present disclosure; Figure 5 A schematic structural diagram of an image generating device provided in one embodiment of the present disclosure; Figure 6 A schematic structural diagram of an electronic device provided in one embodiment of the present disclosure. DETAILED DESCRIPTION

[0011] In order to enable those skilled in the art to better understand the technical solutions in one or more embodiments of the present disclosure, the technical solutions in one or more embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in one or more embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. Based on one or more embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present disclosure.

[0012] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the information involved in this disclosure should be informed to the relevant parties and their authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0013] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.

[0014] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0015] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0016] Taking into account the related art, when users post comment information on e-books, the built-in emoticon packages provided by input methods or reading applications are difficult to meet the users' emoticon package reference needs, which makes it more difficult for users to obtain matching emoticon packages and reduces the efficiency of users posting comment information. Based on this, the embodiments of the present disclosure provide an image generation method, device and related products, which can generate a matching second image for the first text in a scenario where a user posts comment information on the first text in an e-book. The second image can be carried in the comment information corresponding to the first text, which is equivalent to the emoticon package in the comment information corresponding to the first text, thereby achieving the effect of generating emoticon packages according to user needs, reducing the difficulty for users to obtain matching emoticon packages, and improving the efficiency of users posting comment information.

[0017] Among them, the image generation method can be applied to and executed by a terminal device, which includes but is not limited to various types of user terminals such as laptops, tablet computers, desktop computers, set-top boxes, mobile devices (for example, mobile phones, portable music players, personal digital assistants, dedicated messaging devices, portable gaming devices), smart phones, smart speakers, smart watches, smart TVs, and car-mounted terminals.

[0018] Figure 1 A flow chart of an image generation method provided by an embodiment of the present disclosure is shown as follows: Figure 1 As shown, the method includes: Step S102 , in response to a comment request for a first text in the e-book, obtaining a first image matching the first text; Step S104, obtaining an image text matching the first image according to the image content of the first image; Step S106: generating a second image corresponding to the first text based on the first image and the image text; Step S108: publishing comment information corresponding to the first text; the comment information carries the second image.

[0019] In the disclosed embodiment, first, in response to a comment request for the first text in the e-book, a first image matching the first text is obtained, then, based on the image content of the first image, an image text matching the first image is obtained, then, based on the first image and the image text, a second image corresponding to the first text is generated, and finally, a comment message corresponding to the first text is published, and the comment message carries the second image. It can be seen that through this embodiment, in the scenario where the user publishes a comment message for the first text in the e-book, a matching second image can be generated for the first text, and the second image can be carried in the comment message corresponding to the first text, which is equivalent to an emoticon package in the comment message corresponding to the first text, thereby achieving the effect of generating an emoticon package according to user needs, reducing the difficulty for users to obtain matching emoticon packages, and improving the efficiency of users publishing comment information.

[0020] In the above step S102, the user can post comment information on the first text in the e-book while reading the e-book. The first text can be a text of any length in the e-book, for example, the first text is one or more sentences in the e-book, or the first text is a paragraph in the e-book, or the first text is a chapter in the e-book.

[0021] The user can perform a comment operation on the first text by triggering the comment component associated with the first text. For example, the first text is a sentence or multiple sentences in an e-book. After the user selects the sentence or multiple sentences, the "post comment" control associated with the sentence or multiple sentences is triggered, thereby performing a comment operation on the sentence or multiple sentences. In this case, it can be considered that the comment information posted by the user is the comment information on the sentence or multiple sentences. For another example, the first text is a paragraph of text in an e-book. After the user selects all or part of the characters of the paragraph of text, the "post paragraph comment" control associated with the paragraph of text is triggered, thereby performing a comment operation on the paragraph of text. In this case, it can be considered that the comment information posted by the user is the comment information on the paragraph of text. For another example, the first text is a chapter in an e-book. The user can trigger the comment box corresponding to the chapter at the end of the chapter, thereby performing a comment operation on the chapter. In this case, it can be considered that the comment information posted by the user is the comment information on the chapter.

[0022] The terminal device generates a comment request for the first text according to a comment operation performed by a user on the first text, and in response to the comment request, obtains a first image matching the first text. The first image is an image matching the text content of the first text.

[0023] In one embodiment, obtaining a first image matching a first text includes: searching for a second text associated with the first text in the e-book; the first text and the second text are used to describe the same content in the e-book; generating, based on the first text and the second text, at least one candidate image matching the first text, and displaying the candidate image; In response to a selection instruction for a candidate image, the selected candidate image is determined as a first image.

[0024] In this embodiment, first, a second text associated with the first text is searched in the e-book. This association means that the first and second texts describe the same content in the e-book. For example, if the first text describes the appearance of a character in the e-book, the second text can also describe the appearance or posture of the character, and the same content here refers to the character. For another example, if the first text describes a scene in the e-book, such as a party scene, the second text can also describe the party scene, and the same content here refers to the party scene.

[0025] Then, based on the first and second texts, at least one candidate image matching the first text is generated, and each candidate image is displayed. The first and second texts can be used as image generation prompts, which are input into a generative model with image generation capabilities. The generative model then generates at least one candidate image matching the first and second texts based on the first and second texts. The terminal device can display each generated candidate image.

[0026] Finally, the user can select a candidate image from the candidate images. In response to the user's selection instruction for the candidate image, the candidate image selected by the user is determined as the first image. Since the first image is selected from the candidate images, and the candidate image is generated based on the first text and the second text, and the first text and the second text are used to describe the same content in the e-book, the image content of each candidate image matches the text content of the first text, and the image content of the first image also matches the text content of the first text. It can be understood that in this embodiment, the terminal device also displays the first image.

[0027] In this embodiment, after determining the first text, the terminal device can send the first text to the server corresponding to the reading application where the e-book is located, triggering the server to search for a second text that is associated with the first text in the e-book, and generate at least one candidate image matching the first text based on the first text and the second text, and return each candidate image to the terminal device, so that the terminal device displays each candidate image and determines the first image based on the user's selection instruction.

[0028] Figure 2aThis is a schematic diagram of publishing comment information for a first text provided by an embodiment of the present disclosure. Figure 2b A schematic diagram of a comment page provided in an embodiment of the present disclosure. Figure 2c A schematic diagram of a candidate image provided by an embodiment of the present disclosure, Figure 2d Schematic diagram of a first image provided by an embodiment of the present disclosure. Figure 2a As shown, taking the example of a user posting a paragraph comment for a paragraph of text in an e-book, after the user selects all the characters in the paragraph, the terminal device can display the operation bar associated with the paragraph, and the user can trigger the "Post a paragraph comment" control associated with the paragraph in the operation bar. In response to the triggering operation, the terminal device displays the following Figure 2b The user can trigger the "AI emoticon package" control provided in the comment page. In response to the trigger operation, the terminal device uses the text of the previous comment as the first text, and through the above process, generates multiple candidate images that match the first text, and Figure 2c Each candidate image can be displayed on the same page or switched by sliding the screen horizontally. Figure 2c The example of the horizontal screen sliding display of each candidate image is used as an example. Figure 2d As shown, the user can select the first image from various candidate images. Obviously, the terminal device displays the first image. Figure 2d As shown, the user can click the "Next" component to continue the subsequent operation.

[0029] It can be seen that through this embodiment, a second text that is associated with a first text can be searched in an e-book, and the first text and the second text are used to describe the same content in the e-book. Based on the first text and the second text, at least one candidate image matching the first text is generated, and each candidate image is displayed. In response to a selection instruction for a candidate image, the selected candidate image is determined as the first image, thereby accurately generating a first image that matches the text content of the first text, thereby improving the efficiency and accuracy of generating the first image.

[0030] In one embodiment, generating at least one candidate image matching the first text based on the first text and the second text includes: Determining, by a generative model, image content that the candidate image needs to represent based on the first text and the second text; Through the generative model, candidate images are generated according to the image content that the candidate images need to represent.

[0031] In this embodiment, the first and second texts can be used as image generation prompts. These prompts are then input into a generative model capable of generating images. The generative model then determines, based on the first and second texts, the image content that a candidate image is intended to represent. The image content that the candidate image is intended to represent is the content described by the first and second texts, such as the characters or scenes described by the first and second texts. The generative model then generates at least one candidate image based on the image content that the candidate image is intended to represent. The candidate image is used to represent the aforementioned image content.

[0032] For example, the first text is used to describe a character in an e-book, and the second text is also used to describe the character. Then, through the generative model, based on the first text and the second text, it is determined that the image content that the candidate image needs to represent is the character, and through the generative model, multiple candidate images are generated according to the image content that the candidate image needs to represent, and each candidate image is a character image of the character.

[0033] In this embodiment, after determining the first text, the terminal device can send the first text to the server corresponding to the reading application where the e-book is located. This triggers the server to search the e-book for a second text associated with the first text. The server then uses a generative model to determine the image content that the candidate images are intended to represent based on the first and second texts, and generates each candidate image based on the image content that the candidate images are intended to represent. The server also returns each candidate image to the terminal device, which then displays the candidate images and determines the first image based on the user's selection instruction.

[0034] It can be seen that through this embodiment, the generative model can be used to determine the image content that the candidate image needs to represent based on the first text and the second text, and the candidate image can be generated based on the image content that the candidate image needs to represent. The precise image generation capability of the generative model can be utilized to improve the efficiency and accuracy of generating candidate images.

[0035] In one embodiment, generating at least one candidate image matching the first text based on the first text and the second text includes: If there is a to-be-published comment text corresponding to the comment request, then using a generative model to determine the image content that the candidate image needs to represent based on the first text, the second text, and the to-be-published comment text; Through the generative model, the emotional atmosphere that the candidate image needs to represent is determined based on the comment text to be published; Through the generative model, candidate images are generated according to the image content that the candidate images need to represent and the emotional atmosphere that the candidate images need to represent.

[0036] In this embodiment, if there is a comment text to be published corresponding to the above-mentioned comment request, for example, the user enters a comment text in the comment box before triggering the "AI Emoji Package" control, then the comment text is the comment text to be published, and the generative model is used to determine the image content required to be represented by the candidate image based on the first text, the second text and the comment text to be published. The generative model can be used to take the contents represented by the first text, the second text and the comment text to be published as the image content required to be represented by the candidate image. For example, the first text and the second text are used to represent a character in an e-book, and the comment text to be published is used to express the desire to hug the character, then the character and the hugging action for the character are both used as the image content required to be represented by the candidate image.

[0037] Next, a generative model is used to determine the emotional atmosphere that the candidate image needs to represent based on the comment text to be published. The generative model can be used to determine the emotional atmosphere represented by the comment text to be published, and the emotional atmosphere represented by the comment text to be published can be determined as the emotional atmosphere that the candidate image needs to represent. For example, if the comment text to be published is "I really like this scene," the emotional atmosphere represented by the comment text to be published is determined to be happy and joyful, and the emotional atmosphere that the candidate image needs to represent is determined to be happy and joyful. For another example, if the comment text to be published is "I also want to be as motivated as him," the emotional atmosphere represented by the comment text to be published is determined to be inspirational and striving, and the emotional atmosphere that the candidate image needs to represent is determined to be inspirational and striving.

[0038] Finally, the generative model generates at least one candidate image based on the desired image content and the desired emotional atmosphere of the candidate image. The generative model can determine the image style of the candidate image based on the desired emotional atmosphere of the candidate image, thereby generating each candidate image based on the desired image content and the image style of the candidate image. Examples of candidate image styles include comics, cartoons, and ink paintings.

[0039] In this embodiment, after determining the first text and the comment text to be published, the terminal device can send the first text and the comment text to be published to the server corresponding to the reading application where the e-book is located, triggering the server to search the e-book for a second text associated with the first text. The server then uses a generative model to determine the image content that the candidate image needs to represent based on the first text, the second text, and the comment text to be published, and the emotional atmosphere that the candidate image needs to represent based on the comment text to be published. Based on the image content and the emotional atmosphere that the candidate image needs to represent, each candidate image is generated. The server also returns each candidate image to the terminal device, so that the terminal device displays each candidate image and determines the first image based on the user's selection instruction.

[0040] It can be seen that through this embodiment, the situation where there is a comment text to be published corresponding to the above-mentioned comment request is also taken into consideration. The generative model can be used to determine the image content that the candidate image needs to represent based on the first text, the second text and the comment text to be published, and the emotional atmosphere that the candidate image needs to represent is determined based on the comment text to be published. Based on the image content that the candidate image needs to represent and the emotional atmosphere that the candidate image needs to represent, the candidate image is generated, thereby combining the comment text to be published by the user to accurately generate the candidate image and improve the accuracy of the candidate image.

[0041] After generating the candidate images and determining the first image, in step S104, the terminal device retrieves an image text that matches the first image based on its content. The image text matches the content of the first image. The resulting second image, obtained by superimposing the first image and the image text, is the emoticon package that meets the user's emoticon package reference requirements.

[0042] In one embodiment, before obtaining an image text matching the first image based on the image content of the first image, the method further includes: determining a character in the electronic book represented by the first image, and determining the character as image content of the first image; or, A scene in the electronic book represented by the first image is determined, and the scene is determined as image content of the first image.

[0043] In this embodiment, before obtaining the image text, the terminal device also determines whether the first image is used to represent a character or a scene in the e-book. If it is determined that the first image is used to represent a character in the e-book, the character represented by the first image is determined to be, for example, the character Xiao A, and the character Xiao A is determined to be the image content of the first image. If it is determined that the first image is used to represent a scene in the e-book, the scene represented by the first image is determined to be, for example, a party scene, and the party scene is determined to be the image content of the first image.

[0044] In this embodiment, the terminal device can determine the image content of the first image through the above process, or the terminal device can send the image identifier of the first image to the server corresponding to the reading application where the e-book is located, triggering the server to determine the image content of the first image through the above process.

[0045] It can be seen that through this embodiment, the character or scene in the e-book represented by the first image can be determined as the image content of the first image. Therefore, when the image copy matching the first image is determined based on the image content of the first image, the image copy can be matched with the character or scene in the e-book, thereby improving the degree of association between the image copy and the e-book, so that the final generated second image, that is, the emoticon package, is closely associated with the e-book.

[0046] After determining the image content of the first image, in one embodiment, obtaining an image text matching the first image based on the image content of the first image includes: determining at least one text in the electronic book that is related to the image content of the first image, using the determined text as a candidate text that matches the first image, and displaying the candidate text; Based on the candidate texts, determine the image text.

[0047] In this embodiment, the terminal device can determine at least one text in the electronic book that is related to the image content of the first image, use the determined text as a candidate copy matching the first image, display each candidate copy, and determine the image copy based on each candidate copy.

[0048] Of course, the terminal device can also trigger the above-mentioned server to determine at least one text in the e-book related to the image content of the first image, and use the determined text as a candidate copy matching the first image. The server sends each candidate copy to the terminal device, so that the terminal device displays each candidate copy and determines the image copy based on each candidate copy.

[0049] The at least one text in the e-book related to the image content of the first image may be text in the e-book used to introduce the image content of the first image, or classic lines of the character represented by the image content of the first image, or classic quotes from the scene represented by the image content of the first image.

[0050] It can be seen that through this embodiment, at least one text related to the image content of the first image in the e-book can be determined, and the determined text can be used as a candidate copy matching the first image, and each candidate copy is displayed, and based on the candidate copy, the image copy is determined, so that the image copy comes from the e-book and is related to the image content of the first image, thereby improving the degree of association between the image copy and the e-book, improving the degree of association between the image copy and the first image, and improving the accuracy of the image copy.

[0051] As can be seen from the above description, the image content of the first image includes a character or scene in an e-book. Based on this, in one embodiment, determining at least one text in the e-book related to the image content of the first image includes: If the image content of the first image includes a character in the electronic book, determining lines related to the character in the electronic book as text related to the image content of the first image; If the image content of the first image includes a scene in an electronic book, a sentence in the electronic book related to the scene is determined as text related to the image content of the first image.

[0052] In this embodiment, if the first image's image content includes a character from an e-book, then lines related to that character in the e-book are determined as text related to the first image's image content. The lines related to that character can be popular lines from that character. That is, if the first image is a character image, the candidate text can be popular lines from that character in the e-book, thereby combining the character and its popular lines to generate an emoticon package.

[0053] In another case, if the image content of the first image includes a scene from an e-book, sentences related to the scene in the e-book are determined as text related to the image content of the first image. The sentences related to the scene can be classic text from the e-book describing the scene. In other words, when the first image is a scene image, the candidate text can be classic text from the e-book describing the scene, thereby combining the scene and the classic text to generate an emoticon package.

[0054] It can be seen that through this embodiment, the character's lines or sentences describing the scene can be selected from the e-book as candidate text based on the character or scene represented by the first image, so that the image text is closely related to the image content of the first image, so that when the user posts a comment on the first text, a matching emoticon package is generated, thereby improving the accuracy and efficiency of emoticon package generation.

[0055] In this embodiment, in addition to generating candidate copywriting in the above manner, in one case, a generative model can be used to generate candidate copywriting based on the to-be-published comment text corresponding to the comment request. For example, the to-be-published comment text is input into the generative model, and the generative model generates candidate copywriting based on the to-be-published comment text.

[0056] In another case, a generative model can be used to generate candidate texts based on the first text. For example, the first text is input into the generative model, and candidate texts are generated based on the first text. Alternatively, the character or scene represented by the first text is determined, and text describing the character or scene in the e-book is obtained, which may include the first text. The obtained text is input into the generative model, and candidate texts are generated based on the input text.

[0057] In another embodiment, a generative model can be used to generate candidate text based on text in an e-book that describes a character or scene in the first image. For example, a paragraph in the e-book that describes a character or scene in the first image is obtained and input into the generative model. The generative model then generates candidate text based on the input paragraph.

[0058] In each of the above situations, you can input appropriate prompts into the generative model, such as "Please generate a meme from the input text," to facilitate the generative model's generation of candidate texts. The candidate texts are the memes in the emoji package.

[0059] In one embodiment, determining the image text based on candidate texts includes: In response to a selection instruction for a candidate text, determining the selected candidate text as an image text; or, In response to an editing instruction for the candidate text, the edited candidate text is determined as the image text.

[0060] In this embodiment, after displaying each candidate copy, the terminal device can, in response to a user's selection instruction for a candidate copy, determine the selected candidate copy as an image copy, or, in response to a user's editing instruction for a candidate copy, determine the edited candidate copy as an image copy. In this embodiment, the candidate copy can be edited with respect to its content, font, size, color, style, etc.

[0061] It can be seen that through this embodiment, the user can obtain the image copy by selecting a candidate copy or editing the candidate copy, thereby meeting the user's demand for obtaining the image copy.

[0062] After obtaining the image text, in step S106, a second image corresponding to the first text is generated based on the first image and the image text, and the second image is displayed. The terminal device can overlay the image text with the first image to obtain the second image, which is the emoticon package. The terminal device can determine the overlay position of the image text within the first image based on user settings. The terminal device can also display the second image.

[0063] In the above step S108, the terminal device publishes the comment information corresponding to the first text, and the comment information carries the second image. The comment information may also carry or not carry the comment text to be published.

[0064] Figure 3a A schematic diagram of candidate texts provided in one embodiment of the present disclosure. Figure 3b A schematic diagram of text selection provided in an embodiment of the present disclosure. Figure 4a A schematic diagram of a second image provided in an embodiment of the present disclosure, Figure 4b A schematic diagram of comment publishing provided in an embodiment of the present disclosure is provided. Figure 3a Can be Figure 2d Then present. Figure 2d and Figure 3aAs shown, after the user determines the first image among the candidate images, the terminal device can obtain the candidate texts that match the first image through the above process and display the candidate texts. Figure 3b As shown, the user can select a candidate copy from various candidate copies and click "Next" to continue.

[0065] like Figure 4a As shown, after the user selects a candidate copy, the terminal device displays the first image and the selected copy. Figure 4a In the page shown, the user can edit or not edit the selected text. The user can also manually drag the selected text in the first image to get the second image. The second image can display the corresponding "Comment with Picture" component. If the user triggers the component, the user can jump to Figure 4b , users can Figure 4b Enter or not enter comment information, or modify the comment information previously entered, and click the "Publish" component to publish the comment information for the first text.

[0066] Figures 2a to 4b , the component is indicated as triggered by filling it with gray.

[0067] In one embodiment, after selecting a first image, the user can also directly enter an image text in the first image and set the position of the image text to obtain a second image, and then publish comment information about the first text based on the second image. In other words, the terminal device obtains the image text entered by the user in the first image, generates a second image based on the first image and the image text, and publishes comment information about the first text based on the second image.

[0068] Based on the solution of the above embodiment, in one embodiment, the user's record of AI-generated images for each text in the e-book can be counted, and the text with the highest number of generated images in the e-book can be determined based on the record. The determined text will be highlighted in the e-book, prompting the user to comment based on the text. The determined text is the first text mentioned above. When the user requests to post a comment on the first text, a second image can be generated as an AI emoticon package through the above process, and the AI ​​emoticon package can be carried in the comment information posted by the user.

[0069] In one embodiment, in the above step S102, in response to a comment request for a first text in an e-book, obtaining a first image that matches the first text can be as follows: searching the e-book for a second text that has an association with the first text, the first text and the second text are used to describe the same content in the e-book, and based on the first text and the second text, generating at least one candidate image that matches the first text, and using each candidate image as the first image.

[0070] In the above step S104, obtaining an image text matching the first image according to the image content of the first image may include at least one of the following methods: For each first image, if the image content of the first image includes a character in the e-book, the lines related to the character in the e-book are used as the image text matching the first image; For each first image, if the image content of the first image includes a scene in an e-book, a sentence in the e-book that describes the scene is used as the image text matching the first image; For each first image, input the comment text to be published into the generative model, and generate an image copy matching each first image based on the comment text to be published by the generative model; For each first image, input the first text into the generative model, and generate an image copy matching each first image based on the first text by the generative model; For each first image, determining a character or scene represented by the first text, obtaining text from the e-book that describes the character or scene, which may include the first text, inputting the obtained text into a generative model, and generating an image copy matching each first image based on the input text using the generative model; For each first image, a paragraph in the e-book describing the character or scene in the first image is obtained, the paragraph is input into the generative model, and the generative model obtains the image text matching each first image based on the input paragraph.

[0071] In step S106, based on the first image and the image text, a second image corresponding to the first text is generated, which can be achieved by combining each first image with a matching image text to obtain multiple second images. The terminal device also displays each second image as a candidate emoticon package.

[0072] Furthermore, in one case, the terminal device can publish the target emoticon package of the candidate emoticon package selected by the user in the comment information according to the user's selection operation.

[0073] In another scenario, the terminal device may also determine the candidate emoticon package selected by the user based on the user's selection operation and enter the emoticon package editing panel. The user can edit or not edit the candidate emoticon package in the emoticon package editing panel. Editing the candidate emoticon package includes editing either the first image or the image text in the candidate emoticon package. Based on the user's confirmation operation, the terminal device determines the edited candidate emoticon package or the unedited candidate emoticon package as the target emoticon package. The terminal device also publishes the target emoticon package in the comment information.

[0074] The “AI” in the above embodiments of the present disclosure refers to artificial intelligence (AI).

[0075] In summary, through the above embodiments, in a scenario where a user posts comment information on the first text in an e-book, a matching second image can be generated for the first text. The second image can be carried in the comment information corresponding to the first text, which is equivalent to the emoticon package in the comment information corresponding to the first text, thereby achieving the effect of generating emoticon packages according to user needs, reducing the difficulty for users to obtain matching emoticon packages, and improving the efficiency of users posting comment information.

[0076] Figure 5 This is a structural diagram of an image generating device provided by an embodiment of the present disclosure, as shown in FIG. Figure 5 As shown, the device includes: An image acquisition unit 51 is configured to acquire a first image matching the first text in response to a comment request for the first text in the electronic book; A text acquisition unit 52 is configured to acquire an image text matching the first image according to the image content of the first image; An image generating unit 53 is configured to generate a second image corresponding to the first text based on the first image and the image text; The comment publishing unit 54 is configured to publish comment information corresponding to the first text; the comment information carries the second image.

[0077] Optionally, the image acquisition unit 51 is specifically used to: search for a second text that is associated with the first text in the e-book; the first text and the second text are used to describe the same content in the e-book; based on the first text and the second text, generate at least one candidate image that matches the first text, and display the candidate image; in response to a selection instruction for the candidate image, determine the selected candidate image as the first image.

[0078] Optionally, the image acquisition unit 51 is further specifically used to: determine the image content that the candidate image needs to represent based on the first text and the second text through a generative model; and generate the candidate image based on the image content that the candidate image needs to represent through the generative model.

[0079] Optionally, the image acquisition unit 51 is also specifically used to: if there is a comment text to be published corresponding to the comment request, then through the generative model, determine the image content that the candidate image needs to represent based on the first text, the second text and the comment text to be published; through the generative model, determine the emotional atmosphere that the candidate image needs to represent based on the comment text to be published; through the generative model, generate the candidate image based on the image content that the candidate image needs to represent and the emotional atmosphere that the candidate image needs to represent.

[0080] Optionally, the method further includes a content determination unit for determining the character in the e-book represented by the first image and determining the character as the image content of the first image before obtaining the image text matching the first image based on the image content of the first image; or determining the scene in the e-book represented by the first image and determining the scene as the image content of the first image.

[0081] Optionally, the copy acquisition unit 52 is specifically used to: determine at least one text in the electronic book related to the image content of the first image, use the determined text as a candidate copy matching the first image, and display the candidate copy; determine the image copy based on the candidate copy.

[0082] Optionally, the image content of the first image includes a character or scene in the e-book; the text acquisition unit 52 is further specifically used to: if the image content of the first image includes a character in the e-book, determine the lines related to the character in the e-book as text related to the image content of the first image; if the image content of the first image includes a scene in the e-book, determine the sentences related to the scene in the e-book as text related to the image content of the first image.

[0083] Optionally, the copy acquisition unit 52 is further specifically configured to: in response to a selection instruction for the candidate copy, determine the selected candidate copy as the image copy; or, in response to an editing instruction for the candidate copy, determine the edited candidate copy as the image copy.

[0084] Through this embodiment, in a scenario where a user posts a comment on the first text in an e-book, a matching second image can be generated for the first text. The second image can be carried in the comment corresponding to the first text, which is equivalent to an emoticon package in the comment corresponding to the first text, thereby achieving the effect of generating an emoticon package according to user needs, reducing the difficulty for users to obtain matching emoticon packages, and improving the efficiency of users posting comment information.

[0085] The image generating device in the embodiment of the present disclosure can implement each process of the above-mentioned image generating method embodiment and achieve the same effects and functions, which will not be repeated here.

[0086] An embodiment of the present disclosure further provides an electronic device, Figure 6 A schematic diagram of the structure of an electronic device provided in one embodiment of the present disclosure is shown in FIG. Figure 6 As shown, electronic devices can vary significantly due to different configurations or performance. They may include one or more processors 601 and memory 602. Memory 602 may store one or more applications or data. Memory 602 may be either transient or persistent storage. Applications stored in memory 602 may include one or more modules (not shown), each of which may include a series of computer-executable instructions within the electronic device. Furthermore, processor 601 may be configured to communicate with memory 602 to execute the series of computer-executable instructions within memory 602 on the electronic device. The electronic device may also include one or more power supplies 603, one or more wired or wireless network interfaces 604, one or more input or output interfaces 605, one or more keyboards 606, and the like.

[0087] In a specific embodiment, the electronic device includes a processor; and a memory configured to store computer-executable instructions, wherein when the computer-executable instructions are executed, the processor implements the following process: In response to a comment request for a first text in an electronic book, obtaining a first image matching the first text; Acquiring an image text matching the first image according to the image content of the first image; generating a second image corresponding to the first text based on the first image and the image text; Posting comment information corresponding to the first text; the comment information carries the second image.

[0088] Through this embodiment, in a scenario where a user posts a comment on the first text in an e-book, a matching second image can be generated for the first text. The second image can be carried in the comment corresponding to the first text, which is equivalent to an emoticon package in the comment corresponding to the first text, thereby achieving the effect of generating an emoticon package according to user needs, reducing the difficulty for users to obtain matching emoticon packages, and improving the efficiency of users posting comment information.

[0089] The electronic device in the embodiment of the present disclosure can implement each process of the above-mentioned image generation method embodiment and achieve the same effects and functions, which will not be repeated here.

[0090] Another embodiment of the present disclosure further provides a computer-readable storage medium for storing computer-executable instructions. When the computer-executable instructions are executed by a processor, the following process is implemented: In response to a comment request for a first text in an electronic book, obtaining a first image matching the first text; Acquiring an image text matching the first image according to the image content of the first image; generating a second image corresponding to the first text based on the first image and the image text; Posting comment information corresponding to the first text; the comment information carries the second image.

[0091] Through this embodiment, in a scenario where a user posts a comment on the first text in an e-book, a matching second image can be generated for the first text. The second image can be carried in the comment corresponding to the first text, which is equivalent to an emoticon package in the comment corresponding to the first text, thereby achieving the effect of generating an emoticon package according to user needs, reducing the difficulty for users to obtain matching emoticon packages, and improving the efficiency of users posting comment information.

[0092] The computer-readable storage medium in the embodiment of the present disclosure can implement each process of the above-mentioned image generation method embodiment and achieve the same effects and functions, which will not be repeated here.

[0093] Another embodiment of the present disclosure further provides a computer program product, the computer program product including a computer program, which implements the following process when executed by a processor: In response to a comment request for a first text in an electronic book, obtaining a first image matching the first text; Acquiring an image text matching the first image according to the image content of the first image; generating a second image corresponding to the first text based on the first image and the image text; Posting comment information corresponding to the first text; the comment information carries the second image.

[0094] Through this embodiment, in a scenario where a user posts a comment on the first text in an e-book, a matching second image can be generated for the first text. The second image can be carried in the comment corresponding to the first text, which is equivalent to an emoticon package in the comment corresponding to the first text, thereby achieving the effect of generating an emoticon package according to user needs, reducing the difficulty for users to obtain matching emoticon packages, and improving the efficiency of users posting comment information.

[0095] The computer program product in the embodiment of the present disclosure can implement each process of the above-mentioned image generation method embodiment and achieve the same effects and functions, which will not be repeated here.

[0096] In various embodiments of the present disclosure, the computer-readable storage medium includes a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0097] In the 1990s, technological improvements could be clearly distinguished as either hardware improvements (for example, improvements to circuit structures like diodes, transistors, and switches) or software improvements (improvements to process flows). However, with the advancement of technology, many process flow improvements can now be considered direct improvements to hardware circuit structures. Designers almost always create the corresponding hardware circuit structure by programming the improved process flow into the hardware circuit. Therefore, it cannot be said that a process flow improvement cannot be implemented using physical hardware modules. For example, a programmable logic device (PLD), such as a field programmable gate array (FPGA), is an integrated circuit whose logical function is determined by user programming. Designers can "integrate" a digital system on a PLD by programming it themselves, without having to hire a chip manufacturer to design and manufacture a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly performed using software called a "logic compiler." This is similar to the software compilers used during program development. Before compilation, the original code must be written in a specific programming language, called a Hardware Description Language (HDL). There are many types of HDL, including ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art will also understand that simply by programming a method flow in one of these hardware description languages ​​and then programming it into an integrated circuit, a hardware circuit that implements the logic method flow can be easily obtained.

[0098] The controller can be implemented in any suitable manner. For example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the memory control logic. Those skilled in the art will also appreciate that, in addition to implementing the controller purely in computer-readable program code, the controller can also be implemented in the form of logic gates, switches, an application-specific integrated circuit, a programmable logic controller, an embedded microcontroller, etc. by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the means for implementing the various functions included therein can also be considered as structures within the hardware component. Alternatively, the means for implementing the various functions can be considered both a software module implementing the method and a structure within the hardware component.

[0099] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0100] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing the embodiments of the present disclosure, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0101] Those skilled in the art will appreciate that one or more embodiments of the present disclosure may be provided as a method, system, or computer program product. Thus, one or more embodiments of the present disclosure may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, one or more embodiments of the present disclosure may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0102] The present disclosure is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present disclosure. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0103] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0104] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0105] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0106] One or more embodiments of the present disclosure may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. One or more embodiments of the present disclosure may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communications network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0107] The various embodiments of this disclosure are described in a progressive manner. Similar portions between the various embodiments can be referenced to each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiments are generally similar to the method embodiments, so their description is relatively simple. For relevant portions, refer to the description of the method embodiments.

[0108] The foregoing is merely an embodiment of the present disclosure and is not intended to limit the present disclosure. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present disclosure are intended to be included within the scope of the claims of the present disclosure.

Claims

1. An image generation method, characterized in that: include: In response to a comment request for a first text in an electronic book, obtaining a first image matching the first text; Acquiring an image text matching the first image according to the image content of the first image; generating a second image corresponding to the first text based on the first image and the image text; Publishing comment information corresponding to the first text; The comment information carries the second image.

2. The method according to claim 1, characterized in that The acquiring of a first image matching the first text includes: searching the electronic book for a second text associated with the first text, wherein the first text and the second text describe the same content in the electronic book; generating at least one candidate image matching the first text according to the first text and the second text, and displaying the candidate image; In response to a selection instruction for the candidate image, the selected candidate image is determined as the first image.

3. The method according to claim 2, characterized in that Generating at least one candidate image matching the first text according to the first text and the second text includes: Determining, by a generative model, image content that the candidate image needs to represent based on the first text and the second text; The candidate image is generated by the generative model according to the image content that the candidate image needs to represent.

4. The method according to claim 2, characterized in that Generating at least one candidate image matching the first text according to the first text and the second text includes: If there is a to-be-published comment text corresponding to the comment request, determining, by a generative model, the image content that the candidate image needs to represent based on the first text, the second text, and the to-be-published comment text; Determining the emotional atmosphere that the candidate image needs to represent based on the comment text to be published using the generative model; The candidate image is generated by the generative model according to the image content that the candidate image needs to represent and the emotional atmosphere that the candidate image needs to represent.

5. The method according to claim 1, wherein Before acquiring an image text matching the first image according to the image content of the first image, the method further includes: determining a character in the electronic book represented by the first image, and determining the character as image content of the first image; or, A scene in the electronic book represented by the first image is determined, and the scene is determined as image content of the first image.

6. The method according to claim 1, characterized in that The acquiring, based on the image content of the first image, an image text matching the first image includes: determining at least one text in the electronic book that is related to the image content of the first image, using the determined text as a candidate text matching the first image, and displaying the candidate text; The image text is determined based on the candidate texts.

7. The method according to claim 6, characterized in that The image content of the first image includes a character or scene in the electronic book; and determining at least one text in the electronic book related to the image content of the first image includes: If the image content of the first image includes a character in the electronic book, determining lines related to the character in the electronic book as text related to the image content of the first image; If the image content of the first image includes a scene in the electronic book, a sentence in the electronic book related to the scene is determined as text related to the image content of the first image.

8. The method according to claim 6, characterized in that The determining the image text based on the candidate texts includes: In response to a selection instruction for the candidate text, determining the selected candidate text as the image text; or, In response to an editing instruction for the candidate text, the edited candidate text is determined as the image text.

9. An image generating device, characterized in that: include: An image acquiring unit, configured to acquire a first image matching the first text in response to a comment request for the first text in the electronic book; a text acquisition unit, configured to acquire an image text matching the first image according to the image content of the first image; an image generating unit, configured to generate a second image corresponding to the first text based on the first image and the image text; A comment publishing unit, configured to publish comment information corresponding to the first text; The comment information carries the second image.

10. An electronic device, characterized in that: include: processor; as well as, A memory configured to store computer-executable instructions, which, when executed, cause the processor to implement the method according to any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that The computer-readable storage medium is used to store computer-executable instructions, and the computer-executable instructions implement the method according to any one of claims 1 to 8 when executed by a processor.

12. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Information sending method, device and equipment

    CN107612815A

  • Emoji package generation method, device and equipment and medium

    CN111353064A

  • Method and device for generating comment reply and medium

    CN115269828A

  • Cboth recommendation method and device, electronic equipment and storage medium

    CN117668281A

  • Image generation method and device and related product

    CN119991882A