Image generation method, apparatus, and related product
By generating images and text that match the e-book text using a generative model, the problem of existing emojis failing to meet user needs is solved, and the efficiency of publishing comment information is improved.
Patent Information
- Application Number
- CN202510786934.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2045-06-12
AI Technical Summary
The existing emoji sets are insufficient to meet users' needs when posting comments, making it more difficult to obtain matching emoji sets and reducing users' posting efficiency.
Generative models generate matching images based on the text content and related text in the ebook, and combine the image content to obtain matching image text, ultimately generating emoticons carried in the comment information.
It enables the generation of matching emojis based on user needs, reducing the difficulty for users to obtain matching emojis and improving the efficiency of posting comments.
Smart Images

Figure CN120672909B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to an image generation method, apparatus and related products. Background Technology
[0002] In related technologies, users can post comments on paragraphs or chapters of interest while reading ebooks. Users often need to use emojis when posting comments. However, current emojis are mainly built-in emojis provided by input methods or reading applications. These built-in emojis may not match the user's emoji usage needs when posting comments, increasing the difficulty for users to obtain matching emojis and reducing the efficiency of posting comments. Summary of the Invention
[0003] This disclosure provides an image generation method, apparatus, and related products that can generate a matching second image for a first text in an e-book when a user posts a comment on the first text. The second image can be carried in the comment information corresponding to the first text, essentially acting as an emoticon in the comment information corresponding to the first text. This achieves the effect of generating emoticons according to user needs, reducing the difficulty for users to obtain matching emoticons and improving the efficiency of users posting comments.
[0004] In a first aspect, embodiments of this disclosure provide an image generation method, including:
[0005] In response to a comment request for a first text in the ebook, a first image matching the first text is obtained;
[0006] Based on the image content of the first image, obtain the image text that matches the first image;
[0007] Based on the first image and the image text, generate a second image corresponding to the first text;
[0008] The comment information corresponding to the first text is published; the comment information carries the second image.
[0009] Secondly, embodiments of this disclosure provide an image generation apparatus, comprising:
[0010] An image acquisition unit is configured to acquire a first image matching the first text in response to a comment request for the first text in an e-book;
[0011] The text acquisition unit is used to acquire image text that matches the first image based on the image content of the first image;
[0012] An image generation unit is configured to generate a second image corresponding to the first text based on the first image and the image text.
[0013] A comment publishing unit is used to publish comment information corresponding to the first text; the comment information carries the second image.
[0014] Thirdly, embodiments of this disclosure provide an electronic device, including: a processor; and a memory configured to store computer-executable instructions, which, when executed, cause the processor to perform the method described in the first aspect above.
[0015] Fourthly, embodiments of this disclosure provide a computer-readable storage medium for storing computer-executable instructions that, when executed by a processor, implement the method described in the first aspect.
[0016] Fifthly, embodiments of this disclosure provide a computer program product, the computer program product including a computer program, which, when executed by a processor, implements the method described in the first aspect above.
[0017] In one or more embodiments of this disclosure, firstly, in response to a comment request for first text in an e-book, a first image matching the first text is obtained. Then, based on the image content of the first image, image text matching the first image is obtained. Next, based on the first image and the image text, a second image corresponding to the first text is generated. Finally, comment information corresponding to the first text is published, and the comment information carries the second image. Therefore, through this embodiment, in scenarios where a user publishes comment information for first text in an e-book, a matching second image can be generated for the first text. This second image can be carried in the comment information corresponding to the first text, essentially acting as an emoticon within the comment information, thereby achieving the effect of generating emoticons according to user needs, reducing the difficulty for users to obtain matching emoticons, and improving the efficiency of users publishing comment information. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in one or more embodiments or related technologies of this disclosure, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments recorded in this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 A schematic flowchart of an image generation method provided in an embodiment of this disclosure.
[0020] Figure 2a This is a schematic diagram illustrating the posting of comment information to a first text according to an embodiment of the present disclosure;
[0021] Figure 2b A schematic diagram of a comment page provided for an embodiment of this disclosure;
[0022] Figure 2c A schematic diagram of a candidate image provided in an embodiment of this disclosure;
[0023] Figure 2d A schematic diagram of a first image provided for an embodiment of this disclosure;
[0024] Figure 3a A schematic diagram illustrating a candidate document provided in an embodiment of this disclosure;
[0025] Figure 3b A schematic diagram illustrating text selection provided for an embodiment of this disclosure;
[0026] Figure 4a A schematic diagram of a second image provided for an embodiment of this disclosure;
[0027] Figure 4b This is a schematic diagram illustrating the posting of comments according to an embodiment of this disclosure;
[0028] Figure 5 This is a schematic diagram of the structure of an image generation apparatus provided in an embodiment of the present disclosure;
[0029] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure. Detailed Implementation
[0030] To enable those skilled in the art to better understand the technical solutions in one or more embodiments of this disclosure, the technical solutions in one or more embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of the embodiments. Based on one or more embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of this disclosure.
[0031] It is understood that before using the technical solutions disclosed in the embodiments of this disclosure, relevant parties should be informed of the type, scope of use, and usage scenarios of the information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and authorization from the relevant parties should be obtained.
[0032] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0033] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0034] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0035] Considering that in related technologies, when users post comments on e-books, the built-in emoticons provided by input methods or reading applications are insufficient to meet users' emoticon usage needs, increasing the difficulty for users to obtain matching emoticons and reducing the efficiency of users posting comments, this disclosure provides an image generation method, apparatus, and related products. These methods can generate a matching second image for a first text in an e-book, which can be carried within the comment information corresponding to the first text, essentially acting as an emoticon within the comment information. This achieves the effect of generating emoticons according to user needs, reducing the difficulty for users to obtain matching emoticons and improving the efficiency of users posting comments.
[0036] The image generation method can be applied to and executed by terminal devices, including but not limited to laptops, tablets, desktop computers, set-top boxes, mobile devices (e.g., mobile phones, portable music players, personal digital assistants, dedicated messaging devices, portable gaming devices), smartphones, smart speakers, smartwatches, smart TVs, in-vehicle terminals, and other types of user terminals.
[0037] Figure 1 This is a schematic flowchart of an image generation method provided in an embodiment of the present disclosure, as shown below. Figure 1 As shown, the method includes:
[0038] Step S102: In response to a comment request for the first text in the e-book, obtain a first image that matches the first text;
[0039] Step S104: Obtain image text matching the first image based on the image content of the first image;
[0040] Step S106: Based on the first image and the image text, generate the second image corresponding to the first text;
[0041] Step S108: Publish the comment information corresponding to the first text; the comment information carries the second image.
[0042] In this embodiment, firstly, in response to a comment request for the first text in an e-book, a first image matching the first text is obtained. Then, based on the image content of the first image, image text matching the first image is obtained. Next, based on the first image and image text, a second image corresponding to the first text is generated. Finally, comment information corresponding to the first text is published, and the comment information carries the second image. Therefore, this embodiment enables the generation of a matching second image for the first text in scenarios where a user publishes comment information for the first text in an e-book. This second image can be carried in the comment information corresponding to the first text, essentially acting as an emoticon within the comment information, thereby achieving the effect of generating emoticons according to user needs, reducing the difficulty for users to obtain matching emoticons, and improving the efficiency of users publishing comment information.
[0043] In step S102 above, while reading an e-book, the user can post comments on the first text in the e-book. The first text can be any length of text in the e-book. For example, the first text can be one or more sentences in the e-book, or a paragraph in the e-book, or a chapter in the e-book.
[0044] Users can comment on a given text by triggering the comment component associated with that text. For example, if the first text is one or more sentences in an ebook, selecting those sentences triggers the associated "Post Comment" control, thus commenting on those sentences. In this case, the user's comment is considered a comment on that specific sentence or sentence. Similarly, if the first text is a paragraph in an ebook, selecting all or part of that paragraph triggers the associated "Post Paragraph Comment" control, thus commenting on that paragraph. In this case, the user's comment is considered a comment on that paragraph. Furthermore, if the first text is a chapter in an ebook, the user can trigger the corresponding comment box at the end of that chapter, thus commenting on that chapter. In this case, the user's comment is considered a comment on that chapter.
[0045] The terminal device generates a comment request for the first text based on the user's comment action on the first text, and in response to the comment request, retrieves a first image that matches the first text. The first image is an image that matches the text content of the first text.
[0046] In one embodiment, obtaining a first image that matches the first text includes:
[0047] Find a second text in the ebook that is related to the first text; the first and second texts describe the same content in the ebook.
[0048] Based on the first text and the second text, generate at least one candidate image that matches the first text, and display the candidate image;
[0049] In response to the selection instruction for the candidate image, the selected candidate image is determined as the first image.
[0050] In this embodiment, firstly, a second text related to the first text is searched within the e-book. This relationship means that both the first and second texts describe the same content within the e-book. For example, if the first text describes the appearance of a character in the e-book, the second text can also describe the character's appearance or actions; the shared content here refers to that character. Similarly, if the first text describes a scene in the e-book, such as a party scene, the second text can also describe that party scene; the shared content here refers to that party scene.
[0051] Then, based on the first text and the second text, at least one candidate image matching the first text is generated, and each candidate image is displayed. The first text and the second text can be used as image-generating prompts, which are input into a generative model with image-generating capabilities. This generative model then generates at least one candidate image matching the first text and the second text. The terminal device can then display the generated candidate images.
[0052] Finally, the user can select one candidate image from the various candidate images. In response to the user's selection instruction for a candidate image, the candidate image selected by the user is determined as the first image. Since the first image is selected from the various candidate images, and the candidate image is generated based on first text and second text, which describe the same content in the e-book, the image content of each candidate image matches the text content of the first text, and the image content of the first image also matches the text content of the first text. It is understood that in this embodiment, the terminal device also displays the first image.
[0053] In this embodiment, after determining the first text, the terminal device can send the first text to the server corresponding to the reading application where the e-book is located, triggering the server to search for a second text in the e-book that is related to the first text. Based on the first text and the second text, at least one candidate image matching the first text is generated, and each candidate image is returned to the terminal device. The terminal device then displays each candidate image and determines the first image based on the user's selection instruction.
[0054] Figure 2a This is a schematic diagram illustrating the posting of comment information to a first text in accordance with an embodiment of this disclosure. Figure 2b This is a schematic diagram of a comment page provided according to an embodiment of the present disclosure. Figure 2c This is a schematic diagram of a candidate image provided in one embodiment of the present disclosure. Figure 2d This is a schematic diagram of a first image provided for an embodiment of this disclosure. (See diagram below.) Figure 2a As shown, taking the example of a user posting a comment on a section of text in an e-book, after the user selects all the characters in the text section, the terminal device can display the operation bar associated with that text section. The user can trigger the "Post Comment" control associated with that text section in the operation bar, and the terminal device will respond to the triggered operation by displaying, as shown below. Figure 2b On the comment page shown, users can trigger the "AI emoji" control provided on the page. In response to this trigger, the terminal device uses the text of the previous comment as the first text, and through the above process, generates multiple candidate images matching the first text, and then... Figure 2c The image displayed shows the candidate images. These candidate images can be displayed on the same page, or you can swipe horizontally to switch between them. Figure 2c The example shown is a landscape-style swipe-to-swipe display of candidate images. Figure 2d As shown, the user can select the first image from among the candidate images, and obviously, the terminal device displays the first image. Figure 2d As shown, users can click the "Next" component to continue with the next steps.
[0055] As can be seen, through this embodiment, it is possible to find a second text that is related to the first text in the e-book. The first text and the second text are used to describe the same content in the e-book. Based on the first text and the second text, at least one candidate image that matches the first text is generated and each candidate image is displayed. In response to the selection instruction for the candidate image, the selected candidate image is determined as the first image, thereby accurately generating a first image that matches the text content of the first text, improving the efficiency and accuracy of generating the first image.
[0056] In one embodiment, generating at least one candidate image matching the first text based on the first text and the second text includes:
[0057] The generative model determines the image content that the candidate image needs to represent based on the first and second texts.
[0058] Generative models are used to generate candidate images based on the image content that the candidate images are intended to represent.
[0059] In this embodiment, the first text and the second text can be used as image-generating prompts. These prompts are input into a generative model with image-generating capabilities. Based on the first and second texts, the generative model determines the image content that a candidate image should represent. The image content that the candidate image should represent is the content described by the first and second texts, such as the character or scene described in the first and second texts. Then, the generative model generates at least one candidate image based on the image content that the candidate image should represent. This candidate image is used to represent the aforementioned image content.
[0060] For example, if the first text describes a character in an e-book and the second text also describes the same character, then the generative model determines the image content that the candidate image needs to represent as the character based on the first and second texts. The generative model then generates multiple candidate images based on the image content that the candidate images need to represent, and each candidate image is a character image of the character.
[0061] In this embodiment, after determining the first text, the terminal device sends the first text to the server corresponding to the reading application where the e-book is located. This triggers the server to search for a second text related to the first text in the e-book. Using a generative model, the server determines the image content to be represented by candidate images based on the first and second texts, and generates candidate images according to this content. The server then returns each candidate image to the terminal device, allowing the terminal device to display the candidate images and determine the first image based on the user's selection.
[0062] As can be seen, through this embodiment, the generative model can determine the image content that the candidate image needs to represent based on the first text and the second text, and generate the candidate image based on the image content that the candidate image needs to represent. By utilizing the accurate image generation capability of the generative model, the efficiency and accuracy of generating candidate images can be improved.
[0063] In one embodiment, generating at least one candidate image matching the first text based on the first text and the second text includes:
[0064] If there exists a text for a comment to be published corresponding to the above comment request, then the generative model is used to determine the image content that the candidate image should represent based on the first text, the second text, and the text for the comment to be published.
[0065] Generative models are used to determine the emotional atmosphere that candidate images are intended to represent, based on the text of the comment to be published.
[0066] Generative models are used to generate candidate images based on the image content and emotional atmosphere that the candidate images are intended to represent.
[0067] In this embodiment, if there exists a comment text to be published corresponding to the aforementioned comment request—for example, if the user entered comment text in the comment box before triggering the "AI emoji" control—then this comment text is the comment text to be published. The generative model then determines the image content to be represented by the candidate image based on the first text, the second text, and the comment text to be published. The generative model can use the content represented by each of the first text, the second text, and the comment text to be published as the image content to be represented by the candidate image. For example, if the first text and the second text represent a character in an e-book, and the comment text to be published expresses the desire to hug that character, then both the character and the hugging action towards that character are used as the image content to be represented by the candidate image.
[0068] Next, a generative model is used to determine the emotional atmosphere that the candidate images should represent, based on the text of the comment to be published. The generative model can determine the emotional atmosphere expressed by the text of the comment to be published, and then use this emotional atmosphere as the emotional atmosphere that the candidate images should represent. For example, if the text of the comment to be published is "I really like this scene," then the emotional atmosphere expressed by the text is determined to be happiness and joy, and thus the emotional atmosphere that the candidate images should represent is happiness and joy. Similarly, if the text of the comment to be published is "I also want to be as ambitious as him," then the emotional atmosphere expressed by the text is determined to be inspirational and striving, and thus the emotional atmosphere that the candidate images should represent is inspirational and striving.
[0069] Finally, a generative model generates at least one candidate image based on the image content and emotional atmosphere that the candidate image is intended to represent. The generative model can determine the image style of the candidate image based on the emotional atmosphere it is intended to represent, and thus generate each candidate image based on its image content and style. Examples of candidate image styles include comic book style, cartoon style, and ink painting style.
[0070] In this embodiment, after determining the first text and the text to be published as a comment, the terminal device sends the first text and the text to be published as a comment to the server corresponding to the reading application where the e-book is located. This triggers the server to search for a second text in the e-book that is related to the first text. Then, using a generative model, based on the first text, the second text, and the text to be published as a comment, the server determines the image content that the candidate images should represent, and based on the text to be published as a comment, determines the emotional atmosphere that the candidate images should represent. Based on the image content and the emotional atmosphere that the candidate images should represent, the server generates various candidate images. The server then returns each candidate image to the terminal device, which displays the candidate images and determines the first image based on the user's selection instruction.
[0071] As can be seen, this embodiment also takes into account the situation where there is a comment text to be published corresponding to the above-mentioned comment request. A generative model can be used to determine the image content that the candidate image needs to represent based on the first text, the second text, and the comment text to be published. Based on the comment text to be published, the emotional atmosphere that the candidate image needs to represent can be determined. Based on the image content and the emotional atmosphere that the candidate image needs to represent, a candidate image is generated. Thus, combined with the comment text that the user wants to publish, the candidate image is accurately generated, improving the accuracy of the candidate image.
[0072] After generating various candidate images and determining the first image, in step S104 above, the terminal device obtains image text that matches the first image based on its image content. The image text matches the image content of the first image. The second image obtained by overlaying the first image and the image text is the emoticon that satisfies the user's emoticon referencing needs.
[0073] In one embodiment, before obtaining the image text matching the first image based on the image content of the first image, the method further includes:
[0074] Identify the character in the e-book represented by the first image, and define that character as the image content of the first image;
[0075] or,
[0076] Determine the scene in the e-book represented by the first image, and define that scene as the image content of the first image.
[0077] In this embodiment, before acquiring the image text, the terminal device further determines whether the first image is used to represent a character or a scene in the e-book. If it is determined that the first image is used to represent a character in the e-book, then the character represented by the first image, such as character A, is determined, and character A is determined as the image content of the first image. If it is determined that the first image is used to represent a scene in the e-book, then the scene represented by the first image, such as a party scene, is determined, and the party scene is determined as the image content of the first image.
[0078] In this embodiment, the terminal device can determine the image content of the first image through the above process, or the terminal device can send the image identifier of the first image to the server corresponding to the reading application where the e-book is located, triggering the server to determine the image content of the first image through the above process.
[0079] As can be seen, through this embodiment, the characters or scenes in the e-book represented by the first image can be determined as the image content of the first image. Thus, when determining the image text that matches the first image based on the image content of the first image, the image text can be matched with the characters or scenes in the e-book, improving the correlation between the image text and the e-book, so that the final generated second image, i.e., the emoticon, is closely related to the e-book.
[0080] After determining the image content of the first image, in one embodiment, based on the image content of the first image, image text matching the first image is obtained, including:
[0081] Identify at least one text in the ebook that is related to the image content of the first image, select the identified text as candidate text that matches the first image, and display the candidate text.
[0082] Based on the candidate text, determine the text for the image.
[0083] In this embodiment, the terminal device can determine at least one text in the e-book that is related to the image content of the first image, use the determined text as a candidate text that matches the first image, display each candidate text, and determine the image text based on each candidate text.
[0084] Of course, the terminal device can also trigger the aforementioned server to determine at least one text in the e-book that is related to the image content of the first image, and use the determined text as a candidate text to match the first image. The server sends each candidate text to the terminal device, so that the terminal device displays each candidate text and determines the image text based on each candidate text.
[0085] At least one text in the e-book related to the image content of the first image may be text used in the e-book to introduce the image content of the first image, or classic lines of dialogue of the character represented by the image content of the first image, or classic quotes of the scene represented by the image content of the first image.
[0086] As can be seen, through this embodiment, at least one text in the e-book that is related to the image content of the first image can be identified. The identified text is used as a candidate text to match the first image, and each candidate text is displayed. Based on the candidate text, the image text is determined, thereby making the image text come from the e-book and related to the image content of the first image, improving the correlation between the image text and the e-book, improving the correlation between the image text and the first image, and improving the accuracy of the image text.
[0087] As described above, the image content of the first image includes characters or scenes from the e-book. Therefore, in one embodiment, determining at least one piece of text in the e-book related to the image content of the first image includes:
[0088] If the image content of the first image includes a character from the e-book, then the lines in the e-book related to that character are determined as text related to the image content of the first image;
[0089] If the image content of the first image includes a scene from the e-book, then sentences in the e-book related to that scene are identified as text related to the image content of the first image.
[0090] In this embodiment, in one scenario, if the image content of the first image includes a character from an e-book, then the dialogue related to that character in the e-book is determined as text related to the image content of the first image. The dialogue related to that character can be popular lines from that character. That is, when the first image is a character image, the candidate text can be popular lines from that character in the e-book, thereby combining the character and the character's popular lines to generate an emoticon.
[0091] In another scenario, if the content of the first image includes a scene from an ebook, then sentences from the ebook related to that scene are identified as text related to the content of the first image. These sentences can be classic text from the ebook describing the scene. In other words, when the first image is a scene image, the candidate text can be classic text from the ebook describing the scene, thus combining the scene and the classic text describing the scene to generate an emoji.
[0092] As can be seen, through this embodiment, based on the character or scene represented by the first image, lines of dialogue of the character or sentences describing the scene can be selected from the e-book as candidate text, so that the image text is closely related to the image content of the first image, thereby generating matching emoticons when users post comments on the first text, improving the accuracy and efficiency of emoticon generation.
[0093] In this embodiment, in addition to generating candidate text through the methods described above, in another scenario, a generative model can be used to generate candidate text based on the text of the comment to be published corresponding to the comment request. For example, the text of the comment to be published is input into the generative model, and the generative model generates candidate text based on the text of the comment to be published.
[0094] In another scenario, a generative model can be used to generate candidate text based on the first text. For example, the first text can be input into a generative model, which then generates candidate text based on it. Alternatively, the role or scene represented by the first text can be determined, and text describing that role or scene in an e-book can be obtained. This text may include the first text. The obtained text can then be input into a generative model, which generates candidate text based on the input text.
[0095] In another scenario, a generative model can be used to generate candidate text based on the text in the e-book that describes the characters or scenes in the first image. For example, obtain the paragraph in the e-book that describes the characters or scenes in the first image, input the paragraph into the generative model, and the generative model will generate candidate text based on the input paragraph.
[0096] In all the above scenarios, appropriate prompts such as "Please generate text snippets from emojis based on the input text" can be input into the generative model to facilitate the generation of candidate texts. The candidate texts are the text snippets from emojis.
[0097] In one embodiment, determining image text based on candidate text includes:
[0098] In response to the selection instruction for candidate text, the selected candidate text is determined as image text;
[0099] or,
[0100] In response to an editing instruction for the candidate text, the edited candidate text is determined as the image text.
[0101] In this embodiment, after displaying each candidate text, the terminal device can, in response to the user's selection instruction for the candidate text, determine the selected candidate text as image text, or, in response to the user's editing instruction for the candidate text, determine the edited candidate text as image text. In this embodiment, the text content, font, font size, color, style, etc., of the candidate text can be edited.
[0102] As can be seen, through this embodiment, users can obtain image text by selecting candidate text or editing candidate text, thus meeting users' needs for obtaining image text.
[0103] After obtaining the image text, in step S106 above, a second image corresponding to the first text is generated based on the first image and the image text, and the second image is displayed. The terminal device can overlay the image text onto the first image to obtain the second image, which is the emoticon. The terminal device can determine the overlay position of the image text within the first image according to user settings. The terminal device can also display the second image.
[0104] In step S108 above, the terminal device publishes comment information corresponding to the first text, which carries a second image. This comment information may or may not carry the comment text to be published.
[0105] Figure 3a This is a schematic diagram illustrating a candidate document provided in an embodiment of this disclosure. Figure 3b This is a schematic diagram illustrating text selection provided in one embodiment of the present disclosure. Figure 4a This is a schematic diagram of a second image provided according to an embodiment of the present disclosure. Figure 4b This is a schematic diagram illustrating the posting of comments according to an embodiment of this disclosure. Figure 3a It is possible Figure 2d Then it will appear. For example... Figure 2d and Figure 3a As shown, after the user selects the first image from the candidate images, the terminal device can obtain and display the candidate texts that match the first image through the above process. For example... Figure 3b As shown, users can select a candidate copy from the various candidate copy options and click "Next" to continue.
[0106] like Figure 4a As shown, after the user selects a candidate text, the terminal device displays the first image and the selected text. Figure 4a On the page shown, users can choose to edit or not edit the selected text. Users can also manually drag and drop the selected text within the first image to create a second image. The second image may display a corresponding "Comment with Image" component; if the user triggers this component, they will be redirected to... Figure 4b Users can Figure 4bEnter or leave the comment blank, or modify the previously entered comment, and click the "Publish" component to publish the comment for the first text.
[0107] Figures 2a to 4b In this context, a component is indicated to be triggered by filling it with gray.
[0108] In one embodiment, after selecting the first image, the user can directly input image text and set the position of the image text in the first image to obtain a second image, and then publish comment information on the first text based on the second image. That is, the terminal device obtains the image text input by the user in the first image, generates a second image based on the first image and the image text, and publishes comment information on the first text based on the second image.
[0109] Based on the solutions described above, in one embodiment, the system can record the number of AI-generated images created by users for various texts in an e-book. Based on this record, texts with a high frequency of AI-generated images can be identified and highlighted in the e-book, prompting the user to comment on those texts. This identified text is the first text mentioned above. When a user requests to post a comment on the first text, a second image can be generated as an AI emoticon through the above process, and this AI emoticon is included in the user's comment.
[0110] In one embodiment, in step S102 above, in response to a comment request for the first text in the e-book, obtaining a first image matching the first text can be: searching for a second text in the e-book that is related to the first text, where the first text and the second text are used to describe the same content in the e-book; generating at least one candidate image matching the first text based on the first text and the second text; and using each candidate image as the first image.
[0111] In step S104 above, obtaining image text matching the first image based on the image content of the first image can include at least one of the following methods:
[0112] For each first image, if the image content of the first image includes a character from the e-book, then the lines from the e-book related to that character will be used as the image text to match the first image.
[0113] For each first image, if the image content of the first image includes a scene from the e-book, then the sentence in the e-book that describes the scene is used as the image text that matches the first image.
[0114] For each first image, the text to be published for the comment is input into the generative model, and the generative model generates image text matching each first image based on the text to be published for the comment.
[0115] For each first image, the first text is input into the generative model, and the generative model generates image text matching each first image based on the first text;
[0116] For each first image, determine the character or scene represented by the first text, obtain the text in the e-book used to describe the character or scene, which may include the first text, input the obtained text into the generative model, and generate image text matching each first image based on the input text.
[0117] For each first image, obtain the paragraph in the e-book that describes the character or scene in the first image, input the paragraph into the generative model, and obtain the image text matching each first image based on the input paragraph.
[0118] In step S106 above, generating a second image corresponding to the first text based on the first image and the image text can be achieved by combining each first image with the matched image text to obtain multiple second images. The terminal device also displays each second image as a candidate emoticon.
[0119] Furthermore, in one scenario, the terminal device can, based on the user's selection, display the target emoji selected by the user in the comment information.
[0120] In another scenario, the terminal device can also determine the candidate emojis selected by the user based on their actions and enter the emoji editing panel. In this panel, the user can choose to edit or not edit the candidate emojis. Editing a candidate emoji includes editing either the first image or the text accompanying the image. Based on the user's confirmation, the terminal device will either select the edited candidate emoji or the unedited candidate emoji as the target emoji. The terminal device will then post the target emoji in the comment section.
[0121] In the above embodiments of this disclosure, "AI" refers to Artificial Intelligence.
[0122] In summary, through the above embodiments, when a user posts a comment on the first text in an e-book, a matching second image can be generated for the first text. This second image can be carried in the comment information corresponding to the first text, which is equivalent to an emoticon in the comment information corresponding to the first text. This achieves the effect of generating emoticons according to user needs, reduces the difficulty for users to obtain matching emoticons, and improves the efficiency of users posting comments.
[0123] Figure 5This is a schematic diagram of the structure of an image generation apparatus provided in an embodiment of the present disclosure, as shown below. Figure 5 As shown, the device includes:
[0124] Image acquisition unit 51 is configured to acquire a first image matching the first text in response to a comment request for the first text in the e-book;
[0125] The text acquisition unit 52 is used to acquire image text matching the first image based on the image content of the first image;
[0126] Image generation unit 53 is used to generate a second image corresponding to the first text based on the first image and the image text;
[0127] The comment publishing unit 54 is used to publish comment information corresponding to the first text; the comment information carries the second image.
[0128] Optionally, the image acquisition unit 51 is specifically configured to: search for a second text in the e-book that is related to the first text; the first text and the second text are used to describe the same content in the e-book; generate at least one candidate image that matches the first text based on the first text and the second text, and display the candidate image; and, in response to a selection instruction for the candidate image, determine the selected candidate image as the first image.
[0129] Optionally, the image acquisition unit 51 is further configured to: determine the image content that the candidate image is required to represent based on the first text and the second text using a generative model; and generate the candidate image based on the image content that the candidate image is required to represent using the generative model.
[0130] Optionally, the image acquisition unit 51 is further configured to: if there is a comment text to be published corresponding to the comment request, determine the image content to be represented by the candidate image based on the first text, the second text, and the comment text to be published using a generative model; determine the emotional atmosphere to be represented by the candidate image based on the comment text to be published using the generative model; and generate the candidate image based on the image content to be represented by the candidate image and the emotional atmosphere to be represented by the candidate image using the generative model.
[0131] Optionally, it further includes a content determination unit, configured to determine the character in the e-book represented by the first image and determine the character as the image content of the first image before obtaining the image text matching the first image based on the image content of the first image; or, to determine the scene in the e-book represented by the first image and determine the scene as the image content of the first image.
[0132] Optionally, the text acquisition unit 52 is specifically configured to: determine at least one text in the e-book that is related to the image content of the first image, use the determined text as a candidate text that matches the first image, and display the candidate text; and determine the image text based on the candidate text.
[0133] Optionally, the image content of the first image includes characters or scenes from the e-book; the text acquisition unit 52 is further specifically configured to: if the image content of the first image includes characters from the e-book, then determine the lines in the e-book related to the characters as text related to the image content of the first image; if the image content of the first image includes scenes from the e-book, then determine the sentences in the e-book related to the scenes as text related to the image content of the first image.
[0134] Optionally, the text acquisition unit 52 is further configured to: in response to a selection instruction for the candidate text, determine the selected candidate text as the image text; or, in response to an editing instruction for the candidate text, determine the edited candidate text as the image text.
[0135] This embodiment enables the generation of a matching second image for a first text in an e-book when a user posts a comment. This second image can be carried in the comment corresponding to the first text, essentially acting as an emoticon in the comment, thereby achieving the effect of generating emoticons according to user needs, reducing the difficulty for users to obtain matching emoticons, and improving the efficiency of users posting comments.
[0136] The image generation apparatus in this embodiment can implement the various processes of the above-described image generation method embodiments and achieve the same effects and functions, which will not be repeated here.
[0137] One embodiment of this disclosure also provides an electronic device. Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure, as shown below. Figure 6As shown, electronic devices can vary considerably due to differences in configuration or performance. They may include one or more processors 601 and memories 602, with the memory 602 storing one or more application programs or data. The memory 602 can be temporary or persistent storage. The application programs stored in the memory 602 may include one or more modules (not shown), each module including a series of computer-executable instructions within the electronic device. Furthermore, the processor 601 may be configured to communicate with the memory 602, executing the series of computer-executable instructions stored in the memory 602 on the electronic device. The electronic device may also include one or more power supplies 603, one or more wired or wireless network interfaces 604, one or more input or output interfaces 605, one or more keyboards 606, etc.
[0138] In one specific embodiment, the electronic device includes a processor; and a memory configured to store computer-executable instructions, which, when executed, cause the processor to perform the following process:
[0139] In response to a comment request for a first text in the ebook, a first image matching the first text is obtained;
[0140] Based on the image content of the first image, obtain the image text that matches the first image;
[0141] Based on the first image and the image text, generate a second image corresponding to the first text;
[0142] The comment information corresponding to the first text is published; the comment information carries the second image.
[0143] This embodiment enables the generation of a matching second image for a first text in an e-book when a user posts a comment. This second image can be carried in the comment corresponding to the first text, essentially acting as an emoticon in the comment, thereby achieving the effect of generating emoticons according to user needs, reducing the difficulty for users to obtain matching emoticons, and improving the efficiency of users posting comments.
[0144] The electronic device in this embodiment can implement the various processes of the above-described image generation method embodiments and achieve the same effects and functions, which will not be repeated here.
[0145] Another embodiment of this disclosure also provides a computer-readable storage medium for storing computer-executable instructions that, when executed by a processor, implement the following process:
[0146] In response to a comment request for a first text in the ebook, a first image matching the first text is obtained;
[0147] Based on the image content of the first image, obtain the image text that matches the first image;
[0148] Based on the first image and the image text, generate a second image corresponding to the first text;
[0149] The comment information corresponding to the first text is published; the comment information carries the second image.
[0150] This embodiment enables the generation of a matching second image for a first text in an e-book when a user posts a comment. This second image can be carried in the comment corresponding to the first text, essentially acting as an emoticon in the comment, thereby achieving the effect of generating emoticons according to user needs, reducing the difficulty for users to obtain matching emoticons, and improving the efficiency of users posting comments.
[0151] The computer-readable storage medium in this disclosure embodiment can implement the various processes of the above-described image generation method embodiment and achieve the same effects and functions, which will not be repeated here.
[0152] Another embodiment of this disclosure also provides a computer program product, the computer program product including a computer program, which, when executed by a processor, implements the following process:
[0153] In response to a comment request for a first text in the ebook, a first image matching the first text is obtained;
[0154] Based on the image content of the first image, obtain the image text that matches the first image;
[0155] Based on the first image and the image text, generate a second image corresponding to the first text;
[0156] The comment information corresponding to the first text is published; the comment information carries the second image.
[0157] This embodiment enables the generation of a matching second image for a first text in an e-book when a user posts a comment. This second image can be carried in the comment corresponding to the first text, essentially acting as an emoticon in the comment, thereby achieving the effect of generating emoticons according to user needs, reducing the difficulty for users to obtain matching emoticons, and improving the efficiency of users posting comments.
[0158] The computer program product in this disclosure embodiment can implement the various processes of the above-described image generation method embodiment and achieve the same effects and functions, which will not be repeated here.
[0159] In various embodiments of this disclosure, the computer-readable storage medium includes read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk, etc.
[0160] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many methodological improvements today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that a methodological improvement cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program and "integrate" a digital system onto a PLD themselves, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must also be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also understand that by simply performing some logic programming on the method flow using one of these hardware description languages and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.
[0161] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.
[0162] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.
[0163] For ease of description, the above apparatus is described by dividing it into various functional units. Of course, in implementing the embodiments of this disclosure, the functions of each unit can be implemented in one or more software and / or hardware.
[0164] Those skilled in the art will understand that one or more embodiments of this disclosure can be provided as a method, system, or computer program product. Therefore, one or more embodiments of this disclosure can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, one or more embodiments of this disclosure can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0165] This disclosure is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the flowchart illustrations and / or block diagrams. Figure One One or more processes and / or boxes Figure One A device that provides the functions specified in one or more boxes.
[0166] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure One One or more processes and / or boxes Figure One The function specified in one or more boxes.
[0167] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure One One or more processes and / or boxes Figure One The steps of the function specified in one or more boxes.
[0168] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0169] One or more embodiments of this disclosure can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a particular task or implement a particular abstract data type. One or more embodiments of this disclosure can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can reside in local and remote computer storage media, including storage devices.
[0170] The various embodiments in this disclosure are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0171] The above description is merely an embodiment of this disclosure and is not intended to limit the scope of this disclosure. Various modifications and variations can be made to this disclosure by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the scope of the claims of this disclosure.
Claims
1. An image generation method, characterized in that, include: In response to a comment request for a first text in the ebook, a first image matching the first text is obtained; Based on the image content of the first image, obtain the image text that matches the first image; Based on the first image and the image text, generate a second image corresponding to the first text; Publish the comment information corresponding to the first text; The comment information contains the second image; the second image includes an emoji. The image content of the first image includes characters or scenes from the e-book; the step of obtaining image text matching the first image based on its image content includes: Identify at least one text in the e-book that is related to the image content of the first image, and use the identified text as a candidate text that matches the first image; based on the candidate text, determine the image text.
2. The method according to claim 1, characterized in that, The step of obtaining the first image that matches the first text includes: The system searches for a second text in the e-book that is related to the first text; the first text and the second text are used to describe the same content in the e-book. Based on the first text and the second text, generate at least one candidate image that matches the first text, and display the candidate image; In response to the selection instruction for the candidate image, the selected candidate image is determined as the first image.
3. The method according to claim 2, characterized in that, The step of generating at least one candidate image matching the first text based on the first text and the second text includes: Using a generative model, the image content that the candidate image needs to represent is determined based on the first text and the second text; The generative model generates the candidate image based on the image content that the candidate image is required to represent.
4. The method according to claim 2, characterized in that, The step of generating at least one candidate image matching the first text based on the first text and the second text includes: If there exists a comment text to be published corresponding to the comment request, then the generative model is used to determine the image content that the candidate image needs to represent based on the first text, the second text, and the comment text to be published; The generative model is used to determine the emotional atmosphere that the candidate image is required to represent, based on the text of the comment to be published. The generative model generates candidate images based on the image content and emotional atmosphere that the candidate images are intended to represent.
5. The method according to claim 1, characterized in that, Before obtaining the image text matching the first image based on the image content of the first image, the method further includes: Identify the character in the e-book represented by the first image, and define the character as the image content of the first image; or, The scene in the e-book represented by the first image is determined, and the scene is determined as the image content of the first image.
6. The method according to claim 1, characterized in that, Determining at least one text in the e-book related to the image content of the first image includes: If the image content of the first image includes a character from the e-book, then the lines in the e-book related to the character are determined as text related to the image content of the first image; If the image content of the first image includes a scene from the e-book, then sentences in the e-book related to the scene are identified as text related to the image content of the first image.
7. The method according to claim 1, characterized in that, Determining the image text based on the candidate text includes: In response to the selection instruction for the candidate text, the selected candidate text is determined as the image text; or, In response to the editing instruction for the candidate text, the edited candidate text is determined as the image text.
8. An image generation apparatus, characterized in that, include: An image acquisition unit is configured to acquire a first image matching the first text in response to a comment request for the first text in an e-book; The text acquisition unit is used to acquire image text that matches the first image based on the image content of the first image; An image generation unit is configured to generate a second image corresponding to the first text based on the first image and the image text. The comment publishing unit is used to publish comment information corresponding to the first text; The comment information contains the second image; the second image includes an emoji. The image content of the first image includes characters or scenes from the e-book; the text acquisition unit is specifically used to: determine at least one text in the e-book that is related to the image content of the first image, and use the determined text as a candidate text that matches the first image; and determine the image text based on the candidate text.
9. An electronic device, characterized in that, include: processor; as well as, A memory configured to store computer-executable instructions, which, when executed, cause the processor to perform the method described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store computer-executable instructions that, when executed by a processor, implement the method described in any one of claims 1-7.
11. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the method described in any one of claims 1-7.
Citation Information
Patent Citations
Information sending method, device and equipment
CN107612815A
Image generation method and device and related product
CN119991882A