Multimedia content generation method and device, electronic equipment, medium and product

By receiving guidance information, the system automatically identifies the target object and generates multimedia content, solving the problem of low efficiency in generating multimedia content during reading. This enables the efficient generation of multimedia content corresponding to the book's content during reading, thus improving the user experience.

CN121505080APending Publication Date: 2026-02-10DOUYIN VISION CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511602552.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-04
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

When reading e-books, current technology requires users to leave the reading page to generate multimedia content, which interrupts the reading rhythm and the generation process is not efficient.

Method used

By receiving guidance information, the system automatically identifies target objects in books and generates multimedia content, guiding users to efficiently generate multimedia content corresponding to the book content during the reading process. It also utilizes machine learning models to filter target objects and generate multimedia content.

Benefits of technology

While maintaining a sense of immersion in reading, the system efficiently generates multimedia content that matches the book's content, improving the user's reading experience and operational efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121505080A_ABST
    Figure CN121505080A_ABST
Patent Text Reader

Abstract

The invention relates to a multimedia content generation method and device, electronic equipment, a medium and a product, and relates to the field of computers. The multimedia content generation method comprises the steps that guide information is received, the guide information is generated according to a target object and used for guiding a user to generate multimedia content, the target object is from a first text of an electronic book and used for forming a story line in the first text, and the multimedia content is used for representing the guide information; displaying the guide information; and in response to triggering of the guide information, displaying the generated multimedia content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computers, and in particular to a method, apparatus, electronic device, medium, and product for generating multimedia content. Background Technology

[0002] Currently, more and more users are using electronic devices to read books. For example, they read e-books through websites or e-book applications on their mobile phones, computers, and tablets. With the help of electronic devices, users can also perform interactive operations such as saving, annotating, or commenting on content of interest in the book, greatly enhancing the interactivity of the reading process. Summary of the Invention

[0003] According to some embodiments of this disclosure, a method for generating multimedia content is provided, comprising: receiving guidance information, the guidance information being generated based on a target object and used to guide a user to generate multimedia content, the target object being derived from a first text of an e-book and used to constitute a storyline in the first text, the multimedia content being used to characterize the guidance information; displaying the guidance information; and displaying the generated multimedia content in response to triggering the guidance information.

[0004] According to some other embodiments of this disclosure, a multimedia content generation apparatus is provided, comprising: a receiving module configured to receive guidance information generated based on a target object for guiding a user to generate multimedia content, the target object being derived from a first text in an e-book and used to constitute a storyline in the first text, the multimedia content being used to characterize the guidance information; a first display module configured to display the guidance information; and a second display module configured to display the generated multimedia content in response to triggering the guidance information.

[0005] According to some embodiments of the present disclosure, an electronic device is provided, including: a memory; and a processor coupled to the memory, the processor being configured to execute a method for generating multimedia content according to any embodiment of the present disclosure based on instructions stored in the memory.

[0006] According to some embodiments of the present disclosure, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, performs a method for generating multimedia content according to any embodiment of the present disclosure.

[0007] According to some embodiments of this disclosure, a computer program product is provided that, when the computer program product is run on a computer, causes the computer to implement the multimedia content generation method of any embodiment of this disclosure.

[0008] Other features, aspects, and advantages of this disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description

[0009] Embodiments of this disclosure are described below with reference to the accompanying drawings. It should be understood that the drawings described below are merely illustrative of some embodiments of this disclosure and are not intended to limit the scope of this disclosure. In the drawings:

[0010] Figure 1 A flowchart illustrating a method for generating multimedia content according to some embodiments of the present disclosure is shown.

[0011] Figure 2 A flowchart illustrating a method for determining a target object according to some embodiments of the present disclosure is shown.

[0012] Figure 3 A flowchart illustrating a method for determining a target section according to some embodiments of the present disclosure is shown.

[0013] Figure 4 A flowchart illustrating a method for determining guidance information according to some embodiments of the present disclosure is shown.

[0014] Figure 5 A flowchart illustrating a method for generating multimedia content according to some embodiments of the present disclosure is shown.

[0015] Figure 6 A schematic diagram of a reading interface according to some embodiments of the present disclosure is shown.

[0016] Figure 7 A schematic diagram of a reading interface according to other embodiments of the present disclosure is shown.

[0017] Figure 8 A schematic diagram of the structure of a multimedia content generation apparatus according to some embodiments of the present disclosure is shown.

[0018] Figure 9 A block diagram of an electronic device according to some embodiments of the present disclosure is shown.

[0019] Figure 10 Block diagrams of electronic devices according to other embodiments of the present disclosure are shown. Detailed Implementation

[0020] The technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. It should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein.

[0021] It should be understood that the various steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect. Unless otherwise specifically stated, the relative arrangement of components and steps set forth in these embodiments should be interpreted as merely exemplary and does not limit the scope of this disclosure.

[0022] As used in this disclosure, the term "comprising" and its variations are open-ended terms that include at least the following elements / features but do not exclude other elements / features, i.e., "including but not limited to". The term "based on" means "at least partially based on".

[0023] It should be noted that the concepts of "first," "second," etc., used in this disclosure are used only to distinguish different devices, modules, or units, and are not intended to define the order of functions performed by these devices, modules, or units or their interdependencies. Unless otherwise specified, the concepts of "first," "second," etc., are not intended to imply that the objects described herein must be in a given temporal, spatial, rank, or any other given order.

[0024] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0025] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0026] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data shall comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals shall be provided for users to choose to authorize or refuse.

[0027] The embodiments of this disclosure are described in detail below with reference to the accompanying drawings; however, this disclosure is not limited to these specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. Furthermore, in one or more embodiments, specific features, structures, or characteristics can be combined in any suitable manner that will be apparent to those skilled in the art from this disclosure.

[0028] Books typically describe characters, objects, scenes, locations, and plots through text. Readers sometimes try to visualize the scenes depicted in the author's words. However, this can be difficult for some users.

[0029] With the development of artificial intelligence technology, especially the rapid evolution of large-scale modeling, some applications can already generate images from user-inputted text. However, if a user wants to complete this operation while reading, they need to leave the reading page, access an application or website for image generation, input the relevant descriptions from the book and their own needs into the image generation page, and then execute the generation process. This entire process requires complete user intervention, disrupting the user's reading rhythm. Therefore, how to provide users with an efficient way to generate multimedia content without disrupting their reading experience, and ensure that the generated multimedia content corresponds to the reading content, has become an urgent problem to be solved.

[0030] This disclosure provides a method for generating multimedia content. In this method, the system can automatically determine which object in the book will generate multimedia content based on its content, and then guide the user to trigger the generation of the multimedia content, thereby efficiently completing the generation process.

[0031] Figure 1 A flowchart illustrating a method for generating multimedia content according to some embodiments of the present disclosure is shown. Figure 1 As shown, the steps of this embodiment include steps S11 to S13. This method can be executed on the user equipment side, such as a user's client or user terminal.

[0032] In step S11, guidance information is received. The guidance information is generated based on the target object and is used to guide the user to generate multimedia content. The target object comes from the first text of the e-book and is used to form the storyline in the first text. The multimedia content is used to represent the guidance information.

[0033] The first text of an e-book can refer to one or more chapters of the e-book, or it can be a portion of a chapter of the e-book.

[0034] The target object can be considered as a constituent element of the first text, such as characters, objects, locations, plot points, character combinations, etc. That is, the target object can be the plot itself, or an object that has an impact on the plot.

[0035] The guidance information can include information about the target object, such as the name of the target object or a descriptive text of the target object in the first text. Thus, through the guidance information, users can clearly understand that multimedia content is generated based on the target object.

[0036] The guidance information can include guidance text and control information. The guidance text describes the target object, allowing the user to clearly understand which object's multimedia content is to be generated. The controls are used to trigger the multimedia content generation process.

[0037] The guiding information can be a direct quote from the content related to the target object in the initial text, or it can describe the target object in the form of a question. For example, the guiding text could be a question asking the user if they would like to see what the target object looks like.

[0038] In addition, the guidance information can also include more content, such as character images used to guide users, etc., which will be described in more detail later.

[0039] An exemplary process for determining guidance information is as follows. First, the storyline of a first text and the objects involved in that storyline are determined. This process can be implemented using a machine learning model, such as a generative model like a large language model. In some embodiments, the first text can be input into the model along with a first cue message, which instructs the model to identify the target object from the first text. After identifying the target object, the generative model can be used again to determine guidance information. For example, information about the target object (e.g., name, identifier, etc.) can be input into the model along with a second cue message, which instructs the model to generate guidance information. The model can be pre-trained using labeled text, target objects, and corresponding guidance information to make accurate predictions.

[0040] In step S12, guidance information is displayed.

[0041] The guidance information can be displayed at any position within the first text, such as at the end of the first text. For example, if the first text is the chapter the user is currently reading, the guidance information would be displayed at the end of that chapter. The displayed guidance information can include text, as well as images, controls, and so on.

[0042] In step S13, in response to the triggering of the guidance information, the generated multimedia content is displayed.

[0043] In response to user-triggered guidance, multimedia content can be generated locally on the user's device, or the server can perform the generation of multimedia content and return the result to the user's device. After receiving the generated multimedia content, the user can publish it on the network, for example, as a first-text comment, or the user can save it locally.

[0044] An exemplary process involves, in response to a triggering of guidance information, inputting the target object associated with the guidance information, the text associated with the target object in the first text, and a third prompt into a model for generating multimedia content. The third prompt can instruct the model to generate multimedia content for the target object based on the input information, such as generating an image or video of the target object. The information and text of the target object associated with the guidance information can be stored on the user device side or on the server side, and can be obtained by searching the storage unit using the identifier of the guidance information.

[0045] The above embodiments can determine, based on the first text, which object in the book to generate multimedia content for during the user's reading process, and guide the user to execute the generation process. This allows for the efficient generation of visual content matching key objects in the book while maintaining the user's reading immersion, thus improving the user's reading experience.

[0046] Since the first text may involve multiple objects, these objects can be filtered to determine the target object. That is, from the multiple objects involved in the first text, the objects that the user is more likely to want to visualize are filtered out. In some embodiments, the target object is determined from one or more candidate objects based on at least one of the priority of the dimension to which each candidate object belongs and the dimensions involved in the interactive information of the first text; the candidate objects are determined from the multiple objects involved in the first text based on the content of the objects involved in at least one of the first text and the second text of the e-book, wherein the second text includes the first text and is different from the first text, and the multiple objects belong to one or more dimensions.

[0047] Figure 2 A flowchart illustrating a method for determining a target object according to some embodiments of the present disclosure is shown. Figure 2 As shown, the method of this embodiment includes steps S21 to S22.

[0048] In step S21, candidate objects are determined from multiple objects involved in the first text based on the content of objects involved in at least one of the first text and the second text of the e-book.

[0049] The second text includes the first text, but covers a broader range of content. For example, the second text may include content that precedes or follows the first text, or both.

[0050] One or more dimensions of an object can be considered, such as one or more object types, like people, items, locations, plots, character relationships, etc. Objects with specified dimensions can be identified from the first text using named entity recognition, or by processing the first text using a natural language model to determine objects with one or more dimensions.

[0051] Several methods for identifying candidate objects are described below as examples.

[0052] In some embodiments, candidate objects are determined from multiple objects based on at least one of the following: the amount of content involved in each object in the first text and the degree of association between each object and a first role. The first role is the main role in the first text. The amount of content involved in an object can be determined based on the content in the text that relates to that object. For example, it could be the number of words in the text, or it could be determined by analyzing the text using a natural language processing model and providing an evaluation result of the amount of content. By filtering candidate objects based on the amount of content, objects that occupy a relatively large proportion of the first text can be selected.

[0053] The primary character can also be determined based on the amount of content the character is involved in within the first text. Alternatively, the protagonist of the book can be directly used as the primary character, and this can be determined from the book's synopsis. In this way, objects with a stronger connection to the main character can be identified as the target objects.

[0054] Of course, the two methods mentioned above can also be combined. For example, a weighted average can be calculated based on the content volume score and the score related to the primary character, and then the target object can be determined comprehensively from the two aspects mentioned above based on the weighted result. Of course, those skilled in the art can also use other methods to determine the target object, which will not be elaborated here.

[0055] By using the above method, more important objects can be objectively selected as candidate objects based on the descriptions of each object in the text.

[0056] In some embodiments, the candidate object is the object that first appears in the first text. Let the second text include the first text and the text preceding the first text in the book. By determining the position of the candidate object in the second text, it can be determined whether a candidate object first appears in the first text. For example, the name of the candidate object can be searched in the second text, and its position in the text can be marked. Since the descriptive information of an object appearing for the first time is relatively limited, using the object appearing for the first time as a candidate object increases the probability of triggering guiding information. Furthermore, it helps users understand the visual information of the object appearing for the first time, thereby improving the user's reading fluency.

[0057] In some embodiments, candidate objects are those that influence the content following the first text, and are determined based on the amount of content each object involves in the content following the first text. The first text and the content following it can be considered as the second text. The names of each object can be searched within the second text, and the content associated with that name can be determined. The amount of content each object involves in the second text can be determined based on word count or a text processing model. For example, if an object occupies a significant portion of the second text—that is, its content size ranking is greater than a specified value, or its content size exceeds a specified threshold—it indicates that the object will still have a significant impact on the plot after the first text. By using such objects as candidate objects, users can visualize subsequent important objects during reading, thereby improving reading fluency.

[0058] In step S22, the target object is determined from one or more candidate objects based on at least one of the priority of the dimension to which each candidate object belongs and the dimensions involved in the interaction information of the first text.

[0059] Different priorities can be pre-set; for example, the priority of the character dimension can be set higher than that of the item dimension. This priority can be determined based on the type of book or the proportion of each dimension in the generated multimedia content. Therefore, when the initial text involves multiple candidate objects across different dimensions, the user can be guided to generate multimedia content for objects with higher priority dimensions.

[0060] Interactive information in the first text refers to reader feedback on the content. This includes comments, tags, likes, and so on. For example, semantic analysis of comments can be performed to determine which dimensions readers want to visualize. If many readers express a desire to see comments about certain characters in the text, then the character dimension can be given higher priority. Furthermore, the dimensions involved in multimedia content posted in the comments can be statistically analyzed, and the priority of each dimension can be determined based on the proportion of multimedia content in each dimension.

[0061] In some embodiments, the target object belongs to the target dimension, which is determined based on the dimensions involved in the comments content of the first text and the number of interactions involved in each dimension within the comments content. That is, the dimensions involved in the comments content of the first text can be determined in advance using text processing or multimedia content recognition models. Then, based on the number of comments, likes, replies, and other interactions involved in each dimension, the most popular dimension is further determined as the target dimension. The most popular dimension is, for example, the dimension with the highest number of interactions, or the dimension with the highest weighted sum of the number of interactions of various types. Thus, after determining the candidate objects that meet the requirements based on the content level of the book, the target object can be further determined based on the needs of the readers.

[0062] The above embodiments can increase the probability of users triggering the multimedia content generation function, improve the availability of guidance information, and also improve the user experience and the efficiency of users generating multimedia content.

[0063] Some books have numerous chapters, and certain chapters may be related. These chapters may appear consecutively or intermittently. Clustering can be used to group related chapters into a single category, and further, it can be determined which chapter guides the user to generate multimedia content, and which target object the guidance information in that chapter is associated with. In some embodiments, the first text is the target chapter, and the target object is a candidate object belonging to a specified dimension within the target chapter. The target chapter is determined from multiple chapters based on at least one of the specified dimension and the state of the target object across multiple chapters, with multiple chapters belonging to the same category. Multiple chapters belonging to the same category are obtained through clustering, which is performed based on objects of a specified dimension in each chapter of the e-book. Candidate objects are determined based on the amount of content involved in the objects of the specified dimension across multiple chapters.

[0064] Figure 3 A flowchart illustrating a method for determining a target section according to some embodiments of this disclosure is shown. Figure 3 As shown, the method of this embodiment includes steps S31 to S33.

[0065] In step S31, the chapters of the e-book are clustered to obtain multiple categories, each category including multiple chapters. The clustering is performed based on objects of a specified dimension in each chapter.

[0066] The specified dimension can be any one of multiple dimensions. That is, the clustering can be performed multiple times based on multiple dimensions.

[0067] For example, a chapter can be represented by information about objects within a specified dimension. One example is constructing a vector to represent a chapter, where each element represents information about an object, such as its frequency of appearance or the length of the book. For instance, when representing chapters by characters, the vector can have N (N is a positive integer) elements pre-defined to represent N characters in the book, and then the values ​​of each element in the vector can be determined based on the information about each character in a given chapter. After determining the representation vectors for each chapter, these vectors can be clustered. Thus, in the clustering results, chapters belonging to the same cluster are those that are similar in the specified dimension.

[0068] In step S32, for one or more chapters within any category, the objects involved in the chapters of that category are determined as candidate objects based on the amount of content involved in the specified dimension within those chapters. For example, the object occupying the largest amount of content in that category can be identified as a candidate object.

[0069] In step S33, for multiple chapters in each category, based on the specified dimension and at least one of the states of the candidate object in the multiple chapters of that category, a target chapter is determined from the multiple chapters, and the candidate object is determined as the target object of the target chapter. That is, the target chapter is designated as the first text.

[0070] Objects in different dimensions have different characteristics. Therefore, when determining the target chapter, we can decide which chapter to use as the target chapter based on the type of the specified dimension or the state of the object. In other words, we can determine which chapter of that category can best highlight the importance of the object in that dimension.

[0071] The state of a candidate object can be determined based on its appearance description in the book. For example, the same location may have different appearance characteristics in different seasons or eras, which can be considered different states. For the character dimension, the state can also be determined based on the character's different identities. For the item dimension, the state can be determined based on changes in the item's function. For example, a treasure may have a decorative function in some chapters, while possessing magical functions in others. When determining the state of an object, the chapter text, the object's name, and information from the specified dimensions can be input into a text processing model to obtain the output of the text processing model, which represents the object's state in each chapter. In some embodiments, in response to any one of the specified dimensions being a single character, location, or item, the target chapter is the chapter where the candidate object first appears or its state changes; in response to at least one of the specified dimensions being a character group or plot, the target chapter is the last chapter belonging to the category. That is, for single characters, locations, items, etc., when they first appear or their state changes, multimedia content can be generated to help users quickly present their visual content. For groups of characters who may have had a specific story, multimedia content is generated in the last chapter of that category. This avoids spoilers and helps users understand the relationships between the characters.

[0072] This embodiment allows for the identification of target chapters within related chapters that play a crucial role in that dimension from various perspectives, and the placement of guidance information within those chapters. This enables timely and efficient visual guidance for users seeking content of interest.

[0073] In guiding users to generate multimedia content for a target object, other characters in the book can be used as "introducers." In some embodiments, receiving guidance information includes receiving guidance information comprising a first virtual object and first spoken text, wherein the first virtual object represents a guiding role, the guiding role is associated with the target object and is determined based on the text associated with the target object in the first text, and the first spoken text is the spoken text of the guiding role and is determined based on the target object. For example, if a sword appears for the first time in a chapter, and the chapter describes the sword's appearance in great detail, then the character who discovered the sword in that chapter can be used as the guiding role, and the guiding text could include, "Would you like to see what this sword looks like?"

[0074] Figure 4 A flowchart illustrating a method for determining guidance information according to some embodiments of this disclosure is shown. Figure 4 As shown, the determination method in this embodiment includes steps S41 to S44.

[0075] In step S41, a guiding role is determined based on the text in the first text that is associated with the target object.

[0076] The text associated with the target object can be a paragraph containing that object, or it can include several adjacent paragraphs. Of course, the associated text can also be text containing a certain number of words related to the target object, and so on. The characters appearing in this associated text can then serve as guiding characters. If multiple characters appear, the character with the highest frequency in the text can be identified.

[0077] In some embodiments, the text may also be input into a text processing model, and prompts may be input into the model to instruct the model to determine the role with the greatest relevance to the target object.

[0078] In step S42, the first spoken text for the guiding role is determined based on the guiding role and the target object. That is, the first spoken text is "spoken" by the guiding role and includes the target object.

[0079] In some embodiments, the first spoken text is generated based on the facilitator's language style and the target audience, with the language style determined according to the description of the facilitator in the e-book. Thus, even if the first spoken text is not the original text from the e-book, it can still match the facilitator's style. Language style can include tone of voice, filler words, etc. The language style can be determined by classifying the character's spoken text in the first book.

[0080] For example, for a lively character, the first spoken text could be, "I've never seen such a fun treasure before. Do you want to see it? But don't try to take it from me!"; while for a careful and gentle character, the first spoken text could be, "Let's take a look at this sword together. Be careful not to cut your hand; it's very sharp."

[0081] In some embodiments, the first spoken text is extracted from a first text and relates to the content of the target object. That is, the original text from the e-book can be directly adopted as guiding information. For example, when the target object is a plot, the first spoken text could be the speech content of an important character who drives the plot. This helps maintain the user's immersion.

[0082] In step S43, a first virtual object representing the guiding role is obtained. The first virtual object can be multimedia content representing the guiding role, such as an image, or an object that can be interacted with.

[0083] In some embodiments, the first virtual object is generated based on a base virtual avatar and a first text description of the guiding character. The base virtual avatar is pre-generated. For example, one or more base virtual avatars for each character can be pre-set, and each different avatar can correspond to multiple stages in the book, such as a character's youth, middle age, old age, etc. Then, during the generation of the first virtual object, a first virtual object that better matches the first text is generated based on the first text and the base virtual avatar. Thus, the base virtual avatar can constrain the generated virtual object to a certain extent, avoiding the generation of low-quality images and improving the efficiency of image generation. This process can be implemented using a model for generating images based on text and images.

[0084] In step S44, guidance information is generated based on the first virtual object and the first spoken text. This generated guidance information can guide the user through both character and text methods, increasing the probability of triggering the multimedia content generation function.

[0085] For dimensions such as characters, locations, objects, and plots, the generated multimedia content can allow users to describe the appearance of target objects in these dimensions. For groups of characters, in addition to generating multimedia content representing their appearance, it can also generate multimedia content representing their relationships, assisting users in understanding complex character relationships.

[0086] In some embodiments, the target object is a group of characters, including multiple characters from an e-book. The guidance information includes a second virtual object and guidance text. The second virtual object represents the image of the character group, and the guidance text guides the user to generate relationships between the multiple characters. For example, if the target object is a group of eight characters, the second virtual object could be a character image including these eight characters, and the guidance information could be, "We've finished our introductions. Do you understand the relationships between us now?"

[0087] To avoid prematurely revealing characters who haven't yet appeared, thus affecting the user's reading experience, one embodiment includes displaying guidance information on the end page of the first text in response to its display. This allows users to be guided to generate multimedia content only after the entire group of characters has appeared. Currently, guidance information can also be displayed on the end page of the first text for target characters in other dimensions.

[0088] Multimedia content can be generated based on text. For example, a prompt message can be generated and input into a model used to generate an image, resulting in multimedia content output by the model. Figure 5 A flowchart illustrating a method for generating multimedia content according to some embodiments of the present disclosure is shown. Figure 5 As shown, the generation method of this embodiment includes steps S51 to S53.

[0089] In step S51, in response to the triggering of the guidance information, a prompt message is generated based on the text in the e-book that involves the target object.

[0090] In some embodiments, in response to the target object appearing for the first time in the first text, a prompt message is generated based on the text in the first text that relates to the target object; or, in response to the target object not appearing for the first time in the first text, a summary of the content relating to the target object in the content preceding the first text is generated, and a prompt message is generated based on the text in the first text that relates to the target object and the summary.

[0091] In other words, for a target character appearing for the first time in the first text, the basis for generating multimedia content can be obtained solely from the first text. However, for characters that do not appear for the first time, while using the content in the first text as the basis for generation, the summary of previous texts can also be referenced, thereby avoiding the generated multimedia content matching the first text but not conforming to the descriptions in other parts of the e-book.

[0092] When generating prompts, the text following the first text can be omitted. This avoids prematurely revealing events that haven't occurred to the user, thus preserving the reading experience.

[0093] In step S52, multimedia content for the target object is generated based on the prompt information.

[0094] Multimedia content can be "generated with a single click" in response to user-triggered guidance information. That is, after the user triggers the guidance information, the generated multimedia content can be displayed without requiring any additional user action or displaying the generation basis to the user, which offers high operational efficiency.

[0095] Alternatively, a prompt message can be displayed before generating multimedia content so that users understand the basis for generating the multimedia content and can make modifications if necessary.

[0096] Therefore, the prompt message can be displayed to the user and allowed to edit it, or it can be hidden from the user.

[0097] In step S53, multimedia content is displayed.

[0098] Therefore, the text and image processing capabilities of generative models can be used to generate multimedia content that matches the books, thus improving the user's reading experience.

[0099] Figure 6 and Figure 7 Two schematic diagrams of reading interfaces are shown as examples, and the reading interfaces include two different forms of guidance information.

[0100] exist Figure 6 In the middle, interface 6 includes the book's text 61 and guidance information 62. Guidance information 62 includes the first virtual image of the guide character 621, guidance text 622, and control 623. The guidance text 622 is "This sword is amazing, want to see what it looks like?" In response to the triggering of control 623, a detailed image of the sword can be generated based on the description of the sword in the current chapter.

[0101] exist Figure 7 In the interface 7, the text 71 of the book and the guidance information 72 are included. The guidance information 72 includes a description of the target object 721, "Highlight Scene," guidance text 722, and a control 723. The guidance text 722 is a passage from the book: "The sky has finally cleared, but the world in XX's heart is still raining." In response to the triggering of the control 723, an image of the character XX in a clear sky scene can be generated based on the description of the highlight scene in the current chapter, and the character's emotions can be expressed through facial expressions.

[0102] The methods of various embodiments of this disclosure have been described above. The apparatus for implementing the methods of each embodiment is described below.

[0103] Figure 8 A schematic diagram of the structure of a multimedia content generation apparatus according to some embodiments of the present disclosure is shown. For example... Figure 8 As shown, the generation device 8 in this embodiment includes: a receiving module 81 configured to receive guidance information, the guidance information being generated based on a target object and used to guide the user to generate multimedia content, the target object being a first text from an e-book and used to constitute a storyline in the first text, and the multimedia content being used to represent the guidance information; a first display module 82 configured to display the guidance information; and a second display module 83 configured to display the generated multimedia content in response to the triggering of the guidance information.

[0104] In some embodiments, the target object is determined from one or more candidate objects based on at least one of the priority of the dimension to which each candidate object belongs and the dimensions involved in the interaction information of the first text; the candidate object is determined from multiple objects involved in the first text based on the content of the object involved in at least one of the first text and the second text of the e-book, wherein the second text includes the first text and is different from the first text, and the multiple objects belong to one or more dimensions.

[0105] In some embodiments, candidate objects are determined from a plurality of objects based on at least one of the amount of content involved in each object in the first text and the degree of association between each object and the first role, wherein the first role is the main role in the first text.

[0106] In some embodiments, the candidate object is the object that first appears in the first text.

[0107] In some embodiments, candidate objects are objects that affect the content following the first text, and are determined based on the amount of content each object is involved in within the content following the first text.

[0108] In some embodiments, the target object belongs to the target dimension, which is determined based on the dimensions involved in the comment content of the first text and the number of interactions involved in each dimension in the comment content.

[0109] In some embodiments, the first text is the target chapter, and the target object is a candidate object belonging to a specified dimension in the target chapter; the target chapter is determined from multiple chapters based on at least one of the specified dimension and the state of the target object in multiple chapters, and the multiple chapters belong to the same category; the multiple chapters belonging to the same category are obtained by clustering, and the clustering is performed based on the objects of the specified dimension in each chapter of the e-book; the candidate object is determined based on the amount of content involved by the objects of the specified dimension in multiple chapters.

[0110] In some embodiments, in response to any one of the specified dimensions being a single character, location, or item, the target chapter is the chapter where the candidate object first appears or its state changes; in response to at least one of the specified dimensions being a character group or plot, the target chapter is the last chapter among the chapters belonging to the category.

[0111] In some embodiments, the receiving module 81 is further configured to receive guidance information including a first virtual object and a first spoken text, wherein the first virtual object is used to represent a guidance role, the guidance role is associated with a target object and is determined based on the text associated with the target object in the first text, and the first spoken text is the spoken text of the guidance role and is determined based on the target object.

[0112] In some embodiments, the first spoken text is generated based on the facilitator's language style and the target object, with the language style determined according to the description of the facilitator in the e-book; or the first spoken text is extracted from a first text and involves content related to the target object.

[0113] In some embodiments, the first virtual object is generated based on a base virtual avatar and a first text description of the guiding character, the base virtual avatar being pre-generated.

[0114] In some embodiments, the target object is a group of characters, which includes multiple characters from an e-book. The guidance information includes a second virtual object and guidance text. The second virtual object is used to represent the image of the group of characters, and the guidance text is used to guide the user to generate relationships between the multiple characters.

[0115] In some embodiments, the first display module 82 is further configured to display guidance information on the end page in response to the end page of the first text being displayed.

[0116] In some embodiments, the second display module 83 is further configured to: generate prompt information based on the text of the target object in the e-book in response to the triggering of the guidance information; generate multimedia content of the target object according to the prompt information; and display the multimedia content.

[0117] In some embodiments, the second display module 83 is further configured to: generate a prompt message based on the text in the first text that relates to the target object when the target object first appears in the first text; or, generate a summary of the content related to the target object in the content preceding the first text when the target object does not first appear in the first text, and generate a prompt message based on the text in the first text that relates to the target object and the summary.

[0118] According to some embodiments of the present disclosure, an electronic device is provided, including: a memory; and a processor coupled to the memory, the processor being configured to execute a method for generating multimedia content according to any embodiment of the present disclosure based on instructions stored in the memory.

[0119] Figure 9 A block diagram of an electronic device according to some embodiments of the present disclosure is shown.

[0120] Memory 91 is used to store one or more computer-readable instructions. Memory 91 may include any combination of various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory, including but not limited to random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), read-only memory (ROM), and flash memory. Memory 91 may, for example, store operating systems, application programs, bootloaders, databases, and other programs, as well as various application programs and various data.

[0121] The processor 92 is configured to execute computer-readable instructions to implement the method described in any of the foregoing embodiments. Specific implementations of each step of the method can be found in the above embodiments; repeated details will not be elaborated upon here.

[0122] The processor 92 can be configured to perform the steps of the foregoing embodiments. The processor 92 can be embodied in various processing devices, such as a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The central processing unit (CPU) can be an x86 or ARM architecture, etc.

[0123] The processor 92 and the memory 91 can communicate with each other directly or indirectly. For example, the processor 92 and the memory 91 can communicate via a network. The network can include wireless networks, wired networks, and / or any combination of wireless and wired networks. The processor 92 and the memory 91 can also communicate with each other via a system bus, which is not limited in this disclosure.

[0124] It should be noted that Figure 9 The components of the electronic device 9 shown are merely exemplary and not limiting. The electronic device 9 may have other components depending on the specific application requirements. The processor 92 can control other components in the electronic device 9 to perform desired functions.

[0125] Electronic device 9 can be implemented by software, firmware and / or hardware, and can be integrated into a device with the relevant application installed.

[0126] Figure 10 Block diagrams of electronic devices according to other embodiments of the present disclosure are shown.

[0127] Figure 10 The electronic device 10 shown can be a computer system with a dedicated hardware structure, which can perform corresponding functions when the relevant application is installed.

[0128] Electronic devices include, but are not limited to, mobile terminals such as smartphones, laptops, personal digital assistants (PDAs), tablet computers (PCs), PMPs (portable multimedia players), in-vehicle terminals (such as in-vehicle navigation terminals), wearable devices, and fixed terminals such as digital televisions and desktop computers.

[0129] like Figure 10As shown, the Central Processing Unit (CPU) 101 performs various processes based on programs stored in the Read-Only Memory (ROM) 102 or programs loaded from the storage section 108 into the Random Access Memory (RAM) 103. The RAM 103 stores data required as needed when the CPU 101 performs various processes, etc. The CPU is merely exemplary; it could also be other types of processors, such as the various processors described above. The ROM 102, RAM 103, and storage section 108 can be various forms of computer-readable storage media. It should be noted that although... Figure 10 The image shows ROM 102, RAM 103 and storage section 108, but one or more of them may be combined or located in the same or different memory or storage modules.

[0130] CPU 101, ROM 102 and RAM 103 are interconnected via bus 104. Input / output interface 105 is also connected to bus 104.

[0131] The following components are connected to the input / output interface 105: input section 106, such as a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output section 107, including displays such as cathode ray tube (CRT), liquid crystal display (LCD), speakers, vibrators, etc.; storage section 108, including hard disks, magnetic tapes, etc.; and communication section 109, including network interface cards such as LAN cards, modems, etc. Communication section 109 allows communication processing via a network such as the Internet. It is easy to understand that, although... Figure 10 The portion of the electronic device 10 shown communicates via bus 104, but it may also communicate via a network or other means, wherein the network may include a wireless network, a wired network, and / or any combination of wireless and wired networks.

[0132] As needed, drive 1010 is also connected to input / output interface 105. Removable media 1011, such as disks, optical disks, magneto-optical disks, semiconductor memories, etc., are installed on drive 1010 as needed, so that computer programs read from them can be installed into storage section 108 as needed.

[0133] When the above series of processes are implemented through software, the program constituting the software can be installed from a network such as the Internet or a storage medium such as a removable medium 1011.

[0134] According to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this disclosure include a computer program product that, when run on a computer, causes the computer to perform the methods described in any of the foregoing embodiments. The computer program product includes computer instructions carried on a computer-readable medium, containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer instructions can be downloaded and installed from a network via communication section 109, or installed from storage section 108, or installed from ROM 102. When the computer program is executed by CPU 101, the methods of the embodiments of this disclosure are performed.

[0135] According to some embodiments of the present disclosure, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, performs a method for generating multimedia content according to any embodiment of the present disclosure.

[0136] According to some embodiments of this disclosure, a computer program product is provided that, when the computer program product is run on a computer, causes the computer to implement the multimedia content generation method of any embodiment of this disclosure.

[0137] It should be noted that, in the context of this disclosure, a computer-readable medium can be a tangible medium that may contain or store programs for use by or in conjunction with an instruction execution system, apparatus, or device.

[0138] A computer-readable medium may be a computer-readable storage medium, a computer-readable signal medium, or any combination thereof.

[0139] Computer-readable storage media include, but are not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. Computer instructions are stored on the computer-readable storage medium that, when executed by a processor, implement the methods described in any of the foregoing embodiments.

[0140] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0141] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0142] In some embodiments, a computer program is also provided, comprising: instructions that, when executed by a processor, cause the processor to perform the methods described in any of the foregoing embodiments. For example, the instructions may be embodied in computer program code.

[0143] In embodiments of this disclosure, computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof. These programming languages ​​include, but are not limited to, object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network (including a local area network (LAN) or a wide area network (WAN)), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0144] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0145] The functions described above can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary hardware logic components that can be used include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.

[0146] While specific embodiments of this disclosure have been described in detail by way of example, those skilled in the art should understand that the examples are for illustrative purposes only and not intended to limit the scope of this disclosure. Those skilled in the art should understand that modifications can be made to the above embodiments without departing from the scope and spirit of this disclosure. The scope of this disclosure is defined by the appended claims.

Claims

1. A method for generating multimedia content, comprising: The system receives guidance information, which is generated based on a target object and used to guide the user to generate multimedia content. The target object comes from the first text of an e-book and is used to form the storyline in the first text. The multimedia content is used to represent the guidance information. Display the aforementioned guidance information; In response to the triggering of the guidance information, the generated multimedia content is displayed.

2. The generation method according to claim 1, wherein: The target object is determined from one or more candidate objects based on at least one of the following: the priority of the dimension to which each candidate object belongs, and the dimensions involved in the interaction information of the first text. The candidate objects are determined from a plurality of objects mentioned in the first text based on the content of the object in at least one of the first text and the second text of the e-book, wherein the second text includes the first text and is different from the first text, and the plurality of objects belong to one or more dimensions.

3. The generation method according to claim 2, wherein, The candidate objects are determined from the plurality of objects based on at least one of the following: the amount of content involved in each object in the first text and the degree of association between each object and the first role, wherein the first role is the main role in the first text.

4. The generation method according to claim 2, wherein, The candidate object is the object that appears for the first time in the first text.

5. The generation method according to claim 2, wherein, The candidate objects are those that affect the content following the first text, and are determined based on the amount of content each object is involved in within the content following the first text.

6. The generation method according to claim 2, wherein, The target object belongs to the target dimension, which is determined based on the dimensions involved in the comment content of the first text and the number of interactions involved in the comment content for each dimension.

7. The generation method according to claim 1, wherein: The first text is the target chapter, and the target object is a candidate object belonging to a specified dimension within the target chapter; The target chapter is determined from the multiple chapters based on at least one of the specified dimension and the state of the target object in multiple chapters, wherein the multiple chapters belong to the same category; The multiple chapters belonging to the same category are obtained by clustering, and the clustering is performed based on objects of a specified dimension in each chapter of the e-book; The candidate objects are determined based on the amount of content involved in the objects of the specified dimensions in the multiple chapters.

8. The generation method according to claim 7, wherein: In response to any one of the specified dimensions being a single character, location, or item, the target chapter is the chapter in which the candidate object first appears or its status changes; In response to the specified dimension being at least one of character group or plot, the target chapter is the last chapter among the chapters belonging to the category.

9. The generation method according to claim 1, wherein, The received guidance information includes: The system receives guidance information including the first virtual object and the first spoken text, wherein the first virtual object is used to represent a guidance role, the guidance role is associated with the target object and is determined based on the text associated with the target object in the first text, and the first spoken text is the spoken text of the guidance role and is determined based on the target object.

10. The generation method according to claim 9, wherein: The first spoken text is generated based on the facilitator's language style and the target object, wherein the language style is determined according to the description of the facilitator in the e-book; or The first spoken text is extracted from the first text and relates to the content of the target object.

11. The generation method according to claim 9, wherein, The first virtual object is generated based on the basic virtual image and the first text describing the guide character, and the basic virtual image is pre-generated.

12. The generation method according to claim 1, wherein, The target object is a group of characters, which includes multiple characters from the e-book. The guidance information includes a second virtual object and guidance text. The second virtual object is used to represent the image of the group of characters, and the guidance text is used to guide the user to generate relationships between the multiple characters.

13. The generation method according to any one of claims 1 to 12, wherein, The display of the guidance information includes: In response to displaying the end page of the first text, the guidance information is displayed on the end page.

14. The generation method according to any one of claims 1 to 12, wherein, The multimedia content displayed in response to the triggering of the guidance information includes: In response to the triggering of the guidance information, a prompt message is generated based on the text in the e-book that relates to the target object; Based on the prompt information, generate multimedia content for the target object; Display the multimedia content.

15. The generation method according to claim 14, wherein, The step of generating prompt information based on the text in the e-book that relates to the target object includes: In response to the target object's first appearance in the first text, the prompt message is generated based on the text in the first text that relates to the target object; or, In response to the fact that the target object does not appear for the first time in the first text, a summary of the content involving the target object in the content preceding the first text is generated, and the prompt information is generated based on the text involving the target object in the first text and the summary.

16. A multimedia content generation apparatus, comprising: The receiving module is configured to receive guidance information, which is generated based on a target object and used to guide the user to generate multimedia content. The target object comes from the first text of the e-book and is used to form the storyline in the first text. The multimedia content is used to represent the guidance information. The first display module is configured to display the guidance information; The second display module is configured to display the generated multimedia content in response to the triggering of the guidance information.

17. An electronic device comprising: Memory; as well as A processor coupled to the memory, the processor being configured to perform a method for generating multimedia content as described in any one of claims 1 to 15, based on instructions stored in the memory.

18. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for generating multimedia content according to any one of claims 1 to 15.

19. A computer program product, when run on a computer, causes the computer to implement the method for generating multimedia content according to any one of claims 1 to 15.

Citation Information

Patent Citations

  • Data processing method and device, electronic equipment and storage medium

    CN112819933A

  • Content display method and device, electronic equipment and storage medium

    CN116975330A

  • Book information processing method and device, equipment and storage medium

    CN117314585A

  • Image generation method and device and related product

    CN119991882A

  • Storyline visualization

    US20130157234A1