Content processing method and device, electronic equipment, medium and program product
By using machine learning models to perform semantic understanding of e-book content from multiple dimensions, key segments are identified and matched with media content to generate a media content set. This solves the problems of scattered and inconsistent media content in e-books and improves the user's reading experience.
Patent Information
- Application Number
- CN202511384417.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-25
- Publication Date
- 2026-01-09
AI Technical Summary
The distribution of user-created or AI-generated media content in existing e-books is scattered and of varying quality, making it inconvenient for users to find and affecting the reading experience.
By using machine learning models to perform semantic understanding of e-book content from multiple dimensions, key segments are identified and matched with corresponding media content to generate media content sets in various dimensions, filtering out low-quality content and providing users with convenient viewing.
It has improved the utilization efficiency of e-book media content, enriched reading methods, and enhanced the user's reading experience.
Smart Images

Figure CN121301593A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to a content processing method, apparatus, electronic device, medium, and program product. Background Technology
[0002] The emergence of e-books has made reading more convenient for users. Traditional e-books are usually presented in plain text format, but as users' reading needs have increased, images, videos, and other media content have been created for specific passages in e-books to help users understand the text and improve the reading experience. Summary of the Invention
[0003] According to some embodiments of this disclosure, a content processing method is provided, comprising: performing semantic understanding on the content of an e-book from one or more dimensions to determine key segments corresponding to each dimension; determining media content matching each key segment from multiple media contents corresponding to the e-book, wherein each media content includes at least one of images and videos, and each media content is generated based on one or more segments in the e-book; and generating a media content set for each dimension based on the key segments corresponding to each dimension and the media content matching each key segment.
[0004] According to other embodiments of this disclosure, a content processing apparatus is provided, comprising: a first determining module configured to perform semantic understanding on the content of an e-book from one or more dimensions, and determine key segments corresponding to each dimension; a second determining module configured to determine media content matching each key segment from multiple media contents corresponding to the e-book, based on the content of each key segment, wherein each media content includes at least one of images and videos, and each media content is generated based on one or more segments in the e-book; and a generating module configured to generate a media content set for each dimension based on the key segments corresponding to each dimension and the media content matching each key segment.
[0005] According to further embodiments of the present disclosure, an electronic device is provided, including: a processor; and a memory coupled to the processor for storing instructions that, when executed by the processor, cause the processor to perform a content processing method as described in any embodiment of the present disclosure.
[0006] According to further embodiments of the present disclosure, a computer-readable storage medium is provided having a computer program stored thereon, wherein, when executed by a processor, the program causes the processor to perform a content processing method according to any embodiment of the present disclosure.
[0007] According to further embodiments of the present disclosure, a computer program product is provided, comprising: instructions, wherein when executed by a processor, the instructions cause the processor to perform a content processing method according to any embodiment of the present disclosure.
[0008] Other features, aspects, and advantages of this disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description
[0009] Embodiments of this disclosure are described below with reference to the accompanying drawings. It should be understood that the drawings described below are merely illustrative of some embodiments of this disclosure and are not intended to limit the scope of this disclosure. In the drawings:
[0010] Figure 1 A flowchart illustrating content processing methods according to some embodiments of this disclosure is shown;
[0011] Figures 2-6 A schematic diagram illustrating the displayed interface of some embodiments of this disclosure;
[0012] Figure 7 This invention discloses a schematic diagram of the structure of a content processing apparatus according to some embodiments of the present disclosure;
[0013] Figure 8 This invention discloses schematic diagrams of the structure of electronic devices according to some embodiments of the present disclosure;
[0014] Figure 9 A schematic diagram of the structure of an electronic device according to other embodiments of the present disclosure is shown. Detailed Implementation
[0015] The technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. It should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein.
[0016] It should be understood that the various steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect. Unless otherwise specifically stated, the relative arrangement of steps set forth in these embodiments should be interpreted as merely exemplary and does not limit the scope of this disclosure.
[0017] As used in this disclosure, the term "comprising" and its variations are open-ended terms that include at least the following elements / features but do not exclude other elements / features, i.e., "including but not limited to". The term "based on" means "at least partially based on".
[0018] It should be noted that the concepts of "first," "second," etc., used in this disclosure are used only to distinguish different devices, modules, or units, and are not intended to define the order of functions performed by these devices, modules, or units or their interdependencies. Unless otherwise specified, the concepts of "first," "second," etc., are not intended to imply that the objects described herein must be in a given temporal, spatial, rank, or any other given order.
[0019] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0020] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0021] The embodiments of this disclosure are described in detail below with reference to the accompanying drawings; however, this disclosure is not limited to these specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. Furthermore, in one or more embodiments, specific features, structures, or characteristics can be combined in any suitable manner that will be apparent to those skilled in the art from this disclosure.
[0022] Media content such as images and videos generated from certain excerpts of e-books can be user-generated, AI-generated, or author-created. The distribution of this media content is relatively scattered; for example, user-generated and AI-generated content is posted in the comment section, while author-created content may appear within the e-book. This makes it inconvenient for users to find specific media content, and some user-generated or AI-generated content does not match the excerpts in the e-book or is of low quality. Therefore, how to effectively utilize these media resources to improve the user's reading experience is a problem that needs to be addressed.
[0023] To address the aforementioned issues, this disclosure proposes a content processing method that can automatically perform semantic understanding of e-book content from one or more dimensions, identify key segments corresponding to each dimension, and match multiple media contents corresponding to the e-book with the content of each key segment to determine the matching media content for each key segment. Then, for each dimension, a media content set for that dimension is generated based on the key segments corresponding to that dimension and the matching media content for each key segment. Through the filtering of key segments and the matching of key segments with media content, some low-quality media content can be filtered out, while more suitable key segments and media content that assist users in understanding the e-book content and that users are more concerned with are selected. The media content set for each dimension can be viewed by users, allowing them to quickly view the media content of each key segment without having to search through different distribution locations. When multiple dimensions of media content sets are generated, users can also conveniently view media content sets under different dimensions. This content processing method effectively utilizes the media content corresponding to e-books, facilitates user viewing of media content, assists users in understanding the e-book content, enriches users' reading methods, and enhances the user's reading experience.
[0024] The following is combined Figures 1-6 This disclosure describes a content processing method. The content processing method can be executed by a content processing apparatus or an electronic device, and the content processing apparatus can be implemented by software or a combination of software and hardware.
[0025] Figure 1 Flowcharts illustrating some embodiments of the content processing method disclosed herein. For example... Figure 1 As shown, the content processing method of this embodiment includes steps S102 to S106.
[0026] In step S102, the content of the e-book is semantically understood from one or more dimensions to determine the key segments corresponding to each dimension.
[0027] Machine learning models can be used to semantically understand the text content of e-books from one or more dimensions, identifying key segments corresponding to each dimension. Each dimension can correspond to multiple key segments, and each key segment can include one or more text paragraphs. E-books can be narrative books, such as novels or screenplays. For example, key segments can include one or more of the following: segments corresponding to the core theme of the e-book, segments that drive the story forward, segments that shape character traits, segments that reflect character emotions, and segments that depict character growth; these are not limited to the examples given.
[0028] For example, one or more dimensions may include at least one of the following: chapter dimension (or story dimension) and character dimension. The key segments corresponding to different dimensions are not exactly the same.
[0029] In step S104, from the multiple media contents corresponding to the e-book, the media content matching each key segment is determined according to the content of each key segment.
[0030] Each piece of media content includes at least one of images and videos. Media content can also be in the form of image galleries, comics, etc., as long as it is presented in image or video format. Each piece of media content is generated based on one or more segments from an e-book. The source of the media content is not limited; it can be created by users or authors, or automatically generated by AI. It is important to note that the media content is used only with the legal authorization of the creator.
[0031] Multiple media contents corresponding to an e-book can be retrieved from different distribution locations and stored in a specific storage area of the database. For example, the media contents corresponding to an e-book can be retrieved from the e-book's comment section. If the e-book includes images, videos, etc., the media contents can also be retrieved directly from the e-book, and this is not limited to the examples given.
[0032] In some embodiments, in response to a user selecting one or more segments in the e-book reading interface and triggering the media content generation function, media content is generated based on the one or more segments and used as the media content corresponding to the e-book. Users can automatically generate media content using the generation function and can publish the generated media content in the comment area. The generated media content can be stored.
[0033] In step S106, a media content set for each dimension is generated based on the key segments corresponding to each dimension and the media content matched with each key segment.
[0034] For each dimension, each key segment and its matching media content are combined to generate a media content set for that dimension. Each key segment can be summarized, and this summary information can be combined with its matching media content to generate the media content set for that dimension. Presenting the media content and text of the key segments to the user makes reading and viewing easier and improves the user experience.
[0035] The content processing method described in the above embodiments can automatically perform semantic understanding of the content of e-books from one or more dimensions, identify key segments corresponding to each dimension, match multiple media contents corresponding to the e-book with the content of each key segment, and determine the media content matched for each key segment. Then, for each dimension, a media content set for that dimension is generated based on the key segments corresponding to that dimension and the media content matched for each key segment. Through the filtering of key segments and the matching of key segments with media content, some low-quality media content can be filtered out, and key segments and media content that are more suitable for assisting users in understanding the content of the e-book and that users are more concerned with can be selected. The media content set for each dimension can be viewed by users, who can quickly view the media content of each key segment without having to search through different distribution locations. When multiple dimensions of media content sets are generated, users can also conveniently view media content sets under different dimensions. The content processing method described in the above embodiments can effectively utilize the media content corresponding to e-books, facilitate users in viewing media content, assist users in understanding the content of e-books, enrich users' reading methods, and improve users' reading experience.
[0036] The following describes how to determine the key segments corresponding to each dimension and how to generate the media content set for each dimension.
[0037] In some embodiments, semantic understanding of the content of an e-book from one or more dimensions, and determining the key segments corresponding to each dimension, includes, for each dimension: determining the key segments corresponding to that dimension based on the results of semantic understanding of the content of the e-book from that dimension and the interaction data corresponding to each segment in the e-book.
[0038] Machine learning models can be used to semantically understand the content of e-books from different dimensions, obtaining the first key segment for each dimension. For each dimension, a second key segment can be determined based on the interaction data corresponding to each segment. The union of the first and second key segments is then used to determine the key segment corresponding to that dimension.
[0039] For example, interaction data includes at least one of the following: comment data, feedback data, and tagging data. For example, comment data includes comment data for each paragraph, comment data for each chapter, etc. Segments with a comment count greater than a first threshold can be selected as the second key segment; chapters with a comment count greater than a second threshold can be selected, and segments from those chapters with a comment count greater than a third threshold can be selected as the second key segment. For example, feedback data includes positive and negative feedback data. Positive feedback is, for example, "like," and negative feedback is, for example, "dislike." Segments with a positive feedback count greater than a fourth threshold can be selected as the second key segment; chapters with a positive feedback count greater than a fifth threshold can be selected, and segments from those chapters with a positive feedback count greater than a sixth threshold can be selected as the second key segment. Users can tag certain paragraphs or sentences, such as underlining. Segments with a tag count greater than a seventh threshold can be selected as the second key segment. If multiple interaction data are referenced, the second key segment can be determined based on each type of interaction data, or the interaction values for each segment can be weighted and summed to obtain the interaction value for each segment, and the segment with the interaction value greater than an eighth threshold can be selected as the second key segment.
[0040] For each dimension, the second key segment determined based on the interaction data can be further matched with that dimension, selecting the second key segment that matches that dimension. For example, if the dimension is a role, when determining the key segment corresponding to a certain role, for each second key segment determined based on the interaction data, the second key segment is matched with that role to determine whether it belongs to a segment in which that role appears. If it does not belong, the second key segment is filtered out.
[0041] The first key segment identified based on the content of the e-book is an important segment that embodies the content of the e-book. Interaction data can reflect the user's attention to each segment in the e-book and the popularity of each segment. Combining interaction data to select key segments can better match the user's needs, thereby improving the effectiveness of the subsequently generated media content set.
[0042] Users reading ebooks can post media content related to the ebooks, such as in the comments section. Other users can then comment on and provide feedback on the posted content. Alternatively, users can comment on and provide feedback on the media content within the ebook itself.
[0043] Interactive data includes at least one of the following: comment data, feedback data, and viewing data related to e-book media content. One or more media content pieces can be selected based on at least one of these three data points, and the corresponding segments of the selected content pieces will be used as the third key segment. Selecting one or more media content pieces based on at least one of these three data points can refer to the method for selecting the second key segment in the foregoing embodiments, and will not be repeated here. Viewing data may include the number of views.
[0044] For each dimension, the union of the first and third key segments can be used as the key segment for that dimension. Alternatively, for each dimension, the union of the first, second, and third key segments can be used as the key segment for that dimension.
[0045] In some embodiments, one or more dimensions include a chapter dimension. Semantic understanding of the content of an e-book from one or more dimensions and determination of key segments corresponding to each dimension includes: using a machine learning model to semantically understand the content of the e-book from the chapter dimension and determine multiple candidate chapters that drive the plot development from multiple chapters of the e-book; performing semantic understanding on the content of each candidate chapter and determining one or more key segments in each candidate chapter as key segments corresponding to the chapter dimension.
[0046] The chapter dimension, also known as the story dimension, involves understanding the story content of an e-book, selecting multiple candidate chapters that drive the story's development from a range of chapters, and then selecting one or more key segments from each candidate chapter, thus obtaining multiple key segments corresponding to the chapter dimension.
[0047] By selecting key segments from chapter or story dimensions, these key segments can be strung together to reflect the development of the entire story. Subsequently, a media content set is generated based on these key segments. Users can quickly understand the key nodes in the story's development by viewing the media content set, and can also quickly view key segments in different chapters and the media content corresponding to those key segments.
[0048] In some embodiments, one or more dimensions include a chapter dimension, and semantic understanding is performed on the content of each chapter to determine one or more key segments in the candidate chapters as the key segments corresponding to the chapter dimension.
[0049] One or more key segments can be extracted from each chapter as key segments corresponding to the chapter dimension. This results in a richer and more comprehensive media content set, containing information and media content from key segments of each chapter.
[0050] In some embodiments, generating a media content set for each dimension based on the key segments corresponding to each dimension and the media content matched by each key segment includes: determining the descriptive text of each key segment corresponding to the chapter dimension; and aggregating the descriptive text of the key segments corresponding to the chapter dimension and the matched media content according to the chapter order to generate a media content set for the chapter dimension.
[0051] For each key segment, the descriptive text can be the original text of the key segment or a summary of the key segment using a machine learning model. The descriptive text derived from the summary of key segments takes up less space and is easier for users to read. The descriptive text for each key segment can also include its section identifier (number, name).
[0052] By aggregating descriptive texts of multiple key segments and matching media content according to chapter order, the resulting media content set is more in line with users' reading habits, more convenient for users to view, and improves user experience.
[0053] If an ebook contains multiple storylines, one or more dimensions include the storyline dimension. In some embodiments, a machine learning model is used to determine the multiple storylines of the ebook based on a semantic understanding of the ebook's content. For each chapter included in each storyline, semantic understanding of the content of each chapter is performed to determine one or more key segments in each chapter, thereby obtaining the key segments corresponding to each storyline.
[0054] In some embodiments, for each key segment of each storyline, a descriptive text for that key segment is determined, and the descriptive text of the key segments corresponding to each storyline and the matching media content are aggregated in chapter order to generate a media content set for each storyline.
[0055] Media content sets can be generated for each storyline, allowing users to view content sets from any storyline, matching different user needs, facilitating user viewing, and improving user experience.
[0056] One or more dimensions may also include role dimensions. The following describes how to identify key segments for role dimensions.
[0057] In some embodiments, one or more dimensions include a role dimension. Semantic understanding of the content of an e-book from one or more dimensions to determine key segments corresponding to each dimension includes: using a machine learning model to semantically understand the content of the e-book from the role dimension, determining multiple parts related to each role in the e-book, wherein the parts are chapters, tasks, or branch stories; for each role, semantic understanding of each part related to the role, determining one or more key segments in each part, obtaining key segments corresponding to each role, as key segments corresponding to the role dimension.
[0058] Semantic understanding of e-book content from a character perspective can identify the chapters in which the character appears, the chapters in which the character appears as a main character, or the chapters in which the character's highlight moments occur. Furthermore, for each chapter, one or more key segments related to the character can be identified based on the semantic understanding of that chapter.
[0059] Some ebooks drive the story by having characters perform tasks. For this particular type of ebook, multiple tasks can be identified for each character, and one or more key segments can be identified from each task. If each task corresponds to multiple chapters, one or more key segments can be identified for each chapter; alternatively, one or more key chapters can be identified for each task, and one or more key segments can be identified from each key chapter.
[0060] If an ebook contains branch stories (storylines) corresponding to a certain character, for example, one can identify the branch stories of that character and then determine one or more key segments from those branch stories. For instance, one or more key segments can be determined for each chapter in a branch story, or one or more key chapters can be identified for a branch story, and one or more key segments can be determined from each key chapter.
[0061] The method for identifying key segments for each character can be determined based on the type of ebook. For example, a chapter-based approach applies to all ebooks; for ebooks where the story progresses through characters performing tasks, key segments can be identified by chapter or by task; for ebooks with branching narratives, key segments can be identified by character in addition to chapter-based identification. If an ebook both uses character-driven tasks and includes branching narratives, key segments can be identified separately by chapter, task, and branching narratives, creating different media content sets to match different user needs, facilitate user viewing, and improve user experience.
[0062] In some embodiments, generating a media content set for each dimension based on the key segments corresponding to each dimension and the media content matched for each key segment includes: determining the descriptive text of each key segment for each character, wherein the key segment includes at least one of the first appearance segment and the exit segment; and for each character, arranging the descriptive text of the key segments corresponding to the character and the matched media content in the order of story development, and aggregating them with the character's descriptive information to generate a media content set for the character.
[0063] Machine learning models can be used to generate descriptive text for key segments, which will not be elaborated further. Each character's key segments include at least one of their first appearance and departure segments, and may also include at least one of highlight segments or segments where they reappear after several chapters. A media content set can be generated for each character, including not only descriptive text for the character's key segments and matching media content, but also character description information. For example, character setting information or character introduction information such as the character's name, age, background, origins, and personality.
[0064] The descriptive text of key segments for each character and the matching media content are arranged in the order of the story's development and aggregated with the character's descriptive information to generate a media content set for each character. Users can view relevant content for characters they are interested in, which helps users understand the characters, matches different user needs, and improves the user experience.
[0065] The following describes how to determine the media content that matches each key segment.
[0066] In some embodiments, determining the media content matching each key segment from multiple media contents corresponding to an e-book, based on the content of each key segment, includes: determining one or more media contents corresponding to each key segment from multiple media contents, wherein each media content is labeled with a corresponding segment; and determining one or more media contents matching each key segment based on the semantic information of each key segment and the image recognition results of one or more media contents corresponding to each key segment.
[0067] If a user generates media content using the generation function, one or more segments selected by the user can be retrieved, and the corresponding segments can be labeled on the media content. For media content within an e-book, one or more segments corresponding to the media content can be directly identified and labeled accordingly. For user-published media content, one or more segments can be labeled based on the relevant descriptive information of the published media content or the segments corresponding to the comments section where the media content was posted. For example, if a user posts media content in the comments section corresponding to a specific segment, it can be determined that the media content corresponds to that segment.
[0068] A multimodal machine learning model can be used to understand each key segment and media content, determining one or more media content pairs that match each key segment. If a key segment does not have corresponding media content, media content can be automatically generated based on that key segment, or the key segment can be deleted.
[0069] The above methods can be used to filter out media content that better matches each key segment, thereby improving the accuracy and effectiveness of the subsequently generated media content sets and enhancing the user experience.
[0070] In some embodiments, determining one or more media contents that match each key segment based on the semantic information of each key segment and the image recognition results of one or more media contents corresponding to each key segment includes: determining one or more candidate media contents that match each key segment based on the semantic information of each key segment and the image recognition results of one or more media contents corresponding to each key segment; and in response to determining multiple candidate media contents that match each key segment, selecting candidate media contents as media contents that match each key segment based on the interaction data of the multiple candidate media contents.
[0071] If multiple candidate media content matches a given key segment or one or more key segments, the media content matching each key segment can be further filtered based on the interaction data of these candidate media content. For example, the interaction data for each candidate media content includes at least one of comment data, feedback data, and view data. For example, for each key segment, the candidate media content with the most comments or the candidate media content with the number of comments exceeding a corresponding threshold can be selected. For example, for each key segment, the candidate media content with the most positive feedback or the candidate media content with the number of positive feedback exceeding a corresponding threshold can be selected. For example, for each key segment, the candidate media content with the most views or the candidate media content with the number of views exceeding a corresponding threshold can be selected. For example, for each key segment, the interaction value of each candidate key segment can be obtained by weighted summing of the number of comments, the number of positive feedback, and the number of views, and the candidate media content with the highest interaction value or the candidate media content with the interaction value exceeding a corresponding threshold can be selected.
[0072] The method described above determines the media content corresponding to each key segment not only based on the matching results of key segments and media content, but also based on the interaction data of the media content. This can improve the accuracy of determining the media content for each key segment, while making the determined media content more in line with the user's needs and enhancing the user experience.
[0073] The following describes how to display a media content set when a user interacts with a media content set in a certain dimension.
[0074] In some embodiments, the content processing method further includes: in response to a user's operation of viewing a media content set of a target dimension in one or more dimensions, displaying an interface of the media content set of the target dimension, and displaying descriptive text of key segments and matching media content in the interface of the media content set of the target dimension.
[0075] This system can display information about one or more dimensions of media content, such as controls or entry points. Users can trigger the display of the target dimension's media content set by activating the information within that dimension. For example, information about one or more dimensions of media content can be displayed in the reading interface, comment area, recommendation interface, or other interfaces or areas of an e-book. Information about one or more dimensions of media content can be displayed in every reading interface of an e-book, or in the reading interface at the end of each chapter. Information about one or more dimensions of media content can be displayed in the comment area of each paragraph and / or chapter of an e-book. For example, information (controls, entry points, or other information) in the comment area can be associated with each paragraph and / or chapter. In response to a user triggering the comment area information, the comment area for that paragraph is displayed, and information about one or more dimensions of media content is displayed in the comment area. The display location of one or more dimensions of media content is not limited to the examples given; users can view one or more dimensions of media content from multiple locations, and users can select any dimension of media content to view.
[0076] By displaying descriptive text and matching media content for key segments within a target dimension's media content set, users can quickly understand the core content of the story within that target dimension. Users can utilize the media content set for quick review, fast reading, and skipping of uninteresting segments, aiding in comprehension of e-book content, facilitating user operation, and enhancing the user experience.
[0077] Different display strategies can be adopted for information in media content sets of different dimensions. For example, information (controls, entry points, etc.) for media content sets at the chapter level can be displayed in the e-book's reading interface, comment area, recommendation interface, and other interfaces or areas, and can be displayed in multiple interfaces, areas, and multiple locations. Information for media content sets for each character can be displayed in the reading interface and comment area corresponding to the chapters and paragraphs in which that character appears.
[0078] The interface for media content sets can be a newly added interface or an existing interface can be used to display the media content sets. For example, a media content set can be displayed in the comments interface, that is, the comments information and the media content set are displayed on the same interface. This is not limited to the examples given.
[0079] In some embodiments, displaying descriptive text and matching media content of key segments in a target dimension of media content set in the interface includes: displaying descriptive text of a first key segment, media content matching the first key segment, and title information corresponding to the first key segment in the interface; in response to a user's triggering operation on the title information corresponding to the first key segment, displaying an e-book reading interface and displaying the content of the first key segment in the reading interface; and / or displaying descriptive text of a first key segment, media content matching the first key segment, and comment information corresponding to the first key segment in the interface; in response to a user's triggering operation on the comment information corresponding to the first key segment, displaying a comment interface and displaying the comment information corresponding to the triggering operation at a specified location in the comment interface.
[0080] The interface for the target dimension's media content set can display descriptive text and matching content for one or more key segments. This depends on the screen size, the display size of the key segment's descriptive text and the matching content, and can be adjusted according to actual needs; no restrictions are imposed here. The title information of the first key segment includes at least one of the following: the title information of the chapter corresponding to the first key segment, or the title information summarized from the first key segment.
[0081] like Figure 2 As shown, the interface for the media content set at the target dimension displays the title information 201 corresponding to the first key segment, the descriptive text 202 of the first key segment, and the media content 203 matching the first key segment. If the same chapter corresponds to multiple key segments and uses chapter title information, only one title information for this chapter can be displayed. The interfaces for media content sets at both the role and chapter dimensions can use... Figure 2 The methods of display are as follows, but are not limited to the examples given.
[0082] For targets with a character-based dimension, different display methods can be used. For example... Figure 3 As shown, the interface for a media content set of a specific character can display the title information 301 corresponding to the first key segment, the descriptive text 302 of the first key segment, and the media content 303 matching the first key segment. It can also display descriptive information about the character, such as the character's name 304 and a brief introduction 305.
[0083] In response to a user's trigger action on the title information corresponding to the first key segment, the system can return to the page containing the first key segment in the e-book's reading interface. The first key segment can be highlighted in a specified way, such as by highlighting, underlining, or changing the font, and is not limited to the examples given.
[0084] The interface for the target dimension's media content set can also display commentary information corresponding to the first key segment. For example... Figure 2 As shown, the system can display partial comment information 204 corresponding to the first key segment. In response to a user triggering a comment, the system displays the comment interface. The user-triggered comment can be displayed in a specified location within the interface, such as being the first comment displayed (pinned to the top), or displayed in the middle area, etc., and is not limited to the examples given. The comment can also be highlighted in a specific way, allowing users to quickly locate it. Furthermore, the system can display feedback information 205 corresponding to comment information 204, which can include positive feedback.
[0085] In some embodiments, the target dimension is a role dimension. In response to a user's operation of viewing the media content set of the target dimension in one or more dimensions, the interface for displaying the media content set of the target dimension includes: for each role, in response to the user reading the key segment corresponding to the role, displaying the viewing entry point of the role's media content set, wherein the key segment corresponding to the role includes at least one of the role's first appearance segment and exit segment; and in response to the user triggering the viewing entry point, displaying the interface for displaying the role's media content set.
[0086] Key segments for a character can also include highlight clips of the character, clips where the character reappears after one or more interruptions, etc. These key segments display an entry point to the character's media content set, allowing users to easily view relevant character information after reading the text, aiding comprehension of the current content, and providing a review of the character's backstory, thus enhancing the user experience.
[0087] like Figure 3 As shown, the interface for a character's media content set can be the comment interface itself. That is, the comment interface displays the character's media content set, one or more comment messages 306 corresponding to the first key segment, user information 307 corresponding to the comment message 306, and feedback information 308 corresponding to the comment message 306. For example, the information from the comment interface can be displayed in association with the first key segment in the reading interface. In response to a user triggering the comment interface, the comment interface is displayed, showing the character's media content set corresponding to the first key segment. The comment interface also displays the title information corresponding to the first key segment, the descriptive text of the first key segment, and the media content matched by the first key segment.
[0088] The method described in the above embodiments allows users to access a reading interface and read the first key segment and related content by triggering the title information corresponding to that segment. This enables the linkage between the media content set and the e-book, allowing users to easily and quickly locate the original text, improving operational efficiency and convenience, and enhancing the user experience. Users can also view comment information corresponding to the first key segment, achieving linkage between the media content set and comment information, further improving operational efficiency and convenience, and enhancing the user experience.
[0089] In some embodiments, displaying descriptive text and matching media content of key segments in a target dimension's media content set in the interface includes: in response to a user's selection of one or more target key segments in the target dimension's media content set, displaying descriptive text and matching media content of one or more target key segments; in response to a user completing the viewing of the descriptive text and matching media content of one or more target key segments, displaying an e-book reading interface, and displaying content following the chapter containing one or more target key segments in the reading interface.
[0090] Users can quickly understand the content of one or more target key passages by selecting and viewing descriptive text and media content, and skip the original text of one or more target key passages to view the text content following the chapter containing those passages. This allows users to skip sections of text they don't want to read, quickly understand those sections, and connect them with subsequent text, enriching the reading experience and enhancing the overall reading experience.
[0091] Users control the display of descriptive text and matching media content for different key segments by switching between them. In some embodiments, displaying descriptive text and matching media content for key segments in a target dimension of media content set on the interface includes: displaying descriptive text for a first key segment and media content matching the first key segment on the interface; and, in response to the user's switching operation, displaying descriptive text for a second key segment and media content matching the second key segment on the interface in the order of the key segments in the media content set.
[0092] For example, users can trigger the switching display of descriptive text and media content for different key segments by using switching controls or swiping in preset directions. Users can quickly browse and view the media content corresponding to different key segments through these switching actions, improving the user experience.
[0093] In some embodiments, in response to a user's switching operation, displaying the description text of the second key segment and the media content matching the second key segment in the interface of the media content set according to the order of key segments in the media content set includes: in response to a user's switching operation, determining the second key segment according to the order of key segments in the media content set; in response to the user's reading progress not reaching the chapter corresponding to the second key segment, hiding the description text of the second key segment and the media content matching the second key segment in the interface of the media content set, and displaying a continue viewing control and guidance information, wherein the guidance information is used to prompt the user that the chapter corresponding to the second key segment has not been read; and in response to the user triggering the continue viewing control, displaying the description text of the second key segment and the media content matching the second key segment in the interface of the media content set.
[0094] If the user's switch action corresponds to the next key segment (second key segment), which belongs to a chapter the user has not yet read, the description text of the second key segment and the media content matching the second key segment can be hidden initially. If the user wants to continue viewing, they can trigger a continue viewing control, which will then display the description text of the second key segment and the media content matching the second key segment.
[0095] like Figure 4 As shown, the "Continue Viewing" control 401 and guidance information 402 are displayed in the media content set interface. For example, a mask can be used to cover the description text of the second key segment and the media content matched by the second key segment, and this is not limited to the example given.
[0096] The method described in the above embodiments allows users to choose whether to view the corresponding media content for unread chapters, thereby improving the user experience.
[0097] The interface of the media content set can also be displayed in different modes. In some embodiments, displaying descriptive text and matching media content of key segments in the target dimension media content set in the interface includes: displaying descriptive text and matching media content of key segments in the target dimension media content set in a first mode and displaying a switching control for a second mode, wherein the layout and switching method of the descriptive text and matching media content of key segments in the target dimension media content set are different in the first mode and the second mode; in response to the user triggering the switching control, displaying descriptive text and matching media content of key segments in the target dimension media content set in the second mode and displaying an autoplay control; in response to the user triggering the autoplay control, automatically switching the display of descriptive text and matching media content of key segments in the target dimension media content set and playing the audio corresponding to the descriptive text.
[0098] like Figure 2As shown, the interface for the media content set in the target dimension can display the current mode 206 and a switching control 207. In response to the user triggering the switching control 207, the following display will appear: Figure 5 The interface shown.
[0099] like Figure 5 As shown, the interface displays the media content set of the target dimension in the second mode. This interface displays the current mode 501, switching controls 502, description information of key segments 503, media content matching the key segments 504, and may also display an autoplay control 505, as well as forward switching controls 506 and backward switching controls 507. Forward switching controls 506 and backward switching controls 507 are used to switch the description text of the key segments and the matching media content.
[0100] In response to a user triggering the autoplay control, the system automatically switches between displaying descriptive text for key segments within the target media content set and the matching media content, and plays the audio corresponding to the descriptive text. If the media content is video, playback of the audio corresponding to the descriptive text can be stopped, or the video and audio can be played sequentially.
[0101] The methods described in the above embodiments provide users with a wider variety of display options, meet the needs of different users, enhance the user's visual experience, and improve the ease of user operation.
[0102] The display style of the media content set interface can vary. In some embodiments, displaying descriptive text of key segments and matching media content of a target dimension in the interface includes: displaying descriptive text of a first key segment, media content matching the first key segment, and a background image matching the first key segment in the interface, wherein the background image is generated based on the content of the first key segment.
[0103] like Figure 6 As shown, different layout methods can be used to create a more immersive user experience. The index axis for key segments can be designed in an S-shape, with the layout of the key segment's description information and matching media content matching the shape of the index axis. For example... Figure 2 As shown, the index axis of key segments can be designed as a straight line. Furthermore, as... Figure 6 As shown, a background image 604 can be displayed, which can be generated based on the content of the currently displayed first key segment.
[0104] The methods described above can provide users with a more immersive experience when browsing media content sets, enhancing the visual experience.
[0105] This disclosure also provides a content processing apparatus, which is described below in conjunction with... Figure 7 Describe it.
[0106] Figure 7 These are structural diagrams of some embodiments of the content processing apparatus disclosed herein. Figure 7 As shown, the content processing device 70 of this embodiment includes: a first determining module 710, a second determining module 720, and a generating module 730.
[0107] The first determining module 710 is configured to perform semantic understanding of the content of the e-book from one or more dimensions, and determine the key segments corresponding to each dimension.
[0108] The second determining module 720 is configured to determine media content matching each key segment from multiple media contents corresponding to the e-book, based on the content of each key segment, wherein each media content includes at least one of images and videos, and each media content is generated based on one or more segments in the e-book.
[0109] The generation module 730 is configured to generate a media content set for each dimension based on the key segments corresponding to each dimension and the media content matched by each key segment.
[0110] The content processing device described in the above embodiments can automatically perform semantic understanding of the content of e-books from one or more dimensions, determine key segments corresponding to each dimension, match multiple media contents corresponding to the e-book with the content of each key segment, and determine the media content matched for each key segment. Then, for each dimension, a media content set for that dimension is generated based on the key segments corresponding to that dimension and the media content matched for each key segment. Through the filtering of key segments and the matching of key segments with media content, some low-quality media content can be filtered out, and key segments and media content that are more suitable for assisting users in understanding the content of the e-book and that users are more concerned with can be selected. The media content set for each dimension can be viewed by the user, who can quickly view the media content of each key segment without having to search through different distribution locations. When multiple dimensions of media content sets are generated, users can also conveniently view media content sets under different dimensions. The content processing device described in the above embodiments can effectively utilize the media content corresponding to e-books, facilitate users in viewing media content, assist users in understanding the content of e-books, enrich users' reading methods, and improve users' reading experience.
[0111] In some embodiments, the second determining module 720 is configured to determine one or more media contents corresponding to each key segment from multiple media contents, wherein each media content is labeled with a corresponding segment; and to determine one or more media contents matching each key segment based on the semantic information of each key segment and the image recognition results of one or more media contents corresponding to each key segment.
[0112] In some embodiments, the second determining module 720 is configured to determine one or more candidate media contents that match each key segment based on the semantic information of each key segment and the image recognition results of one or more media contents corresponding to each key segment; in response to determining multiple candidate media contents that match each key segment, the module selects candidate media contents as media contents that match each key segment based on the interaction data of the multiple candidate media contents.
[0113] In some embodiments, the first determining module 710 is configured to, for each dimension, determine the key segment corresponding to the dimension based on the result of semantic understanding of the content of the e-book from the dimension and the interaction data corresponding to each segment in the e-book.
[0114] In some embodiments, one or more dimensions include a chapter dimension. The first determining module 710 is configured to use a machine learning model to perform semantic understanding of the content of the e-book from the chapter dimension, determine multiple candidate chapters that drive the story development among multiple chapters of the e-book; perform semantic understanding of the content of each candidate chapter, and determine one or more key segments in each candidate chapter as key segments corresponding to the chapter dimension.
[0115] In some embodiments, the generation module 730 is configured to determine the descriptive text of each key segment corresponding to the chapter dimension; and to aggregate the descriptive text of the key segments corresponding to the chapter dimension and the matching media content in chapter order to generate a media content set for the chapter dimension.
[0116] In some embodiments, one or more dimensions include a role dimension. The first determining module 710 is configured to use a machine learning model to perform semantic understanding of the content of the e-book from the role dimension, determine multiple parts related to each role in the e-book, wherein the parts are chapters, tasks or branch stories; for each role, perform semantic understanding of each part related to the role, determine one or more key segments in each part, and obtain the key segments corresponding to each role as the key segments corresponding to the role dimension.
[0117] In some embodiments, the generation module 730 is configured to determine the descriptive text of each key segment for each character, wherein the key segment includes at least one of the first appearance segment and the exit segment; and for each character, arrange the descriptive text of the key segment corresponding to the character and the matching media content in the order of story development, and aggregate them with the character's descriptive information to generate a media content set for the character.
[0118] In some embodiments, the content processing apparatus further includes a display module 740 configured to display an interface of the media content set of the target dimension in response to a user's operation of viewing a media content set of a target dimension in one or more dimensions, and to display descriptive text of key segments and matching media content in the media content set of the target dimension in the interface.
[0119] In some embodiments, the display module 740 is configured to, in response to a user's selection of one or more target key segments in a media content set of a target dimension, display descriptive text of one or more target key segments and matching media content; and in response to the user completing the viewing of the descriptive text of one or more target key segments and matching media content, display the reading interface of the e-book, and display the content following the chapter where one or more target key segments are located in the reading interface.
[0120] In some embodiments, the display module 740 is configured to display a description text of a first key segment, media content matching the first key segment, and title information corresponding to the first key segment in an interface; in response to a user's trigger operation on the title information corresponding to the first key segment, display a reading interface for an e-book and display the content of the first key segment in the reading interface; and / or display a description text of the first key segment, media content matching the first key segment, and comment information corresponding to the first key segment in an interface; in response to a user's trigger operation on the comment information corresponding to the first key segment, display a comment interface and display the comment information corresponding to the trigger operation at a specified location in the comment interface.
[0121] In some embodiments, the display module 740 is configured to display the description text of the first key segment and the media content matching the first key segment in the interface; and in response to the user's switching operation, to display the description text of the second key segment and the media content matching the second key segment in the interface according to the order of the key segments in the media content set.
[0122] In some embodiments, the display module 740 is configured to, in response to a user's switching operation, determine a second key segment according to the order of key segments in the media content set; in response to the user's reading progress not reaching the chapter corresponding to the second key segment, hide the description text of the second key segment and the media content matching the second key segment in the interface of the media content set, and display a continue viewing control and guidance information, wherein the guidance information is used to prompt the user that the chapter corresponding to the second key segment has not been read; in response to the user triggering the continue viewing control, display the description text of the second key segment and the media content matching the second key segment in the interface.
[0123] In some embodiments, the display module 740 is configured to display descriptive text of key segments and matching media content in a target dimension media content set in a first mode on the interface, and display a switching control for a second mode, wherein the layout and switching method of the descriptive text of key segments and matching media content in the target dimension media content set are different in the first mode and the second mode; in response to the user triggering the switching control, the descriptive text of key segments and matching media content in the target dimension media content set are displayed in the second mode, and an autoplay control is displayed; in response to the user triggering the autoplay control, the display of descriptive text of key segments and matching media content in the target dimension media content set is automatically switched, and the audio corresponding to the descriptive text is played.
[0124] In some embodiments, the target dimension is the role dimension, and the display module 740 is configured to, for each role, display the viewing entry of the role's media content set in response to the user reading the key segment corresponding to the role, wherein the key segment corresponding to the role includes at least one of the role's first appearance segment and exit segment; and display the interface of the role's media content set in response to the user triggering the viewing entry.
[0125] In some embodiments, the display module 740 is configured to display in the interface a descriptive text of a first key segment, media content matching the first key segment, and a background image matching the first key segment, wherein the background image is generated based on the content of the first key segment.
[0126] This disclosure also provides an electronic device, which is described below in conjunction with... Figure 8 and 9 Describe it. Figure 8 A block diagram of an electronic device according to some embodiments of the present disclosure is shown.
[0127] like Figure 8 As shown, the electronic device 8 includes a processor 82; and a memory 81 coupled to the processor 82 for storing instructions, which, when executed by the processor 82, cause the processor 82 to perform a content processing method as described in any embodiment of this disclosure.
[0128] The electronic device described in the above embodiments can automatically perform semantic understanding of the content of e-books from one or more dimensions, identify key segments corresponding to each dimension, match multiple media contents corresponding to the e-book with the content of each key segment, and determine the media content matched for each key segment. Then, for each dimension, a media content set for that dimension is generated based on the key segments corresponding to that dimension and the media content matched for each key segment. Through the filtering of key segments and the matching of key segments with media content, some low-quality media content can be filtered out, and key segments and media content that are more suitable for assisting users in understanding the content of the e-book and that users are more concerned with can be selected. The media content set for each dimension can be viewed by the user, allowing the user to quickly view the media content of each key segment without having to search through different distribution locations. When multiple dimensions of media content sets are generated, users can also conveniently view media content sets under different dimensions. The electronic device described in the above embodiments can effectively utilize the media content corresponding to e-books, facilitate users in viewing media content, assist users in understanding the content of e-books, enrich users' reading methods, and improve users' reading experience.
[0129] Memory 81 is used to store one or more computer-readable instructions. Memory 81 may include any combination of various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory, including but not limited to random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), read-only memory (ROM), and flash memory. Memory 81 may, for example, store operating systems, application programs, bootloaders, databases, and other programs, as well as various application programs and various data.
[0130] The processor 82 is configured to execute computer-readable instructions to implement the method of any of the foregoing embodiments. Specific implementations of each step of the method can be found in the above embodiments, and repeated details will not be elaborated here.
[0131] Processor 82 can be configured to execute Figure 1 The processor 82 can be various processing devices, such as a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The central processing unit (CPU) can be an x86 or ARM architecture, etc.
[0132] The processor 82 and the memory 81 can communicate with each other directly or indirectly. For example, the processor 82 and the memory 81 can communicate via a network. The network can include wireless networks, wired networks, and / or any combination of wireless and wired networks. The processor 82 and the memory 81 can also communicate with each other via a system bus, which is not limited in this disclosure.
[0133] It should be noted that Figure 8 The components of the electronic device 8 shown are merely exemplary and not limiting. The electronic device 8 may have other components depending on the specific application requirements. The processor 82 can control other components in the electronic device 8 to perform desired functions.
[0134] Electronic device 8 can be implemented by software, firmware and / or hardware, and can be integrated into a device with the relevant application installed.
[0135] Figure 9 Block diagrams of electronic devices according to other embodiments of the present disclosure are shown.
[0136] Figure 9 The electronic device 9 shown can be a computer system with a dedicated hardware structure, which can perform corresponding functions when the relevant application is installed.
[0137] Electronic devices include, but are not limited to, mobile terminals such as smartphones, laptops, personal digital assistants (PDAs), tablet PCs, portable multimedia players (PMPs), in-vehicle terminals (such as in-vehicle navigation terminals), wearable devices, and fixed terminals such as digital televisions and desktop computers.
[0138] like Figure 9 As shown, the Central Processing Unit (CPU) 91 performs various processes based on programs stored in the Read-Only Memory (ROM) 92 or programs loaded from the storage section 98 into the Random Access Memory (RAM) 93. The RAM 93 stores data required as needed when the CPU 91 performs various processes, etc. The CPU is merely exemplary; it could also be other types of processors, such as the various processors described above. The ROM 92, RAM 93, and storage section 98 can be various forms of computer-readable storage media. It should be noted that although... Figure 9 The diagram shows ROM 92, RAM 93 and storage section 98, but one or more of them may be combined or located in the same or different memory or storage modules.
[0139] CPU 91, ROM 92 and RAM 93 are interconnected via bus 94. Input / output interface 95 is also connected to bus 94.
[0140] The following components are connected to the input / output interface 95: input section 96, such as a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output section 97, including displays such as cathode ray tube (CRT), liquid crystal display (LCD), speakers, vibrators, etc.; storage section 98, including hard disk, magnetic tape, etc.; and communication section 99, including network interface cards such as LAN cards, modems, etc. The communication section 99 allows communication processing via a network such as the Internet. It is easy to understand that, although... Figure 9 The portion of the electronic device 9 shown communicates via bus 94, but it may also communicate via a network or other means, wherein the network may include a wireless network, a wired network, and / or any combination of wireless and wired networks.
[0141] As needed, drive 910 is also connected to input / output interface 95. Removable media 911, such as disks, optical disks, magneto-optical disks, semiconductor memories, etc., are installed on drive 910 as needed, so that computer programs read from them can be installed into storage section 98 as needed.
[0142] When the above series of processes are implemented through software, the program constituting the software can be installed from a network such as the Internet or a storage medium such as a removable medium 911.
[0143] According to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. This disclosure also provides a computer program product, including: instructions, wherein when executed by a processor, the instructions cause the processor to perform the content processing method of any embodiment of this disclosure.
[0144] The disclosed computer program product can automatically perform semantic understanding of e-book content from one or more dimensions, identify key segments corresponding to each dimension, match multiple media contents corresponding to the e-book with the content of each key segment, and determine the media content matched for each key segment. Then, for each dimension, based on the key segments corresponding to that dimension and the media content matched for each key segment, a media content set for that dimension is generated. Through the filtering of key segments and the matching of key segments with media content, some low-quality media content can be filtered out, and more suitable key segments and media content that are more relevant to the user's understanding of the e-book content can be selected. The media content set for each dimension can be viewed by the user, allowing them to quickly view the media content of each key segment without having to search through different distribution locations. When multiple dimensions of media content sets are generated, users can also conveniently view media content sets under different dimensions. The disclosed computer program product can effectively utilize the media content corresponding to e-books, facilitate user viewing of media content, assist users in understanding the content of e-books, enrich users' reading methods, and enhance their reading experience.
[0145] For example, some embodiments of this disclosure include a computer program product that, when run on a computer, causes the computer to perform the methods of any of the foregoing embodiments. The computer program product includes computer instructions carried on a computer-readable medium, comprising program code for performing the methods shown in the flowchart. In such embodiments, the computer instructions can be downloaded and installed from a network via communication section 99, or installed from storage section 98, or installed from ROM 92. When the computer program is executed by CPU 91, the methods of embodiments of this disclosure are performed.
[0146] This disclosure also provides a computer-readable storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the processor performs a content processing method according to any embodiment of this disclosure.
[0147] This disclosed computer-readable storage medium can automatically perform semantic understanding of e-book content from one or more dimensions, identify key segments corresponding to each dimension, match multiple media contents corresponding to the e-book with the content of each key segment, and determine the media content matched for each key segment. Then, for each dimension, based on the key segments corresponding to that dimension and the media content matched for each key segment, a media content set for that dimension is generated. Through the filtering of key segments and the matching of key segments with media content, some low-quality media content can be filtered out, and key segments and media content that are more suitable for assisting users in understanding the e-book content and that users are more concerned with can be selected. The media content set for each dimension can be viewed by users, who can quickly view the media content of each key segment without having to search through different distribution locations. When multiple dimensions of media content sets are generated, users can also conveniently view media content sets under different dimensions. This disclosed computer-readable storage medium can effectively utilize the media content corresponding to e-books, facilitate users in viewing media content, assist users in understanding the content of e-books, enrich users' reading methods, and improve users' reading experience.
[0148] It should be noted that, in the context of this disclosure, a computer-readable medium can be a tangible medium that may contain or store programs for use by or in conjunction with an instruction execution system, apparatus, or device.
[0149] A computer-readable medium may be a computer-readable storage medium, a computer-readable signal medium, or any combination thereof.
[0150] Computer-readable storage media include, but are not limited to, systems, apparatuses, or devices that are electrical, magnetic, optical, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. Computer instructions are stored on the computer-readable storage medium that, when executed by a processor, implement the methods described in any of the foregoing embodiments.
[0151] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0152] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0153] In some embodiments, a computer program is also provided, comprising: instructions that, when executed by a processor, cause the processor to perform the methods described in any of the foregoing embodiments. For example, the instructions may be embodied in computer program code.
[0154] In embodiments of this disclosure, computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof. These programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network (including a local area network (LAN) or a wide area network (WAN)), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0155] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0156] The functions described above can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary hardware logic components that can be used include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0157] While specific embodiments of this disclosure have been described in detail by way of example, those skilled in the art should understand that the examples are for illustrative purposes only and not intended to limit the scope of this disclosure. Those skilled in the art should understand that modifications can be made to the above embodiments without departing from the scope and spirit of this disclosure. The scope of this disclosure is defined by the appended claims.
Claims
1. A content processing method, comprising: Semantic understanding of the content of e-books is performed from one or more dimensions to identify key segments corresponding to each dimension; From the multiple media contents corresponding to the e-book, media content matching each key segment is determined based on the content of each key segment, wherein each media content includes at least one of images and videos, and each media content is generated based on one or more segments in the e-book; Based on the key segments corresponding to each dimension and the media content matching each key segment, a media content set for each dimension is generated.
2. The content processing method according to claim 1, wherein, The step of determining the media content matching each key segment from multiple media contents corresponding to the e-book, based on the content of each key segment, includes: From the plurality of media content, one or more media content corresponding to each key segment are determined, wherein each media content is labeled with the corresponding segment; Based on the semantic information of each key segment and the image recognition results of one or more media contents corresponding to each key segment, one or more media contents that match each key segment are determined.
3. The content processing method according to claim 2, wherein, The step of determining one or more media contents that match each key segment based on the semantic information of each key segment and the image recognition results of one or more media contents corresponding to each key segment includes: Based on the semantic information of each key segment and the image recognition results of one or more media contents corresponding to each key segment, determine one or more candidate media contents that match each key segment; In response to determining multiple candidate media contents that match each key segment, the candidate media contents are selected as the media contents that match each key segment based on the interaction data of the multiple candidate media contents.
4. The content processing method according to claim 1, wherein, The semantic understanding of the content of the e-book from one or more dimensions, and the determination of key segments corresponding to each dimension, includes, for each dimension: Based on the semantic understanding results of the content of the e-book from the said dimension, and the interactive data corresponding to each segment in the e-book, the key segments corresponding to the said dimension are determined.
5. The content processing method according to claim 1, wherein, The one or more dimensions include chapter dimensions, and the semantic understanding of the e-book content from one or more dimensions to determine the key segments corresponding to each dimension includes: Using a machine learning model, semantic understanding of the content of the e-book is performed from the chapter dimension to identify multiple candidate chapters that drive the story forward from multiple chapters of the e-book; Semantic understanding is performed on the content of each candidate chapter to identify one or more key segments in each candidate chapter as the key segments corresponding to the chapter dimension.
6. The content processing method according to claim 5, wherein, The step of generating a media content set for each dimension based on the key segments corresponding to each dimension and the media content matched with each key segment includes: For each key segment corresponding to the chapter dimension, determine the descriptive text of the key segment; The descriptive text of the key segments corresponding to the chapter dimension and the matching media content are aggregated in chapter order to generate the media content set of the chapter dimension.
7. The content processing method according to claim 1, wherein, The one or more dimensions include a role dimension, and the semantic understanding of the e-book content from one or more dimensions to determine the key segments corresponding to each dimension includes: Using a machine learning model, the content of the e-book is semantically understood from the perspective of the character, and multiple parts related to each character in the e-book are identified, wherein the parts are chapters, tasks or branch stories; For each role, semantic understanding is performed on each part related to the role to determine one or more key segments in each part, and the key segments corresponding to each role are obtained as the key segments corresponding to the role dimension.
8. The content processing method according to claim 7, wherein, The step of generating a media content set for each dimension based on the key segments corresponding to each dimension and the media content matched with each key segment includes: For each key segment corresponding to each character, a descriptive text for the key segment is determined, wherein the key segment includes at least one of the first appearance segment and the exit segment; For each character, the descriptive text of the key segments corresponding to the character and the matching media content are arranged in the order of the story development, and then aggregated with the character's descriptive information to generate the character's media content set.
9. The content processing method according to any one of claims 1-8, further comprising: In response to a user's action of viewing a media content set of a target dimension among the one or more dimensions, an interface for the media content set of the target dimension is displayed, and descriptive text of key segments and matching media content are displayed in the interface.
10. The content processing method according to claim 9, wherein, The description text and matching media content of key segments in the media content set of the target dimension displayed on the interface include: In response to the user's selection of one or more target key segments in the media content set of the target dimension, the descriptive text of the one or more target key segments and the matching media content are displayed; In response to the user completing the viewing of the descriptive text and matching media content of one or more target key segments, the reading interface of the e-book is displayed, and the content following the chapter where the one or more target key segments are located is displayed in the reading interface.
11. The content processing method according to claim 9, wherein, The description text and matching media content of key segments in the media content set of the target dimension displayed on the interface include: The interface displays a description of the first key segment, media content matching the first key segment, and title information corresponding to the first key segment. In response to the user's triggering operation on the title information corresponding to the first key segment, the e-book's reading interface is displayed, and the content of the first key segment is displayed on the reading interface; and / or The interface displays a description of the first key segment, media content matching the first key segment, and comment information corresponding to the first key segment. In response to the user's trigger operation on the comment information corresponding to the first key segment, a comment interface is displayed, and the comment information corresponding to the trigger operation is displayed at a specified position in the comment interface.
12. The content processing method according to claim 9, wherein, The description text and matching media content of key segments in the media content set of the target dimension displayed on the interface include: The interface displays the descriptive text of the first key segment and the media content matched by the first key segment; In response to the user's switching operation, the description text of the second key segment and the media content matching the second key segment are displayed on the interface in the order of the key segments in the media content set.
13. The content processing method according to claim 12, wherein, In response to the user's switching operation, the description text of the second key segment and the media content matching the second key segment are displayed on the interface in the order of the key segments in the media content set, including: In response to the user's switching operation, the second key segment is determined according to the order of key segments in the media content set; In response to the user's reading progress not reaching the chapter corresponding to the second key segment, the description text of the second key segment and the media content matching the second key segment are hidden in the interface of the media content set, and a continue viewing control and guidance information are displayed, wherein the guidance information is used to prompt the user that the chapter corresponding to the second key segment has not been read; In response to the user triggering the continue viewing control, the description text of the second key segment and the media content matching the second key segment are displayed in the interface.
14. The content processing method according to claim 9, wherein, The description text and matching media content of key segments in the media content set of the target dimension displayed on the interface include: The interface displays the descriptive text of key segments and matching media content in the target dimension's media content set in a first mode, and displays a switching control for a second mode. The layout and switching method of the descriptive text of key segments and matching media content in the target dimension's media content set are different in the first mode and the second mode. In response to the user triggering the switching control, the descriptive text of key segments and matching media content in the media content set of the target dimension are displayed in a second mode, and an autoplay control is displayed; In response to the user triggering the autoplay control, the system automatically switches between displaying the descriptive text of key segments in the media content set of the target dimension and the matching media content, and plays the audio corresponding to the descriptive text.
15. The content processing method according to claim 9, wherein, The target dimension is a role dimension, and the interface for displaying the media content set of the target dimension in response to a user's operation of viewing the media content set of the target dimension among one or more dimensions includes: For each character, in response to the user reading the key segment corresponding to the character, an entry point for viewing the media content set of the character is displayed, wherein the key segment corresponding to the character includes at least one of the character's first appearance segment and exit segment; In response to the user triggering the viewing entry, an interface displaying the media content set of the character is shown.
16. The content processing method according to claim 9, wherein, The description text and matching media content of key segments in the media content set of the target dimension displayed on the interface include: The interface displays a description text of a first key segment, media content matching the first key segment, and a background image matching the first key segment, wherein the background image is generated based on the content of the first key segment.
17. A content processing apparatus, comprising: The first determination module is configured to perform semantic understanding of the content of the e-book from one or more dimensions and determine the key segments corresponding to each dimension. The second determining module is configured to determine media content matching each key segment from multiple media contents corresponding to the e-book, based on the content of each key segment, wherein each media content includes at least one of images and videos, and each media content is generated based on one or more segments in the e-book; The generation module is configured to generate a media content set for each dimension based on the key segments corresponding to each dimension and the media content matched by each key segment.
18. An electronic device comprising: processor; as well as A memory coupled to the processor is used to store instructions that, when executed by the processor, cause the processor to perform the content processing method as described in any one of claims 1 to 16.
19. A computer-readable storage medium having a computer program stored thereon, wherein, When the program is executed by a processor, the processor performs the content processing method as described in any one of claims 1 to 16.
20. A computer program product comprising: Instructions, wherein when executed by a processor, the processor performs the content processing method as described in any one of claims 1 to 16.