Audio playing method and device, computer readable storage medium and product

Displaying highlight audio through the audio recommendation interface solves the problem of long search paths for users in audio book applications, enables quick access to the playback interface of content of interest, and improves operational efficiency and resource utilization.

CN120596055APending Publication Date: 2025-09-05DOUYIN VISION CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510756689.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Users have to navigate long paths to find books in audiobook apps or websites, repeatedly jumping between different pages, which consumes client resources and reduces operational efficiency.

Method used

Highlight audio is displayed through the audio recommendation interface, including a recommendation stream generated based on the original content and transition content of multiple chapters, allowing users to quickly switch to the playback interface of the audiobooks they are interested in, reducing interface jumps.

Benefits of technology

It improves the efficiency of user information acquisition and the success rate of book finding, saves client resources and improves operational efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120596055A_ABST
    Figure CN120596055A_ABST
Patent Text Reader

Abstract

The invention relates to an audio playing method and device, a computer readable storage medium and a product, and relates to the technical field of computers. The audio playing method comprises the steps that an audio recommendation interface is displayed, the recommendation interface comprises first audio, the first audio is generated according to contents of multiple chapters of a first audio book and comprises original contents extracted from the multiple chapters and transition contents, and the transition contents are generated based on the original contents; playing the first audio; in response to a first operation on the recommendation interface, playing a second audio different from the first audio in the audio recommendation stream; and in response to a second operation on the recommendation interface, playing the original audio of the first audio book.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular to an audio playback method, device, computer-readable storage medium, and product. Background Art

[0002] With the development of multimedia application technology, users can consume a wide variety of content. Audiobooks are content that carries books through sound. Audiobooks are generally produced by voice actors reading the book text aloud, or automatically through text-to-speech (TTS) technology.

[0003] When users search for a book of interest on an app or website offering audiobooks, they typically find an audiobook on a recommended page, such as the homepage, and then navigate to its details. They then randomly select a chapter from the catalog and play that chapter on the book's play page. If interested, they'll listen to the audiobook from the beginning. Summary of the Invention

[0004] According to some embodiments of the present disclosure, an audio playback method is provided, including: displaying a recommendation interface for audio, the recommendation interface including first audio, wherein the first audio is generated based on the content of multiple chapters of a first audio book and includes original content extracted from the multiple chapters and transition content, and the transition content is generated based on the original content; playing the first audio; in response to a first operation on the recommendation interface, playing second audio in an audio recommendation stream that is different from the first audio; and in response to a second operation on the recommendation interface, playing the original audio of the first audio book.

[0005] According to other embodiments of the present disclosure, an audio playback device is provided, including: a display module configured to display an audio recommendation interface, the recommendation interface including a first audio, wherein the first audio is generated based on the contents of multiple chapters of a first audio book and includes original content extracted from the multiple chapters and transition content, and the transition content is generated based on the original content; a first playback module configured to play the first audio; a second playback module configured to play a second audio different from the first audio in an audio recommendation stream in response to a first operation on the recommendation interface; and a third playback module configured to play the original audio of the first audio book in response to a second operation on the recommendation interface.

[0006] According to some embodiments of the present disclosure, an audio playback device is provided, comprising: a memory; and a processor coupled to the memory, wherein the processor is configured to execute the audio playback method of any embodiment of the present disclosure based on instructions stored in the memory.

[0007] According to some embodiments of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the audio playback method of any embodiment of the present disclosure is performed.

[0008] According to some embodiments of the present disclosure, a computer program product is provided. When the computer program product is run on a computer, the computer implements the audio playback method of any embodiment of the present disclosure.

[0009] Other features, aspects and advantages of the present disclosure will become apparent from the following detailed description of exemplary embodiments of the present disclosure with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The following describes embodiments of the present disclosure with reference to the accompanying drawings. It should be understood that the drawings described below only relate to some embodiments of the present disclosure and do not constitute a limitation to the present disclosure. In the accompanying drawings:

[0011] Figure 1 A flowchart of an audio playback method according to some embodiments of the present disclosure is shown.

[0012] Figure 2 A schematic flow chart of a method for generating a first audio according to some embodiments of the present disclosure is shown.

[0013] Figure 3 A schematic diagram showing a recommendation interface according to some embodiments of the present disclosure is shown.

[0014] Figure 4 Schematic diagrams showing recommendation interfaces according to other embodiments of the present disclosure.

[0015] Figure 5 A schematic structural diagram of an audio playback device according to some embodiments of the present disclosure is shown.

[0016] Figure 6 A block diagram of an electronic device according to some embodiments of the present disclosure is shown.

[0017] Figure 7 A block diagram of an electronic device according to some other embodiments of the present disclosure is shown.

[0018] It should be understood that, for ease of description, the dimensions of the various parts shown in the drawings are not necessarily drawn to scale. The same or similar reference numerals are used throughout the drawings to indicate the same or similar parts. Therefore, once an item is defined in one drawing, it may not be discussed further in subsequent drawings. DETAILED DESCRIPTION

[0019] The following will be combined with the accompanying drawings in the embodiments of the present disclosure to clearly and completely describe the technical solutions in the embodiments of the present disclosure. It should be understood that the present disclosure can be implemented in various forms and should not be interpreted as being limited to the embodiments described herein.

[0020] It should be understood that the various steps described in the method embodiments of the present disclosure can be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect. Unless otherwise specifically stated, the relative arrangement of the components and steps described in these embodiments should be interpreted as being merely exemplary and not limiting the scope of the present disclosure.

[0021] The term “including” and its variations used in the present disclosure are open terms that include at least the following elements / features but do not exclude other elements / features, that is, “including but not limited to.” The term “based on” means “at least in part based on.”

[0022] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules, or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules, or units. Unless otherwise specified, concepts such as "first" and "second" are not intended to imply that the objects described in such a manner must be in a given order in time, space, ranking, or any other manner.

[0023] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".

[0024] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0025] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0026] The following detailed description of the embodiments of the present disclosure is provided in conjunction with the accompanying drawings, but the present disclosure is not limited to these specific embodiments. The following specific embodiments may be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments. In addition, in one or more embodiments, specific features, structures, or characteristics may be combined in any suitable manner that will be apparent to those skilled in the art from this disclosure.

[0027] In the related art, users face a long search path within audiobook apps or websites. During this process, they must repeatedly navigate between the website or app's homepage, the details page, and the playback page for each audiobook. This means that the search process triggers the display of numerous pages, consuming significant client resources. Furthermore, user operation efficiency is relatively low.

[0028] The present disclosure provides an audio playback method. In this method, a recommended audio stream is displayed through an audio recommendation interface, allowing users to listen to the highlights of audio books in the recommended stream. Thus, if the user is interested in a highlight audio, they can quickly access the playback interface of the audio book corresponding to the highlight audio.

[0029] Figure 1 FIG. 1 shows a flow chart of an audio playback method according to some embodiments of the present disclosure. Figure 1 As shown, the audio playing method of this embodiment includes steps S11 to S14.

[0030] In step S11, an audio recommendation interface is displayed, where the recommendation interface includes a first audio, wherein the first audio is generated based on contents of multiple chapters of a first audio book and includes original content extracted from the multiple chapters and transition content, where the transition content is generated based on the original content.

[0031] The recommendation interface is used to play the highlight audio of one or more audio books, such as the first audio, the second audio, etc. The recommendation interface can display the relevant information of multiple audios at the same time, or it can display the relevant information of only one audio at a time.

[0032] The first audio contains both original content from the audiobook and transition content. The transition content can be a connection between original content from different chapters or an introduction to the original content in a certain chapter.

[0033] The extracted segments of original content are non-adjacent in the first audiobook. That is, the first audio is not a complete paragraph extracted directly from the first audiobook, but is instead a splicing of multiple non-adjacent paragraphs with transitional content. This allows the first audio to more comprehensively present the content of the first audiobook.

[0034] In step S12, the first audio is played.

[0035] The first audio may be manually triggered to play by a user or automatically played. In the automatic playback scenario, the first audio may be automatically played in response to content corresponding to the first audio being displayed on the recommendation interface.

[0036] In step S13 , in response to the first operation on the recommendation interface, a second audio different from the first audio in the audio recommendation stream is played.

[0037] The first operation may be an operation for switching audio. In response to the first operation, the first audio stops playing and the second audio starts playing.

[0038] In step S14 , in response to the second operation on the recommendation interface, the original audio of the first audio book is played.

[0039] The second action is used to trigger the playback of the original content of the audiobook corresponding to the first audio. For example, on the recommendation interface, a control similar to "Listen to the Full Text" can be set, and the second action is the operation used to trigger this control. Of course, the second action can also be implemented in other ways, which will not be detailed here.

[0040] Through the above embodiments, the present disclosure can quickly display the highlight audio of multiple audio books to users through the audio recommendation interface. The highlight audio is generated based on the original content and transition content of multiple chapters of the audio book, so as to efficiently deliver the key content of the audio book to the user. Moreover, if the user is interested, the original audio of the audio book can be played through the second operation. Thus, the user does not need to jump back and forth between different interfaces in the process of finding a book. Through the first operation or the second operation, the user can switch to the highlight audio of other audio books, or start listening to the original audio directly. In this way, the client does not need to display a large number of irrelevant interfaces unnecessarily, and the user does not need to trigger multiple interfaces through complex operations, thereby saving the client's resources and improving the user's operating efficiency.

[0041] In the process of generating the first audio, it is necessary to select which chapters to extract the original content from to add to the highlight audio. The target chapters for providing the original content can be determined according to the target dimension of each audiobook. Figure 2 FIG. 1 shows a flow chart of a method for generating a first audio signal according to some embodiments of the present disclosure. Figure 2 As shown, the method of this embodiment includes steps S21 to S23.

[0042] In step S21 , a target dimension of the first audio book is determined according to the original content of the first audio book.

[0043] The target dimension is used to characterize the characteristics of an audiobook. This dimension can include at least one of the following: the identity of a key character, the relationship status of multiple related characters, plot turning points, or knowledge points. For example, for an audiobook focused on character development, the identity of a key character (e.g., the protagonist) can be the target dimension; for an audiobook focused on group stories, the relationship status of multiple related characters can be the target dimension; for an audiobook focused on suspense, plot turning points can be the target dimension; for an audiobook focused on knowledge, knowledge points can be the target dimension, and so on.

[0044] In some embodiments, a machine learning model can be used to semantically understand the original content of the first audio book to determine its target dimension. For example, the machine learning model can be pre-trained using the original content of the audio book annotated with the target dimension. After receiving the original content of the audio book as input, the trained model can output the target dimension of the audio book. Alternatively, the type of the audio book can be determined based on the original content of the audio book, and the target dimension can be determined based on the correspondence between the type and the target dimension.

[0045] The machine learning model can be a neural network model, such as a large language model (LLM) or a basic model.

[0046] Of course, some audiobooks may include multiple target dimensions, and multiple highlight videos may be generated based on the different target dimensions of a particular audiobook.

[0047] In step S22 , a plurality of target chapters associated with the target dimension in the original content of the first audio book are determined.

[0048] For example, the original content of the first audiobook may be searched for chapters that include the target dimension.

[0049] In step S23, a first audio corresponding to the target dimension is generated according to the target chapter.

[0050] For example, the target original text associated with the target dimension can be determined from the target chapter, where the target original text can be all or part of the content of the target chapter; then, transition content is generated based on the target original text; finally, the first audio corresponding to the target dimension is generated based on the target original text and the transition content.

[0051] Through the above embodiment, the generated highlight audio can better match the content of the audio book, thereby improving the user's information acquisition efficiency and the success rate of finding books, and thus can also save client resources.

[0052] The following uses several target dimensions as examples to exemplarily describe the method of determining the target dimension and generating the first audio.

[0053] First, the target dimension is described as the identity of the key role as an example.

[0054] An exemplary method for determining the target dimension is as follows: determining the target dimension of the first audiobook based on the original content of the first audiobook includes: determining a key character of the first audiobook based on the original content of the first audiobook; and determining the identity of the key character as the target dimension of the first audiobook in response to the key character having different identities in multiple chapters.

[0055] A key role can be a role whose ranking for the amount of content involved is greater than a threshold, or a role whose ranking for the number of chapters involved is greater than a specified ranking, etc. If the role is a key role in the audiobook and its identity changes as the book content progresses, the identity of the role can be determined as the target dimension.

[0056] The identities of key characters can be determined by searching the original content for identity-related keywords, or by processing the original content using a machine learning model. The machine learning model can be a pre-trained model that outputs the identity of the character based on the input content and the character.

[0057] An exemplary method for generating the first audio is as follows. Generating the first audio corresponding to the target dimension based on the target chapter includes: extracting multiple dialogues representing multiple identities of key characters from the target chapter; generating dialogue scene description information as transition content based on the multiple dialogues and the context of the multiple dialogues; and generating the first audio based on the multiple dialogues and the transition content.

[0058] Specifically, dialogues that characterize each role can be extracted from the target chapter corresponding to that role. These dialogues can be generated by the key role or by roles other than the key role. This extraction process can be implemented using a machine learning model. For example, the original content of the target chapter and the role are input into the model, along with processing instructions to extract dialogues that characterize the role from the original content, and then the model outputs the dialogues.

[0059] The dialogue scene description information may be a description of the environment, characters, and events during the dialogue process. This information can be obtained by obtaining the paragraphs containing the dialogue in the original content and extracting key information from these paragraphs.

[0060] For example, an audio novel describes the story of a person who grows from a vendor in a vegetable market to the boss of a large enterprise. The target chapter includes the chapters involving the character A as a vendor and the chapters involving the character A as a boss. The dialogues extracted from the target chapter include, for example, dialogue 1 and dialogue 2. Dialogue 1 corresponds to the identity of the vendor, and the content includes: "Mrs. Liu, are you coming to buy vegetables? Today's vegetables are fresh!" "A, you should clean yourself up, the girl from that family is willing to marry you." Dialogue 2 corresponds to the identity of the boss, and the content includes: "Mr. A, the other party in the merger and acquisition case wants a 20% premium and has a tough attitude..." "Cut it down to 15%, and if that doesn't work, build your own factory." The transition content is, for example, City B, today is xx month of xx year, and it is raining outside. On the top floor of the tallest office building in Block C, A stares at the street view outside the window with a serious expression. Thus, the content of the generated first audio is, for example:

[0061] (Dialogue) "Mrs. Liu, are you coming to buy some groceries? Today's vegetables are so fresh!" "Ah, you should clean yourself up. Which girl would be willing to marry you?"

[0062] (Narration) City B, today is xx month of xx year, and it's raining outside. On the top floor of the tallest office building in Block C, A gazes out the window at the street scene with a serious expression.

[0063] (Dialogue) "Mr. A, the other party in the merger and acquisition is demanding a 20% premium and is being very aggressive..." "Let's get it down to 15%. If that doesn't work, we'll build our own factory."

[0064] Next, the relationship status of multiple associated roles is described as an example.

[0065] An exemplary method for determining a target dimension is as follows: determining the target dimension of a first audio book based on original content of the first audio book includes: determining, based on the original content of the first audio book, multiple associated characters in the first audio book; and, in response to the multiple associated characters having different relationship statuses in multiple chapters, determining the relationship statuses of the multiple associated characters as the target dimension of the first audio book.

[0066] For example, the associated characters include characters that appear in the same scene in the audiobook. Considering that some audiobooks have many characters, multiple groups of associated characters can be sorted according to the frequency of appearance in the same scene to select multiple associated characters with stronger correlation.

[0067] Relationship states are used to describe the relationships between multiple characters, such as classmates, friends, couples, colleagues, and so on. After determining the associated characters, the original content of the chapters where the multiple associated characters appear, along with processing instructions, can be input into the machine learning model. The processing instructions indicate that multiple relationship states of the multiple associated characters are determined based on the original content, and then the multiple relationship states are output by the model.

[0068] The change in the relationship status between related characters reflects the change in the content of the story, so it can be used as the target dimension.

[0069] An exemplary method for generating the first audio is as follows. Generating the first audio corresponding to the target dimension based on the target chapter includes: extracting, from the target chapter, multiple dialogues representing multiple relationship states of multiple related characters; generating dialogue scene description information as transition content based on the multiple dialogues and the context of the multiple dialogues; determining the timbre of each character based on relevant content of each character in the original content of the first audiobook; and generating the first audio based on the dialogues, the transition content, and the timbre of each character.

[0070] Specifically, dialogues that can represent each relationship state can be extracted from the target chapter corresponding to each relationship state. These dialogues can be dialogues between the multiple associated roles, or between a role in the multiple associated roles and other roles outside the associated roles. This extraction process can be implemented by a machine learning model. For example, the original content and relationship state of the target chapter are input into the model, as well as processing instructions that indicate extracting dialogues that can represent the relationship state from the original content, and then obtaining dialogues output by the model.

[0071] In dialogue scenes involving multiple linked characters, different timbres can be assigned to different characters. The timbre of each character in the first audio can be the same as the timbre of the same character in the original audiobook content. For example, the timbre of a character in the first audio can be determined based on the character's voice characteristics in the original content. Transitional content can use a different timbre for each of the multiple linked characters. This timbre can be the timbre of the narration in the original content or a newly created timbre.

[0072] Below, we describe the target dimension as a plot turning point as an example.

[0073] An exemplary manner of determining the target dimension is as follows: determining the target dimension of the first audio book based on the original content of the first audio book includes: in response to the first audio book including multiple plot turning points, determining the plot turning points as the target dimension of the first audio book.

[0074] Plot turning points represent content or events where the plot undergoes significant changes. For example, the original content and processing instructions can be input into a machine learning model. The processing instructions instruct the model to extract plot turning points, and the model then outputs the content or location of these turning points in the original text. Alternatively, the machine learning model can be used to divide the original content into multiple parts based on the plot, and the beginning or end of each part can be used as a plot turning point.

[0075] An exemplary method for generating the first audio is as follows. Generating the first audio corresponding to the target dimension based on the target chapter includes: extracting key content related to plot turning points from the target chapter; re-determining the order of multiple plot turning points; generating transition content based on the key content and the order; and generating the first audio based on the key content and the transition content.

[0076] Key content related to plot turning points can be the turning point itself, or it can be a description or summary of the preceding events, preserving the most exciting content while users are reading the original content. To enhance the appeal of highlight audio, plot turning points can be reordered, for example, placing more exciting points closer to the front. Then, use transitional content to connect the reordered key content.

[0077] For example, the plot of an audiobook is that character D meets character E. Three years later, D learns that E is missing. The content of the first audio generated can be as follows:

[0078] (Original) "What? E is missing? Impossible! She was here yesterday!"

[0079] (Narrator) D was shocked. In his mind, E couldn't possibly be someone who would disappear at will. The time came three years ago.

[0080] (Original) "Hello, D! My name is E, and I'll be your deskmate from now on!"

[0081] Below, the target dimension is used as an example for description.

[0082] An exemplary manner of determining the target dimension is as follows: According to the original content of the first audio book, determining the target dimension of the first audio book includes: in response to the first audio book being of the knowledge type, determining knowledge points as the target dimension of the first audio book.

[0083] Knowledge points can be important concepts, key interpretations, key sentences, and so on. Machine learning models can be used to extract knowledge points from the original content. Audiobooks that use knowledge points as a target dimension can include books on science, philosophy, psychology, and other genres.

[0084] An exemplary method for generating the first audio is as follows: Generating the first audio corresponding to the target dimension based on the target chapter includes: determining a knowledge point sentence from the target chapter; generating questions about the knowledge point sentence based on the knowledge point sentence as transition content; and generating the first audio based on the knowledge point sentence and the transition content.

[0085] When generating questions, a machine learning model can be used. For example, the extracted knowledge point sentences and processing instructions can be input into the model, where the processing instructions represent the questions that generated the knowledge point sentences.

[0086] An example using knowledge points as the target dimension is as follows. The generated first audio may include:

[0087] (Question) In future interstellar exploration, astronauts will need to communicate with remote command centers on Earth in real time. How can information be transmitted instantaneously, transcending distance limitations?

[0088] (Original) Quantum entanglement can realize ultra-long-distance information transmission and is theoretically not limited by distance.

[0089] The above describes several target dimensions and corresponding methods for generating the first audio. Those skilled in the art should be aware that, as needed, the above methods can be used alone or in combination. For example, for an audio book with complex character relationships and a dramatic plot, a variety of different highlight audios can be generated from several different target dimensions, such as the identity of the characters, the relationship status between multiple characters, and plot turning points. In addition, other target dimensions can also be used, which will not be elaborated in this disclosure.

[0090] When generating the first audio, the text of the first audio can be determined based on the target original text and the transitional content. This text can then be processed using speech-to-text technology to generate the first audio. For example, a text processing model can be used to process the text of the original content extracted from multiple chapters to generate the text of the transitional content. The original content text and the transitional content text can then be converted into audio to obtain the first audio. When generating the text of the transitional content, an instruction can be input to the text processing model to generate the text of the transitional content based on the original content. This text processing model can be, for example, a "text-generating text model" (a model that generates text based on text).

[0091] Alternatively, the audio of the transitional content can be directly generated from the audio of the target original text in the original audio, and then these audios can be spliced ​​together to generate the first audio. For example, an audio processing model can be used to process the audio of the original content extracted from multiple chapters to generate the audio of the transitional content; the first audio can be generated based on the audio of the original content and the audio of the transitional content. The audio processing model can be a model that generates audio from audio, that is, both the input and output of the model include audio. This allows the text step to be skipped and audio can be generated directly from audio.

[0092] The following describes two layout methods of the recommendation interface.

[0093] In one exemplary embodiment, the recommendation interface displays information about the currently playing audio in full screen. In some embodiments, in response to the recommendation interface loading a first audio, the first audio is played, wherein the first operation includes a sliding operation and the second operation is a triggering operation of a specified control.

[0094] Figure 3 Schematic diagram of the recommendation interface of some embodiments of the present disclosure is shown. Figure 3 As shown in part (a) of the figure, interface 3 displays information about the first audio in full screen, for example, player 31 for the first audio, which is the highlight audio of the first audio book. Player 31 may include the title of the first audio book, the text of the first audio, etc. In response to loading the first audio in interface 3, player 31 automatically begins playing the first audio.

[0095] In response to the user performing an upward sliding operation in interface 3, the interface may be switched to Figure 3 The interface shown in part (b) of FIG3 is a player 32 that displays the second audio in the recommended stream in the interface 3. The second audio can be the highlight audio of the second audio book or the original audio of the second audio book. In response to the second audio being loaded in the interface 3, the player 32 automatically starts playing the second audio. In response to the user's first operation, such as a sliding operation, the second audio can also be switched from the second audio to the original audio.

[0096] exist Figure 3 In the example embodiment, the player 31 may include a control “listen to original text” 310. In response to triggering the control 310, a playback interface (not shown) of the original content of the first audiobook may be displayed.

[0097] Through the above embodiment, the user can conveniently switch between different audios in the audio recommendation stream through a sliding operation.

[0098] Another exemplary embodiment is that the recommendation interface displays cards corresponding to multiple audio tracks. In some embodiments, the recommendation interface includes cards for multiple audio tracks, and playing a first audio track includes: playing the first audio track in response to triggering a play control on the card for the first audio track, the first operation being a triggering operation on the play control on the card for the second audio track, and the second operation being a triggering operation on an area on the card for the second audio track other than the play control.

[0099] Figure 4 Schematic diagrams of recommendation interfaces of other embodiments of the present disclosure are shown. Figure 4 As shown, interface 4 includes multiple audio cards. Each card can include information about the audiobook corresponding to the audio, such as the cover, title, and completion status. For example, card 41 corresponds to the first audio, and card 42 corresponds to the second audio. Each card includes audio information and playback controls. For example, card 41 includes a playback control "Listen to the Key Points" 410, and card 42 includes a playback control "Listen to the Key Points" 420.

[0100] Taking card 41 as an example, in response to the user triggering the play control 410, the first audio is played; in response to the user triggering the area in card 41 except where the control 410 is located, the play interface of the first audio book corresponding to the first audio is displayed.

[0101] Interface 4 can be an independent recommendation page in an application or website, or a sub-interface within the playback interface of an audiobook. Alternatively, it can be displayed when the audiobook a user is listening to is about to end (for example, when the proportion of the remaining unplayed content in the entire book is less than a threshold, or the remaining unplayed duration is less than a specified duration).

[0102] In the embodiments of the present disclosure, the second audio can be the original audio of the second audiobook, or audio generated based on the content of multiple chapters of the second audiobook (i.e., the highlight audio of the second audiobook). Therefore, the recommendation interface can be an interface dedicated to playing the highlight audio of each audiobook, or a recommendation interface that mixes the highlight audio and the original audio.

[0103] The above describes the methods of various embodiments of the present disclosure. The following describes the apparatuses for executing the methods of the various embodiments described above.

[0104] Figure 5 FIG. 1 shows a schematic diagram of the structure of an audio playback device according to some embodiments of the present disclosure. Figure 5As shown, the audio playback device 5 of this embodiment includes: a display module 51, configured to display an audio recommendation interface, the recommendation interface includes a first audio, wherein the first audio is generated based on the contents of multiple chapters of a first audio book, and includes original content extracted from the multiple chapters and transition content, and the transition content is generated based on the original content; a first playback module 52, configured to play the first audio; a second playback module 53, configured to play a second audio different from the first audio in the audio recommendation stream in response to a first operation on the recommendation interface; a third playback module 54, configured to play the original audio of the first audio book in response to a second operation on the recommendation interface.

[0105] In some embodiments, the audio playback device 5 further includes an audio generation module 55 .

[0106] In some embodiments, the audio generation module 55 is configured to: determine the target dimension of the first audio book based on the original content of the first audio book, the target dimension including at least one of the identity of a key character, the relationship status of multiple related characters, plot turning points, and knowledge points; determine multiple target chapters associated with the target dimension in the original content of the first audio book; and generate a first audio corresponding to the target dimension based on the target chapters.

[0107] In some embodiments, the audio generation module 55 is further configured to: determine a key character of the first audio book based on the original content of the first audio book; and determine the identity of the key character as a target dimension of the first audio book in response to the key character having different identities in multiple chapters.

[0108] In some embodiments, the audio generation module 55 is further configured to: extract multiple dialogues representing multiple identities of key characters from the target chapter; generate dialogue scene description information as transition content based on the multiple dialogues and the context of the multiple dialogues; and generate a first audio based on the multiple dialogues and the transition content.

[0109] In some embodiments, the audio generation module 55 is further configured to: determine multiple associated characters in the first audio book based on the original content of the first audio book; and in response to the multiple associated characters having different relationship statuses in multiple chapters, determine the relationship statuses of the multiple associated characters as the target dimension of the first audio book.

[0110] In some embodiments, the audio generation module 55 is further configured to: extract multiple dialogues representing multiple relationship states of multiple related characters from the target chapter; generate dialogue scene description information as transition content based on the multiple dialogues and the context of the multiple dialogues; determine the timbre of each character based on the relevant content of each character in the original content of the first audio book; and generate the first audio based on the dialogues, transition content and the timbre of each character.

[0111] In some embodiments, the audio generation module 55 is further configured to: in response to the first audio book including a plurality of plot turning points, determine the plot turning points as a target dimension of the first audio book.

[0112] In some embodiments, the audio generation module 55 is further configured to: extract key content related to the plot turning point from the target chapter; re-determine the order of multiple plot turning points; generate transition content based on the key content and the order; and generate a first audio based on the key content and the transition content.

[0113] In some embodiments, the audio generation module 55 is further configured to: in response to the first audio book being of the knowledge type, determine the knowledge points as the target dimensions of the first audio book.

[0114] In some embodiments, the audio generation module 55 is further configured to: determine knowledge point sentences from the target chapter; generate questions about the knowledge point sentences as transition content based on the knowledge point sentences; and generate a first audio based on the knowledge point sentences and the transition content.

[0115] In some embodiments, the audio generation module 55 is configured to: utilize a text processing model to process the text of the original content extracted from multiple chapters to generate the text of the transition content; and convert the text of the original content and the text of the transition content into audio to obtain a first audio.

[0116] In some embodiments, the audio generation module 55 is configured to: use an audio processing model to process the audio of the original content extracted from multiple chapters to generate audio of the transition content; and generate a first audio based on the audio of the original content and the audio of the transition content.

[0117] In some embodiments, the first playback module 52 is further configured to play the first audio in response to the recommendation interface loading the first audio, wherein the first operation includes a sliding operation, and the second operation is a triggering operation of a designated control.

[0118] In some embodiments, the recommendation interface includes cards for multiple audios, and the first playback module 52 is further configured to: play the first audio in response to triggering the playback control on the card of the first audio, the first operation is a triggering operation on the playback control on the card of the second audio, and the second operation is a triggering operation on the area on the card of the second audio except the playback control.

[0119] In some embodiments, the second audio is original audio of the second audio book, or audio generated based on contents of multiple chapters of the second audio book.

[0120] The present disclosure can quickly display highlight audio of multiple audio books to users through an audio recommendation interface. The highlight audio is generated based on the original content and transition content of multiple chapters of the audio book, so as to efficiently deliver the key content of the audio book to the user. Moreover, if the user is interested, the original audio of the audio book can be played through a second operation. Thus, the user does not need to jump back and forth between different interfaces in the process of finding a book. Through the first operation or the second operation, the user can switch to the highlight audio of other audio books, or start listening to the original audio directly. In this way, the client does not need to display a large number of irrelevant interfaces unnecessarily, and the user does not need to trigger multiple interfaces through complex operations, thereby saving the client's resources and improving the user's operating efficiency.

[0121] According to some embodiments of the present disclosure, an audio playback device is provided, comprising: a memory; and a processor coupled to the memory, wherein the processor is configured to execute the audio playback method of any embodiment described in the present disclosure based on instructions stored in the memory.

[0122] According to some embodiments of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the audio playback method of any embodiment described in the present disclosure is performed.

[0123] According to some embodiments of the present disclosure, a computer program product is provided. When the computer program product is run on a computer, the computer is enabled to implement the audio playback method of any embodiment described in the present disclosure.

[0124] Figure 6 A block diagram of an electronic device according to some embodiments of the present disclosure is shown.

[0125] Memory 61 is used to store one or more computer-readable instructions. Memory 61 may include any combination of various forms of computer-readable storage media, such as volatile and / or non-volatile memory, including but not limited to random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), read-only memory (ROM), and flash memory. Memory 61 may store, for example, an operating system, applications, a boot loader, databases, and other programs, as well as various applications and data.

[0126] The processor 62 is configured to execute computer-readable instructions to implement the method described in any of the aforementioned embodiments. Detailed implementations of the various steps of the method can be found in the aforementioned embodiments, and repetitive details are omitted here.

[0127] The processor 62 can be configured to execute the steps of the aforementioned embodiments. The processor 62 can be embodied as various processing devices, such as a central processing unit (CPU), a network processor (NP), or a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The central processing unit (CPU) can be based on an X86 or ARM architecture, for example.

[0128] The processor 62 and the memory 61 can communicate with each other directly or indirectly. For example, the processor 62 and the memory 61 can communicate via a network. The network can include a wireless network, a wired network, and / or any combination of wireless and wired networks. The processor 62 and the memory 61 can also communicate with each other via a system bus, which is not limited in this disclosure.

[0129] It should be noted that Figure 6 The components of the electronic device 6 shown are merely exemplary and non-limiting. The electronic device 6 may further include other components according to actual application requirements. The processor 62 may control the other components in the electronic device 6 to perform desired functions.

[0130] The electronic device 6 may be implemented by software, firmware and / or hardware, and may be integrated into a device installed with relevant application programs.

[0131] Figure 7 A block diagram of an electronic device according to some other embodiments of the present disclosure is shown.

[0132] Figure 7 The electronic device 7 shown may be a computer system with a dedicated hardware structure, which can execute corresponding functions when a relevant application program is installed.

[0133] Electronic devices include, but are not limited to, mobile terminals such as smartphones, laptops, personal digital assistants (PDAs), tablet personal computers (Tablet PCs), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), wearable devices, etc., as well as fixed terminals such as digital televisions and desktop computers, etc.

[0134] like Figure 7 As shown, the central processing unit (CPU) 71 executes various processes according to the program stored in the read-only memory (ROM) 72 or the program loaded from the storage part 78 to the random access memory (RAM) 73. In the RAM 73, data required when the CPU 71 executes various processes is stored as needed. The central processing unit is merely an example, and it can also be other types of processors, such as the various processors described above. The ROM 72, RAM 73 and the storage part 78 can be various forms of computer-readable storage media. It should be noted that although Figure 7 ROM 72, RAM 73 and storage portion 78 are shown separately in FIG, but one or more of them may be combined or located in the same or different memory or storage modules.

[0135] The CPU 71, the ROM 72, and the RAM 73 are connected to one another via a bus 74. To the bus 74, an input / output interface 75 is also connected.

[0136] The following components are connected to the input / output interface 75: an input portion 76 such as a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output portion 77 including a display such as a cathode ray tube (CRT), a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage portion 78 including a hard disk, a magnetic tape, etc.; and a communication portion 79 including a network interface card such as a LAN card, a modem, etc. The communication portion 79 allows communication processing to be performed via a network such as the Internet. It is easy to understand that although Figure 7 Parts of the electronic device 7 are shown to communicate via a bus 74, but they may also communicate via a network or other means, where the network may include a wireless network, a wired network, and / or any combination of a wireless network and a wired network.

[0137] A drive 710 is also connected to the input / output interface 75 as needed. A removable medium 711 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is mounted on the drive 710 as needed so that a computer program read therefrom is installed in the storage section 78 as needed.

[0138] When the series of processing described above is implemented by software, the program constituting the software can be installed from a network such as the Internet or a storage medium such as the removable medium 711 .

[0139] According to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, some embodiments of the present disclosure include a computer program product that, when the computer program product is run on a computer, causes the computer to implement the method described in any of the aforementioned embodiments. The computer program product includes computer instructions carried on a computer-readable medium, containing program code for executing the method shown in the flowchart. In such an embodiment, the computer instructions can be downloaded and installed from the network through the communication part 79, or installed from the storage part 78, or installed from the ROM 72. When the computer program is executed by the CPU 71, the method of the embodiment of the present disclosure is executed.

[0140] It should be noted that, in the context of the present disclosure, a computer-readable medium may be a tangible medium that may contain or store a program for use by an instruction execution system, apparatus, or device or for use in conjunction with an instruction execution system, apparatus, or device.

[0141] The computer readable medium may be a computer readable storage medium, or a computer readable signal medium, or any combination of the two.

[0142] Computer-readable storage media include, but are not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or device. The computer-readable storage medium stores computer instructions that, when executed by a processor, implement the method described in any of the aforementioned embodiments.

[0143] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, embodying computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. Program code embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wire, optical cable, RF (radio frequency), or any suitable combination thereof.

[0144] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0145] In some embodiments, a computer program is further provided, comprising: instructions, which, when executed by a processor, cause the processor to perform the method described in any of the aforementioned embodiments. For example, the instructions may be embodied as computer program codes.

[0146] In embodiments of the present disclosure, computer program code for performing the operations of the present disclosure may be written in one or more programming languages ​​or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0147] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0148] The functions described above may be at least partially performed by one or more hardware logic components. For example, and without limitation, exemplary hardware logic components that may be used include: a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), an application specific standard product (ASSP), a system on a chip (SOC), a complex programmable logic device (CPLD), and the like.

[0149] Although some specific embodiments of the present disclosure have been described in detail by way of examples, those skilled in the art will appreciate that the above examples are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Those skilled in the art will appreciate that modifications may be made to the above embodiments without departing from the scope and spirit of the present disclosure. The scope of the present disclosure is defined by the appended claims.

Claims

1. An audio playback method, comprising: Displaying an audio recommendation interface, the recommendation interface including a first audio, wherein the first audio is generated based on contents of a plurality of chapters of a first audiobook and includes original contents extracted from the plurality of chapters and transition contents generated based on the original contents; Playing the first audio; In response to a first operation on the recommendation interface, playing a second audio in the audio recommendation stream that is different from the first audio; In response to a second operation on the recommendation interface, the original audio of the first audio book is played.

2. The audio playback method according to claim 1, further comprising: determining, based on the original content of the first audio book, a target dimension of the first audio book, the target dimension comprising at least one of the identity of a key character, relationship status of multiple related characters, plot turning points, and knowledge points; determining a plurality of target chapters associated with the target dimension in the original content of the first audio book; A first audio corresponding to the target dimension is generated according to the target chapter.

3. The audio playback method according to claim 2, wherein: Determining the target dimension of the first audio book according to the original content of the first audio book includes: determining a key character of the first audio book based on the original content of the first audio book; In response to the key character having different identities in the multiple chapters, the identity of the key character is determined as a target dimension of the first audio book.

4. The audio playback method according to claim 3, wherein: Generating the first audio corresponding to the target dimension according to the target chapter includes: Extracting, from the target chapter, a plurality of dialogues representing the multiple identities of the key character; generating, according to the plurality of dialogues and the contexts of the plurality of dialogues, dialogue scene description information as the transition content; The first audio is generated based on the multiple dialogues and the transition content.

5. The audio playback method according to claim 2, wherein: Determining the target dimension of the first audio book according to the original content of the first audio book includes: determining, based on original content of the first audio book, a plurality of associated characters in the first audio book; In response to the plurality of associated characters having different relationship statuses in the plurality of chapters, the relationship statuses of the plurality of associated characters are determined as a target dimension of the first audio book.

6. The audio playback method according to claim 5, wherein: Generating the first audio corresponding to the target dimension according to the target chapter includes: Extracting, from the target chapter, a plurality of dialogues representing a plurality of relationship states of the plurality of associated characters; generating, according to the plurality of dialogues and the contexts of the plurality of dialogues, dialogue scene description information as the transition content; determining a timbre for each character based on relevant content of each character in the original content of the first audio book; The first audio is generated based on the dialogue, the transition content, and the timbre of each character.

7. The audio playback method according to claim 2, wherein: Determining the target dimension of the first audio book according to the original content of the first audio book includes: In response to the first audio book including a plurality of plot turning points, the plot turning points are determined as target dimensions of the first audio book.

8. The audio playback method according to claim 7, wherein: Generating the first audio corresponding to the target dimension according to the target chapter includes: extracting key content related to the plot turning point from the target chapter; Re-determining the order of the plurality of plot turning points; generating the transition content according to the key content and the sequence; The first audio is generated according to the key content and the transition content.

9. The audio playback method according to claim 2, wherein: Determining the target dimension of the first audio book according to the original content of the first audio book includes: In response to the first audio book being of the knowledge type, knowledge points are determined as target dimensions of the first audio book.

10. The audio playback method according to claim 9, wherein: Generating the first audio corresponding to the target dimension according to the target chapter includes: Determining a knowledge point sentence from the target chapter; Based on the knowledge point sentence, generating a question about the knowledge point sentence as the transition content; The first audio is generated according to the knowledge point sentence and the transition content.

11. The audio playback method according to any one of claims 1 to 10, further comprising: Using a text processing model, processing the text of the original content extracted from the plurality of chapters to generate the text of the transitional content; The text of the original content and the text of the transition content are converted into audio to obtain the first audio.

12. The audio playback method according to any one of claims 1 to 10, further comprising: Using an audio processing model, processing the audio of the original content extracted from the plurality of chapters to generate the audio of the transition content; The first audio is generated according to the audio of the original content and the audio of the transition content.

13. The audio playback method according to any one of claims 1 to 10, wherein: Playing the first audio includes: In response to the recommendation interface loading the first audio, the first audio is played, wherein the first operation includes a sliding operation, and the second operation is a triggering operation of a designated control.

14. The audio playback method according to any one of claims 1 to 10, wherein: The recommendation interface includes a plurality of audio cards, and playing the first audio includes: In response to triggering the play control on the card of the first audio, the first audio is played, the first operation is a triggering operation of the play control on the card of the second audio, and the second operation is a triggering operation of the area on the card of the second audio except the play control.

15. The audio playback method according to any one of claims 1 to 10, wherein: The second audio is original audio of the second audio book, or audio generated according to contents of multiple chapters of the second audio book.

16. An audio playback device, comprising: a display module configured to display an audio recommendation interface, the recommendation interface including a first audio, wherein the first audio is generated based on contents of a plurality of chapters of a first audio book and includes original contents extracted from the plurality of chapters and transition contents, the transition contents being generated based on the original contents; A first playing module is configured to play the first audio; a second playing module configured to play a second audio different from the first audio in the audio recommendation stream in response to a first operation on the recommendation interface; The third playing module is configured to play the original audio of the first audio book in response to a second operation on the recommendation interface.

17. An audio playback device, comprising: Memory; as well as A processor coupled to the memory, wherein the processor is configured to execute the audio playback method according to any one of claims 1 to 15 based on instructions stored in the memory.

18. A computer-readable storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the audio playback method according to any one of claims 1 to 15 is implemented.

19. A computer program product, which, when executed on a computer, enables the computer to implement the audio playback method according to any one of claims 1 to 15.

Citation Information

Patent Citations

  • Linkage display method for audio book playing progress in page and electronic equipment

    CN110209320A

  • Information recommendation method and device, electronic equipment and storage medium

    CN117311574A

  • Video generation method and device, computer equipment and storage medium

    CN117544834A

  • Book information display method and device, electronic equipment and readable storage medium

    CN117763138A

  • Video generation method and apparatus, computer device, and storage medium

    US20250173917A1