Audio program content playback control method, device, equipment and storage medium
By generating and playing the summary content of the continuous listening content in the audio program content sharing platform, the problem of users forgetting the prospect story after a long pause is solved, and a better understanding of the storyline and narrative rhythm of the subsequent paragraphs is achieved, and the listening efficiency of the audio program content is improved.
Patent Information
- Application Number
- CN202110541007.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-05-18
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2041-05-18
AI Technical Summary
In the audio program content sharing platform, when users pause for a long time and do not continue listening, the foreground story is forgotten, affecting the understanding of the plot and narrative rhythm of subsequent paragraphs.
Provides a playback control method for audio program content, which helps users review the core ideas of the played part and undertake it when listening to the continuous listening. The method includes displaying the listening control area when the user resumes playback, and playing the corresponding listening summary content. The summary content is the summary information generated based on the audio content of the played part.
By intelligently generating a review summary and converting it into audio playback, users can help recall the content of the program they have heard, enhance their understanding of the content of the continuous listening, reduce the repeated replay due to forgetting, and improve the listening efficiency of the audio program content.
Smart Images

Figure CN113761268B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to the field of machine learning technology, and provides a playback control method, device, equipment and storage medium for audio program content. Background Art
[0002] With the rapid development of Internet technology, various social software has emerged. Among them, audio program content sharing platforms are attracting more and more attention and popularity. Podcast platforms are a very common audio program content sharing platform. Many netizens like to record and share audio programs through podcast platforms, or listen to audio programs shared by other users, including audio novels, storytelling, crosstalk, talk shows, etc.
[0003] However, in various audio program content sharing platforms in the related art, when playing audio program content, if the user pauses the program for a long time and does not resume listening to the program for a long time, when resuming listening, directly continuing to listen to the program at the position where it was paused last time, it will cause the user to be unable to connect well with the foreground story that has been forgotten, resulting in the user's thinking being unable to quickly keep up with the storyline and narrative rhythm of the subsequent paragraphs of the program, and causing varying degrees of forgetfulness of the content that has been listened to. Summary of the invention
[0004] Embodiments of the present application provide a method, apparatus, device, and storage medium for controlling the playback of audio program content, so as to improve the listening efficiency of audio program content.
[0005] The first audio program content playback control method provided in the embodiment of the present application includes:
[0006] In response to a pause operation triggered on the target audio program content, pausing the playing of the target audio program content;
[0007] In response to a resume operation triggered for the target audio program content, a resume listening control area is displayed in the playback control interface, and resume listening summary content corresponding to the target audio program content is played, wherein the resume listening summary content is summary information generated for audio content corresponding to the played portion of the target audio program content;
[0008] After the playback of the resume listening summary content is finished, the unplayed portion of the target audio program content continues to be played.
[0009] The second method for controlling the playback of audio program content provided in the embodiment of the present application includes:
[0010] After receiving a pause request for the target audio program content sent by the client, the corresponding pause time is recorded;
[0011] After receiving the recovery request for the target audio program content sent by the client, recording the corresponding continued listening time;
[0012] Based on the time interval between the pause time and the resume listening time, resume listening summary content for the target audio program content is generated, and the resume listening summary content is fed back to the client, so that the client displays a resume listening control area in a playback control interface and plays the resume listening summary content, wherein the resume listening summary content is summary information generated for the audio content corresponding to the played portion of the target audio program content.
[0013] The first audio program content playback control device provided in the embodiment of the present application includes:
[0014] a pause unit, configured to pause playing of the target audio program content in response to a pause operation triggered on the target audio program content;
[0015] A resume unit is used to respond to a resume operation triggered for the target audio program content, display a resume control area in a playback control interface, and play resume summary content corresponding to the target audio program content, wherein the resume summary content is summary information generated for the audio content corresponding to the played portion of the target audio program content; after the resume summary content is played, continue to play the unplayed portion of the target audio program content.
[0016] Optionally, the continued listening control area includes a summary control widget, and the continued playing unit is further used for:
[0017] Before the playback of the resumed listening summary content ends, in response to a closing operation triggered on the summary control widget, the playback of the resumed listening summary content is closed, and the audio content corresponding to the unplayed portion of the target audio program content continues to be played.
[0018] Optionally, the device further comprises:
[0019] A setting unit is used to set the replay permission for the target object in response to the setting operation of the replay permission control in the permission setting interface before the replay unit responds to the playback operation of resuming the playback of the target audio program content, and send the corresponding replay permission information to the server, so that the server associates the replay permission information with the identification information of the target object and saves it.
[0020] Optionally, the replay unit is further used for:
[0021] In response to the recovery operation triggered by the target object for the target audio program content, if it is determined that the target object has the resume play permission based on the resume play permission information associated with the target object, a resume listening control area is displayed in the playback control interface, and the resume listening summary content corresponding to the target audio program content is played.
[0022] Optionally, the resume unit is further configured to determine the resume listening summary content in the following manner:
[0023] Determine the corresponding review duration based on the time interval between the pause time corresponding to the pause operation and the continued listening time corresponding to the resume operation, and the played duration corresponding to the played portion of the target audio program content;
[0024] Based on the review duration, selecting a segment of audio content with a playback duration equal to the review duration from the audio content corresponding to the played portion as the audio content to be reviewed;
[0025] Convert the audio content to be reviewed into text information, and generate a summary content text for the text information based on a text summarization technology;
[0026] The summary content text is converted into audio to obtain the continued listening summary content.
[0027] Optionally, the replay unit is specifically used for:
[0028] Determine a corresponding first review duration based on the time interval and the played duration corresponding to the played portion of the target audio program content;
[0029] Determining a corresponding second review duration based on the content difficulty level corresponding to the target program content, wherein the greater the content difficulty level, the longer the second review duration;
[0030] The sum of the first review duration and the second review duration is used as the corresponding review duration.
[0031] Optionally, the replay unit is specifically used for:
[0032] If the target audio program content includes a sound of an object, based on the object sound, the summary content text is converted into audio to obtain the continued listening summary content;
[0033] If the target audio program content includes sounds of multiple objects, the sound with the highest proportion is determined by extracting features from the sounds of the multiple objects, and based on the sound with the highest proportion, the summary content text is converted into audio to obtain the continued listening summary content.
[0034] The second audio program content playback control device provided in the embodiment of the present application includes:
[0035] A first recording unit, configured to record a corresponding pause time after receiving a pause request for a target audio program content sent by a client;
[0036] A second recording unit, configured to record a corresponding continued listening time after receiving a recovery request for the target audio program content sent by the client;
[0037] A feedback unit is used to generate a resume listening summary content for the target audio program content based on the time interval between the pause time and the resume listening time, and feed back the resume listening summary content to the client, so that the client displays a resume listening control area in a playback control interface and plays the resume listening summary content, wherein the resume listening summary content is summary information generated for the audio content corresponding to the played portion of the target audio program content.
[0038] Optionally, the device further comprises:
[0039] A determination unit is used to determine whether the target audio program content satisfies at least one of the following target conditions:
[0040] The played time corresponding to the played part of the target audio program content is not less than the first time threshold;
[0041] The time interval between the pause time and the continued listening time corresponding to the target audio program content is not less than a second duration threshold.
[0042] Optionally, the feedback unit is specifically used for:
[0043] Determine a corresponding review duration based on the time interval and the played duration corresponding to the played portion of the target audio program content;
[0044] Based on the review duration, selecting a segment of audio content with a playback duration equal to the review duration from the audio content corresponding to the played portion as the audio content to be reviewed;
[0045] Convert the audio content to be reviewed into text information, and generate a summary content text for the text information based on a text summarization technology;
[0046] The summary content text is converted into audio to obtain the continued listening summary content.
[0047] Optionally, the feedback unit is specifically used for:
[0048] If the time interval is not greater than the preset interval threshold, the product of the broadcast duration and the first preset ratio value is used as the review duration;
[0049] If the time interval is greater than the preset interval threshold, then each time the time interval increases by a set duration, the first preset ratio value is increased by a first set step to obtain a first ratio value, and the product of the played duration and the first ratio value is used as the review duration.
[0050] Optionally, the feedback unit is specifically used for:
[0051] Determine a corresponding first review duration based on the time interval and the played duration corresponding to the played portion of the target audio program content;
[0052] Determining a corresponding second review duration based on the content difficulty level corresponding to the target program content, wherein the greater the content difficulty level, the longer the second review duration;
[0053] The sum of the first review duration and the second review duration is used as the corresponding review duration.
[0054] Optionally, the feedback unit is specifically used for:
[0055] If the time interval is not greater than the preset interval threshold, the product of the broadcast duration and the second preset ratio value is used as the first review duration;
[0056] If the time interval is greater than the preset interval threshold, then every time the time interval increases by a set time, the second preset ratio value is increased by a second set step to obtain a second ratio value, and the product of the broadcast time and the second ratio value is used as the first review time.
[0057] Optionally, the feedback unit is specifically used for:
[0058] If the difficulty level of the content is not greater than the preset level threshold, the product of the broadcast time and the third preset ratio value is used as the second review time;
[0059] If the content difficulty level is greater than the preset level threshold, the third preset ratio value is increased by a third preset step size every time the content difficulty level increases by a set level to obtain a third ratio value, and the product of the broadcast time and the third ratio value is used as the second review time.
[0060] Optionally, the feedback unit is specifically used for:
[0061] If the target audio program content includes a sound of an object, based on the object sound, the summary content text is converted into audio to obtain the continued listening summary content;
[0062] If the target audio program content includes sounds of multiple objects, the sound with the highest proportion is determined by extracting features from the sounds of the multiple objects, and based on the sound with the highest proportion, the summary content text is converted into audio to obtain the continued listening summary content.
[0063] Optionally, the device further comprises:
[0064] The associating unit is used to obtain the replay permission information associated with the target object after receiving the setting request for the replay permission control in the permission setting interface sent by the client, and associate the replay permission information with the identification information of the target object for storage.
[0065] An electronic device provided in an embodiment of the present application includes a processor and a memory, wherein the memory stores program code, and when the program code is executed by the processor, the processor executes the steps of any one of the above-mentioned methods for controlling the playback of audio program content.
[0066] The embodiment of the present application provides a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device performs the steps of any of the above-mentioned audio program content playback control methods.
[0067] An embodiment of the present application provides a computer-readable storage medium, which includes a program code. When the program product is run on an electronic device, the program code is used to enable the electronic device to execute the steps of any one of the above-mentioned methods for controlling the playback of audio program content.
[0068] The beneficial effects of this application are as follows:
[0069] The embodiment of the present application provides a playback control method, device, equipment and storage medium for audio program content. Since the present application supports intelligent generation of a review summary when a user clicks to continue listening to audio, and converts it into audio for playback, it can help users review the core ideas of the program content they have listened to before, and then connect with the content they continue listening to, enhance the user's understanding of the program content, reduce the situation where users repeatedly replay the audio program content because they have forgotten it, and improve the listening efficiency of the audio program content.
[0070] Other features and advantages of the present application will be described in the following description, and partly become apparent from the description, or be understood by practicing the present application. The purpose and other advantages of the present application can be realized and obtained by the structures specifically pointed out in the written description, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0072] Figure 1 A schematic diagram of a method for resuming broadcasting an audio program in the related art;
[0073] Figure 2 An optional schematic diagram of an application scenario in an embodiment of the present application;
[0074] Figure 3 This is a flowchart of the implementation of the first audio program content playback control method in the embodiment of the present application;
[0075] Figure 4 A schematic diagram of a playback control interface in an embodiment of the present application;
[0076] Figure 5 A schematic diagram of another playback control interface in an embodiment of the present application;
[0077] Figure 6 A schematic diagram of a permission setting interface in an embodiment of the present application;
[0078] Figure 7 A schematic diagram of a method for resuming playing an audio program in an embodiment of the present application;
[0079] Figure 8 is a flowchart of the implementation of the second method for controlling the playback of audio program content in the embodiment of the present application;
[0080] Fig. 9 A flow chart of a method for generating summary content of continued listening in an embodiment of the present application;
[0081] Fig.10 A schematic diagram of a language recognition process in an embodiment of the present application;
[0082] Fig.11 A schematic diagram of a model structure in an embodiment of the present application;
[0083] Fig. 12A A schematic diagram of a parameter-based speech synthesis process in an embodiment of the present application;
[0084] Fig. 12B A schematic diagram of a specific process of text analysis in an embodiment of the present application;
[0085] Fig.13A This is a flow chart of a method for controlling the playback of audio program content based on a client and a server in an embodiment of the present application;
[0086] Fig. 13B This is a timing diagram of interaction between a client and a server in an embodiment of the present application;
[0087] Fig.14 It is a schematic diagram of the composition structure of the first audio program content playback control device in the embodiment of the present application;
[0088] Fig.15 It is a schematic diagram of the composition structure of the second audio program content playback control device in the embodiment of the present application;
[0089] Fig.16 A schematic diagram of a hardware structure of an electronic device to which an embodiment of the present application is applied;
[0090] Fig.17 A schematic diagram of the hardware structure of another electronic device to which the embodiments of the present application are applied. DETAILED DESCRIPTION
[0091] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the technical solution of the present application, rather than all of the embodiments. Based on the embodiments recorded in the application documents, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the technical solution of the present application.
[0092] The following is an introduction to some concepts involved in the embodiments of the present application.
[0093] Audio products: The audio products covered in this application include audiobooks, podcasts, and other products that deliver content only through the interactive form of sound broadcasting.
[0094] Program: refers to the unit of content delivered in an interactive form of sound broadcast in audio products.
[0095] Audio program content: The audio program content in the embodiments of the present application refers to audio programs shared on instant messaging software or podcast platforms, such as audio novels, crosstalk, storytelling, talk shows, radio stations, etc. Among them, audio novels refer to general audio program content files. When it is played in the player, the playback speed can be adjusted, and the playback stop time can be automatically remembered to facilitate "reading". The audio program content in the embodiments of the present application may refer to audio content containing audio data (referring to text that can be obtained through voice recognition, not pure music).
[0096] Resume listening: also known as resume broadcast, refers to when a user interrupts listening to a certain episode of a program and then clicks play again to continue listening.
[0097] Client or user end: refers to the program that corresponds to the server and provides local services to customers. Except for some applications that only run locally, they are generally installed on ordinary client computers and need to cooperate with the server to run. After the development of the Internet, the more commonly used user ends include web browsers used for the World Wide Web, email clients for sending and receiving emails, and client software for instant messaging. For this type of application, there needs to be a corresponding server and service program in the network to provide corresponding services, such as database services, email services, etc. In this way, a specific communication connection needs to be established between the client and the server to ensure the normal operation of the application.
[0098] Application operation interface: It is the medium for interaction and information exchange between the application system and the user. It realizes the conversion between the internal form of information and the form acceptable to humans. The purpose is to enable users to operate the application conveniently and efficiently to achieve two-way interaction and complete the work they want to complete with the help of the application. In the embodiment of the present application, the application operation interface includes human-computer interaction and graphical user interface. The specific application operation interface includes permission setting interface, playback control interface, etc. Among them, different application operation interfaces are used to display different content to users and realize different information interactions between users and applications.
[0099] Audio program content sharing platform and podcast: Audio program content sharing platform is a kind of digital broadcasting technology, which can be used to record the audio program content of online broadcast or similar online audio programs. Netizens can download online radio programs to their own players to listen to them on the go, without having to sit in front of a computer or listen in real time, enjoying the freedom of anytime and anywhere. In addition, users can also make their own audio programs and upload them to the Internet through the podcast platform to share with netizens. It can be understood as a client that plays audio program content and video. At present, there are many applications for audio program content sharing platforms, such as podcasts.
[0100] Playback control page: A user-oriented page used to control the playback of audio program content. The number of playback control pages set on an audio program content sharing platform may be one or more as needed, and multiple playback control pages are jumped according to the set logic. In the embodiment of the present application, the playback control page mainly refers to the page used to control the playback of audio program content, including a playback control area and a document display area, wherein the playback control area is mainly used to control the playback of audio program content, including the control of playback speed, playback progress, etc.; the document display area is mainly used to display the document content corresponding to the audio content in the currently playing audio program content. In addition, the playback control page may also include a video playback area, a document overview area, etc.
[0101] Text-To-Speech (TTS) is a technology that generates artificial speech through mechanical and electronic methods. TTS technology (also known as text-to-speech technology) belongs to speech synthesis. It is a technology that converts text information generated by the computer itself or input from the outside into understandable and fluent Chinese spoken output.
[0102] Artificial intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making.
[0103] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics and other technologies. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0104] Machine Learning (ML) is a multi-disciplinary subject that involves probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. It specializes in studying how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all areas of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning.
[0105] With the research and advancement of artificial intelligence technology, artificial intelligence technology has been studied and applied in many fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, unmanned driving, automatic driving, drones, robots, smart medical care, smart customer service, etc. I believe that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.
[0106] The solution provided in the embodiments of this application involves technologies such as artificial intelligence and machine learning. The methods of acoustic models, language models, generative neural network models, etc. proposed in the embodiments of this application can be divided into two parts, including a training part and an application part; among them, the training part involves the technical field of machine learning. In the training part, the above-mentioned model is trained by machine learning technology, and the model parameters are continuously adjusted by optimization algorithms; the application part is used to use the acoustic model, speech model, etc. trained in the training part for speech recognition, and the generative neural network model trained in the training part is used to generate summary content, etc.
[0107] The following is a brief introduction to the design concept of the embodiment of the present application:
[0108] When voice is the only input channel, the user's efficiency in receiving information is much lower than that of multi-modal interactive input methods such as voice, vision, and touch. In audio programs, podcasts are mostly 1 to 3 hours long, and audiobooks are tens of hours long, which makes it impossible for most users to listen to the entire program in one go. When users do not continue listening to a program for a long time, continuing to listen directly at the last pause point will result in the user being unable to connect well with the forgotten foreground story, and the user's thinking will be unable to quickly keep up with the storyline and narrative rhythm of the subsequent sections of the program.
[0109] That is to say, the audio product experience in the related art is that when the user pauses the listening process in the middle and clicks the play button again, the listening starts from the last paused position, such as Figure 1As shown, it is a schematic diagram of a method for resuming playing an audio program in the related art. However, when a user pauses a program for a long time, he or she will forget the content that has been listened to to varying degrees.
[0110] In view of this, the embodiments of the present application propose a playback control method, device, equipment and storage medium for audio program content. The present application supports intelligent generation of a review summary when a user clicks to continue listening to audio, and converts it into audio for playback, which can help users review the core ideas of the program content they have listened to before, and then connect with the content they continue listening to, enhance the user's understanding of the program content, reduce the situation where users repeatedly replay the audio program content because they have forgotten it, and improve the listening efficiency of the audio program content.
[0111] The preferred embodiments of the present application are described below in conjunction with the drawings in the specification. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application, and are not used to limit the present application. In addition, the embodiments and features in the embodiments of the present application may be combined with each other if there is no conflict.
[0112] like Figure 2 As shown, it is a schematic diagram of an application scenario of an embodiment of the present application. The application scenario diagram includes two terminal devices 210 and a server 230, and the terminal device 210 can log in to the relevant interface 220 of the target business execution. The terminal device 210 and the server 230 can communicate through a communication network.
[0113] In the embodiment of the present application, the interface 220 may be a playback control interface, a permission setting interface, etc. The user may log in to the interface 220 through the terminal device 210, the terminal device 210 responds to the operation triggered by the user on the interface 220, and sends a relevant request to the server 230, the server 230 feeds back relevant information to the terminal device, etc. For example: the terminal device 210 responds to the recovery operation triggered for the target audio program content, sends a recovery request to the server 230, the server 230 also generates a resume summary content based on the request, and feeds back to the terminal device 210, the terminal device 210 displays the resume control area in the playback control interface, and plays the resume summary content corresponding to the target audio program content, etc., which will not be listed one by one here, and will be described in detail below.
[0114] In an optional implementation, the communication network is a wired network or a wireless network.
[0115] In the embodiment of the present application, the terminal device 210 is an electronic device used by the user, which can be a personal computer, a mobile phone, a tablet computer, a notebook, an e-book reader, or other computer devices with certain computing capabilities and running instant messaging software and websites or social software and websites. Each terminal device 210 is connected to the server 230 via a wireless network, and the server 230 is a server cluster or a cloud computing center composed of one server or several servers, or a virtualization platform.
[0116] In the embodiment of the present application, a client related to the audio program content is installed on the terminal device 210, and the client can be software, such as instant messaging software, podcast software, or a small program, a web page, etc., which is not specifically limited here. Correspondingly, the server is a server corresponding to the software, web page, small program, etc.
[0117] Among them, users can directly search for and play the audio program content they like to listen to through podcast software, or listen to the audio program content shared by friends in software such as instant messaging, or search or listen to the audio program content in public accounts, mini-programs, etc. It should be noted that the audio program content in the embodiment of the present application refers to the audio recorded by the user. For example, the user narrates each chapter in a novel and records the corresponding audio files. After that, the user shares the recorded audio files to the podcast platform for everyone to listen to, that is, listening to the book. In this scenario, the audio program content refers to the audio recorded by the user. The user who listens to the audio program content can play the corresponding summary content of the audio program content in the playback control page through the method in the embodiment of the present application. The summary content of the continued listening is generated by the client or the server for the audio content corresponding to the played part of the audio program content uploaded by the user. In addition, in addition to novels, it can also be crosstalk, storytelling, etc. recorded by the user, which is not specifically limited here.
[0118] It should be noted that Figure 2 The figure is only an example. In fact, the number of terminal devices and servers is not limited and is not specifically limited in the embodiments of the present application.
[0119] See also Figure 3 As shown, it is an implementation flow chart of the first audio program content playback control method provided in the embodiment of the present application, which is applied to a terminal device. The specific implementation process of the method is as follows:
[0120] S31: the terminal device pauses playing the target audio program content in response to a pause operation triggered on the target audio program content;
[0121] S32: the terminal device displays a resume control area in the playback control interface in response to the resume operation triggered for the target audio program content, and plays the resume summary content corresponding to the target audio program content, wherein the resume summary content is summary information generated for the audio content corresponding to the played portion of the target audio program content;
[0122] S33: After the playback of the resume listening summary content is finished, the terminal device continues to play the unplayed portion of the target audio program content.
[0123] like Figure 4 As shown in FIG. 4 , it is a schematic diagram of a playback control interface in an embodiment of the present application, and both interface 41 and interface 42 in the figure are playback control interfaces. The user can trigger a pause operation or a resume operation for the audio program content by clicking the pause / play control S410 in interface S41. When the control is in the state shown in interface 41, it indicates that the target audio program content is currently paused. When the control is in the state shown in interface 42, it indicates that the target audio program content is currently resumed.
[0124] As shown in the dashed box S420 in the interface 42, which is the continue listening control area in the embodiment of the present application, intelligent continue listening is being performed, that is, the continue listening summary content is played, and the continue listening summary content is generated based on the played part (the part before 22:22). After the continue listening summary content is played, the target audio program content can be played normally.
[0125] like Figure 5 As shown, it is a schematic diagram of another playback control interface in an embodiment of the present application, indicating that after the resume listening summary content is played, the resume listening control area S420 will no longer be displayed, and the target audio program content will continue to be played.
[0126] In the above implementation mode, it supports intelligent generation of a review summary when the user clicks on the button to continue listening to the audio, and converts it into audio for playback, which can help the user review the core ideas of the program content that he has listened to before, and then connect it with the content that he continues to listen to, thereby enhancing the user's understanding of the program content, reducing the situation where the user repeatedly replays the audio program content due to forgetting the audio program content that he has listened to, improving the listening efficiency of the audio program content, and optimizing the user experience of audio products.
[0127] In an optional implementation, the continue listening control area includes a summary control widget. Figure 4 As shown in the middle interface 42, the "skip" in S420 is a summary control control in the embodiment of the present application. The user can click "skip" to end the playback of the summary content.
[0128] Specifically, before the playback of the resumed summary content ends, if the user clicks "skip", the terminal device responds to the closing operation triggered by the summary control, closes the playback of the resumed summary content, and continues to play the audio content corresponding to the unplayed part of the target audio program content, and displays the following: Figure 5 The playback control interface shown no longer displays the resume listening control area, directly skips the playback of the resume listening summary content, and continues to play the target audio program content, that is, continues to play from 22:22.
[0129] In the above implementation, the user can skip the playback of the resumed summary content based on the summary control widget. In addition, the user can also adjust the playback speed of the resumed summary content by doubling the speed, etc., to improve the playback efficiency of the audio program content.
[0130] Optionally, the embodiment of the present application also supports turning on the "smart resume listening" function when the user clicks to continue listening to audio. When the user turns on this function, after clicking to continue listening, the resume listening control area is displayed in the playback control interface, and the resume listening summary content is played.
[0131] For example Figure 6 As shown, it is a permission setting interface in an embodiment of the present application, wherein the dotted box S60 is a resume permission control. The user can turn on or off "smart resume listening" by clicking on the control. Figure 6 The display shows that "Smart Resume Listening" is turned on.
[0132] When the user clicks Figure 6 When the resume play permission control shown turns on "smart resume listening", the terminal device responds to the setting operation of the resume play permission control in the permission setting interface, sets the resume play permission for the target object, and sends the corresponding resume play permission information to the server, so that the server associates the resume play permission information with the identification information of the target object and saves it.
[0133] In the embodiment of the present application, when setting the resume permission control, the user can turn on or off "smart resume listening". Therefore, in an optional implementation, when the terminal device responds to the recovery operation triggered by the target object (referring to the user or user account) for the target audio program content, it is necessary to further determine whether the target object has the resume permission. When the user turns on "smart resume listening", the target object has the resume permission, otherwise, the target object does not have the resume permission.
[0134] Specifically, if it is determined that the target object has the resume play permission according to the resume play permission information associated with the target object, a resume listening control area is displayed in the play control interface, and the resume listening summary content corresponding to the target audio program content is played.
[0135] The replay permission information associated with the target object may be stored locally on the terminal device when the user sets the permission, or may be requested by the terminal device to the server and returned by the server.
[0136] In the above implementation, by adding an "intelligent continue listening" switch, after turning it on, users can enjoy the function of intelligently generating a review summary, which can effectively improve the product user experience.
[0137] See also Figure 7 As shown, it is a schematic diagram of a method for resuming playing an audio program in an embodiment of the present application. Figure 1 Compared with the schematic diagram of the audio program resuming method in the related art shown, the present application provides an “intelligent resume listening” function. When the “intelligent resume listening” function is turned on, after pausing the program, when you click the play button again, it does not directly resume playback at the pause position, but generates intelligent review content audio (i.e., resume listening summary content), and plays the review content, and then resumes playback at the pause position to avoid users forgetting the content they have listened to after pausing the program for a long time. The present application can help users review the content of programs they have listened to before and enhance their understanding of the content they continue listening to.
[0138] See also Figure 8 As shown, it is an implementation flow chart of the second audio program content playback control method provided in the embodiment of the present application, which is applied to the server. The specific implementation process of the method is as follows:
[0139] S81: After receiving the pause request for the target audio program content sent by the client, the server records the corresponding pause time;
[0140] S82: After receiving the recovery request for the target audio program content sent by the client, the server records the corresponding continued listening time;
[0141] S83: The server generates a resume listening summary content for the target audio program content based on the time interval between the pause time and the resume listening time, and feeds back the resume listening summary content to the client, so that the client displays the resume listening control area in the playback control interface and plays the resume listening summary content, wherein the resume listening summary content is summary information generated for the audio content corresponding to the played portion of the target audio program content.
[0142] Among them, the client is installed on the terminal device, and the communication between the client and the server is the communication between the terminal device and the server.
[0143] In an embodiment of the present application, when the user clicks to pause playback, the terminal responds to the pause operation triggered for the target audio program content, pauses the playback of the target audio program content, and sends a pause request to the server, which may carry the corresponding pause time. When the server receives the pause request, it records the pause time, for example, recorded as T1. Similarly, when the user clicks to resume playback, the terminal responds to the resume operation triggered for the target audio program content, resumes the playback of the target audio program content, and sends a resume request to the server, which may carry the corresponding resume time. When the server receives the resume request, it records the resume time, for example, recorded as T2. Furthermore, based on the time interval y=T2-T1 between the pause time and the resume listening time, a resume listening summary content is generated.
[0144] In an optional implementation, before generating the continued listening summary content for the target audio program content based on the time interval between the pause time and the continued listening time, it is necessary to further determine whether the conditions for generating the continued listening summary content are met. Only when the conditions are met can the continued listening summary content for the target audio program content be generated based on the time interval between the pause time and the continued listening time.
[0145] Specifically, the target conditions include at least one of the following:
[0146] Condition 1: the played duration of the played portion of the target audio program content is not less than the first duration threshold;
[0147] Condition 2: The time interval between the pause time and the resume listening time corresponding to the target audio program content is not less than the second duration threshold.
[0148] That is, when the target audio program content satisfies at least one of the above target conditions, the condition for generating the continued listening summary content is satisfied.
[0149] Specifically, when the user turns on the "Smart Resume Listening" function, if the user pauses listening during the listening process and then clicks the play button again, the client will upload the corresponding broadcast time of the target audio program content to the server, and store the time when the user pauses and resumes playing the same program in the background and upload it to the server as a basis for judgment.
[0150] Assume that the first duration threshold is 2 minutes and the second duration threshold is 5 hours. Then, the time information T1 and T2 when the user pauses and resumes the same program is uploaded to the server, and the server calculates the time interval between T1 and T2. When the program has been played for less than 2 minutes, the content summary is not generated; or, when the listening time interval y is within 5 hours, the content summary is not generated; or, when the program has been played for less than 2 minutes and the listening time interval is within 5 hours, the content summary is not generated.
[0151] It should be noted that the above is only an example. Of course, the above judgment may not be performed and the content summary may be directly generated. No specific limitation is made here.
[0152] For the situation where the summary content of the continued listening needs to be generated, an optional implementation method is as follows: Fig. 9 The flowchart shown implements S83, which is a flowchart of a method for generating a summary of continued listening content in an embodiment of the present application, including the following steps:
[0153] S901: The server determines a corresponding review duration based on the time interval between the pause time and the resume listening time, and the played duration of the played portion of the target audio program content;
[0154] S902: The server selects, based on the review duration, a segment of audio content with a playback duration equal to the review duration from the audio content corresponding to the played portion as the audio content to be reviewed;
[0155] S903: The server converts the audio content to be reviewed into text information, and generates a summary content text for the text information based on a text summarization technology;
[0156] S904: The server converts the summary content text into audio to obtain the summary content for continued listening.
[0157] That is, in step S902, it is necessary to determine the paragraph range that needs to be reviewed from the played audio content as the audio content to be reviewed. By extracting the summary of the part of the content, generating a summary content text, and then converting the summary content text into audio, the summary content for continued listening can be obtained.
[0158] It should be noted that the method for generating the summary content of the continued listening in the embodiment of the present application can be executed by the server alone, by the terminal device alone, or by the server and the terminal device together. That is, the summary content of the continued listening in the embodiment of the present application can be generated only by the server side, can be generated only by the client side installed on the terminal device, or can be jointly generated based on the interaction between the server and the client.
[0159] Among them, when it is executed by the terminal device alone, that is, the terminal device determines the corresponding review duration based on the time interval between the pause time corresponding to the pause operation and the continued listening time corresponding to the resume operation, and the played time corresponding to the played part of the target audio program content; based on the review time, a section of audio content with a playing time equal to the review time is selected from the audio content corresponding to the played part as the audio content to be reviewed; the audio content to be reviewed is converted into text information, and a summary content text for the text information is generated based on the text summary technology; the summary content text is converted into audio to obtain the continued listening summary content.
[0160] For example, when the terminal device and the server are jointly executed, the client can determine the audio content to be reviewed and notify the server through the terminal device, and the server generates the summary content of the continued listening based on the audio content to be reviewed, etc., which is not specifically limited here.
[0161] like Fig. 9 When the method for generating the resume listening summary content shown is executed by a terminal device alone, the client installed on the terminal device can also determine whether the conditions for generating the resume listening summary content for the target audio program content are met before generating the resume listening summary content based on the time interval between the pause time and the resume listening time. For the specific judgment process and conditions, please refer to the above embodiment, and the repeated parts will not be repeated.
[0162] Similarly, the following methods for determining the review time and converting the summary content text into audio can be performed by the server alone, by the terminal device alone, or by the server and the terminal device together. The following mainly uses the example of the server alone performing the method for illustration.
[0163] The following is a detailed description of the process of determining the length of the review and converting the summary content text into audio:
[0164] In the embodiment of the present application, if the range of the review section needs to be determined, the corresponding review duration can be determined through the above step S901. Specifically, the review duration (also referred to as the review section range duration) is determined according to the user's listening time interval, content difficulty, etc., denoted as A1, and the determination method is as follows:
[0165] Determination method 1: Determine only based on the user's listening time interval.
[0166] In an optional implementation, if the time interval is not greater than a preset interval threshold, the product of the broadcast time and a first preset ratio value is used as the review time; if the time interval is greater than the preset interval threshold, the first preset ratio value is increased by a first set step size every time the time interval increases by a set time to obtain a first ratio value, and the product of the broadcast time and the first ratio value is used as the review time.
[0167] That is, first determine whether the listening time interval is greater than the preset interval threshold, assuming that the preset interval threshold is 5 hours. When the user's listening time interval is ≤ 5 hours, it can be determined that the basic review content is 20% of the listened segment, and the first preset ratio is 20%, then A1 = the duration of the listened content (i.e. the broadcast duration) * 20%.
[0168] That is, when y≤5, A1=x*20%. Among them, the duration of the content listened to=x, the time interval between two listenings=y, and the duration of the review section range=A1.
[0169] Assuming that the set duration is 1 hour, the first set step is 1%. When the user listens for a long time interval, greater than 5 hours, the review section range will also increase accordingly. For every 1 hour increase in the time interval, the review time will increase by 1%.
[0170] That is, when y>5, A1=x*[20%+(y-5)*1%], and the first ratio value is 20%+(y-5)*1%.
[0171] The second determination method is to determine based on the user's listening time interval and the difficulty of the content.
[0172] Compared with the first determination method, that is, when calculating the duration, the difficulty of the program content also needs to be considered. Specifically, in another optional implementation, the corresponding first review duration can be determined based on the time interval and the corresponding broadcast duration of the broadcasted portion of the target audio program content; the corresponding second review duration can be determined based on the content difficulty level corresponding to the target program content, wherein the greater the content difficulty level, the longer the second review duration; and then, the sum of the first review duration and the second review duration is used as the corresponding review duration.
[0173] Wherein, when determining the corresponding first review duration based on the time interval, it is similar to the determination method 1:
[0174] If the time interval is not greater than the preset interval threshold, the product of the broadcast time and the second preset ratio value is used as the first review time; if the time interval is greater than the preset interval threshold, the second preset ratio value is increased by a second set step every time the time interval increases by a set time to obtain a second ratio value, and the product of the broadcast time and the second ratio value is used as the first review time.
[0175] For example, the preset interval threshold is 5 hours, the duration of the content listened to = x, the time interval between two listenings = y, and the duration of the first review is A11.
[0176] Then, when y≤5, A11=x*20% (wherein the second preset ratio is 20%);
[0177] When y>5, A11=x*[20%+(y-5)*1%] (wherein the set time is 1 hour and the second set step is 1%), and the second ratio value is 20%+(y-5)*1%.
[0178] It should be noted that the first preset ratio value and the second preset ratio value in the embodiment of the present application can be the same or different, and no specific limitation is made here. Similarly, the first set step size and the second set step size can be the same or different, and no specific limitation is made here.
[0179] Wherein, when determining the corresponding second review duration based on the content difficulty level corresponding to the target program content, the specific process is as follows:
[0180] If the content difficulty level is not greater than the preset level threshold, the product of the broadcast time and the third preset ratio value is used as the second review time; if the content difficulty level is greater than the preset level threshold, the third preset ratio value is increased by a third set step for each increase in the content difficulty level to obtain a third ratio value, and the product of the broadcast time and the third ratio value is used as the second review time.
[0181] Assuming that the second review time is A12, the preset level threshold is 1, the set level is 1, the third preset ratio is 0%, and the third set step is 5%, then:
[0182] When z=1, A12=x*0%=0;
[0183] When z>1, A12=x*[(z-1)*5%].
[0184] That is, the higher the difficulty level of the content, the longer the corresponding second review time. If the content difficulty level is level 1, the second review time is +0%, and thereafter, the second review time increases by 5% for each level increase in difficulty.
[0185] It should be noted that the above is illustrated by taking the third preset ratio value of 0% as an example. The third preset ratio value is actually a non-negative number. In addition to the 0% listed above, it can also be 1%, 2%, 3%, etc. The specific value can be set according to the actual situation and is not specifically limited here.
[0186] In the embodiment of the present application, the field to which the program belongs is determined based on the program label, and then the content difficulty level is determined based on the field. Of course, other methods for determining the content difficulty level are also applicable and are not specifically limited here. See Table 1, which is an example of the relationship between the content difficulty level and the field in the embodiment of the present application.
[0187] Table 1
[0188]
[0189] The above description is based on three content difficulty levels as an example, and therefore, the second review duration corresponding to the highest content difficulty level is increased by 10%.
[0190] A1=A11+A12, in summary, it can be expressed as:
[0191] When y≤5, A1=x*[20%+(z-1)*5%]
[0192] That is, when y>5, A1=x*[20%+(y-5)*1%+(z-1)*5%].
[0193] It should be noted that if the final review duration A1 exceeds 100% of the duration of the content already listened to for the program, it is recorded as A1 = 100% of the duration of the content already listened to.
[0194] In the above implementation, the scope of content that needs to be helped to recall when intelligently continuing to listen is determined by judging the user's historical listening behavior, the difficulty of the audio program content, and other information, and speech recognition and automatic summarization technology are used to generate summary content for continued listening. Finally, the audio content is synthesized through speech synthesis technology, and the "smart continue listening" function is supported to be turned on when the user clicks to continue listening to the audio. By playing the summary audio content, the user can recall the audio content he has heard before and better connect with the continued listening content.
[0195] In the embodiment of the present application, after determining the review time, the paragraph range to be reviewed can be converted into text information. Specifically, firstly, the audio program content to be reviewed is uploaded to the server, and the audio content is converted into text information mainly by using the automatic speech recognition (Automatic Speech Recognition, ASR) language recognition technology. Among them, the ASR language recognition process is as follows: Fig.10 As shown, the specific process is as follows:
[0196] First, the audio content to be reviewed is uploaded to the server (i.e. Fig.10 Then, the important information reflecting the speech characteristics is extracted from the speech waveform of the audio content to be reviewed, and relatively irrelevant information (such as background noise) is removed, and this information is converted into a set of discrete parameter vectors (i.e. Fig.10 Encoding (feature extraction) in .
[0197] If there are multiple sound features in the audio content to be reviewed, it is also necessary to separate the speaking voice and determine the sound feature with the highest proportion. Specifically, the voice information is first preprocessed by performing voice activity detection (VAD) and framing on the audio content, and obtaining a sound waveform diagram. The conversion from time domain to frequency domain is then completed through Fourier transform, that is, Fourier transform is performed on each frame, and the characteristic parameter Mel Frequency Cepstral Coefficent (MFCC) is used to obtain the spectrum of each frame, and finally summarized as a spectrogram. This application uses this method to remove background noise, irrelevant human voices, etc. from program audio.
[0198] After feature extraction is completed, feature recognition and character generation (i.e. Fig.10 Decoding in), usually this application calls each pronunciation a "phoneme", which is the smallest unit in speech, such as vowels and consonants in Mandarin pronunciation. The speech is divided into frames through the acoustic model, which mainly handles pronunciation-related work. The output of the acoustic model includes the basic phoneme state and probability of utterance, covering the acoustic characteristics of the target language, and identifying the smallest "phoneme" in the speech. The system finds the currently spoken phoneme from each frame, and then composes words from multiple phonemes, and then composes text sentences from words. In the process, by judging which phoneme has the highest probability, the frame belongs to which phoneme. The system then composes words from multiple phonemes, and then composes text sentences from words. The language model training set helps the system combine semantic scenarios and contexts to achieve the best recognition effect.
[0199] Finally, the text output corresponding to the audio content to be reviewed is obtained through decoding.
[0200] Furthermore, the server generates a summary of the program continuation listening using a generative text summarization technique, hereinafter referred to as "summary".
[0201] In order to provide a better summary review experience, this application limits the playback time of the program content review summary finally played for the user to no more than 90 seconds, and the corresponding text content length shall not exceed 1000 words. The summary content length limit can be input into the system server.
[0202] Then, the generative text summarization technology (abstractive) is used to generate a content summary from the review paragraphs of audio products.
[0203] In the embodiment of the present application, the generative summary is based on the natural language generation (NLG) technology, which is a natural language description generated by the algorithm model according to the content of the source document, rather than extracting the original sentence. The generative text summary is mainly implemented by a deep neural network structure, also known as an encoder and decoder (Encoder, Decoder) architecture. The natural language processing (NLP) natural semantic recognition technology is used to establish an abstract semantic representation. After machine semantic recognition of the article content, the corresponding paragraph summary is generated according to the requirement of providing the summary length.
[0204] The generative summarization technology used in this application is based on the seq2seq (Sequence-to-Sequence) model in deep learning, and is implemented by adding an attention mechanism. The basic model structure is as follows Fig.11 As shown in the figure, the basic structure of the generative neural network model is mainly composed of an encoder and a decoder, and both encoding and decoding are implemented by neural networks.
[0205] The encoder is responsible for encoding the input original text into a vector C (Context), which is a representation of the original text and contains the text context. The decoder is responsible for extracting important information from this vector, obtaining semantic processing clips, and generating a text summary.
[0206] For example, the original text is “The XX XX became the largest tech…”, and the generated text summary is “XX tech…”, where XX is the abbreviation of XX XX.
[0207] In addition, considering that summaries generated from long texts have problems such as incoherent generation and repeated words and sentences in the field of text summarization, this application combines the intra-attention mechanism to solve the above problems, namely: 1) the classic decoder-encoder attention mechanism (Intra-temporal attention); 2) the attention mechanism inside the decoder (Intra-decoder attention).
[0208] Specifically, Intra-temporal attention enables the decoder to dynamically and on-demand obtain information from the input when generating results. It acts on the Encoder and calculates weights for each word in the input text (input), so that the generated content information can cover the original text. In the process of calculating the weights of Intra-temporal attention, the present application adopts a method to punish words with higher weights in the input to prevent the word from being given a high weight again in the subsequent decoding process. Intra-decoder attention enables the model to pay attention to the generated words, helping to solve the problem of repeating the same words and sentences when generating long sentences. It acts on the Decoder and also calculates weights for the generated words, so as to avoid generating repeated content. Then the two are spliced together for decoding to generate the next word. For each decoding step t, the sequence generated by the present application in the first decoding step is empty. This method is simpler and more widely applicable to other types of recursive networks.
[0209] In an optional implementation, when converting the summary content text into audio to obtain the continued listening summary content, the TTS of the continued listening summary content can be generated and played by learning the audio sound in the target audio program content. Of course, the TTS of the continued listening summary content can also be generated by some other sounds, such as fixed female voices, male voices, cartoon character voices, etc.
[0210] The following is a detailed description of the process of generating the TTS of the continued listening summary content by learning the audio sounds in the target audio program content:
[0211] Specifically, if the target audio program content contains the sound of an object, then based on the object sound, the summary content text is converted into audio to obtain the summary content for continued listening. If the target audio program content contains the sounds of multiple objects, then by performing feature extraction on the sounds of multiple objects, the sound with the highest proportion (i.e., the sound with the highest proportion) is determined, and based on the sound with the highest proportion, the summary content text is converted into audio to obtain the summary content for continued listening. That is to say, when the target audio program content contains the sounds of multiple objects, that is, when there are multiple sound features, the speaking voice can be separated to obtain the sound with the highest proportion (the sound with the highest proportion) in the audio, and the summary content text is converted into audio by performing acoustic feature learning of the audio features of the sound with the highest proportion using speech synthesis technology to obtain the summary content for continued listening.
[0212] Among them, considering the small amount of data in the audio program content, this application adopts a parameter-based speech synthesis method. This method uses a statistical model to generate speech parameters at any time and converts the parameters into sound waveforms. This process is actually a process of abstracting a text into phonetic features, using a statistical model to learn the correspondence between the phonetic features and their acoustic features, and then restoring the predicted acoustic features into a waveform. Use mainstream neural networks to predict, and then use a vocoder to generate a waveform to achieve the last step of converting features to waveforms.
[0213] See also Fig. 12A As shown, it is a schematic diagram of a parameter method speech synthesis process in an embodiment of the present application, which can be summarized as: audio feature extraction (parameter extraction) -> Hidden Markov Model (HMM) modeling -> parameter synthesis -> waveform reconstruction process. Fig. 12A The above processes are introduced in detail respectively:
[0214] First, it is necessary to extract audio features from the speech signal of the target audio program content.
[0215] For the target audio program content, this application mainly extracts its melspectrogram audio features. MFCC is a commonly used audio feature. For sound, it is actually a one-dimensional time domain signal, and it is difficult to intuitively see the change pattern of the frequency domain. Considering that the use of Fourier transform can obtain its frequency domain information, but the time domain information is lost, and it is impossible to see the change of the frequency domain with the time domain, so it is impossible to describe the sound well. In order to solve this problem, many time-frequency analysis methods have emerged, such as short-time Fourier, wavelet, Wigner distribution, etc. are all commonly used time-frequency domain analysis methods. Short-time Fourier is used in the embodiment of the present application.
[0216] Among them, short-time Fourier transform (STFT) refers to Fourier transform of short-time signal, which is obtained by framing long-time signal, and is suitable for analyzing stable signal. In the embodiment of the present application, it is assumed that the transformation of speech signal is flat within a shorter time span, and then fast Fourier transform (FFT) is performed on each frame by framing and windowing, and finally the result of each frame is stacked along another dimension to obtain a two-dimensional signal form similar to a picture. If the original signal of the present application is a sound signal, the two-dimensional signal obtained by STFT expansion is the so-called spectrogram.
[0217] The spectrogram is often a large image. In order to obtain the sound features of appropriate size, it is often transformed into a mel spectrum through a mel-scale filter bank. The mel cepstrum is obtained by performing cepstrum analysis (taking logarithm and performing discrete cosine transform) on the mel spectrum.
[0218] Based on the Mel-frequency cepstrum, parameters such as fundamental frequency parameters, speech parameters, etc. can be extracted.
[0219] Furthermore, HMM modeling is performed. Specifically, a continuous density hidden Markov model (CD-HMM) set is used to model speech parameters, and the output state of each HMM state is represented by a single Gaussian function (Gaussian) or a mixed Gaussian function (GaussianMixed Model, GMM, also known as Gaussian mixture model), and the goal of its parameter generation algorithm is to calculate a speech parameter sequence with a maximum likelihood function under the premise of a given Gaussian distribution sequence.
[0220] The above two processes correspond to Fig. 12A The training module in the above process can train the context-related HMM model, and then perform speech synthesis based on the model, that is, corresponding to Fig. 12A The synthesis module in .
[0221] After audio feature extraction and HMM modeling, it is necessary to perform parameter synthesis and waveform reconstruction on the summary content text.
[0222] Specifically, firstly, the summary content text needs to be input into the synthesis module (corresponding to Fig. 12A ), and then perform text analysis on the text, extract context features, and then generate a state sequence based on the context-related HMM model obtained by the above process, and then generate speech parameters. Finally, based on the parameter synthesizer, the speech parameters are converted into an acoustic waveform (i.e., parameter synthesis, waveform reconstruction), and the speech is output (i.e., the summary content of the continuation of listening).
[0223] Among them, when the audio features of the target audio program content are learned through speech synthesis technology, it specifically means that the phonemes, word segmentation, part of speech acquisition, sentence meaning understanding, and rhythm prediction and pinyin prediction in the audio features of the target audio program content are disassembled through end-to-end speech synthesis technology. Fig. 12B As shown, it is a schematic diagram of a specific process of text analysis in an embodiment of the present application, including the steps of input, sentence structure analysis, text regularization, text conversion to phonemes, and rhythm prediction.
[0224] After the text is input, it is necessary to perform sentence structure analysis on the text, including language identification and sentence segmentation. When performing sentence segmentation, this application is implemented based on a statistical word segmentation method:
[0225] From a formal point of view, a word is a stable combination of characters. Therefore, in the context, the more times adjacent characters appear at the same time, the more likely they are to form a word. Therefore, the frequency or probability of adjacent co-occurrence of characters can better reflect the feasibility of forming a word. The frequency of the expected combination of adjacent co-occurring characters can be counted to calculate their mutual occurrence information. The formula for calculating the mutual occurrence information of Chinese characters X and Y is M(X, Y) = lg(P(X, Y) / P(X)P(Y)). Among them, P(X, Y) is the adjacent co-occurrence probability of Chinese characters X and Y, and P(X) and P(Y) are the frequencies of X and Y appearing in the corpus respectively. The mutual occurrence information reflects the closeness of the relationship between Chinese characters. When the closeness is higher than a certain threshold, it can be considered that this character group may constitute a word. This method only needs to count the frequency of character groups in the corpus, and does not require a segmentation dictionary, so it is also called dictionary-free word segmentation or statistical word extraction method.
[0226] In the text regularization part, text regularization classification and rule replacement are required. In the text conversion to phoneme part, language identification is also required first, and then part of speech prediction and text conversion to phoneme are performed.
[0227] Among them, part-of-speech prediction is part-of-speech tagging. Among them, part-of-speech tagging is also called word class tagging or simply tagging, which refers to the procedure of marking a correct part of speech for each word in the word segmentation result, that is, the process of determining whether each word is a noun, verb, adjective or other part of speech. Assist this application in syntactic analysis preprocessing. This application can perform part-of-speech tagging based on the HMM model, and the model can be trained using a large corpus with labeled data, and labeled data refers to text in which each word is assigned a correct part-of-speech tag. In addition, sentence meaning can also be understood through syntactic analysis, and syntactic analysis refers to the basic task of determining the syntactic structure of a sentence or the dependency relationship between words in a sentence. This step can be completed by constructing a grammar tree.
[0228] Finally, the rhythm prediction part mainly refers to rhythm prediction, which is the key to speech synthesis.
[0229] In summary, after the above Fig. 12A , Fig. 12B In the above process listed, the server outputs the summary content TTS to the client, and after the user clicks the "play" button, the client preferentially plays the audio of the review content, thereby achieving the effect of helping users review the historical listening content in this application.
[0230] The above is the method for determining the review time and converting the summary content text into audio listed in the embodiment of the present application. The method can also be executed by the terminal device alone, or by the terminal device and the server together. The process is similar for these two methods, and the repeated parts will not be repeated.
[0231] In an optional implementation, after receiving a setting request for a resume permission control in a permission setting interface sent by a client, the server obtains resume permission information associated with a target object, and associates the resume permission information with identification information of the target object and saves the association.
[0232] Specific, for example Figure 6 As shown, the user can set the replay permission through the permission setting interface, and the client sends a setting request to the server, the request carries the identification information of the target object and the related replay permission information, which is associated and saved by the server.
[0233] In the above implementation, it supports turning on the "smart resume listening" function when the user clicks on the resume listening audio. By playing the summary audio content, it helps the user recall the audio content that he has listened to before and better connect with the resume listening content.
[0234] In summary, the playback control method of the audio program content in this application supports the intelligent generation of a review summary when the user clicks to continue listening to the audio, recommends a quick recall function to the user, and helps the user review the content of the episode that he has listened to before. This part of the content can be better connected with the content to be listened to, and enhances the user's understanding of the content to be listened to.
[0235] See also Fig.13A As shown, it is a flow chart of a method for implementing audio program content playback control based on a client and a server in an embodiment of the present application. The implementation process of the method is as follows:
[0236] On the client side: First, the "Smart Continue Listening" function is turned on; the user pauses the program (i.e., the target audio program content); the user clicks the play button for the program;
[0237] Based on the user's pause and play, the target audio program content is resumed. At this time, the client needs to first analyze the program's playing time, which can be divided into two cases: the program's playing time is less than 2 minutes and the program's playing time is greater than or equal to 2 minutes:
[0238] If the program has been broadcast for less than 2 minutes, no summary content will be generated;
[0239] If the program has been broadcast for >= 2 minutes, the server will continue to determine the user's listening time interval;
[0240] On the server side: if the user's listening interval is less than 5 hours, no resume listening summary content will be generated;
[0241] If the user's listening time interval is greater than or equal to 5 hours, the review segment range is determined; the specific determination method may refer to the determination method 1, determination method 2, etc. listed in the above embodiment, and the repeated parts will not be repeated.
[0242] Then, the server converts the audio of the segment to be reviewed into text information; generates a summary content text; determines the main sound in the program (i.e. the sound with the highest proportion); and generates a summary content for continued listening based on the sound.
[0243] The specific implementation method of the above process can be found in the examples in the relevant part above, and the repeated parts will not be repeated here.
[0244] Finally, the server feeds back the resume listening summary content to the client, and the client plays the resume listening summary content.
[0245] Based on the above introduction, taking the program broadcast time >= 2 minutes and the user listening time interval >= 5 hours as an example, the following Fig. 13B This section describes the interaction process between the client and the server in detail. Fig. 13B As shown, it is an interaction timing diagram between a client and a server in an embodiment of the present application, which specifically includes the following steps:
[0246] Step S1301: the client pauses playing the target audio program content in response to the pause operation triggered for the target audio program content, and sends a pause request to the server;
[0247] Step S1302: the server records the corresponding pause time;
[0248] Step S1303: the client responds to the recovery operation triggered for the target audio program content and sends a recovery request to the server;
[0249] Step S1304: the server records the corresponding continued listening time;
[0250] Step S1305: The server determines that the target audio program content meets the target condition;
[0251] Step S1306: the server generates a resume listening summary content for the target audio program content based on the time interval between the pause time and the resume listening time, and feeds back the resume listening summary content to the client;
[0252] Step S1307: the client plays the continued listening summary content corresponding to the target audio program content, and, after the continued listening summary content is played, continues to play the unplayed portion of the target audio program content.
[0253] Based on the same inventive concept, the present application embodiment also provides a playback control device for audio program content. Fig.14 As shown, it is a structural diagram of the playback control device 1400 for audio program content, which may include:
[0254] The pause unit 1401 is used to pause the playing of the target audio program content in response to the pause operation triggered for the target audio program content;
[0255] The resume unit 1402 is used to respond to the recovery operation triggered for the target audio program content, display the resume listening control area in the playback control interface, and play the resume listening summary content corresponding to the target audio program content, wherein the resume listening summary content is the summary information generated for the audio content corresponding to the played part of the target audio program content; after the resume listening summary content is played, continue to play the unplayed part of the target audio program content.
[0256] Optionally, the continue listening control area includes a summary control widget, and the continue playing unit 1402 is further used to:
[0257] Before the playback of the resumed listening summary content ends, in response to a closing operation triggered on the summary control widget, the playback of the resumed listening summary content is closed, and the audio content corresponding to the unplayed portion of the target audio program content continues to be played.
[0258] Optionally, the device further comprises:
[0259] The setting unit 1403 is used to set the replay permission for the target object in response to the setting operation of the replay permission control in the permission setting interface before the replay unit 1402 responds to the playback operation of resuming the playback of the target audio program content, and send the corresponding replay permission information to the server, so that the server associates the replay permission information with the identification information of the target object and saves it.
[0260] Optionally, the replay unit 1402 is further configured to:
[0261] In response to the recovery operation triggered by the target object for the target audio program content, if it is determined that the target object has the resume play permission based on the resume play permission information associated with the target object, the resume listening control area is displayed in the playback control interface, and the resume listening summary content corresponding to the target audio program content is played.
[0262] Optionally, the replay unit 1402 is further configured to determine the replay summary content in the following manner:
[0263] Determine the corresponding review duration based on the time interval between the pause time corresponding to the pause operation and the continued listening time corresponding to the resume operation, and the played duration corresponding to the played portion of the target audio program content;
[0264] Based on the review duration, a segment of audio content having a playback duration equal to the review duration is selected from the audio content corresponding to the played portion as the audio content to be reviewed;
[0265] Convert the audio content to be reviewed into text information, and generate a summary text for the text information based on text summarization technology;
[0266] The summary content text is converted into audio to obtain the summary content for continued listening.
[0267] Optionally, the replay unit 1402 is specifically used for:
[0268] Determine a corresponding first review duration based on the time interval and the corresponding broadcast duration of the broadcast portion of the target audio program content;
[0269] Determine the corresponding second review duration based on the content difficulty level corresponding to the target program content, wherein the greater the content difficulty level, the longer the second review duration;
[0270] The sum of the first review duration and the second review duration is taken as the corresponding review duration.
[0271] Optionally, the replay unit 1402 is specifically used for:
[0272] If the target audio program content includes the sound of an object, based on the object sound, the summary content text is converted into audio to obtain the continued listening summary content;
[0273] If the target audio program content contains sounds of multiple objects, the sound with the highest proportion is determined by extracting features from the sounds of the multiple objects, and based on the sound with the highest proportion, the summary content text is converted into audio to obtain the continued listening summary content.
[0274] Based on the same inventive concept, the present application embodiment also provides another apparatus for controlling the playback of audio program content. Fig.15 As shown, it is a schematic diagram of the structure of the playback control device 1500 for audio program content, which may include:
[0275] The first recording unit 1501 is configured to record the corresponding pause time after receiving a pause request for the target audio program content sent by the client;
[0276] The second recording unit 1502 is configured to record the corresponding continued listening time after receiving a resume request for the target audio program content sent by the client;
[0277] Feedback unit 1503 is used to generate a resume listening summary content for the target audio program content based on the time interval between the pause time and the resume listening time, and feed back the resume listening summary content to the client, so that the client displays the resume listening control area in the playback control interface and plays the resume listening summary content, wherein the resume listening summary content is summary information generated for the audio content corresponding to the played part of the target audio program content.
[0278] Optionally, the device further comprises:
[0279] The determination unit 1504 is configured to determine whether the target audio program content satisfies at least one of the following target conditions:
[0280] The played time corresponding to the played part of the target audio program content is not less than the first time threshold;
[0281] The time interval between the pause time and the resume listening time corresponding to the target audio program content is not less than the second duration threshold.
[0282] Optionally, the feedback unit 1503 is specifically used for:
[0283] Determine a corresponding review duration based on the time interval and the corresponding broadcast duration of the broadcasted portion of the target audio program content;
[0284] Based on the review duration, a segment of audio content having a playback duration equal to the review duration is selected from the audio content corresponding to the played portion as the audio content to be reviewed;
[0285] Convert the audio content to be reviewed into text information, and generate a summary text for the text information based on text summarization technology;
[0286] The summary content text is converted into audio to obtain the summary content for continued listening.
[0287] Optionally, the feedback unit 1503 is specifically used for:
[0288] If the time interval is not greater than the preset interval threshold, the product of the broadcast time and the first preset ratio value is used as the review time;
[0289] If the time interval is greater than the preset interval threshold, the first preset ratio value is increased by a first set step length each time the time interval increases by a set time to obtain a first ratio value, and the product of the played time and the first ratio value is used as the review time.
[0290] Optionally, the feedback unit 1503 is specifically used for:
[0291] Determine a corresponding first review duration based on the time interval and the corresponding broadcast duration of the broadcast portion of the target audio program content;
[0292] Determine the corresponding second review duration based on the content difficulty level corresponding to the target program content, wherein the greater the content difficulty level, the longer the second review duration;
[0293] The sum of the first review duration and the second review duration is taken as the corresponding review duration.
[0294] Optionally, the feedback unit 1503 is specifically used for:
[0295] If the time interval is not greater than the preset interval threshold, the product of the broadcast duration and the second preset ratio value is used as the first review duration;
[0296] If the time interval is greater than the preset interval threshold, the second preset ratio is increased by a second preset step size every time the time interval increases by a set time to obtain a second ratio, and the product of the played time and the second ratio is used as the first review time.
[0297] Optionally, the feedback unit 1503 is specifically used for:
[0298] If the difficulty level of the content is not greater than the preset level threshold, the product of the broadcast time and the third preset ratio value is used as the second review time;
[0299] If the content difficulty level is greater than the preset level threshold, the third preset ratio value is increased by a third set step length each time the content difficulty level increases by a set level to obtain a third ratio value, and the product of the broadcast time and the third ratio value is used as the second review time.
[0300] Optionally, the feedback unit 1503 is specifically used for:
[0301] If the target audio program content includes the sound of an object, based on the object sound, the summary content text is converted into audio to obtain the continued listening summary content;
[0302] If the target audio program content contains sounds of multiple objects, the sound with the highest proportion is determined by extracting features from the sounds of the multiple objects, and based on the sound with the highest proportion, the summary content text is converted into audio to obtain the continued listening summary content.
[0303] Optionally, the device further comprises:
[0304] The associating unit 1505 is used to obtain the replay permission information associated with the target object after receiving the setting request for the replay permission control in the permission setting interface sent by the client, and associate the replay permission information with the identification information of the target object and save it.
[0305] For the convenience of description, the above parts are divided into modules (or units) according to their functions and described separately. Of course, when implementing this application, the functions of each module (or unit) can be implemented in the same or multiple software or hardware.
[0306] Those skilled in the art will appreciate that various aspects of the present application may be implemented as a system, method, or program product. Therefore, various aspects of the present application may be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software, which may be collectively referred to as a "circuit", "module", or "system" herein.
[0307] Based on the same inventive concept as the above method embodiment, an electronic device is also provided in the embodiment of the present application. The electronic device can be used for playing and controlling audio program content. In one embodiment, the electronic device can be a terminal device, such as Figure 2 The terminal device 210 shown may be an electronic device such as a smart phone, a tablet computer, a laptop or a PC.
[0308] Please refer to Fig.16 The terminal device 210 includes a display unit 1640, a processor 1680 and a memory 1620, wherein the display unit 1640 includes a display panel 1641, which is used to display information input by a user or information provided to a user and various object selection interfaces of the terminal device 210, and in the embodiment of the present application, is mainly used to display interfaces and shortcut windows of applications installed in the terminal device 210. Optionally, the display panel 1641 can be configured in the form of a liquid crystal display (LCD) or an organic light-emitting diode (OLED).
[0309] The processor 1680 is used to read the computer program and then execute the method defined by the computer program. For example, the processor 1680 reads the social application program, thereby running the application on the terminal device 210 and displaying the interface of the application on the display unit 1640. The processor 1680 may include one or more general-purpose processors and may also include one or more digital signal processors (DSP) to perform related operations to implement the technical solutions provided in the embodiments of the present application.
[0310] The memory 1620 generally includes internal memory and external memory. The internal memory may be a random access memory (RAM), a read-only memory (ROM), and a cache (CACHE), etc. The external memory may be a hard disk, an optical disk, a USB disk, a floppy disk, or a tape drive, etc. The memory 1620 is used to store computer programs and other data. The computer program includes an application corresponding to the application, etc. The other data may include data generated after the operating system or the application is run, and the data includes system data (such as configuration parameters of the operating system) and user data. In the embodiment of the present application, program instructions are stored in the memory 1620, and the processor 1680 executes the program instructions stored in 1620 to implement the playback control method of the audio program content discussed above, or to implement the function of the adaptation application discussed above.
[0311] In addition, the terminal device 210 may also include a display unit 1640 for receiving input digital information, character information or contact touch operation / contactless gesture, and generating signal input related to user settings and function control of the terminal device 210. Specifically, in the embodiment of the present application, the display unit 1640 may include a display panel 1641. The display panel 1641, such as a touch screen, can collect the user's touch operation on or near it (such as the player using any suitable object or accessory such as a finger, stylus, etc. on the display panel 1641 or on the display panel 1641) and drive the corresponding connection device according to a pre-set program. Optionally, the display panel 1641 may include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch position of the user, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact point coordinates, and then sends it to the processor 1680, and can receive and execute commands sent by the processor 1680. In an embodiment of the present application, if the user triggers a recovery operation on the target audio program content by clicking, the touch detection device in the display panel 1641 detects the touch operation, and sends a signal corresponding to the detected touch operation to the touch controller. The touch controller converts the signal into touch point coordinates and sends them to the processor 1680. The processor 1680 determines whether the user's operation is successful based on the received touch point coordinates.
[0312] The display panel 1641 may be implemented in various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the display unit 1640, the terminal device 210 may further include an input unit 1630, which may include but is not limited to one or more of a physical keyboard, a function key (such as a volume control key, a switch key, etc.), a trackball, a mouse, and a joystick. Fig.16 In the figure, the input unit 1630 includes an image input device 1631 and other input devices 1632 as an example.
[0313] In addition to the above, the terminal device 210 may also include a power supply 1690 for supplying power to other modules, an audio circuit 1660, a near field communication module 1670, and an RF circuit 1610. The terminal device 210 may also include one or more sensors 1650, such as an accelerometer, a light sensor, a pressure sensor, etc. The audio circuit 1660 specifically includes a speaker 1661 and a microphone 1662, etc. For example, the user can use voice control, and the terminal device 210 can collect the user's voice through the microphone 1662, can be controlled by the user's voice, and when the user needs to be prompted, the corresponding prompt tone is played through the speaker 1661.
[0314] Based on the same inventive concept as the above method embodiment, an electronic device is also provided in the embodiment of the present application. The electronic device can be used for playing control of audio program content. In one embodiment, the electronic device can be a server, such as Figure 2 In this embodiment, the structure of the electronic device can be as follows: Fig.17 As shown, it includes a memory 1701 , a communication module 1703 and one or more processors 1702 .
[0315] The memory 1701 is used to store computer programs executed by the processor 1702. The memory 1701 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system and programs required for running the instant messaging function, etc.; the data storage area may store various instant messaging information and operation instruction sets, etc.
[0316] The memory 1701 may be a volatile memory, such as a random-access memory (RAM); the memory 1701 may also be a non-volatile memory, such as a read-only memory, a flash memory, a hard disk drive (HDD) or a solid-state drive (SSD); or the memory 1701 may be any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 1701 may be a combination of the above memories.
[0317] The processor 1702 may include one or more central processing units (CPU) or a digital processing unit, etc. The processor 1702 is used to implement the above-mentioned audio program content playback control method when calling the computer program stored in the memory 1701 .
[0318] The communication module 1703 is used to communicate with terminal devices and other servers.
[0319] The specific connection medium between the memory 1701, the communication module 1703 and the processor 1702 is not limited in the embodiment of the present application. Fig.17 In the embodiment, the memory 1701 and the processor 1702 are connected via a bus 1704. The bus 1704 is connected to the processor 1702 via a bus 1704. Fig.17 The connections between other components are shown in bold lines, which are only for illustration and are not intended to be limiting. Bus 1704 can be divided into address bus, data bus, control bus, etc. For ease of illustration, Fig.17 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0320] The memory 1701 stores a computer storage medium, and the computer storage medium stores computer executable instructions, and the computer executable instructions are used to implement the playback control method of the audio program content of the embodiment of the present application. The processor 1702 is used to execute the playback control method of the audio program content, such as Figure 8 shown.
[0321] In some possible implementations, various aspects of the method for controlling the playback of audio program content provided by the present application may also be implemented in the form of a program product, which includes program code. When the program product is run on a computer device, the program code is used to enable the computer device to execute the steps of the method for controlling the playback of audio program content according to various exemplary implementations of the present application described above in this specification. For example, the computer device may execute the following steps: Figure 3 or Figure 8 Follow the steps shown in .
[0322] The program product may use any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0323] The program product of the embodiment of the present application can adopt a portable compact disk read-only memory (CD-ROM) and include program code, and can be run on a computing device. However, the program product of the present application is not limited to this. In this document, a readable storage medium can be any tangible medium containing or storing a program, which can be used by or in combination with a command execution system, device or device.
[0324] The readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, wherein the readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable signal medium may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with a command execution system, apparatus, or device.
[0325] The program code embodied on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the foregoing.
[0326] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present application.
[0327] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.
Claims
1. A method for controlling the playback of audio program content, It is characterized in that The method includes: In response to a pause operation triggered on the target audio program content, pausing the playing of the target audio program content; In response to the resume operation triggered for the target audio program content, a resume listening control area is displayed in the playback control interface, and the resume listening summary content corresponding to the target audio program content is played, wherein the resume listening summary content is summary information generated for the audio content corresponding to the played portion of the target audio program content; the resume listening summary content is determined in the following manner: based on the review duration, a section of audio content is selected from the audio content corresponding to the played portion as the audio content to be reviewed; the summary content generated based on the audio content to be reviewed is used as the resume listening summary content; the review duration is determined in the following manner: based on the pause time corresponding to the pause operation and the resume listening time corresponding to the resume operation, a time interval is determined; if the time interval is greater than a preset interval threshold, then each time the time interval increases by a set duration, a first preset ratio value is increased by a first set step length to obtain a first ratio value, and the product of the played duration corresponding to the played portion of the target audio program content and the first ratio value is used as the review duration; After the playback of the resume listening summary content is finished, the unplayed portion of the target audio program content continues to be played.
2. The method according to claim 1, It is characterized in that The continue listening control area includes a summary control widget, and the method further includes: Before the playback of the resumed listening summary content ends, in response to a closing operation triggered on the summary control widget, the playback of the resumed listening summary content is closed, and the audio content corresponding to the unplayed portion of the target audio program content continues to be played.
3. The method according to claim 1, It is characterized in that Before the resume operation triggered for the target audio program content is displayed in the playback control interface in response to the resume operation, and the resume summary content corresponding to the target audio program content is played, the method further includes: In response to the setting operation of the replay permission control in the permission setting interface, the replay permission for the target object is set, and the corresponding replay permission information is sent to the server, so that the server associates the replay permission information with the identification information of the target object and saves it.
4. The method according to claim 3, It is characterized in that In response to the recovery operation triggered for the target audio program content, displaying a continue listening control area in the playback control interface and playing the continue listening summary content corresponding to the target audio program content includes: In response to the recovery operation triggered by the target object for the target audio program content, if it is determined that the target object has the resume play permission based on the resume play permission information associated with the target object, a resume listening control area is displayed in the playback control interface, and the resume listening summary content corresponding to the target audio program content is played.
5. The method according to any one of claims 1 to 4, It is characterized in that The step of generating summary content based on the audio content to be reviewed as the resume listening summary content includes: Convert the audio content to be reviewed into text information, and generate a summary content text for the text information based on a text summarization technology; The summary content text is converted into audio to obtain the continued listening summary content.
6. The method according to claim 1, It is characterized in that The method for determining the review time also includes: The currently determined review duration is used as the first review duration; Determining a corresponding second review duration based on the content difficulty level corresponding to the target program content, wherein the greater the content difficulty level, the longer the second review duration; The sum of the first review duration and the second review duration is used as the corresponding review duration.
7. The method according to claim 5, It is characterized in that The step of converting the summary content text into audio to obtain the continued listening summary content includes: If the target audio program content includes a sound of an object, based on the object sound, the summary content text is converted into audio to obtain the continued listening summary content; If the target audio program content includes sounds of multiple objects, the sound with the highest proportion is determined by extracting features from the sounds of the multiple objects, and based on the sound with the highest proportion, the summary content text is converted into audio to obtain the continued listening summary content.
8. A method for controlling the playback of audio program content, It is characterized in that The method includes: After receiving a pause request for the target audio program content sent by the client, the corresponding pause time is recorded; After receiving the recovery request for the target audio program content sent by the client, recording the corresponding continued listening time; Based on the review duration, a segment of audio content is selected from the audio content corresponding to the played portion as the audio content to be reviewed; the summary content generated based on the audio content to be reviewed is used as the summary content of the continued listening of the target audio program content; the review duration is determined in the following manner: based on the pause time and the continued listening time, a time interval is determined; if the time interval is greater than a preset interval threshold, each time the time interval increases by a set time, a first preset ratio value is increased by a first set step to obtain a first ratio value, and the product of the played time corresponding to the played portion of the target audio program content and the first ratio value is used as the review duration; The resume listening summary content is fed back to the client so that the client displays a resume listening control area in a playback control interface and plays the resume listening summary content, wherein the resume listening summary content is summary information generated for the audio content corresponding to the played portion of the target audio program content.
9. The method according to claim 8, It is characterized in that In the step of selecting a segment of audio content from the audio content corresponding to the played portion based on the review duration as the audio content to be reviewed; Before using the summary content generated based on the audio content to be reviewed as the summary content for continued listening of the target audio program content, the method further includes: Determine that the target audio program content satisfies at least one of the following target conditions: The played time corresponding to the played part of the target audio program content is not less than the first time threshold; The time interval between the pause time and the continued listening time corresponding to the target audio program content is not less than a second duration threshold.
10. The method according to claim 8, It is characterized in that The step of using the summary content generated based on the audio content to be reviewed as the summary content for continued listening of the target audio program content includes: Convert the audio content to be reviewed into text information, and generate a summary content text for the text information based on a text summarization technology; The summary content text is converted into audio to obtain the continued listening summary content.
11. The method according to claim 8, It is characterized in that The review period is also determined as follows: If the time interval is not greater than the preset interval threshold, the product of the broadcast duration and the first preset ratio value is used as the review duration.
12. The method according to claim 8, It is characterized in that The method for determining the review time also includes: The currently determined review duration is used as the first review duration; Determining a corresponding second review duration based on the content difficulty level corresponding to the target program content, wherein the greater the content difficulty level, the longer the second review duration; The sum of the first review duration and the second review duration is used as the corresponding review duration.
13. The method according to claim 12, It is characterized in that The determining of the corresponding second review duration based on the content difficulty level corresponding to the target program content specifically includes: If the difficulty level of the content is not greater than the preset level threshold, the product of the broadcast time and the third preset ratio value is used as the second review time; If the content difficulty level is greater than the preset level threshold, the third preset ratio value is increased by a third preset step size every time the content difficulty level increases by a set level to obtain a third ratio value, and the product of the broadcast time and the third ratio value is used as the second review time.
14. The method according to claim 10, It is characterized in that The step of converting the summary content text into audio to obtain the continued listening summary content includes: If the target audio program content includes a sound of an object, based on the object sound, the summary content text is converted into audio to obtain the continued listening summary content; If the target audio program content includes sounds of multiple objects, the sound with the highest proportion is determined by extracting features from the sounds of the multiple objects, and based on the sound with the highest proportion, the summary content text is converted into audio to obtain the continued listening summary content.
15. The method according to any one of claims 8 to 13, It is characterized in that The method further comprises: After receiving the setting request for the resume permission control in the permission setting interface sent by the client, the resume permission information associated with the target object is obtained, and the resume permission information is associated with the identification information of the target object and saved.
16. A playback control device for audio program content, It is characterized in that include: A pause unit, configured to pause playing of the target audio program content in response to a pause operation triggered on the target audio program content; a resume unit, for displaying a resume control area in a playback control interface in response to a resume operation triggered for the target audio program content, and playing resume summary content corresponding to the target audio program content, wherein the resume summary content is summary information generated for the audio content corresponding to the played portion of the target audio program content; the resume summary content is determined in the following manner: based on the review duration, selecting a segment of audio content from the audio content corresponding to the played portion as the audio content to be reviewed; using the summary content generated based on the audio content to be reviewed as the resume summary content; the review duration is determined in the following manner: determining a time interval based on a pause time corresponding to the pause operation and a resume time corresponding to the resume operation; if the time interval is greater than a preset interval threshold, then increasing a first preset ratio value by a first set step length each time the time interval increases by a set duration to obtain a first ratio value, and using the product of the played duration corresponding to the played portion of the target audio program content and the first ratio value as the review duration; and continuing to play the unplayed portion of the target audio program content after the playback of the resume summary content ends.
17. A playback control device for audio program content, It is characterized in that include: A first recording unit, configured to record a corresponding pause time after receiving a pause request for a target audio program content sent by a client; A second recording unit, configured to record a corresponding continued listening time after receiving a recovery request for the target audio program content sent by the client; A feedback unit is used to select a segment of audio content from the audio content corresponding to the played portion based on the review time as the audio content to be reviewed; and use the summary content text generated based on the audio content to be reviewed as the summary content of the continued listening of the target audio program content; the review time is determined in the following manner: based on the pause time and the continued listening time, a time interval is determined; if the time interval is greater than a preset interval threshold, then when the time interval increases by a set time, a first preset ratio value is increased by a first set step to obtain a first ratio value, and the product of the played time corresponding to the played portion of the target audio program content and the first ratio value is used as the review time; The resume listening summary content is fed back to the client so that the client displays a resume listening control area in a playback control interface and plays the resume listening summary content, wherein the resume listening summary content is summary information generated for the audio content corresponding to the played portion of the target audio program content.
18. An electronic device, It is characterized in that It includes a processor and a memory, wherein the memory stores a program code, and when the program code is executed by the processor, the processor executes the steps of any method described in claims 1 to 7 or the steps of any method described in claims 8 to 15.
19. A computer-readable storage medium, It is characterized in that It includes program code. When the program product is run on an electronic device, the program code is used to enable the electronic device to execute the steps of any method described in claims 1 to 7 or the steps of any method described in claims 8 to 15.
Citation Information
Patent Citations
System and method for generating a visual summary of previously viewed multimedia content
CN103168297A
Providing a summary of a multimedia document in a session
CN110325982A
Systems and methods for recommending a pause position and for resuming playback of media content
WO2019084181A1