Audio program content playing control method and device, equipment and storage medium

By generating summary content for continued listening and providing intelligent review functions, the problem of user forgetfulness in audio program content sharing platforms is solved, and listening efficiency and user experience are improved.

CN120687631APending Publication Date: 2025-09-23TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510803797.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2021-05-18
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

On audio program content sharing platforms, when users pause for a long time and do not resume listening, they are unable to connect the previous and subsequent storylines well, resulting in forgetfulness and repeated playback.

Method used

After the audio program is paused, a resume listening summary is generated based on the user's historical listening behavior and content difficulty level, and a resume listening control area is displayed in the playback interface, providing intelligent review and playback functions.

Benefits of technology

It improves the listening efficiency of audio program content, reduces repeated playback caused by forgetfulness, and enhances users' understanding of program content and immersive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120687631A_ABST
    Figure CN120687631A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computers, and particularly provides a playing control method and device for audio program content, equipment and a storage medium, which are used for improving the listening efficiency of the audio program content. The method comprises the following steps: in response to a pause operation triggered for target audio program content, pausing playing of the target audio program content; in response to a recovery operation triggered for the target audio program content, displaying a continuous listening control area in the playing control interface, and playing continuous listening summary content corresponding to the target audio program content; the continuous listening summary content is summary information generated for the audio content corresponding to the played part in the target audio program content; the review duration corresponding to the continuous listening summary content is determined according to at least one of the following: the historical listening behavior of the object and the content difficulty level corresponding to the target audio program content; and after the continuous listening summary content is played, continuously playing the unplayed part in the target audio program content from the current pause position.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application. The application number of the original application is 202110541007.4, the original application date is May 18, 2021, and the name of the original application is "Playback control method, device, equipment and storage medium for audio program content". The entire content of the original application is incorporated into this application by reference. Technical Field

[0002] The present application relates to the field of computer technology, and in particular to the field of machine learning technology, and provides a playback control method, apparatus, device, and storage medium for audio program content. Background Art

[0003] With the rapid development of internet technology, various social networking apps have emerged. Among them, audio content sharing platforms are gaining increasing attention and popularity. Podcast platforms are a common type of audio content sharing platform. Many users use podcast platforms to record and share audio programs, or listen to audio programs shared by other users, including audio novels, storytelling, crosstalk, talk shows, and more.

[0004] However, in various audio program content sharing platforms in the related art, when playing audio program content, if the user pauses the program for a long time and does not resume listening to the program for a long time, when resuming listening, the user directly resumes listening at the position where the last pause was made, which will cause the user to be unable to connect well with the foreground story that has been forgotten, resulting in the user's thinking being unable to quickly keep up with the storyline and narrative rhythm of the subsequent paragraphs of the program, and causing varying degrees of forgetfulness of the content that has been listened to. Summary of the Invention

[0005] Embodiments of the present application provide a method, apparatus, device, and storage medium for controlling the playback of audio program content, so as to improve the listening efficiency of audio program content.

[0006] The first method for controlling the playback of audio program content provided in an embodiment of the present application includes:

[0007] During the playback of the target audio program, in response to a pause operation triggered on the target audio program content, pausing the playback of the target audio program content;

[0008] In response to a resume operation triggered for the target audio program content, a continue-listening control area is displayed in the playback control interface, and a continue-listening summary content corresponding to the target audio program content is played; wherein the continue-listening summary content is summary information generated for the audio content corresponding to the played portion of the target audio program content; the review duration corresponding to the continue-listening summary content is determined based on at least one of the following: the subject's historical listening behavior and the content difficulty level corresponding to the target audio program content; the continue-listening playback control area is used to control the playback status of the continue-listening summary content;

[0009] After the playback of the resume listening summary content is finished, the unplayed portion of the target audio program content is continued to be played from the current pause position.

[0010] The second method for controlling the playback of audio program content provided in the embodiment of the present application includes:

[0011] During playback of the target audio program, upon receiving a pause request and a resume request for the target audio program content from a client, determining a review duration based on at least one of the subject's historical listening behavior and a content difficulty level corresponding to the target audio program content;

[0012] Based on the review duration, generating a resume listening summary content according to the audio content corresponding to the played portion of the target audio program content; the resume listening summary content is summary information of the played portion;

[0013] The resume listening summary content is fed back to the client, so that the client displays a resume listening control area in a playback control interface, plays the resume listening summary content, and continues to play the unplayed portion of the target audio program content from the current pause position after the resume listening summary content is finished playing.

[0014] The first audio program content playback control device provided in an embodiment of the present application includes:

[0015] a pause unit, configured to pause the playing of the target audio program content in response to a pause operation triggered on the target audio program content during the playing of the target audio program;

[0016] A resume unit is configured to, in response to a resume operation triggered for the target audio program content, display a resume listening control area in the playback control interface and play the resume listening summary content corresponding to the target audio program content; wherein the review duration corresponding to the resume listening summary content is determined based on at least one of the following: the subject's historical listening behavior, the content difficulty level corresponding to the target audio program content; the resume listening playback control area is configured to control the playback status of the resume listening summary content; after the resume listening summary content is finished playing, the unplayed portion of the target audio program content is continued to be played from the current pause position.

[0017] Optionally, the resume listening control area includes a summary control widget, and the resume playing unit is further configured to:

[0018] Before the playback of the continue listening summary content ends, in response to a closing operation triggered on the summary control, the playback of the continue listening summary content is closed, and the audio content corresponding to the unplayed portion of the target audio program content continues to be played.

[0019] Optionally, the replay unit is used to:

[0020] In response to a resume operation triggered for the target audio program content, a resume control area containing resume prompt information is displayed in the playback control interface to prompt the subject that intelligent resume listening is currently in progress, and resume listening summary content corresponding to the target audio program content is played;

[0021] Furthermore, the replay unit is further configured to:

[0022] The continue listening control area is no longer displayed in the playback control interface.

[0023] Optionally, the device further includes:

[0024] a setting unit configured to, in response to a setting operation on a resume permission control in a permission setting interface, set a resume permission for a target object in response to a resume operation triggered on the target audio program content, before the resume unit displays a resume listening control area in the playback control interface and plays the resume listening summary content corresponding to the target audio program content, thereby turning on or off a smart resume listening mode; the smart resume listening mode refers to a smart playback control function that plays customized resume listening summary content according to the object's operation after playback of the audio program content is paused;

[0025] And, the corresponding resume permission information is sent to the server, so that the server associates the resume permission information with the identification information of the target object and stores it.

[0026] Optionally, the replay unit is further configured to:

[0027] In response to a resume operation triggered for the target audio program content, if it is determined, based on the resume permission information associated with the target object, that the target object has the resume permission, determining that the current state is in the smart resume listening mode;

[0028] Furthermore, a continue-listening control area is displayed in the play control interface, and the continue-listening summary content corresponding to the target audio program content is played.

[0029] Optionally, the replay unit is further configured to determine the resume listening summary content in the following manner:

[0030] Based on the review duration, selecting a segment of audio content from the audio content corresponding to the played portion as the audio content to be reviewed;

[0031] Converting the audio content to be reviewed into text information, and generating a summary text for the text information based on text summarization technology;

[0032] The summary content text is converted into audio according to key sound features in the audio content to be reviewed to obtain the continued listening summary content.

[0033] Optionally, the replay unit is specifically configured to:

[0034] determining a review duration corresponding to the resume listening summary content based on a time interval between a pause time corresponding to the pause operation and a resume listening time corresponding to the resume operation, and a played duration corresponding to a played portion of the target audio program content;

[0035] The review duration is positively correlated with both the time interval and the broadcast duration.

[0036] Optionally, the replay unit is specifically configured to:

[0037] Determining a first review duration based on a time interval between the pause operation and the resume operation and an elapsed time duration corresponding to a played portion of the target audio program content; wherein the first review duration is positively correlated with both the time interval and the elapsed time duration;

[0038] determining a corresponding second review duration based on a content difficulty level corresponding to the target program content, wherein the second review duration is positively correlated with the content difficulty level;

[0039] The sum of the first review duration and the second review duration is used as the corresponding review duration.

[0040] Optionally, the replay unit is specifically configured to:

[0041] If the target audio program content includes sounds of multiple objects, determining the sound with the highest proportion by extracting features from the multiple object sounds;

[0042] Based on the sound with the highest proportion, the summary content text is converted into audio to obtain the continued listening summary content.

[0043] The second audio program content playback control device provided in the embodiment of the present application includes:

[0044] a determination unit configured to, upon receiving a pause request and a resume request for the target audio program content sent by a client during playback of the target audio program, determine a review duration based on at least one of the subject's historical listening behavior and a content difficulty level corresponding to the target audio program content;

[0045] a generating unit configured to generate, based on the review duration and according to the audio content corresponding to the played portion of the target audio program content, a resume listening summary content; the resume listening summary content is summary information of the played portion;

[0046] The feedback unit is configured to feed back the resume listening summary content to the client, so that the client displays a resume listening control area in a playback control interface, plays the resume listening summary content, and continues to play the unplayed portion of the target audio program content from the current pause position after the resume listening summary content is finished.

[0047] Optionally, the device further includes:

[0048] The determining unit is configured to determine, before the generating unit generates the continued listening summary content based on the review duration and the audio content corresponding to the played portion of the target audio program content, whether the target audio program content satisfies at least one of the following target conditions:

[0049] The played duration corresponding to the played portion of the target audio program content is not less than a first duration threshold;

[0050] The time interval between the pause time and the continued listening time corresponding to the target audio program content is not less than a second duration threshold.

[0051] Optionally, the generating unit is specifically configured to:

[0052] Based on the review duration, selecting a segment of audio content from the audio content corresponding to the played portion as the audio content to be reviewed;

[0053] Converting the audio content to be reviewed into text information, and generating a summary text for the text information based on text summarization technology;

[0054] The summary content text is converted into audio according to key sound features in the audio content to be reviewed to obtain the continued listening summary content.

[0055] Optionally, the determining unit is specifically configured to:

[0056] determining a review duration corresponding to the resume listening summary content based on a time interval between a pause time corresponding to the pause operation and a resume listening time corresponding to the resume operation, and a played duration corresponding to a played portion of the target audio program content;

[0057] The review duration is positively correlated with both the time interval and the broadcast duration.

[0058] Optionally, the determining unit is specifically configured to:

[0059] If the time interval is not greater than the preset interval threshold, determining the review duration according to the broadcast duration; wherein the review duration is positively correlated with the broadcast duration;

[0060] If the time interval is greater than a preset interval threshold, the review duration is determined according to the broadcast duration and the time interval; wherein the review duration is positively correlated with both the time interval and the broadcast duration.

[0061] Optionally, the determining unit is specifically configured to:

[0062] If the time interval is not greater than the preset interval threshold, the product of the broadcast duration and the first preset ratio value is used as the review duration;

[0063] If the time interval is greater than the preset interval threshold, the first preset proportional value is increased by a first set step size each time the time interval increases by a set duration to obtain a first proportional value, and the product of the played duration and the first proportional value is used as the review duration.

[0064] Optionally, the determining unit is specifically configured to:

[0065] Determining a first review duration based on a time interval between the pause operation and the resume operation and an elapsed time duration corresponding to a played portion of the target audio program content; wherein the first review duration is positively correlated with both the time interval and the elapsed time duration;

[0066] determining a corresponding second review duration based on a content difficulty level corresponding to the target program content, wherein the second review duration is positively correlated with the content difficulty level;

[0067] The sum of the first review duration and the second review duration is used as the corresponding review duration.

[0068] Optionally, the determining unit is specifically configured to:

[0069] If the difficulty level of the content is not greater than the preset level threshold, determining the second review time according to the broadcast time; wherein the review time is positively correlated with the broadcast time;

[0070] If the content difficulty level is greater than a preset level threshold, the second review time is determined according to the played time and the content difficulty level; wherein the review time is positively correlated with both the played time and the content difficulty level.

[0071] Optionally, the determining unit is specifically configured to:

[0072] If the time interval is not greater than the preset interval threshold, the product of the broadcast duration and the second preset ratio value is used as the first review duration;

[0073] If the time interval is greater than the preset interval threshold, the second preset ratio value is increased by a second set step size every time the time interval increases by a set duration to obtain a second ratio value, and the product of the played duration and the second ratio value is used as the first review duration.

[0074] Optionally, the feedback unit is specifically configured to:

[0075] If the difficulty level of the content is not greater than the preset level threshold, the product of the broadcast duration and the third preset ratio value is used as the second review duration;

[0076] If the content difficulty level is greater than a preset level threshold, the third preset ratio value is increased by a third preset step size for each increase in the content difficulty level to obtain a third ratio value, and the product of the broadcast duration and the third ratio value is used as the second review duration.

[0077] Optionally, the generating unit is specifically configured to:

[0078] If the target audio program content includes sounds of multiple objects, determining the sound with the highest proportion by extracting features from the multiple object sounds;

[0079] Based on the sound with the highest proportion, the summary content text is converted into audio to obtain the continued listening summary content.

[0080] Optionally, the device further includes:

[0081] An association unit is configured to, after receiving a setting request for a resume permission control in a permission setting interface sent by the client, obtain the resume permission information associated with the target object, and associate and save the resume permission information with the identification information of the target object; wherein the resume permission control is used to turn on or off the smart resume listening mode; the smart resume listening mode refers to: after the playback of the audio program content is paused, an intelligent playback control function is used to play customized resume listening summary content according to the operation of the object.

[0082] An embodiment of the present application provides an electronic device, including a processor and a memory, wherein the memory stores program code, and when the program code is executed by the processor, the processor performs the steps of any one of the above-mentioned methods for controlling the playback of audio program content.

[0083] Embodiments of the present application provide a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the electronic device to perform the steps of any of the aforementioned methods for controlling playback of audio program content.

[0084] An embodiment of the present application provides a computer-readable storage medium including a program code. When the program product is run on an electronic device, the program code is used to enable the electronic device to execute the steps of any one of the above-mentioned methods for controlling the playback of audio program content.

[0085] The beneficial effects of this application are as follows:

[0086] The embodiments of the present application provide a method, apparatus, device, and storage medium for controlling the playback of audio program content. First, the method can dynamically adjust the review duration and summary content generation method based on the user's actual listening experience (such as historical behavior, program content difficulty, etc.), thereby enhancing the personalized experience. Compared to the traditional single choice of "replay from the beginning" or "directly resume", the present application provides users with a more natural and efficient content connection method.

[0087] Secondly, the resume summary is presented in audio format, eliminating the need for users to actively read or search for information, maintaining the inherent immersive experience of the audio program. This mechanism is particularly suitable for audio programs with complex content and strong logical structures, such as knowledge podcasts, audiobooks, and lectures. It helps users quickly refresh their memories and understand the context, effectively reducing repeated playback due to forgetfulness, thereby improving overall listening efficiency and content absorption.

[0088] In addition, a resume listening control area is provided in the playback control interface, further enhancing the user's control over the playback status, allowing users to choose to skip, pause or replay the summary content according to their own needs, improving interactive flexibility and ease of use.

[0089] In summary, since this application supports intelligently generating a review summary when the user clicks to continue listening to audio, and converting it into audio for playback, it can help users review the core ideas of the program content they have listened to before, and then connect with the content they continue listening to, enhance the user's understanding of the program content, reduce the situation where users repeatedly replay the audio program content due to forgetting it, and improve the listening efficiency of audio program content.

[0090] Other features and advantages of the present application will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present application. The purposes and other advantages of the present application can be realized and obtained by the structures particularly pointed out in the written description, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0091] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0092] Figure 1 A schematic diagram of a method for resuming broadcasting an audio program in the related art;

[0093] Figure 2 This is an optional schematic diagram of an application scenario in an embodiment of the present application;

[0094] Figure 3 This is a flowchart of an implementation of a first method for controlling playback of audio program content in an embodiment of the present application;

[0095] Figure 4 A schematic diagram of a playback control interface in an embodiment of the present application;

[0096] Figure 5 This is a schematic diagram of another playback control interface in an embodiment of the present application;

[0097] Figure 6 A schematic diagram of a permission setting interface in an embodiment of the present application;

[0098] Figure 7 Schematic diagram of a method for resuming broadcasting an audio program in an embodiment of the present application;

[0099] Figure 8 This is a flowchart of an implementation of the second method for controlling playback of audio program content in an embodiment of the present application;

[0100] Figure 9 This is a flow chart of a method for generating summary content of continued listening in an embodiment of the present application;

[0101] Figure 10 This is a schematic diagram of a language recognition process in an embodiment of the present application;

[0102] Figure 11 This is a schematic diagram of a model structure in an embodiment of the present application;

[0103] Figure 12A A schematic diagram of a parametric speech synthesis process in an embodiment of the present application;

[0104] Figure 12B This is a schematic diagram of a specific process of text analysis in an embodiment of the present application;

[0105] Figure 13A This is a flow chart of a method for controlling the playback of audio program content based on a client and a server in an embodiment of the present application;

[0106] Figure 13B This is a timing diagram of interaction between a client and a server in an embodiment of the present application;

[0107] Figure 14 This is a schematic diagram of the structure of the first audio program content playback control device in an embodiment of the present application;

[0108] Figure 15 Schematic diagram of the structure of the second audio program content playback control device in the embodiment of the present application;

[0109] Figure 16 A schematic diagram of the hardware structure of an electronic device to which an embodiment of the present application is applied;

[0110] Figure 17 A schematic diagram of the hardware structure of another electronic device to which an embodiment of the present application is applied. DETAILED DESCRIPTION

[0111] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of the technical solutions of this application, but not all of them. Based on the embodiments described in this application document, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the technical solutions of this application.

[0112] The following is an introduction to some concepts involved in the embodiments of this application.

[0113] Audio products: The audio products covered in this application include, but are not limited to, audiobooks, podcasts, and other products that deliver content only through interactive sound broadcasting.

[0114] Program: refers to the unit of content delivered in an audio product through the interactive form of sound broadcast.

[0115] Audio program content: The audio program content in the embodiments of this application refers to audio programs shared on instant messaging software or podcast platforms, such as audio novels, crosstalk, storytelling, talk shows, radio stations, etc. Among them, audio novels refer to general audio program content files. When playing in a player, the playback speed can be adjusted and the playback stop time can be automatically remembered to facilitate "reading". The audio program content in the embodiments of this application can refer to audio content containing audio data (referring to audio content that can be converted into text through voice recognition, not pure music).

[0116] Resume listening: also known as resume broadcast, refers to when a user interrupts listening to a program and then clicks play again to continue listening.

[0117] A client, or user end, is a program that interfaces with a server and provides local services to clients. Aside from some applications that run only locally, it's typically installed on a regular client computer and requires interaction with the server. Since the development of the Internet, common client applications have included web browsers for the World Wide Web, email clients for sending and receiving emails, and instant messaging client software. These applications require corresponding servers and service programs on the network to provide the corresponding services, such as databases and email. Therefore, a specific communication connection must be established between the client and server to ensure the proper operation of the application.

[0118] Application operation interface: It is the medium for interaction and information exchange between the application system and the user. It realizes the conversion between the internal form of information and the form acceptable to humans. The purpose is to enable users to operate the application conveniently and efficiently to achieve two-way interaction and complete the work they want to complete with the help of the application. In the embodiment of this application, the application operation interface includes human-computer interaction and graphical user interface. The specific application operation interface includes the permission setting interface, the playback control interface, etc. Among them, different application operation interfaces are used to display different content to the user, realizing different information interactions between the user and the application.

[0119] Audio content sharing platforms and podcasts: Audio content sharing platforms are a type of digital broadcasting technology that can be used to record audio content from online radio broadcasts and similar online audio programs. Users can download online broadcasts to their own players and listen to them on the go, eliminating the need to be physically seated at a computer or listen in real time, enjoying the freedom of anytime, anywhere. Furthermore, users can create their own audio programs and upload them online through podcast platforms to share with the public. These platforms can be considered clients for playing audio and video content. Currently, audio content sharing platforms have many applications, such as podcasts.

[0120] Playback control interface: A user-facing page used to control the playback of audio program content. An audio program content sharing platform can have one or more playback control interfaces as needed, with multiple playback control interfaces being redirected based on pre-defined logic. In this embodiment of the present application, the playback control interface primarily refers to a page used to control the playback of audio program content, including a resume playback control area, which is primarily used to control the playback status of the resume summary content.

[0121] Text-to-speech (TTS) technology is a technology that produces artificial speech through mechanical and electronic means. TTS (also known as text-to-speech) is a type of speech synthesis technology that converts computer-generated or externally input text into understandable, fluent spoken Chinese.

[0122] Artificial intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also involves studying the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.

[0123] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0124] Machine learning (ML) is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is at the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of AI. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning.

[0125] With the research and advancement of artificial intelligence technology, artificial intelligence technology has been studied and applied in many fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, unmanned driving, autonomous driving, drones, robots, smart medical care, smart customer service, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0126] The solutions provided in the embodiments of this application involve technologies such as artificial intelligence and machine learning. The methods for the acoustic model, language model, and generative neural network model proposed in the embodiments of this application can be divided into two parts, including a training part and an application part. The training part involves the technical field of machine learning. In the training part, the above-mentioned models are trained using machine learning technology, and the model parameters are continuously adjusted through optimization algorithms. The application part is used to perform speech recognition using the acoustic model, speech model, etc. trained in the training part, and to generate summary content using the generative neural network model trained in the training part.

[0127] The following is a brief introduction to the design concept of the embodiment of this application:

[0128] When voice is the sole input channel, users' efficiency in receiving information is far lower than when using multimodal interactive input methods such as voice, vision, and touch. Among audio programs, podcasts tend to be one to three hours long, and audiobooks can run for dozens of hours, making it impossible for most users to listen to the entire program in one sitting. If users pause for a while and then resume listening right where they left off, they lose connection to the previously forgotten story, making it difficult for them to quickly catch up with the plot and narrative rhythm of the subsequent sections.

[0129] That is to say, the audio product experience in the related art is that when the user pauses listening in the middle of the listening process and clicks the play button again, the listening starts from the last paused position, such as Figure 1As shown in FIG, it is a schematic diagram of a method for resuming an audio program in the related art. However, when a user pauses a program for a long time, the user will forget the content that has been listened to to varying degrees.

[0130] In view of this, embodiments of the present application propose a playback control method, apparatus, device, and storage medium for audio program content. This application supports intelligently generating a review summary when a user clicks to continue listening to audio, and converting it into audio for playback. This helps users review the core ideas of previously listened program content, thereby connecting with the continued listening content, enhancing the user's understanding of the program content, reducing the situation where users repeatedly replay audio program content due to forgetting it, and improving the listening efficiency of audio program content.

[0131] The preferred embodiments of the present application are described below in conjunction with the drawings in the specification. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application and are not used to limit the present application. In addition, the embodiments and features in the embodiments of the present application can be combined with each other if there is no conflict.

[0132] like Figure 2 As shown, it is a schematic diagram of an application scenario of an embodiment of the present application. The application scenario diagram includes two terminal devices 210 and a server 230. The terminal device 210 can log in to the relevant interface 220 of the target business execution. The terminal device 210 and the server 230 can communicate through a communication network.

[0133] In the embodiment of the present application, the interface 220 may be a playback control interface, a permission setting interface, etc. The user may log in to the interface 220 through the terminal device 210. The terminal device 210 responds to the operation triggered by the user on the interface 220 and sends a relevant request to the server 230. The server 230 then feeds back relevant information to the terminal device, etc. For example, in response to the recovery operation triggered for the target audio program content, the terminal device 210 sends a recovery request to the server 230. The server 230 also generates a resume summary content based on the request and feeds it back to the terminal device 210. The terminal device 210 displays the resume control area in the playback control interface and plays the resume summary content corresponding to the target audio program content, etc. These will not be listed here one by one, and will be described in detail below.

[0134] In an optional implementation, the communication network is a wired network or a wireless network.

[0135] In the embodiments of the present application, terminal device 210 is an electronic device used by a user. This electronic device can be a personal computer, mobile phone, tablet computer, laptop, e-book reader, or other computer device with certain computing capabilities and running instant messaging software and websites or social networking software and websites. Each terminal device 210 is connected to server 230 via a wireless network. Server 230 is a single server or a server cluster composed of multiple servers, a cloud computing center, or a virtualization platform.

[0136] In the embodiment of the present application, a client associated with the audio program content is installed on the terminal device 210. The client can be software, such as instant messaging software, podcast software, or a mini-program, webpage, etc., without specific limitation herein. Correspondingly, the server is a server corresponding to the software, webpage, mini-program, etc.

[0137] Among them, users can directly search for and play the audio program content they like to listen to through podcast software, or listen to the audio program content shared by friends in instant messaging software, or search or listen to audio program content in public accounts, mini-programs, etc. It should be noted that the audio program content in the embodiment of the present application refers to the audio recorded by the user. For example, the user narrates each chapter in a novel and records the corresponding audio files. The user then shares the recorded audio files to the podcast platform for everyone to listen to, which is to listen to the book. In this scenario, the audio program content refers to the audio recorded by the user. The user who listens to the audio program content can use the method in the embodiment of the present application to play the corresponding continued listening summary content of the audio program content in the playback control interface. The continued listening summary content is the summary information generated by the client or server for the audio content corresponding to the played part of the audio program content uploaded by the user. In addition, in addition to novels, it can also be crosstalk, storytelling, etc. recorded by the user, which is not specifically limited here.

[0138] It should be noted that Figure 2 The examples shown are just for illustration. In fact, the number of terminal devices and servers is not limited and is not specifically limited in the embodiments of this application.

[0139] In an embodiment of the present application, when there are multiple servers, the multiple servers can be combined into a blockchain, and the servers are nodes on the blockchain; as disclosed in the embodiment of the present application, the playback control method of the audio program content, wherein the relevant data involved can be stored on the blockchain, for example, pause time, resume time, content difficulty level, review time, playback time, etc.

[0140] In addition, the embodiments of the present application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, assisted driving and other scenarios.

[0141] It should also be emphasized that in the specific implementation of this application, subject-related data, such as historical listening behavior and other operational behaviors listed herein, is involved. When the above embodiments of this application are applied to specific products or technologies, the subject's permission or consent must be obtained, and the collection, use, and processing of the relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions.

[0142] See Figure 3 As shown, it is an implementation flow chart of the first audio program content playback control method provided in the embodiment of the present application, which is applied to a terminal device, specifically a client in the terminal device. The specific implementation process of the method is as follows S31 to S33:

[0143] S31: During the playback of the target audio program, the terminal device pauses the playback of the target audio program content in response to a pause operation triggered on the target audio program content;

[0144] S32: In response to the resume operation triggered for the target audio program content, the terminal device displays a resume control area in the playback control interface and plays resume summary content corresponding to the target audio program content; wherein the resume summary content is summary information generated for the audio content corresponding to the played portion of the target audio program content; the review duration corresponding to the resume summary content is determined based on at least one of the following: the subject's historical listening behavior and the content difficulty level corresponding to the target audio program content; the resume playback control area is used to control the playback status of the resume summary content;

[0145] In an embodiment of the present application, the historical listening behavior of an object refers to the behavioral characteristics exhibited by the object in the process of listening to audio program content in the past, such as the length of time listened to, the time related to the pause operation, the time related to the resume operation, etc. In this application, the user's understanding and memory of the played content can be intelligently judged based on these data, thereby adaptively adjusting the review time; and the content difficulty level is based on the semantic complexity, professional terminology density, speech speed and other factors of the audio program content. The level is divided after evaluation. High-difficulty content usually requires a longer review time to help users better connect with the context. The review time determined in the above manner can better meet the actual needs of users and improve the effectiveness of information review.

[0146] In an embodiment of the present application, the resume playback control area is used to control the playback status of the resume summary content, such as skipping, speed increase, etc., so that users can flexibly control the review process according to their own understanding, thereby enhancing the interactive experience and content absorption effect.

[0147] S33: After the playback of the resumed listening summary content is finished, the terminal device continues to play the unplayed portion of the target audio program content from the current pause position.

[0148] like Figure 4 , which is a schematic diagram of a playback control interface in an embodiment of the present application. Both interfaces 41 and 42 in the figure are playback control interfaces. A user can trigger a pause or resume operation on the audio program content by clicking the pause / play control S410 in interface S41. When the control is in the state shown in interface 41, playback of the target audio program content is currently paused. When the control is in the state shown in interface 42, playback of the target audio program content is currently resumed.

[0149] Optionally, after the terminal device responds to the recovery operation triggered for the target audio program content, it can also display a resume listening control area containing resume listening prompt information in the playback control interface to prompt the object that it is currently resuming listening intelligently and play the resume listening summary content corresponding to the target audio program content.

[0150] Still with the above Figure 4 For example, Figure 4 The dotted box S420 in the middle interface 42 is the continue listening control area in the embodiment of the present application. In this area, "intelligent continue listening is in progress", which is a continue listening prompt information in the embodiment of the present application, and at the same time, the continue listening summary content is played. The continue listening summary content is generated based on the played part (the part before 22:22). In this way, the user can intuitively know that they are currently in the intelligent continue listening state, and improve their perception and control ability of the playback process. At the same time, combined with the playback of the continue listening summary content, it helps users quickly review key information, enhance the understanding of content coherence, and improve the overall listening experience and learning efficiency of the audio program.

[0151] After the resume listening summary content is played, the target audio program content can be resumed normally.

[0152] Optionally, after the terminal device responds to the close operation triggered for the summary control widget, in addition to closing the playback of the resume listening summary content and continuing to play the audio content corresponding to the unplayed part of the target audio program content, the resume listening control area may no longer be displayed in the playback control interface.

[0153] like Figure 5 As shown in FIG, which is a schematic diagram of another playback control interface in an embodiment of the present application, it shows that after the resume listening summary content is played, the resume listening control area S420 is no longer displayed and the target audio program content continues to play. This can maintain the simplicity of the playback interface, avoid unnecessary information interference, and allow users to focus more on the main audio content currently being played, improving the overall operation smoothness and user experience.

[0154] In the above-mentioned implementation mode, it supports intelligent generation of a review summary when the user clicks to continue listening to the audio, and converts it into audio for playback, which can help the user review the core ideas of the program content that he has listened to before, and then connect with the content that he continues to listen to, thereby enhancing the user's understanding of the program content, reducing the situation where the user repeatedly replays the audio program content due to forgetting the audio program content that he has listened to, improving the listening efficiency of the audio program content, and optimizing the user experience of audio products.

[0155] In an optional embodiment, the continue listening control area includes a summary control widget. Figure 4 As shown in the middle interface 42, the "skip" in S420 is a summary control control in the embodiment of the present application. The user can click "skip" to end the playback of the summary content.

[0156] Specifically, before the playback of the resumed summary content ends, if the user clicks "Skip", the terminal device responds to the close operation triggered by the summary control control, closes the playback of the resumed summary content, and continues to play the audio content corresponding to the unplayed part of the target audio program content, and displays the following: Figure 5 The playback control interface shown no longer displays the resume listening control area, directly skips the playback of the resume listening summary content, and continues to play the target audio program content, that is, continues to play from 22:22.

[0157] In the above embodiment, the user can skip the playback of the resumed summary content based on the summary control. In addition, the user can also adjust the playback speed of the resumed summary content by doubling the speed, etc., to improve the playback efficiency of the audio program content.

[0158] Optionally, the embodiment of the present application also supports turning on the "smart resume listening" function when the user clicks to continue listening to audio. When the user turns on this function, after clicking to continue listening, the resume listening control area will be displayed in the playback control interface, and the resume listening summary content will be played.

[0159] For example Figure 6 As shown, it is a permission setting interface in an embodiment of the present application, wherein the dotted box S60 is a resume permission control. The user can turn on or off "smart resume listening" by clicking on the control. Figure 6 The display shows that "Smart Resume Listening" is turned on.

[0160] When the user clicks Figure 6 When the replay permission control shown turns on "smart resume listening", the terminal device responds to the setting operation of the replay permission control in the permission setting interface, sets the replay permission for the target object, and sends the corresponding replay permission information to the server, so that the server associates the replay permission information with the identification information of the target object and saves it.

[0161] In an embodiment of the present application, when setting the resume permission control, the user can turn on or off "smart resume listening". Therefore, in an optional implementation, when the terminal device responds to the recovery operation triggered by the target object (referring to the user or user account) for the target audio program content, it is also necessary to further determine whether the target object has the resume permission, that is, whether the object currently has the smart resume listening mode turned on or turned off (i.e., not turned on) the smart resume listening mode. When the user turns on "smart resume listening", he or she has the resume permission, otherwise he or she does not have the resume permission.

[0162] Specifically, if it is determined that the target object has the resume play permission according to the resume play permission information associated with the target object, a resume listening control area is displayed in the play control interface, and the resume listening summary content corresponding to the target audio program content is played.

[0163] The replay permission information associated with the target object may be stored locally on the terminal device when the user sets the permission, or may be requested by the terminal device to the server and returned by the server.

[0164] In the above embodiment, by adding an "intelligent continue listening" switch, after turning it on, users can enjoy the function of intelligently generating a review summary, which can effectively improve the product user experience.

[0165] See Figure 7 As shown, it is a schematic diagram of an audio program replay method in an embodiment of the present application. Figure 1 Compared with the schematic diagram of the audio program resuming method in the related art shown, the present application provides an "intelligent resume listening" function. When the "intelligent resume listening" function is turned on, after pausing the program playback, when you click the play button again, it does not directly resume playback at the pause position, but generates intelligent review content audio (i.e., resume listening summary content), and plays the review content, and then resumes playback at the pause position to avoid users forgetting the content they have listened to after pausing the program for a long time. The present application can help users review the program content they have listened to before and enhance their understanding of the content they continue listening to.

[0166] See Figure 8 As shown in FIG. 1 , it is an implementation flow chart of the second method for controlling the playback of audio program content provided in an embodiment of the present application, which is applied to a server. The specific implementation process of the method is as follows S81 to S83:

[0167] S81: During playback of a target audio program, the server receives a pause request and a resume request for the target audio program content from the client, and determines a review duration based on at least one of the subject's historical listening behavior and a content difficulty level corresponding to the target audio program content;

[0168] S82: The server generates a resume listening summary based on the review duration and the audio content corresponding to the played portion of the target audio program content; the resume listening summary is summary information of the played portion;

[0169] S83: The server feeds back the resume listening summary content to the client, so that the client displays the resume listening control area in the playback control interface and plays the resume listening summary content. After the resume listening summary content is finished playing, the client continues to play the unplayed part of the target audio program content from the current pause position.

[0170] Among them, the client is installed on the terminal device, and the communication between the client and the server is the communication between the terminal device and the server.

[0171] In an embodiment of the present application, when the user clicks to pause playback, the terminal responds to the pause operation triggered for the target audio program content, pauses the playback of the target audio program content, and sends a pause request to the server, which may carry the corresponding pause time. When the server receives the pause request, it records the pause time, for example, as T1. Similarly, when the user clicks to resume playback, the terminal responds to the resume operation triggered for the target audio program content, resumes the playback of the target audio program content, and sends a resume request to the server, which may carry the corresponding resume time. When the server receives the resume request, it records the resume time, for example, as T2. Furthermore, based on the time interval y=T2-T1 between the pause time and the resume listening time, a resume listening summary content is generated.

[0172] In an optional embodiment, before generating the continued listening summary content for the target audio program content, it is necessary to further determine whether the conditions for generating the continued listening summary content are met. Only if the conditions are met can the continued listening summary content for the target audio program content be generated based on the time interval between the pause time and the continued listening time.

[0173] Specifically, the target conditions include at least one of the following:

[0174] Condition 1: The played duration of the played portion of the target audio program content is not less than the first duration threshold;

[0175] Condition 2: The time interval between the pause time and the resume time corresponding to the target audio program content is not less than the second duration threshold.

[0176] That is, when the target audio program content satisfies at least one of the above target conditions, the condition for generating the continued listening summary content is satisfied.

[0177] Specifically, when the user turns on the "Smart Resume Listening" function, if the user pauses listening during the listening process and then clicks the play button again, the client will upload the corresponding playback time of the target audio program content to the server, and store the time when the user pauses and resumes playing the same program in the background and upload it to the server as a basis for judgment.

[0178] Assume that the first duration threshold is 2 minutes and the second duration threshold is 5 hours. Then, the time information T1 and T2 when the user pauses and resumes the same program is uploaded to the server, and the server calculates the time interval between T1 and T2. If the program has been played for less than 2 minutes, no content summary is generated; or if the listening interval y is within 5 hours, no content summary is generated; or if the program has been played for less than 2 minutes and the listening interval is within 5 hours, no content summary is generated.

[0179] It should be noted that the above is only an example. Of course, the above judgment can be omitted and the content summary can be directly generated. No specific limitation is made here.

[0180] For the case where the summary content of continued listening needs to be generated, an optional implementation method is as follows: Figure 9 The flowchart shown in step S82 is a flowchart of a method for generating a resume listening summary in an embodiment of the present application, including the following steps S901 to S903:

[0181] S901: The server selects, based on the review duration, a segment of audio content corresponding to the played portion, the playback duration of which is equal to the review duration, as the audio content to be reviewed;

[0182] S902: The server converts the audio content to be reviewed into text information, and generates a summary text for the text information based on text summarization technology;

[0183] S903: The server converts the summary content text into audio according to the key sound features in the audio content to be reviewed, and obtains the resume listening summary content.

[0184] That is, in step S901, it is necessary to determine the paragraph range that needs to be reviewed from the played audio content as the audio content to be reviewed. By extracting the summary of this part of the content, generating a summary content text, and then converting the summary content text into audio, you can get the summary content that can be listened to again.

[0185] It should be noted that the method for generating the resume listening summary content in the embodiments of the present application can be executed by the server alone, by the terminal device alone, or by both the server and the terminal device. That is, the resume listening summary content in the embodiments of the present application can be generated solely by the server side, solely by the client side installed on the terminal device, or jointly generated based on the interaction between the server and the client.

[0186] Among them, when executed by a terminal device alone, that is, the terminal device determines the review duration based on the subject's historical listening behavior and at least one of the content difficulty levels corresponding to the target audio program content; based on the review duration, a section of audio content with a playback time equal to the review duration is selected from the audio content corresponding to the played part as the audio content to be reviewed; the audio content to be reviewed is converted into text information, and a summary content text for the text information is generated based on text summarization technology; then, based on key sound features in the audio content to be reviewed, the summary content text is converted into audio to obtain a summary content for continued listening.

[0187] For example, when executed jointly by the terminal device and the server, the client may determine the audio content to be reviewed and notify the server via the terminal device, and the server may generate a summary of the continued listening based on the audio content to be reviewed, etc., which is not specifically limited here.

[0188] like Figure 9 When the method for generating the resume listening summary content shown is executed by a terminal device alone, the client installed on the terminal device can also determine whether the conditions for generating the resume listening summary content for the target audio program content are met before generating the resume listening summary content based on the time interval between the pause time and the resume listening time. The specific judgment process and conditions are shown in the above embodiment, and the repeated parts are not repeated here.

[0189] Similarly, the methods for determining the review duration and converting the summary content text into audio, listed below, can be performed by the server alone, by the terminal device alone, or by both the server and the terminal device. The following examples primarily use the server alone as an example.

[0190] The following is a detailed description of the process for determining the length of the review and converting the summary text into audio:

[0191] In the embodiment of the present application, if it is necessary to determine the range of the review section, the corresponding review duration can be determined through the above step S901. Specifically, the review duration (also referred to as the review section range duration) is determined based on at least one of the user's historical listening behavior (such as the listening time interval) and the content difficulty, denoted as A1, and the determination method is as follows:

[0192] Determination method 1: Determine based only on the user's historical listening behavior.

[0193] In this manner, the review duration corresponding to the resumed listening summary content can be determined based on the time interval between the pause time corresponding to the pause operation and the resumed listening time corresponding to the resume operation, as well as the played time corresponding to the played portion of the target audio program content; wherein the review duration is positively correlated with both the time interval and the played time, that is, the longer the time interval, the longer the review duration; similarly, the longer the played time, the longer the review duration.

[0194] In an optional embodiment, if the time interval is no greater than a preset interval threshold, it indicates that the time it takes for the user to resume playback after pausing the video is short, and the user's memory of the content that has been played is still relatively clear. In this case, the review duration can be determined based solely on the duration of playback; wherein the review duration is positively correlated with the duration of playback. The review duration is positively correlated with the duration of playback. For example, if the user's playback duration is 10 minutes, approximately 1 minute of review content can be generated proportionally; if the playback duration increases to 30 minutes, the review duration will be correspondingly increased to approximately 3 minutes to help the user recall key content more comprehensively. This design allows the length of the review content to adaptively match the user's listening progress, improving the effect of information connection.

[0195] If the time interval is greater than the preset interval threshold, it indicates that the user paused for a long time and may have forgotten some key content, requiring a more comprehensive review to assist understanding. In this case, the review time can be determined based on the elapsed playback time and the time interval; the review time is positively correlated with both the time interval and the elapsed playback time. For example, the longer the time interval, the longer the user's pause and resume time, and the higher the degree of forgetfulness. Therefore, a longer review time is required to help the user connect the content. At the same time, the longer the elapsed playback time, the more content has been played, and the amount of information that needs to be reviewed also increases, requiring a correspondingly longer review time. By comprehensively considering these two factors, the length of the review content can be dynamically adjusted to make the review effect more in line with the user's actual memory status and content comprehension needs.

[0196] Specifically, if the time interval is no greater than the preset threshold, the review duration only needs to be determined based on the elapsed time, and the review duration is positively correlated with the elapsed time. To ensure content continuity while avoiding redundant interference to the user during the review process, the product of the elapsed time and a first preset ratio can be used as the review duration. If the time interval is greater than the preset threshold, the user's memory decay worsens as the time interval increases, and the review intensity needs to be increased to improve content cohesion. In this case, the first preset ratio can be increased by a first set step size for each increase in the time interval, resulting in a first ratio value. The product of the elapsed time and the first ratio value is then used as the review duration. This allows the review duration to be dynamically adjusted based on the user's actual memory decay, improving comprehension while also taking into account playback efficiency.

[0197] That is, first determine whether the listening interval is greater than a preset interval threshold, assuming the preset interval threshold is 5 hours. If the user's listening interval is ≤ 5 hours, the basic review content can be determined to be 20% of the listened segment. The first preset ratio is 20%, so A1 = the duration of the listened content (i.e., the broadcast duration) * 20%.

[0198] That is, when y≤5, A1=x*20%, where the duration of the content listened to=x, the time interval between two listenings=y, and the duration of the review section range=A1.

[0199] Assuming the duration is set to 1 hour, the first setting step is 1%. If the user listens for a longer time interval, greater than 5 hours, the review segment range will also increase accordingly. For every additional 1 hour of time interval, the review time will increase by 1%.

[0200] That is, when y>5, A1=x*[20%+(y-5)*1%], and the first proportional value is 20%+(y-5)*1%.

[0201] The second determination method is based on the user's historical listening behavior and content difficulty.

[0202] Compared to determination method 1, the difficulty of the program content also needs to be considered when calculating the duration. Specifically, in another optional implementation, a first review duration can be determined based on the time interval and the elapsed duration of the played portion of the target audio program content, where the first review duration is positively correlated with both the time interval and the elapsed duration. A second review duration can be determined based on the difficulty level of the target program content, where the greater the difficulty level, the longer the second review duration. Furthermore, the sum of the first and second review durations is used as the corresponding review duration.

[0203] Among them, when determining the corresponding first review duration based on the time interval, it is similar to the determination method 1 and will not be repeated here.

[0204] If the time interval is not greater than the preset interval threshold, the review time only needs to be determined based on the broadcast time, and the review time is positively correlated with the broadcast time. In order to avoid redundant interference to the user during the review process while ensuring content continuity, the product of the broadcast time and the second preset ratio value can be used as the first review time; if the time interval is greater than the preset interval threshold, as the time interval increases, the user's memory decay increases, and the review intensity needs to be increased to improve the content connection effect. At this time, it can be set that the second preset ratio value is increased by a second set step every time the time interval increases by a set time length to obtain a second ratio value, and the product of the broadcast time and the second ratio value is used as the first review time.

[0205] For example, the preset interval threshold is 5 hours, the duration of the content listened to = x, the time interval between two listenings = y, and the duration of the first review is A11.

[0206] Then, when y≤5, A11=x*20% (wherein the second preset ratio is 20%);

[0207] When y>5, A11=x*[20%+(y-5)*1%] (wherein the set time is 1 hour and the second set step is 1%), and the second proportional value is 20%+(y-5)*1%.

[0208] It should be noted that the first preset ratio value and the second preset ratio value in the embodiment of the present application can be the same or different, and are not specifically limited here. Similarly, the first set step size and the second set step size can be the same or different, and are not specifically limited.

[0209] The specific process of determining the second review duration based on the difficulty level of the target program content is as follows:

[0210] If the content difficulty level is not greater than the preset level threshold, it indicates that the current audio program content is relatively simple or easy to understand, and the user's memory and understanding of the played content are relatively good. At this time, the second review time can be determined only based on the played time; among which, the review time is positively correlated with the played time. The specific relationship between the review time and the played time can be found in the above embodiment and will not be repeated here.

[0211] If the content difficulty level is greater than the preset level threshold, the current audio program content is more complex or highly professional, and the user may need a longer time to review and understand the played content after resuming playback. The second review time can be determined based on the playback time and the content difficulty level; among them, the review time is positively correlated with the playback time and the content difficulty level. For example: when the user has played for 15 minutes, if the program content is an easy-to-understand life podcast (low difficulty level), the system can generate about 1.5 minutes of review content; if the program content is a high-difficulty professional course (such as quantum physics explanation), the system will increase the review time according to the content difficulty level, for example, generating a 3-minute review content. In this way, the system can dynamically adjust the review time according to the complexity of different content to help users better understand and connect the audio content.

[0212] Specifically, if the content difficulty level is not greater than the preset level threshold, then the review time only needs to be determined based on the broadcast time, and the review time is positively correlated with the broadcast time. In order to avoid redundant interference to the user during the review process while ensuring content continuity, the product of the broadcast time and the third preset ratio value can be used as the second review time; if the content difficulty level is greater than the preset level threshold, then as the time interval increases, the user's memory decay increases, and the review intensity needs to be increased to improve the content connection effect. At this time, the third preset ratio value can be increased by a third set step for each increase in the set level of the content difficulty level to obtain a third ratio value, and the product of the broadcast time and the third ratio value can be used as the second review time.

[0213] Assuming that the second review duration is A12, the preset level threshold is 1, the set level is 1, the third preset ratio is 0%, and the third set step is 5%, then:

[0214] When z=1, A12=x*0%=0;

[0215] When z>1, A12=x*[(z-1)*5%].

[0216] That is, the higher the difficulty level of the content, the longer the corresponding second review time. If the content difficulty level is level 1, the second review time is +0%, and for each subsequent level increase in difficulty, the second review time is increased by 5%.

[0217] It should be noted that the above is illustrated by taking the third preset ratio value of 0% as an example. The third preset ratio value is actually a non-negative number. In addition to the 0% listed above, it can also be 1%, 2%, 3%, etc. The specific value can be set according to actual conditions and is not specifically limited here.

[0218] In the embodiment of the present application, the field to which the program belongs is determined based on the program label, and then the content difficulty level is determined based on the field. Of course, other ways of determining the content difficulty level are also applicable, such as language complexity analysis based on audio text, such as vocabulary professionalism, sentence structure complexity, etc., or a comprehensive evaluation based on user feedback data, such as the user's listening completion rate, number of repeated playbacks, and other behavioral indicators, etc.; this article does not make specific limitations here. Refer to Table 1, which is an example of the relationship between content difficulty level and field in an embodiment of the present application.

[0219] Table 1

[0220]

[0221] The above description is based on three levels of content difficulty. Therefore, the second review duration corresponding to the highest level of content difficulty is increased by 10%.

[0222] A1=A11+A12. In summary, it can be expressed as:

[0223] When y≤5, A1=x*[20%+(z-1)*5%]

[0224] That is, when y>5, A1=x*[20%+(y-5)*1%+(z-1)*5%].

[0225] It should be noted that if the final review duration A1 exceeds 100% of the duration of the program already listened to, it is recorded as A1 = 100% of the duration of the program already listened to.

[0226] Determination method three: Determine based on the difficulty of the audio program content. The specific implementation method can be found in the second review duration determination method listed above, and will not be repeated here.

[0227] In the above-mentioned embodiment, the scope of content that needs to be helped to be recalled by intelligently continuing to listen is determined by judging information such as the user's historical listening behavior and the difficulty of the audio program content, and speech recognition and automatic summarization technology are used to generate summary content for continued listening. Finally, the audio content is synthesized through speech synthesis technology, and the "smart continue listening" function is supported to be turned on when the user clicks to continue listening to the audio. By playing the summary audio content, the user is helped to recall the audio content he has heard before and to better connect with the continued listening content.

[0228] In the embodiment of the present application, after determining the review time, the paragraph range to be reviewed can be converted into text information. Specifically, firstly, the audio program content to be reviewed is uploaded to the server, and the automatic speech recognition (Automatic Speech Recognition, ASR) language recognition technology is mainly used to convert the audio content into text information. Among them, the ASR language recognition process is as follows: Figure 10 The specific process is as follows:

[0229] First, the audio content to be reviewed is uploaded to the server (i.e. Figure 10 Then, the important information reflecting the speech characteristics is extracted from the speech waveform of the audio content to be reviewed, and relatively irrelevant information (such as background noise) is removed. Then, this information is converted into a set of discrete parameter vectors (i.e. Figure 10 Encoding (feature extraction) in .

[0230] If there are multiple sound features in the audio content to be reviewed, it is also necessary to separate the speaking voice and extract the key sound features, such as judging the sound features with the highest proportion. Specifically, the voice information is first preprocessed, and the audio content is subjected to voice activity detection (VAD) and framing, and a sound waveform is obtained. Then, the conversion from time domain to frequency domain is completed through Fourier transform, that is, Fourier transform is performed on each frame, and the characteristic parameter Mel Frequency Cepstral Coefficent (MFCC) is used to obtain the spectrum of each frame, and finally summarized as a spectrum diagram. This application uses this method to remove background noise, irrelevant human voices, etc. from the program audio.

[0231] After the feature extraction is completed, the feature recognition and character generation phase (i.e. Figure 10 Decoding in), usually each pronunciation is called a "phoneme" in this application, which is the smallest unit in speech, such as vowels and consonants in Mandarin pronunciation. The speech is divided into frames through the acoustic model, which mainly handles pronunciation-related work. The output of the acoustic model includes the basic phoneme state and probability of utterance, covering the acoustic characteristics of the target language, and identifying the smallest "phoneme" in the speech. The system finds the currently spoken phoneme from each frame, and then composes words from multiple phonemes, and then composes text sentences from words. In the process, by judging which phoneme has the highest probability, the frame belongs to which phoneme. The system then composes words from multiple phonemes, and then composes text sentences from words. The language model training set helps the system combine semantic scenes and context to achieve the best recognition effect.

[0232] Finally, the text output corresponding to the audio content to be reviewed is obtained through decoding.

[0233] Furthermore, the server generates a summary of the program continuation listening by using generative text summarization technology, hereinafter referred to as "summary".

[0234] In order to provide a better summary review experience, this application limits the playback time of the program content review summary finally played for the user to no more than 90 seconds, and the corresponding text content length shall not exceed 1000 words. The summary content length limit can be input into the system server.

[0235] Subsequently, the generative text summarization technology (abstractive) is used to generate a content summary from the review paragraphs of audio products.

[0236] In the embodiment of the present application, the generative summary is based on the natural language generation (NLG) technology, which is a natural language description generated by the algorithm model according to the content of the source document, rather than extracting the original sentence. The generative text summary is mainly implemented by the deep neural network structure, also known as the encoder and decoder (Encoder, Decoder) architecture. The natural language processing (NLP) natural semantic recognition technology is used to establish an abstract semantic representation. After the machine semantic recognition of the article content, the corresponding paragraph summary is generated according to the requirement of providing the summary length.

[0237] The generative summarization technology used in this application is based on the seq2seq (Sequence-to-Sequence) model in deep learning, and is implemented by adding an attention mechanism. The basic model structure is as follows Figure 11 As shown in the figure, the basic structure of the generative neural network model mainly consists of an encoder and a decoder, and both encoding and decoding are implemented by neural networks.

[0238] The encoder encodes the input text into a vector C (Context), which represents the original text and includes its context. The decoder extracts important information from this vector, obtains semantic processing snippets, and generates a text summary.

[0239] For example, the original text is “The XX XX became the largest tech…”, and the generated text summary is “XX tech…”, where XX is the abbreviation of XX XX.

[0240] In addition, considering that summaries generated from long texts have problems such as incoherent generation and repeated words and sentences in the field of text summarization, this application combines the intra-attention mechanism to solve the above problems, namely: 1) the classic decoder-encoder attention mechanism (Intra-temporal attention); 2) the attention mechanism within the decoder (Intra-decoder attention).

[0241] Specifically, Intra-temporal attention enables the decoder to obtain information from the input end dynamically and on demand when generating results. It acts on the Encoder and calculates weights for each word in the input text (input), so that the generated content information can cover the original text. In the process of calculating the weights of Intra-temporal attention, this application adopts a method to punish words with higher weights in the input to prevent the word from being given a high weight again in the subsequent decoding process. Intra-decoder attention enables the model to pay attention to the generated words, helping to solve the problem of repeating the same words when generating long sentences. It acts on the Decoder and also calculates weights for the generated words, so as to avoid generating repeated content. The two are then spliced ​​together for decoding to generate the next word. For each decoding step t, the sequence generated by this application in the first decoding step is empty. This method is simpler and more widely applicable to other types of recursive networks.

[0242] In an optional embodiment, when converting the summary content text into audio to obtain the resume listening summary content, the TTS of the resume listening summary content can be generated and played by learning the audio sound in the target audio program content. Of course, the TTS of the resume listening summary content can also be generated by some other voices, such as a fixed female voice, male voice, cartoon character voice, etc.

[0243] The following details the process of generating a TTS summary of the continued listening content by learning the audio sounds in the target audio program content:

[0244] Specifically, if the target audio program content contains the sound of an object, then based on the object sound, the summary content text is converted into audio to obtain the continued listening summary content. If the target audio program content contains the sounds of multiple objects, then by extracting features from the multiple object sounds, the sound with the highest proportion (i.e., the sound with the highest proportion) is determined, and based on the sound with the highest proportion, the summary content text is converted into audio to obtain the continued listening summary content. That is to say, when the target audio program content contains the sounds of multiple objects, that is, when there are multiple sound features, the speaking voice can be separated to obtain the sound with the highest proportion (the sound with the highest proportion) in the audio, and the acoustic feature learning of the speech synthesis technology is performed on the audio features of the sound with the highest proportion, and the summary content text is converted into audio to obtain the continued listening summary content.

[0245] Considering the relatively small amount of data in audio program content, this application adopts a parameter-based speech synthesis method. This method uses a statistical model to generate speech parameters at any time and converts the parameters into sound waveforms. This process is actually a process of abstracting text into phonetic features, using a statistical model to learn the correspondence between the phonetic features and their acoustic features, and then restoring the predicted acoustic features into a waveform. The final step of converting features to waveforms is achieved by using mainstream neural networks for prediction and then using a vocoder to generate waveforms.

[0246] See Figure 12A As shown, it is a schematic diagram of a parameter method speech synthesis process in an embodiment of the present application, which can be summarized as: audio feature extraction (parameter extraction) -> Hidden Markov Model (HMM) modeling -> parameter synthesis -> waveform reconstruction process. Figure 12A The above processes are introduced in detail:

[0247] First, audio features need to be extracted from the speech signal of the target audio program content.

[0248] For the target audio program content, this application mainly extracts its melspectrogram audio features. MFCC is a relatively common audio feature. For sound, it is actually a one-dimensional time domain signal, and it is difficult to intuitively see the change pattern of the frequency domain. Considering that the use of Fourier transform can obtain its frequency domain information, but loses the time domain information, it is impossible to see the change of the frequency domain with the time domain, so it is impossible to describe the sound well. In order to solve this problem, many time-frequency analysis methods have emerged, such as short-time Fourier, wavelet, Wigner distribution, etc., which are all commonly used time-frequency domain analysis methods. Short-time Fourier is used in the embodiment of this application.

[0249] Among them, short-time Fourier transform (STFT) refers to the Fourier transform of a short-time signal. The short-time signal is obtained by framing a long-time signal and is suitable for analyzing a stable signal. In an embodiment of the present application, it is assumed that the transformation of the speech signal is flat within a shorter time span. By framing and windowing, each frame is subjected to a fast Fourier transform (FFT), and finally the result of each frame is stacked along another dimension to obtain a two-dimensional signal form similar to a picture. If the original signal of the present application is a sound signal, the two-dimensional signal obtained by STFT expansion is the so-called spectrogram.

[0250] Spectrograms are often large images. To obtain sound features of appropriate size, they are often transformed into mel-spectra using mel-scale filter banks. Cepstrum analysis (taking the logarithm and performing discrete cosine transform) on the mel-spectrogram yields the mel-spectrogram.

[0251] Based on the Mel-frequency cepstrum, parameters such as fundamental frequency parameters and speech parameters can be extracted.

[0252] Furthermore, HMM modeling is performed. Specifically, a continuous density hidden Markov model (CD-HMM) set is used to model speech parameters. The output state of each HMM state is represented by a single Gaussian function (Gaussian) or a mixture of Gaussian functions (Gaussian Mixed Model, GMM, also known as Gaussian Mixture Model). The goal of its parameter generation algorithm is to calculate the speech parameter sequence with the maximum likelihood function under the premise of a given Gaussian distribution sequence.

[0253] The above two processes correspond to Figure 12A The training module in the above process can train the context-related HMM model, and then perform speech synthesis based on the model, which corresponds to Figure 12A The synthesis module in .

[0254] After audio feature extraction and HMM modeling, it is necessary to perform parameter synthesis and waveform reconstruction on the summary content text.

[0255] Specifically, first, the summary content text needs to be input into the synthesis module (corresponding to Figure 12A ), and then perform text analysis on the text, extract context features, and then generate a state sequence based on the context-related HMM model obtained by the above process modeling, and then generate speech parameters. Finally, based on the parameter synthesizer, the speech parameters are converted into acoustic waveforms (i.e. parameter synthesis and waveform reconstruction), and the speech is output (i.e., the summary content of the continued listening).

[0256] Among them, when the audio features of the target audio program content are learned through speech synthesis technology, it specifically refers to the use of end-to-end speech synthesis technology to disassemble the audio features of the target audio program content into phonemes, word segmentation, part of speech acquisition, sentence meaning understanding, and perform rhythm prediction, pinyin prediction, etc. Figure 12B The figure shows a specific flow chart of text analysis in an embodiment of the present application, including the steps of input, sentence structure analysis, text regularization, text conversion to phonemes, and rhythm prediction.

[0257] After the text is input, it is necessary to perform sentence structure analysis on the text, including language identification and sentence segmentation. When performing sentence segmentation, this application is based on a statistical word segmentation method:

[0258] From a formal perspective, a word is a stable combination of characters. Therefore, in a context, the more times adjacent characters appear simultaneously, the more likely they are to form a word. Therefore, the frequency or probability of adjacent co-occurrence of characters can better reflect the feasibility of forming a word. The frequency of each combination of characters that are expected to co-occur adjacently can be counted to calculate their mutual occurrence information. The formula for calculating the mutual occurrence information of Chinese characters X and Y is M(X, Y) = lg(P(X, Y) / P(X)P(Y)). P(X, Y) is the probability of adjacent co-occurrence of Chinese characters X and Y, and P(X) and P(Y) are the frequencies of occurrence of X and Y in the corpus, respectively. Mutual occurrence information reflects the closeness of the relationship between Chinese characters. When the closeness is higher than a certain threshold, it can be considered that this character group may form a word. This method only requires counting the frequency of character groups in the corpus and does not require a segmentation dictionary. Therefore, it is also called dictionary-free word segmentation or statistical word extraction method.

[0259] In the text regularization part, text regularization classification and rule replacement are required. In the text-to-phoneme part, language identification is also required first, followed by part-of-speech prediction and text-to-phoneme conversion.

[0260] Among them, part-of-speech prediction is part-of-speech tagging. Among them, part-of-speech tagging is also called part-of-speech tagging or simply tagging, which refers to the process of marking a correct part of speech for each word in the word segmentation result, that is, the process of determining whether each word is a noun, verb, adjective or other part of speech. Assist this application in syntactic analysis preprocessing. This application can perform part-of-speech tagging based on the HMM model. The model can be trained using a large corpus of labeled data, and labeled data refers to text in which each word is assigned a correct part-of-speech tag. In addition, sentence meaning can also be understood through syntactic analysis. Syntactic analysis refers to the fact that its basic task is to determine the syntactic structure of a sentence or the dependency relationship between words in a sentence. This step can be completed by constructing a grammar tree.

[0261] Finally, the rhythm prediction part mainly refers to rhythm prediction, which is the key to speech synthesis.

[0262] In summary, after the above Figure 12A , Figure 12B In the above process listed, the server outputs the resume summary content TTS to the client, and after the user clicks the "play" button, the client gives priority to playing the audio of the review content, thereby achieving the effect of helping users review the historical listening content in this application.

[0263] The above is the method for determining the review time and converting the summary content text into audio listed in the embodiment of the present application. This method can also be executed by the terminal device alone, or by the terminal device and the server together. The process is similar for these two methods, and the repeated parts will not be repeated.

[0264] In an optional implementation, after receiving the setting request for the resume permission control in the permission setting interface sent by the client, the server obtains the resume permission information associated with the target object, and associates the resume permission information with the identification information of the target object and saves it.

[0265] Specifically, for example Figure 6 As shown, the user can set the replay permission through the permission setting interface, and the client sends a setting request to the server, which carries the identification information of the target object and the relevant replay permission information, which is associated and saved by the server.

[0266] In the above implementation, it supports turning on the "smart resume listening" function when the user clicks to continue listening to the audio. By playing the summary audio content, it helps the user recall the audio content he has listened to before and better connect with the continued listening content.

[0267] In summary, the playback control method for audio program content in this application supports intelligently generating a review summary when the user clicks to continue listening to the audio, recommending a quick recall function to the user, and helping the user review the content of the previously listened episode. This part of the content can better connect with the content being listened to, enhancing the user's understanding of the content being listened to.

[0268] See Figure 13A As shown in FIG, it is a flow chart of a method for controlling the playback of audio program content based on a client and a server in an embodiment of the present application. The implementation process of the method is as follows:

[0269] On the client side: First, the "Smart Resume" function is enabled; the user pauses the program (i.e., the target audio program content); the user clicks the play button for the program;

[0270] Based on the user's pause and play, the target audio program content is resumed. At this time, the client needs to first analyze the program's playing time, which can be divided into two cases: the program's playing time is less than 2 minutes and the program's playing time is greater than or equal to 2 minutes.

[0271] If the program has been played for less than 2 minutes, no resume summary content will be generated;

[0272] If the program has been played for more than 2 minutes, the server will continue to determine the user's listening time interval;

[0273] On the server side: If the user's listening interval is less than 5 hours, no resume summary content will be generated;

[0274] If the user listening time interval is greater than or equal to 5 hours, the review segment range is determined. The specific determination method can refer to the determination method 1 and determination method 2 listed in the above embodiment, and the repeated parts are not repeated here.

[0275] Then, the server converts the audio of the segment to be reviewed into text information; generates a summary content text; determines the main sound in the program (i.e. the sound with the highest proportion); and generates a summary content for continued listening based on the sound.

[0276] The specific implementation of the above process can be found in the examples in the relevant parts above, and the repeated parts will not be repeated.

[0277] Finally, the server feeds back the resume listening summary content to the client, and the client plays the resume listening summary content.

[0278] Based on the above introduction, taking the program broadcast time>=2min and the user listening time interval>=5h as an example, the following Figure 13B This section describes the interaction between the client and the server in detail. Figure 13B As shown, it is a timing diagram of interaction between a client and a server in an embodiment of the present application, which specifically includes the following steps:

[0279] Step S1301: The client pauses playing the target audio program content in response to a pause operation triggered on the target audio program content, and sends a pause request to the server;

[0280] Step S1302: The server records the corresponding pause time;

[0281] Step S1303: The client responds to the recovery operation triggered for the target audio program content and sends a recovery request to the server;

[0282] Step S1304: The server records the corresponding continued listening time;

[0283] Step S1305: The server determines whether the target audio program content meets the target condition;

[0284] Step S1306: The server generates a resume listening summary for the target audio program content based on the time interval between the pause time and the resume listening time, and feeds the resume listening summary back to the client;

[0285] Step S1307: the client plays the continued listening summary content corresponding to the target audio program content, and, after the continued listening summary content is played, continues to play the unplayed portion of the target audio program content.

[0286] Based on the same inventive concept, the embodiment of the present application also provides a playback control device for audio program content. Figure 14 As shown, it is a structural diagram of the playback control device 1400 for audio program content, which may include:

[0287] The pause unit 1401 is configured to pause the playing of the target audio program content in response to a pause operation triggered on the target audio program content during the playing of the target audio program;

[0288] The resume unit 1402 is used to display a resume control area in the playback control interface in response to a resume operation triggered for the target audio program content, and play the resume summary content corresponding to the target audio program content; wherein the review duration corresponding to the resume summary content is determined based on at least one of the following: the subject's historical listening behavior, the content difficulty level corresponding to the target audio program content; the resume playback control area is used to control the playback status of the resume listening summary content; after the resume listening summary content is finished playing, the unplayed part of the target audio program content continues to be played from the current pause position.

[0289] Optionally, the continue listening control area includes a summary control widget, and the resume playing unit 1402 is further configured to:

[0290] Before the playback of the continue listening summary content ends, in response to a closing operation triggered on the summary control, the playback of the continue listening summary content is closed, and the audio content corresponding to the unplayed portion of the target audio program content continues to be played.

[0291] Optionally, the replay unit 1402 is configured to:

[0292] In response to a resume operation triggered for the target audio program content, a resume control area containing resume prompt information is displayed in the playback control interface to prompt the subject that intelligent resume listening is currently in progress, and resume listening summary content corresponding to the target audio program content is played;

[0293] Furthermore, the replay unit 1402 is further configured to:

[0294] The continue listening control area is no longer displayed in the playback control interface.

[0295] Optionally, the device further includes:

[0296] A setting unit 1403 is configured to, in response to a resume operation triggered by the resume unit 1402 for the target audio program content, display a resume listening control area in the playback control interface and, before playing the resume listening summary content corresponding to the target audio program content, set the resume listening permission for the target object in response to a setting operation on the resume listening permission control in the permission setting interface to enable or disable an intelligent resume listening mode; the intelligent resume listening mode is an intelligent playback control function that plays customized resume listening summary content according to the object's operation after playback of the audio program content is paused;

[0297] And, the corresponding resume permission information is sent to the server, so that the server associates the resume permission information with the identification information of the target object and stores it.

[0298] Optionally, the replay unit 1402 is further configured to:

[0299] In response to a resume operation triggered for the target audio program content, if it is determined, based on the resume permission information associated with the target object, that the target object has the resume permission, determining that the current state is in the smart resume listening mode;

[0300] Furthermore, a continue-listening control area is displayed in the play control interface, and the continue-listening summary content corresponding to the target audio program content is played.

[0301] Optionally, the replay unit 1402 is further configured to determine the resume listening summary content in the following manner:

[0302] Based on the review duration, selecting a segment of audio content from the audio content corresponding to the played portion as the audio content to be reviewed;

[0303] Converting the audio content to be reviewed into text information, and generating a summary text for the text information based on text summarization technology;

[0304] The summary content text is converted into audio according to key sound features in the audio content to be reviewed to obtain the continued listening summary content.

[0305] Optionally, the replay unit 1402 is specifically configured to:

[0306] determining a review duration corresponding to the resume listening summary content based on a time interval between a pause time corresponding to the pause operation and a resume listening time corresponding to the resume operation, and a played duration corresponding to a played portion of the target audio program content;

[0307] The review duration is positively correlated with both the time interval and the broadcast duration.

[0308] Optionally, the replay unit 1402 is specifically configured to:

[0309] Determining a first review duration based on a time interval between the pause operation and the resume operation and an elapsed time duration corresponding to a played portion of the target audio program content; wherein the first review duration is positively correlated with both the time interval and the elapsed time duration;

[0310] determining a corresponding second review duration based on a content difficulty level corresponding to the target program content, wherein the second review duration is positively correlated with the content difficulty level;

[0311] The sum of the first review duration and the second review duration is used as the corresponding review duration.

[0312] Optionally, the replay unit 1402 is specifically configured to:

[0313] If the target audio program content includes sounds of multiple objects, determining the sound with the highest proportion by extracting features from the multiple object sounds;

[0314] Based on the sound with the highest proportion, the summary content text is converted into audio to obtain the continued listening summary content.

[0315] Based on the same inventive concept, the embodiment of the present application also provides another apparatus for controlling the playback of audio program content. Figure 15 As shown, it is a structural diagram of the playback control device 1500 for audio program content, which may include:

[0316] The determining unit 1501 is configured to, upon receiving a pause request and a resume request for the target audio program content from a client during playback of the target audio program, determine a review duration based on at least one of the subject's historical listening behavior and a content difficulty level corresponding to the target audio program content;

[0317] Generating unit 1502, configured to generate a resume listening summary content based on the review duration and the audio content corresponding to the played portion of the target audio program content; the resume listening summary content is summary information of the played portion;

[0318] Feedback unit 1503 is used to feed back the resume listening summary content to the client, so that the client displays the resume listening control area in the playback control interface, plays the resume listening summary content, and after the resume listening summary content is finished playing, continues to play the unplayed part of the target audio program content from the current pause position.

[0319] Optionally, the device further includes:

[0320] The determination unit 1504 is configured to determine, before the generation unit 1502 generates the continued listening summary content based on the review duration and the audio content corresponding to the played portion of the target audio program content, whether the target audio program content satisfies at least one of the following target conditions:

[0321] The played duration corresponding to the played portion of the target audio program content is not less than a first duration threshold;

[0322] The time interval between the pause time and the continued listening time corresponding to the target audio program content is not less than a second duration threshold.

[0323] Optionally, the generating unit 1502 is specifically configured to:

[0324] Based on the review duration, selecting a segment of audio content from the audio content corresponding to the played portion as the audio content to be reviewed;

[0325] Converting the audio content to be reviewed into text information, and generating a summary text for the text information based on text summarization technology;

[0326] The summary content text is converted into audio according to key sound features in the audio content to be reviewed to obtain the continued listening summary content.

[0327] Optionally, the determining unit 1501 is specifically configured to:

[0328] determining a review duration corresponding to the resume listening summary content based on a time interval between a pause time corresponding to the pause operation and a resume listening time corresponding to the resume operation, and a played duration corresponding to a played portion of the target audio program content;

[0329] The review duration is positively correlated with both the time interval and the broadcast duration.

[0330] Optionally, the determining unit 1501 is specifically configured to:

[0331] If the time interval is not greater than the preset interval threshold, determining the review duration according to the broadcast duration; wherein the review duration is positively correlated with the broadcast duration;

[0332] If the time interval is greater than a preset interval threshold, the review duration is determined according to the broadcast duration and the time interval; wherein the review duration is positively correlated with both the time interval and the broadcast duration.

[0333] Optionally, the determining unit 1501 is specifically configured to:

[0334] If the time interval is not greater than the preset interval threshold, the product of the broadcast duration and the first preset ratio value is used as the review duration;

[0335] If the time interval is greater than the preset interval threshold, the first preset proportional value is increased by a first set step size each time the time interval increases by a set duration to obtain a first proportional value, and the product of the played duration and the first proportional value is used as the review duration.

[0336] Optionally, the determining unit 1501 is specifically configured to:

[0337] Determining a first review duration based on a time interval between the pause operation and the resume operation and an elapsed time duration corresponding to a played portion of the target audio program content; wherein the first review duration is positively correlated with both the time interval and the elapsed time duration;

[0338] determining a corresponding second review duration based on a content difficulty level corresponding to the target program content, wherein the second review duration is positively correlated with the content difficulty level;

[0339] The sum of the first review duration and the second review duration is used as the corresponding review duration.

[0340] Optionally, the determining unit 1501 is specifically configured to:

[0341] If the difficulty level of the content is not greater than the preset level threshold, determining the second review time according to the broadcast time; wherein the review time is positively correlated with the broadcast time;

[0342] If the content difficulty level is greater than a preset level threshold, the second review time is determined according to the played time and the content difficulty level; wherein the review time is positively correlated with both the played time and the content difficulty level.

[0343] Optionally, the determining unit 1501 is specifically configured to:

[0344] If the time interval is not greater than the preset interval threshold, the product of the broadcast duration and the second preset ratio value is used as the first review duration;

[0345] If the time interval is greater than the preset interval threshold, the second preset ratio value is increased by a second set step size every time the time interval increases by a set duration to obtain a second ratio value, and the product of the played duration and the second ratio value is used as the first review duration.

[0346] Optionally, the feedback unit 1503 is specifically configured to:

[0347] If the difficulty level of the content is not greater than the preset level threshold, the product of the broadcast duration and the third preset ratio value is used as the second review duration;

[0348] If the content difficulty level is greater than a preset level threshold, the third preset ratio value is increased by a third preset step size for each increase in the content difficulty level to obtain a third ratio value, and the product of the broadcast duration and the third ratio value is used as the second review duration.

[0349] Optionally, the generating unit 1502 is specifically configured to:

[0350] If the target audio program content includes sounds of multiple objects, determining the sound with the highest proportion by extracting features from the multiple object sounds;

[0351] Based on the sound with the highest proportion, the summary content text is converted into audio to obtain the continued listening summary content.

[0352] Optionally, the device further includes:

[0353] The association unit 1505 is used to obtain the resume permission information associated with the target object after receiving the setting request for the resume permission control in the permission setting interface sent by the client, and associate the resume permission information with the identification information of the target object and save it; wherein the resume permission control is used to turn on or off the smart resume listening mode; the smart resume listening mode refers to: after the audio program content is paused, an intelligent playback control function is used to play customized resume listening summary content according to the object's operation.

[0354] For the convenience of description, the above parts are divided into modules (or units) according to their functions and described separately. Of course, when implementing this application, the functions of each module (or unit) can be implemented in the same or multiple software or hardware.

[0355] Those skilled in the art will appreciate that various aspects of the present application may be implemented as systems, methods, or program products. Therefore, various aspects of the present application may be specifically implemented in the following forms: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation that combines hardware and software aspects, which may be collectively referred to herein as a "circuit," "module," or "system."

[0356] Based on the same inventive concept as the above method embodiment, an electronic device is also provided in the embodiment of the present application. The electronic device can be used for playing and controlling audio program content. In one embodiment, the electronic device can be a terminal device, such as Figure 2 The terminal device 210 shown may be an electronic device such as a smart phone, a tablet computer, a laptop computer or a PC.

[0357] Please refer to Figure 16 The terminal device 210 includes a display unit 1640, a processor 1680, and a memory 1620. The display unit 1640 includes a display panel 1641 for displaying information input by a user or information provided to a user, as well as various object selection interfaces of the terminal device 210. In the embodiment of the present application, the display panel 1641 is mainly used to display interfaces and shortcut windows of applications installed in the terminal device 210. Optionally, the display panel 1641 can be configured in the form of a liquid crystal display (LCD) or an organic light-emitting diode (OLED).

[0358] The processor 1680 is configured to read a computer program and then execute the method defined by the computer program. For example, the processor 1680 reads a social networking application, thereby running the application on the terminal device 210 and displaying the application interface on the display unit 1640. The processor 1680 may include one or more general-purpose processors and may also include one or more digital signal processors (DSPs) to perform related operations to implement the technical solutions provided in the embodiments of the present application.

[0359] The memory 1620 generally includes internal memory and external memory. The internal memory may be a random access memory (RAM), a read-only memory (ROM), and a cache (CACHE), etc. The external memory may be a hard disk, an optical disk, a USB disk, a floppy disk, or a tape drive, etc. The memory 1620 is used to store computer programs and other data. The computer program includes an application corresponding to the application, etc. Other data may include data generated after the operating system or the application is run, and the data includes system data (such as configuration parameters of the operating system) and user data. In the embodiment of the present application, program instructions are stored in the memory 1620, and the processor 1680 executes the program instructions stored in 1620 to implement the playback control method of the audio program content discussed above, or to implement the function of the adaptation application discussed above.

[0360] In addition, the terminal device 210 may also include a display unit 1640 for receiving input digital information, character information or contact touch operations / contactless gestures, and generating signal inputs related to user settings and function control of the terminal device 210. Specifically, in an embodiment of the present application, the display unit 1640 may include a display panel 1641. The display panel 1641, such as a touch screen, can collect user touch operations on or near it (such as operations performed by a player using a finger, stylus, or any other suitable object or accessory on or on the display panel 1641) and drive the corresponding connection device according to a pre-set program. Optionally, the display panel 1641 may include two parts: a touch detection device and a touch controller. The touch detection device detects the user's touch direction and detects the signal generated by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device and converts it into touch point coordinates, which are then sent to the processor 1680. It can also receive commands sent by the processor 1680 and execute them. In an embodiment of the present application, if the user triggers the recovery operation of the target audio program content by clicking, the touch detection device in the display panel 1641 detects the touch operation, and sends a signal corresponding to the detected touch operation to the touch controller. The touch controller converts the signal into touch coordinates and sends them to the processor 1680. The processor 1680 determines whether the user's operation is successful based on the received touch coordinates.

[0361] The display panel 1641 can be implemented using various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the display unit 1640, the terminal device 210 can also include an input unit 1630, which can include but is not limited to one or more of a physical keyboard, function keys (such as volume control keys, power keys, etc.), a trackball, a mouse, and a joystick. Figure 16 In the figure, the input unit 1630 includes an image input device 1631 and other input devices 1632 as an example.

[0362] In addition to the above, the terminal device 210 may also include a power supply 1690 for powering other modules, an audio circuit 1660, a near-field communication module 1670, and an RF circuit 1610. The terminal device 210 may also include one or more sensors 1650, such as an accelerometer 1651, a distance sensor 1652, a fingerprint sensor 1653, a temperature sensor 1654, etc. The audio circuit 1660 specifically includes a speaker 1661 and a microphone 1662. For example, a user can use voice control. The terminal device 210 can collect the user's voice through the microphone 1662, can be controlled by the user's voice, and when a prompt is required, a corresponding prompt tone is played through the speaker 1661.

[0363] Based on the same inventive concept as the above method embodiment, an electronic device is also provided in the embodiment of the present application. The electronic device can be used for playing and controlling audio program content. In one embodiment, the electronic device can be a server, such as Figure 2 In this embodiment, the structure of the electronic device can be as follows: Figure 17 As shown, it includes a memory 1701 , a communication module 1703 and one or more processors 1702 .

[0364] Memory 1701 is used to store computer programs executed by processor 1702. Memory 1701 may mainly include a program storage area and a data storage area. The program storage area may store an operating system and programs required for running instant messaging functions, while the data storage area may store various instant messaging messages and operating instruction sets.

[0365] Memory 1701 may be a volatile memory, such as random-access memory (RAM); a non-volatile memory, such as read-only memory, flash memory, a hard disk drive (HDD), or a solid-state drive (SSD); or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but is not limited thereto. Memory 1701 may be a combination of the aforementioned memories.

[0366] The processor 1702 may include one or more central processing units (CPUs) or digital processing units, etc. The processor 1702 is configured to implement the aforementioned audio program playback control method when calling the computer program stored in the memory 1701 .

[0367] The communication module 1703 is used to communicate with terminal devices and other servers.

[0368] The specific connection medium between the memory 1701, the communication module 1703 and the processor 1702 is not limited in the embodiment of the present application. Figure 17 In the embodiment, the memory 1701 and the processor 1702 are connected via a bus 1704. Figure 17 The connections between the other components are shown in bold lines, which are only for illustration and are not intended to be limiting. The bus 1704 can be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, Figure 17 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0369] The memory 1701 stores a computer storage medium, which stores computer executable instructions. The computer executable instructions are used to implement the playback control method of the audio program content of the embodiment of the present application. The processor 1702 is used to execute the playback control method of the audio program content, such as Figure 8 shown.

[0370] In some possible implementations, various aspects of the audio program content playback control method provided in the present application may also be implemented in the form of a program product, which includes program code. When the program product is run on a computer device, the program code is used to cause the computer device to execute the steps of the audio program content playback control method according to various exemplary embodiments of the present application described above in this specification. For example, the computer device may execute the following steps: Figure 3 or Figure 8 Follow the steps shown in .

[0371] The program product may employ any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0372] The program product of the embodiment of the present application may be a portable compact disc read-only memory (CD-ROM) and include program code, and can be run on a computing device. However, the program product of the present application is not limited thereto. In this document, a readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with a command execution system, device, or device.

[0373] A readable signal medium may include a data signal transmitted in baseband or as part of a carrier wave, which carries readable program code. Such a transmitted data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium that can transmit, propagate, or transfer a program for use by or in conjunction with a command execution system, apparatus, or device.

[0374] The program code embodied on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0375] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.

[0376] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.

Claims

1. A method for controlling the playback of audio program content, characterized in that: The method includes: During the playback of the target audio program, in response to a pause operation triggered on the target audio program content, pausing the playback of the target audio program content; In response to a resume operation triggered for the target audio program content, a continue-listening control area is displayed in the playback control interface, and a continue-listening summary content corresponding to the target audio program content is played; wherein the continue-listening summary content is summary information generated for the audio content corresponding to the played portion of the target audio program content; the review duration corresponding to the continue-listening summary content is determined based on at least one of the following: the subject's historical listening behavior and the content difficulty level corresponding to the target audio program content; the continue-listening playback control area is used to control the playback status of the continue-listening summary content; After the playback of the resume listening summary content is finished, the unplayed portion of the target audio program content is continued to be played from the current pause position.

2. The method according to claim 1, wherein The continue listening control area includes a summary control widget, and the method further includes: Before the playback of the continue listening summary content ends, in response to a closing operation triggered on the summary control, the playback of the continue listening summary content is closed, and the audio content corresponding to the unplayed portion of the target audio program content continues to be played.

3. The method according to claim 2, wherein The method of displaying a resume control area in a playback control interface in response to a resume operation triggered for the target audio program content and playing a resume summary corresponding to the target audio program content includes: In response to a resume operation triggered for the target audio program content, a resume control area containing resume prompt information is displayed in the playback control interface to prompt the subject that intelligent resume listening is currently in progress, and resume listening summary content corresponding to the target audio program content is played; And, in response to the closing operation triggered on the summary control, closing the playback of the continued listening summary content and continuing to play the audio content corresponding to the unplayed portion of the target audio program content, further comprising: The continue listening control area is no longer displayed in the playback control interface.

4. The method according to claim 1, wherein Before displaying the resume control area in the playback control interface in response to the resume operation triggered for the target audio program content and playing the resume summary content corresponding to the target audio program content, the method further includes: In response to a setting operation for the resume permission control in the permission setting interface, the resume permission is set for the target object to enable or disable the smart resume listening mode; the smart resume listening mode is an intelligent playback control function that plays customized resume listening summary content according to the object's operation after the audio program content is paused; And, the corresponding resume permission information is sent to the server, so that the server associates the resume permission information with the identification information of the target object and stores it.

5. The method according to claim 4, wherein The method of displaying a resume control area in a playback control interface in response to a resume operation triggered for the target audio program content and playing a resume summary corresponding to the target audio program content includes: In response to a resume operation triggered for the target audio program content, if it is determined, based on the resume permission information associated with the target object, that the target object has the resume permission, determining that the current state is in the smart resume listening mode; Furthermore, a continue-listening control area is displayed in the play control interface, and the continue-listening summary content corresponding to the target audio program content is played.

6. The method according to any one of claims 1 to 5, wherein: The resume listening summary content is determined in the following manner: Based on the review duration, selecting a segment of audio content from the audio content corresponding to the played portion as the audio content to be reviewed; Converting the audio content to be reviewed into text information, and generating a summary text for the text information based on text summarization technology; The summary content text is converted into audio according to key sound features in the audio content to be reviewed to obtain the continued listening summary content.

7. The method according to any one of claims 1 to 5, wherein: Determining the review duration according to the historical listening behavior includes: determining a review duration corresponding to the resume listening summary content based on a time interval between a pause time corresponding to the pause operation and a resume listening time corresponding to the resume operation, and a played duration corresponding to a played portion of the target audio program content; The review duration is positively correlated with both the time interval and the broadcast duration.

8. The method according to any one of claims 1 to 5, wherein: Determining the review duration according to the historical listening behavior and the content difficulty level includes: Determining a first review duration based on a time interval between the pause operation and the resume operation and an elapsed time duration corresponding to a played portion of the target audio program content; wherein the first review duration is positively correlated with both the time interval and the elapsed time duration; determining a corresponding second review duration based on a content difficulty level corresponding to the target program content, wherein the second review duration is positively correlated with the content difficulty level; The sum of the first review duration and the second review duration is used as the corresponding review duration.

9. The method according to claim 6, wherein The step of converting the summary content text into audio based on key sound features in the audio content to be reviewed to obtain the continued listening summary content includes: If the target audio program content includes sounds of multiple objects, determining the sound with the highest proportion by extracting features from the multiple object sounds; Based on the sound with the highest proportion, the summary content text is converted into audio to obtain the continued listening summary content.

10. A method for controlling the playback of audio program content, characterized in that: The method includes: During playback of the target audio program, upon receiving a pause request and a resume request for the target audio program content from a client, determining a review duration based on at least one of the subject's historical listening behavior and a content difficulty level corresponding to the target audio program content; Based on the review duration, generating a resume listening summary content according to the audio content corresponding to the played portion of the target audio program content; the resume listening summary content is summary information of the played portion; The resume listening summary content is fed back to the client, so that the client displays a resume listening control area in a playback control interface, plays the resume listening summary content, and continues to play the unplayed portion of the target audio program content from the current pause position after the resume listening summary content is finished playing.

11. The method according to claim 10, wherein Before generating the resume listening summary content based on the review duration and the audio content corresponding to the played portion of the target audio program content, the method further includes: Determine that the target audio program content meets at least one of the following target conditions: The played duration corresponding to the played portion of the target audio program content is not less than a first duration threshold; The time interval between the pause time and the continued listening time corresponding to the target audio program content is not less than a second duration threshold.

12. The method according to claim 10, wherein The generating of the resume listening summary content based on the review duration and the audio content corresponding to the played portion of the target audio program content includes: Based on the review duration, selecting a segment of audio content from the audio content corresponding to the played portion as the audio content to be reviewed; Converting the audio content to be reviewed into text information, and generating a summary text for the text information based on text summarization technology; The summary content text is converted into audio according to key sound features in the audio content to be reviewed to obtain the continued listening summary content.

13. The method according to claim 10, wherein Determining the review duration according to the historical listening behavior includes: determining a review duration corresponding to the resume listening summary content based on a time interval between a pause time corresponding to the pause operation and a resume listening time corresponding to the resume operation, and a played duration corresponding to a played portion of the target audio program content; The review duration is positively correlated with both the time interval and the broadcast duration.

14. The method according to claim 13, wherein The determining, based on a time interval between a pause time corresponding to the pause operation and a continued listening time corresponding to the resume operation, and a played duration corresponding to a played portion of the target audio program content, of a review duration corresponding to the continued listening summary content includes: If the time interval is not greater than the preset interval threshold, determining the review duration according to the broadcast duration; wherein the review duration is positively correlated with the broadcast duration; If the time interval is greater than a preset interval threshold, the review duration is determined according to the broadcast duration and the time interval; wherein the review duration is positively correlated with both the time interval and the broadcast duration.

15. The method according to claim 14, wherein If the time interval is not greater than the preset interval threshold, determining the review duration according to the broadcast duration includes: If the time interval is not greater than the preset interval threshold, the product of the broadcast duration and the first preset ratio value is used as the review duration; If the time interval is greater than a preset interval threshold, determining the review duration according to the broadcast duration and the time interval includes: If the time interval is greater than the preset interval threshold, the first preset proportional value is increased by a first set step size each time the time interval increases by a set duration to obtain a first proportional value, and the product of the played duration and the first proportional value is used as the review duration.

16. The method according to claim 10, wherein Determining the review duration according to the historical listening behavior and the content difficulty level includes: Determining a first review duration based on a time interval between the pause operation and the resume operation and an elapsed time duration corresponding to a played portion of the target audio program content; wherein the first review duration is positively correlated with both the time interval and the elapsed time duration; determining a corresponding second review duration based on a content difficulty level corresponding to the target program content, wherein the second review duration is positively correlated with the content difficulty level; The sum of the first review duration and the second review duration is used as the corresponding review duration.

17. The method according to claim 16, wherein The determining of the corresponding second review duration based on the content difficulty level corresponding to the target program content includes: If the difficulty level of the content is not greater than the preset level threshold, determining the second review time according to the broadcast time; wherein the review time is positively correlated with the broadcast time; If the content difficulty level is greater than a preset level threshold, the second review time is determined according to the played time and the content difficulty level; wherein the review time is positively correlated with both the played time and the content difficulty level.

18. The method according to claim 17, wherein If the difficulty level of the content is not greater than the preset level threshold, determining the second review duration according to the broadcast duration includes: If the difficulty level of the content is not greater than the preset level threshold, the product of the broadcast duration and the third preset ratio value is used as the second review duration; If the content difficulty level is greater than a preset level threshold, determining the second review time according to the broadcast time and the content difficulty level includes: If the content difficulty level is greater than a preset level threshold, the third preset ratio value is increased by a third preset step size for each increase in the content difficulty level to obtain a third ratio value, and the product of the broadcast duration and the third ratio value is used as the second review duration.

19. The method according to claim 12, wherein The step of converting the summary content text into audio based on key sound features in the audio content to be reviewed to obtain the continued listening summary content includes: If the target audio program content includes sounds of multiple objects, determining the sound with the highest proportion by extracting features from the multiple object sounds; Based on the sound with the highest proportion, the summary content text is converted into audio to obtain the continued listening summary content.

20. The method according to any one of claims 10 to 19, wherein: The method further comprises: After receiving the setting request for the resume play permission control in the permission setting interface sent by the client, the resume play permission information associated with the target object is obtained, and the resume play permission information is associated with the identification information of the target object and saved; wherein, the resume play permission control is used to turn on or off the smart resume listening mode; the smart resume listening mode refers to: after the audio program content is paused, an intelligent playback control function is used to play customized resume listening summary content according to the object's operation.

21. A device for controlling the playback of audio program content, characterized in that: include: a pause unit, configured to pause the playing of the target audio program content in response to a pause operation triggered on the target audio program content during the playing of the target audio program; A resume unit is configured to, in response to a resume operation triggered for the target audio program content, display a resume listening control area in the playback control interface and play the resume listening summary content corresponding to the target audio program content; wherein the review duration corresponding to the resume listening summary content is determined based on at least one of the following: the subject's historical listening behavior, the content difficulty level corresponding to the target audio program content; the resume listening playback control area is configured to control the playback status of the resume listening summary content; after the resume listening summary content is finished playing, the unplayed portion of the target audio program content is continued to be played from the current pause position.

22. A device for controlling the playback of audio program content, characterized in that: include: a determination unit configured to, upon receiving a pause request and a resume request for the target audio program content sent by a client during playback of the target audio program, determine a review duration based on at least one of the subject's historical listening behavior and a content difficulty level corresponding to the target audio program content; A generating unit, configured to generate a resume listening summary content based on the review duration and the audio content corresponding to the played portion of the target audio program content; The resume listening summary content is summary information of the played portion; The feedback unit is configured to feed back the resume listening summary content to the client, so that the client displays a resume listening control area in a playback control interface, plays the resume listening summary content, and continues to play the unplayed portion of the target audio program content from the current pause position after the resume listening summary content is finished.

23. An electronic device, characterized in that: It includes a processor and a memory, wherein the memory stores program code, and when the program code is executed by the processor, the processor executes the steps of any method described in claims 1 to 9 or any method described in claims 10 to 20.

24. A computer-readable storage medium, characterized in that The method comprises a program code, and when the program code is run on an electronic device, the program code is used to enable the electronic device to execute the steps of any method described in claims 1 to 9 or the steps of any method described in claims 10 to 20.

25. A computer program product, characterized in that The method comprises a computer program stored in a computer-readable storage medium; when a processor of an electronic device reads the computer program from the computer-readable storage medium, the processor executes the computer program, so that the electronic device performs the steps of the method described in any one of claims 1 to 9 or the steps of the method described in any one of claims 10 to 20.