Live broadcast content display method and device

By obtaining and displaying the audio content during live broadcast pause and converting it into a text list, the problem that users cannot view part of the live broadcast content in real time is solved, and the live viewing experience and information acquisition efficiency is improved.

CN120091150APending Publication Date: 2025-06-03SHANGHAI BILIBILI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510300085.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The existing live broadcast software cannot continue to play the content when it is paused after being paused in the middle. When the user briefly leaves the live broadcast and returns to the live broadcast room, he can only see the latest live broadcast content, and cannot view the unwatched content in real time, which affects the user's live broadcast viewing experience.

Method used

A live content display method is provided, which obtains the target audio content through a pause operation, converts it into the target text content, and displays the live content during the pause in the form of a text list when the user replays the live broadcast.

Benefits of technology

When users re-watch the live broadcast, they can quickly understand the live broadcast content during their departure, improve the efficiency of users to obtain live broadcast content, and thus improve the user's viewing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120091150A_ABST
    Figure CN120091150A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a live broadcast content display method and device, computer equipment, a computer readable storage medium and a computer program product, and relates to the technical field of video live broadcast. The method comprises the steps of playing live broadcast content of a target live broadcast room; in response to a pause operation, pausing playing of the live broadcast content and acquiring target audio content of the target live broadcast room; wherein the target audio content comes from a live broadcast segment output by the target live broadcast room during the playing pause period; obtaining target text content according to the target audio content; and displaying the target text content in the target live broadcast room when a re-playing operation for the live broadcast content is monitored. According to the technical scheme provided by the embodiment of the invention, the user who leaves the live broadcast for a short time can quickly know the live broadcast content in the leaving process when watching the live broadcast again, so that the efficiency of acquiring the live broadcast content by the user is improved, and the watching experience of the user is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the technical field of video live streaming, and in particular, to a live content display method, device, computer device, computer-readable storage medium, and computer program product. Background Art

[0002] With the rapid popularization of Internet technology, webcasting is being received and loved by more and more people. Webcasting generally involves: a live streaming platform, a host terminal, and an audience terminal. Among them, the host terminal can provide multimedia content (such as video content) to the audience terminal via the live streaming platform, and can also receive multimedia content (such as comment content) provided by the audience terminal via the live streaming platform, thus achieving the effect of interacting while live streaming. Due to the strong on-site and interactive nature of webcasting, it has been favored by more and more audiences and hosts.

[0003] However, live streaming software cannot resume playing the content at the pause point. When a user briefly leaves the live stream and then returns to the live stream room, they can only see the latest live content. After that, the user can only re-watch the unviewed part of the live content by viewing the live stream replay, and the process of finding the unviewed part is also time-consuming and laborious, affecting the user's live streaming viewing experience.

[0004] It should be noted that the above content is not necessarily prior art and does not limit the patent protection scope of the present application. Summary of the Invention

[0005] Embodiments of the present application provide a live content display method, device, computer device, computer-readable storage medium, and computer program product to solve or alleviate one or more of the above technical problems.

[0006] One aspect of the embodiments of the present application provides a live content display method, the method comprising: Playing the live content of a target live stream room; In response to a pause operation, pausing the playback of the live content and obtaining target audio content of the target live stream room; wherein, the target audio content comes from the live stream segment output by the target live stream room during the pause of playback; Obtaining target text content according to the target audio content; When a re-play operation for the live content is detected, displaying the target text content in the target live stream room.

[0007] Optionally, in response to a pause operation, pausing the playback of the live content and obtaining target audio content of the target live stream room, includes: When the pause operation is detected, determining the pause time; Cache live segments, where the live segments are the live content of the target live room within a preset time range after the pause time; and Obtain the target audio content according to the cached live segments.

[0008] Optionally, the target audio content includes one or more target audio segments, and the time length of each target audio segment is a preset time length. The target text content includes one or more target text segments. Obtaining the target text content according to the target audio content includes: Generate a corresponding target text segment each time a target audio segment is cached until one or more target text segments corresponding to one or more target audio segments are obtained.

[0009] Optionally, the target audio content further includes remaining audio segments, and the time length of the remaining audio segments is less than the preset time length. Obtaining the target text content according to the target audio content includes: Determine the remaining text segment corresponding to the remaining audio segment according to the remaining audio segment; Obtain the target text content according to one or more target text segments and the remaining text segment.

[0010] Optionally, the target text content is displayed in a list form in chronological order. The method further includes: Determine multiple target statements according to the target text content; Generate a text list according to multiple target statements; Wherein, in the text list, one target statement corresponds to one time node, and multiple target statements are sorted in chronological order.

[0011] Optionally, the method further includes: In response to selecting one of the multiple target statements, determine the target time node corresponding to the selected target statement; Display the live video corresponding to the target time node.

[0012] Another aspect of the embodiments of the present application provides a live content display device, and the device includes: A playback module for playing the live content of the target live room; A response module for responding to a pause operation, pausing the playback of the live content and obtaining the target audio content of the target live room; wherein, the target audio content comes from the live segments output by the target live room during the pause of playback; An acquisition module, configured to acquire target text content according to the target audio content; A display module, configured to display the target text content in the target live broadcast room when a re-play operation for the live broadcast content is detected.

[0013] Another aspect of the embodiments of the present application provides a computer device, including: At least one processor; and A memory communicatively connected to the at least one processor; Wherein: the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method as described above.

[0014] Another aspect of the embodiments of the present application provides a computer-readable storage medium, in which computer instructions are stored, and when the computer instructions are executed by a processor, the method as described above is implemented.

[0015] Another aspect of the embodiments of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method as described above is implemented.

[0016] The embodiments of the present application adopting the above technical solutions may include the following advantages: When re-playing after a live broadcast is paused, at least part of the live broadcast content during the pause is displayed in text form. Thus, users who briefly leave the live broadcast can quickly understand the live broadcast content during their absence when re-watching the live broadcast, improving the efficiency of users to obtain the live broadcast content, and thus improving the viewing experience of users. BRIEF DESCRIPTION OF THE DRAWINGS The drawings exemplarily show embodiments and form a part of the specification, and are used together with the written description of the specification to explain the exemplary embodiments of the embodiments. The shown embodiments are only for illustrative purposes and do not limit the scope of the claims. In all the drawings, the same reference numerals refer to similar but not necessarily identical elements.

[0017] Figure 1 Schematically shows an operating environment diagram of the live broadcast content display method according to Embodiment 1 of the present application; Figure 2 Schematically shows a flowchart of the live broadcast content display method according to Embodiment 1 of the present application; Figure 3 Schematically shows Figure 2 The sub-step flowchart of step S202 in Figure 4 Schematically shows Figure 2Sub - step flowchart of step S204; Figure 5 Schematically shows the new flowchart of the live content display method according to Embodiment 1 of the present application; Figure 6 Schematically shows another new flowchart of the live content display method according to Embodiment 1 of the present application; Figure 7 Schematically shows the application example flowchart of the live content display method according to Embodiment 1 of the present application; Figure 8 Schematically shows the block diagram of the live content display device according to Embodiment 2 of the present application; Figure 9 Schematically shows the hardware architecture diagram of the computer device according to Embodiment 3 of the present application; and Figure 10 Schematically shows the implementation effect diagram of the live content display method according to Embodiment 1 of the present application. Detailed implementation manners

[0018] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts belong to the scope of protection of the present application.

[0019] It should be noted that the descriptions involving "first", "second", etc. in the embodiments of the present application are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In addition, the technical solutions between various embodiments can be combined with each other, but it must be based on the fact that those of ordinary skill in the art can implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present application.

[0020] In the description of the present application, it should be understood that the numerical labels before the steps do not identify the sequence of execution of the steps, but are only used to facilitate the description of the present application and distinguish each step, and thus cannot be understood as a limitation to the present application.

[0021] First, provide the term explanations involved in the present application: ASR (Automatic Speech Recognition): is a technology that converts speech signals into text. By analyzing the acoustic features of speech, combining language models and acoustic models, human speech input is accurately translated into understandable text output.

[0022] Video coding format: a technical specification for compressing and storing video data. It uses a specific algorithm to remove redundant information, reduce the size of video files, and preserve visual quality as much as possible. Common video coding formats include H.264 (AVC), H.265 (HEVC), AV1, etc.

[0023] H.265: Also known as HEVC, High Efficiency Video Coding, is an advanced video compression standard designed to improve the efficiency of video encoding. Compared with the previous H.264 standard, H.265 can reduce the amount of data by about 50% at the same video quality, thereby significantly reducing the bandwidth and storage space required for video transmission.

[0024] Live streaming: refers to the process of encoding and processing audio and video data from the live source (such as the anchor's camera, microphone or other acquisition equipment) and transmitting it to the live server through the network. This process compresses and encodes the collected original audio and video data, converts it into a format suitable for network transmission (such as H.264 video encoding and AAC audio encoding), and then pushes the data to the server through protocols such as RTMP (Real Time Messaging Protocol) and HLS (HTTP Live Streaming).

[0025] RTMP (Real-Time Messaging Protocol): is a protocol for efficiently transmitting audio, video and data on the Internet. It supports low-latency real-time streaming transmission and can realize functions such as live video, online games, and real-time communication. It uses the TCP protocol for reliable transmission to ensure data integrity and sequence, and supports encoding of multiple media formats, such as H.264 video and AAC audio.

[0026] NLP (Natural Language Processing): is an important branch of artificial intelligence that aims to enable computers to understand, generate and process human language. By combining linguistics, computer science and machine learning techniques, machines can parse the grammar, semantics and pragmatic rules of language, thereby achieving natural language interaction with humans.

[0027] Secondly, in order to facilitate those skilled in the art to understand the technical solutions provided in the embodiments of the present application, the relevant technologies are described below: In current live streaming software, it is basically impossible to pause during real-time live streaming. When a user briefly leaves the software and then returns to the live streaming room, what they see is the latest live content, and the content during the period of leaving is invisible. Most live streaming software provides a way to view past live content: live replay. After the live streaming ends, users can view the complete replay of the host's live content during this time. Users can find the parts they haven't watched in the replay video. However, this method cannot view the content during the brief leave in real time. It takes some time after the live streaming ends, and users need to wait for a long time. Moreover, it is not easy to find the missed part from the live content that can be dozens of minutes or even several hours long. A small number of software also provides a method. That is, after leaving or when the live streaming lags, it will continue to cache the live content and play it back at a multiple speed after coming back, and then catch up with the normal live streaming progress. However, because the video is played back at a multiple speed too fast, it is difficult to understand the host's voice.

[0028] For this reason, the embodiments of this application provide a technical solution for displaying live content. In this technical solution, the live content during the period when the user pauses the live streaming can be cached, and when the user replays the live streaming, the live content during the pause period is displayed in the form of a text list, so as to try to restore the content of what the host said during the user's leave, so that the user can still obtain the most core information of the live streaming room during this time period. See the following text for details.

[0029] Finally, for the convenience of understanding, an exemplary operating environment is provided below.

[0030] As Figure 1 shown, the operating environment diagram includes: service platform 2, host terminals (4A, 4B,..., 4M), and user terminals (6A, 6B,..., 6N).

[0031] In the live streaming scenario, the host terminals (4A, 4B,..., 4M) log in to the service platform 2, and the live data is pushed to the user terminals (6A, 6B,..., 6N) in real time through the service platform 2.

[0032] The service platform 2 can provide live streaming room services, and it can be a single server, a server cluster, or a cloud computing service center.

[0033] The host terminals (4A, 4B,..., 4M) are used to generate live data in real time and perform the operation of pushing the live data. The live data can include audio data or video data. The host terminal can be an electronic device such as a smart phone or a tablet computer. Of course, the host terminal can be a virtual computing instance in the service platform 2.

[0034] The client terminals (6A, 6B, …, 6N) can be configured to receive the live data of the host terminal in real time. The client terminals (6A, 6B, …, 6N) can be any type of computing device, such as a smart phone, a tablet device, a laptop computer, a smart TV, an in-vehicle terminal, etc. The client terminals (6A, 6B, …, 6N) can be built with a browser or a dedicated program, and receive the live data through the browser or the dedicated program to output content to the user. The content can include video, audio, comments, text data, and / or the like.

[0035] The client terminals (6A, 6B, …, 6N) can include a player. The player outputs (e.g., displays, presents) content to the user. The content can include video, audio, comments, text data, and / or the like. The client terminals (6A, 6B, …, 6N) can include an interface, and the interface can include an input element (touch screen). For example, the input element can be configured to receive user instructions, and the user instructions can enable the client terminals (6A, 6B, …, 6N) to perform various operations, such as sending bullet comments, inputting comments, giving gifts, etc.

[0036] The host terminals (4A, 4B, …, 4M), the client terminals (6A, 6B, …, 6N), and the service platform 2 can be connected through a network. The network can include various network devices, such as routers, switches, multiplexers, hubs, modems, bridges, repeaters, firewalls, and / or proxy devices, etc. The network can include physical links, such as coaxial cable links, twisted pair cable links, optical fiber links, and their combinations and / or the like. The network can include wireless links, such as cellular links, satellite links, Wi-Fi links, and / or the like.

[0037] It should be noted that the numbers of the host terminals and the client terminals in the figure are only illustrative and are not used to limit the patent protection scope of the present application. According to the actual situation, there can be any number of host terminals and client terminals.

[0038] The following takes the client terminal 6A as the execution subject and introduces the technical solutions of the present application through multiple embodiments. It should be noted that these embodiments can be implemented in various different forms and should not be construed as being limited only to the embodiments described herein.

[0039] Embodiment 1 Figure 2 The flowchart of the live content display method according to Embodiment 1 of the present application is schematically shown.

[0040] As Figure 2 shown, the live content display method can include steps S200 to S206, which are for the client terminal, where: Step S200, play the live content of the target live room.

[0041] Step S202, in response to a pause operation, pause playing the live content and obtain the target audio content of the target live room; wherein, the target audio content comes from the live clip output by the target live room during the pause of playing.

[0042] Step S204, obtain target text content according to the target audio content.

[0043] Step S206, when a re-play operation for the live content is detected, display the target text content in the target live room.

[0044] The live content display method provided in this embodiment displays the live content during the pause in the form of text when re-playing after the live is paused. Thus, users who briefly leave the live can quickly understand the live content during their absence when re-watching the live, improving the efficiency of users to obtain the live content and thus enhancing the viewing experience of users.

[0045] The following will be combined with Figure 2 to elaborate in detail on each step in steps S200 to S206 and other optional steps.

[0046] Step S200 , play the live content of the target live room.

[0047] In some embodiments, the target live room can be a game live room, an entertainment live room, an education live room, etc. In other embodiments, the target live room can also be a voice podcast or a meeting live room (such as an online meeting room), etc.

[0048] When the target live room is a live room based on the host's streaming, the live content can include the audio-visual content pushed by the host.

[0049] When the target live room is a multi-person meeting live room, the live content can be the audio-visual content obtained by synthesizing the multi-person streaming.

[0050] Step S202 , in response to a pause operation, pause playing the live content and obtain the target audio content of the target live room; wherein, the target audio content comes from the live clip output by the target live room during the pause of playing.

[0051] In some embodiments, the pause operation can be that the user manually clicks the pause button. In other embodiments, the pause operation can also be triggered by means of voice commands, gesture recognition, device sensors (such as the user leaving the device by a certain distance), etc.

[0052] The target audio content can be the speech content of the host. When multiple roles (such as the host and the connected audience, or multiple participants in an online meeting room) are speaking in the target live broadcast room, the speech content of different roles can also be recognized through speech recognition or other means. At this time, only the speech content of the host can be used as the target audio content, or the speech content of all roles can be used as the target audio content and the speaking roles can be marked.

[0053] After obtaining the target audio content, preprocessing can also be performed on the target audio content, such as noise reduction, enhancement, format conversion, etc. on the audio.

[0054] After the user triggers the pause operation, continuously obtain the target audio content during the pause playback period. Thus, it is possible to retain the key content that may appear during the pause for the user after the user temporarily leaves the target live broadcast room, thereby providing support for subsequent text content extraction.

[0055] In the actual application process, the target audio content can be obtained in various ways. The following provides an exemplary way.

[0056] In an alternative embodiment, as Figure 3 shown, step S202 includes: S300, when the pause operation is detected, determine the pause time.

[0057] S302, cache the live broadcast segment, where the live broadcast segment is the live broadcast content of the target live broadcast room within a preset time range after the pause time.

[0058] S304, obtain the target audio content according to the cached live broadcast segment.

[0059] In some embodiments, the preset time range can be fixedly set to 2 minutes, 5 minutes, 8 minutes, etc. In other embodiments, the preset time range can also be dynamically set according to the network condition or the storage pressure.

[0060] Format optimization can be performed on the cached live broadcast segment, for example, using a more efficient coding format (such as H.265) to reduce the storage space and bandwidth occupancy. A cache cleaning mechanism can also be set to regularly clean the cached live broadcast segments. If the live broadcast segment is lost during the caching period (such as network interruption), the missing part can also be requested to be retransmitted from the host side through a live stream protocol (such as RTMP).

[0061] Cache the live broadcast segment within a preset time range after the pause operation is triggered to obtain the target audio content. Thus, it is possible to prevent excessive cache occupation of storage while ensuring that the key live broadcast content during the pause is not missed, saving storage resources.

[0062] Step S204 , according to the target audio content, obtain the target text content.

[0063] The target audio content can be identified through ASR automatic speech recognition technology to obtain the target text content. It is also possible to combine multiple speech recognition models (such as neural network models based on deep learning, traditional statistical models, etc.) to improve the accuracy and robustness of speech recognition through model fusion. The language of the target text content can be Chinese or the system language set by the user.

[0064] After obtaining the target text content, it can automatically correct errors and optimize it by combining the context semantics and the live broadcast scene knowledge base (such as common terms used by anchors and live broadcast thesaurus). It can also conduct real-time review of the target text content and filter or replace illegal content (such as false, sensitive words, etc.).

[0065] In this embodiment, the live broadcast content during the pause is extracted in the form of text. Thus, when the user re-watches the live broadcast, he or she can quickly understand the key information said by the anchor during the pause, thereby improving the viewing experience and information acquisition efficiency.

[0066] As mentioned above, the target audio content is extracted from the cached live segment, so the target text content can be obtained from all the cached target audio content at once after the cache is completed, or it can be gradually obtained from part of the cached target audio content during the cache process. The following provides an exemplary method for obtaining the target text content during the cache process.

[0067] In an optional embodiment, the target audio content includes one or more target audio segments, the time length of each target audio segment is a preset time length, and the target text content includes one or more target text segments, and step S204 includes: Each time a target audio segment is cached, a corresponding target text segment is generated until one or more target text segments corresponding to one or more target audio segments are obtained.

[0068] The preset time length can be set to a fixed 10s, 20s, etc., or it can be dynamically set according to the device performance of the user end or the pressure of the speech recognition interface (such as the ASR interface). It can also be dynamically adjusted according to the complexity of the audio content (such as speech speed, background noise, etc.).

[0069] In this embodiment, the target audio content is parsed in the form of audio segments to generate the target text content. Thus, it is possible to avoid processing too long target audio content at one time, thereby reducing the complexity of speech recognition and the consumption of computing resources, thereby ensuring the quality and efficiency of generating the target text content and improving the user viewing experience.

[0070] In the actual application process, it may occur that the user returns to the live broadcast room before the cached live broadcast segment reaches the preset time range. At this time, the cached audio content may not reach the preset time length, that is, it is not a complete target audio segment. In this case, this part of the audio content can be directly omitted, or the text content of this part of the audio content can be continuously generated. The following provides an exemplary solution.

[0071] In an alternative embodiment, the target audio content further includes a remaining audio segment, and the time length of the remaining audio segment is less than the preset time length. As Figure 4 shown, step S204 includes: S400, determining a remaining text segment corresponding to the remaining audio segment according to the remaining audio segment.

[0072] S402, obtaining the target text content according to one or more of the target text segments and the remaining text segment.

[0073] When determining the remaining text segment, the remaining audio segment can be semantically associated with the obtained target text segments, avoiding isolated understanding and ensuring the correctness of the remaining text segment.

[0074] In this embodiment, the remaining audio segment with a length less than the preset time length is parsed to generate a corresponding text segment. Thereby, it is possible to ensure that all the audio content during the pause is obtained, enabling the user to understand all the live broadcast content during the pause and improving the user's viewing experience.

[0075] Step S206 , in the case of detecting a re-play operation for the live broadcast content, the target text content is displayed in the target live broadcast room.

[0076] The style of the text (such as font, color, animation effect, etc.) can be adaptively set according to the user's preference or the theme of the target live broadcast room. The target text content can be displayed in the form of a floating window, a separate pane, or a bullet screen. For example, the target text content in the floating window can be displayed automatically or scrolled. Additionally, the display position can be a fixed position or a movable position.

[0077] In some embodiments, it is also possible to support the user to search for the displayed target text content, and the search results are highlighted to facilitate the user to quickly find the interested part. And it supports the user to directly select the target text content for sharing or reporting, etc.

[0078] In some embodiments, keywords such as technical terms and popular words in the target text content can also be automatically recognized, and corresponding associated links can be set on the display interface to support users to learn about the specific content of these keywords through the links.

[0079] In this embodiment, when the user replays the live broadcast, the text content during the pause is displayed. Thereby, the user can intuitively and quickly understand the live broadcast content during the pause, improving the efficiency of the user to obtain the missed live broadcast content and the viewing experience of the live broadcast.

[0080] As described above, the target text content can be displayed in various forms. The following provides an exemplary display method.

[0081] In an alternative embodiment, the target text content is displayed in a list form in chronological order, as Figure 5 shown, the method further includes: S500, determining a plurality of target statements according to the target text content.

[0082] S502, generating a text list according to the plurality of target statements, wherein in the text list, one target statement corresponds to one time node, and the plurality of target statements are sorted in chronological order.

[0083] The text list can have a fixed display form, or can be adaptively adjusted according to user needs or the theme of the target live broadcast room, such as adjusting the font size, color, list style, etc.

[0084] The text list can include language options to translate the target text content according to the language selected by the user. In some embodiments, it can also include a close button. When the user clicks the close button, the display of the text list will be closed.

[0085] In some embodiments, sentiment analysis can also be performed on the target text content, and the sentiment tendency of the anchor can be marked in the list. The text list can also include a key content summary of the target text content extracted through natural language processing technology (NLP), etc.

[0086] In some embodiments, the time node can be the absolute time of the target statement (such as Beijing time). In other embodiments, the time node can also be a relative time relative to the start or pause time of the live broadcast. When the time node is a relative time relative to the pause time, the specific pause time can also be displayed in the text list.

[0087] In this embodiment, the target text content is generated into a text list according to the time sequence and displayed. Thus, users can quickly browse the key information missed during the pause, and at the same time can quickly locate the parts they are interested in according to the time nodes, improving the efficiency of information acquisition.

[0088] In addition to the above functions, the text list can also have a variety of other functions. The following provides an exemplary function.

[0089] In an alternative embodiment, as Figure 6 shown, the method further includes: S600, in response to selecting one of the multiple target statements, determining the target time node corresponding to the selected target statement.

[0090] S602, displaying the live broadcast screen corresponding to the target time node.

[0091] With the user's authorization, according to the user's viewing history and preferences, statements that may be of interest can be automatically recommended and highlighted. The user can also be allowed to select multiple statements at once and view the time nodes corresponding to these statements, facilitating the user to quickly browse multiple key points.

[0092] In some embodiments, the live broadcast screen corresponding to the target time node can be directly displayed in the main window of the live broadcast room. In other embodiments, the live broadcast screen corresponding to the target time node can also be displayed in a pop-up window, picture-in-picture, etc. The target statement can also be displayed in the form of subtitles on the live broadcast screen at the same time.

[0093] After the live broadcast screen corresponding to the target time node finishes playing, the live broadcast screen corresponding to the subsequent target statement can be continued to play, or it can be directly paused, and the user can be asked whether to continue playing the subsequent screen or return to the latest live broadcast screen.

[0094] In this embodiment, precise positioning of the live broadcast content is achieved by selecting the target statement, which can facilitate users to quickly find the content they are interested in, saving the users' time. By combining text and pictures, a richer and more efficient live broadcast viewing experience is provided.

[0095] To make the present application easier to understand, the following combines Figure 7 and Figure 10 to provide an exemplary application, where Figure 10 1 in is the real-time live broadcast screen of the live broadcast room, 2 is the text list, and 3 is the close button of the text list. Among them: S11, the user enters the classroom live broadcast room O (i.e., the target live broadcast room) through the mobile phone P (i.e., the user terminal), and the live broadcast room O is conducting a classroom live broadcast (i.e., the live broadcast content); S12. The user suddenly has something urgent and needs to leave. The user clicks the pause button on the mobile phone P (i.e., the pause operation). The mobile phone P pauses the live broadcast screen and records the pause time as 10 minutes and 11 seconds after the classroom live broadcast room O starts the live broadcast; S13. The mobile phone P continues to pull the live stream of the classroom live broadcast room O and caches the subsequent live video (i.e., the live clip); S14. Every 10s (i.e., the preset time length) the mobile phone P caches the live video, it performs ASR automatic speech recognition on the audio track of the cached live video (i.e., the target audio clip), translates the content spoken by the anchor into Chinese, and obtains the Chinese text (i.e., the target text clip); S15. The user comes back after 1 minute and 5 seconds and clicks the play button on the mobile phone P (i.e., the re-play operation). The mobile phone P continues to play the latest live broadcast screen; S16. The mobile phone P performs ASR automatic speech recognition on the audio track of the last 5 seconds of the cached live video (i.e., the remaining audio clip), translates the content spoken by the anchor into Chinese, and obtains the Chinese text (i.e., the remaining text clip); S17. The mobile phone P combines all the obtained texts and performs clause segmentation according to semantics to obtain multiple sentences (i.e., the target statements); S18. The mobile phone P displays a text list on the right side of the live broadcast screen. The time of this pause, "10:11", is displayed above the list. The multiple obtained sentences and the relative time of each sentence with respect to the pause time are sequentially displayed in the list according to time. A close button is displayed below the list; S19. The user clicks one of the sentences, and the relative time of this sentence is "00:36". The mobile phone P automatically jumps the live broadcast screen to 10 minutes and 47 seconds after the live broadcast starts.

[0096] Embodiment 2 Figure 8 The block diagram of the live content display device according to Embodiment 2 of the present application is schematically shown. The device can be divided into one or more program modules. One or more program modules are stored in a storage medium and executed by one or more processors to complete the embodiments of the present application. The program modules referred to in the embodiments of the present application refer to a series of computer program instruction segments that can complete specific functions. The following description will specifically introduce the functions of each program module in this embodiment. As Figure 8 shown, the device 1000 may include: a playback module 1100, a response module 1200, an acquisition module 1300, and a display module 1400, where: The playback module 1100 is configured to play the live content of the target live broadcast room; A response module 1200, configured to pause the playback of the live content and obtain target audio content of the target live broadcast room in response to a pause operation; wherein, the target audio content comes from a live segment output by the target live broadcast room during the paused playback; An acquisition module 1300, configured to obtain target text content according to the target audio content; A display module 1400, configured to display the target text content in the target live broadcast room when a re-play operation for the live content is detected.

[0097] As an optional embodiment, the response module 1200 is further configured to: Determine a pause time when the pause operation is detected; Cache a live segment, where the live segment is the live content of the target live broadcast room within a preset time range after the pause time; and Obtain the target audio content according to the cached live segment.

[0098] As an optional embodiment, the acquisition module 1300 is further configured to: Generate a corresponding target text segment every time a target audio segment is cached until one or more target text segments corresponding to one or more target audio segments are obtained.

[0099] As an optional embodiment, the acquisition module 1300 is further configured to: Determine a remaining text segment corresponding to the remaining audio segment according to the remaining audio segment; Obtain the target text content according to one or more target text segments and the remaining text segment.

[0100] As an optional embodiment, the apparatus 1000 further includes a text list generation module, configured to: Determine a plurality of target statements according to the target text content; Generate a text list according to the plurality of target statements; Wherein, in the text list, one target statement corresponds to one time node, and the plurality of target statements are sorted in chronological order.

[0101] As an optional embodiment, the apparatus 1000 further includes a live video display module, configured to: Determine a target time node corresponding to the selected target statement in response to selecting one of the plurality of target statements; Display the live video corresponding to the target time node.

[0102] Embodiment III Figure 9 Schematically shown is a hardware architecture diagram of a computer device 10000 suitable for implementing the live content display method according to Embodiment 3 of the present application. In some embodiments, the computer device 10000 may be a terminal device such as a smart phone, a wearable device, a tablet computer, a personal computer, a vehicle-mounted terminal, a game console, a virtual device, a workbench, a digital assistant, a set-top box, a robot, etc. In other embodiments, the computer device 10000 may be a rack server, a blade server, a tower server, or a cabinet server (including an independent server or a server cluster composed of multiple servers), etc. As Figure 9 shown, the computer device 10000 includes, but is not limited to: a memory 10010, a processor 10020, and a network interface 10030 that can be communicatively linked to each other through a system bus. Among them: The memory 10010 includes at least one type of computer-readable storage medium. The readable storage medium includes flash memory, a hard disk, a multimedia card, a card-type memory (such as an SD or DX memory), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory 10010 may be an internal storage module of the computer device 10000, such as the hard disk or memory of the computer device 10000. In other embodiments, the memory 10010 may also be an external storage device of the computer device 10000, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc., equipped on the computer device 10000. Of course, the memory 10010 may also include both the internal storage module and the external storage device of the computer device 10000. In this embodiment, the memory 10010 is generally used to store the operating system and various application software installed on the computer device 10000, such as the program code of the live content display method. In addition, the memory 10010 may also be used to temporarily store various data that have been output or will be output.

[0103] In some embodiments, the processor 10020 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other chips. The processor 10020 is generally used to control the overall operation of the computer device 10000, such as performing control and processing related to data interaction or communication with the computer device 10000. In this embodiment, the processor 10020 is used to run the program code stored in the memory 10010 or process data.

[0104] The network interface 10030 may include a wireless network interface or a wired network interface, which is generally used to establish a communication link between the computer device 10000 and other computer devices. For example, the network interface 10030 is used to connect the computer device 10000 to an external terminal through a network, and establish a data transmission channel and a communication link between the computer device 10000 and the external terminal. The network may be a wireless or wired network such as an enterprise intranet (Intranet), the Internet, Global System of Mobile communication (GSM for short), Wideband Code Division Multiple Access (WCDMA for short), 4G network, 5G network, Bluetooth, Wi-Fi, etc.

[0105] It should be noted that Figure 9 only the computer device with components 10010 - 10030 is shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components may be implemented alternatively.

[0106] In this embodiment, the live content display method stored in the memory 10010 may also be divided into one or more program modules and executed by one or more processors (such as the processor 10020) to complete the embodiments of the present application.

[0107] Embodiment 4 The embodiments of the present application also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the live content display method in the embodiments are implemented.

[0108] In this embodiment, the computer-readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the computer-readable storage medium may be an internal storage unit of a computer device, such as the hard disk or memory of the computer device. In other embodiments, the computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc., equipped on the computer device. Of course, the computer-readable storage medium may also include both the internal storage unit and the external storage device of the computer device. In this embodiment, the computer-readable storage medium is generally used to store the operating system installed on the computer device and various application software, such as the program code of the live content display method in the embodiment. In addition, the computer-readable storage medium can also be used to temporarily store various data that have been output or will be output.

[0109] Embodiment 5 The embodiment of the present application also provides a computer program product, including a computer program, which implements the method in the above embodiment when executed by a processor.

[0110] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the embodiments of the present application can be implemented by a general-purpose computer device. They can be concentrated on a single computer device or distributed on a network composed of multiple computer devices. Optionally, they can be implemented by program codes executable by the computer device, so that they can be stored in a storage device and executed by the computer device. And in some cases, the steps shown or described can be executed in a different order from here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module to implement. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.

[0111] It should be noted that the above are only the preferred embodiments of the present application, and do not limit the patent protection scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.

Claims

1. A method for displaying live content, characterized in that: For a user terminal, the method includes: Play the live content of the target live broadcast room; In response to the pause operation, the live broadcast content is paused and the target audio content of the target live broadcast room is acquired; wherein the target audio content comes from the live broadcast segment output by the target live broadcast room during the pause period; According to the target audio content, obtaining target text content; When a replay operation on the live broadcast content is detected, the target text content is displayed in the target live broadcast room.

2. The method according to claim 1, characterized in that In response to the pause operation, pausing the playing of the live content and acquiring the target audio content of the target live room, including: In case the pausing operation is detected, determining the pausing time; caching a live broadcast segment, where the live broadcast segment is the live broadcast content of the target live broadcast room within a preset time range after the pause time; and The target audio content is obtained according to the cached live broadcast segment.

3. The method according to claim 2, characterized in that The target audio content includes one or more target audio segments, each of which has a preset time length, and the target text content includes one or more target text segments; and obtaining the target text content according to the target audio content includes: Each time a target audio segment is cached, a corresponding target text segment is generated until one or more target text segments corresponding to one or more target audio segments are obtained.

4. The method according to claim 3, characterized in that The target audio content also includes a remaining audio segment, and the time length of the remaining audio segment is less than the preset time length; and obtaining the target text content according to the target audio content includes: Determining, according to the remaining audio segments, remaining text segments corresponding to the remaining audio segments; The target text content is acquired according to one or more of the target text segments and the remaining text segments.

5. The method according to claim 1, characterized in that The target text content is displayed in chronological order and in list form; the method further includes: Determining a plurality of target sentences according to the target text content; Generate a text list according to the plurality of target sentences; Among them, in the text list, one target sentence corresponds to one time node, and multiple target sentences are sorted in chronological order.

6. The method according to claim 5, characterized in that The method further comprises: In response to selecting one of the target sentences from the plurality of target sentences, determining a target time node corresponding to the selected target sentence; Display the live screen corresponding to the target time node.

7. A live content display device, characterized in that: The device comprises: The playback module is used to play the live content of the target live broadcast room; A response module, configured to, in response to a pause operation, pause the playing of the live content and obtain target audio content of the target live room; wherein the target audio content comes from a live segment output by the target live room during the pause period; An acquisition module, used to acquire target text content according to the target audio content; A display module is used to display the target text content in the target live broadcast room when a replay operation on the live broadcast content is detected.

8. A computer device, characterized in that: include: at least one processor; and a memory communicatively coupled to the at least one processor; wherein: The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.