Method and apparatus for checking audiovisual content, and device and storage medium

By receiving selections for multiple text clips and creating continuous audio-visual content, the problem of users having difficulty quickly positioning specific speakers in audio-visual content is solved, and more efficient content acquisition and sharing is achieved.

WO2024093442A9PCT designated stage expired Publication Date: 2025-06-05BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2023/113406
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-10-31
Filing Date
2023-08-16
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

When playing audio-visual content, it is difficult for users to quickly locate the corresponding part of a specific speaker in the video or audio record, especially during long meetings, speeches or online classes.

Method used

A method and device are provided to create segment audio-visual content by receiving selections for multiple text segments, such that discontinuous parts become continuous in segment audio-visual content, and present a sharing portal for sharing these segment audio-visual content.

Benefits of technology

Users can easily view parts of specific speakers, thereby improving the efficiency of obtaining expected content and supporting quick sharing of audio-visual content, improving the efficiency and flexibility of audio-visual content sharing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2023113406_05062025_PF_FP_ABST
    Figure CN2023113406_05062025_PF_FP_ABST
Patent Text Reader

Abstract

According to the embodiments of the present disclosure, provided are a method and apparatus for checking audiovisual content, and a device and a storage medium. The method comprises: providing a check interface for audiovisual content, wherein the check interface comprises a playing control for playing the audiovisual content; and presenting at least one speaking time axis in the playing control, wherein the at least one speaking time axis is used for indicating the temporal distribution of spoken content of at least one speaker associated with the audiovisual content. In this way, by means of the embodiments of the present disclosure, a user is aided in acquiring speaking time distribution information of each speaker in audiovisual content, thereby helping the user to acquire required information more conveniently.
Need to check novelty before this filing date? Find Prior Art

Description

Method, apparatus, device and storage medium for viewing audiovisual content

[0001] This application claims priority to the Chinese invention patent application entitled “Methods, devices, equipment and storage media for viewing audiovisual content” filed on October 31, 2022, with application number 202211352393.3, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] Example embodiments of the present disclosure relate generally to the field of computers, and more particularly to methods, apparatuses, devices, and computer-readable storage media for viewing audiovisual content. Background Art

[0003] With the development of computer technology, the Internet has become the main platform for people to obtain and share content. For example, people can use the Internet to publish a variety of content or receive content shared by other users.

[0004] Sharing audiovisual content (e.g., audio or video) has become one of the most common forms of internet-based content sharing. For example, users can use a player to play a video or audio recording of a speech or meeting shared by another user. However, during such playback, it is difficult to quickly locate the corresponding portion of a specific speaker in the video or audio recording.

[0005] Summary of the Invention

[0006] In a first aspect of the present disclosure, a method for viewing audiovisual content is provided. The method includes: receiving a selection of multiple text segments, the multiple text segments corresponding to multiple sections in target audiovisual content, the multiple sections including at least a first section and a second section that are discontinuous in the target audiovisual content; creating segmented audiovisual content based on at least the multiple sections of the target audiovisual content, wherein the first section and the second section are contiguous in the segmented audiovisual content; and presenting a sharing portal for sharing the segmented audiovisual content.

[0007] In a second aspect of the present disclosure, a device for viewing audiovisual content is provided. The device includes a receiving module configured to receive selections of multiple text segments, the multiple text segments corresponding to multiple sections in target audiovisual content, the multiple sections including at least a first section and a second section that are discontinuous in the target audiovisual content; a control module configured to create segmented audiovisual content based on at least the multiple sections of the target audiovisual content, wherein the first section and the second section are continuous in the segmented audiovisual content; and a presentation module configured to present a sharing portal for sharing the segmented audiovisual content.

[0008] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the method of the first aspect.

[0009] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein a computer program is stored on the medium, and when the program is executed by a processor, the method of the first aspect is implemented.

[0010] In a fifth aspect of the present disclosure, a playback system is provided. The playback system includes: a main timeline indicating at least the current playback position of audiovisual content; and at least one speech timeline indicating the temporal distribution of speech content of at least one speaker associated with the audiovisual content.

[0011] It should be understood that the contents described in the summary of the present invention are not intended to limit the key features or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:

[0013] FIG1 shows a schematic diagram of a conventional audio-visual content player;

[0014] 2A to 2C illustrate schematic diagrams of example playback systems according to some embodiments of the present disclosure;

[0015] 3A and 3B illustrate example viewing interfaces for audiovisual content according to some embodiments of the present disclosure;

[0016] 4A and 4B are schematic diagrams showing sharing of audiovisual content segments according to some embodiments of the present disclosure;

[0017] FIG5 illustrates a flow chart of an example process for viewing audiovisual content according to some embodiments of the present disclosure;

[0018] FIG6 shows a block diagram of an apparatus for viewing audiovisual content according to some embodiments of the present disclosure; and

[0019] FIG7 shows a block diagram of a device capable of implementing various embodiments of the present disclosure. DETAILED DESCRIPTION

[0020] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0021] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, i.e., "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may be included below.

[0022] As discussed above, people can use players to obtain audio-visual content. Figure 1 shows a schematic diagram of a traditional audio-visual content player 100. As shown in Figure 1, in the player 100, people usually need to drag the time axis control to locate the desired playback moment.

[0023] However, such playback control is inefficient. For example, in the example of FIG1 , the audiovisual content has a length of more than 1 hour, which makes it difficult for the user to quickly locate the desired playback position through the time axis.

[0024] This situation is particularly evident in the review of audiovisual content such as conferences, lectures, or online classes, which often involve multiple speakers, and people may want to quickly locate the part of a particular speaker's speech.

[0025] Embodiments of the present disclosure provide a system for playing audiovisual content (audio or video). The system may include a main timeline to indicate at least the current playback position of the audiovisual content. Furthermore, the system may include at least one speech timeline to indicate the temporal distribution of speech content by at least one speaker associated with the audiovisual content.

[0026] Furthermore, embodiments of the present disclosure provide a solution for viewing audiovisual content. This solution provides a viewing interface for the audiovisual content, wherein the viewing interface includes playback controls for playing the audiovisual content. Furthermore, at least one speech timeline can be presented within the playback controls, indicating the temporal distribution of the speech content of at least one speaker associated with the audiovisual content.

[0027] Based on this approach, embodiments of the present disclosure can provide a speech timeline within a playback system or playback controls to provide a temporal distribution of the speech content of speakers associated with the audiovisual content. This makes it easier for users to view the portion corresponding to a specific speaker, thereby improving the efficiency of users in accessing desired content.

[0028] Hereinafter, exemplary solutions according to embodiments of the present disclosure will be described in detail with reference to the accompanying drawings.

[0029] Sample playback system

[0030] In some embodiments, embodiments of the present disclosure can utilize a timeline to provide richer information about audio-visual content.

[0031] Figure 2A shows a schematic diagram 200A of an example playback system 205 according to some embodiments of the present disclosure. As shown in Figure 2A, playback system 205 (also referred to as player 205 or playback control 205) can be used to play corresponding audiovisual content. Playback system 205 can be provided by, for example, an appropriate electronic device. Examples of such electronic devices include, but are not limited to, desktop computers, laptop computers, smartphones, tablet computers, personal digital assistants, or smart wearable devices.

[0032] In some embodiments, the audiovisual content may include audio and video files locally stored in the audiovisual system 205, audio and video files stored in the cloud, or audio and video streams. Such audio and video streams may include, for example, playback streams of recorded audiovisual content (e.g., meeting minutes) or live streams of live audiovisual content.

[0033] Main Timeline

[0034] As shown in Figure 2A, the playback system 205 may include a main timeline 210. In some embodiments, the main timeline 210 may indicate the current playback position of the audio-visual content, that is, the playback progress. For example, the main timeline 210 may include a playback position indicator 215 to indicate the time point at which the audio-visual content is currently being played.

[0035] For example, if the audio-visual content to be played is a recorded content, the length information of the audio-visual content is fixed, and the total length of the time axis can correspond to the total duration of the time content. In addition, the play position indicator 215 can be set accordingly according to the corresponding relationship between the position and the play time.

[0036] In yet other examples, if the audiovisual content to be played is live content with increasing duration, the play position indicator 215 may, for example, always be set to the far right of the main timeline 210. In addition, if the user desires to replay specific content that has already been broadcast live, the user may, for example, move the play position indicator 215 to jump back to the corresponding time point.

[0037] In some embodiments, as shown in FIG1 , the main timeline 210 can also present graphical information corresponding to the audio waveform of the audiovisual content. In this way, users can more easily understand which parts of the audiovisual content are worth paying attention to and which parts, such as those with less audio waveforms, can be temporarily ignored. Thus, such a playback system 205 can improve user access to content.

[0038] In some embodiments, the audiovisual content may be, for example, a recording of an online conference. Accordingly, as shown in FIG1 , the main timeline 210 may also present, for example, an interaction identifier 220 corresponding to an interaction behavior in the online conference.

[0039] Such an interaction mark 220 can be set at a corresponding position of the main time axis 210 to indicate that a corresponding interaction behavior has occurred at a corresponding moment. In some embodiments, different graphics of the interaction mark 220 can correspond to different interaction behaviors.

[0040] In some embodiments, the main timeline 210 may include, for example, an interactive identifier 220 for indicating file sharing in an online meeting. Accordingly, the interactive identifier 220 may include, for example, a graphic corresponding to the format of the shared file, for example, a thumbnail of the shared file.

[0041] In some embodiments, when a user selects the interactive marker 220, the playback system 205 may, for example, guide the user to obtain descriptive information about the shared file. For example, when the user hovers over the interactive marker 220 with a mouse, the playback system 205 may, for example, display information about the shared file, such as the file name, format, size, sharer, etc., in a floating window. In another example, if the user clicks the interactive marker 220, the playback system 205 may guide the user to obtain the content of the shared file, for example, by directing the user to an online viewing interface for the file.

[0042] In some embodiments, the main timeline 210 may include, for example, an interaction identifier 220 for indicating an online chat in an online conference. Online chat herein refers to any appropriate chat based on text, emoticons, images, and / or audio, for example, conducted using an instant messaging tool for the online conference. Accordingly, the graphical identifier of the interaction identifier 220 may be determined, for example, based on the content of the online chat. Alternatively, the graphical identifier of the interaction identifier 220 may be determined, for example, based on the graphical identifiers (e.g., avatars) of the users participating in the chat.

[0043] In some embodiments, when a user selects the interactive marker 220, the playback system 205 may, for example, guide the user to obtain a description of the online chat. For example, when the user hovers over the interactive marker 220 with a mouse, the playback system 205 may display information about the online chat, such as the participants in the online chat, the chat content, etc., in a floating window. In another example, if the user clicks the interactive marker 220, the playback system 205 may guide the user to obtain the full content of the previous chat, for example, directing the user to jump to the viewing interface of the in-conference chat content.

[0044] In some embodiments, the main timeline 210 may include, for example, an interactive marker 220 for indicating comments in the online meeting. Comments here may include, for example, any appropriate comments based on text, emoticons, images, and / or audio. For example, a user's "like" can also be understood as a comment on the corresponding content. Accordingly, the graphical identifier of the interactive marker 220 may be determined based on the content and / or type of the comment. For example, if the comment is based on an emoticon, the graphical representation of the interactive marker 220 may be generated based on the emoticon.

[0045] In some embodiments, when a user selects the interactive marker 220, the playback system 205 may, for example, guide the user to obtain descriptive information about the comment. For example, when the user hovers the mouse over the interactive marker 220, the playback system 205 may display information about the comment, such as the commenter, comment time, and replies to the comment, in a floating window. In another example, if the user clicks the interactive marker 220, the playback system 205 may guide the user to a comment viewing interface to obtain more comprehensive information about the comment.

[0046] In some embodiments, the main timeline 210 may also present speaker information indicating the temporal distribution of speech content of at least one speaker associated with the audiovisual content. In this case, the main timeline 210 may also be considered a type of speech timeline.

[0047] For example, the main timeline 210 can assign a corresponding color mark to each speaker. Accordingly, the color distribution on the main timeline 210 can be used to indicate which one or more speakers correspond to the corresponding time period. It should be understood that other appropriate formats can also be used to use the main timeline 210 to indicate the temporal distribution of the speech content of the speakers.

[0048] Speech Timeline

[0049] 2A , the playback system 205 may further include a viewing portal 230 for viewing a speech timeline. In some embodiments, the viewing portal 230 may indicate a graphic identifier (eg, avatar) of one or more speakers associated with the audiovisual content.

[0050] After receiving the user's selection of the viewing portal 230 , as shown in FIG2B , the playback system 205 may present, for example, a speech timeline 240 - 1 and a speech timeline 240 - 2 (individually or collectively referred to as speech timelines 240 ).

[0051] In some embodiments, the speech timeline 240 can be used to indicate the temporal distribution of the speech content of at least one speaker associated with the audiovisual content. For example, if the speaker spoke at the corresponding time, the speech timeline 240 may be filled with a first graph; conversely, if the speaker did not speak at the corresponding time, the speech timeline 240 may be filled with a second graph. This allows users to intuitively understand when each speaker spoke.

[0052] In some embodiments, as shown in FIG2B , speech timeline 240 can also similarly present graphical information corresponding to the audio waveform of the portion of audiovisual content corresponding to the speaker. This allows the user to intuitively understand when the speaker was silent and when they spoke frequently. This information further helps users quickly access desired content.

[0053] In some embodiments, the number of speech timelines 240 may be determined based on the number of speakers participating in the audiovisual content. In some embodiments, the number of speakers may be determined based on the number of terminals participating in the online conference. For example, multiple conference participants may access the online conference through the same terminal (or using the same account), and such multiple participants may be identified as the same speaker, even though the speaker may include multiple different speakers.

[0054] In some embodiments, the number of such speakers may be determined based on the number of speakers in the audiovisual content. It should be understood that any appropriate speaker recognition technology may be used to determine the corresponding speakers in the audiovisual content, and this disclosure is not intended to be limiting in this regard.

[0055] In some embodiments, upon receiving a selection of the viewing portal 230, the playback system 205 may present a speech timeline 240 corresponding to all speakers of the audiovisual content. Taking FIG2B as an example, the audiovisual content may include two speakers ("Speaker 1" and "Speaker 2"). Accordingly, the order in which the corresponding speech timelines 240-1 and 240-2 are presented in the playback system 205 may be determined based on the speaker information.

[0056] In some examples, the presentation order of the speech timeline can be determined based on, for example, text identifiers of the speakers, such as user names or nicknames of the speakers, and the presentation order of the speech timeline can be based on the order of the text identifiers of the speakers.

[0057] In yet other examples, the order in which speech timelines are presented may be determined based on the proportion of speech content by the speakers. For example, if the speech content proportion of "Speaker 1" reaches "70%", which is greater than the speech content proportion of "Speaker 2" (30%), speech timeline 240-1 may be presented in priority over speech timeline 240-2.

[0058] In yet other examples, the order in which speech timelines are presented can be determined based on the start time of each speaker's speech. For example, the speech of "Speaker 1" may start three minutes into the meeting. Accordingly, speech timeline 240-1 may be presented in priority over speech timeline 240-2.

[0059] It should be understood that other appropriate sorting strategies may be used to sort the multiple speech timelines 240 to facilitate users to obtain desired content more efficiently.

[0060] In some embodiments, the playback system 205 may also present descriptive information of the corresponding speaker in association with the speech timeline 240. For example, the speech timeline 240-1 may include a text identifier (e.g., a username or nickname) of the corresponding speaker. Alternatively, the speech timeline 240-1 may also include a graphical identifier (e.g., an avatar) of the corresponding speaker.

[0061] In some embodiments, the playback system 205 may also present the percentage of speech content of the corresponding speaker in association with at least one speech timeline. The speech timeline 240-1 may include the percentage of speech content of "Speaker 1" as "XX%".

[0062] In some embodiments, similar to the main timeline 210 , the speech timeline 240 may also present an interaction identifier (not shown in FIG. 2B ) for indicating an interaction behavior associated with a corresponding speaker in the online conference.

[0063] In some embodiments, such interactive behaviors refer to corresponding interactive behaviors in which the corresponding speaker participates, such as the file sharing, online chatting, or commenting behaviors discussed above. The interaction logic of the interaction identifiers presented on the speech timeline 240 can be similar to the interaction identifiers 220 discussed above, and will not be described in detail here.

[0064] In some embodiments, the speech timeline 240 - 1 may also be automatically collapsed or expanded in response to the user selecting the viewing portal 230 . For example, the playback system may always provide the speech timeline for all speakers by default, regardless of the selection of the viewing portal 230 .

[0065] In some embodiments, the playback system 205 may further provide a search portal 250 for a speech timeline. Using the search portal 250, a user may initiate a viewing request associated with a specific speaker.

[0066] In some embodiments, upon receiving a selection of the viewing portal 250, the playback system 205 may present visual elements associated with all speakers associated with the audiovisual content. Such visual elements may include, for example, text identifiers (e.g., user names or nicknames) or graphic identifiers (e.g., avatars) of the speakers.

[0067] Furthermore, the playback system 205 may receive a user's selection of a specific visual element from among the multiple visual elements to determine that the user desires to view the speech timeline of the speaker corresponding to the selected visual element. For example, the user may click on the avatar of "Speaker 1" to cause the playback system 205 to only display the speech timeline 240-1 corresponding to "Speaker 1" and not display the speech timeline 240-2.

[0068] As another example, the user may also provide input indicating a target speaker by viewing the portal 250. For example, the user may enter at least part of the nickname or username of "Speaker 1" to automatically match "Speaker 1" and cause the playback system 205 to accordingly present the speech timeline 240-1 corresponding to "Speaker 1" instead of presenting the speech timeline 240-2.

[0069] In some embodiments, the search entry 250 may be provided independently of the view entry 230. For example, when the view entry 230 is not selected, the playback system 205 may also provide a search entry 250 for viewing a specific speaker.

[0070] Alternatively, the search portal 250 may be provided dependent on the viewing portal 230. That is, only when the viewing portal 230 is activated and the speech timelines of all speakers are presented, the search portal 250 is provided accordingly for quickly filtering or locating a specific speech timeline.

[0071] In some embodiments, the speech timeline 240 may also support various types of user interactions. For example, as shown in FIG2C , the user may click on a position 260 in the speech timeline 240 - 1 to indicate that the audiovisual content is expected to be played starting from that position.

[0072] Accordingly, the playback system 205 can play the audio-visual content from the time point 270 corresponding to the position 260. In some embodiments, the playback system 205 can play the audio-visual content continuously from the time point 270. For example, if the time point 270 is "5 minutes and 30 seconds", the audio-visual content will be played continuously from "5 minutes and 30 seconds" until the end.

[0073] Alternatively, the playback system 205 may also play the portion of the audiovisual content corresponding to "Speaker 1" starting at time point 270. That is, the playback system 205 may play only the portion of the audiovisual content of "Speaker 1" corresponding to the speech timeline 240-1, and play it starting at time point 270, thereby achieving the effect of only listening to a specific speaker.

[0074] In some embodiments, if the user performs a preset operation on the speech timeline 240-1 (for example, double-clicking the speech timeline 240-1), the playback system 205 can make the part of the audio-visual content corresponding to "Speaker 1" play from the beginning, that is, only play the part of the audio-visual content corresponding to "Speaker 1" in the audio-visual content.

[0075] It should be understood that, for descriptive purposes, although various examples of playback systems have been discussed above in conjunction with Figures 2A to 2C , the various features discussed above (e.g., provision of audio waveforms, provision of interactive identifiers, provision of speech timelines, style of speech timelines, interaction of speech timelines, etc.) can be provided independently or in a combination different from that shown in Figures 2A to 2C . For example, if the playback system provides the feature of interactive identifiers, the timeline of the playback system can be a graphical style similar to the timeline of a traditional playback system, and does not necessarily have to be used to indicate an audio waveform.

[0076] Furthermore, while the examples shown in Figures 2A to 2C are for reviewing recorded content, the playback system 205 can also be used to play real-time audiovisual content (e.g., live audio and video streams). Accordingly, the speech timeline discussed above can be used to indicate, for example, the temporal distribution of historical speech content of at least one speaker associated with a historical portion of the real-time audiovisual content. For example, the speech timeline can graphically present the temporal distribution of each speaker's historical speech content from the start of the live broadcast to the current moment.

[0077] Sample viewing interface

[0078] In some embodiments, embodiments of the present disclosure may also provide an interface for viewing audiovisual content. Such an interface may be, for example, a playback interface for recorded content or a live interface for real-time content. It should be understood that for the sake of convenience, the following uses the "meeting minutes" scenario as an example for viewing audiovisual content, but such a scenario is merely exemplary, and embodiments of the present disclosure may also be applied to other appropriate scenarios.

[0079] FIG3A illustrates an example viewing interface 300 according to some embodiments of the present disclosure. As shown in FIG3A , viewing interface 300 may include playback controls 310. Playback controls 310 may, for example, be implemented using playback system 205 as discussed above. As shown in FIG3A , playback controls 310 may, for example, include a main timeline 312 and speech timelines 314-1 and 314-2 (individually or collectively referred to as speech timelines 314).

[0080] In some embodiments, the viewing interface 300 further includes a text control 320 for presenting text content corresponding to the audiovisual content. In some embodiments, the text content may be generated based on the audio of the audiovisual content. For example, if the audiovisual content is a meeting transcript, the text content may be generated based on speech recognition of the audio of each speaker in the meeting. For example, if the audiovisual content is a live broadcast, the text content may be generated based on real-time speech recognition of each speaker.

[0081] In some embodiments, as shown in FIG3B , the user may select the speech timeline 314 - 1 , for example. Accordingly, the text content 322 corresponding to “Speaker 1 ” may be adjusted to be highlighted in the text control 320 relative to other text content 324 of other speakers.

[0082] In some embodiments, emphasizing the text content 322 relative to other text content 324 may include, for example, increasing the prominence of the text content 322 displayed in the text control 320. For example, the display style (e.g., text color, background color, boldness, font size, underline) of the text content 322 may be adjusted to be more prominent. For example, the text content 322 may be bolded or highlighted.

[0083] Alternatively, emphasizing text content 322 relative to text content 324 may include, for example, reducing the prominence of other text content 324 displayed in the text control. For example, the display style (e.g., text color, background color, boldness, font size, underline) of text content 324 may be adjusted to be less prominent. For example, as shown in FIG3B , the text color of other text content 324 may be gray to contrast with the black text content 322.

[0084] 2C , the user may also select a specific position in the speech timeline to trigger the audiovisual content to be played starting at the corresponding moment. Alternatively or additionally, when a specific position in the speech timeline is selected, the text content corresponding to the specific position may also be adjusted to be highlighted in the text control 320.

[0085] For example, the text content presented in the text control 320 always corresponds to the moment at which the audiovisual content is currently playing. When a user selects a specific moment in the speech timeline, the text content corresponding to that moment (for example, the text corresponding to a certain passage spoken by the speaker) can be adjusted to the top of the text display 320 for prominent presentation. Alternatively or additionally, the display style of the text content can also be adjusted to highlight the text content. For example, one or more words corresponding to that time point can be highlighted for prominent presentation.

[0086] Sharing of audio-visual content

[0087] In some embodiments, embodiments of the present disclosure may also support sharing of audiovisual content segments based on speech timelines. As shown in FIG4A , a user may select one or more speech timelines (e.g., speech timeline 430 - 1 ) from among the multiple speech timelines in playback control 410 (or playback system 410 ) for sharing.

[0088] After receiving the selection, the audiovisual content corresponding to the speech timeline 430-1 can be generated for sharing. Taking Figure 4A as an example, after the user selects the speech timeline 430-1 and clicks the sharing entry 420 (i.e., issuing a sharing request), the entire speech content of "Speaker 1" can be used to generate independent audiovisual content, for example, to be shared with other users or organizations.

[0089] As another example, as shown in Figure 4B, the user can also select one or more time segments in the speech timelines 430-1 and 430-2, such as time segment 440-1, time segment 440-2, and time segment 440-3. Accordingly, after the user clicks on the sharing entry 420 (i.e., issues a sharing request), multiple discrete audio-visual content segments corresponding to time segment 440-1, time segment 440-2, and time segment 440-3 can be combined to generate independent segmented audio-visual content, for example, for sharing with other users or organizations.

[0090] Based on this approach, the embodiments of the present disclosure can support users to more efficiently share audiovisual content segments by selecting a speech timeline or time segment, thereby improving the efficiency of audiovisual content sharing and the efficiency of information acquisition for those being shared. In addition, the embodiments of the present disclosure also support users to select non-contiguous segments to create, which further increases the flexibility of sharing audiovisual content segments.

[0091] Example Process

[0092] FIG5 illustrates a flow chart of an example process 500 for viewing audiovisual content according to some embodiments of the present disclosure. Process 500 may be implemented on a suitable electronic device. Examples of such electronic devices may include, but are not limited to, desktop computers, laptop computers, smartphones, tablet computers, personal digital assistants, or smart wearable devices.

[0093] As shown in FIG. 5 , in block 510 , the electronic device provides a viewing interface for audio-visual content, where the viewing interface includes a play control for playing the audio-visual content.

[0094] In box 520, the electronic device presents at least one speech timeline in the playback control, where the at least one speech timeline is used to indicate the temporal distribution of speech content of at least one speaker associated with the audio-visual content.

[0095] In some embodiments, the viewing interface further includes a text control for presenting text content corresponding to the audiovisual content, where the text content is generated based on the audio of the audiovisual content.

[0096] In some embodiments, the method also includes: in response to the selection of a first speech timeline in at least one speech timeline, causing the first text content corresponding to the first speaker in the text content to be highlighted in the text control relative to the second text content of other speakers, wherein the first speech timeline corresponds to the first speaker.

[0097] In some embodiments, first text content corresponding to a first speaker in the text content is highlighted in a text control relative to second text content of other speakers: the prominence of the first text content displayed in the text control is increased; and / or the prominence of the second text content displayed in the text control is reduced.

[0098] In some embodiments, the method further includes: receiving a selection of a first position in a first speech timeline of at least one speech timeline; and causing text content corresponding to the first position in the text content to be highlighted in the text control.

[0099] In some embodiments, presenting at least one speech timeline in the playback control includes: presenting a viewing entry in the playback control for viewing the speech timeline; and presenting at least one speech timeline in the playback control in response to selection of the viewing entry.

[0100] In some embodiments, at least one speech timeline includes multiple speech timelines, and the presentation order of the multiple speech timelines in the playback control is determined based on at least one of the following: text identifiers of the multiple speakers corresponding to the multiple speech timelines, the proportion of the speech content of the multiple speakers, or the start time of the speech content of the multiple speakers.

[0101] In some embodiments, presenting at least one speech timeline in the playback control includes: receiving a viewing request associated with a target speaker; and presenting a target speech timeline corresponding to the target speaker, the target speech timeline being used to indicate the temporal distribution of the target speech content of the target speaker.

[0102] In some embodiments, receiving a viewing request associated with a target speaker includes: presenting multiple visual elements associated with multiple speakers associated with audio-visual content; and receiving a viewing request associated with the target speaker based on a preset operation of a target visual element corresponding to the target speaker among the multiple visual elements.

[0103] In some embodiments, receiving the viewing request associated with the target speaker includes receiving the viewing request associated with the target speaker based on input indicating the target speaker.

[0104] In some embodiments, the method further includes: receiving a selection of a second position in a first speech timeline of the at least one speech timeline; and causing the corresponding portion of the audiovisual content to be played from a time point corresponding to the second position.

[0105] In some embodiments, the first speech timeline corresponds to the first speaker, and causing at least part of the audio-visual content to be played from the time point corresponding to the second position includes: causing the audio-visual content to be played continuously from the time point; or causing part of the audio-visual content corresponding to the first speaker in the audio-visual content to be played from the time point.

[0106] In some embodiments, the method further includes: presenting description information of the corresponding speaker in association with the at least one speech timeline, where the description information is generated based on a text identifier and / or a graphic identifier of the speaker.

[0107] In some embodiments, the method further includes: presenting information on the proportion of speech content of the corresponding speaker in association with at least one speech timeline.

[0108] In some embodiments, the playback control further includes a main timeline for presenting graphical information corresponding to an audio waveform of the audiovisual content.

[0109] In some embodiments, the audio-visual content is an audio-visual recording of an online meeting, and the playback control further includes a main timeline, which is used to present a first interaction identifier corresponding to a first interaction behavior in the online meeting.

[0110] In some embodiments, at least one speech timeline further presents a second interaction identifier for indicating a second interaction behavior associated with the corresponding speaker in the online conference.

[0111] In some embodiments, the first interactive behavior and / or the second interactive behavior includes at least one of the following: file sharing, online chatting, and commenting.

[0112] In some embodiments, the method further includes: presenting first description information for the first interaction behavior in response to a first selection of the first interaction identifier; and / or presenting second description information for the second interaction behavior in response to a second selection of the second interaction identifier.

[0113] In some embodiments, the method further includes: receiving a selection of at least one time segment in at least one speech timeline; and generating a first segment of audio-visual content corresponding to the at least one time segment for sharing based on a first sharing request associated with the at least one time segment.

[0114] In some embodiments, the method also includes: receiving a selection of a group of speech timelines in at least one speech timeline, the group of timelines including one or more speech timelines; and based on a second sharing request associated with the group of timelines, generating a second segment of audio-visual content corresponding to the group of speech timelines for sharing.

[0115] In some embodiments, the audiovisual content includes real-time audiovisual content, and the at least one speech timeline is used to indicate temporal distribution of historical speech content of at least one speaker associated with a historical portion of the real-time audiovisual content.

[0116] Example devices and equipment

[0117] The embodiments of the present disclosure also provide corresponding devices for implementing the above methods or processes. Figure 6 shows a schematic structural block diagram of a device 600 for viewing audio-visual content according to some embodiments of the present disclosure.

[0118] As shown in FIG6 , the apparatus 600 includes a providing module 610 configured to provide a viewing interface for audio-visual content, where the viewing interface includes a play control for playing the audio-visual content.

[0119] In addition, the apparatus 600 further includes a presentation module 620 configured to present at least one speech timeline in the playback control, where the at least one speech timeline is used to indicate the temporal distribution of speech content of at least one speaker associated with the audio-visual content.

[0120] In some embodiments, the viewing interface further includes a text control for presenting text content corresponding to the audiovisual content, where the text content is generated based on the audio of the audiovisual content.

[0121] In some embodiments, the presentation module 620 is further configured to: in response to selection of a first speech timeline in at least one speech timeline, cause the first text content corresponding to the first speaker in the text content to be highlighted in the text control relative to the second text content of other speakers, wherein the first speech timeline corresponds to the first speaker.

[0122] In some embodiments, first text content corresponding to a first speaker in the text content is highlighted in a text control relative to second text content of other speakers: the prominence of the first text content displayed in the text control is increased; and / or the prominence of the second text content displayed in the text control is reduced.

[0123] In some embodiments, the presentation module 620 is further configured to: receive a selection of a first position in a first speech timeline of at least one speech timeline; and highlight the text content corresponding to the first position in the text control.

[0124] In some embodiments, the presentation module 620 is further configured to: present a viewing entry for viewing the speech timeline in the playback control; and present at least one speech timeline in the playback control in response to selection of the viewing entry.

[0125] In some embodiments, at least one speech timeline includes multiple speech timelines, and the presentation order of the multiple speech timelines in the playback control is determined based on at least one of the following: text identifiers of the multiple speakers corresponding to the multiple speech timelines, the proportion of the speech content of the multiple speakers, or the start time of the speech content of the multiple speakers.

[0126] In some embodiments, the presentation module 620 is further configured to: receive a viewing request associated with a target speaker; and present a target speech timeline corresponding to the target speaker, the target speech timeline being used to indicate the temporal distribution of the target speech content of the target speaker.

[0127] In some embodiments, the presentation module 620 is further configured to: present multiple visual elements associated with multiple speakers associated with the audio-visual content; and receive a viewing request associated with a target speaker based on a preset operation of a target visual element corresponding to the target speaker among the multiple visual elements.

[0128] In some embodiments, presentation module 620 is further configured to receive a viewing request associated with a target speaker based on the input indicating the target speaker.

[0129] In some embodiments, the presentation module 620 is further configured to: receive a selection of a second position in a first speech timeline of the at least one speech timeline; and cause the corresponding portion of the audiovisual content to be played from a time point corresponding to the second position.

[0130] In some embodiments, the first speech timeline corresponds to the first speaker, and the presentation module 620 is further configured to: enable the audiovisual content to be played continuously from a point in time; or enable the portion of the audiovisual content corresponding to the first speaker to be played from a point in time.

[0131] In some embodiments, the presentation module 620 is further configured to present description information of the corresponding speaker in association with at least one speech timeline, where the description information is generated based on a text identifier and / or a graphic identifier of the speaker.

[0132] In some embodiments, the presentation module 620 is further configured to present information on the proportion of speech content of the corresponding speaker in association with at least one speech timeline.

[0133] In some embodiments, the playback control further includes a main timeline for presenting graphical information corresponding to an audio waveform of the audiovisual content.

[0134] In some embodiments, the audio-visual content is an audio-visual recording of an online meeting, and the playback control further includes a main timeline, which is used to present a first interaction identifier corresponding to a first interaction behavior in the online meeting.

[0135] In some embodiments, at least one speech timeline further presents a second interaction identifier for indicating a second interaction behavior associated with the corresponding speaker in the online conference.

[0136] In some embodiments, the first interactive behavior and / or the second interactive behavior includes at least one of the following: file sharing, online chatting, and commenting.

[0137] In some embodiments, the presentation module 620 is further configured to: present first descriptive information for the first interactive behavior in response to a first selection of the first interaction identifier; and / or present second descriptive information for the second interactive behavior in response to a second selection of the second interaction identifier.

[0138] In some embodiments, the presentation module 620 is further configured to: receive a selection of at least one time segment in at least one speech timeline; and based on a first sharing request associated with the at least one time segment, generate a first segment of audio-visual content corresponding to the at least one time segment for sharing.

[0139] In some embodiments, the presentation module 620 is further configured to: receive a selection of a group of speech timelines in at least one speech timeline, the group of timelines including one or more speech timelines; and based on a second sharing request associated with the group of timelines, generate a second segment of audio-visual content corresponding to the group of speech timelines for sharing.

[0140] In some embodiments, the audiovisual content includes real-time audiovisual content, and the at least one speech timeline is used to indicate temporal distribution of historical speech content of at least one speaker associated with a historical portion of the real-time audiovisual content.

[0141] The units included in the device 600 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units can be implemented using software and / or firmware, such as machine executable instructions stored on a storage medium. In addition to or as an alternative to machine executable instructions, some or all of the units in the device 600 can be implemented at least in part by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that can be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0142] Figure 7 shows a block diagram of a computing device / server 700 in which one or more embodiments of the present disclosure may be implemented. It should be understood that the computing device / server 700 shown in Figure 7 is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein.

[0143] As shown in FIG7 , computing device / server 700 is in the form of a general-purpose computing device. Components of computing device / server 700 may include, but are not limited to, one or more processors or processing units 710, memory 720, storage device 730, one or more communication units 740, one or more input devices 760, and one or more output devices 760. Processing unit 710 may be a real or virtual processor and is capable of performing various processes according to a program stored in memory 720. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to increase the parallel processing capabilities of computing device / server 700.

[0144] The computing device / server 700 typically includes a plurality of computer storage media. Such media can be any accessible media that the computing device / server 700 can obtain, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 720 can be a volatile memory (e.g., register, cache, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory) or some combination thereof. The storage device 730 can be a removable or non-removable medium and can include a machine-readable medium, such as a flash drive, a disk or any other medium, which can be used to store information and / or data (e.g., training data for training) and can be accessed within the computing device / server 700.

[0145] The computing device / server 700 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 7 , a disk drive for reading from or writing to a removable, non-volatile disk (e.g., a “floppy disk”) and an optical drive for reading from or writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. The memory 720 may include a computer program product 725 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.

[0146] The communication unit 740 enables communication with other computing devices via a communication medium. Additionally, the functionality of the components of the computing device / server 700 can be implemented as a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the computing device / server 700 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or other network nodes.

[0147] Input device 750 may be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 760 may be one or more output devices, such as a display, speaker, printer, etc. Computing device / server 700 may also communicate with one or more external devices (not shown) via communication unit 740, as needed, such as storage devices, display devices, etc., with one or more devices that allow users to interact with computing device / server 700, or with any device that allows computing device / server 700 to communicate with one or more other computing devices (e.g., a network card, modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).

[0148] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which one or more computer instructions are stored, wherein the one or more computer instructions are executed by a processor to implement the method described above.

[0149] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products implemented according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0150] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0151] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0152] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple implementations of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part for a module, program segment or instruction, and a part for a module, program segment or instruction comprises one or more executable instructions for realizing the logical function of the specification. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.

[0153] While various implementations of the present disclosure have been described above, the foregoing description is intended to be illustrative, non-exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is selected to best explain the principles of the implementations, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the implementations disclosed herein.

Claims

1. A method for viewing audiovisual content, comprising: providing a viewing interface for the audiovisual content, the viewing interface including playback controls for playing the audiovisual content; and presenting at least one speech timeline in the playback controls, the at least one speech timeline being configured to indicate the temporal distribution of the speech content of at least one speaker associated with the audiovisual content.

2. The method according to claim 1, wherein the viewing interface further includes a text control for presenting text content corresponding to the audiovisual content, the text content being generated based on the audio of the audiovisual content.

3. The method according to claim 2, further comprising: in response to selection of a first speech timeline among the at least one speech timeline, causing first text content corresponding to a first speaker in the text content to be highlighted relative to second text content of other speakers in the text control, wherein the first speech timeline corresponds to the first speaker.

4. The method according to claim 3, wherein causing first text content corresponding to a first speaker in the text content to be highlighted relative to second text content of other speakers in the text control: increases the prominence of the first text content displayed in the text control; and / or decreases the prominence of the second text content displayed in the text control.

5. The method according to claim 2, further comprising: receiving a selection of a first position in a first speech timeline of the at least one speech timeline; and causing text content corresponding to the first position in the text content to be prominently presented in the text control.

6. The method according to claim 1, wherein presenting at least one speech timeline in the playback controls comprises: presenting a viewing entry for viewing the speech timeline in the playback controls; and in response to selection of the viewing entry, presenting the at least one speech timeline in the playback controls.

7. The method according to claim 1, wherein the at least one speech timeline includes a plurality of speech timelines, and the presentation order of the plurality of speech timelines in the playback controls is determined based on at least one of the following: the text identifiers of a plurality of speakers corresponding to the plurality of speech timelines, the proportion of the speech content of the plurality of speakers, or the start times of the speech content of the plurality of speakers.

8. The method according to claim 1, wherein presenting at least one speech timeline in the playback controls comprises: receiving a viewing request associated with a target speaker; and presenting a target speech timeline corresponding to the target speaker, the target speech timeline being configured to indicate the temporal distribution of the target speech content of the target speaker.

9. The method according to claim 8, wherein receiving a viewing request associated with a target speaker comprises: presenting a plurality of visual elements associated with a plurality of speakers associated with the audiovisual content; and Receiving a viewing request associated with the target speaker based on a preset operation on the target visual element corresponding to the target speaker among the multiple visual elements.

10. The method according to claim 8, wherein receiving a viewing request associated with the target speaker comprises: Receiving a viewing request associated with the target speaker based on an input indicating the target speaker.

11. The method according to claim 1, further comprises: Receiving a selection of a second position in a first speech timeline among the at least one speech timeline; and Causing a corresponding part of the audiovisual content to be played from the time point corresponding to the second position.

12. The method according to claim 11, wherein the first speech timeline corresponds to a first speaker, and causing at least part of the audiovisual content to be played from the time point corresponding to the second position comprises: Causing the audiovisual content to be continuously played from the time point; or Causing a part of the audiovisual content corresponding to the first speaker in the audiovisual content to be played from the time point.

13. The method according to claim 1, further comprises: Presenting description information of the corresponding speaker associated with the at least one speech timeline, the description information being generated based on the text identifier and / or graphic identifier of the speaker.

14. The method according to claim 1, further comprises: Presenting proportion information of the speech content of the corresponding speaker associated with the at least one speech timeline.

15. The method according to claim 1, wherein the playback control further comprises a main timeline for presenting graphic information corresponding to the audio waveform of the audiovisual content.

16. The method according to claim 1, wherein the audiovisual content is an audiovisual record of an online meeting, and the playback control further comprises a main timeline for presenting a first interaction identifier corresponding to a first interaction behavior in the online meeting.

17. The method according to claim 16, wherein the at least one speech timeline further presents a second interaction identifier for indicating a second interaction behavior associated with the corresponding speaker in the online meeting.

18. The method according to claim 16 or 17, wherein the first interaction behavior and / or the second interaction behavior comprises at least one of the following: file sharing, online chat, and comment.

19. The method according to claim 16 or 17, further comprises: Presenting first description information for the first interaction behavior in response to a first selection of the first interaction identifier; and / or Presenting second description information for the second interaction behavior in response to a second selection of the second interaction identifier.

20. The method according to claim 1, further comprises: Receiving a selection of at least one time segment among the at least one speech timeline; and Generating first segment audiovisual content corresponding to the at least one time segment for sharing based on a first sharing request associated with the at least one time segment.

21. The method according to claim 1, further comprises: Receiving a selection for a group of speech timelines in the at least one speech timeline, the group of timelines including one or more speech timelines; And Based on a second sharing request associated with the group of timelines, causing second segment audiovisual content corresponding to the group of speech timelines to be generated for sharing.

22. The method according to claim 1, wherein the audiovisual content includes real-time audiovisual content, and the at least one speech timeline is used to indicate the temporal distribution of the historical speech content of at least one speaker associated with a historical part of the real-time audiovisual content.

23. A device for viewing audiovisual content, Comprising: A providing module configured to provide a viewing interface for the audiovisual content, the viewing interface including playback controls for playing the audiovisual content; And A presenting module configured to present at least one speech timeline in the playback controls, the at least one speech timeline being used to indicate the temporal distribution of the speech content of at least one speaker associated with the audiovisual content.

24. A playback system, Comprising: A main timeline that at least indicates the current playback position of the audiovisual content; And At least one speech timeline that is used to indicate the temporal distribution of the speech content of at least one speaker associated with the audiovisual content.

25. The playback system according to claim 24, wherein the main timeline further presents graphical information corresponding to the audio waveform of the audiovisual content.

26. The playback system according to claim 24, wherein the audiovisual content is an audiovisual record of an online meeting, and the main timeline further presents a first interaction identifier corresponding to a first interaction behavior in the online meeting.

27. The playback system according to claim 26, wherein the at least one speech timeline further presents a second interaction identifier for indicating a second interaction behavior associated with the corresponding speaker in the online meeting.

28. The playback system according to claim 26 or 27, wherein the first interaction behavior and / or the second interaction behavior includes at least one of the following: file sharing, online chat, and comment.

29. The playback system according to claim 26 or 27, Wherein: A first selection of the first interaction identifier is used to trigger first description information for the first interaction behavior; And / or A second selection of the second interaction identifier is used to trigger the presentation of second description information for the second interaction behavior.

30. The playback system according to claim 24, wherein the at least one speech timeline includes a plurality of speech timelines, and the presentation order of the plurality of speech timelines in the playback system is determined based on at least one of the following: The text identifiers of the plurality of speakers corresponding to the plurality of speech timelines, The proportion of the speech content of the plurality of speakers, or The start time of the speech content of the plurality of speakers.

31. An electronic device, Comprising: At least one processing unit; And At least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, cause the device to perform the method according to any one of claims 1 to 22.

32. A computer-readable storage medium having stored thereon a computer program, which when executed by a processor implements the method according to any one of claims 1 to 22.