Methods, apparatus, devices, and storage media for checking audiovisual content
The playback system with a main and speech time axis improves navigation in audiovisual content by graphically representing speaker interactions, enabling efficient access and sharing of specific segments.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- BEIJING ZITIAO NETWORK TECH CO LTD
- Filing Date
- 2023-08-16
- Publication Date
- 2026-04-22
AI Technical Summary
Existing audiovisual content playback systems are inefficient for quickly locating specific speaker segments, particularly in lengthy recordings like meetings or online classes, due to the difficulty in navigating long timelines.
A playback system with a main time axis and speech time axis is introduced, providing graphical representations of speaker interactions and speech timelines to facilitate quick access to specific speaker segments, along with features for highlighting relevant text content and sharing audiovisual segments.
Enhances user efficiency in accessing desired speaker content by offering intuitive navigation and highlighting, allowing for precise selection and sharing of audiovisual segments.
Smart Images

Figure 2026513006000001_ABST
Abstract
Description
Technical Field
[0001] This application claims the priority of a Chinese patent application filed on October 31, 2022, with the title "Method, Apparatus, Device, and Storage Medium for Checking Audio-Visual Content", and the application number 202211352393.3, the entire content of which is incorporated herein by reference.
[0002] Exemplary embodiments of the present disclosure generally relate to the field of computers, and in particular, to a method, apparatus, device, and computer-readable storage medium for checking audio-visual content.
Background Art
[0003] With the development of computer technology, the Internet has become a major platform for people to access and share content. For example, people can use the Internet to post various contents or receive contents shared by other users.
[0004] In Internet-based content sharing, the sharing of audio-visual content (such as audio content or video content) has become one of the most important forms. People can, for example, use a player to play the speech shared by other users or the video or audio recording of a certain meeting. However, during such playback, it is difficult to quickly find the corresponding part of a specific speaker in such a video or audio recording.
Summary of the Invention
[0005] A first aspect of this disclosure provides a method for checking audiovisual content. The method includes receiving a selection for a plurality of text segments, creating an audiovisual content segment based on at least a plurality of parts of a target audiovisual content, and presenting a sharing entry for sharing the audiovisual content segment, wherein the plurality of text segments correspond to a plurality of parts in the target audiovisual content, and the plurality of parts include at least a first and second part that are not contiguous in the target audiovisual content, and the first and second parts are contiguous in the audiovisual content segment.
[0006] A second aspect of the present disclosure provides an apparatus for checking audiovisual content. The apparatus comprises a receiving module configured to receive selections for a plurality of text segments; a control module configured to create an audiovisual content segment based on at least a plurality of parts of a target audiovisual content; and a presentation module configured to present a shared entry for sharing an audiovisual content segment, wherein the plurality of text segments correspond to a plurality of parts in the target audiovisual content, and the plurality of parts include at least a first and second part that are not contiguous in the target audiovisual content, and the first and second parts are contiguous in the audiovisual content segment.
[0007] A third aspect of this disclosure provides an electronic device comprising at least one processing unit and at least one memory coupled to the at least one processing unit and storing instructions to be executed by the at least one processing unit. When the instructions are executed by the at least one processing unit, the device causes the device to perform the method of the first aspect.
[0008] A fourth aspect of this disclosure provides a computer-readable storage medium. The medium stores a computer program that, when executed by a processor, performs the method of the first aspect.
[0009] A fifth aspect of this disclosure provides a playback system, which includes a main time axis indicating at least the current playback position of audiovisual content, and at least one speech time axis used to indicate the temporal distribution of speech content of at least one speaker related to the audiovisual content.
[0010] It should be understood that the contents described in the section on means for solving this problem are not intended to limit the main or important features of the embodiments of this disclosure, nor do they limit the scope of this disclosure. Other features of this disclosure will be readily apparent from the following description. [Brief explanation of the drawing]
[0011] The above and other features, advantages, and aspects of each embodiment disclosed herein will become clearer upon further reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals indicate the same or similar elements. [Figure 1] A schematic diagram of a conventional audiovisual content player is shown. [Figure 2A] Figure 2A shows a schematic diagram of an exemplary regeneration system according to some embodiments of the present disclosure. [Figure 2B] Figure 2B shows a schematic diagram of an exemplary regeneration system according to some embodiments of the present disclosure. [Figure 2C] Figure 2C shows a schematic diagram of an exemplary regeneration system according to some embodiments of the present disclosure. [Figure 3A] Figure 3A shows an exemplary checking interface for audiovisual content according to some embodiments of the present disclosure. [Figure 3B]Figure 3B shows an exemplary checking interface for audiovisual content according to some embodiments of the present disclosure. [Figure 4A] Figure 4A shows a schematic diagram illustrating the sharing of audiovisual content segments according to some embodiments of the present disclosure. [Figure 4B] Figure 4B shows a schematic diagram illustrating the sharing of audiovisual content segments according to some embodiments of the present disclosure. [Figure 5] A flowchart illustrating an exemplary process for checking audiovisual content according to some embodiments of this disclosure is shown. [Figure 6] A block diagram of an apparatus for checking audiovisual content according to some embodiments of this disclosure is shown. [Figure 7] A block diagram of an apparatus capable of carrying out several embodiments of this disclosure is shown. [Modes for carrying out the invention]
[0012] The embodiments of this disclosure will be described in more detail below with reference to the accompanying drawings. While the accompanying drawings show several embodiments of this disclosure, it should be understood that this disclosure can be realized in various forms and should not be construed as being limited to the embodiments described herein. Rather, these embodiments are provided to allow for a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0013] In the description of the embodiments of this disclosure, the term “including” and similar terms should be understood as non-restrictive inclusion, i.e., “including, but not limited to.” The term “based on” should be understood as “based at least in part.” The term “one embodiment” or “this embodiment” should be understood as “at least one embodiment.” The term “several embodiments” should be understood as “at least several embodiments.” The following may include other explicit and implicit definitions.
[0014] As mentioned above, people can obtain audiovisual content using a player. Figure 1 shows a schematic diagram of a conventional audiovisual content player 100. As shown in Figure 1, with player 100, people typically need to drag and control the time axis to position it at the desired playback time.
[0015] However, such playback control is inefficient. For example, in the example in Figure 1, the audiovisual content is more than one hour long, making it difficult for the user to quickly position themselves at the desired playback point along the timeline.
[0016] This situation is particularly pronounced when playing back audiovisual content such as meetings, speeches, or online classes. In such scenarios, there are usually multiple speakers, and it is expected that the portion of a particular speaker's dialogue will be quickly identified.
[0017] Embodiments of the present disclosure provide a playback system for audiovisual content (audio content or video content). The system may include at least a main time axis for indicating the current playback position of the audiovisual content. Furthermore, the system may further include at least one speech time axis for indicating the temporal distribution of speech content of at least one speaker related to the audiovisual content.
[0018] In addition, embodiments of the present disclosure provide an approach for checking audio-visual content. According to this approach, for audio-visual content, a check interface including a playback controller for playing the audio-visual content may be provided. Further, at least one speech timeline may be presented to the playback controller, and the at least one speech timeline is used to show the time distribution of the speech content of at least one speaker related to the audio-visual content.
[0019] In this way, embodiments of the present disclosure can provide a speech timeline to a playback system or a playback controller so as to provide a time distribution corresponding to the speech content of a speaker related to the audio-visual content. Thereby, the implementation of the present disclosure can improve the efficiency for a user to access desired content because it facilitates the user to check a portion corresponding to a specific speaker.
[0020] Hereinafter, exemplary approaches according to embodiments of the present disclosure will be described in detail in conjunction with the accompanying drawings.
[0021] Exemplary Regeneration System In some embodiments, embodiments of the present disclosure can use a timeline to provide richer information about audio-visual content.
[0022] FIG. 2A shows a schematic diagram 200A of an exemplary playback system 205 according to some embodiments of the present disclosure. As shown in FIG. 2A, the playback system 205 (also referred to as player 2,05 or playback controller 205) can be used to play corresponding audio-visual content. The playback system 205 may be provided by, for example, a suitable electronic device, and examples of such electronic devices include, but are not limited to, desktop computers, notebook computers, smartphones, tablet computers, personal digital assistants, or smart wearable devices.
[0023] In some embodiments, the audiovisual content may include local audio / video files in the audiovisual system 205, audio / video files stored in the cloud, or audio / video streams. Such audio / video streams may include, for example, a playback stream of recorded audiovisual content (e.g., meeting minutes) or a live streaming stream of live-streamed audiovisual content.
[0024] Main timeline As shown in Figure 2A, the playback system 205 may include a main time axis 210. In some embodiments, the main time axis 210 may indicate the current playback position of the audiovisual content, i.e., the playback progress. For example, the main time axis 210 may include a playback position indicator 215 that indicates the point in time when the audiovisual content is currently being played.
[0025] For example, if the audiovisual content being played is recorded content, the length information of this audiovisual content is constant, and the total length of this time axis may correspond to the total duration of that time content. Furthermore, the playback position indicator 215 may be set correspondingly based on the correspondence between position and playback time.
[0026] In some other examples, if the audiovisual content being played is live-streamed content whose duration is still increasing, the playback position indicator 215 may, for example, always be set to the far right of the main timeline 210. Also, if the user wants to play back a particular piece of live-streamed content, they can jump back to the corresponding point in time by, for example, moving the playback position indicator 215.
[0027] In some embodiments, as shown in Figure 1, the main time axis 210 may further present graphical information corresponding to the audio waveform of the audiovisual content. In this way, the user can more easily understand which parts of the audiovisual content are noteworthy and which parts, for example, the audio waveform, can be less frequently and temporarily ignored. As a result, such a playback system 205 can improve the efficiency of user access to the content.
[0028] In some embodiments, the audiovisual content may be, for example, recorded content relating to an online meeting. Correspondingly, as shown in Figure 1, the main timeline 210 may further present interaction indicators 220 corresponding to interaction actions in an online meeting, for example.
[0029] Such interaction indicators 220 may be set at corresponding positions on the main timeline 210 to indicate that the corresponding interaction action occurred at the corresponding time. In some embodiments, different graphics of the interaction indicators 220 may correspond to different interaction actions.
[0030] In some embodiments, the main timeline 210 may include, for example, an interaction indicator 220 for indicating file sharing in an online meeting. Correspondingly, this interaction indicator 220 may have, for example, a graphic corresponding to the format of the file being shared, such as a thumbnail of the file being shared.
[0031] In some embodiments, when a user selects the interaction indicator 220, the playback system 205 may guide the user to, for example, obtain descriptive information about the shared file. For example, when a user hovers the cursor over the interaction indicator 220, for example via a mouse, the playback system 205 may display information about the shared file, such as the file name, format, size, and who shared it, in a floating window. In yet another embodiment, when a user clicks the interaction indicator 220, the playback system 205 may guide the user to access the content of the shared file, for example, by guiding the user to jump to an online viewing interface for the file.
[0032] In some embodiments, the main timeline 210 may include, for example, an interaction indicator 220 for indicating online chat in an online meeting. Online chat here refers to any appropriate chat based on text, emojis, images, and / or audio, conducted using, for example, an instant communication tool for an online meeting. Correspondingly, the graphical indicator of this interaction indicator 220 may be determined, for example, based on the content of the online chat. Alternatively, the graphical indicator of this interaction indicator 220 may be determined, for example, by the graphical indicator (e.g., avatar) of a user who is able to participate in the chat.
[0033] In some embodiments, when a user selects the interaction indicator 220, the playback system 205 may guide the user to, for example, obtain descriptive information about the online chat. For example, when a user hovers the cursor over the interaction indicator 220 with the mouse, the playback system 205 may, for example, display information about the online chat, such as the participants of the online chat and the chat content, in a floating window. In yet another embodiment, when a user clicks the interaction indicator 220, the playback system 205 may, for example, guide the user to access the full content of the previous chat, or, for example, guide the user to jump to an interface for viewing the chat content in the meeting.
[0034] In some embodiments, the main timeline 210 may include, for example, an interaction indicator 220 for indicating comments in an online meeting. The comments here may include any appropriate comments based on, for example, text, emojis, images, and / or audio. For example, a user's "like" may also be understood as one of the comments on the corresponding content. Correspondingly, the graphical representation of this interaction indicator 220 may be determined, for example, based on the content and / or type of the comment. For example, in the case of an emoji-based comment, the graphical representation of this interaction indicator 220 may be generated based on the emoji.
[0035] In some embodiments, when a user selects the interaction indicator 220, the playback system 205 may guide the user to retrieve descriptive information about the comment. For example, when a user hovers the cursor over the interaction indicator 220 with the mouse, the playback system 205 may display information about the comment, such as the commenter, the time of the comment, and replies to the comment, in a floating window. In yet another embodiment, when a user clicks the interaction indicator 220, the playback system 205 may guide the user to jump to an interface to view the comment in order to obtain richer information about the comment.
[0036] In some embodiments, the main timeline 210 may further present speaker information to show the temporal distribution of speech content from at least one speaker related to the audiovisual content. In such cases, the main timeline 210 may be recognized as a type of speech timeline.
[0037] For example, the main timeline 210 may be a corresponding color mark assigned to each speaker. Correspondingly, the distribution of colors on the main timeline 210 may be used to indicate which speaker or which speakers correspond to which time period. It should be understood that other appropriate forms may be used to show the temporal distribution of speaker utterances using the main timeline 210.
[0038] Timeline of statements In some embodiments, as shown in Figure 2A, the playback system 205 may further include, for example, a check entry 230 for viewing the speech timeline. In some embodiments, the check entry 230 may display, for example, a graphical representation (e.g., an avatar) of one or more speakers associated with the audiovisual content.
[0039] Upon receiving a user's selection to the check entry point 230, the playback system 205 may, as shown in Figure 2B, present, for example, speech timeline 240-1 and speech timeline 240-2 (referred to individually or collectively as speech timeline 240).
[0040] In some embodiments, the speech timeline 240 may be used to show the temporal distribution of speech content from at least one speaker related to audiovisual content. For example, if this speaker spoke at a corresponding time, the speech timeline 240 may be filled with a first graphic, and conversely, if this speaker did not speak at a corresponding time, the speech timeline 240 may be filled with a second graphic. Thus, the user can intuitively understand when each speaker spoke.
[0041] In some embodiments, as shown in Figure 2B, the speech timeline 240 may also present graphical information corresponding to the audio waveforms of some audiovisual content associated with the speaker. In this way, the user can intuitively understand when the speaker did not speak and when the speaker spoke frequently. Such information helps the user quickly access the desired content.
[0042] In some embodiments, the number of speaking time points 240 may be determined based on the number of speakers participating in the audiovisual content. In some embodiments, the number of such speakers may be determined by the number of devices participating in the online meeting. For example, multiple meeting participants may access the online meeting via the same device (or using the same account), in which case these multiple participants may be identified as the same speaker, even though they may include multiple different speakers.
[0043] In some embodiments, the number of such speakers may be determined based on the number of speakers in the audiovisual content. Any suitable speaker identification technique may be employed to determine the corresponding speakers in the audiovisual content, and it should be understood that this disclosure is not intended to limit this.
[0044] In some embodiments, upon receiving a selection for the check entry 230, the playback system 205 may present speech timelines 240 corresponding to all speakers in the audiovisual content. As an example in Figure 2B, the audiovisual content may include, for example, two speakers ("Speaker 1" and "Speaker 2"). Correspondingly, the presentation order of the corresponding speech timelines 240-1 and 240-2 in the playback system 205 may be determined, for example, based on speaker information.
[0045] In some examples, the presentation order of the utterance timeline may be determined based on, for example, the speaker's text marker. Such text markers may include, for example, the speaker's username or nickname, and the presentation order of the utterance timeline may be based on, for example, the rearrangement of the speaker's text marker.
[0046] In some other examples, the order in which the speech timelines are presented may be determined, for example, based on the proportion of each speaker's speech content. For example, if the proportion of "Speaker 1's" speech content reaches "70%" and is greater than the proportion of "Speaker 2's" speech content, which is "30%", then speech timeline 240-1 may be presented with priority over speech timeline 240-2, for example.
[0047] In some other examples, the order in which the speech timelines are presented may be determined based on, for example, the start time of each speaker's speech content. For example, the start time of "Speaker 1's" speech content is, for example, 1 minute after the start of the meeting, which is earlier than the start time of "Speaker 2's" speech content, for example, 3 minutes after the start of the meeting. Therefore, speech timeline 240-1 may be presented with priority over speech timeline 240-2, for example.
[0048] Please understand that you may rearrange multiple 240-minute utterance timelines by employing other appropriate sorting methods to facilitate users' efficient access to desired content.
[0049] In some embodiments, the playback system 205 may further present descriptive information of the corresponding speaker in relation to the speech timeline 240. For example, the speech timeline 240-1 may have a text marker for the corresponding speaker (e.g., username or nickname). Alternatively, the speech timeline 240-1 may further have a graphical marker for the corresponding speaker (e.g., avatar).
[0050] In some embodiments, the playback system 205 may further present percentage information of the spoken content of a corresponding speaker in relation to at least one spoken time axis. Spoken time axis 240-1 may include the percentage "XX%" of spoken content of "speaker 1".
[0051] In some embodiments, the speech timeline 240, like the main timeline 210, may further present interaction markers (not shown in Figure 2B) to indicate interaction behaviors related to the corresponding speaker in an online meeting.
[0052] In some embodiments, such interaction actions are corresponding interaction actions involving the corresponding speaker, such as the file sharing, online chat, or comment described above. The interaction logic of the interaction indicator presented on the speech time axis 240 may be the same as that of the interaction indicator 220 described above and is not described in detail herein.
[0053] In some embodiments, the speech timeline 240-1 may be automatically collapsed and expanded in response to the user's selection for the check entry 230. For example, the playback system may, by default, always provide speech timelines for all speakers, regardless of the selection for the check entry 230.
[0054] In some embodiments, the playback system 205 may further provide a search entry point 250 relating to the speech time axis. Using the search entry point 250, the user may issue a search request related to a specific speaker.
[0055] In some embodiments, upon receiving a selection for the check entry 250, the playback system 205 may present visual elements related to all speakers associated with the audiovisual content. Such visual elements may include, for example, a text marker for the speaker (e.g., username or nickname) or a graphical marker (e.g., avatar).
[0056] Furthermore, the playback system 205 may receive that the user has selected a specific visual element from among multiple visual elements and determine that the user wishes to see the speaker's speech timeline corresponding to the selected visual element. For example, the user may click on the avatar of "Speaker 1" and cause the playback system 205 to present only speech timeline 240-1 corresponding to "Speaker 1" and not present speech timeline 240-2.
[0057] As another example, the user may provide input indicating a target speaker, for example, through a check entry 250. For example, the user may enter at least part of the nickname or username of "Speaker 1" to automatically match with "Speaker 1" and cause the playback system 205 to present the speech timeline 240-1 corresponding to "Speaker 1," but not the speech timeline 240-2.
[0058] In some embodiments, the search entry point 250 may be provided independently of, for example, the check entry point 230. For example, if the check entry point 230 is not selected, the playback system 205 may appropriately provide a search entry point 250 for looking at a specific speaker.
[0059] Alternatively, the search entry 250 may be provided, for example, depending on the check entry 230. That is, the search entry 250 is provided correspondingly to quickly filter or find a specific utterance timeline only when the check entry 230 is triggered and the utterance timelines of all speakers are presented.
[0060] In some embodiments, the speech timeline 240 may support various types of user interaction. For example, as shown in Figure 2C, the user may indicate that clicking on position 260 in the speech timeline 240-1 will cause playback of this audiovisual content to begin from that position.
[0061] Correspondingly, the playback system 205 can play audiovisual content from time 270, which corresponds to position 260. In some embodiments, the playback system 205 may play the audiovisual content continuously from time 270. For example, if time 270 is "5 minutes 30 seconds," the audiovisual content will be played continuously from "5 minutes 30 seconds" until the end.
[0062] Alternatively, the playback system 205 may play the portion of the audiovisual content corresponding to "speaker 1" starting from time 270. That is, by playing only the portion of the audiovisual content of "speaker 1" corresponding to the speaking time axis 240-1, and starting from time 270, the playback system 205 can achieve the effect of listening to only a specific speaker.
[0063] In some embodiments, when a user performs a predetermined operation on the speech timeline 240-1 (for example, double-clicking on this speech timeline 240-1), the playback system 205 may play the portion of the audiovisual content corresponding to "speaker 1" from the beginning, that is, it may play only the portion of the audiovisual content corresponding to "speaker 1".
[0064] For explanatory purposes, various examples of playback systems have been described above in conjunction with Figures 2A to 2C. However, it should be understood that the various features mentioned above (e.g., provision of audio waveforms, provision of interaction indicators, provision of speech time axes, manner of speech time axes, interaction on speech time axes, etc.) may be provided independently or in combinations other than those shown in Figures 2A to 2C. For example, if a playback system provides the characteristic of interaction indicators, the time axis of the playback system may be in the same graphical manner as the time axis of conventional playback systems and does not necessarily have to be used to show audio waveforms.
[0065] Furthermore, while the examples shown in Figures 2A to 2C relate to the playback of recorded content, the playback system 205 may also be used to play real-time audiovisual content (e.g., an audio / video live stream). Correspondingly, the speech timeline described above can be used, for example, to show the temporal distribution of historical speech content of at least one speaker related to the historical portion of real-time audiovisual content. For example, the speech timeline can graphically present the temporal distribution of historical speech content for each speaker from the start time of the live stream to the current time.
[0066] Exemplary check interface In some embodiments, embodiments of the present disclosure may further provide an audiovisual content checking interface. Such a checking interface may be, for example, a playback interface for recorded content or a live streaming interface for real-time content. Below, for the purpose of convenience, a “meeting minutes” scenario is used as an example of checking audiovisual content, but it should be understood that such a scenario is merely illustrative, and embodiments of the present disclosure may be applied to other suitable scenarios.
[0067] Figure 3A shows an exemplary check interface 300 according to some embodiments of the present disclosure. As shown in Figure 3A, the check interface 300 may include a playback controller 310. The playback controller 310 may be implemented, for example, using the playback system 205 described above. As shown in Figure 3A, the playback controller 310 may include, for example, a main time axis 312 and speech time axes 314-1 and 314-2 (referred to individually or collectively as speech time axis 314).
[0068] In some embodiments, the check interface 300 further includes a text controller 320 used to present text content corresponding to the audiovisual content. In some embodiments, this text content may be generated based on the audio of the audiovisual content. For example, if the audiovisual content is meeting minutes, an example of text content may be generated based on speech recognition of the audio spoken by each speaker in the meeting. For example, if the audiovisual content is real-time live streaming content, an example of this text content may be generated, for example, based on identifying the real-time voice of each speaker.
[0069] In some embodiments, as shown in Figure 3B, the user may, for example, select a speech time axis 314-1, and accordingly, the text content 322 corresponding to "speaker 1" may be adjusted in the text controller 320 to be highlighted compared to other text content 32 of other speakers.
[0070] In some embodiments, highlighting text content 322 compared to other text content 324 may include, for example, increasing the visibility of text content 322 displayed in the text controller 320. For example, the display characteristics of text content 322 (e.g., text color, background color, degree of boldness, font size, underlining, etc.) may be adjusted to make it more prominent. For example, text content 322 may be made bold or highlighted.
[0071] Alternatively, highlighting text content 322 compared to text content 324 may be achieved, for example, by reducing the prominence of other text content 324 displayed in the text controller. For example, the display characteristics of text content 324 (e.g., text color, background color, degree of boldness, font size, underline) may be adjusted to make it less conspicuous. For example, as shown in Figure 3B, the text color of other text content 324 may be changed to gray to create a contrast with the black text content 322.
[0072] As illustrated with reference to Figure 2C, the user may trigger playback of audiovisual content from a corresponding time by selecting a specific position within the speech timeline. Alternatively or additionally, when a specific position within the speech timeline is selected, the text content corresponding to this specific position may also be adjusted to be highlighted in the text controller 320.
[0073] For example, the text content presented to the text controller 320 always corresponds to the time the audiovisual content is currently playing. When a user selects a specific time within the speech timeline, a segment of text content corresponding to this time (e.g., text corresponding to a specific statement made by the speaker) may be positioned at the top of the text disclosure 320 to make it more prominent. Alternatively or additionally, the display manner of this segment of text content may also be adjusted to make it more prominent. For example, one or more words corresponding to that time may be highlighted to make them more prominent.
[0074] Sharing audiovisual content segments In some embodiments, embodiments of the present disclosure may also support, for example, the sharing of audiovisual content segments based on a speech time axis. As shown in Figure 4A, the user may select one or more speech time axes (e.g., speech time axis 430-1) from among multiple speech time axes in the playback controller 410 (or playback system 410) for sharing.
[0075] Upon receiving this selection, the audiovisual content segment corresponding to this speech time axis 430-1 can be generated for sharing. Using Figure 4A as an example, after a user selects speech time axis 430-1 and clicks the share entry 420 (i.e., sends a share request), the entire speech content of "Speaker 1" is used for sharing, for example, with other users or organizations by generating an independent segment of audiovisual content.
[0076] As another example, as shown in Figure 4B, the user may select, for example, one or more time segments within the speech time axes 430-1 and 430-2, for example, time segment 440-1, time segment 440-2, and time segment 440-3. Correspondingly, after the user clicks the sharing entry point 420 (i.e., after issuing a sharing request), the multiple individual audiovisual content segments associated with time segments 440-1, time segment 440-2, and time segment 440-3 are combined, for example, to generate independent audiovisual content segments, which are then used for sharing with other users or organizations.
[0077] In this way, embodiments of the present disclosure can support the efficient sharing of audiovisual content segments by allowing users to select a speech timeline or time segment, thereby improving the efficiency of sharing audiovisual content and the efficiency of access to the information by those who will receive it. Furthermore, embodiments of the present disclosure also support users selecting and creating discontinuous segments, thereby further improving the flexibility of sharing audiovisual content segments.
[0078] Exemplary process Figure 5 shows a flowchart of an exemplary process 500 for checking audiovisual content according to some embodiments of the present disclosure. Process 500 may be carried out using appropriate electronic devices. Examples of such electronic devices include, but are not limited to, desktop computers, notebook computers, smartphones, tablet computers, personal digital assistants, or smart wearable devices.
[0079] As shown in Figure 5, in box 510, the electronic equipment provides a check interface for audiovisual content, and the check interface includes a playback controller for playing the audiovisual content. In box 520, the electronic device presents at least one speech time axis in the playback controller, the at least one speech time axis is used to show the temporal distribution of speech content of at least one speaker related to audiovisual content.
[0080] In some embodiments, the check interface further includes a text controller, which is used to present text content corresponding to audiovisual content, and the text content is generated based on the audio of the audiovisual content.
[0081] In some embodiments, the method includes, in response to a selection for a first speech timeline within at least one speech timeline, causing a text controller to highlight first text content corresponding to a first speaker in the text content relative to second text content of other speakers, where the first speech timeline corresponds to the first speaker. In some embodiments, highlighting the first text content corresponding to the first speaker in the text content relative to the second text content of other speakers in the text controller increases the prominence of the first text content displayed in the text controller and / or decreases the prominence of the second text content displayed in the text controller.
[0082] In some embodiments, the method further includes receiving a selection for a first position in a first utterance time axis of at least one utterance time axis, and highlighting the text content corresponding to the first position in the text content in a text controller.
[0083] In some embodiments, presenting at least one utterance timeline in the playback controller includes presenting a check entry to the playback controller for checking the utterance timeline, and presenting at least one utterance timeline in the playback controller in response to a selection of the check entry.
[0084] In some embodiments, at least one speech timeline includes multiple speech timelines, and the presentation order of the multiple speech timelines in the playback controller is determined based on at least one of the following: text markers for multiple speakers corresponding to the multiple speech timelines, the proportion of speech content from multiple speakers, or the start times of speech content from multiple speakers.
[0085] In some embodiments, presenting at least one utterance timeline in the playback controller includes receiving a check request related to a target speaker and presenting a target utterance timeline corresponding to the target speaker, the target utterance timeline being used to show the temporal distribution of the target utterance content of the target speaker.
[0086] In some embodiments, receiving a check request related to a target speaker includes presenting multiple visual elements related to multiple speakers related to audiovisual content, and receiving a check request related to a target speaker based on a predetermined operation on a target visual element corresponding to a target speaker within the multiple visual elements.
[0087] In some embodiments, receiving a check request related to a target speaker includes receiving a check request related to a target speaker based on an input indicating a target speaker.
[0088] In some embodiments, the method further includes receiving a selection for a second position in a first speech timeline of at least one speech timeline, and playing a corresponding portion of audiovisual content from the point in time corresponding to the second position.
[0089] In some embodiments, the first speech timeline corresponds to the first speaker, and playing at least a portion of the audiovisual content from a point in time corresponding to the second position includes playing the audiovisual content continuously from that point in time, or playing the portion of the audiovisual content corresponding to the first speaker from that point in time.
[0090] In some embodiments, the method further includes presenting corresponding speaker descriptive information in relation to at least one utterance time axis, the descriptive information being generated based on speaker text markers and / or graphical markers.
[0091] In some embodiments, the method further includes presenting percentage information of the utterance content of a corresponding speaker in relation to at least one utterance time axis.
[0092] In some embodiments, the playback controller further includes a main time axis, which is used to present graphical information corresponding to the audio waveform of the audiovisual content.
[0093] In some embodiments, the audiovisual content is an audiovisual recording of an online meeting, and the playback controller further includes a main timeline, which is used to present a first interaction indicator corresponding to a first interaction action in the online meeting.
[0094] In some embodiments, at least one speech timeline further presents a second interaction marker to indicate a second interaction behavior related to the corresponding speaker in the online meeting.
[0095] In some embodiments, the first interaction action and / or the second interaction action includes at least one of file sharing, online chat, and commenting.
[0096] In some embodiments, the method further includes presenting first descriptive information relating to a first interaction action in response to a first selection for a first interaction indicator, and / or presenting second descriptive information relating to a second interaction action in response to a second selection for a second interaction indicator.
[0097] In some embodiments, the method further includes receiving a selection for at least one time segment within at least one speech time axis, and generating a first audiovisual content segment corresponding to at least one time segment for sharing based on a first sharing request related to at least one time segment.
[0098] In some embodiments, the method further includes receiving a selection for a set of speech timelines, which comprises one or more speech timelines, within at least one speech timeline; and generating a second audiovisual content segment corresponding to the speech timeline set for sharing, based on a second sharing request associated with the timeline set.
[0099] In some embodiments, the audiovisual content includes real-time audiovisual content, and at least one utterance time axis is used to show the temporal distribution of historical utterance content of at least one speaker related to the historical portion of the real-time audiovisual content.
[0100] Exemplary devices and equipment The embodiments of this disclosure also provide corresponding apparatus for carrying out the methods or processes described above. Figure 6 shows a schematic block diagram of an apparatus 600 for checking audiovisual content according to some embodiments of this disclosure.
[0101] As shown in Figure 6, the device 600 includes a providing module 610 configured to provide a check interface for audiovisual content, the check interface including a playback controller for playing the audiovisual content.
[0102] Furthermore, the device 600 includes a presentation module 620 configured to present at least one speech time axis in the playback controller, the at least one speech time axis being used to show the temporal distribution of speech content of at least one speaker related to audiovisual content.
[0103] In some embodiments, the check interface further includes a text controller, which is used to present text content corresponding to audiovisual content, and the text content is generated based on the audio of the audiovisual content.
[0104] In some embodiments, the presentation module 620 is further configured to cause the text controller to highlight a first text content corresponding to a first speaker in the text content relative to a second text content of another speaker, in response to a selection for a first utterance time axis within at least one utterance time axis, where the first utterance time axis corresponds to the first speaker.
[0105] In some embodiments, highlighting the first text content within the text content corresponding to the first speaker in the text controller relative to the second text content of other speakers increases the prominence of the first text content displayed in the text controller and / or decreases the prominence of the second text content displayed in the text controller.
[0106] In some embodiments, the presentation module 620 is further configured to receive a selection for a first position in a first utterance time axis of at least one utterance time axis, and to highlight the text content corresponding to the first position in the text content in the text controller.
[0107] In some embodiments, the presentation module 620 is further configured to present a check entry to the playback controller for checking the speech time axis, and to present at least one speech time axis to the playback controller in response to a selection of the check entry.
[0108] In some embodiments, at least one speech timeline includes multiple speech timelines, and the presentation order of the multiple speech timelines in the playback controller is determined based on at least one of the following: text markers for multiple speakers corresponding to the multiple speech timelines, the proportion of speech content from multiple speakers, or the start times of the speech content from multiple speakers.
[0109] In some embodiments, the presentation module 620 is further configured to receive a check request related to a target speaker and to present a target utterance timeline corresponding to the target speaker, the target utterance timeline being used to show the temporal distribution of the target utterance content of the target speaker.
[0110] In some embodiments, the presentation module 620 is further configured to present multiple visual elements related to multiple speakers associated with audiovisual content, and to receive check requests related to a target speaker based on predetermined operations on a target visual element corresponding to a target speaker within the multiple visual elements.
[0111] In some embodiments, the presentation module 620 is further configured to receive check requests related to a target speaker based on an input indicating a target speaker. In some embodiments, the presentation module 620 is further configured to receive a selection for a second position in a first speech timeline of at least one speech timeline, and to play the corresponding portion of audiovisual content from the point in time corresponding to the second position.
[0112] In some embodiments, the first utterance timeline corresponds to the first speaker, and the presentation module 620 is further configured to play audiovisual content continuously from a given time, or to play portions of the audiovisual content corresponding to the first speaker from a given time.
[0113] In some embodiments, the presentation module 620 is further configured to present corresponding speaker descriptive information in relation to at least one utterance time axis, the descriptive information being generated based on the speaker's text and / or graphical markers.
[0114] In some embodiments, the presentation module 620 is further configured to present percentage information of the corresponding speaker's utterance content in relation to at least one utterance time axis.
[0115] In some embodiments, the playback controller further includes a main time axis, which is used to present graphical information corresponding to the audio waveform of the audiovisual content.
[0116] In some embodiments, the audiovisual content is an audiovisual recording of an online meeting, and the playback controller further includes a main timeline, which is used to present a first interaction indicator corresponding to a first interaction action in the online meeting.
[0117] In some embodiments, at least one speech timeline further presents a second interaction marker to indicate a second interaction behavior related to the corresponding speaker in the online meeting.
[0118] In some embodiments, the first interaction action and / or the second interaction action includes at least one of file sharing, online chat, and commenting.
[0119] In some embodiments, the presentation module 620 is further configured to present first descriptive information relating to a first interaction action in response to a first selection for a first interaction indicator, and / or to present second descriptive information relating to a second interaction action in response to a second selection for a second interaction indicator.
[0120] In some embodiments, the presentation module 620 is further configured to receive a selection for at least one time segment within at least one speech time axis and to generate a first audiovisual content segment corresponding to the at least one time segment for sharing, based on a first sharing request related to the at least one time segment.
[0121] In some embodiments, the presentation module 620 is further configured to receive a selection for a set of speech timelines, which includes one or more speech timelines, within at least one speech timeline, and to generate a second audiovisual content segment corresponding to the speech timeline set for sharing, based on a second sharing request related to the timeline set.
[0122] In some embodiments, the audiovisual content includes real-time audiovisual content, and at least one utterance time axis is used to show the temporal distribution of historical utterance content of at least one speaker related to the historical portion of the real-time audiovisual content.
[0123] The units included in the device 600 may be implemented using a variety of means, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units may be implemented using software and / or firmware, for example, using machine-executable instructions stored in a storage medium. In addition to, or instead of, machine-executable instructions, some or all of the units in the device 600 may be implemented at least partially by one or more hardware logic components. Non-limiting examples of usable and exemplary types of hardware logic components include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), and composite programmable logic devices (CPLDs).
[0124] Figure 7 shows a block diagram of a computing device / server 700 that can implement one or more embodiments of the present disclosure. It should be understood that the computing device / server 700 shown in Figure 7 is merely illustrative and should not constitute any limitation on the functionality and scope of the embodiments described herein.
[0125] As shown in Figure 7, the computing device / server 700 is in the form of a general-purpose computing device. The components of the computing device / server 700 may include, but are not limited to, one or more processors or processing units 710, memory 720, storage devices 730, one or more communication units 740, one or more input devices 760, and one or more output devices 760. The processing unit 710 may be an actual processor or a virtual processor and is capable of performing various processes based on a program stored in memory 720. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to enhance the parallel processing capability of the computing device / server 700.
[0126] The computing device / server 700 typically includes multiple computer storage media. Such media may include, but are not limited to, volatile and non-volatile media, removable and non-removable media, and may be any available media accessible to the computing device / server 700. Memory 720 may include volatile memory (e.g., registers, caches, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or a combination thereof). Storage devices 730 may include machine-readable media such as flash drives, disks, or any other media that are removable or non-removable and can be used to store information and / or data (e.g., training data for training) and are accessible within the computing device / server 700.
[0127] The computing device / server 700 may further include other removable / non-removable, volatile / non-volatile storage media. Not shown in Figure 7, a magnetic disk drive for reading and writing to removable non-volatile magnetic disks (e.g., “floppy disks”) and an optical disk drive for reading and writing to removable non-volatile optical disks may be provided. In these cases, each drive may be connected to a bus (not shown) via one or more data media interfaces. The memory 720 may include a computer program product 725 having one or more program modules configured to perform various methods or operations of various embodiments of the present disclosure.
[0128] The communication unit 740 enables communication with other computing devices via a communication medium. Furthermore, the functionality of the components of the computing device / server 700 may be implemented as a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the computing device / server 700 can use logical connections to one or more other servers, networked personal computers (PCs), or other network nodes to operate in a networked environment.
[0129] The input device 750 may be one or more input devices, such as a mouse, keyboard, or tracking ball. The output device 760 may be one or more output devices, such as a monitor, speaker, or printer. The computing device / server 700 may, if necessary, communicate with one or more external devices (not shown), such as a storage device or display device, via the communication unit 740, with one or more devices that enable a user to interact with the computing device / server 700, or with any device (e.g., a network card or modem) that enables the computing device / server 700 to communicate with one or more other computing devices. Such communication may be performed via an input / output (I / O) interface (not shown).
[0130] In accordance with exemplary embodiments of the present disclosure, a computer-readable storage medium is provided which stores one or more computer instructions, and which are executed by a processor to carry out the method described above.
[0131] Each aspect of this disclosure is described herein with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products implemented in accordance with this disclosure. It should be understood that each box in the flowcharts and / or block diagrams, and any combination thereof, can be implemented by computer-readable program instructions.
[0132] These computer-readable program instructions, when provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, generate a machine such that, when the instructions are executed by the processing unit of the computer or other programmable data processing device, it produces a device that performs the functions / operations specified in one or more boxes of a flowchart and / or block diagram. Furthermore, by storing these computer-readable program instructions, which cause computers, programmable data processing devices, and / or other devices to function in a particular manner, on a computer-readable storage medium, the computer-readable medium containing the instructions has a product containing instructions that perform each of the functions / operations specified in one or more boxes of a flowchart and / or block diagram.
[0133] When computer-readable program instructions are loaded onto a computer, other programmable data processing device, or other device, a series of operational steps are executed on the computer, other programmable data processing device, or other device to generate a computer implementation process, thereby enabling the instructions executed on the computer, other programmable data processing device, or other device to implement a function / operation specified in one or more boxes of a flowchart and / or block diagram.
[0134] The flowcharts and block diagrams in the accompanying drawings illustrate architectures, functions, and operations that may be implemented in several implemented systems, methods, and computer program products relating to this disclosure. In this regard, each box in the flowchart or block diagram may represent a module, program segment, or part of an instruction, and a module, program segment, or part of an instruction may contain one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions attached to the boxes may occur in a different order than those attached to the accompanying drawings. For example, two consecutive boxes may actually be executed substantially in parallel, or in reverse order depending on the functions involved. Also note that each box in the block diagrams and / or flowcharts, and combinations of boxes in the block diagrams and / or flowcharts, may be implemented in a dedicated hardware-based system that performs a given function or operation, or in a combination of dedicated hardware and computer instructions.
[0135] The above descriptions of the various implementations of this disclosure are illustrative, not exhaustive, and not limited to the disclosed implementations. Many modifications and changes will be apparent to an ordinary art engineer without departing from the scope and spirit of the described implementations. The choice of terms used herein is intended to best describe the principles, practical applications, or improvements in the technology in the market of each implementation, or to enable other ordinary art engineers in the art to understand each implementation disclosed herein.
Claims
1. To provide a check interface for audiovisual content, which includes a playback controller for playing the audiovisual content, The playback controller includes presenting at least one speech time axis, A method for checking audiovisual content, wherein the at least one speech time axis is used to show the temporal distribution of speech content of at least one speaker related to the audiovisual content.
2. The method according to claim 1, wherein the check interface further includes a text controller, the text controller is used to present text content corresponding to the audiovisual content, and the text content is generated based on the audio of the audiovisual content.
3. The text controller further includes, in response to a selection for a first speech time axis within the at least one speech time axis, highlighting the first text content corresponding to the first speaker in the text content relative to the second text content of other speakers, The method according to claim 2, wherein the first speaking time axis corresponds to the first speaker.
4. To highlight the first text content corresponding to the first speaker in the aforementioned text content in the text controller relative to the second text content of other speakers is: To increase the prominence of the first text content displayed in the text controller, and / or The objective is to reduce the prominence of the second text content displayed in the text controller. The method according to claim 3.
5. Receiving a selection for a first position in the first speech time axis of the at least one speech time axis, The text content corresponding to the first position within the text content is highlighted in the text controller, The method according to claim 2, further comprising:
6. In the aforementioned playback controller, presenting at least one speech time axis is: The aforementioned playback controller provides a check entry point for checking the speech timeline, In response to the selection of the check entry point, the playback controller presents the at least one speech time axis, The method according to claim 1, including the method described in claim 1.
7. The aforementioned at least one speech time axis includes multiple speech time axes, The method according to claim 1, wherein the presentation order of the plurality of speech timelines in the playback controller is determined based on at least one of the following: the text indicators of the plurality of speakers corresponding to the plurality of speech timelines, the proportion of speech content of the plurality of speakers, or the start time of the speech content of the plurality of speakers.
8. In the aforementioned playback controller, presenting at least one speech time axis is: Receiving check requests related to the target speaker, This includes presenting a timeline of target statements corresponding to the aforementioned target speaker, The method according to claim 1, wherein the target utterance time axis is used to show the temporal distribution of the target utterance content of the target speaker.
9. Receiving a check request related to the target speaker means To present multiple visual elements related to multiple speakers related to the aforementioned audiovisual content, Based on a predetermined operation on the target visual element corresponding to the target speaker within the plurality of visual elements, a check request related to the target speaker is received. The method according to claim 8, including the method described in claim 8.
10. The method of claim 8, wherein receiving a check request related to a target speaker includes receiving a check request related to a target speaker based on an input indicating a target speaker.
11. Receiving a selection for a second position on the first speech time axis of at least one speech time axis, The corresponding portion of the audiovisual content is played back from the point in time corresponding to the second position, The method according to claim 1, further comprising:
12. The aforementioned first utterance time axis corresponds to the first speaker, Playing at least a portion of the aforementioned audiovisual content from a point in time corresponding to the second position means Playing the aforementioned audiovisual content continuously from the aforementioned time, or To play back the portion of the audiovisual content corresponding to the first speaker within the aforementioned audiovisual content from that point in time, The method according to claim 11, including the method described in claim 11.
13. The method further includes presenting corresponding speaker descriptive information in relation to at least one of the aforementioned speech timelines, The method according to claim 1, wherein the descriptive information is generated based on the speaker's text marker and / or graphical marker.
14. The method according to claim 1, further comprising presenting percentage information of the spoken content of a corresponding speaker in relation to the at least one spoken time axis.
15. The method according to claim 1, wherein the playback controller further includes a main time axis, the main time axis being used to present graphical information corresponding to the audio waveform of the audiovisual content.
16. The method according to claim 1, wherein the audiovisual content is an audiovisual recording of an online meeting, and the playback controller further includes a main timeline, the main timeline being used to present a first interaction indicator corresponding to a first interaction action in the online meeting.
17. The method according to claim 16, wherein the at least one speech time axis further presents a second interaction marker for indicating a second interaction action related to the corresponding speaker in the online meeting.
18. The method according to claim 16 or 17, wherein the first interaction action and / or the second interaction action includes at least one of file sharing, online chat, and commenting.
19. In response to a first selection for the first interaction indicator, present first descriptive information relating to the first interaction operation, and / or In response to the second selection for the second interaction indicator, present second descriptive information relating to the second interaction operation. The method according to claim 16 or 17, further comprising:
20. Receiving a selection for at least one time segment in the aforementioned at least one speech time axis, To generate a first audiovisual content segment corresponding to the at least one time segment for sharing, based on a first sharing request related to the at least one time segment, The method according to claim 1, further comprising:
21. Receiving a selection for a set of speech timelines in at least one speech timeline, the set of speech timelines including one or more speech timelines, To generate a second audiovisual content segment corresponding to the speech timeline set for sharing, based on a second sharing request related to the aforementioned timeline set, The method according to claim 1, further comprising:
22. The method according to claim 1, wherein the audiovisual content includes real-time audiovisual content, and the at least one speech time axis is used to show the temporal distribution of historical speech content of at least one speaker related to the historical portion of the real-time audiovisual content.
23. A checking interface for audiovisual content, comprising a providing module configured to provide the checking interface, which includes a playback controller for playing the audiovisual content, The playback controller comprises a presentation module configured to present at least one speech time axis, An apparatus for checking audiovisual content, wherein the at least one speech time axis is used to show the temporal distribution of speech content of at least one speaker related to the audiovisual content.
24. A main timeline that at least indicates the current playback position of the audiovisual content, At least one speech time axis used to show the temporal distribution of speech content of at least one speaker related to the audiovisual content, A playback system that includes this.
25. The playback system according to claim 24, wherein the main time axis further presents graphical information corresponding to the audio waveform of the audiovisual content.
26. The playback system according to claim 24, wherein the audiovisual content is an audiovisual recording of an online meeting, and the main time axis further presents a first interaction marker corresponding to a first interaction action in the online meeting.
27. The playback system according to claim 26, wherein the at least one speech time axis further presents a second interaction marker for indicating a second interaction action related to the corresponding speaker in the online meeting.
28. The playback system according to claim 26 or 27, wherein the first interaction operation and / or the second interaction operation includes at least one of file sharing, online chat, and commenting.
29. The first selection for the first interaction indicator is used to trigger first descriptive information relating to the first interaction operation, and / or The playback system according to claim 26 or 27, wherein the second selection for the second interaction indicator is used to trigger the presentation of second descriptive information relating to the second interaction operation.
30. The aforementioned at least one speech time axis includes multiple speech time axes, The playback system according to claim 24, wherein the presentation order of the plurality of speech time axes in the playback system is determined based on at least one of the following: the text markers of the plurality of speakers corresponding to the plurality of speech time axes, the proportion of the speech content of the plurality of speakers, or the start time of the speech content of the plurality of speakers.
31. At least one processing unit, The system comprises at least one memory coupled to the at least one processing unit, which stores instructions to be executed by the at least one processing unit, When the instruction is executed by the at least one processing unit, the electronic device performs the method according to any one of claims 1 to 22.
32. A computer-readable storage medium that stores a computer program and, when the computer program is executed by a processor, implements the method according to any one of claims 1 to 22.
Citation Information
Patent Citations
Recording method and device based on recorder program, equipment and storage medium
CN112151041A
Conference information intelligent retrieval method
CN113326387A