Method and apparatus for viewing media content, device and storage medium
By presenting speech time and content information in the media content viewing interface, it solves the problem that users find it difficult to quickly locate and understand the speech content of specific speakers, and improves the efficiency of viewing media content.
Patent Information
- Application Number
- PCT/CN2024/126190
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-30
- Filing Date
- 2024-10-21
- Publication Date
- 2025-05-08
AI Technical Summary
When viewing media content, it is difficult for users to quickly locate and understand the speech content of a specific speaker, resulting in inefficiency.
In the viewing interface of media content, speech time information of the speakers associated with media content is presented, and speech content information is presented in association with speech time information, so that users can quickly locate and understand the speech content of interest.
By presenting speech time and content information, users can quickly understand the speech content of specific speakers, improving the efficiency of viewing media content.
Smart Images

Figure CN2024126190_08052025_PF_FP_ABST
Abstract
Description
Method, apparatus, device and storage medium for viewing media content
[0001] This application claims priority to the Chinese invention patent application entitled “Method, apparatus, device and storage medium for viewing media content” and application number 2023114250459, filed on October 30, 2023, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] Example embodiments of the present disclosure relate generally to the field of computers, and more particularly to methods, apparatuses, devices, and computer-readable storage media for viewing media content. Background Art
[0003] With the development of computer technology, the Internet has become the primary platform for people to access and share content. For example, people can use the Internet to publish a wide variety of content or receive content shared by other users. Sharing media content (e.g., audio or video) has become one of the most common forms of Internet-based content sharing. For example, people can use a player to play a video or audio recording of a speech or meeting shared by other users.
[0004] Summary of the Invention
[0005] In a first aspect of the present disclosure, a method for viewing media content is provided. The method comprises: presenting, in a media content viewing interface, speech time information of one or more speakers associated with the media content, the speech time information indicating the location of the speech time period of the one or more speakers in the media content; and presenting, in the viewing interface, speech content information related to the speech content of the one or more speakers in association with the speech time information.
[0006] In a second aspect of the present disclosure, a device for viewing media content is provided. The device includes: a time information presentation module configured to present, in a media content viewing interface, speech time information of one or more speakers associated with the media content, where the speech time information indicates the location of the speech time period of the one or more speakers in the media content; and a content information presentation module configured to present, in the viewing interface, speech content information related to the speech content of the one or more speakers in association with the speech time information.
[0007] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the method of the first aspect.
[0008] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein a computer program is stored on the computer-readable storage medium, and the computer program can be executed by a processor to implement the method of the first aspect.
[0009] It should be understood that the content described in this summary section is not intended to limit the key features or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:
[0011] FIG1 shows a schematic diagram of an example environment in which embodiments of the present disclosure can be implemented;
[0012] FIG2A illustrates an example viewing interface for media content according to some embodiments of the present disclosure;
[0013] 2B-2E illustrate portions of example viewing interfaces for media content according to some embodiments of the present disclosure;
[0014] FIG3 illustrates an example viewing interface for media content according to some embodiments of the present disclosure;
[0015] FIG4 illustrates a flow chart of a process for viewing media content according to some embodiments of the present disclosure;
[0016] FIG5 shows a block diagram of an apparatus for viewing media content according to some embodiments of the present disclosure; and
[0017] FIG6 shows a block diagram of a device capable of implementing various embodiments of the present disclosure. DETAILED DESCRIPTION
[0018] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0019] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.
[0020] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0021] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0022] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.
[0023] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0024] It should be noted that the titles of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and any type of embodiment may be included under any section / subsection. Furthermore, the embodiments described in any section / subsection may be combined in any manner with any other embodiments described in the same section / subsection and / or in different sections / subsections.
[0025] Herein, unless explicitly stated otherwise, executing a step “in response to A” does not mean executing the step immediately after “A” but may include one or more intermediate steps.
[0026] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may be included below. The terms "first", "second", etc. may refer to different or the same objects. Other explicit and implicit definitions may be included below.
[0027] Sample Environment
[0028] Figure 1 shows a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. In this example environment 100, an application 120 is installed on a terminal device 110. A user 140 can interact with the application 120 via the terminal device 110 and / or a device attached to the terminal device 110. In some embodiments, the application 120 may be an application that provides a single primary function, also known as a single-function application. In some embodiments, the application 120 may also be an integrated suite application that provides multiple functions to the user 140. To this end, the suite application may include multiple components for implementing these functions, in which case each component can be considered to correspond to a sub-application. Suite applications in the office field are sometimes also referred to as "office applications," "collaborative office platforms," etc. As examples, the components integrated in the suite application may include, but are not limited to, one or more of the following: a chat component (also known as an instant messaging (IM) component), a document component, an audio and video conferencing component, an email component, a calendar component, a schedule component, a task component, etc. In some embodiments, the suite application can also be downloaded and installed as a single application on the terminal device 110.
[0029] In the environment 100 of Figure 1 , if application 120 is started, terminal device 110 may present an interface 150 of application 120 to user 140. Interface 150 may include entry controls and access portals for various components provided by application 120, as well as interactive interfaces associated with the components, such as a conversation interface presenting chat content, a video conferencing interface, a file sharing interface, and so on.
[0030] Interface 150 can include a viewing interface for media content, through which user 140 can play media content and view information associated with the media content. Such media content can be, for example, video or audio. In some embodiments, media content can include audio and video files locally on terminal device 110, audio and video files stored in the cloud (e.g., server 130), or audio and video streams. Such audio and video streams can, for example, include a playback stream of recorded media content (e.g., meeting minutes), or a live stream of live content.
[0031] In some embodiments, the terminal device 110 communicates with the server 130 to enable the supply of services to the application 120. The terminal device 110 can be any type of mobile terminal, fixed terminal or portable terminal, including a mobile phone, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a television receiver, a radio broadcast receiver, an e-book device, a gaming device or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, the terminal device 110 can also support any type of interface for the user (such as a "wearable" circuit, etc.). The server 130 can be various types of computing systems / servers that can provide computing power, including but not limited to mainframes, edge computing nodes, computing devices in cloud environments, and the like.
[0032] It should be understood that the structure and function of the various elements in the environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of the present disclosure.
[0033] As briefly mentioned above, people can use a player to play a video or audio recording of a meeting. During playback, they may want to locate the portion of the video or audio recording corresponding to a specific speaker (for example, the main speaker or the speaker of primary interest) and understand the content of that speaker's speech. Traditional solutions require viewing the entire media content or the entire transcript of the meeting, which is time-consuming and inefficient.
[0034] To this end, an embodiment of the present disclosure provides a solution for viewing media content. According to an embodiment of the present disclosure, in a viewing interface for media content, the speaking time information of one or more speakers associated with the media content is presented to indicate the position of the speaking time period of one or more speakers in the media content. In the viewing interface, speech content information related to the speech content of one or more speakers is also presented, which can be presented in association with the speaking time information. In this way, the speaking time and speech summary of the speaker are presented to the user at the same time, thereby facilitating a quick understanding of the speech content of the speaker of interest. This can advantageously improve the efficiency of viewing media content.
[0035] The following describes embodiments of the present disclosure with continued reference to the accompanying drawings.
[0036] Sample viewing interface
[0037] Fig. 2 A shows the example of the viewing interface for media content, which can be presented, for example, by terminal device 110. The viewing interface includes a zone 201 for playing media content. As mentioned above, media content can be pre-recorded or being recorded. Media content can be, for example, video or audio. In certain embodiments, media content can be a media file for recording a session (for example, a meeting).
[0038] Region 202 of the viewing interface is used to present information about one or more speakers (e.g., speakers) associated with the media content. As shown in FIG2A , region 202 presents speech time information for speakers A, B, and C, indicating the location of their respective speech time periods within the media content.
[0039] The speaker associated with the media content described herein may refer to the entity that utters the voice in the media content. Exemplarily, if the media content is a record of a conversation, the speaker may be a participant in the conversation. In some embodiments, a speaker may include multiple speakers. For example, such a speaker may be determined based on the terminals participating in the online conference. For example, multiple conference participants may access the online conference through the same terminal (or use the same account), then such multiple participants may be identified as the same speaker, although they may include multiple different speakers. Alternatively, in some embodiments, a speaker may correspond to one speaker. It should be understood that any appropriate speaker recognition technology may be used to determine the corresponding speaker in the media content, and the present disclosure is not intended to be limited to this.
[0040] In some embodiments, information about all speakers associated with the media content may be presented in area 202. Alternatively, in some embodiments, information about some speakers associated with the media content may be presented in area 202. For example, a user may select which speakers' information is presented in area 202. For another example, information about a certain number of speakers may be presented.
[0041] In some embodiments, the speaking time information can be presented to the user in the form of a timeline. For example, for the first speaker, the terminal device 110 can present a first timeline of the media content, and display the first speaking time period of the first speaker in the first timeline in a first style. Correspondingly, for the second speaker, the terminal device 110 can present a second timeline of the media content, and display the second speaking time period of the second speaker in the second timeline in a second style different from the first style. The timelines presented for the speakers (for example, the first timeline and the second timeline mentioned above) can be collectively referred to as speaking timelines.
[0042] Continuing with the example of Figure 2A , area 202 displays a timeline 210-1 for speaker A, a timeline 220-2 for speaker B, and a timeline 220-3 for speaker C, collectively or individually referred to as timelines 220. On each timeline 220, the corresponding speaker's speaking time period is highlighted, with the highlighting pattern varying for each speaker. This allows users to intuitively understand when a speaker was not speaking and when they were speaking frequently. This information helps users quickly access desired content.
[0043] 2A is merely exemplary and not intended to be limiting. Different speakers may be distinguished in any suitable manner, such as by displayed color, displayed line width, and the like.
[0044] Furthermore, representing the speaking time periods using different time axes is merely exemplary, and the speaking time information may be presented in any suitable manner and is not limited to being presented in area 202. For example, the corresponding speaking time periods of each speaker may be indicated in different styles in the playback progress bar of the media content.
[0045] In some embodiments, the presented speech timeline can also support various types of user interactions. As an example, if the terminal device 110 receives a selection of a position in a timeline, the media content can be played starting from the moment corresponding to the selected position. For example, as shown in Figure 2A, the user can click on position 250 in the speech timeline 210-1 to indicate that the media content is desired to be played starting from that position. Accordingly, the terminal device 110 can start playing the media content from the time point corresponding to the position 250.
[0046] In some embodiments, a timeline presented in association with the media content may be displayed in the viewing interface, also referred to as a fourth timeline or a playback timeline. A play mark indicating the current playback moment is displayed on the playback timeline. It should be understood that the "current playback moment" indicates the moment corresponding to the image or frame displayed in the media content display area, and the media content may be in a state of being played or paused. An example is described with reference to FIG2B . As shown in FIG2B , a playback timeline 260 is displayed in the viewing interface, on which a play mark 261 is displayed. The current playback moment is in the speaking time period of speaker A. Accordingly, a position mark 262 is displayed on the speaking timeline 210-1 of speaker A.
[0047] In some embodiments, the play icon on the playback timeline can jump in response to a trigger operation on the speech timeline. In response to a trigger operation (e.g., a user click) on a first position in the speech timeline, the play icon can be displayed at a second position on the playback timeline corresponding to the first position. In other words, the play icon jumps.
[0048] Comparing Figures 2B and 2C , the user triggers (e.g., clicks) position 263 in speech timeline 210-1. Accordingly, the display positions of play mark 261 and position mark 262 change. Another example is described with reference to Figure 2D . The user triggers (e.g., clicks) position 264 in speech timeline 210-2. Accordingly, the display position of play mark 261 changes. Position mark 262 is displayed on speech timeline 210-2 instead of speech timeline 210-1.
[0049] In some embodiments, whether a position within the speech timeline is triggerable depends on whether the speaker corresponding to the speech timeline is speaking at the moment corresponding to the position. Specifically, a position on the speech timeline corresponding to a speaker that is within the speech time period of the speaker is triggerable. For example, a position on the speech timeline 210-1 that is within the speech time period of speaker A is triggerable, and a position on the speech timeline 210-2 that is within the speech time period of speaker B is triggerable. For example, in the example of Figure 2E, position 265 on the speech timeline 210-2 is not within the speech time period of speaker B and is therefore not triggerable. If the user attempts to trigger the position, a prohibition sign 266 may be presented.
[0050] In some embodiments, the terminal device 110 may also display speech ratio information for one or more speakers in the viewing interface, indicating the ratio of the corresponding speech time period of the one or more speakers to the duration of the media content. For example, in the example of FIG2A , the ratios "YY%," "ZZ%," and "WW%" are displayed for speakers A, B, and C, respectively.
[0051] In some embodiments, the order in which speech timelines are presented may be determined based on the speech percentage of the speakers. For example, if the speech percentage of "Speaker A" reaches "70%", which is greater than the speech percentage of "Speaker B" of "30%", speech timeline 210-1 may be presented in priority over speech timeline 210-2, as shown in FIG2A .
[0052] Reference is made to Figure 3, which shows an example of a viewing interface for media content, which can be presented, for example, by the terminal device 110. In this viewing interface, speech content information related to the speech content of the speaker is presented. The speech content information may include any suitable information related to the speech content of the speaker. For example, the speech content information of a certain speaker may include the important speech content or excerpts of the speech content of the speaker. In some embodiments, the speech content information may include speech summary information to summarize the speech content of one or more speakers. For example, a speech summary of the speech content of each speaker may be presented.
[0053] Speech content information can be presented in association with speech time information. For example, each speaker's speech content information can be presented in association with that speaker's speech time information. Specifically, Figure 3 shows speech content information 310-1 for speaker A and speech content information 310-2 for speaker B, collectively or individually referred to as speech content information 310. It should be understood that, due to limitations on the display area of the interface, speaker C's speech content information could be displayed by sliding the screen, for example.
[0054] In some embodiments, each speaker's speech content information can be presented adjacent to the speaker's speech timeline. For example, as shown in Figure 3, speaker A's speech content information 310-1 is adjacent to speech timeline 210-1, specifically below speech timeline 210-1; speaker B's speech content information 310-2 is adjacent to speech timeline 210-2, specifically below speech timeline 210-2.
[0055] In some embodiments, speech content information may be presented by default in the viewing interface. In some embodiments, speech content information may be presented in response to user input. For example, a first control for displaying speech content information may be presented in the viewing interface. If the first control is triggered by the user, the terminal device 110 may present the speech content information in the viewing interface. For example, FIG2A shows a first control 220. In response to the user clicking on the first control 220, the speech content information 310 shown in FIG3 is presented.
[0056] In some embodiments, in response to the presentation of speech content information, terminal device 110 may present a second control in the viewing interface for hiding the speech content information. If the second control is triggered by the user, the presentation of the speech content information in the viewing interface may cease. For example, after the speech summary 310 is presented, the first control 220 in FIG. 2A transforms into the second control 320 shown in FIG. 3 . If the user clicks on the second control 320, the speech content information 310 may be retracted.
[0057] In the examples described above, the controls for presenting and hiding speech content information are shared by different speakers. In other examples, separate controls can be set for each speaker to present and hide that speaker's speech content information. In this way, users can learn more about the speeches of speakers they are interested in.
[0058] In some embodiments, the speaking time period of a speaker may include multiple sub-time periods. Accordingly, speech content information related to the speech content in at least one of these sub-time periods, such as a speech summary, may be presented. Such a sub-time period may be any sub-time period, or may be a time period whose length exceeds a threshold length. For example, as shown in the timeline 210-1 for speaker A, the speaking time period of speaker A includes multiple sub-time periods. If the length of one or more of the time periods exceeds the threshold, a summary of the speech content of such time periods may be presented. In this way, the user can understand the speech content of the speaker at a refined time granularity, thereby more accurately locating the speech content of interest or concern.
[0059] In some embodiments, a speech summary corresponding to a sub-time period may be presented. In such embodiments, if the speech summary is selected, a location marker may be displayed on the speaker's speech timeline at the sub-time period corresponding to the selected speech summary. This further facilitates viewing speech content of interest or concern.
[0060] It should be understood that, for descriptive purposes only, although examples of viewing interfaces are discussed above in conjunction with Figures 2A to 3 , the various aspects discussed above (e.g., the provision of a speech timeline, the style of a speech timeline, the interaction of a speech timeline, the style of controls, the presentation position of various elements, etc.) may be provided independently or in a combination different from that shown in Figures 2A to 3 .
[0061] Furthermore, while the examples shown in Figures 2A to 3 are for reviewing recorded content, embodiments of the present disclosure may also be used to view real-time media content (e.g., live audio and video streams). Accordingly, the speech timeline discussed above may be used to indicate the temporal distribution of historical speech content of at least one speaker associated with a historical portion of real-time media content.
[0062] Example Process
[0063] 4 shows a flow chart of a process 400 for viewing media content according to some embodiments of the present disclosure. The process 400 may be implemented at the terminal device 110. The process 400 is described below with reference to the figure.
[0064] In block 410, the terminal device 110 presents the speaking time information of one or more speakers associated with the media content in the viewing interface of the media content. The speaking time information indicates the position of the speaking time period of the one or more speakers in the media content.
[0065] In block 410 , the terminal device 110 presents, in a viewing interface, speech content information related to speech content of the one or more speakers in association with speech time information.
[0066] In some embodiments, one or more speakers include a first speaker and a second speaker, and presenting the speaking time information of one or more speakers includes: for the first speaker, presenting a first timeline of media content, and the first speaking time period of the first speaker is displayed in a first style in the first timeline; and for the second speaker, presenting a second timeline of media content, and the second speaking time period of the second speaker is displayed in a second style different from the first style in the second timeline.
[0067] In some embodiments, process 400 further includes receiving a selection of a location in the first timeline; and starting playing the media content from a time corresponding to the selected location.
[0068] In some embodiments, presenting the speech content information in association with the speech time information includes: presenting the first speech content information of the first speaker adjacent to the first timeline; and presenting the second speech content information of the second speaker adjacent to the second timeline.
[0069] In some embodiments, process 400 further includes: presenting speech ratio information of one or more speakers in the viewing interface, where the speech ratio information indicates the ratio of corresponding speech time periods of the one or more speakers to the duration of the media content.
[0070] In some embodiments, presenting speech content information of one or more speakers includes: presenting a first control for displaying speech content information in a viewing interface; and presenting the speech content information in the viewing interface in response to a triggering operation on the first control.
[0071] In some embodiments, process 400 further includes: in response to the presentation of the speech content information, presenting a second control for hiding the speech content information in the viewing interface; and in response to a triggering operation on the second control, stopping presenting the speech content information in the viewing interface.
[0072] In some embodiments, the speaking time period of a third speaker among one or more speakers includes multiple sub-time periods, and presenting the speech content information includes: for at least one sub-time period among the multiple sub-time periods, presenting third speech content information related to the speech content of the third speaker in at least one sub-time period.
[0073] In some embodiments, presenting the third speech content information includes presenting at least one speech content summary corresponding to at least one sub-time period. Process 400 also includes, in response to selecting a speech content summary from the at least one speech content summary, presenting a position marker at the sub-time period corresponding to the selected speech content summary on the third timeline, where the third timeline is used to indicate the speech time period of the third speaker.
[0074] In some embodiments, a fourth timeline presented in association with the media content is displayed in the viewing interface, a play mark indicating the current playback moment is displayed on the fourth timeline, and process 400 further includes: in response to a trigger operation on a first position in the first timeline or the second timeline, displaying a play mark at a second position corresponding to the first position on the fourth timeline.
[0075] In some embodiments, a position on the first timeline that is within the first speech time period is triggerable, and a position on the second timeline that is within the second speech time period is triggerable.
[0076] Example devices and equipment
[0077] 5 shows a schematic structural block diagram of an apparatus 500 for viewing media content according to certain embodiments of the present disclosure. Apparatus 500 may be implemented as or included in terminal device 110. Each module / component in apparatus 500 may be implemented by hardware, software, firmware, or any combination thereof.
[0078] As shown, apparatus 500 includes a time information presentation module 510 configured to present, in a media content viewing interface, speech time information of one or more speakers associated with the media content. The speech time information indicates the location of the corresponding speech time period of the one or more speakers within the media content. Apparatus 500 also includes a content information presentation module 520 configured to present speech content information related to the speech content of the one or more speakers in association with the speech time information.
[0079] In some embodiments, one or more speakers include a first speaker and a second speaker, and the time information presentation module 510 is further configured to: for the first speaker, present a first timeline of media content, and the first speaking time period of the first speaker is displayed in a first style in the first timeline; and for the second speaker, present a second timeline of media content, and the second speaking time period of the second speaker is displayed in a second style different from the first style in the second timeline.
[0080] In some embodiments, the apparatus 500 further includes: a selection receiving module configured to receive a selection of a position in the first timeline; and a content playing module configured to start playing the media content from a time corresponding to the selected position.
[0081] In some embodiments, the content information presentation module 520 is further configured to: present a first speech summary of the first speaker adjacent to the first timeline; and present a second speech summary of the second speaker adjacent to the second timeline.
[0082] In some embodiments, the device 500 also includes: a proportion information presentation module, configured to present the speech proportion information of one or more speakers in the viewing interface, where the speech proportion information indicates the proportion of the corresponding speech time period of one or more speakers to the duration of the media content.
[0083] In some embodiments, the content information presentation module 520 is further configured to: present a first control for displaying speech content information in the viewing interface; and present the speech content information in the viewing interface in response to a triggering operation on the first control.
[0084] In some embodiments, the content information presentation module 520 is further configured to: in response to the presentation of the speech content information, present a second control for hiding the speech content information in the viewing interface; and in response to a triggering operation on the second control, stop presenting the speech content information in the viewing interface.
[0085] In some embodiments, the speech time period of a third speaker among one or more speakers includes multiple sub-time periods, and the speech information presentation module 520 is further configured to: for at least one sub-time period among the multiple sub-time periods, present third speech content information related to the speech content of the third speaker in the at least one sub-time period.
[0086] In some embodiments, the speech information presentation module 520 is further configured to present at least one speech content summary corresponding to at least one sub-time period. The apparatus 500 further includes a position marker presentation module configured to, in response to selection of a speech content summary from the at least one speech content summary, present a position marker at the sub-time period corresponding to the selected speech content summary on the third timeline, the third timeline being used to indicate the speech time period of the third speaker.
[0087] In some embodiments, a fourth timeline presented in association with the media content is displayed in the viewing interface, a play mark indicating the current play time is displayed on the fourth timeline, and the device 500 also includes: a play mark display module, configured to display a play mark at a second position corresponding to the first position on the fourth timeline in response to a trigger operation on the first position in the first timeline or the second timeline.
[0088] In some embodiments, a position on the first timeline that is within the first speech time period is triggerable, and a position on the second timeline that is within the second speech time period is triggerable.
[0089] FIG6 shows a block diagram of an electronic device 600 in which one or more embodiments of the present disclosure may be implemented. It should be understood that the electronic device 600 shown in FIG6 is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. The electronic device 600 shown in FIG6 can be used to implement the electronic device 110 of FIG1 .
[0090] As shown in FIG6 , electronic device 600 is a general-purpose electronic device. Components of electronic device 600 may include, but are not limited to, one or more processors or processing units 610, memory 620, storage device 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660. Processing unit 610 may be a real or virtual processor and is capable of performing various processes according to programs stored in memory 620. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to enhance the parallel processing capabilities of electronic device 600.
[0091] The electronic device 600 typically includes a plurality of computer storage media. Such media can be any accessible media that can be obtained by the electronic device 600, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 620 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 630 can be a removable or non-removable medium and can include a machine-readable medium, such as a flash drive, a disk, or any other medium that can be used to store information and / or data and can be accessed within the electronic device 600.
[0092] The electronic device 600 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG6 , a disk drive for reading or writing from a removable, non-volatile disk (e.g., a “floppy disk”) and an optical drive for reading or writing from a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. The memory 620 may include a computer program product 625 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.
[0093] The communication unit 640 enables communication with other electronic devices via a communication medium. Additionally, the functions of the components of the electronic device 600 can be implemented in a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the electronic device 600 can operate in a networked environment using a logical connection with one or more other servers, a network personal computer (PC), or another network node.
[0094] The input device 650 may be one or more input devices, such as a mouse, keyboard, or trackball. The output device 660 may be one or more output devices, such as a display, a speaker, or a printer. The electronic device 600 may also communicate with one or more external devices (not shown) through the communication unit 640 as needed, such as a storage device, a display device, or the like, with one or more devices that allow a user to interact with the electronic device 600, or with any device that allows the electronic device 600 to communicate with one or more other electronic devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).
[0095] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above.
[0096] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0097] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0098] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0099] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple implementations of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part for a module, program segment or instruction, and a part for a module, program segment or instruction comprises one or more executable instructions for realizing the logical function of the specification. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.
[0100] While various implementations of the present disclosure have been described above, the foregoing description is intended to be illustrative, not exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is selected to best explain the principles of the implementations, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. A method for viewing media content, comprising: In a viewing interface of the media content, presenting speech time information of one or more speakers associated with the media content, wherein the speech time information indicates the position of the speech time period of the one or more speakers in the media content; as well as In the viewing interface, speech content information related to speech content of the one or more speakers is presented in association with the speech time information.
2. The method according to claim 1, wherein the one or more speakers include a first speaker and a second speaker, and presenting the speech time information of the one or more speakers comprises: For the first speaker, present a first timeline of the media content, wherein a first speaking time period of the first speaker is displayed in a first style in the first timeline; as well as For the second speaker, a second timeline of the media content is presented, and a second speaking time period of the second speaker is displayed in a second style different from the first style in the second timeline.
3. The method according to claim 2, further comprising: receiving a selection of a location in the first timeline; as well as The media content is played starting from a time corresponding to the selected position.
4. The method according to claim 2, wherein presenting the speech content information in association with the speech time information comprises: presenting first speech content information of the first speaker adjacent to the first timeline; as well as The second speech content information of the second speaker is presented adjacent to the second time axis.
5. The method according to claim 1, further comprising: In the viewing interface, speech ratio information of the one or more speakers is presented, and the speech ratio information indicates the ratio of the corresponding speech time periods of the one or more speakers to the duration of the media content.
6. The method according to claim 1, wherein presenting the speech content information of the one or more speakers comprises: In the viewing interface, presenting a first control for displaying the speech content information; as well as In response to a triggering operation on the first control, the speech content information is presented in the viewing interface.
7. The method according to claim 6, further comprising: In response to the presentation of the speech content information, presenting a second control for hiding the speech content information in the viewing interface; as well as In response to the triggering operation on the second control, the presentation of the speech content information in the viewing interface is stopped. interest.
8. The method according to claim 1, wherein the speech time period of the third speaker among the one or more speakers includes a plurality of sub-time periods, and presenting the speech content information includes: For at least one sub-time period among the multiple sub-time periods, third speech content information related to speech content of the third speaker in the at least one sub-time period is presented.
9. The method according to claim 8, wherein presenting the third speech content information comprises: presenting at least one speech content summary corresponding to the at least one sub-time period, and The method further comprises: In response to selection of a speech content summary from the at least one speech content summary, a position mark is presented at a sub-time period corresponding to the selected speech content summary on a third time axis, and the third time axis is used to indicate a speech time period of the third speaker.
10. The method according to claim 2, wherein the viewing interface displays a fourth timeline associated with the media content, a play mark indicating a current play time is displayed on the fourth timeline, and the method further comprises: In response to a trigger operation on a first position in the first timeline or the second timeline, the play mark is displayed at a second position corresponding to the first position on the fourth timeline. 11 . The method according to claim 10 , wherein the position on the first time axis within the first speech time period is triggerable, and the position on the second time axis within the second speech time period is triggerable.
12. An apparatus for viewing media content, comprising: A time information presentation module, configured to present, in a viewing interface of the media content, speech time information of one or more speakers associated with the media content, wherein the speech time information indicates the position of the speech time period of the one or more speakers in the media content; as well as The speech information presenting module is configured to present, in the viewing interface, speech content information related to speech content of the one or more speakers in association with the speech time information.
13. An electronic device, comprising: at least one processing unit; as well as At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to any one of claims 1 to 11 when executed by the at least one processing unit.
14. A computer-readable storage medium having a computer program stored thereon, wherein the computer program can be executed by a processor to implement the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Method and device for viewing media content, equipment and storage medium
CN119922363A
Recording method and device based on recorder program, equipment and storage medium
CN112151041A
Conference recording method, terminal equipment and conference recording system
CN116193179A
Method, device and equipment for viewing audiovisual content and storage medium
CN117956233A
Voice recording and reproducing device
JP2003122397A