Conference recording method, terminal device and conference recording system
By providing a meeting review interface with the dimensions and time dimensions of participants after the online meeting, audio identification and data playback, the problem of difficulty in finding information in online meetings is solved, and fast and intuitive information positioning is achieved.
Patent Information
- Application Number
- CN202111424519.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-26
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2041-11-26
AI Technical Summary
In the online meeting minutes, due to the differentiation of oral expression and written expression, it is difficult for users to quickly and effectively locate the required information, and the existing technology cannot effectively solve this problem.
It provides a meeting record method, through a meeting review interface that displays the dimensions and time dimensions of the participants after the online meeting, the audio identifier identifies the participants' speech time period, and supports the playback and filtering of audio data and video data, allowing participants to filter the required information through participants and time nodes.
It quickly and intuitively locates the target information in the conference record, improves the efficiency of finding and locates the target information in the conference record, and solves the problem of difficulty in finding information in online meetings.
Smart Images

Figure CN116193179B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of online conference technology, and in particular to a conference recording method, terminal equipment, and conference recording system. Background Art
[0002] With the rapid development of Internet technology, online conference applications (APPs) have been widely used. Different users in different locations can use online conference applications to participate in the same conference, which greatly facilitates different users in different locations to participate in the conference.
[0003] Online conferencing applications can provide users with the ability to conduct audio or video conferences. However, these applications also present numerous challenges. For example, because participants using these applications are in different locations and environments, online meetings lack the formality of formal meetings. Participants may not be as engaged as in face-to-face meetings, leading to missed key points. Alternatively, participants may be taking notes during the meeting, missing other key points during the recording process. Therefore, the value of meeting minutes is becoming increasingly prominent for online meetings, as users can use the minutes to implement the conclusions after the meeting. Online conferencing applications typically generate meeting minutes by automatically converting participant audio into text, generating a transcript of the entire online meeting conversation. This transcript serves as a meeting record for users to review afterward. When a user needs to find target information during the meeting, the application detects keywords entered by the user, uses these keywords to search for the target information within the meeting record generated using the aforementioned method, and then outputs the target information for the user to review.
[0004] However, due to the differences between spoken and written expressions, as well as the fact that different users may express the same intent differently, and even the same user may express it differently at different times, the keywords entered by the user may not match the target information in the meeting minutes. Consequently, online meeting applications that use this method to obtain meeting minutes can make it more difficult for users to find the information they need after the meeting, preventing them from quickly and effectively locating the target information. Summary of the Invention
[0005] The embodiments of the present application provide a meeting recording method, terminal device and meeting recording system, which can provide a meeting recording function that can quickly and effectively locate the required target information, which is beneficial to improving the efficiency of finding and locating target information in meeting records.
[0006] In a first aspect, an embodiment of the present application provides a meeting recording method, which is applied to a client that provides an online meeting function. The online meeting includes an online audio conference or an online video conference. The method may include: detecting a trigger operation of an online meeting end option or a trigger operation of an online meeting record viewing option; in response to the trigger operation of the online meeting end option or the trigger operation of the online meeting record viewing option, displaying an online meeting review interface, the online meeting review interface including a timeline of multiple participants and multiple audio identifiers distributed along the timeline, each of the multiple audio identifiers being used to identify a segment of audio data of the participant on the timeline where the audio identifier is located, the starting position of the audio identifier on the timeline is the starting time of the segment of audio data of the participant, and the ending position of the audio identifier on the timeline is the ending time of the segment of audio data of the participant, and the segment of audio data of the participant is generated by recording the conference speech voice signal of the participant in the time period between the start time and the end time.
[0007] Thus, in the first aspect of the embodiments of the present application, after an online meeting concludes, the online meeting review interface can provide participants with meeting record information in both participant and time dimensions. Participants can filter required information by participant and / or time node, and independently reproduce all or part of the meeting scene. This online meeting review interface can provide a meeting record function that can quickly and effectively locate the desired target information, which is beneficial to improving the efficiency of finding and locating target information in the meeting record.
[0008] A participant's timeline can have one or more audio identifiers. The number of audio identifiers and the length of each audio identifier are related to the participant's speaking time in the online meeting. The start time of an audio identifier on the participant's timeline is the start time of the participant's speech in the online meeting, and the end time of the audio identifier on the participant's timeline is the end time of the participant's speech.
[0009] In one possible design, the method further includes: detecting a first trigger operation of an audio identifier among the multiple audio identifiers, and displaying keywords of audio data corresponding to the audio identifier in response to the first trigger operation.
[0010] A possible design of an embodiment of the present application is to enable participants to filter audio data by keywords by displaying keywords of audio data corresponding to audio identifiers, thereby improving the efficiency of finding and locating target information in meeting records.
[0011] In one possible design, the method further includes: detecting a second triggering operation of an audio identifier among the multiple audio identifiers, and playing audio data corresponding to the audio identifier in response to the second triggering operation.
[0012] A possible design of an embodiment of the present application provides an online meeting review interface. After the online meeting, participants can click on the corresponding audio identifier on the online meeting review interface according to the participants and the time node, thereby playing the audio data corresponding to the audio identifier and reproducing the participant's speech in audio form, so as to know the specific speech content of the participant at the corresponding time point. This can locate the required information more simply, intuitively and quickly, which is beneficial to improving the efficiency of finding and locating target information in the meeting minutes. Replacing text with audio data can, on the one hand, solve the problems of being unable to find or difficult to understand caused by the inconsistency between oral expression and text expression. On the other hand, the listening method can be understood by the brain faster than the reading method, and can compare and filter information faster.
[0013] In one possible design, the method also includes: detecting a third trigger operation of a first audio identifier among the multiple audio identifiers from the first time axis to the second time axis, the first audio identifier is distributed on the first time axis, and at least one second audio identifier is distributed on the second time axis; in response to the third trigger operation, displaying the first audio identifier on the second time axis, and associating the first audio data with the second audio data corresponding to at least one second audio identifier based on the start time and end time of the first audio data corresponding to the first audio identifier; detecting a fourth trigger operation of the first audio identifier or the at least one second audio identifier, and in response to the fourth trigger operation, playing the first audio data and the second audio data corresponding to the at least one second audio identifier.
[0014] A possible design of the embodiment of the present application is that if different participants make a joint speech within a period of time, you can listen to the speech content of each participant separately by clicking on the audio identifier on the timeline of each participant, and you can also associate multiple audio identifiers through the above-mentioned third trigger operation to simultaneously play the speech content of the participant corresponding to each of the multiple audio identifiers, thereby reproducing the meeting discussion scene at that time. In particular, for content that is hotly discussed in the meeting, due to the intense discussion, it may be impossible to grasp everyone's point of view. After the meeting, the speech content of different participants can be split, integrated, and repeatedly analyzed through audio identifiers to grasp the point of view of different participants in the meeting.
[0015] In one possible design, the time period of the first audio identifier overlaps with the time period of at least one second audio identifier, and playing the first audio data and the second audio data corresponding to the at least one second audio identifier includes: according to the time position of the first audio identifier on the second time axis and the time position of the at least one second audio identifier on the second time axis, playing at least two of the non-intersecting parts of the first audio data, the intersecting parts of the first audio data and the second audio data corresponding to the at least one second audio identifier, or the non-intersecting parts of the second audio data corresponding to the at least one second audio identifier.
[0016] In one possible design, the method also includes: detecting a fifth trigger operation of the first audio identifier, displaying the first audio identifier on the first timeline in response to the fifth trigger operation, and disassociating the first audio data and the second audio data corresponding to at least one second audio identifier, wherein the disassociation is used to independently play the first audio data and the second audio data corresponding to the second audio identifier.
[0017] A possible design of an embodiment of the present application is to cancel the association of multiple audio identifiers through the above-mentioned fifth trigger operation, so as to independently play the speech content of the participants corresponding to the multiple audio identifiers, so that the audio data of a participant can be flexibly selected for playback, so as to flexibly reproduce part of the information of the meeting discussion scene at that time, and improve the efficiency of finding and locating target information in the meeting records.
[0018] In one possible design, before detecting the trigger operation of the online meeting end option or the trigger operation of the online meeting record viewing option, the method also includes: recording and generating at least one audio data of the participant using the client, and recording the start time and end time of each of the at least one audio data on the participant's timeline; sending the at least one audio data and the start time and end time of each of the at least one audio data on the participant's timeline to the server; wherein, the at least one audio data and the start time and end time of each of the at least one audio data on the participant's timeline are used to generate an online meeting review interface.
[0019] For example, if a participant using the client makes three speeches in an online meeting, three audio data points of the participant are recorded, and the start and end times of each audio data point on the participant's timeline are recorded. For example, the start time of one of the three audio data points on the participant's timeline is the start time of the participant's speech, and the end time of the audio data on the participant's timeline is the end time of the participant's speech.
[0020] In one possible design, the at least one audio data and the at least one audio data are respectively stored in at least one storage unit at the start time and the end time of the participant's time axis; the at least one storage unit is connected in series through the participant's time pointer.
[0021] In one possible design, the method also includes: detecting a sixth trigger operation of an audio identifier among the multiple audio identifiers, and in response to the sixth trigger operation, displaying a thumbnail of the video data corresponding to the audio identifier, where the video data corresponding to the audio identifier is generated by the main interface screen during the time period between the start time and the end time of recording the audio data corresponding to the audio identifier.
[0022] One possible design of an embodiment of the present application involves recording the main interface video and associating the audio and video data of each participant based on time information. When the online meeting review interface is displayed, a thumbnail of the video data corresponding to the audio identifier is displayed based on the sixth trigger operation detected. By displaying thumbnails of the video data associated with the audio data, participants can filter the audio and video data using the thumbnails, thereby improving the efficiency of finding and locating target information in the meeting record.
[0023] In one possible design, the method further includes: detecting a seventh trigger operation of the thumbnail, and in response to the seventh trigger operation, playing the video data and audio data corresponding to the audio identifier.
[0024] A possible design of an embodiment of the present application is to play the video data and audio data corresponding to the audio identifier in response to the seventh trigger operation, so as to completely reproduce the meeting scene at that time through a combination of audio and video, thereby speeding up the comparison and determination process of the target information.
[0025] In one possible design, the online meeting review interface also includes at least one annotation marker, which is distributed on the timeline of at least one participant. Each of the at least one annotation marker is used to identify the participant on the timeline where the annotation marker is located, and the annotation at the time point where the annotation marker is located.
[0026] A possible design of an embodiment of the present application is to present the annotation marks generated by the annotation actions of the participants at the corresponding time points on the online meeting review interface, so that the participants can quickly locate the time nodes of the key content through the annotation marks after the meeting.
[0027] At least one annotation identifier can include the identifier of each participant as a speaker. This ensures that the annotation identifier has corresponding annotation content in the meeting minutes. At least one annotation identifier can also include the annotation identifier of the participant using the client of this embodiment. This allows different clients to obtain personalized annotated meeting minutes for independent viewing.
[0028] In one possible design, the method further includes: detecting an eighth trigger operation of one of the at least one annotation identifier, and in response to the eighth trigger operation, playing the audio data, or the audio data and video data, at the time point where the annotation identifier is located.
[0029] In one possible design, the timelines of the multiple participants and the multiple audio identifiers distributed along the timelines are located in the timeline area of the online conference review interface, and the online conference review interface also includes a video display area.
[0030] Detecting an operation of reducing the timeline area, and in response to the operation of reducing the timeline area, reducing the timeline area, reducing the distance between the timelines of the multiple participants, and increasing the video display area; or detecting an operation of increasing the timeline area, and in response to the operation of increasing the timeline area, increasing the timeline area and reducing the video display area.
[0031] One possible design of an embodiment of the present application reduces the size of the timeline area and increases the size of the video area, allowing for clearer viewing of video content. By increasing the size of the timeline area and decreasing the size of the video area, the timelines of each participant can be separated, allowing for flexible playback of audio data and / or video data corresponding to the audio identifiers of different participants. This allows for flexible adjustment of the timeline area and video display area to meet user needs.
[0032] In one possible design, when the distance between the time axes of the multiple participants is reduced until the time axes of the multiple participants completely overlap, at least two audio identifiers among the multiple audio identifiers overlap with each other, and the method further includes:
[0033] A tenth trigger operation of detecting two audio identifiers that overlap each other is detected. In response to the tenth trigger operation, one of the two audio identifiers is displayed on a timeline, and the other audio identifier is displayed above or below the timeline so that the two audio identifiers do not overlap.
[0034] One possible design of an embodiment of the present application is to merge the timelines of multiple participants into one timeline by reducing the timeline area and associating the audio data of multiple participants. The associated multiple audio data can be used by the user to review the entire online meeting as a whole. After merging into one timeline, the tenth trigger operation can also be used to make the two audio identifiers non-overlapping, so that the audio data corresponding to different audio identifiers can be played independently. In this way, flexible reproduction of the meeting content can be achieved.
[0035] In a second aspect, an embodiment of the present application provides a terminal device, comprising: a processor, a memory, and a display screen, wherein the memory and the display screen are coupled to the processor, the memory being used to store computer program code, the computer program code comprising computer instructions for a client providing an online conference function, wherein the online conference includes an online audio conference or an online video conference, and when the processor reads the computer instructions from the memory, the terminal device performs the following operations:
[0036] Detecting the triggering operation of the online meeting end option or the online meeting record viewing option;
[0037] In response to the trigger operation of the online meeting end option or the trigger operation of the online meeting record viewing option, the online meeting review interface is displayed. The online meeting review interface includes a timeline of multiple participants and multiple audio identifiers distributed along the timeline. Each of the multiple audio identifiers is used to identify a segment of audio data of the participant on the timeline where the audio identifier is located. The starting position of the audio identifier on the timeline is the starting time of the segment of audio data of the participant, and the ending position of the audio identifier on the timeline is the ending time of the segment of audio data of the participant. The segment of audio data of the participant is generated by recording the conference speech voice signal of the participant in the time period between the start time and the end time.
[0038] In one possible design, the terminal device further performs: detecting a first trigger operation of an audio identifier among the multiple audio identifiers, and displaying a keyword of the audio data corresponding to the audio identifier in response to the first trigger operation.
[0039] In one possible design, the terminal device further performs: detecting a second triggering operation of an audio identifier among the multiple audio identifiers, and playing audio data corresponding to the audio identifier in response to the second triggering operation.
[0040] In one possible design, the terminal device also performs: detecting a third trigger operation of the first audio identifier among the multiple audio identifiers from the first time axis to the second time axis, the first audio identifier is distributed on the first time axis, and at least one second audio identifier is distributed on the second time axis; in response to the third trigger operation, displaying the first audio identifier on the second time axis, and associating the first audio data with the second audio data corresponding to the at least one second audio identifier based on the start time and end time of the first audio data corresponding to the first audio identifier; detecting a fourth trigger operation of the first audio identifier or the at least one second audio identifier, and in response to the fourth trigger operation, playing the first audio data and the second audio data corresponding to the at least one second audio identifier.
[0041] In one possible design, the time period of the first audio identifier overlaps with the time period of the at least one second audio identifier, and the playing of the first audio data and the second audio data corresponding to the at least one second audio identifier includes: according to the time position of the first audio identifier on the second time axis and the time position of the at least one second audio identifier on the second time axis, playing at least two of the non-intersecting part of the first audio data, the intersecting part of the first audio data and the second audio data corresponding to the at least one second audio identifier, or the non-intersecting part of the second audio data corresponding to the at least one second audio identifier.
[0042] In one possible design, the terminal device also performs: detecting a fifth trigger operation of the first audio identifier; in response to the fifth trigger operation, displaying the first audio identifier on the first timeline; and disassociating the first audio data and the second audio data corresponding to the at least one second audio identifier; the disassociation is used for the first audio data and the second audio data corresponding to the second audio identifier to be played independently.
[0043] In one possible design, before detecting the trigger operation of the online meeting end option or the trigger operation of the online meeting record viewing option, the terminal device also performs the following: recording and generating at least one audio data of a participant using the client, and recording the start time and end time of each of the at least one audio data on the participant's timeline; sending the at least one audio data and the start time and end time of each of the at least one audio data on the participant's timeline to the server; wherein, the at least one audio data and the at least one audio data are each used to generate the online meeting review interface at the start time and end time of the participant's timeline.
[0044] In one possible design, the at least one audio data and the at least one audio data are respectively stored in at least one storage unit at the start time and the end time of the participant's time axis; the at least one storage unit is connected in series through the time pointer of the participant.
[0045] In one possible design, the terminal device also performs: detecting a sixth trigger operation of an audio identifier among the multiple audio identifiers, and in response to the sixth trigger operation, displaying a thumbnail of the video data corresponding to the audio identifier, where the video data corresponding to the audio identifier is generated by the main interface screen during the time period between the start time and the end time of recording the audio data corresponding to the audio identifier.
[0046] In one possible design, the terminal device further performs: detecting a seventh trigger operation of the thumbnail, and in response to the seventh trigger operation, playing the video data and audio data corresponding to the audio identifier.
[0047] In one possible design, the online meeting review interface also includes at least one annotation marker, which is distributed on the timeline of at least one participant. Each of the at least one annotation marker is used to identify the participant on the timeline where the annotation marker is located, and the annotation at the time point where the annotation marker is located.
[0048] In one possible design, the terminal device further performs: detecting an eighth trigger operation of a marking identifier among the at least one marking identifier, and in response to the eighth trigger operation, playing the audio data, or the audio data and video data, at the time point where the marking identifier is located.
[0049] In one possible design, the timelines of the multiple participants and the multiple audio identifiers distributed along the timelines are located in the timeline area of the online meeting review interface, which also includes a video display area; an operation of shrinking the timeline area is detected, and in response to the operation of shrinking the timeline area, the timeline area is shrunk, the distance between the timelines of the multiple participants is reduced, and the video display area is enlarged; or, an operation of enlarging the timeline area is detected, and in response to the operation of enlarging the timeline area, the timeline area is enlarged and the video display area is reduced.
[0050] In one possible design, when the distance between the time axes of the multiple participants is reduced until the time axes of the multiple participants completely overlap, at least two audio identifiers among the multiple audio identifiers overlap with each other, and the terminal device also performs: detecting a tenth trigger operation of the two overlapping audio identifiers, and in response to the tenth trigger operation, displaying one of the two audio identifiers on the time axis, and displaying the other audio identifier above or below the time axis, so that the two audio identifiers do not overlap.
[0051] In a third aspect, embodiments of the present application provide an apparatus, included in a terminal device, that implements the terminal device behavior described in the first aspect or any of the possible implementations of the first aspect. This functionality can be implemented in hardware or by hardware executing corresponding software. The hardware or software includes at least one module or unit corresponding to the aforementioned functionality. For example, a communication module or unit, a processing module or unit, etc.
[0052] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium comprising computer instructions. When the computer instructions are executed on a terminal device, the terminal device executes the method described in the first aspect or any possible implementation of the first aspect.
[0053] In a fifth aspect, an embodiment of the present application provides a computer program product, which, when executed on a computer, enables the computer to execute the method described in the first aspect and any possible implementation thereof.
[0054] In the sixth aspect, an embodiment of the present application provides a meeting recording system, which may include a server and multiple clients, the server establishes communication connections with the multiple clients respectively, and the multiple terminal devices are each used to execute the method described in the first aspect or any possible implementation method of the first aspect.
[0055] Among them, the terminal equipment, device, computer storage medium, computer program product, or conference recording system provided in the embodiments of the present application are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods provided above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 A schematic diagram of the architecture of an online conference system provided in an embodiment of the present application;
[0057] Figure 2 A schematic diagram of a four-person (N=4) online conference system provided in an embodiment of the present application;
[0058] Figure 3 A schematic diagram of the structure of a terminal device 300 (e.g., a mobile phone) provided in an embodiment of the present application;
[0059] Figure 4 A schematic diagram of the structure of the server 400 provided in an embodiment of the present application;
[0060] Figure 5 A flowchart of a conference recording method provided in an embodiment of the present application;
[0061] Figure 6AA schematic diagram of a user interface for an online meeting provided in an embodiment of the present application;
[0062] Figure 6B A schematic diagram of a user interface for an online meeting provided in an embodiment of the present application;
[0063] Figure 7 A schematic diagram of a user interface for an online meeting provided in an embodiment of the present application;
[0064] Figure 8 A schematic diagram of a storage method for a conference recording method provided in an embodiment of the present application;
[0065] Figure 9 A flowchart of a conference recording method provided in an embodiment of the present application;
[0066] Figure 10 A schematic diagram of a user interface for an online meeting provided in an embodiment of the present application;
[0067] Figure 11 A schematic diagram of a user interface for an online meeting provided in an embodiment of the present application;
[0068] Figure 12 A flowchart of a conference recording method provided in an embodiment of the present application;
[0069] Figure 13 A schematic diagram of a user interface for an online meeting provided in an embodiment of the present application;
[0070] Figure 14 A flowchart of a conference recording method provided in an embodiment of the present application;
[0071] Figure 15 A schematic diagram of a user interface for an online meeting provided in an embodiment of the present application;
[0072] Figure 16 A flowchart of a conference recording method provided in an embodiment of the present application;
[0073] Figure 17 A schematic diagram of a user interface for an online meeting provided in an embodiment of the present application;
[0074] Figure 18 A schematic diagram of a user interface for an online meeting provided in an embodiment of the present application;
[0075] Figure 19 A schematic diagram of a user interface for an online meeting provided in an embodiment of the present application;
[0076] Figure 20 A schematic diagram of the composition of a conference recording device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0077] The following describes in detail a conference recording method, terminal device, and conference recording system provided in an embodiment of the present application in conjunction with the accompanying drawings.
[0078] The terms "first" and "second" and the like in the specification and drawings of this application are used to distinguish different objects, or to distinguish different processing of the same object, rather than to describe a specific order of objects.
[0079] Furthermore, the terms "including," "having," and any variations thereof, as used in the description of this application are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or units is not limited to the listed steps or units, but may optionally include other steps or units not listed, or may optionally include other steps or units inherent to the process, method, product, or apparatus.
[0080] It should be noted that in the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be interpreted as being more preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0081] In this application, unless otherwise specified, "plurality" means two or more. "And / or" in this document is simply a description of an association between related objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone.
[0082] Figure 1 This is a schematic diagram of the architecture of an online conference system provided in an embodiment of the present application. The online conference system may include a server and multiple clients. The server may communicate with multiple clients. Multiple clients may establish communication connections through the server. Multiple clients may be Figure 1 , client 1, client 2, ..., client N-1, and client N, etc. N is any integer greater than 1.
[0083] Through the server, client 1, client 2, ..., client N-1 and client N, an online conference function can be provided to N users (hereinafter referred to as participants).
[0084] The present embodiment of the application uses an example in which one client corresponds to one participant, i.e., one participant uses one client to access the online meeting. Participants do not need to use different clients to access the online meeting. Audio or video data is synchronized between the various clients accessing the online meeting. Of course, it is understood that one client can also correspond to multiple participants. For example, multiple participants in a group can use one client to access the online meeting. The present embodiment of the application is not limited to this.
[0085] Synchronizing audio or video data between clients connected to an online meeting means that the client used by a participant acting as a speaker can synchronize the speaker's voice or video data with other clients connected to the online meeting. Participants acting as speakers can switch roles during the meeting.
[0086] The aforementioned clients 1, 2, ..., N-1, and N can be distributed in different physical locations. For example, client 1 is located in office 1 in city 1, client 2 is located in office 2 in city 1, and client N-1 is located in office 3 in city 3, etc. This embodiment of the application does not provide a detailed description of these examples.
[0087] Client 1, Client 2, ..., Client N-1, and Client N can each serve as a user interface for an online conference. Participants can access the online conference by activating their own client, i.e., establishing a communication connection with the server and invoking the online conference service provided by the server.
[0088] The server can provide online conference services to client 1, client 2, ..., client N-1, and client N. Taking client 1 as an example, the server can use client 1 to call client 1's microphone / display to record audio / video content, create a timeline for each participant, so that the audio / video content of different participants is linked to the participant's timeline, and provide an online conference review interface after the meeting for participants to review the meeting scene.
[0089] In the embodiment of the present application, any client can be hardware (such as a terminal device) or software (such as an APP). For example, if the client is a terminal device, the terminal device (also referred to as a user equipment (UE)) is a device with wireless transceiver function, which can be deployed on land, including indoors or outdoors, handheld or vehicle-mounted; it can also be deployed on the water surface (such as a ship, etc.); it can also be deployed in the air (for example, on an airplane, a balloon, and a satellite, etc.). The terminal can be a mobile phone, a tablet computer (pad), a computer with wireless transceiver function, a virtual reality (VR) terminal, an augmented reality (AR) terminal, a wireless terminal in industrial control, a wireless terminal in self-driving, a wireless terminal in remote medical, a wireless terminal in smart grid, a wireless terminal in transportation safety, a wireless terminal in a smart city, a wireless terminal in a smart home, and a wireless terminal in the Internet of Things (IoT), etc. If the client is an APP, the APP can be deployed in any of the above-mentioned terminal devices. For example, the APP can be deployed in a mobile phone, tablet, personal computer (PC), smart bracelet, audio, television, smart watch, or other terminal devices. The embodiments of this application do not specifically limit the specific form of the terminal device.
[0090] The above server can be one or more physical servers ( Figure 1 It can also be a computer cluster, a virtual machine in a cloud computing scenario, or a cloud server, etc.
[0091] It should be noted that in the embodiment of the present application, the above-mentioned APP can be an application built into the terminal device itself, or it can be an application provided by a third-party service provider and installed by the user. There is no specific limitation on this.
[0092] The following embodiments are described using an APP that provides an online conference function as an example.
[0093] Take N=4 as an example, Figure 2 A schematic diagram of a four-person (N=4) online conference system provided in an embodiment of the present application is shown in FIG. Figure 2As shown, the four-person online conference system includes four clients (211, 212, 213, and 214) and server 220. The four clients (211, 212, 213, and 214) can be deployed on different terminal devices. The participant of client 211 (i.e., the user who joins the online conference using client 211) is participant 001, the participant of client 212 is participant 002, the participant of client 213 is participant 003, and the participant of client 214 is participant 004.
[0094] When participant 001 is the speaker, client 211 can use the microphone of its own terminal device to collect participant 001's speech voice signal, obtaining a segment of audio data from participant 001. Client 211 then transmits participant 001's audio data to clients 212, 213, and 214 via server 220. Clients 212, 213, and 214 can play the audio data, allowing participants 002, 003, and 004 to receive participant 001's speech voice signal.
[0095] In some embodiments, client 211 may also schedule the display screen of its own terminal device to capture the display content of participant 001's display screen, thereby obtaining a segment of video data (also referred to as screen recording data) of participant 001. Client 211 transmits the segment of video data of participant 001 to client 212, client 213, and client 214 via server 220. Client 212, client 213, and client 214 may play the video data, so that participant 002, participant 003, and participant 004 can receive the display content of participant 001's display screen.
[0096] The server of the embodiment of the present application can assign a timeline to each participant, so that the audio / video data of different participants are linked to their respective timelines, and provide an online meeting review interface after the meeting for participants to review the meeting scene. Figure 2Server 220 in the example can assign a timeline to each participant 001, participant 002, participant 003, and participant 004. For the aforementioned audio data segment of participant 001, server 220 can save the audio data, its start time, and end time, and associate the audio data, its start time, and end time with the participant 001 timeline to generate an online meeting review interface after the online meeting. This online meeting review interface allows participant 001, participant 002, participant 003, or participant 004 to independently replay part or all of the meeting scene after the online meeting. Participant 001, participant 002, participant 003, or participant 004 can play the participant's conference speech voice signal by operating on this online meeting review interface.
[0097] Similarly, for the aforementioned video data of participant 001, server 220 may also save the video data, the start time and the end time of the video data, and associate the video data and the start time and the end time of the video data with the timeline of participant 001, so as to generate an online meeting review interface after the online meeting. The online meeting review interface allows participant 001, participant 002, participant 003, or participant 004 to independently reproduce part or all of the meeting scene after the online meeting. Participant 001, participant 002, participant 003, or participant 004 can play the content displayed on the participant's display screen by operating on the online meeting review interface.
[0098] It should be noted that the participant's timeline can be implemented by the participant's identification information and time information. The participant's identification information can be the participant's identity document (ID), the participant's terminal device identifier, or the participant's mobile phone number, etc., and this embodiment of the application does not specifically limit this. The time information can be time information in any time zone, such as Beijing time.
[0099] Therefore, unlike the method of recording the entire conversation of the meeting in text and then searching for information by keywords, the meeting recording method of the embodiment of the present application can save the audio / video data of the participants from the participant dimension and time dimension. In this way, after the online meeting is over, the meeting record information of the participant dimension and time dimension can be provided to the participants through the online meeting review interface. The participants can filter the required information by the participants and / or time nodes and independently reproduce all or part of the meeting scene. The meeting recording method of the embodiment of the present application can provide a meeting recording function that can quickly and effectively locate the required target information, which is beneficial to improving the efficiency of finding and locating target information in the meeting record. Its specific implementation method can be found in the explanation of the following embodiment.
[0100] Figure 3 This is a schematic diagram of the structure of a terminal device 300 (such as a mobile phone) provided in an embodiment of the present application. It should be understood that Figure 3 The structure shown does not constitute a specific limitation on the terminal device 300. In other embodiments of the present application, the terminal device 300 may include Figure 3 The structures shown may have more or fewer components, or some components may be combined or separated, or the components may be arranged differently. Figure 3 The various components shown in the drawings may be implemented in hardware, software, or a combination of hardware and software, including one or more signal processing and / or application specific integrated circuits.
[0101] Exemplarily, the terminal device 300 may include a processor 310, an external memory interface 320, an internal memory 321, a universal serial bus (USB) interface 330, a charging management module 340, a power management module 341, a battery 342, an antenna 1, an antenna 2, a mobile communication module 350, a wireless communication module 360, an audio module 370, a speaker 370A, a receiver 370B, a microphone 370C, an earphone interface 370D, a sensor 380, a button 390, a motor 391, an indicator 392, a camera 393, a display 394, and a subscriber identification module (SIM) card interface 395. It will be understood that the structure illustrated in this embodiment does not constitute a specific limitation on the terminal device 300. In other embodiments of the present application, the terminal device 300 may include more or fewer components than shown, or combine certain components, or split certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0102] The processor 310 may include one or more processing units. For example, the processor 310 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). The different processing units may be independent devices or integrated into one or more processors. In some embodiments, the terminal device 300 may also include one or more processors 310. The controller may be the nerve center and command center of the terminal device 300. The controller may generate operation control signals based on instruction opcodes and timing signals to control instruction fetching and execution. The processor 310 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 310 is a high-speed cache memory. This memory may store instructions or data that have just been used or are being recycled by the processor 310. If the processor 310 needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 310, and thus improves the efficiency of the terminal device 300 system.
[0103] In some embodiments, the processor 310 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface. The USB interface 330 is an interface that complies with USB standards and may be a Mini USB interface, a Micro USB interface, a USB Type-C interface, or the like. The USB interface 130 may be used to connect a charger to charge the terminal device 300, or to transmit data between the terminal device 300 and peripheral devices. It may also be used to connect headphones to play audio through the headphones.
[0104] It is understood that the interface connection relationship between the modules illustrated in the embodiment of the present invention is merely an illustrative illustration and does not constitute a structural limitation on the terminal device 300. In other embodiments of the present application, the terminal device 300 may also adopt a different interface connection method from the above embodiment, or a combination of multiple interface connection methods.
[0105] The charging management module 340 is configured to receive charging input from a charger. The charger can be either a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 340 can receive charging input from the wired charger via the USB interface 330. In some wireless charging embodiments, the charging management module 340 can receive wireless charging input via the wireless charging coil of the terminal device 300. While charging the battery 342, the charging management module 340 can also provide power to the terminal device 300 via the power management module 341.
[0106] The power management module 341 is used to connect the battery 342, the charging management module 340, and the processor 310. The power management module 341 receives input from the battery 342 and / or the charging management module 340 and provides power to the processor 310, the internal memory 321, the display 394, the camera 393, and the wireless communication module 360. The power management module 341 can also be used to monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage, impedance). In some other embodiments, the power management module 341 can also be set in the processor 310. In other embodiments, the power management module 341 and the charging management module 340 can also be set in the same device.
[0107] The wireless communication functionality of terminal device 300 can be implemented using antenna 1, antenna 2, mobile communication module 350, wireless communication module 360, a modem processor, and a baseband processor. Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in terminal device 300 can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In other embodiments, the antennas can be used in conjunction with a tuning switch.
[0108] The mobile communication module 350 can provide solutions for wireless communications including 2G / 3G / 4G / 5G, etc., applied to the terminal device 300. The mobile communication module 150 may include at least one filter, a switch, a power amplifier, a low-noise amplifier, etc. The mobile communication module 350 can receive electromagnetic waves from the antenna 1, filter, amplify, and process the received electromagnetic waves, and transmit them to the modem processor for demodulation. The mobile communication module 350 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation through the antenna 1. In some embodiments, at least some of the functional modules of the mobile communication module 350 can be set in the processor 310. In some embodiments, at least some of the functional modules of the mobile communication module 350 can be set in the same device as at least some of the modules of the processor 310.
[0109] The modem processor may include a modulator and a demodulator. The modulator is used to modulate the low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After being processed by the baseband processor, the low-frequency baseband signal is passed to the application processor. The application processor outputs a sound signal through an audio device (not limited to the speaker 370A, the receiver 370B, etc.) or displays an image or video through the display screen 394. In some embodiments, the modem processor may be an independent device. In other embodiments, the modem processor may be independent of the processor 310 and be provided in the same device as the mobile communication module 350 or other functional modules.
[0110] The wireless communication module 360 can provide wireless communication solutions including wireless local area networks (WLAN), Bluetooth, global navigation satellite system (GNSS), frequency modulation (FM), NFC, infrared technology (IR), etc. applied to the terminal device 300. The wireless communication module 360 can be one or more devices integrating at least one communication processing module. The wireless communication module 360 receives electromagnetic waves via the antenna 2, frequency modulates and filters the electromagnetic wave signals, and sends the processed signals to the processor 310. The wireless communication module 360 can also receive the signal to be transmitted from the processor 310, frequency modulate it, amplify it, and convert it into electromagnetic waves for radiation through the antenna 2.
[0111] In some embodiments, the antenna 1 of the terminal device 300 is coupled to the mobile communication module 150, and the antenna 2 is coupled to the wireless communication module 360, so that the terminal device 300 can communicate with the network and other devices through wireless communication technology. The wireless communication technology may include GSM, GPRS, CDMA, WCDMA, TD-SCDMA, LTE, GNSS, WLAN, NFC, FM, and / or IR technology. The above-mentioned GNSS may include a global positioning system (GPS), a global navigation satellite system (GLONASS), a Beidou navigation satellite system (BDS), a quasi-zenith satellite system (QZSS) and / or a satellite-based augmentation system (SBAS).
[0112] The terminal device 300 implements display functionality through a GPU, display screen 394, and an application processor. The GPU is a microprocessor for image processing that connects the display screen 394 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 310 may include one or more GPUs that execute instructions to generate or modify display information.
[0113] Display screen 394 is used to display images, videos, etc. Display screen 394 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-oLed, or a quantum dot light-emitting diode (QLED). In some embodiments, terminal device 300 may include one or N display screens 394, where N is a positive integer greater than one.
[0114] The terminal device 300 can implement the shooting function through an ISP, one or more cameras 393, a video codec, a GPU, one or more display screens 394, and an application processor.
[0115] The NPU is a neural network (NN) computing processor. Drawing on the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it rapidly processes input information and can continuously self-learn. The NPU enables intelligent cognitive applications in the terminal device 300, such as image recognition, face recognition, speech recognition, and text comprehension.
[0116] The external memory interface 320 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the terminal device 300. The external memory card communicates with the processor 310 via the external memory interface 320 to implement data storage functions. For example, data files such as music, photos, and videos can be stored on the external memory card.
[0117] The internal memory 321 can be used to store one or more computer programs, which include instructions. The processor 310 can execute the above instructions stored in the internal memory 321, so that the terminal device 300 performs the meeting recording method provided in some embodiments of the present application, as well as various functional applications and data processing. The internal memory 321 may include a program storage area and a data storage area. The program storage area may store an operating system; the program storage area may also store one or more applications (such as a gallery, contacts, etc.). The data storage area may store data created during the use of the terminal device 300 (such as photos, contacts, etc.). In addition, the internal memory 321 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc. In some embodiments, the processor 310 can execute the instructions stored in the internal memory 321 and / or the instructions stored in the memory provided in the processor 310, so that the terminal device 300 performs the meeting recording method provided in the embodiments of the present application, as well as various functional applications and data processing.
[0118] The terminal device 300 can implement audio functions such as music playback and recording through an audio module 370, a speaker 370A, a receiver 370B, a microphone 370C, a headphone jack 370D, and an application processor. The audio module 370 is used to convert digital audio information into analog audio signal outputs, and also to convert analog audio inputs into digital audio signals. The audio module 370 can also be used to encode and decode audio signals. In some embodiments, the audio module 370 can be provided within the processor 310, or some functional modules of the audio module 370 can be provided within the processor 310. The speaker 370A, also known as a "speaker," is used to convert audio electrical signals into sound signals. The terminal device 300 can listen to music or make hands-free calls through the speaker 370A. The receiver 370B, also known as a "handset," is used to convert audio electrical signals into sound signals. When the terminal device 300 receives a call or voice message, the voice can be heard by holding the receiver 170B close to the ear. Microphone 370C, also known as "microphone" or "microphone", is used to convert sound signals into electrical signals. When making a call or sending a voice message, the user can speak by putting their mouth close to the microphone 370C to input the sound signal into the microphone 370C. The terminal device 300 can be provided with at least one microphone 370C. In other embodiments, the terminal device 300 can be provided with two microphones 370C, which can not only collect sound signals but also realize noise reduction function. In other embodiments, the terminal device 300 can also be provided with three, four or more microphones 370C to realize sound signal collection, noise reduction, sound source identification, directional recording function, etc. The headphone jack 370D is used to connect wired headphones. The headphone jack 370D can be a USB interface 330, or a 3.5mm open mobile terminal platform (OMTP) standard interface, or a cellular telecommunications industry association of the USA (CTIA) standard interface.
[0119] The sensor 380 may include a pressure sensor, a gyro sensor, an air pressure sensor, a magnetic sensor, an acceleration sensor, a distance sensor, a proximity light sensor, a fingerprint sensor, a temperature sensor, a touch sensor, an ambient light sensor, a bone conduction sensor, and the like.
[0120] The buttons 390 include a power button, a volume button, etc. The buttons 390 may be mechanical buttons or touch buttons. The terminal device 300 may receive button inputs and generate key signal inputs related to user settings and function control of the terminal device 300.
[0121] The SIM card interface 395 is used to connect a SIM card. The SIM card can be connected to and disconnected from the terminal device 300 by inserting or removing it from the SIM card interface 395. The terminal device 300 can support one or N SIM card interfaces, where N is a positive integer greater than one. The SIM card interface 395 can support Nano SIM cards, Micro SIM cards, SIM cards, and the like. Multiple cards can be inserted into the same SIM card interface 395 simultaneously. The multiple cards can be of the same or different types. The SIM card interface 395 can also be compatible with different types of SIM cards. The SIM card interface 395 can also be compatible with external memory cards. The terminal device 300 interacts with the network through the SIM card to implement functions such as calls and data communications. In some embodiments, the terminal device 300 uses an eSIM, or embedded SIM card. The eSIM card can be embedded in the terminal device 300 and cannot be separated from the terminal device 300.
[0122] Figure 4 A structural diagram of the server 400 provided in the embodiment of the present application is shown as follows: Figure 4 As shown, the server 400 may be Figure 1 The server 400 in the illustrated embodiment includes a processor 401 , a memory 402 (one or more computer-readable storage media), and a communication interface 403 . These components can communicate with each other via one or more buses 404 .
[0123] The processor 401 may be one or more CPUs. In the case where the processor 401 is a CPU, the CPU may be a single-core CPU or a multi-core CPU.
[0124] The memory 402 can be connected to the processor 401 via a bus 404 or can be coupled to the processor 401 to store various program codes and / or multiple sets of instructions, as well as data (e.g., audio data, video data, etc.). In a specific implementation, the memory 402 includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), or portable compact disc read-only memory (CD-ROM).
[0125] The communication interface 403 is used to communicate with other devices, for example, to receive data (eg, requests, audio data, video data, etc.) sent by a terminal device, and to send data (eg, audio data, video data, etc.) to the terminal device.
[0126] It should be understood that Figure 4 The server 400 shown is only an example provided in the embodiment of the present application. The server 400 may also have more components than shown in the figure, and the embodiment of the present application does not make specific limitations on this.
[0127] In the embodiment of the present application, the processor 401 executes various functional applications and data processing of the server 400 by running the program code stored in the memory 402.
[0128] Figure 5 This is a flowchart of a conference recording method provided in an embodiment of the present application. The conference recording method can be applied to a client that provides an online conference function, where the online conference includes an online audio conference or an online video conference. In other words, the execution subject of this embodiment can be the above-mentioned Figure 1 Any client in . Figure 5 As shown, the method of this embodiment may include:
[0129] Step 501: Detect a triggering operation of an online conference ending option or a triggering operation of an online conference record viewing option.
[0130] Taking the client as an APP on the terminal device as an example, participants can use the client to access the online meeting, for example Figure 2 Participant 001 shown can use client 211 to access the online conference. Other participants can use their own clients, for example, participant 002 can use client 212 to access the online conference, participant 003 can use client 213 to access the online conference, and so on. A communication connection is established between different clients accessing the online conference through the server. Audio data and / or video data can be transmitted between different clients through the communication connection. For example, when participant 001 is speaking, client 211 can call the microphone of the terminal device to collect the audio data of participant 001, and then transmit the audio data to other clients through the communication connection, for example, Figure 2The client 212, client 213 and client 214 shown, client 212, client 213 and client 214 can each play the audio data so that participant 002, participant 003 and participant 004 receive the voice signal of participant 001. When other participants speak, the transmission method of voice data is similar to the above method, which will not be repeated here. During the online meeting, different participants can each speak as a speaker in different time periods or the same time period. The client used by the speaker can call the microphone of the terminal device to collect the speaker's audio data and transmit it to the client used by other participants. The embodiment of the present application can also save the audio data of different participants in different time periods or the same time period, as well as the start time and end time of different time periods or the same time period, so as to display the online meeting review interface in step 502 below after the online meeting ends.
[0131] In one achievable manner, the triggering method for displaying the online meeting review interface as in step 502 below may be the detection of a triggering operation of an online meeting end option. The triggering operation of the online meeting end option may be a participant's finger, a stylus, or other control object that can be detected by the touch display screen of the terminal device, acting as a triggering operation on the online meeting end option. The online meeting end option may be a control displayed on the user interface for exiting the online meeting. It should be noted that the triggering operation of the online meeting end option may also be a participant's use of other control objects connected to the terminal device to trigger the online meeting end option, such as a mouse, keyboard, etc., which will not be illustrated one by one in the embodiments of the present application.
[0132] For example, Figure 6A This is a schematic diagram of a user interface for an online meeting provided in an embodiment of the present application. As an example and not a limitation, Figure 6A As shown, the display component of the terminal device displays a possible current user interface of the client providing the online conference function. The user interface is a main interface 601 during the online conference. The main interface 601 can display the following: Figure 6A The main interface 601 may specifically include a mute or unmute control 6011, a video start or stop control 6012, a screen share control 6013, a member control 6014, a more control 6015, and an end meeting control 6016. It should be understood that the main interface 601 may also include other more or less display content, which is not limited in this embodiment of the present application. Among them, the end meeting control 6016 can be used as the above-mentioned online meeting end option. Participants perform the following Figure 6A As shown, a click operation is performed on the end meeting control 6016. In response to the click operation, the display component of the terminal device displays the following online meeting review interface.
[0133] In another implementation, the triggering method for displaying the online meeting review interface in step 502 below may be detecting a triggering operation of an online meeting record viewing option. The triggering operation of the online meeting record viewing option may be a participant's finger, a stylus, or other control object detectable by the touch screen of the terminal device, acting to trigger the online meeting record viewing option. The online meeting record viewing option may be a control displayed on the user interface for displaying the online meeting review interface.
[0134] For example, Figure 6B This is a schematic diagram of a user interface for an online meeting provided in an embodiment of the present application. As an example and not a limitation, Figure 6B As shown, the display component of the terminal device displays a possible current user interface of the client providing the online conference function. The user interface is a main interface 602 after the online conference ends. The main interface 602 can display the following: Figure 6B The main interface 602 may specifically include a participant control 6021 and a meeting record control 6022. It should be understood that the main interface 602 may also include other more or less display content, and this embodiment of the application does not limit this. Among them, the meeting record control 6022 can be used as the above-mentioned online meeting record viewing option. Participants perform the following Figure 6B As shown, a click operation is performed on the meeting record control 6022. In response to the click operation, the display component of the terminal device displays the following online meeting review interface.
[0135] Step 502: In response to the triggering operation of the online meeting end option or the triggering operation of the online meeting record viewing option, an online meeting review interface is displayed, where the online meeting review interface includes a timeline of multiple participants and multiple audio identifiers distributed along the timeline.
[0136] Each of the multiple audio identifiers is used to identify a segment of audio data for a participant in the timeline on which the audio identifier is located. The start position of each audio identifier on the timeline represents the start time of the audio data segment for that participant, and the end position of each audio identifier on the timeline represents the end time of the audio data segment for that participant. The audio data segment for that participant is generated by recording the voice signal of the participant speaking in the conference during the period between the start time and the end time.
[0137] A participant's timeline can have one or more audio tags. The number of audio tags and the length of each audio tag on a participant's timeline are related to the participant's speaking time in the online meeting.
[0138] For example, taking an online meeting with five participants as an example, the online meeting review interface may include the timelines of the five participants and multiple audio identifiers distributed along the timelines. The timelines of the five participants are respectively the timeline of participant 1, the timeline of participant 2, the timeline of participant 3, the timeline of participant 4, and the timeline of participant 5. The lengths of the five participants' timelines may be the same or different. In one example, the lengths of the five participants' timelines are the same, the starting positions of the five participants' timelines are the time when the online meeting starts, and the ending positions of the five participants' timelines are the time when the online meeting ends. The time when the online meeting starts may be the scheduled start time of a scheduled online meeting, or the initiation time of a temporary online meeting, or the access time of the first client among the five clients to access the online meeting, etc. The time when the online meeting ends may be the scheduled end time of a scheduled online meeting, or the exit time of the last client among the five clients to exit the online meeting, etc.
[0139] For example, Figure 7 This is a schematic diagram of a user interface for an online meeting provided in an embodiment of the present application. As an example and not a limitation, Figure 7 As shown, the display component of the terminal device displays a possible current user interface of the client providing the online conference function. The user interface is a main interface 701 after the online conference ends. The main interface 701 can display the following Figure 7The content shown. The main interface 701 may specifically include the timelines of five participants and multiple audio identifiers distributed along the timelines. The timelines of the five participants are respectively the timeline 7011 of participant 1, the timeline 7012 of participant 2, the timeline 7013 of participant 3, the timeline 7014 of participant 4 and the timeline 7015 of participant 5. The lengths of the timeline 7011 of participant 1, the timeline 7012 of participant 2, the timeline 7013 of participant 3, the timeline 7014 of participant 4 and the timeline 7015 of participant 5 are the same. The starting position of the timelines of the five participants is the time when the online meeting starts, for example, 14:00 Beijing time. The starting position of the timelines of the five participants is the time when the online meeting ends, for example, 18:00 Beijing time. After the online meeting starts, participant 1 speaks first and obtains a segment of audio data of participant 1, the audio identifier of which is A05. The starting position of A05 on the timeline of participant 1 is the starting time of this segment of audio data of participant 1, that is, the starting time of participant 1's speech. The ending position of A05 on the timeline of participant 1 is the ending time of this segment of audio data of participant 1, that is, the ending time of participant 1's speech. Afterwards, participant 2 speaks, and a segment of audio data of participant 2 is obtained, and the audio identifier of this segment of audio data is A08. The starting position of A08 on the timeline of participant 2 is the starting time of this segment of audio data of participant 2, that is, the starting time of participant 2's speech. The ending position of A08 on the timeline of participant 2 is the ending time of this segment of audio data of participant 2, that is, the ending time of participant 2's speech. Similarly, different participants speak, and one or more segments of audio data of different participants are obtained. Each segment of audio data corresponds to an audio identifier, and the length of each audio identifier can be the length of a participant's speech, for example, 30 minutes. In this way, after the online meeting, the following can be generated Figure 7 The online conference review interface. On the timeline 7011 of participant 1 of the online conference review interface, there are 4 audio identifiers (A05, A06, A07 and A03), on the timeline 7012 of participant 2, there are 3 audio identifiers (A08, A09 and A04), on the timeline 7013 of participant 3, there are 2 audio identifiers (A01 and A02), on the timeline 7014 of participant 4, there is 1 audio identifier (A10), and on the timeline of participant 5, there are 2 audio identifiers (A11 and A12). The online conference review interface allows users to operate on the online conference review interface and reproduce part or all of the conference scene in audio form. It should be understood that the main interface 602 may also include other more or less display content, which is not limited in the embodiment of the present application. As an example, Figure 7The audio identifier is displayed in the form of a rectangular icon. It is understandable that it can also be in other icon forms, and the present embodiment does not give examples one by one. In addition, the present embodiment does not specifically limit the height of the rectangular icon.
[0140] Optionally, the online meeting review interface may also include the total speaking time of each participant in the entire meeting. Figure 7 As shown, the rightmost side of each participant's timeline displays the total duration of the corresponding participant's speech. Taking participant 1 as an example, the total duration of participant 1's speech is 2 hours, 12 minutes, and 21 seconds (2:12:21).
[0141] In this embodiment, by detecting the trigger operation of the online meeting end option or the trigger operation of the online meeting record viewing option, in response to the trigger operation of the online meeting end option or the trigger operation of the online meeting record viewing option, an online meeting review interface is displayed. The online meeting review interface includes a timeline of multiple participants and a plurality of audio identifiers distributed along the timeline. Each of the multiple audio identifiers is used to identify a segment of audio data of the participant on the timeline where the audio identifier is located. The starting position of each audio identifier on the timeline where it is located is the starting time of a segment of audio data of a participant, and the ending position of each audio identifier on the timeline where it is located is the ending time of a segment of audio data of a participant. The segment of audio data of the participant is generated by recording the voice signal of the conference speech of the participant in the time period between the starting time and the ending time. The online meeting review interface can be used by the user to operate on the online meeting review interface after the online meeting ends, and reproduce part or all of the meeting scene in audio form. After the online meeting, participants can access the online meeting review interface, which provides them with meeting records by participant and time. Participants can filter required information by participant and / or time point, and independently replay all or part of the meeting scene. This online meeting review interface allows for quick and efficient locating of desired meeting records, improving the efficiency of finding and locating target information within meeting records.
[0142] For example, if a participant needs to find the specific details of the preparatory work assigned by Participant 1 at around 3:00 PM after an online meeting, recording the entire meeting conversation in text and then searching for information using keywords makes it difficult to locate the information from the text. For example, after the online meeting, the participant enters the keyword "preparatory work." However, because Participant 1 spoke "preparatory work" during the online meeting, the text recorded after the speech is converted to text. Therefore, the keyword "preparatory work" entered by the participant after the online meeting cannot be matched with the "preparatory work" recorded in the text. Furthermore, even if audio data of the entire meeting conversation is stored, since the audio data and the text are independent, if the participant cannot locate the information they are looking for using keywords, they often need to play the audio data to find the information they need. However, they cannot accurately locate the audio data containing the required information, making it difficult for the participant to quickly find the information they need. Unlike this method, the conference record method of the embodiment of the present application can provide an online conference review interface. After the online conference, the participants can click on the audio identifier A06 on the online conference review interface according to the participant (i.e., participant 1 here) and the time node (15:00), thereby playing the audio data corresponding to the audio identifier A06, reproducing the speech of participant 1 in audio form, and thus knowing the specific content of the preparatory work assigned by participant 1 around 15:00. This can more simply, intuitively, and quickly locate the required information, which is beneficial to improving the efficiency of finding and locating target information in the conference record. Using audio data instead of text can, on the one hand, solve the problems of being unable to find or difficult to understand caused by the inconsistency between oral expression and text expression. On the other hand, the listening method can be understood by the brain faster than the reading method, and can compare and filter information faster.
[0143] The audio data segment identified by each of the above multiple audio identifiers may be recorded and generated by the client or by the service.
[0144] Method 1: The client records and generates one or more audio data. Specifically, the client can record and generate one or more audio data of the participants using the client, and record the start time and end time of each of the one or more audio data. The client sends one or more audio data and the start time and end time of each of the one or more audio data to the server, and the server stores the one or more audio data and the start time and end time of each of the one or more audio data to one or more storage units, so that the server can provide the above-mentioned online meeting review interface to the client. Each storage unit includes a segment of audio data, and the start time and end time of the audio data. The one or more storage units here are connected in series by using a set of time pointers of the participants of the client. The audio data in each storage unit can be an audio file.
[0145] As Figure 7 Taking the client of participant 3 as an example, the client of participant 3 can record and generate two audio data of participant 3, and record the start time and end time of each of the two audio data, and send them to the server. The server stores the two audio data and the start time and end time of each of the two audio data in two storage units, and obtains the following based on this. Figure 7 The online conference review interface shown shows two audio identifiers (A01 and A02) distributed along the timeline of participant 3.
[0146] For example, the two storage units of participant 3 may be as follows Figure 8 The structure shown, Figure 8 This is a schematic diagram of a storage method for a conference record method provided in an embodiment of the present application. As an example and not a limitation, Figure 8As shown, the two storage units of participant 3 are storage unit 1 and storage unit 2. Storage unit 1 includes a time pointer 11, a start time 1, audio data 1, an end time 1, and a time pointer 12. Time pointer 11 can point to a start time 1. Start time 1 can be the start time of participant 3's first speech. Audio data 1 includes the audio data of participant 3's first speech. End time 1 can be the end time of participant 3's first speech. Time pointer 12 can point to a time pointer 21 of storage unit 2. Storage unit 2 includes a time pointer 21, a start time 2, audio data 2, an end time 2, and a time pointer 22. Time pointer 21 can point to a start time 2. Start time 2 can be the start time of participant 3's second speech. Audio data 2 includes the audio data of participant 3's second speech. End time 2 can be the end time of participant 3's second speech. Thus, the two storage units can be connected in series through a set of time pointers for participant 3. This set of time pointers includes the aforementioned time pointer 11, time pointer 12, time pointer 21, and time pointer 22.
[0147] Method 2: The server records and generates multiple audio data segments. Specifically, the server can record and generate one or more audio data segments of participants using each client, record the start time and end time of each of the one or more audio data segments, and store the one or more audio data segments of participants using each client and the start time and end time of each of the one or more audio data segments in one or more storage units of each client, so that the server can provide the above-mentioned online meeting review interface to each client. The specific form of each storage unit can be the same as that of Method 1 above and will not be repeated here.
[0148] Figure 9 This is a flowchart of a conference recording method provided in an embodiment of the present application. The conference recording method can be applied to a system that provides an online conference function, where the online conference includes an online audio conference or an online video conference. Figure 9 As shown, this embodiment is illustrated by taking the system including five clients (i.e., client 1 of participant 1, client 2 of participant 2, client 3 of participant 3, client 4 of participant 4, and client 5 of participant 5) and one server as an example. The method of this embodiment may include:
[0149] Step 901: Client 1 accesses the online conference, client 2 accesses the online conference, client 3 accesses the online conference, client 4 accesses the online conference, and client 5 accesses the online conference.
[0150] Client 1, Client 2, Client 3, Client 4, and Client 5 each access the same online conference. The online conference can be identified by a conference ID. For example, Participant 1, Participant 2, Participant 3, Participant 4, and Participant 5 each access the same online conference at the same time or at different times within a time period by clicking a conference link or entering the conference ID.
[0151] In one possible implementation, the online meeting can be scheduled. For example, Figure 9 As shown by the dotted line, before step 901, client 1 may send an online meeting reservation request to the server. The online meeting reservation request may include the meeting topic, meeting start time, meeting end time and multiple participant information. The meeting start time is later than the time when the online meeting reservation request is sent. Taking the application scenario of this embodiment as an example, the multiple participant information may include the identification information of participant 2, the identification information of participant 3, the identification information of participant 4 and the identification information of participant 5. The server allocates an online meeting ID and / or an online meeting link based on the online meeting reservation request. The online meeting ID and / or online meeting link are used for the clients used by different participants to access the online meeting. The server sends an online meeting reservation response to client 1. The online meeting reservation response may include the online meeting ID and / or the online meeting link. Client 1 may send the online meeting ID and / or the online meeting link to client 2, client 3, client 4 and client 5. In addition, the server may also store the aforementioned meeting topic, meeting start time, meeting end time, identification information of participant 1, identification information of participant 2, identification information of participant 3, identification information of participant 4, and identification information of participant 5. Of course, it is understandable that the online meeting reservation request may also include other information such as the meeting agenda, and this embodiment of the present application does not specifically limit this.
[0152] In another possible implementation, the online meeting can also be temporary, that is, not scheduled in advance. For example, client 1 can temporarily initiate an online meeting, that is, the online meeting has already started at the same time as it is initiated. After client 1 accesses the online meeting, it can send an online meeting ID and / or an online meeting link to client 2, client 3, client 4 and client 5 respectively, and client 2, client 3, client 4 and client 5 can access the online meeting according to the online meeting ID or the online meeting link respectively. In addition, the server can also store the meeting start time (that is, the time when client 1 initiates the online meeting) and multiple participant information. Taking the application scenario of this embodiment as an example, the multiple participant information can include the identification information of participant 1, the identification information of participant 2, the identification information of participant 3, the identification information of participant 4 and the identification information of participant 5.
[0153] Step 902: The server allocates the timeline of participant 1, the timeline of participant 2, the timeline of participant 3, the timeline of participant 4, and the timeline of participant 5.
[0154] A detailed explanation of the timeline for each participant can be found in Figure 5 The detailed explanation of step 502 of the illustrated embodiment will not be repeated here.
[0155] One possible implementation method is that for a scheduled online meeting or a temporary online meeting, the server responds to client 1 accessing the online meeting by allocating a timeline for participant 1, the server responds to client 2 accessing the online meeting by allocating a timeline for participant 2, the server responds to client 3 accessing the online meeting by allocating a timeline for participant 3, the server responds to client 4 accessing the online meeting by allocating a timeline for participant 4, and the server responds to client 5 accessing the online meeting by allocating a timeline for participant 5.
[0156] Another possible implementation method is that for a scheduled online meeting, the server responds to the arrival of the start time of the scheduled online meeting or detects that any client (i.e., client 1, client 2, client 3, client 4 or client 5) has accessed the online meeting. The server allocates the timeline of participant 1, the timeline of participant 2, the timeline of participant 3, the timeline of participant 4 and the timeline of participant 5 based on the saved identification information of the participants of the online meeting.
[0157] It should be noted that the order of the above steps 901 and 902 is not limited by the size of the sequence number, and it can also be other orders. For example, when client 1 joins the online meeting, the server can allocate the timeline of participant 1 without waiting for other clients to join the online meeting.
[0158] The server can store the audio data of each participant along the timeline in subsequent online meetings.
[0159] Specifically, the server may establish a set of time pointers for each participant, and connect multiple storage units of each participant in series through the time pointer of each participant, so that the series-connected storage units point to the same participant. The series-connected storage units are used to store the audio data of the participant.
[0160] Step 904: The server records the voice audio of each client to generate multiple audio data segments, records the start time and end time of the voice audio of each client, and associates the multiple audio data segments to the timeline of the corresponding participant based on the start time and end time of the voice audio of each client.
[0161] During a meeting, the server records the audio data of each participant's speech and stores each segment in a separate storage unit. Each storage unit stores the start time of the speech, the audio data for that segment, and the end time of the speech. By connecting multiple storage units containing audio data via time pointers, one or more speeches by a participant during the meeting can be associated with that participant's timeline.
[0162] The specific form of the storage unit can be found in Figure 8 For detailed explanation, please refer to Figure 8 The explanation of the illustrated embodiment will not be repeated here.
[0163] It should be noted that one or more audio data segments can be stored in the following different ways:
[0164] 1) Request storage of audio data based on microphone permissions. For example, using a participant's microphone on a client, the microphone's on-time is used as the start time of a speech segment, and the microphone's end time is used as the end time of the speech segment. During the microphone's on-time, the voice audio picked up by the microphone is recorded to obtain a segment of audio data, which is then stored in a storage unit of the participant. This method is suitable for online meeting scenarios where participants turn on their microphones at the beginning of their speech and turn them off at the end.
[0165] 2) Storing audio data based on the microphone's uplink data stream. For example, when a participant's microphone detects that the voice audio picked up by the microphone is uploaded to the server, the upload start time is used as the start time of a speech segment, and the upload end time is used as the end time of the speech segment. The voice audio picked up by the microphone is recorded to obtain a segment of audio data, which is then stored in a storage unit on the participant.
[0166] In this way, audio data is stored in real time, and the server can record the online meeting and obtain multiple audio data segments in real time. Furthermore, there is no need for participants to manage their microphones; they can keep their microphones on throughout the meeting.
[0167] Optionally, if the time interval between the end time of the first speech and the start time of the second speech of a participant is less than a preset threshold, the two adjacent audio data segments can be stored in the same storage unit. The preset threshold can be flexibly set according to needs, for example, the preset threshold can be 10 seconds.
[0168] 3) Relink after the meeting. During the online meeting, it is sufficient to ensure that the microphone can record all the speeches of the participants. If the participants turn on the microphone throughout the meeting, all the audio data of the participants will be stored in a storage unit during the meeting. After the online meeting, all the audio data of each participant obtained during the meeting are obtained in turn. Then all the audio data of each participant are processed to obtain one or more audio data of the actual speech of the participant. For example, after removing the silent segments and performing noise reduction processing, one or more audio data of the actual speech of the participant can be obtained. By storing one or more audio data of the actual speech separately and relinking them with a time pointer, one or more audio data of the participant can be distributed according to time.
[0169] When relinking and storing all the audio data of each participant in segments, in addition to segmenting from the perspective of time, segmentation can also be performed from the perspective of semantic understanding. If two adjacent speeches belong to related contexts, the audio data of the two adjacent speeches can be treated as a continuous audio data segment. This is because the participant may be discussing the same topic with other participants. Therefore, the adjacent speeches that belong to the same topic in semantic understanding can be treated as a continuous audio data segment without removing the silent segments in the middle. The relevance of the same topic can be judged based on semantic understanding. When the relevance is higher than the set value, it can be determined to be the same topic. Semantic understanding and the judgment of the relevance of different paragraphs can directly use existing speech processing technology, and the embodiments of the present application will not be repeated here.
[0170] For example, suppose the following recordings are recorded from Participant A, Participant B, and Participant C during a meeting:
[0171] {
[0172] Participant A's timeline:
[0173] [2:00:30 PM - 2:02:30 PM] Person A: "...Now, please allow Person B to present the XX project."
[0174] Participant B's timeline:
[0175] [2:02:35-2:02:55] B: "Now, let me report on the progress of Project XX, a joint development project between our company and XX University."
[0176] [14:03:10-14:03:48] B: "Yes, Project XX is a key technical research project jointly undertaken by the company and Professor XX. It focuses on XX, an area where the company's technology is relatively weak. Professor XX is a top expert in this field in the industry."
[0177] [14:04:01~14:35:55] B: "Okay, then I'll continue with my presentation on Project XX. Feel free to interrupt if you have any questions. Last week, Project XX achieved a milestone in its experiments. The lab test results compared well with the company's existing XX series products... That's all about the progress of Project XX. Please feel free to ask any questions, judges."
[0178] Participant C's timeline:
[0179] [2:03:00-2:03:05 PM] C: "Excuse me, is this a collaborative project with Professor XX?"
[0180] [14:03:53~14:03:56] C: "Okay, please continue."
[0181] }
[0182] Based on semantic understanding, the three speeches of participant B are all highly relevant to the report of the XX project. Therefore, the three speech segments of participant B can be collected into the audio data of one storage unit, and the speech time of participant C in the middle is a silent segment.
[0183] Step 905: Client 1 detects the triggering operation of the online meeting end option, and sends online meeting end indication information to the server.
[0184] For details on how to trigger the online meeting end option, see Figure 5 The explanation of step 502 of the illustrated embodiment will not be repeated here.
[0185] After detecting the triggering operation of the online meeting end option, the client 1 may send online meeting end indication information to the server. The online meeting end indication information is used to instruct the server to generate an online meeting review interface.
[0186] Step 906: In response to the online meeting end indication information, the server generates an online meeting review interface and sends the online meeting review interface to client 1. The online meeting review interface includes a timeline of each participant and multiple audio identifiers distributed along the timeline.
[0187] The server generates an online conference review interface based on the multiple audio data segments distributed along the time axis of each participant, as well as the start time and end time of each audio data segment. For example, the server generates an online conference review interface based on the multiple storage units described above. The online conference review interface may include the time axis of each participant and the multiple audio identifiers distributed along the time axis. For example, the online conference review interface may be as follows: Figure 7 shown.
[0188] Step 907 : Client 1 displays an online meeting review interface in response to the triggering operation of the online meeting end option.
[0189] The online meeting review interface allows users to operate on the online meeting review interface and reproduce part or all of the meeting scene in audio form. The specific explanation of step 907 can be found in Figure 5 The explanation of step 502 of the illustrated embodiment will not be repeated here.
[0190] In this embodiment, during an online meeting, the voice audio of each client is recorded to generate multiple audio data segments, the start time and end time of the voice audio of each client are recorded, and the multiple audio data segments are associated with the corresponding participant's timeline based on the start time and end time of the voice audio of each client. After the online meeting is over, based on the recorded multiple audio data segments and the start time and end time of each audio data segment, an online meeting review interface is provided to the client. The online meeting review interface includes multiple participant timelines and multiple audio identifiers distributed along the timeline. Each of the multiple audio identifiers is used to identify a segment of audio data of the participant on the timeline where the audio identifier is located. The starting position of each audio identifier on the timeline where it is located is the start time of a segment of audio data of a participant, and the ending position of each audio identifier on the timeline where it is located is the end time of a segment of audio data of a participant. The online meeting review interface can be used by users to operate on the online meeting review interface after the online meeting is over, and reproduce part or all of the meeting scene in audio form. After the online meeting, participants can access the online meeting review interface, which provides them with meeting records by participant and time. Participants can filter required information by participant and / or time point, and independently replay all or part of the meeting scene. This online meeting review interface allows for quick and efficient locating of desired meeting records, improving the efficiency of finding and locating target information within meeting records.
[0191] In addition, by combining time nodes with audio data to filter target information, it supports users in searching for vague memories.
[0192] After providing an online meeting review interface using any of the above-described meeting recording methods, the entire or partial meeting scene can be presented by detecting any of the following user actions on the online meeting review interface. Detecting any of the following user actions on the online meeting review interface includes, but is not limited to, hovering, clicking, double-clicking, and dragging. Presenting the entire or partial meeting scene includes, but is not limited to, presenting keywords, playing audio data, playing video data, and the like.
[0193] Scenario 1: First trigger operation.
[0194] The client displays the online meeting review interface as described above, detects a first trigger operation of one of the multiple audio identifiers in the online meeting review interface, and in response to the first trigger operation, displays the keywords of the audio data corresponding to the audio identifier. The audio data corresponding to the audio identifier refers to the audio data identified by the audio identifier. The keywords of the audio data can be obtained by processing the audio data. For example, the audio data is processed by voice-to-text conversion, and then the keywords are identified based on semantic understanding.
[0195] For example, the first trigger operation may be a hover operation of a mouse or other control object. When the client detects that the mouse or other control object is hovering over an audio identifier of multiple audio identifiers, the keyword of the audio data corresponding to the audio identifier is displayed. Figure 7 The audio logo A06 shown in FIG, then the following is displayed above the audio logo Figure 7 The keyword "press conference preparation work" is shown. Of course, it is understandable that the keyword can also be located below the audio logo or other locations, and this embodiment of the application does not specifically limit this.
[0196] Scenario 2: Second trigger operation.
[0197] The client displays the online meeting review interface as described above, detects a second trigger operation of one of the multiple audio identifiers on the online meeting review interface, and plays audio data corresponding to the audio identifier in response to the second trigger operation.
[0198] For example, the second trigger operation can be a click / double-click operation. When the client detects a click / double-click on an audio identifier, it plays the audio data corresponding to the audio identifier. In this way, the user can review the meeting content from two dimensions: participants and time, and locate the meeting content. They can quickly locate the time node of the target information they need to find, and can know what each participant said at that time node, helping to recreate the meeting scene after the meeting.
[0199] Scenario 3: The third trigger operation.
[0200] The client displays the online meeting review interface as described above, and detects a third trigger operation of the first audio identifier from the multiple audio identifiers in the online meeting review interface from the first time axis to the second time axis. The first audio identifier is distributed on the first time axis, and one or more second audio identifiers are distributed on the second time axis. The first audio identifier and the at least one second audio identifier are different audio identifiers located on different time axes, and the audio data identified by the first audio identifier and the at least one second audio identifier are different. In response to the third trigger operation, the client displays the first audio identifier on the second time axis, and associates the first audio data corresponding to the first audio identifier with the second audio data corresponding to the at least one second audio identifier based on the start time and end time of the first audio data corresponding to the first audio identifier. That is, the first audio data is associated with the second time axis. Afterwards, the client detects the fourth trigger operation of the first audio identifier or the at least one second audio identifier, and in response to the fourth trigger operation, plays the first audio data and the second audio data corresponding to the at least one second audio identifier.
[0201] The position of the first audio identifier on the second time axis may be the same as the position of the first audio identifier on the first time axis.
[0202] For example, the third trigger operation can be a drag operation. When the client detects that the first audio identifier is dragged from the first timeline to the second timeline, the first audio identifier is displayed on the second timeline, and the first audio data corresponding to the first audio identifier is associated with the second audio data corresponding to at least one second audio identifier. Thus, by dragging the first audio identifier to the second timeline, the first audio data can be associated with the timeline of another participant (i.e., the second timeline here), thereby combining the speeches of different participants.
[0203] In one case, the time period of the first audio identifier overlaps with the time period of at least one second audio identifier, that is, the participant on the time axis where the first audio identifier is located and the participant on the time axis where the second audio identifier is located have at least one or more speeches in the same time period. The client detects the fourth trigger operation of the first audio identifier or the at least one second audio identifier, and in response to the fourth trigger operation, plays at least two of the non-intersecting part of the first audio data, the intersecting part of the first audio data and the at least one second audio data, or the non-intersecting part of at least one second audio data according to the time position of the first audio identifier on the second time axis and the time position of the at least one second audio identifier on the second time axis.
[0204] Therefore, taking the overlapping storage of a first audio identifier and a second audio identifier as an example, if the dragged first audio identifier and a second audio identifier on the second timeline overlap (also called an intersection), by superimposing and merging the first audio identifier and the second audio identifier according to time, the corresponding audio files are also associated, so that the associated audio files can be synchronously controlled and played with one click.
[0205] Optionally, audio files with different audio identifiers can be merged into one audio file based on time. Audio files of different participants are added to different audio tracks. The merged audio file corresponds to the merged audio identifier. You can click the merged audio identifier to play the merged audio file. The merged audio file includes the speeches of multiple participants.
[0206] In this way, if different participants make joint speeches over a period of time, you can listen to each participant's speech individually by clicking on the audio identifier on each participant's timeline. You can also drag the audio identifier to overlay and merge multiple audio identifiers into a single audio identifier, playing the speeches of multiple participants simultaneously, thus recreating the meeting discussion scene at that time. In particular, for content that was hotly debated in the meeting, it may be difficult to grasp everyone's point of view due to the intense discussion. After the meeting, the speeches of different participants can be split, merged, and repeatedly analyzed using audio identifiers to grasp the different participants' point of view.
[0207] Optionally, the client may further detect a fifth trigger operation of the first audio identifier, and in response to the fifth trigger operation, display the first audio identifier on the first timeline, and disassociate the first audio data and the second audio data, so that the first audio data and the second audio data are played independently.
[0208] The fifth trigger operation can be the reverse operation of the third trigger operation. For example, a drag operation in the opposite direction of the above-mentioned drag operation. It is understandable that the fifth trigger operation can also be other operations, such as double-clicking, etc., which is not specifically limited in the embodiments of the present application.
[0209] An example, Figure 10 This is a schematic diagram of a user interface for an online meeting provided in an embodiment of the present application. As an example and not a limitation, Figure 10 As shown, in Figure 7 The following steps are performed on the online meeting review interface shown in the figure: Figure 10The operation shown drags the audio identifier A03 on the timeline of participant 1 to the timeline of participant 3. Since the position of audio identifier A03 on the timeline of participant 1 overlaps with the position of audio identifier A02 on the timeline of participant 3. Therefore, the client can associate the audio files identified by audio identifier A03 and audio identifier A02. Afterwards, when the client detects that the user clicks on the associated audio identifier A03 or audio identifier A02, the audio files of audio identifier A03 and audio identifier A02 can be played synchronously. Similarly, audio identifier A04 on the timeline of participant 2 can be dragged to the timeline of participant 3. Since the position of audio identifier A04 on the timeline of participant 2 overlaps with the position of audio identifier A02 on the timeline of participant 3. Therefore, the client can associate the audio files identified by audio identifier A03, audio identifier A02, and audio identifier A04. Afterwards, when the client detects that the user clicks on the associated audio identifier A03, audio identifier A02, or audio identifier A04, the audio files of audio identifier A03, audio identifier A02, and audio identifier A04 can be played synchronously. In this way, the speeches of participant 1, participant 2, and participant 3 can be played.
[0210] Another example, Figure 11 This is a schematic diagram of a user interface for an online meeting provided in an embodiment of the present application. As an example and not a limitation, Figure 11 As shown, the online meeting review interface includes the timelines of participant A, participant B, and participant C. There is one audio identifier A01 on participant A's timeline, three audio identifiers (A02, A03, and A04) on participant B's timeline, and two audio identifiers (A05 and A06) on participant C's timeline.
[0211] Among them, the speech segment corresponding to audio identifier A02 is [14:02:35~14:02:55], and participant B said: "Now let me report on the progress of the XX project. The XX project is a joint development project between the company and XX University." The speech segment corresponding to audio identifier A03 is [14:03:10~14:03:48], and participant B said: "Yes, the XX project is a key technical research project jointly developed by the company and Professor XX. It mainly focuses on XX, which is the company's weak point in technology. Professor XX is a top expert and scholar in this field in the industry." The speech segment corresponding to audio identifier A04 is [14:04:01~14: [35:55], Participant B: "Okay, then I'll continue with my presentation on Project XX. Feel free to interrupt if you have any questions. Last week, Project XX achieved a milestone in its experiments. The lab test results compared favorably with the company's existing XX series products... That's all about the progress of Project XX. Judges, please ask any questions." The speech segment corresponding to audio identifier A05 is [14:03:00-14:03:05]. Participant C: "Excuse me, is this a collaborative project with Professor XX?" The speech segment corresponding to audio identifier A06 is [14:03:53-14:03:56]. Participant C: "Okay, please continue."
[0212] The client detects that Figure 11 The operation shown is to drag the audio identifier A05 and the audio identifier A06 of participant C to the timeline of participant B. According to the respective positions of the audio identifier A05 and the audio identifier A06 on the timeline of participant C, the moved audio identifier A05 and the audio identifier A06 fall at the corresponding time nodes on the timeline of participant B. The moved audio identifier A05 and the audio identifier A06 overlap with the audio identifier of participant B, and they can be merged into a new audio identifier, as shown in FIG. Figure 11 The audio marker on the timeline of participant B is shown as longer. After dragging both audio markers on participant C's timeline to participant B's timeline, the speech segments corresponding to the merged new audio markers are:
[0213] {
[0214] B Timeline:
[0215] [14:02:35~14:02:55] B: Let me report on the progress of the XX project, which is a joint development project between our company and XX University.
[0216] [14:03:00~14:03:05] C: Excuse me, is this a project you are collaborating on with Teacher XX?
[0217] [14:03:10~14:03:48] B: Yes, the XX project is a key technical research project jointly undertaken by the company and Professor XX. It mainly focuses on XX, which is where the company's technology is relatively weak. Professor XX is a top expert and scholar in this field in the industry.
[0218] [14:03:53~14:03:56] C: Okay, please continue.
[0219] [14:04:01-14:35:55] B: Okay, I'll continue with my presentation on Project XX. Feel free to interrupt if you have any questions. Last week, Project XX achieved a milestone in its experiments. The lab test results compared to the company's existing XX product line... That's all about Project XX's progress. Please feel free to ask any questions, judges.
[0220] }
[0221] Therefore, when it is detected that the user clicks on a new audio identifier on the timeline of participant B, the speech segment described above can be played.
[0222] Similarly, you can perform the reverse operation to split the merged audio marker and disassociate the audio files. For example, double-click the original audio marker or drag the audio marker back to its original position. After splitting, the audio files will be disassociated.
[0223] In this embodiment, by displaying the keywords of the audio data corresponding to the audio identifier on the online meeting review interface, the user can preliminarily filter the audio data based on the keywords, which is beneficial to improving the efficiency of finding and locating target information in the meeting records. By dragging the audio identifier, the audio data corresponding to multiple audio identifiers can be freely associated to play the speeches of multiple participants at the same time, thereby reproducing the meeting discussion scene at that time. For the content that was discussed intensely in the meeting, after the meeting, the audio data corresponding to the audio identifier can be associated or disassociated by fusing or splitting the audio identifiers on the online meeting review interface, so as to play the associated audio data at the same time or play the disassociated audio data independently, and repeatedly analyze it, which is helpful to grasp the viewpoints of different participants during the intense discussion stage after the meeting.
[0224] Figure 12 This is a flowchart of a conference recording method provided in an embodiment of the present application. The conference recording method can be applied to a client or a server that provides an online conference function. The online conference includes an online audio conference or an online video conference. In other words, the execution subject of this embodiment can be the above-mentioned Figure 1 Any client in , or Figure 1 The server in. Figure 12 As shown, the method of this embodiment may include:
[0225] Step 1201: Record the main interface screen to generate video data.
[0226] The video data of this embodiment may also be referred to as screen recording data. The main interface screen refers to the content displayed on the main interface of the client that provides the online conference function.
[0227] In one implementation, the main interface screen of all clients accessing the online meeting is the same. For example, the main interface screen can be a shared desktop or file, which can include but is not limited to slides (PPT), text documents, etc. The shared desktop or file can be a desktop or file shared by any client accessing the online meeting. Optionally, the main interface screen can also include shared annotations. The shared annotations can be annotations of the client used by the speaker. In this way, the recording of the main interface screen is a public recording. After the online meeting ends, the video data generated by the public recording of the main interface screen can be shared by multiple clients.
[0228] Another possible implementation method is that the main interface screens of different clients accessing the online conference can be different. For example, the main interface screen includes not only the shared desktop or files, but also the annotations of the participants. In this way, recording the main interface screen is recording separately. For example, a client records its own main interface screen and generates its own video data. Among them, the client provides an annotation function. Participants using the client can use the annotation function to annotate the content displayed on their own screens, and when the participant is not the speaker, the content annotated by the participant can only be seen by himself. The annotation function can be used to annotate the content that the participant considers important.
[0229] If different participants can be individually tagged, the main interface screens of different clients will have different participant tags. Each client can record its own main interface screen and obtain its own video data. This allows for personalized meeting records that include participant tags and belong to each participant. Each participant's meeting record will be unique due to the different tag content.
[0230] For separate recording, the execution entity of this embodiment can be a client, which can record its own main interface screen to generate video data of the client, that is, video data of the participants using the client. For separate recording, the execution entity of this embodiment can also be a server, which can record the main interface screen of each client connected to the online meeting to generate video data of each client, that is, video data of each participant.
[0231] You can start recording the main interface screen at any of the following times: when a client connects to an online meeting; when a scheduled online meeting starts; when a client shares their desktop or file; or when a participant speaks. You can set the specific time to start recording the main interface screen based on your needs.
[0232] Step 1202: Associate at least one audio data and video data distributed on the time axis of each participant according to the recording time of the video data.
[0233] According to the recording time of the video data, the audio data on the timeline of each participant is associated with the recorded video data, so that when the audio data is played, the video data of the corresponding time is played synchronously. The audio data played can be the audio data on the timeline of a single participant, or the audio data associated by the user through dragging operations. Similarly, when the audio data playback ends, the video data playback stops.
[0234] For the video data obtained from public recording, all audio data on the timeline of each participant can be associated with the video data at the corresponding time according to the time progress of the video data, so that when the audio data on the timeline of different participants are played, the video data at the corresponding time will be played synchronously.
[0235] For the video data recorded separately, all audio data on the timeline of the participant using a client can be associated with the video data of the corresponding time according to the time progress of the video data of the client, so that when the audio data on the timeline of the participant is played, the video data of the corresponding time will be played synchronously.
[0236] Taking the audio data identified by an audio identifier on the timeline of a participant as an example, the video data associated with the audio data is the video data obtained by recording the main interface screen during the time period between the start time and the end time of the audio data.
[0237] Correspondingly, when the online conference review interface is displayed, the online conference review interface may further include a video display area, and the video display area is used to play video data.
[0238] After the client provides the online meeting review interface, the following operations performed by the user on the online meeting review interface can be detected to present all or part of the meeting scene.
[0239] Scenario 4, the sixth trigger operation.
[0240] The client displays the online meeting review interface described above and detects a sixth trigger operation of one of the multiple audio identifiers in the online meeting review interface. In response to the sixth trigger operation, the client displays a thumbnail of the video data corresponding to the audio identifier. The video data corresponding to the audio identifier refers to the video data generated by the main interface screen during the time period between the start time and the end time of the audio data identified by the audio identifier. The thumbnail of the video data can be a key frame in the video data.
[0241] For example, the sixth trigger operation may be a hover operation of a mouse or other control object. When the client detects that the mouse or other control object is hovering over an audio identifier among the multiple audio identifiers, the keyword of the audio data corresponding to the audio identifier and the thumbnail of the video data corresponding to the audio identifier are displayed.
[0242] The client detects a seventh trigger operation of the thumbnail or the audio identifier, and plays the video data and audio data corresponding to the audio identifier in response to the seventh trigger operation.
[0243] For example, the seventh trigger operation may be a click operation of a mouse or other control object. When the client detects that the mouse or other control object clicks an audio identifier, the audio data corresponding to the audio identifier and the video data corresponding to the audio identifier are played. The played video data may be presented in the video display area.
[0244] An example, Figure 13 This is a schematic diagram of a user interface for an online meeting provided in an embodiment of the present application. As an example and not a limitation, Figure 13 As shown, in Figure 7 Based on the online conference review interface shown in FIG, the online conference review interface of this embodiment may further include a video display area B01. When the client detects that the mouse or other control object is hovering over the video display area B02, the video display area B03 is displayed. Figure 13 The audio logo A06 shown in FIG, then the following is displayed above the audio logo Figure 7 The thumbnail B02 of the video data corresponding to the keyword "press conference preparation work" and the audio identifier is shown. Of course, it is understandable that the thumbnail can also be placed at other locations such as below the audio identifier, and this embodiment of the application does not impose specific restrictions on this. When the client detects that the mouse or other control object is clicked, such as Figure 13 The audio identifier A06 shown is played, and the audio data identified by the audio identifier A06 is played, and the corresponding video data is played in the video display area B01.
[0245] In this embodiment, by recording the video data of the main interface screen, the audio data on the timeline of each participant and the video data of the recorded main interface screen are associated according to the recording time of the video data, so that when the audio data is played, the video data at the same moment is played synchronously. By combining audio playback and video playback, the meeting scene at that time can be completely reproduced, further accelerating the information comparison and determination process.
[0246] The main interface can be recorded either publicly or individually. Individual recording allows you to capture attendee annotations, creating personalized video data unique to each attendee. Ultimately, each attendee receives a uniquely annotated meeting video.
[0247] Figure 14 This is a flowchart of a conference recording method provided in an embodiment of the present application. The conference recording method can be applied to a client that provides an online conference function, where the online conference includes an online audio conference or an online video conference. In other words, the execution subject of this embodiment can be the above-mentioned Figure 1 Any client in . Figure 14 As shown, the method of this embodiment may include:
[0248] Step 1401: Detect the operation of the marking function, and generate a marking mark in response to the operation of the marking function.
[0249] The server and client can provide annotation functionality. During a meeting, participants can use this feature to annotate key information or simply mark time points. The client monitors participants' annotation operations in real time and, based on these operations, creates new annotation tags and records the time points.
[0250] Step 1402: Associate the annotation identifier with the timeline of the corresponding participant according to the time point of the operation of the annotation function.
[0251] In this embodiment, the client is executed by a client. Based on the time of the participant's annotation operation using the client, the client can associate the annotation identifier with the participant's timeline. The annotation identifier and the audio data recording do not interfere with each other. At the same time, annotation and speech can be performed simultaneously without affecting their association with the corresponding participant's timeline.
[0252] Optionally, after the above steps 1401 and 1402, the client of this embodiment may further synchronize the annotation identifiers associated with the participants using the client to the server.
[0253] The aforementioned marking mark can be used when the participant is acting as a speaker or when the participant is not acting as a speaker. The marking mark when acting as a speaker will be displayed in the online meeting review interface of other participants, while the marking mark when not acting as a speaker will only be displayed in the participant's online meeting review interface and will not be visible in the online meeting review interface of other participants.
[0254] After the online meeting, the online meeting review interface will include all speaker markers. All speaker markers are distributed on the corresponding participant's timeline. Each marker is used to identify the participant in the timeline where the marker is located as a speaker, and the marker is used to mark the time point where the marker is located.
[0255] In some embodiments, the online meeting review interface may also include attendee identification on the client displaying the online meeting review interface. This includes identification of attendees when they are not presenting. This allows attendees to independently annotate their notes when they are not presenting, and review them independently after the meeting, resulting in a personalized annotated meeting record.
[0256] After the client provides the online meeting review interface, the following operations performed by the user on the online meeting review interface can be detected to present all or part of the meeting scene.
[0257] Scene 5, the eighth trigger operation.
[0258] The client displays the online meeting review interface as described above, detects an eighth trigger operation of a marking mark on the online meeting review interface, and responds to the eighth trigger operation by playing the audio data, or the audio data and video data, at the time point where the marking mark is located.
[0259] For example, the eighth trigger operation may be a click operation, etc. When the client detects that the user clicks on a marked mark, the audio data, or the audio data and video data, at the time point where the marked mark is located are played.
[0260] An example, Figure 15 This is a schematic diagram of a user interface for an online meeting provided in an embodiment of the present application. As an example and not a limitation, Figure 15 As shown, in Figure 7 Based on the online conference review interface shown in FIG, the online conference review interface of this embodiment may further include a marking mark C01 and two marking marks C02. When the client detects that a click is made as shown in FIG. Figure 15 The marked mark C01 shown in the figure plays the audio data, or the audio data and video data at the time point where the marked mark is located. Figure 15The online meeting review interface shown is the interface displayed by the client used by participant 1, as shown in FIG. Figure 15 As shown, for participant 1, the annotation mark C01 indicates that participant 3 has made an annotation when acting as the speaker, while the two annotation marks C02 indicate that participant 1 has made the annotation himself. These two annotation marks C02 are only provided to participant 1. Of course, it is understandable that these two annotation marks C02 can also be provided to other participants, and this embodiment of the application does not specifically limit this.
[0261] In this embodiment, by detecting the operation of the annotation function of the participant, a corresponding annotation identifier is generated and linked to the timeline of the corresponding participant. After the online meeting is over, the time node of the key content can be quickly located through the annotation identifier, and the audio data, or audio data and video data of the time node can be directly played.
[0262] Figure 16 This is a flowchart of a conference recording method provided in an embodiment of the present application. The conference recording method can be applied to a client that provides an online conference function, where the online conference includes an online audio conference or an online video conference. In other words, the execution subject of this embodiment can be the above-mentioned Figure 1 Through the above embodiment, an online conference review interface is obtained, which may include a video display area and a timeline area. The timeline area includes the timelines of multiple participants and multiple audio identifiers distributed along the timeline. Afterwards, the online conference review interface can be adjusted by the method of this embodiment. Figure 16 As shown, the method of this embodiment may include:
[0263] Step 1601: Detect the operation of reducing or increasing the time axis area.
[0264] After providing an online meeting review interface using any of the above-described meeting recording methods, the online meeting review interface can be adjusted to present all or part of the meeting scene by detecting any of the following user operations in the video display area or timeline area. Detecting any of the following user operations in the video display area or timeline area includes, but is not limited to, sliding, dragging, and other operations.
[0265] An example, Figure 17 This is a schematic diagram of a user interface for an online meeting provided in an embodiment of the present application. As an example and not a limitation, Figure 17 As shown, the online meeting review interface may include a video display area B01 and a timeline area D01. The video display area B01 is located above the timeline area D01. Of course, it is understandable that the video display area B01 may also be located below the timeline area D01. Figure 17The interface shown is used as an example. The operation of reducing the timeline area can be that the user's finger, stylus or other control object that can be detected by the touch display screen of the terminal device acts on the video display area and slides downward; or, the user's finger, stylus or other control object that can be detected by the touch display screen of the terminal device acts on the timeline area and slides upward; or, the user acts on the video display area with a mouse and scrolls the wheel downward; or, the user acts on the timeline area with a mouse and scrolls the wheel upward, etc. The embodiments of this application do not give examples one by one. Similarly, the operation of increasing the timeline area can be the reverse operation of reducing the timeline area. For example, the operation of increasing the timeline area can be that the user's finger, stylus or other control object that can be detected by the touch display screen of the terminal device acts on the video display area and slides upward, etc.
[0266] Step 1602: In response to the operation of reducing or increasing the timeline area, adjust the video display area and the timeline area.
[0267] When an operation to shrink the timeline area is detected, in response to the operation to shrink the timeline area, the timeline area is shrunk, the distance between the timelines of multiple participants is reduced, and the video display area is increased. By increasing the video display area, a clearer video data review experience of the meeting record can be provided to the user.
[0268] The timeline area can be reduced to the point where the timelines of multiple participants completely overlap, that is, the timelines of multiple participants are merged into one timeline. For example, Figure 18 This is a schematic diagram of a user interface for an online meeting provided in an embodiment of the present application. As an example and not a limitation, Figure 18 As shown, the timelines of multiple participants in the timeline area of the online conference review interface overlap, and the online conference review interface displays a merged timeline D05. The merged timeline D05 can be the timeline of any participant, or a pre-set total timeline. According to the position of each audio identifier and each annotation identifier in the timeline of each participant before the merger, each audio identifier and each annotation identifier are displayed on the merged timeline D05. For example, Figure 18 As shown, the audio identifier D31, the audio identifier D32 and the annotation identifier D04 are distributed along the merged time axis. Other audio identifiers and annotation identifiers are not shown here one by one.
[0269] Optionally, the audio data of multiple audio identifiers distributed on the merged timeline D05 can also be associated. In this way, the multi-person discussion scene during the online meeting can be reproduced, which is suitable for reviewing the entire meeting as a whole.
[0270] Among the multiple audio identifiers distributed on the merged timeline D05, at least two audio identifiers overlap with each other. The embodiment of the present application can also detect a tenth trigger operation of the two overlapping audio identifiers. The tenth trigger operation can be an operation such as a click or double-click on any one of the two audio identifiers or the overlapping part of the two audio identifiers. In response to the tenth trigger operation, one of the two audio identifiers is displayed on the merged timeline D05, and the other audio identifier is displayed above or below the merged timeline D05, so that the two audio identifiers do not overlap. For example, Figure 19 This is a schematic diagram of a user interface for an online conference provided in an embodiment of the present application. As an example and not a limitation, audio identifier D31 and audio identifier D32 overlap on the merged timeline D05. When the user clicks on the overlapped portion, audio identifier D31 and audio identifier D32 are separated from each other, as shown in FIG. Figure 19 As shown, in response to the user clicking on the overlapped portion of the two, the client displays Figure 19 In the online meeting review interface shown, audio identifier D32 can be suspended in a direction perpendicular to the merged timeline D05, that is, suspended above audio identifier D31. The user can then operate audio identifier D31 and audio identifier D32 separately to play the audio data corresponding to audio identifier D31 and audio identifier D32 respectively.
[0271] When an operation to increase the timeline area is detected, the timeline area is increased in response to the operation to increase the timeline area, and the distance between the timelines of multiple participants is increased, and the video display area is reduced. By increasing the timeline area, the user can be provided with clearer audio identification and / or annotation identification of the meeting records to quickly locate the audio data and / or annotation identification of different participants at different time points.
[0272] Illustratively, the timeline region may be enlarged until the timelines of multiple participants completely overlap, that is, the timelines of multiple participants are merged into one timeline.
[0273] In this embodiment, by reducing the timeline area, multiple participants' timelines are merged into a single timeline, and the audio data of multiple participants is associated. The associated audio data allows users to review the entire online meeting as a whole. The enlarged video display area provides users with a clearer video data review experience of the meeting recording. This allows for flexible adjustment of the timeline area and video display area to meet user needs.
[0274] Based on the same inventive concept, an embodiment of the present application further provides a conference recording device, which can be a chip or system-on-chip in a terminal device, or a functional module in the terminal device for implementing the method described in any of the above possible implementations. The chip or system-on-chip includes a memory, wherein the memory stores instructions, and when the instructions are called by the system-on-chip or chip, the above method is executed.
[0275] Please refer to Figure 20 , Figure 20 The diagram of the composition of a conference recording device provided by an embodiment of the present application is shown. The conference recording device is used to provide an online conference function, and the online conference includes an online audio conference or an online video conference. Figure 20 As shown, the conference recording device 2000 may include: a processing module 2001 and a display module 2002.
[0276] Processing module 2001, configured to detect a triggering operation of an online conference ending option or a triggering operation of an online conference record viewing option;
[0277] The processing module 2001 is also used to respond to the trigger operation of the online meeting end option or the trigger operation of the online meeting record viewing option, and display the online meeting review interface through the display module 2002. The online meeting review interface includes a timeline of multiple participants and multiple audio identifiers distributed along the timeline. Each of the multiple audio identifiers is used to identify a segment of audio data of the participant on the timeline where the audio identifier is located. The starting position of the audio identifier on the timeline is the starting time of the segment of audio data of the participant, and the ending position of the audio identifier on the timeline is the ending time of the segment of audio data of the participant. The segment of audio data of the participant is generated by recording the conference speech voice signal of the participant in the time period between the start time and the end time.
[0278] Exemplarily, the processing module 2001 and the display module 2002 are used to execute the above Figure 5 or Figure 9 The method steps involved in the client 1 of the method embodiment are shown.
[0279] In some embodiments, the processing module 2001 is further configured to detect a first triggering operation of an audio identifier among the multiple audio identifiers, and in response to the first triggering operation, display a keyword of the audio data corresponding to the audio identifier via the display module 2002 .
[0280] In some embodiments, the conference recording device may further include an audio module 2003. The processing module 2001 is further configured to detect a second triggering operation of an audio identifier among the multiple audio identifiers, and in response to the second triggering operation, play audio data corresponding to the audio identifier through the audio module 2003.
[0281] In some embodiments, the processing module 2001 is also used to detect a third trigger operation of the first audio identifier among the multiple audio identifiers from the first time axis to the second time axis, the first audio identifier is distributed on the first time axis, and at least one second audio identifier is distributed on the second time axis; in response to the third trigger operation, the first audio identifier is displayed on the second time axis through the display module 2002, and the first audio data is associated with the second audio data corresponding to the at least one second audio identifier according to the start time and end time of the first audio data corresponding to the first audio identifier; detect a fourth trigger operation of the first audio identifier or the at least one second audio identifier, and in response to the fourth trigger operation, play the first audio data and the second audio data corresponding to the at least one second audio identifier through the audio module 2003.
[0282] In some embodiments, the time period of the first audio identifier overlaps with the time period of the at least one second audio identifier, and playing the first audio data and the second audio data corresponding to the at least one second audio identifier includes: playing at least two of the non-intersecting parts of the first audio data, the intersecting parts of the first audio data and the second audio data corresponding to the at least one second audio identifier, or the non-intersecting parts of the second audio data corresponding to the at least one second audio identifier, according to the time position of the first audio identifier on the second time axis and the time position of the at least one second audio identifier on the second time axis.
[0283] In some embodiments, the processing module 2001 is also used to detect the fifth trigger operation of the first audio identifier. In response to the fifth trigger operation, the first audio identifier is displayed on the first timeline through the display module 2002, and the first audio data and the second audio data corresponding to the at least one second audio identifier are disassociated. The disassociation is used to allow the first audio data and the second audio data corresponding to the second audio identifier to be played independently.
[0284] In some embodiments, the conference recording device may further include a communication module 2004. The processing module 2001 is further configured to, before detecting a trigger operation for the online conference end option or a trigger operation for the online conference record viewing option, record at least one audio data of a participant using the client through the audio module 2003, and record the start time and end time of each of the at least one audio data on the participant's timeline; and send the at least one audio data and the start time and end time of each of the at least one audio data on the participant's timeline to the server through the communication module 2004; wherein the at least one audio data and the start time and end time of each of the at least one audio data on the participant's timeline are used to generate the online conference review interface.
[0285] In some embodiments, the at least one audio data and the at least one audio data are each stored in at least one storage unit at the start time and end time of the participant's time axis; the at least one storage unit is connected in series through the participant's time pointer.
[0286] In some embodiments, the processing module 2001 is also used to detect the sixth trigger operation of an audio identifier among the multiple audio identifiers. In response to the sixth trigger operation, a thumbnail of the video data corresponding to the audio identifier is displayed through the display module 2002. The video data corresponding to the audio identifier is generated by the main interface screen within the time period between the start time and the end time of recording the audio data corresponding to the audio identifier.
[0287] In some embodiments, the processing module 2001 is further configured to detect a seventh trigger operation of the thumbnail, and in response to the seventh trigger operation, play the video data and audio data corresponding to the audio identifier through the audio module 2003 and the display module 2002 .
[0288] In some embodiments, the online meeting review interface also includes at least one annotation identifier, which is distributed on the timeline of at least one participant. Each of the at least one annotation identifier is used to identify the participant on the timeline where the annotation identifier is located, and the annotation at the time point where the annotation identifier is located.
[0289] In some embodiments, the processing module 2001 is also used to detect the eighth trigger operation of one of the at least one annotation identifier, and in response to the eighth trigger operation, play the audio data at the time point where the annotation identifier is located through the audio module 2002, or play audio data and video data through the audio module 2003 and the display module 2002.
[0290] In some embodiments, the timelines of the multiple participants and the multiple audio identifiers distributed along the timelines are located in the timeline area of the online meeting review interface, which also includes a video display area; the processing module 2001 is also used to: detect an operation to shrink the timeline area, and in response to the operation to shrink the timeline area, shrink the timeline area, shrink the distance between the timelines of the multiple participants, and increase the video display area; or, detect an operation to increase the timeline area, and in response to the operation to increase the timeline area, increase the timeline area and decrease the video display area.
[0291] In some embodiments, when the distance between the time axes of the multiple participants is reduced until the time axes of the multiple participants completely overlap, at least two audio identifiers among the multiple audio identifiers overlap with each other. The processing module 2001 is also used to: detect the tenth trigger operation of the two overlapping audio identifiers, and in response to the tenth trigger operation, display one of the two audio identifiers on the time axis through the display module 2002, and display the other audio identifier above or below the time axis, so that the two audio identifiers do not overlap.
[0292] Of course, the unit modules in the conference recording device include but are not limited to the processing module 2001, the display module 2002, etc. For example, the terminal device may further include a storage module, etc.
[0293] In addition, the processing module is one or more processors. The one or more processors, memory and communication module can be connected together, for example, via a bus. The memory is used to store computer program code, which includes instructions. When the processor executes the instruction, the terminal device can execute the relevant method steps in the above embodiment to implement the method in the above embodiment. The communication module can be a wireless communication unit (such as Figure 3 , or a wireless communication module 350 or 360 as shown in FIG.
[0294] An embodiment of the present application also provides a computer-readable storage medium, which stores computer software instructions. When the computer software instructions are executed in an information processing device, the information processing device can execute the relevant method steps in the above embodiments to implement the methods in the above embodiments.
[0295] An embodiment of the present application further provides a computer program product. When the computer program product is run on a computer, the computer is caused to execute the relevant method steps in the above embodiments to implement the methods in the above embodiments.
[0296] Among them, the terminal device, computer storage medium or computer program product provided in the embodiments of the present application are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods provided above, and will not be repeated here.
[0297] Through the description of the above embodiments, those skilled in the art will clearly understand that for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. The specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0298] In the several embodiments provided in this application, it should be understood that the disclosed methods can be implemented in other ways. For example, the vehicle-mounted terminal embodiments described above are merely schematic. For example, the division of the modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of modules or units, which can be electrical, mechanical or other forms.
[0299] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0300] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0301] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program instructions, such as flash memory, mobile hard disk, read-only memory, random access memory, magnetic disk or optical disk.
[0302] The above is only a specific embodiment of the present application, but the scope of protection of this application is not limited to this. Any changes or substitutions within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A conference recording method, which is applied to a client that provides an online conference function, wherein the online conference includes an online audio conference or an online video conference, and is characterized in that: include: Detecting a triggering operation of the online meeting end option or a triggering operation of the online meeting record viewing option; In response to the trigger operation of the online meeting end option or the trigger operation of the online meeting record viewing option, an online meeting review interface is displayed. The online meeting review interface includes multiple time axes and multiple audio identifiers distributed along the time axes. The multiple time axes correspond one-to-one to multiple participants. Each of the multiple audio identifiers is used to identify a segment of audio data of the first participant on the time axis where the audio identifier is located. The starting position of the audio identifier on the time axis is the starting time of the segment of audio data of the first participant, and the ending position of the audio identifier on the time axis is the ending time of the segment of audio data of the first participant. The segment of audio data of the first participant is generated by recording the conference speech voice signal of the first participant in the time period between the start time and the end time.
2. The method according to claim 1, characterized in that The method further comprises: A first triggering operation of an audio identifier among the plurality of audio identifiers is detected, and in response to the first triggering operation, a keyword of audio data corresponding to the audio identifier is displayed.
3. The method according to claim 1 or 2, characterized in that The method further comprises: A second triggering operation of an audio identifier among the multiple audio identifiers is detected, and in response to the second triggering operation, audio data corresponding to the audio identifier is played.
4. The method according to any one of claims 1 to 3, characterized in that The method further comprises: detecting a third triggering operation of a first audio identifier among the plurality of audio identifiers from a first time axis to a second time axis, the first audio identifier being distributed on the first time axis and at least one second audio identifier being distributed on the second time axis; In response to the third trigger operation, displaying the first audio identifier on the second timeline, and associating the first audio data with the second audio data corresponding to the at least one second audio identifier based on the start time and end time of the first audio data corresponding to the first audio identifier; A fourth triggering operation of detecting the first audio identifier or the at least one second audio identifier is detected, and in response to the fourth triggering operation, the first audio data and the second audio data corresponding to the at least one second audio identifier are played.
5. The method according to claim 4, characterized in that The time period of the first audio identifier and the time period of the at least one second audio identifier overlap, and the playing of the first audio data and the second audio data corresponding to the at least one second audio identifier includes: According to the time position of the first audio identifier on the second time axis and the time position of the at least one second audio identifier on the second time axis, at least two of the non-intersecting parts of the first audio data, the intersecting parts of the first audio data and the second audio data corresponding to the at least one second audio identifier, or the non-intersecting parts of the second audio data corresponding to the at least one second audio identifier are played.
6. The method according to claim 4 or 5, characterized in that The method further comprises: Detecting a fifth trigger operation of the first audio identifier; displaying the first audio identifier on the first timeline in response to the fifth trigger operation; and disassociating the first audio data from the second audio data corresponding to the at least one second audio identifier, wherein the disassociation is used to independently play the first audio data and the second audio data corresponding to the second audio identifier.
7. The method according to any one of claims 1 to 6, characterized in that Before detecting the triggering operation of the online meeting end option or the triggering operation of the online meeting record viewing option, the method further includes: Recording and generating at least one audio data of a participant using the client, and recording the start time and end time of each of the at least one audio data on the timeline of the participant; Sending the at least one audio data and the start time and end time of each of the at least one audio data on the timeline of the participant to the server; The at least one audio data and the at least one audio data are each used to generate the online conference review interface at the start time and end time of the participant's timeline.
8. The method according to claim 7, characterized in that The at least one audio data and the at least one audio data are each stored in at least one storage unit at a start time and an end time of the participant's time axis; The at least one storage unit is connected in series via the time pointer of the participant.
9. The method according to any one of claims 1 to 8, characterized in that The method further comprises: Detect a sixth trigger operation of one of the multiple audio identifiers, and in response to the sixth trigger operation, display a thumbnail of the video data corresponding to the audio identifier, where the video data corresponding to the audio identifier is generated on the main interface screen during the time period between the start time and the end time of recording the audio data corresponding to the audio identifier.
10. The method according to claim 9, characterized in that The method further comprises: A seventh trigger operation of the thumbnail is detected, and in response to the seventh trigger operation, the video data and the audio data corresponding to the audio identifier are played.
11. The method according to any one of claims 1 to 10, characterized in that The online meeting review interface also includes at least one annotation identifier, which is distributed on the timeline of at least one participant. Each of the at least one annotation identifier is used to identify the participant on the timeline where the annotation identifier is located, and the annotation at the time point where the annotation identifier is located.
12. The method according to claim 11, characterized in that The method further comprises: An eighth triggering operation of detecting one of the at least one marking identifier is performed, and in response to the eighth triggering operation, audio data, or audio data and video data at a time point where the marking identifier is located is played.
13. The method according to any one of claims 1 to 12, characterized in that The timelines of the multiple participants and the multiple audio identifiers distributed along the timelines are located in a timeline area of the online conference review interface, and the online conference review interface further includes a video display area; detecting an operation of reducing the timeline area, and in response to the operation of reducing the timeline area, reducing the timeline area, reducing the distance between the timelines of the plurality of participants, and increasing the video display area; or An operation of increasing the time axis area is detected, and in response to the operation of increasing the time axis area, the time axis area is increased and the video display area is reduced.
14. The method according to any one of claims 1 to 13, characterized in that When the distance between the time axes of the multiple participants is reduced until the time axes of the multiple participants completely overlap, at least two audio identifiers among the multiple audio identifiers overlap with each other, and the method further includes: A tenth trigger operation of detecting two audio identifiers that overlap each other is detected, and in response to the tenth trigger operation, one of the two audio identifiers is displayed on a timeline and the other audio identifier is displayed above or below the timeline so that the two audio identifiers do not overlap.
15. A terminal device, characterized in that: include: A processor, a memory, and a display screen, wherein the memory and the display screen are coupled to the processor, the memory being used to store computer program code, the computer program code including computer instructions for a client providing an online conference function, wherein the online conference includes an online audio conference or an online video conference, and when the processor reads the computer instructions from the memory, the terminal device performs the following operations: Detecting a triggering operation of the online meeting end option or a triggering operation of the online meeting record viewing option; In response to the triggering operation of the online meeting end option or the triggering operation of the online meeting record viewing option, an online meeting review interface is displayed, and the online meeting review interface includes multiple time axes and multiple audio identifiers distributed along the time axes. The multiple time axes correspond one-to-one to multiple participants. Each of the multiple audio identifiers is used to identify a segment of audio data of the first participant on the time axis where the audio identifier is located. The starting position of the audio identifier on the time axis is the starting time of the segment of audio data of the first participant, and the ending position of the audio identifier on the time axis is the ending time of the segment of audio data of the first participant. The segment of audio data of the first participant is generated by recording the conference speech voice signal of the first participant in the time period between the start time and the end time.
16. The terminal device according to claim 15, characterized in that The terminal device further executes: A first trigger operation of an audio identifier among the plurality of audio identifiers is detected, and in response to the first trigger operation, keywords of audio data corresponding to the audio identifier are displayed.
17. The terminal device according to claim 15 or 16, characterized in that: The terminal device further executes: A second triggering operation of an audio identifier among the multiple audio identifiers is detected, and in response to the second triggering operation, audio data corresponding to the audio identifier is played.
18. The terminal device according to any one of claims 15 to 17, characterized in that: The terminal device further executes: detecting a third triggering operation of a first audio identifier among the plurality of audio identifiers from a first time axis to a second time axis, the first audio identifier being distributed on the first time axis and at least one second audio identifier being distributed on the second time axis; In response to the third trigger operation, displaying the first audio identifier on the second timeline, and associating the first audio data with the second audio data corresponding to the at least one second audio identifier based on the start time and end time of the first audio data corresponding to the first audio identifier; A fourth triggering operation of detecting the first audio identifier or the at least one second audio identifier is detected, and in response to the fourth triggering operation, the first audio data and the second audio data corresponding to the at least one second audio identifier are played.
19. The terminal device according to claim 18, characterized in that The time period of the first audio identifier and the time period of the at least one second audio identifier overlap, and the playing of the first audio data and the second audio data corresponding to the at least one second audio identifier includes: According to the time position of the first audio identifier on the second time axis and the time position of the at least one second audio identifier on the second time axis, at least two of the non-intersecting parts of the first audio data, the intersecting parts of the first audio data and the second audio data corresponding to the at least one second audio identifier, or the non-intersecting parts of the second audio data corresponding to the at least one second audio identifier are played.
20. The terminal device according to claim 18 or 19, characterized in that: The terminal device further executes: Detecting a fifth trigger operation of the first audio identifier; displaying the first audio identifier on the first timeline in response to the fifth trigger operation; and disassociating the first audio data from the second audio data corresponding to the at least one second audio identifier, wherein the disassociation is used to independently play the first audio data and the second audio data corresponding to the second audio identifier.
21. The terminal device according to any one of claims 15 to 20, characterized in that: Before detecting the triggering operation of the online conference ending option or the triggering operation of the online conference record viewing option, the terminal device further performs: Recording and generating at least one audio data of a participant using the client, and recording the start time and end time of each of the at least one audio data on the timeline of the participant; Sending the at least one audio data and the start time and end time of each of the at least one audio data on the timeline of the participant to the server; The at least one audio data and the at least one audio data are each used to generate the online conference review interface at the start time and end time of the participant's timeline.
22. The terminal device according to claim 21, characterized in that The at least one audio data and the at least one audio data are each stored in at least one storage unit at a start time and an end time of the participant's time axis; The at least one storage unit is connected in series via the time pointer of the participant.
23. The terminal device according to any one of claims 15 to 22, characterized in that: The terminal device further executes: Detect a sixth trigger operation of one of the multiple audio identifiers, and in response to the sixth trigger operation, display a thumbnail of the video data corresponding to the audio identifier, where the video data corresponding to the audio identifier is generated on the main interface screen during the time period between the start time and the end time of recording the audio data corresponding to the audio identifier.
24. The terminal device according to claim 23, characterized in that The terminal device further executes: A seventh trigger operation of the thumbnail is detected, and in response to the seventh trigger operation, the video data and the audio data corresponding to the audio identifier are played.
25. The terminal device according to any one of claims 15 to 24, characterized in that: The online meeting review interface also includes at least one annotation identifier, which is distributed on the timeline of at least one participant. Each of the at least one annotation identifier is used to identify the participant on the timeline where the annotation identifier is located, and the annotation at the time point where the annotation identifier is located.
26. The terminal device according to claim 25, characterized in that The terminal device further executes: An eighth triggering operation of detecting one of the at least one marking identifier is performed, and in response to the eighth triggering operation, audio data, or audio data and video data at a time point where the marking identifier is located is played.
27. The terminal device according to any one of claims 15 to 26, characterized in that: The timelines of the multiple participants and the multiple audio identifiers distributed along the timelines are located in a timeline area of the online conference review interface, and the online conference review interface further includes a video display area; detecting an operation of reducing the timeline area, and in response to the operation of reducing the timeline area, reducing the timeline area, reducing the distance between the timelines of the plurality of participants, and increasing the video display area; or An operation of increasing the time axis area is detected, and in response to the operation of increasing the time axis area, the time axis area is increased and the video display area is reduced.
28. The terminal device according to any one of claims 15 to 27, characterized in that: When the distances between the time axes of the multiple participants are shortened until the time axes of the multiple participants completely overlap, and at least two audio identifiers among the multiple audio identifiers overlap with each other, the terminal device further executes: A tenth trigger operation of detecting two audio identifiers that overlap each other is detected, and in response to the tenth trigger operation, one of the two audio identifiers is displayed on a timeline and the other audio identifier is displayed above or below the timeline so that the two audio identifiers do not overlap.
29. A computer-readable storage medium, characterized in that The method comprises computer instructions, which, when executed on a terminal device, enable the terminal device to execute the conference recording method according to any one of claims 1 to 14.
30. A computer program product, characterized in that When the computer program product runs on a computer, it enables the computer to execute the meeting recording method according to any one of claims 1 to 14.
31. A conference recording system, characterized in that: The conference recording system includes a server and multiple clients. The server establishes communication connections with the multiple clients respectively, and the multiple clients are each used to execute the conference recording method according to any one of claims 1 to 14.
Citation Information
Patent Citations
Method and equipment for presenting event tag information
CN113329237A
Cited By
Conference recording method, terminal device, and conference recording system
EP4750047A2