Method, apparatus and medium for recording viewing based on multi-dimensional conference information
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GUANGZHOU SHIYUAN ELECTRONICS CO LTD
- Filing Date
- 2024-02-23
- Publication Date
- 2026-08-07
AI Technical Summary
[0005]本发明提供了一种基于多维会议信息的记录查看方法、装置、设备和介质,以解决对会议记录中的内容进行查看时的精细度较低,对会议记录中目标位置的定位较慢的技术问题
Smart Images

Figure CN120547295B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of electronic conferencing technology, and more particularly to methods, apparatus, devices, and media for viewing records based on multidimensional conferencing information. Background Technology
[0002] Taking meeting minutes is a frequent activity in office settings. The purpose of meeting minutes is to fully reconstruct the conversations of participants and extract the key points of the communication process. With the development of camera capabilities and voice recognition technology, the methods of recording and presenting meeting minutes have also made great progress.
[0003] When taking meeting minutes in existing electronic device-based meetings, the main methods are recording the meeting in chronological order by either recording video or transcribing audio. In addition to linear recording in chronological order, meeting minutes can be generated by taking one speech by the same speaker as the recording unit.
[0004] When the inventors reviewed meeting minutes generated during electronic device-based meetings, they found that the level of detail in viewing the content of the meeting minutes was low, and the location of target positions in the meeting minutes was slow. Summary of the Invention
[0005] This invention provides a method, apparatus, device, and medium for viewing meeting records based on multidimensional meeting information, in order to solve the technical problems of low precision in viewing the content of meeting records and slow location of target positions in meeting records.
[0006] In a first aspect, embodiments of this application provide a method for viewing recordings based on multidimensional meeting information, the method comprising:
[0007] The meeting minutes are displayed in the meeting minutes viewing interface. The meeting minutes are composed of sub-records. Each sub-record corresponds to the continuous speech of the same user ID on the same screen content data during the meeting. The user ID and screen content data corresponding to the associated record are used as two filtering information.
[0008] In the first viewing state, a first trigger operation is received that acts on the first target information. The record viewing interface displays a first option bar and a second option bar in the first viewing state. The first option bar displays each first option label to present each user identifier in the meeting record. The second option bar displays each second option label to present each screen content data in the meeting record. The first target information is any first option label or second option label.
[0009] In response to the first trigger operation, the first result bar and the third option bar corresponding to the first target information are displayed on the record viewing interface. The first result bar is used to display the record identifier of the first filtered record. The first filtered record includes all sub-records in the meeting record that are associated with the filtered information presented by the first target information. The third option bar is used to display the third option label. The third option label is used to present another filtered information that is associated with the same sub-record as the filtered information presented by the first target information.
[0010] Receive a second trigger operation that acts on the second target information, where the second target information is one of all the third option tabs in the third option bar;
[0011] In response to the second trigger operation, the second result column corresponding to the second target information is displayed on the record viewing interface. The second result column is used to display the record identifier of the second filtered record. The second filtered record includes all sub-records in the first filtered record that are associated with the second target information.
[0012] Based on the characteristics of the close connection between the content of speeches and the focus of communication during the meeting, the switching process of the screen that is the focus of communication during the meeting is recorded. The audio and video recorded during the meeting are further broken down according to the switching process of the screen. Each sub-record in the meeting record corresponds to a finer granularity. When viewing the meeting record, the target position can be located more precisely and quickly according to the screen, thereby improving the user's viewing experience of the meeting record.
[0013] Before displaying the record viewing interface for viewing meeting minutes, the following are also included:
[0014] Acquire audio data, video data, and screen display data during the meeting. The screen display data includes the switching time of each screen transition and the screen content data after the transition. The video data is obtained by capturing images of the meeting venue through a camera. The screen transition refers to the switching of the screen displayed on the electronic devices in the meeting venue.
[0015] Based on video data and / or audio data, identify the user ID and speaking time period corresponding to each speech;
[0016] The segmentation structure of the recording data is confirmed based on the speaking time and the switching time. Each recording segment in the segmentation structure is associated with two filtering information, which includes a user identifier and a screen content data.
[0017] Meeting minutes are generated based on a segmented structure, corresponding to audio and video data. Each sub-record in the meeting minutes corresponds to an audio segment.
[0018] Based on the correlation between the content of speeches and the focus of communication during the meeting, the data collected during the meeting was split according to multiple dimensions. Thus, multiple recording segments were obtained by taking the continuous speeches of users corresponding to the same user ID on the same screen content data as units. Each recording segment generates a sub-record. When viewing the meeting records later, the user ID and screen content data associated with the sub-record can be used to search flexibly and quickly.
[0019] The segmentation structure of the recording data is determined based on the speaking time and the switching time. Each recording segment in the segmentation structure is associated with two filtering information, which includes a user identifier and screen content data, including:
[0020] The initial segmentation structure of the recording data is confirmed based on the speaking time period. Each initial recording segment in each initial segmentation structure is associated with a user identifier and the duration reaches the first preset duration.
[0021] The initial recording segment is split according to the switching time to obtain the recording segment. The recording segment is divided with reference to the end time of the display of the screen content data within the initial recording segment when the continuous display time reaches the second preset time. The user identifier associated with the recording segment is the same as the user identifier of the corresponding initial recording segment, and the user identifier associated with the recording segment is the same as the screen content data corresponding to the end point.
[0022] The above-mentioned method of extracting valid recordings based on recording duration and screen dwell time can eliminate the separate saving of invalid records, minimize the number of sub-records corresponding to each meeting's minutes, and thus improve the efficiency of viewing meeting minutes.
[0023] After displaying the record viewing interface for viewing meeting minutes, it also includes:
[0024] Upon receiving a first view operation, the system enters the first view state in response to the first view operation.
[0025] As mentioned above, users need to select a viewing mode to enter the first viewing state and view the meeting minutes based on the filtered information. This allows users to choose different viewing modes according to their own needs and improves the relevance of their meeting minutes viewing.
[0026] After displaying the record viewing interface for viewing meeting minutes, it also includes:
[0027] Upon receiving a second viewing operation, the system enters the second viewing state in response to the second viewing operation; the recording viewing interface displays the timeline corresponding to the recording data in the second viewing state, and the timeline corresponds to the time segment of the recording segment;
[0028] When a trigger operation is received on the target time segment via the timeline, the sub-record corresponding to the target segment is displayed on the record viewing interface.
[0029] As mentioned above, selecting the viewing mode enters the second viewing state. In the second viewing state, the meeting minutes are viewed by time segments in the timeline, allowing for quick and comprehensive confirmation of the target location.
[0030] The timeline displays time segments corresponding to recording segments differently from time segments without corresponding recording segments.
[0031] The above-mentioned distinction between valid and invalid records in the progress bar effectively reminds users of their priority targets for reviewing meeting minutes, thus improving review efficiency.
[0032] Each sub-record includes corresponding sub-audio recording data, text generated from the sub-audio recording data, sub-video data generated from the video data, and associated screen content data.
[0033] As mentioned above, each sub-record contains diverse content, which helps users obtain comprehensive information about the meeting process.
[0034] When a sub-record is displayed, the corresponding text and screen content data are displayed in separate areas, while the corresponding sub-audio recording data and sub-video data are output synchronously.
[0035] As mentioned above, the synchronized display of the corresponding record content for each sub-record effectively improves the convenience for users to view meeting minutes.
[0036] Secondly, embodiments of this application also provide a recording and viewing device based on multi-dimensional meeting information, the recording and viewing device based on multi-dimensional meeting information comprising:
[0037] The interface display unit is used to display the meeting record viewing interface. The meeting record consists of sub-records. Each sub-record corresponds to the continuous speech of the same user ID to the same screen content data during the meeting. The user ID and screen content data corresponding to the associated record are used as two filtering information.
[0038] The first filtering unit is used to receive a first trigger operation acting on the first target information in the first viewing state. The record viewing interface displays a first option bar and a second option bar in the first viewing state. Each first option label displayed in the first option bar is used to present each user identifier in the meeting record. Each second option label displayed in the second option bar is used to present each screen content data in the meeting record. The first target information is any first option label or second option label.
[0039] The first response unit is used to respond to the first trigger operation and display the first result bar and the third option bar corresponding to the first target information on the record viewing interface. The first result bar is used to display the record identifier of the first filtered record. The first filtered record includes all sub-records in the meeting record that are associated with the filtered information presented by the first target information. The third option bar displays each third option label for presenting another filtered information that is associated with the filtered information presented by the first target information in the same sub-record.
[0040] The second filtering unit is used to receive a second trigger operation applied to the second target information, where the second target information is one of all the third option labels in the third option bar.
[0041] The second response unit is used to respond to the second trigger operation and display the second result column corresponding to the second target information on the record viewing interface. The second result column is used to display the record identifier of the second filtered record. The second filtered record includes all sub-records in the first filtered record that are associated with the second target information.
[0042] The recording and viewing device based on multi-dimensional meeting information also includes:
[0043] The data acquisition unit is used to acquire audio data, video data and screen display data during the meeting. The screen display data includes the switching time of each screen switch and the screen content data after the switch. The video data is obtained by capturing images of the meeting venue through a camera. The screen switch refers to the switching of the screen displayed on the electronic devices in the meeting venue.
[0044] The first splitting unit is used to determine the user identifier and speaking time period corresponding to each speech based on video data and / or audio data;
[0045] The second splitting unit is used to confirm the segmentation structure of the recording data based on the speaking time and the switching time. Each recording segment in the segmentation structure is associated with two filtering information, which includes a user identifier and a screen content data.
[0046] The sub-record generation unit is used to generate meeting records corresponding to audio and video data based on the segmented structure. Each sub-record in the meeting record corresponds to an audio segment.
[0047] The second splitting unit includes:
[0048] The first segmentation module is used to determine the initial segmentation structure of the recording data based on the speaking time period. Each initial recording segment in each initial segmentation structure is associated with a user identifier and the duration reaches a first preset duration.
[0049] The second segmentation module is used to split the initial recording segment according to the switching time to obtain the recording segment. The recording segment is divided with reference to the end time of the display of the screen content data within the initial recording segment that has been continuously displayed for a duration of a second preset duration. The user identifier associated with the recording segment is the same as the user identifier of the corresponding initial recording segment, and the user identifier associated with the recording segment is the same as the screen content data corresponding to the endpoint.
[0050] The recording and viewing device based on multi-dimensional meeting information also includes:
[0051] The first operation receiving unit is used to enter the first viewing state in response to the first viewing operation after displaying the record viewing interface for viewing meeting minutes and receiving the first viewing operation.
[0052] The recording and viewing device based on multi-dimensional meeting information also includes:
[0053] The second operation receiving unit is used to, after displaying the record viewing interface for viewing meeting minutes, enter the second viewing state in response to the second viewing operation when a second viewing operation is received; the record viewing interface displays the timeline corresponding to the recording data in the second viewing state, and the timeline corresponds to the time segment of the recording segment;
[0054] The target display unit is used to display the sub-records corresponding to the target time segment on the record viewing interface when a trigger operation is received through the time axis and applied to the target time segment.
[0055] The timeline displays time segments corresponding to recording segments differently from time segments without corresponding recording segments.
[0056] Each sub-record includes corresponding sub-audio recording data, text generated from the sub-audio recording data, sub-video data generated from the video data, and associated screen content data.
[0057] When a sub-record is displayed, the corresponding text and screen content data are displayed in separate areas, while the corresponding sub-audio recording data and sub-video data are output synchronously.
[0058] Thirdly, embodiments of this application also provide an electronic device, which includes:
[0059] One or more processors;
[0060] Memory, used to store one or more computer programs;
[0061] When one or more computer programs are executed by one or more processors, electronic devices enable the recording and viewing method based on multidimensional conference information, as described in the first aspect.
[0062] Fourthly, embodiments of this application also provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the recording and viewing method based on multidimensional meeting information as described in the first aspect. Attached Figure Description
[0063] Figure 1 A flowchart illustrating a method for viewing records based on multidimensional meeting information, provided in this application embodiment.
[0064] Figure 2 This is a schematic diagram illustrating a process for processing audio recording data, as provided in an embodiment of this application.
[0065] Figure 3 This is a schematic diagram of a first interface for displaying meeting minutes provided in an embodiment of this application.
[0066] Figure 4 In order to be in Figure 3 A schematic diagram illustrating the interface changes when the primary target information is the user identifier.
[0067] Figure 5 In order to be in Figure 3 The first target information is a schematic diagram of interface changes based on the screen content data.
[0068] Figure 6 This is a schematic diagram of a second interface for displaying meeting minutes provided in an embodiment of this application.
[0069] Figure 7 This is a flowchart illustrating a method for generating meeting minutes in a method for viewing multi-dimensional meeting information provided in an embodiment of this application.
[0070] Figure 8 This is a schematic diagram of the structure of a recording and viewing device based on multidimensional conference information provided in an embodiment of this application.
[0071] Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0072] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are for illustrative purposes only and not for limiting the invention. Furthermore, it should be noted that, for ease of description, the accompanying drawings show only the parts relevant to the present invention and not the entire structure.
[0073] It should be noted that, due to space limitations, this application specification does not exhaustively list all possible implementation methods. Those skilled in the art should be able to conceive after reading this application specification that, as long as the technical features do not contradict each other, any combination of technical features can constitute an optional implementation method.
[0074] The embodiments of the present invention will be described in detail below.
[0075] In meeting scenarios, due to the need for a time-based meeting schedule and comprehensive recording of the process, raw meeting minutes are typically generated linearly from multimedia data (video and / or audio). This linearly generated raw data cannot be directly located to correspond to specific meeting events, hindering subsequent review of the meeting proceedings. Further processing, such as text recognition, behavior recognition, and voiceprint recognition, can be performed on the raw data to obtain derived information. This derived information can then be used to structurally break down the raw data for quick viewing. For example, meeting minutes can be generated chronologically, with each speaker's speech as a unit of recording. However, existing raw data and meeting minutes generated after deep processing still offer limited viewing options. Subsequent review can only be done through a progress bar or by selecting a speaker, resulting in coarse filtering. The granularity of structural breakdown is relatively large, leading to low precision in viewing and slow target location within the minutes.
[0076] To address the above technical issues, this application proposes a recording and viewing method based on multi-dimensional meeting information. Based on the characteristic that the content of speeches during the meeting is closely related to the focus of communication, the method records the switching process of the screen that serves as the focus of communication during the meeting. The recorded audio and video are then further subdivided according to the screen switching process, with each sub-record in the meeting minutes corresponding to a finer granularity. When viewing the meeting minutes, the method enables more precise and rapid target location based on the screen, improving the user's experience in viewing the meeting minutes.
[0077] Figure 1 This application provides a flowchart of a method for viewing meeting records based on multi-dimensional meeting information. This method is applicable to various electronic devices capable of processing meeting records and is implemented by electronic devices, such as various meeting devices specifically designed to support meeting development, personal computers with meeting applications installed, and servers that process data generated during meetings. Figure 1 As shown, the method for viewing records based on multidimensional meeting information includes, but is not limited to, steps S110-S150:
[0078] Step S110: Display the record viewing interface for viewing meeting minutes. The meeting minutes consist of sub-records. Each sub-record corresponds to the continuous speech of the same user ID on the same screen content data during the meeting. The user ID and screen content data corresponding to the associated record are used as two filtering information.
[0079] In this embodiment, the viewable meeting minutes are generated by processing data collected during the meeting. The meeting minutes are organized by consecutive remarks from the same participant on the same focus (i.e., the displayed screen content), with each unit corresponding to a sub-record. In a meeting, multiple participants typically speak randomly at any time regarding multiple sequentially displayed screen contents; therefore, the meeting minutes generated for a single meeting will have multiple sub-records corresponding to multiple units. Each sub-record is associated with the participant's user identifier and the screen content data corresponding to the screen content, serving as filtering information for accurate location and quick viewing of the meeting minutes.
[0080] The data collected during the meeting mainly includes audio recordings, which are then processed to produce meeting minutes. Please refer to [link / reference needed]. Figure 2 This diagram illustrates the process of processing a segment of audio data from a meeting. TA represents the mapping of audio changes on the timeline, and TB represents the mapping of switching moments on the timeline. The timeline contains eight recorded moments, T0-T7. These moments correspond to the following events that occurred during the meeting: At time T0, Tom opens the presentation page (a display state of on-screen content data) and begins speaking; Tom switches the presentation page twice, with a very short interval between moments T1 and T2, and continues speaking until moment T3; From moment T3, Tim continues speaking from the existing presentation page until moment T5, switching the presentation page at moment T4; From moment T6, Tom continues speaking from the existing presentation page until moment T7. Based on the definition of sub-records described above, this segment of audio data corresponds to multiple sub-records. These sub-records, in chronological order, correspond to: Tom's speech from time T0 to T1 (the content displayed on the screen starting at time T0), Tom's speech from time T2 to T3 (the content displayed on the screen starting at time T2), Tim's speech from time T3 to T4 (the content displayed on the screen starting at time T2), Tim's speech from time T4 to T5 (the content displayed on the screen starting at time T4), and Tom's speech from time T6 to T7 (the content displayed on the screen starting at time T4). This method of organizing and saving meeting records based on two dimensions of meeting information—the user who spoke and the content displayed on the screen—allows for quick filtering and viewing of sub-records based on one or more dimensions when reviewing the meeting records later, according to a basic impression of the target audience.
[0081] The interface for viewing meeting minutes is defined as a minutes viewing interface. This interface can be an application specifically designed to open the meeting minutes provided in this embodiment, or it can be an application function integrated into a supplementary meeting application (such as a meeting minutes application, video conferencing application, etc.). Users can open the minutes viewing interface from the corresponding entry point (such as the application icon or the application's function entry point) when they need to view the meeting minutes. Viewing meeting minutes can be done while they are being generated during the meeting, or it can be done after the meeting has ended. Generally, as usage time increases, the number of meeting minutes increases, and the minutes viewing interface allows viewing any single meeting minute. Multiple meeting minutes can be browsed and the target meeting minute to be viewed can be selected. The "displaying the minutes viewing interface for viewing meeting minutes" described in this embodiment specifically refers to the state where, after selecting the target meeting minute to view, the specific content of the target meeting minute can be viewed on the minutes viewing interface. The meeting minutes displayed in a minutes viewing interface can be pre-generated meeting minutes or meeting minutes received from other electronic devices.
[0082] Step S120: Receive a first trigger operation acting on the first target information in the first viewing state. The record viewing interface displays a first option bar and a second option bar in the first viewing state. Each first option label displayed in the first option bar is used to present each user identifier in the meeting record. Each second option label displayed in the second option bar is used to present each screen content data in the meeting record. The first target information is any first option label or second option label.
[0083] Step S130: In response to the first trigger operation, the first result bar and the third option bar corresponding to the first target information are displayed on the record viewing interface. The first result bar is used to display the record identifier of the first filtered record. The first filtered record includes all sub-records in the meeting record that are associated with the filtered information presented by the first target information. The third option bar displays each third option label for presenting another filtered information that is associated with the filtered information presented by the first target information in the same sub-record.
[0084] As mentioned earlier, for each sub-record in the generated meeting minutes, there are two associated filter information. For meeting minutes with multiple sub-records, the set of specific filter information in all sub-records can be categorized and displayed for user filtering. The first viewing state is the state where both types of filter information in a meeting minute are displayed in categories. The two types of filter information are displayed in different option bars, each option bar has multiple option labels, and the information presented in each option label corresponds to the filter information appearing in the meeting minutes. Specifically, each filter information can be displayed only once, and the option labels are used to receive user trigger operations. In this embodiment, the option bar that displays the user identifier is defined as the first option bar, and the option labels in the first option bar are defined as first option labels. Each first option label is used to present each user identifier in the meeting minutes; the option bar that displays the screen content data is defined as the second option bar, and the option labels in the second option bar are defined as second option labels. Each second option label is used to present each screen content data in the meeting minutes. The trigger operation received in the first viewing state for a certain option label (i.e., any first option label or second option label) is defined as the first trigger operation with the filter information in that option label as the first target information. In the first viewing state, you can confirm that any filter information associated with all sub-records in a meeting record has been selected.
[0085] It should be noted that, in the embodiments of this application, the terms "first," "second," and "third" in "first option bar," "second option bar," "first option label," "second option label," and "third option label" are only used to distinguish different types or individuals of descriptive objects based on the same technology. For example, "first option bar" and "second option bar" are sets of interactive controls implemented using the same technology, where "first" and "second" are used to distinguish that the types of information displayed by the two are different.
[0086] The first viewing state can be the default state, which improves the speed at which users can view meeting minutes. Alternatively, there can be a corresponding mode switching control, requiring a trigger operation to enter the first viewing state; that is, upon receiving a first viewing operation, the system responds and enters the first viewing state. Both of these methods require users to select a viewing mode before entering the first viewing state, allowing them to view meeting minutes based on filter information. This makes it easier for users to choose different viewing modes according to their needs, improving the targeted nature of their meeting minute viewing.
[0087] For each sub-record associated with two types of filtering information, the sub-records can be filtered twice. Each filtering step identifies the filtered records that match the target information within the corresponding filtering range. In other words, a filtered record refers to the set of filtered sub-records. Upon receiving the first trigger operation, all sub-records in the meeting minutes of a single meeting are filtered according to the first target information corresponding to the first trigger operation. The corresponding filtering result is defined as the first filtered record, which is the sub-record in the meeting minutes associated with the first target information. Each sub-record has a corresponding record identifier, used to present the basic information of the sub-record to the user and serve as the interactive entry point for the user to confirm viewing the sub-record. The record identifier can specifically be text, the first frame of sub-video data, and / or a simplified display of the screen content data. A detailed explanation follows with the accompanying diagram. The record identifiers of all sub-records in the first filtered record are displayed in the first result column.
[0088] After obtaining the first filtered record based on one type of filter information, you can further filter using the first filtered record as the filter scope, based on another type of filter information. The alternative filter information associated with the first filtered record is displayed in the third option bar, i.e., each alternative filter information is presented as a third option label in the third option bar. Based on the first filtering, the alternative filter information associated with the first filtered record may be less than the alternative filter information associated with all sub-records in the initial meeting record. In other words, the number of third option labels in the third option bar is less than the number of option labels used to display the same filter information in the first or second option bars. Clearly, in the first filtered record obtained after the first filtering, each sub-record still has two associated filter information, one of which is the first target information. This is equivalent to the alternative filter information, which is different from the first target information among the filter information associated with these sub-records, and is displayed in the third option bar.
[0089] It should be understood that in the first viewing state, there can also be a display area to show the record identifiers of all sub-records in a meeting record. Regardless of the user identifier displayed in any state, when a trigger operation is received targeting a specific record identifier, the sub-record corresponding to that record identifier can be displayed. That is, in the first viewing state, the user can flexibly select the viewing target from the first result bar and the second result bar to display the corresponding sub-record.
[0090] Step S140: Receive a second trigger operation that acts on the second target information, where the second target information is one of all the third option labels in the third option bar.
[0091] Step S150: In response to the second trigger operation, the second result column corresponding to the second target information is displayed on the record viewing interface. The second result column is used to display the record identifier of the second filtered record. The second filtered record includes all sub-records in the first filtered record that are associated with the second target information.
[0092] Overall, steps S120 and S130 complete the first filtering through a first trigger operation, and steps S140 and S150 complete the second filtering through a second trigger operation. This is equivalent to selecting a target information from two types of filtering information sequentially, with each selection corresponding to one filtering step. Finally, the sub-records that match both target information are filtered and displayed. There is no order in which the two target information selections are made. After the first filtering, the selectable sub-records are updated, and the other type of selectable filtering information is retained for display.
[0093] Next, we will further combine... Figures 3-5 The screening process in steps S140-S150 is described in detail. Figure 3 The example presents a schematic diagram of a first type of interface for displaying meeting minutes. In the specific description, the interface for displaying meeting minutes is defined as the minutes viewing interface 10. Figure 3 In the record viewing interface 10 shown, the first display area 111 displays a collection of screen content data for all sub-records in the meeting record, which is equivalent to displaying the first option bar. The second display area 112 displays a collection of user identifiers for all sub-records in the meeting record, which is equivalent to displaying the second option bar. Each screen content data and user identifier is displayed uniquely. The record viewing interface 10 also includes a third display area 113, which displays the user identifiers, text, and recording durations corresponding to several sub-records sequentially from top to bottom. Users can scroll up and down or use the mouse wheel in the first display area 111 to adjust the displayed sub-records.
[0094] exist Figure 3 Based on the record viewing interface 10 shown, the user can first select a user identifier (e.g., "TOM") in the second display area 112. At this time, "TOM" in the user identifier can be confirmed as the first target information. Then, the record identifiers of all sub-records (i.e., the first filtered records) corresponding to "TOM" are displayed in the first result column, and the filter information that is different from the first target information among the two filter information associated with each sub-record of the first filtered record is displayed in the third option column. When the filter information about the user identifier in the first filtered record is only the first target information, the filter information that is different from the first target information is exactly the same type corresponding to the screen content data. Figure 4 The image shows the first triggering operation in... Figure 3The record viewing interface 10, which is displayed after updating the screen based on the previous one, is in... Figure 4 In the record viewing interface 10 shown, the first target information (i.e., "TOM") selected according to the first trigger operation is highlighted in the second display area 112'. Only the screen content data associated with "TOM" (i.e., the third option bar) is displayed in the first display area 111', and only all sub-records associated with "TOM" are displayed in the third display area 113'. For example... Figure 4 The screen content data associated with "TOM" shown only includes two screen content data points, P1 and P4.
[0095] exist Figure 4 Based on the above, the user can further select a screen content data in the first display area 111' and receive a second trigger operation that acts on the second target information. The specific response strategy is to display the second result column corresponding to the second target information in the record viewing interface. There may be multiple record identifiers presented in the second result column, or there may be only one. Since both the first target information and the second target information are obtained from the range of existing sub-records, there will inevitably be a filtering result.
[0096] If there is only one sub-record in the second filter record, and only one record identifier is displayed in the second result column, then all or part of the data for the corresponding record can be output directly. For example, text and screen content data can be output through the display module, and sub-recording data can be output through the sound module. Since text and sub-recording data have a direct source relationship, both can be output or only one can be output. Annotations generated during the display of screen content data can be recorded and displayed in the original data.
[0097] If the second filter record contains multiple sub-records, and the corresponding second result column displays multiple record identifiers, then all or part of the data for each sub-record in the second filter record can be output sequentially. Alternatively, all or part of the data for the corresponding record can be output based on the user's selection. The specific interface for outputting sub-records can be referenced from the implementation of relevant meeting minutes output. The specific output process is not the focus of this solution and will not be detailed here.
[0098] Figure 5 The image shown is in Figure 3 Based on this, firstly, an operation is performed on selected screen content data (e.g., P2) in the first display area 111. The first trigger operation is received accordingly. At this time, P2 is the first target information. Then, all sub-records corresponding to P2 are displayed (i.e., the record identifiers of all sub-records in the first filter record are displayed in the first result column), as well as other related information about the user identifier in the first filter record (i.e., the user identifiers of all sub-records in the first filter record are displayed in the third option column). Figure 5The image shows the first triggering operation in... Figure 3 The record viewing interface 10, which is displayed after updating the screen based on the previous one, is in... Figure 5 In the record viewing interface 10 shown, the first target information (i.e., P2) selected according to the first trigger operation is highlighted in the first display area 111'', only the user identifier associated with P2 is displayed in the second display area 112'' (i.e., the user identifiers of all sub-records in the second filtered records are displayed in the second result column), and only all sub-records associated with P2 are displayed in the third display area 113''. For example Figure 5 The user identifiers associated with P2 shown are only "TOM" and "TIM". Further filtering operations after the first trigger operation can be found in the section on... Figure 3 and Figure 4 The explanation will not be repeated here.
[0099] It should be understood that Figures 3-5 The layout and interaction changes described are merely exemplary implementations within the overall design framework of this application's embodiments and do not represent an exclusive limitation on the overall interaction effect. For example... Figure 3 The content data in the middle screen is displayed in three parts at a time, with the middle part being larger than the two on either side. Preview switching is achieved by swiping left or right. Alternatively, it can be displayed as a uniformly sized array, with preview switching via swiping up or down; or it can be displayed in partial stacks, with the content data in the top section only fully displayed when switched to the top. User identifiers can be circular or square, and can be filled with user photos or avatars, with the user name displayed inside or outside the circle or square. This application does not limit the specific details of the interaction implementation.
[0100] Based on the characteristics of the close connection between the content of speeches and the focus of communication during the meeting, the switching process of the screen that is the focus of communication during the meeting is recorded. The audio and video recorded during the meeting are further broken down according to the switching process of the screen. Each sub-record in the meeting record corresponds to a finer granularity. When viewing the meeting record, the target position can be located more precisely and quickly according to the screen, thereby improving the user's viewing experience of the meeting record.
[0101] exist Figure 3Based on the above, a second viewing state can be further set according to the segmented structure. After generating meeting minutes corresponding to audio and video data based on the segmented structure, it can also include: upon receiving a second viewing operation, entering the second viewing state in response to the second viewing operation; displaying the timeline corresponding to the audio data in the second viewing state, with the timeline corresponding to the time segments of the audio recording; and displaying the sub-record corresponding to the target time segment when a trigger operation is received through the timeline. The first viewing state is the state in which both types of filtered information in a meeting minute are displayed separately. In the first viewing state, any filtered information can be selected. The second viewing state can have a corresponding mode switching control, which needs to be triggered to enter the second viewing state, that is, upon receiving a second viewing operation, entering the second viewing state in response to the second viewing operation.
[0102] The display of time segments corresponding to recording segments on the timeline can differ from or be the same for time segments without corresponding recording segments. Valid and invalid records are clearly displayed in the progress bar, effectively reminding users of priority targets in the meeting minutes and improving viewing efficiency. Selecting a viewing mode enters a second viewing state, where users can view sub-records of the meeting minutes through time segments on the timeline, allowing for quick and comprehensive identification of target locations.
[0103] In practice, the distinction between the first and second viewing states can be blurred. Instead, the information required for both states can be displayed simultaneously, effectively combining the first and second viewing states into a single, fixed viewing state. For example... Figure 6 The second type of meeting record display interface shown is in the overall viewing state. This record viewing interface 10 includes the first display area 111, the second display area 112, and the third display area 113 described above, as well as a timeline 114. Valid records in the timeline 114 are displayed as black segments. Valid records without intervals are separated by short white intervals, and invalid records are displayed in white corresponding to their duration. Each valid record's corresponding time segment can also display a corresponding user identifier nearby to facilitate user confirmation of the operation target. Receiving a trigger operation in any time segment of the timeline 114 makes that time segment the target time segment, subsequently displaying the relevant data for the corresponding sub-record. The specific display of sub-records can be referred to the display states described above, and will not be repeated here.
[0104] In another specific embodiment, the specific process of generating meeting minutes is described in detail. Please refer to... Figure 7 The specific process of generating this meeting record may include, but is not limited to, steps S210-S240:
[0105] Step S210: Acquire audio data, video data and screen display data during the meeting. The screen display data includes the switching time of each screen switch and the screen content data after the switch. The video data is obtained by capturing images of the meeting venue through a camera. The screen switch is the switching of the screen displayed on the electronic devices in the meeting venue.
[0106] In meetings conducted via electronic devices, the communication content is typically recorded and processed digitally using meeting applications installed on the devices. This communication content refers to the information generated or output by all relevant elements during the meeting, such as participants' speeches, on-site visuals, and images displayed on the electronic devices as the focus of communication. The electronic devices used at the meeting can be users' personal laptops or specialized equipment designed and developed to support the meeting, featuring a large display size and fixedly installed at the meeting venue. Examples include various large-screen devices named interactive whiteboards, conference TVs, or conference tablets.
[0107] In meetings conducted using electronic devices, participants' speeches are primarily recorded via microphones, while video is captured via cameras. The displayed content is obtained by directly recording the electronic device's display and interactive actions. In this embodiment, audio and video data can be acquired using relevant data acquisition methods. Regarding the displayed content, the inventors considered that participants' expressions during meetings are typically based on a specific communication focus, which has clear points of change corresponding to user actions. Therefore, the inventors can record and acquire the time of the switching operation that causes a change in the displayed content, i.e., the moment of the screen switch, and correspondingly record and acquire the displayed content data after the switch.
[0108] The original audio and video data are usually just linear records based on the timeline, providing only macro-level information such as total duration or data source, but lacking specific internal structured information. As mentioned earlier, the inventors discovered that the expressions of participants during a meeting are usually based on a certain communication focus, and the changes in the visual content presenting the communication focus have obvious structured characteristics. When the audio and video data, as well as the changes in the visual content of the communication focus, are all recorded based on the same timeline, the structured characteristics of the changes in the visual content can be bound to the audio and video data. Then, through steps S220-S240, based on the structure of the visual content, the original audio and video data are processed to obtain a finely structured meeting record.
[0109] In the specific processing, considering that the content that serves as the focus of communication can have various data formats, such as courseware specifically edited for meetings or various images prepared before the meeting begins, this application embodiment does not limit the specific data format. Data perceived visually can be statically displayed to the user, and all can be used as the focus of communication. The switching time and the content data of the screen after each screen transition are obtained accordingly. Specifically, this acquisition can be by directly extracting single-page data (such as pages from courseware and captured images), or by taking screenshots of the screen (such as text documents or continuous web pages). The corresponding switching time can be a detected page-turning operation or scroll wheel operation.
[0110] It should also be noted that a meeting based on electronic devices may be an offline meeting in which multiple participants use one electronic device in one scenario, or it may be an online meeting in which multiple electronic devices are used in multiple scenarios (there may be multiple participants in a single scenario). In specific processing, all data can be centralized and processed together, or each electronic device can process the data it directly obtains and then summarize it.
[0111] Step S220: Based on the video data and / or audio data, confirm the user identifier and speaking time period corresponding to each speech.
[0112] Having already collected the raw video and audio data, the audio data can be first broken down chronologically into units of one speech by the same speaker. This will yield at least the user identifier and speaking time for each speaker's speech. The user identifier identifies the speaker, and the speaking time is the start and end time of the speech. For example, in a 15-minute meeting, if Tom speaks for 2 minutes, Tim for 5 minutes, Tom for 1 minute, Jack for 6 minutes, and Tom for 1 minute, then there are a total of 5 speeches in this 15-minute meeting. The user identifier for each speech is the individual identifier of the corresponding user (there are 3 user identifiers in total).
[0113] Specifically, confirming the user identifier and speaking time period corresponding to each speech can be achieved through behavioral recognition of video data to identify which user spoke during a given time period, thus determining the corresponding user identifier and speaking time period. Even if multiple users are present in a single video stream, it is possible to identify which user spoke. Alternatively, it can be achieved through voiceprint recognition of audio data. Based on pre-recorded user voiceprint features corresponding to each user identifier, it identifies the corresponding user identifier and speaking time period when users with the same voiceprint features speak consecutively. Another approach is to comprehensively process video and audio data, making a combined judgment based on behavioral and voiceprint recognition, and accurately confirming the user identifier and speaking time period through comparison. For specific solutions in the field of image recognition for behavioral recognition and corresponding speakers, relevant image recognition solutions can be referenced. Voiceprint feature recognition can also adopt implementation solutions from related fields, which will not be elaborated upon here.
[0114] Based on step S220, a relatively coarse structure of the meeting has been identified with reference to the speaking order. Step S130 can then be used to further refine this coarse structure. During the meeting, each user may speak multiple times, and each speech may correspond to a different meeting focus. Based on the structure identified chronologically in step S220, all of a user's speeches can be roughly filtered out based on user identifiers during subsequent review.
[0115] Step S230: Confirm the segmentation structure of the recording data based on the speaking time period and the switching time. Each recording segment in the segmentation structure is associated with two filtering information, which includes a user identifier and a screen content data.
[0116] The segmentation structure of the recording data is determined based on the speaking time period and the switching time because the switching time does not necessarily correspond strictly to the start and end time of the speaking time period. The continuous speaking of the same user may correspond to multiple screen content data, that is, the continuous speaking of the same user may correspond to different content topics. At this time, the speaking time period can be further divided by the switching time to obtain the smallest granular structure of the recording data in this embodiment of the application, namely the recording segment. Based on the user identifier corresponding to the speaking time period and the screen content data confirmed by the switching time, the user identifier and screen content data of each recording segment can be confirmed.
[0117] Meetings requiring minutes are typically discussion-based or multi-person reporting types. These types of meetings are often characterized by a high degree of flexibility in information generation and the need to achieve desired outcomes, making minutes highly valuable. Furthermore, the amount of visual content data corresponding to the meeting's focus is relatively small. During presentations based on this focus, the visual content data generated can be used to represent the content of the speech. This visual content data presents the information intuitively, and the relatively small amount of data allows for quick viewing.
[0118] When confirming the segmentation structure, the initial segmentation structure of the recording data can be confirmed first based on the speaking time period. Each initial recording segment in each initial segmentation structure is associated with a user identifier and the duration reaches a first preset duration. Then, the initial recording segments are split according to the switching time to obtain recording segments. The recording segments are divided with reference to the end time of the display of the screen content data within the initial recording segment that has been continuously displayed for a duration of a second preset duration. The user identifier associated with the recording segment is the same as the user identifier of the corresponding initial recording segment, and the user identifier associated with the recording segment is the same as the screen content data corresponding to the endpoint.
[0119] The speaking time period is determined from the audio recording data and / or the video data recorded synchronously with the audio recording data. Accordingly, the timeline of the audio recording data can be segmented based on the speaking time period to determine the initial segmentation structure based on the speaking user. In this initial segmentation structure, each initial recording segment is associated with a user identifier, and the duration must reach a first preset duration. That is, only when the continuous speaking time of the same user reaches a certain length is it considered a valid speech and necessary to record. Accordingly, an initial recording segment can be obtained as a valid recording. The specific reference value for the "certain length" is defined as the first preset duration, such as 2 seconds, 3 seconds, etc. Based on the initial recording segments, further segmentation is performed according to the switching time to obtain recording segments. The recording segments are divided with reference to the end time of the display of the screen content data within the initial recording segment that has a continuous display duration reaching a second preset duration. The user identifier associated with the recording segment is the same as the user identifier of the corresponding initial recording segment, and the user identifier associated with the recording segment is the same as the screen content data corresponding to the endpoint. Filtering information is used after generating meeting minutes to filter and accurately display the target viewing area based on user selection.
[0120] Please refer to Figure 2TA represents the mapping of recording changes on the timeline, and TB represents the mapping of switching moments on the timeline. There are eight recorded moments on the timeline, from T0 to T7. The events that occurred during the meeting at these moments are as follows: At T0, Tom opens the presentation page (a display state of screen content data) and begins speaking; Tom switches the presentation page twice with a very short interval between T1 and T2 and continues speaking until T3; Tim continues speaking from the original presentation page from T3 until T5, and switches the presentation page at T4; Tom continues speaking from the original presentation page from T6 until T7. Based on the scheme described above, the data collected during this process confirms three initial recording segments: T0-T3, T3-T5, and T6-T7, with corresponding user identifiers of Tom, Tim, and Tom, respectively. The confirmed switching times include T0, T1, T2, and T4. According to the recording segmentation method and screen content data confirmation method described above, the recording segment T0-T1 corresponds to the user identifier Tom, and the screen content data is the courseware page displayed during T0-T1; the recording segment T1-T3 corresponds to the user identifier Tom, and the screen content data is the courseware page displayed during T2-T4; the recording segment T3-T4 corresponds to the user identifier Tim, and the screen content data is the courseware page displayed during T2-T4; the recording segment T4-T5 corresponds to the user identifier Tim, and the screen content data is the courseware page displayed after T4; the recording segment T6-T7 corresponds to the user identifier Tom, and the screen content data is the courseware page displayed after T4. Extracting valid recordings based on audio duration and screen dwell time eliminates the need for separate saving of invalid records, minimizing the number of sub-records corresponding to each meeting's minutes, thereby improving the efficiency of viewing meeting minutes.
[0121] Step S240: Generate meeting minutes corresponding to audio and video data based on the segmented structure. Each sub-record in the meeting minutes corresponds to an audio segment.
[0122] Based on the previously confirmed timeline-based segmentation structure, multiple sub-records are formed within the complete meeting minutes, each corresponding to an audio segment. After performing text recognition on the audio data, the text recognition results are split according to the segmentation structure, and the split text is mapped to the corresponding sub-record of the audio segment. Video data can be split according to the same segmentation structure within a unified timeline. Specifically, each sub-record includes corresponding sub-audio data, text generated from the sub-audio data, sub-video data generated from the video data, and associated on-screen content data. The diverse content of each sub-record helps users comprehensively obtain information about the meeting process. When displaying sub-records subsequently, the corresponding text and on-screen content data are displayed in separate areas, while the corresponding sub-audio and sub-video data are output synchronously. The synchronous display of the content corresponding to each sub-record effectively improves the convenience for users to view the meeting minutes.
[0123] It should be understood that "meeting minutes generated before the record viewing interface for viewing meeting minutes" means that a meeting minute must be generated locally or remotely before it can be viewed through the record viewing interface. For the generation and viewing of any two different meeting minutes, the generation and viewing of one meeting minute is not interfered with by the generation and viewing process of the other. For example, two meeting minutes might be generated as the meeting progresses, and then either meeting minute can be viewed as needed; moreover, each generated meeting minute can be viewed multiple times.
[0124] Figure 8 This is a schematic diagram of a recording and viewing device based on multi-dimensional meeting information, provided as an embodiment of this application. Figure 8 As shown, the recording and viewing device based on multi-dimensional meeting information includes an interface display unit 210, a first filtering unit 220, a first response unit 230, a second filtering unit 240, and a second response unit 250.
[0125] The interface display unit 210 is used to display a meeting record viewing interface. The meeting record consists of sub-records, each sub-record corresponding to the continuous speech of the same user ID on the same screen content data during the meeting, and the user ID and screen content data corresponding to the record are used as two filtering information. The first filtering unit 220 is used to receive a first trigger operation acting on the first target information in the first viewing state. The record viewing interface displays a first option bar and a second option bar in the first viewing state. The first option bar displays each first option label to present each user ID in the meeting record, and the second option bar displays each second option label to present each screen content data in the meeting record. The first target information is any first option label or second option label. The first response unit 230 is used to respond to the first trigger operation and, in the record... The record viewing interface displays a first result bar and a third option bar corresponding to the first target information. The first result bar displays the record identifier of the first filtered record. The first filtered record includes all sub-records in the meeting record that are associated with the filtered information presented by the first target information. The third option bar displays various third option labels that are associated with another filtered record that is associated with the filtered information presented by the first target information. The second filtering unit 240 is used to receive a second trigger operation applied to the second target information. The second target information is one of all the third option labels in the third option bar. The second response unit 250 is used to respond to the second trigger operation and display the second result bar corresponding to the second target information in the record viewing interface. The second result bar displays the record identifier of the second filtered record. The second filtered record includes all sub-records in the first filtered record that are associated with the second target information.
[0126] The recording and viewing device based on multi-dimensional meeting information also includes:
[0127] The data acquisition unit is used to acquire audio data, video data and screen display data during the meeting. The screen display data includes the switching time of each screen switch and the screen content data after the switch. The video data is obtained by capturing images of the meeting venue through a camera. The screen switch refers to the switching of the screen displayed on the electronic devices in the meeting venue.
[0128] The first splitting unit is used to determine the user identifier and speaking time period corresponding to each speech based on video data and / or audio data;
[0129] The second splitting unit is used to confirm the segmentation structure of the recording data based on the speaking time and the switching time. Each recording segment in the segmentation structure is associated with two filtering information, which includes a user identifier and a screen content data.
[0130] The sub-record generation unit is used to generate meeting records corresponding to audio and video data based on the segmented structure. Each sub-record in the meeting record corresponds to an audio segment.
[0131] The second splitting unit includes:
[0132] The first segmentation module is used to determine the initial segmentation structure of the recording data based on the speaking time period. Each initial recording segment in each initial segmentation structure is associated with a user identifier and the duration reaches a first preset duration.
[0133] The second segmentation module is used to split the initial recording segment according to the switching time to obtain the recording segment. The recording segment is divided with reference to the end time of the display of the screen content data within the initial recording segment that has been continuously displayed for a duration of a second preset duration. The user identifier associated with the recording segment is the same as the user identifier of the corresponding initial recording segment, and the user identifier associated with the recording segment is the same as the screen content data corresponding to the endpoint.
[0134] The recording and viewing device based on multi-dimensional meeting information also includes:
[0135] The first operation receiving unit is used to enter the first viewing state in response to the first viewing operation after displaying the record viewing interface for viewing meeting minutes and receiving the first viewing operation.
[0136] The recording and viewing device based on multi-dimensional meeting information also includes:
[0137] The second operation receiving unit is used to enter the second viewing state in response to a second viewing operation after displaying the record viewing interface for viewing meeting minutes and receiving a second viewing operation. In the second viewing state, the record viewing interface displays the timeline corresponding to the recording data, and the timeline corresponds to the time segment of the recording.
[0138] The target display unit is used to display the sub-records corresponding to the target time segment on the record viewing interface when a trigger operation is received through the time axis and applied to the target time segment.
[0139] The timeline displays time segments corresponding to recording segments differently from time segments without corresponding recording segments.
[0140] Each sub-record includes corresponding sub-audio recording data, text generated from the sub-audio recording data, sub-video data generated from the video data, and associated screen content data.
[0141] When a sub-record is displayed, the corresponding text and screen content data are displayed in separate areas, while the corresponding sub-audio recording data and sub-video data are output synchronously.
[0142] The recording and viewing device based on multidimensional meeting information provided in this application embodiment is included in an electronic device and can be used to execute the corresponding recording and viewing method based on multidimensional meeting information provided in the above embodiment, and has corresponding functions and beneficial effects.
[0143] It is worth noting that in the above embodiments of the recording and viewing device based on multi-dimensional meeting information, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the scope of protection of the present invention.
[0144] Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 9 As shown, the electronic device includes a processor 310 and a memory 320, and may also include an input device 330, an output device 340, and a communication device 350; the number of processors 310 in the electronic device may be one or more. Figure 9 Taking a processor 310 as an example; the processor 310, memory 320, input device 330, output device 340, and communication device 350 in the electronic device can be connected via a bus or other means. Figure 9 Taking the example of a connection between China and Israel via a bus.
[0145] The memory 320, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the multi-dimensional meeting information recording and viewing method in this embodiment. The processor 310 executes various functional applications and data processing of the electronic device by running the software programs, instructions, and modules stored in the memory 320, thereby realizing the aforementioned multi-dimensional meeting information recording and viewing method.
[0146] The memory 320 may primarily include a program storage area and a data storage area. The program storage area may store the operating system and at least one application program required for a given function; the data storage area may store data created based on the use of the electronic device. Furthermore, the memory 320 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some instances, the memory 320 may further include memory remotely located relative to the processor 310, which can be connected to the electronic device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0147] Input device 330 can be used to receive input digital or character information, and to generate signal inputs related to user settings and function control of the electronic device. Output device 340 may include display devices such as a display screen.
[0148] The aforementioned electronic device includes a recording and viewing device based on multidimensional conference information, which can be used to execute any recording and viewing method based on multidimensional conference information, and has corresponding functions and beneficial effects.
[0149] This application also provides a computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program performs relevant operations in the recording and viewing method based on multidimensional meeting information provided in any embodiment of this application, and has corresponding functions and beneficial effects.
[0150] Those skilled in the art will understand that embodiments of this application may be provided as methods, systems, or computer program products.
[0151] Therefore, this application may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application may take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, produce implementations of the flowchart... Figure 1 One or more processes and / or boxes Figure 1 The computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The functions specified in one or more boxes. These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable apparatus for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0152] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory. Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0153] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0154] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0155] Note that the above description is merely a preferred embodiment of the present invention and the technical principles employed. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments, and may include many other equivalent embodiments without departing from the concept of the present invention, the scope of which is determined by the scope of the appended claims.
Claims
1. A method for viewing records based on multidimensional meeting information, characterized in that, include: The meeting record viewing interface is displayed. The meeting record consists of sub-records. Each sub-record corresponds to the continuous speech of the same user ID on the same screen content data during the meeting. The user ID and screen content data corresponding to the associated record are used as two filtering information. In the first viewing state, a first trigger operation is received that acts on the first target information. The record viewing interface displays a first option bar and a second option bar in the first viewing state. Each first option label displayed in the first option bar is used to present each user identifier in the meeting record. Each second option label displayed in the second option bar is used to present each screen content data in the meeting record. The first target information is any one of the first option label or the second option label. In response to the first trigger operation, a first result bar and a third option bar corresponding to the first target information are displayed on the record viewing interface. The first result bar is used to display the record identifier of the first filtered record. The first filtered record includes all sub-records in the meeting record that are associated with the filtered information presented by the first target information. The third option bar is used to display a third option label. The third option label is used to present another filtered information that is associated with the filtered information presented by the first target information in the same sub-record. Receive a second trigger operation that acts on the second target information, wherein the second target information is one of all the third option tabs in the third option bar; In response to the second triggering operation, a second result column corresponding to the second target information is displayed on the record viewing interface. The second result column is used to display the record identifier of the second filtered record. The second filtered record includes all sub-records in the first filtered record that are associated with the second target information.
2. The method for viewing records based on multi-dimensional meeting information according to claim 1, characterized in that, Before displaying the record viewing interface for viewing meeting minutes, the system also includes: The system acquires audio recording data, video data, and screen display data during the meeting. The screen display data includes the switching time of each screen switch and the screen content data after the switch. The video data is obtained by capturing images of the meeting venue through a camera. The screen switch refers to the switching of the display screen of the electronic devices in the meeting venue. Based on the video data and / or audio data, confirm the user identifier and speaking time period corresponding to each speech; The segmentation structure of the recording data is determined based on the speaking time period and the switching time. Each recording segment in the segmentation structure is associated with two filtering information, which includes a user identifier and a screen content data. Based on the segmentation structure, meeting records corresponding to the audio and video data are generated, and each sub-record in the meeting record corresponds to one of the audio segments.
3. The method for viewing records based on multi-dimensional meeting information according to claim 2, characterized in that, The segmentation structure of the recording data is determined based on the speaking time period and the switching time. Each recording segment in the segmentation structure is associated with two filtering information, which includes a user identifier and screen content data, including: The initial segmentation structure of the recording data is determined based on the speaking time period. Each initial recording segment in each initial segmentation structure is associated with a user identifier, and the duration reaches a first preset duration. The initial recording segment is split according to the switching time to obtain a recording segment. The recording segment is divided with reference to the end time of the display of the screen content data that has been continuously displayed for a second preset duration within the initial recording segment. The user identifier associated with the recording segment is the same as the user identifier of the corresponding initial recording segment, and the user identifier associated with the recording segment is the same as the screen content data corresponding to the endpoint.
4. The method for viewing records based on multidimensional meeting information according to any one of claims 2-3, characterized in that, After displaying the record viewing interface for viewing meeting minutes, it also includes: Upon receiving a first view operation, the system enters the first view state in response to the first view operation.
5. The method for viewing records based on multidimensional meeting information according to claim 4, characterized in that, After displaying the record viewing interface for viewing meeting minutes, it also includes: Upon receiving a second viewing operation, the system enters a second viewing state in response to the second viewing operation; the recording viewing interface displays the timeline corresponding to the recording data in the second viewing state, and the timeline corresponds to the recording segment display time segment; When a trigger operation is received on a target time segment via the time axis, the sub-record corresponding to the target time segment is displayed on the record viewing interface.
6. The method for viewing records based on multi-dimensional meeting information according to claim 5, characterized in that, The timeline is displayed differently for time segments corresponding to the recording segments and for time segments without corresponding recording segments.
7. The method for viewing records based on multidimensional meeting information according to any one of claims 2-3, characterized in that, Each sub-record includes corresponding sub-audio recording data, text generated based on the sub-audio recording data, sub-video data generated based on the video data, and associated screen content data.
8. The method for viewing records based on multidimensional meeting information according to claim 7, characterized in that, When the sub-record is displayed, the corresponding text and screen content data are displayed in separate areas, and the corresponding sub-audio recording data and sub-video data are output synchronously.
9. A recording and viewing device based on multi-dimensional meeting information, characterized in that, include: The interface display unit is used to display a record viewing interface for viewing meeting records. The meeting records are composed of sub-records. Each sub-record corresponds to the continuous speech of the same user ID to the same screen content data during the meeting. The user ID and screen content data corresponding to the associated record are used as two filtering information. The first filtering unit is used to receive a first trigger operation acting on the first target information in the first viewing state. The record viewing interface displays a first option bar and a second option bar in the first viewing state. Each first option label displayed in the first option bar is used to present each user identifier in the meeting record. Each second option label displayed in the second option bar is used to present each screen content data in the meeting record. The first target information is any one of the first option label or the second option label. A first response unit is used to respond to the first trigger operation and display a first result bar and a third option bar corresponding to the first target information on the record viewing interface. The first result bar is used to display the record identifier of the first filtered record. The first filtered record includes all sub-records in the meeting record that are associated with the filtered information presented by the first target information. Each third option label displayed in the third option bar is used to present another filtered information that is associated with the filtered information presented by the first target information in the same sub-record. The second filtering unit is used to receive a second trigger operation applied to the second target information, wherein the second target information is one of all the third option labels in the third option bar; The second response unit is used to respond to the second trigger operation and display the second result column corresponding to the second target information on the record viewing interface. The second result column is used to display the record identifier of the second filtered record. The second filtered record includes all sub-records in the first filtered record that are associated with the second target information.
10. An electronic device, characterized in that, include: One or more processors; Memory, used to store one or more computer programs; When the one or more computer programs are executed by the one or more processors, the electronic device enables the recording and viewing method based on multidimensional conference information as described in any one of claims 1-8.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the recording and viewing method based on multidimensional conference information as described in any one of claims 1-8.
Citation Information
Patent Citations
Conference record consulting method and device, computer equipment and storage medium
CN112839195A
Conference recording method, terminal equipment and conference recording system
CN116193179A