Video conference record generation method and computer readable medium

TW202630632AActive Publication Date: 2026-07-16ASUSTEK COMPUTER INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
TW · TW
Patent Type
Applications
Current Assignee / Owner
ASUSTEK COMPUTER INC
Filing Date
2025-01-06
Publication Date
2026-07-16

AI Technical Summary

Technical Problem

Online meetings lack the ability to generate physical interaction, making it difficult to build deep relationships with other participants and accurately judge their preferences and needs, leading to communication inefficiencies and reduced success rates.

Method used

A method for generating video conference transcripts that includes acquiring audio and image data segments, identifying participants, integrating them into text, and adding emotion tags to create a conference record, using a computer-readable recording medium to execute these steps.

Benefits of technology

Enables accurate judgment of participants' preferences and needs, enhancing communication efficiency and success rates by providing emotion tags in the meeting record.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure TWG2TA001067780_001
    Figure TWG2TA001067780_001
  • Figure TWG2TA001067780_002
    Figure TWG2TA001067780_002
  • Figure TWG2TA001067780_003
    Figure TWG2TA001067780_003
Patent Text Reader

Abstract

A video conference record generation method adapted to an electronic device suitable for executing a video conference which allows a plurality of participants to join is provided. The video conference record generation method comprises the following steps. Firstly, a plurality of audio clips is received. Then, the participant corresponds to each of the audio clips is identified. Then, the participant video clip corresponds to each of the audio clips is received and a plurality of emotion marks corresponding to the plurality of video clips respectively is generated. Afterward, the audio clips are integrated and converted into a conference document. Thereafter, the emotion marks are labelled at the corresponding sections of the conference document to generate a conference record. A computer readable medium containing a program for generating the conference record is also provided.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This case relates to the technical field of video conferencing, and in particular to a method for generating video conferencing minutes suitable for generating meeting minutes and a computer-readable recording medium. [Previous Technology]

[0002] The main drawback of online meetings is that they cannot generate physical interaction, it is not easy to build in-depth relationships with other participants, and it is not easy to accurately judge the preferences and needs of other participants (especially important people), which may lead to errors and affect communication efficiency and success rate. [Summary of the Invention]

[0003] This application provides a method for generating video conference records, applicable to an electronic device suitable for executing a video conference, and suitable for providing multiple participants. This method for generating video conference records includes the following steps: First, acquiring multiple audio data segments generated during the video conference. Then, identifying the participants corresponding to each audio data segment. Next, acquiring an image data segment of one participant corresponding to each audio data segment. Next, integrating these audio data segments and converting them into a conference text. Then, generating multiple emotion tags corresponding to each participant's image data segment based on these participant image data segments. Finally, marking the corresponding multiple emotion tags on the conference text to generate a conference record.

[0004] This application also provides a computer-readable recording medium with a built-in program. When the computer loads and executes this program, the following steps can be completed: First, acquire multiple audio data segments generated during the video conference. Then, identify the participants corresponding to each audio data segment. Next, acquire one video data segment of a participant corresponding to each audio data segment. Then, integrate these audio data segments and convert them into a conference text. Then, generate multiple emotion tags corresponding to each participant's video data segment based on these participant video data segments. Finally, mark the corresponding multiple emotion tags on the conference text to generate a conference record.

[0005] The video conference record generation method provided in this case can generate emotion tags corresponding to the text data fragments of each participant, and then merge the meeting text data and these emotion tags into meeting records. This helps users to accurately judge the preferences and needs of other participants and avoid errors that may affect communication efficiency and success rate.

Implementation Method

[0006] The specific embodiments of this case will be described in more detail below with reference to the schematic diagrams. The advantages and features of this case will become clearer based on the following description and the scope of the patent application. It should be noted that the drawings are all in a very simplified form and use non-precise proportions, and are only used to facilitate and clarify the illustration of the embodiments of this case.

[0007] The first figure is a schematic diagram of a video conference recording device 100 provided according to an embodiment of this case. This video conference recording device 100 is applicable to a video conferencing system 10. This video conferencing system 10 includes a server 11 and a plurality of terminal devices 12a, 12b, 12c (only three terminal devices 12a, 12b, 12c are shown in the figure for illustrative purposes). These terminal devices 12a, 12b, 12c are connected to the server 11 via the Internet W, and each participant connects to the server 11 via the terminal devices 12a, 12b, 12c to participate in the video conference.

[0008] The video conference recording device 100 is located on the server 11 and includes an audio data acquisition unit 110, an audio data analysis unit 120, an image data acquisition unit 130, an emotion image generation unit 140, a conference text generation unit 150, and a conference record generation unit 160.

[0009] The audio data acquisition unit 110 is adapted to acquire multiple audio data segments A1, A2, A3 generated during a video conference. In one embodiment, the audio data acquisition unit 110 can acquire audio data segments A1, A2, A3 from each audio data source (i.e., terminal devices 12a, 12b, 12c) according to the chronological order of the conference. In the embodiment shown in the first figure, each terminal device 12a, 12b, 12c generates audio data segments A1, A2, A3 (i.e., the participants corresponding to each terminal device 12a, 12b, 12c all speak), but this embodiment is not limited to this. In practical use, some terminal devices 12a, 12b, 12c may not generate audio data segments A1, A2, A3.

[0010] The audio data analysis unit 120 is adapted to identify the participant corresponding to each audio data segment A1, A2, A3. In one embodiment, the audio data analysis unit 120 can identify the participant corresponding to each audio data segment A1, A2, A3 based on the data source. Specifically, audio data segments A1, A2, A3 from the same terminal devices 12a, 12b, 12c are identified as belonging to the same participant.

[0011] The image data acquisition unit 130 is adapted to acquire the participant image data segments Vd1, Vd2, Vd3 corresponding to each audio data segment A1, A2, A3. In one embodiment, the image data acquisition unit 130 may first record the video of the video conference during the video conference, and then acquire the participant image data segments Vd1, Vd2, Vd3 corresponding to each audio data segment A1, A2, A3 according to the participants and time periods corresponding to each audio data segment A1, A2, A3.

[0012] The emotion image generation unit 140 is adapted to generate a plurality of emotion markers Em1, Em2, Em3 corresponding to each participant's image data fragment Vd1, Vd2, Vd3 based on these participant image data fragments Vd1, Vd2, Vd3. In one embodiment, these emotion markers Em1, Em2, Em3 are emojis, but this invention is not limited to this.

[0013] In one embodiment, the emotion image generation unit 140 includes an emotion classification model 142, which is adapted to analyze the facial features of participants and output emotion classification data (such as anger, happiness, etc.). This emotion classification model 142 is a trained deep learning classification model.

[0014] The conference text generation unit 150 is adapted to integrate these audio data segments A1, A2, A3 and convert them into a conference text D1. In one embodiment, the conference text generation unit 150 includes a speech-to-text model 152. The conference text generation unit 150 can integrate multiple audio data segments A1, A2, A3 generated during a video conference into a single audio record in chronological order, and convert this audio record into conference text D1.

[0015] The meeting record generation unit 160 is adapted to mark the meeting text D1 with corresponding plural emotion markers Em1, Em2, Em3 to generate a meeting record D2.

[0016] In one embodiment, the video conference recording device 100 also includes a summary extraction unit 170. The summary extraction unit 170 is electrically coupled to the conference text generation unit 150 and is adapted to extract a summary from the conference text D1 to generate a summary text D3. The summary extraction unit 170 can analyze the conference text D1 using either an extractive summarization method or an abstractive summarization method to generate the summary text D3.

[0017] In this embodiment, the video conference recording device 100 is located on the server 11 to generate conference recording D2. However, this embodiment is not limited to this. In other embodiments, the video conference recording device 100 may also be located on terminal devices 12a, 12b, 12c. Furthermore, in one embodiment, the video conference recording device 100 may be a software program stored on a non-transitory computer-readable recording medium.

[0018] Please also refer to Figure 2, which is a flowchart of the video conference recording generation method provided according to the first embodiment of this case. This video conference recording generation method is applicable to an electronic device suitable for executing a video conference. This video conference is suitable for providing multiple participants to join. This electronic device can be the server 11 in Figure 1, or the terminal devices 12a, 12b, 12c in Figure 1.

[0019] As shown in the figure, this method for generating video conference records includes the following steps.

[0020] First, as described in step S210, multiple audio data segments A1, A2, A3 generated during the video conference are acquired. In one embodiment, this step involves acquiring audio data segments A1, A2, A3 from each audio data source in chronological order of the conference proceedings.

[0021] Subsequently, as described in step S220, the participants corresponding to each audio data segment A1, A2, A3 are identified.

[0022] Next, as described in step S230, the participant image data segments Vd1, Vd2, Vd3 corresponding to one of the audio data segments A1, A2, A3 are obtained.

[0023] Then, as described in step S240, multiple emotion markers Em1, Em2, Em3 corresponding to each participant's image data segment Vd1, Vd2, Vd3 are generated based on these participant image data segments Vd1, Vd2, Vd3. In one embodiment, these emotion markers Em1, Em2, Em3 can be emojis. In another embodiment, these emotion markers Em1, Em2, Em3 can each correspond to a different emotion color line.

[0024] Then, as described in step S250, these audio data segments A1, A2, A3 are integrated and converted into a conference text D1. In one embodiment, this step involves integrating the multiple audio data segments A1, A2, A3 generated during the video conference into a single audio record in chronological order, and converting this audio record into a conference text D1 using a speech-to-text model.

[0025] Subsequently, as described in step S260, corresponding plural emotion markers Em1, Em2, Em3 are marked on the meeting text D1 to generate a meeting record D2.

[0026] Please also refer to the third figure, which shows one embodiment of step S220 in the second figure.

[0027] First, as described in step S320, analyze the voiceprint features corresponding to each audio data segment A1, A2, A3.

[0028] Then, as described in step S340, the plurality of audio data segments A1, A2, A3 are classified according to voiceprint features to identify the participants corresponding to each audio data segment A1, A2, A3.

[0029] In one embodiment, the aforementioned step S320 can analyze parameters such as speech rate, tone, and sound quality (i.e., voiceprint features) of each audio data segment A1, A2, A3 to confirm the voiceprint features corresponding to each audio data segment A1, A2, A3. Then, voiceprint comparison is used to identify which audio data segments A1, A2, A3 belong to the same participant. Essentially, it involves comparing the degree of difference in voiceprint data to distinguish different participants in the video conference.

[0030] Please also refer to Figure 4, which shows another embodiment of step S220 in Figure 2.

[0031] First, as described in step S420, the source of each audio data segment A1, A2, A3 is confirmed. Step S420 is to confirm which terminal device 12a, 12b, 12c (participant) each audio data segment in the video conference comes from.

[0032] Then, as described in step S440, the participants corresponding to each audio data segment A1, A2, A3 are identified based on the data source. Specifically, audio data segments A1, A2, A3 from the same audio data source (i.e., terminal devices 12a, 12b, 12c or participants) are identified as belonging to the same participant.

[0033] Please also refer to Figure 5, which shows one embodiment of step S240 in Figure 2.

[0034] First, as described in step S520, facial features of the participants are extracted from the participants' image segments to generate facial feature data.

[0035] Subsequently, as described in step S540, an emotion classification model is used to generate emotion classification data based on facial feature data. In one embodiment, this emotion classification model is a trained deep learning classification model.

[0036] Next, as described in step S560, emotion labels Em1, Em2, and Em3 are generated based on the emotion classification data.

[0037] In one embodiment, to simplify computation time, the aforementioned emotion classification model has a plurality of preset emotion types. Step S540 classifies facial feature data according to these preset emotion types to generate emotion classification data (i.e., which preset emotion type it belongs to). In addition, each preset emotion type is pre-set with corresponding emotion labels Em1, Em2, Em3, and step S560 directly outputs the corresponding emotion labels Em1, Em2, Em3 based on the emotion classification data.

[0038] Furthermore, in other embodiments, the aforementioned emotion classification model 142 may have a plurality of preset emotion types. Step S540 is to classify facial feature data according to these preset emotion types to generate emotion classification data (i.e., which preset emotion type it belongs to). In addition, each preset emotion type is pre-set with a corresponding emotion color line. This emotion color line can be used to annotate or display the corresponding text in the meeting text D1.

[0039] Please also refer to Figure 6, which is a flowchart of the video conference recording generation method provided according to the second embodiment of this case. As shown in the figure, this video conference recording generation method includes the following steps.

[0040] First, as described in step S610, acquire multiple audio data segments A1, A2, A3 generated during the video conference.

[0041] Subsequently, as described in step S620, the participants corresponding to each audio data segment A1, A2, A3 are identified.

[0042] Next, as described in step S630, the video data segments Vd1, Vd2, and Vd3 corresponding to one of the audio data segments A1, A2, and A3 are obtained.

[0043] Then, as described in step S640, multiple emotion markers Em1, Em2, Em3 corresponding to each participant's image data fragment Vd1, Vd2, Vd3 are generated based on these participant image data fragments Vd1, Vd2, Vd3.

[0044] Then, as described in step S650, these audio data segments A1, A2, A3 are integrated and converted into a conference text D1.

[0045] Subsequently, as described in step S660, corresponding plural emotion markers Em1, Em2, Em3 are marked on the meeting text D1 to generate a meeting record D2.

[0046] The aforementioned steps S610 to S660 are similar to steps S210 to S260 in the second figure, and will not be described in detail here.

[0047] Next, as described in step S670, a summary is extracted from the conference text D1 to generate a summary text D3. Step S670 can be performed by analyzing the conference text D1 using either extractive summarization or abstractactive summarization to generate the summary text D3.

[0048] Next, as described in step S680, a user input instruction is obtained.

[0049] Subsequently, as described in step S690, a summary text fragment is selected in summary text D3 in response to user input instructions.

[0050] Then, as described in step S695, the corresponding meeting record paragraph in meeting record D2 is identified and presented based on the summary text fragment.

[0051] In one embodiment, during step S670, when extracting the summary from the meeting text D1, the source information of the text in the summary text D3 is retained. Subsequently, in step S695, the corresponding meeting record paragraph in the meeting record D2 can be identified based on this source information.

[0052] Please also refer to Figure 7, which is a flowchart of the video conference recording generation method provided according to the third embodiment of this case. As shown in the figure, this video conference recording generation method includes the following steps.

[0053] First, as described in step S710, acquire multiple audio data segments A1, A2, A3 generated during the video conference.

[0054] Subsequently, as described in step S720, the participants corresponding to each audio data segment A1, A2, A3 are identified.

[0055] Next, as described in step S730, the participant image data segments Vd1, Vd2, Vd3 corresponding to one of the audio data segments A1, A2, A3 are obtained.

[0056] Then, as described in step S740, multiple emotion markers Em1, Em2, Em3 corresponding to each participant's image data fragment Vd1, Vd2, Vd3 are generated based on these participant image data fragments Vd1, Vd2, Vd3.

[0057] Then, as described in step S750, these audio data segments A1, A2, A3 are integrated and converted into a conference text D1.

[0058] Subsequently, as described in step S760, the corresponding plural emotion markers Em1, Em2, Em3 are marked on the meeting text D1 to generate meeting record D2.

[0059] The aforementioned steps S710 to S760 are similar to steps S210 to S260 in the second figure, and will not be described in detail here.

[0060] Compared with the embodiment in the second figure, this embodiment further includes step S770 after generating meeting record D2, analyzing meeting text D1 according to a preset principle, and launching a preset application to execute a preset task when meeting text D1 meets the preset principle.

[0061] In one embodiment, the preset principle may be that the meeting text D1 contains descriptive text such as location, time, and product. The preset application may be a map application, a calendar application, or an image browsing application. For example, when the meeting text D1 contains descriptive text about a location, the map application will be launched, and the location described in the meeting text will be marked on the map application (preset task). When the meeting text D1 contains descriptive text about a product, the image browsing application will be launched, and the corresponding product image will be opened.

[0062] Please also refer to Figure 8, which is a flowchart of the video conference recording generation method provided in the fourth embodiment of this case.

[0063] First, as described in step S810, acquire multiple audio data segments A1, A2, A3 generated during the video conference.

[0064] Subsequently, as described in step S820, the participants corresponding to each audio data segment A1, A2, A3 are identified.

[0065] Next, as described in step S830, the participant image data segments Vd1, Vd2, Vd3 corresponding to one of the audio data segments A1, A2, A3 are obtained.

[0066] Then, as described in step S840, multiple emotion markers Em1, Em2, Em3 corresponding to each participant's image data fragment Vd1, Vd2, Vd3 are generated based on these participant image data fragments Vd1, Vd2, Vd3.

[0067] Then, as described in step S850, these audio data segments A1, A2, A3 are integrated and converted into a conference text D1.

[0068] The aforementioned steps S810 to S850 are similar to steps S210 to S250 in the second figure, and will not be described in detail here.

[0069] Next, as described in step S860, a user input instruction is obtained.

[0070] Subsequently, as described in step S870, a meeting text segment is extracted from the meeting text D1 in response to the user's input command.

[0071] Then, as described in step S880, the corresponding emotion image is output based on the meeting text fragment.

[0072] In one embodiment, step S880 may first identify the audio data segments A1, A2, A3 corresponding to the meeting text segments, and then output the corresponding emotion images (i.e., emotion markers Em1, Em2, Em3). However, this embodiment is not limited to this. In other embodiments, a text-to-image generation model may also be used to generate corresponding images as emotion images based on the meeting text segments.

[0073] This application also provides a non-transitory computer-readable recording medium with a built-in program, suitable for video conferencing to generate video conference records. When the computer loads and executes this program, it can perform the following actions: First, acquire multiple audio data segments A1, A2, A3 generated during the video conference. Then, identify the participants corresponding to each audio data segment A1, A2, A3. Next, acquire video data segments Vd1, Vd2, Vd3 corresponding to one of the participants in each audio data segment A1, A2, A3. Then, generate multiple emotion tags Em1, Em2, Em3 corresponding to each participant's video data segment Vd1, Vd2, Vd3 based on these participant video data segments Vd1, Vd2, Vd3. Then, integrate these audio data segments A1, A2, A3 and convert them into a conference text D1. Next, mark the corresponding multiple emotion tags Em1, Em2, Em3 on the conference text D1 to generate a conference record D2.

[0074] The video conference record generation method provided in this case can generate emotion markers Em1, Em2, and Em3 corresponding to the text data fragments of each participant. The meeting text D1 data and these emotion markers Em1, Em2, and Em3 are then merged into meeting record D2, which helps users accurately judge the preferences and needs of other participants (especially important people) and avoid errors that could affect communication efficiency and success rate.

[0075] The above is merely a preferred embodiment of this case and does not limit the scope of this case in any way. Any equivalent substitution or modification made by a person skilled in the art to the technical means and technical content disclosed in this case without departing from the scope of the technical means of this case shall be deemed as not departing from the technical means of this case and shall still fall within the protection scope of this case. [Simplified Explanation of the Diagram]

[0076] The first figure is a schematic diagram of a video conference recording generation device provided according to an embodiment of the present invention; the second figure is a flowchart of a video conference recording generation method provided according to the first embodiment of the present invention; the third figure shows one embodiment of step S220 in the second figure; the fourth figure shows another embodiment of step S220 in the second figure; the fifth figure shows one embodiment of step S240 in the second figure; the sixth figure is a flowchart of a video conference recording generation method provided according to the second embodiment of the present invention; the seventh figure is a flowchart of a video conference recording generation method provided according to the third embodiment of the present invention; and the eighth figure is a flowchart of a video conference recording generation method provided according to the fourth embodiment of the present invention.

Claims

1. A method for generating video conference recordings, applicable to an electronic device, the electronic device being adapted to execute a video conference, and the video conference being adapted to allow multiple participants to join, the method comprising: acquiring multiple audio data segments generated during the video conference; identifying the participants corresponding to each audio data segment; acquiring an image data segment of one of the participants corresponding to each audio data segment; generating multiple emotion markers corresponding to each participant's image data segment based on the participant's image data segment; integrating the audio data segments and converting them into a conference text; marking the conference text with the corresponding multiple emotion markers to generate a conference record; generating multiple emotion color lines corresponding to each audio data segment based on the participant's image data segment; and marking the text in the conference text with the corresponding emotion color lines.

2. The method for generating video conference records as described in claim 1, wherein, The steps for identifying the participants corresponding to each audio data segment include: analyzing the voiceprint features corresponding to each audio data segment; and classifying the plurality of audio data segments based on the voiceprint features to identify the participants corresponding to each audio data segment.

3. The method for generating video conference records as described in claim 1, wherein, The steps of identifying the participants corresponding to each audio data segment include: identifying the data source of each audio data segment; and identifying the participants corresponding to each audio data segment based on the data source.

4. The method for generating video conference records as described in claim 1, wherein, This emotion marker is an emoji.

5. The video conference recording generation method as described in claim 1 further includes: extracting a summary from the conference text to generate a summary text.

6. The video conference recording generation method as described in claim 1 further includes: analyzing the conference text according to a preset principle, and launching a preset application when the conference text satisfies the preset principle.

7. The method for generating video conference records as described in claim 6, wherein, This default application is a map application.

8. The method for generating video conference records as described in claim 6, wherein, This default application is a calendar application.

9. The method for generating video conference records as described in claim 6, wherein, This default application is an image browsing application.

10. The video conference recording generation method as described in claim 1 further includes: obtaining a user input instruction; extracting a conference text segment from the conference text in response to the user input instruction; and generating an image based on the conference text segment using a text-to-image generation model.

11. The method for generating video conference records as described in claim 1, wherein, The step of generating multiple emotion tags corresponding to each audio data segment based on the video data segments of the participants includes: extracting the facial features of the participants from the video data segments to generate facial feature data; generating emotion classification data based on the facial feature data using an emotion classification model; and generating the emotion tag based on the emotion classification data.

12. A computer-readable recording medium containing a program, which, when loaded and executed by a computer, can perform the video conference recording generation method as described in claim 1.