Method for recording video conference and video conference system thereof

By using portrait recognition and voice processing algorithms in the video conferencing system, the speakers and their voice content in the video conferencing system are automatically recognized and recorded, and the problem of difficulty in automatically recording the conference content in the prior art is solved, real-time and accurate conference record generation is achieved.

CN120017784APending Publication Date: 2025-05-16MERRY ELECTRONICS (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311512855.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-14
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Existing video conferencing systems have difficulty recording the content of the meeting automatically, especially the problem of converting the voices of different speakers into text content and being associated with the speakers.

Method used

By using portrait recognition algorithms and voice processing algorithms in the video conferencing system, the images and sounds of participants are identified, the sound sources are matched, and the sounds are converted into text contents, and the meeting content is displayed and recorded in real time.

Benefits of technology

It realizes automatic recording of video conference content, able to identify speakers in real time and convert their voices into text content, providing a complete and accurate conference record.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120017784A_ABST
    Figure CN120017784A_ABST
Patent Text Reader

Abstract

The invention relates to a method for recording a video conference and a video conference system thereof. The method of recording a video conference includes: providing a user interface to a display device, wherein the user interface includes a first region, a second region, and a timeline; in response to images corresponding to a plurality of conventioneers obtained from the video signal through a portrait recognition algorithm, displaying the image of each conventioneer to a first area; in response to the voice processing algorithm, converting the audio clip of one of the conventioneers obtained from the audio signal into the text content, associating the text content with the corresponding one of the conventioneers, and displaying the text content to the second area based on the speaking sequence; and adjusting the time length of the time shaft along with the recording time of the video conference.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to a video conferencing system and a method for using the same, and in particular to a method for recording a video conference and a video conferencing system thereof. Background Art

[0002] With the development of the internet, the use of online conferencing software has increased significantly, allowing people to conduct video conferences with remote users without having to travel. To avoid forgetting what was discussed after a video conference, manual transcripts are often used, which is time-consuming. Furthermore, transcripts make it difficult to identify the speakers, making it impossible to determine which participants made which remarks upon subsequent viewing. Summary of the Invention

[0003] Based on this, it is necessary to provide a method for recording video conferences and a video conference system thereof, which can automatically convert the voices of different speakers into text content for recording.

[0004] A method for recording a video conference, comprising: utilizing a processor to execute the following steps when a video conference is initiated:

[0005] Providing a user interface to a display device, wherein the user interface includes a first area, a second area, and a timeline;

[0006] In response to obtaining an image corresponding to each of a plurality of participants from a video signal using a portrait recognition algorithm, displaying the image of each participant in the first area;

[0007] In response to converting an audio clip of one of the participants obtained from an audio signal into text content through a voice processing algorithm, the text content is associated with the corresponding one of the participants and displayed in the second area based on a speaking order; and the time length of the time axis is adjusted according to a recording time of the video conference.

[0008] In one embodiment, the first area provides an editing function, and after displaying the image of each participant in the first area, the method further includes:

[0009] Rename the image to a name corresponding to the image through the editing function.

[0010] In one embodiment, the processor further includes executing the following steps when initiating the video conference:

[0011] Identifying each participant included in the video signal through the portrait recognition algorithm and obtaining the relative position of each participant in a conference space;

[0012] Separating a human voice from the audio signal using a voiceprint recognition module;

[0013] Executing a sound source localization algorithm to determine a source position of the human voice in the conference space;

[0014] Matching the human voice with one of the participants corresponding to the human voice based on the relative position and the source position; and

[0015] The audio segment corresponding to the human voice is converted into the text content through the voice processing algorithm.

[0016] In one embodiment, the step of displaying the text content in the second area includes:

[0017] After converting the audio segment corresponding to the human voice into the text content through the voice processing algorithm, extracting the image of the participant matching the human voice and the corresponding name from the first area; and

[0018] The image, the name, the text content, and a receiving time of the audio clip are displayed in the second area.

[0019] In one embodiment, after converting the audio segment corresponding to the human voice into the text content through the voice processing algorithm, the method further includes:

[0020] The time segment corresponding to the audio segment in the timeline is associated with the text content.

[0021] In one embodiment, wherein the user interface provides a tagging function, the method further comprises:

[0022] Based on a time point at which the marking function is enabled, a highlight mark is performed on the text content corresponding to the time point in the second area.

[0023] In one embodiment, wherein the text content presented in the second area has a play function, the method further includes:

[0024] When the play function is enabled, the audio clip corresponding to the text content is played.

[0025] A video conferencing system, comprising:

[0026] a display device;

[0027] a storage device including an application program; and

[0028] a processor coupled to the display device and the storage, and configured to execute the application to initiate a video conference, and when the video conference is initiated, comprising:

[0029] Providing a user interface to the display device, wherein the user interface includes a first area, a second area, and a timeline;

[0030] In response to obtaining an image corresponding to each of a plurality of participants from a video signal using a portrait recognition algorithm, displaying the image of each participant in the first area;

[0031] In response to converting an audio clip of one of the participants obtained from an audio signal into text content through a voice processing algorithm, the text content is associated with the corresponding one of the participants and displayed in the second area based on a speaking order; and the time length of the time axis is adjusted according to a recording time of the video conference.

[0032] In one embodiment, the first area provides an editing function, and the processor is configured to:

[0033] After displaying the image of each participant in the first area, a name corresponding to the image is renamed through the editing function.

[0034] In one embodiment, the storage further includes a portrait recognition module, a voiceprint recognition module, and a speech processing module.

[0035] The processor is configured to:

[0036] Executing the portrait recognition algorithm through the portrait recognition module to identify each of the participants included in the video signal and obtain the relative position of each participant in a conference space;

[0037] Using the voiceprint recognition module, a human voice is separated from an audio signal;

[0038] Executing a sound source localization algorithm through the voiceprint recognition module to determine a source position of the human voice in the meeting space;

[0039] Matching the human voice with one of the participants corresponding to the human voice based on the relative position and the source position; and

[0040] The voice processing algorithm is executed by the voice processing module to convert the audio segment corresponding to the human voice into the text content.

[0041] In one embodiment, the processor is configured to:

[0042] After converting the audio segment corresponding to the human voice into the text content through the voice processing module, extracting a name corresponding to the image of one of the participants matching the human voice from the first area; and

[0043] The name, the text content, and a receiving time of the audio clip are displayed in the second area.

[0044] In one embodiment, the processor is configured to:

[0045] After the audio segment corresponding to the human voice is converted into the text content by the voice processing module, the time segment corresponding to the audio segment in the time axis is associated with the text content.

[0046] In one embodiment, wherein the user interface provides a tagging function, the processor is configured to:

[0047] Based on a time point at which the marking function is enabled, a highlight mark is performed on the text content corresponding to the time point in the second area.

[0048] In one embodiment, wherein the text content presented in the second area has a play function, the processor is configured to:

[0049] When the play function is enabled, the audio clip corresponding to the text content is played.

[0050] The video conferencing system of the above embodiment includes: a display device; a storage device including an application; and a processor coupled to the display device and the storage device and configured to execute the application to initiate a video conference. When the video conference is initiated, the system includes: providing a user interface to the display device, wherein the user interface includes a first area, a second area, and a timeline; in response to obtaining images corresponding to multiple participants from a video signal via a portrait recognition algorithm, displaying the image of each participant in the first area; in response to converting an audio clip of one of the participants obtained from the audio signal via a voice processing algorithm into text content, associating the text content with the corresponding one of the participants and displaying the text content in the second area based on the speaking order; and adjusting the time length of the timeline according to the recording time of the video conference.

[0051] Based on the above, this method can identify the source of speech in real time, convert human voice into text, associate the text with the speaker, and present the identification results through the user interface. This provides a complete meeting record for user reference. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the conventional technology, the following briefly introduces the drawings required for use in the embodiments or the conventional technology descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0053] Figure 1 is a schematic structural diagram of a video conferencing system according to an embodiment;

[0054] Figure 2 is a flowchart of a method for recording a video conference according to an embodiment;

[0055] Figure 3 is a structural diagram of an application example of a video conferencing system according to an embodiment;

[0056] Figure 4 is a schematic diagram of a user interface according to an embodiment.

[0057] Description of reference numerals:

[0058] 100: Video conferencing system

[0059] 110: Processor

[0060] 120: Storage

[0061] 121: Application

[0062] 130: Imaging device

[0063] 140: Radio receiver

[0064] 150: Display device

[0065] 310: Image Module

[0066] 311: Portrait recognition module

[0067] 313: Screenshot module

[0068] 315: Image recording module

[0069] 320: Sound module

[0070] 321: Voiceprint recognition module

[0071] 323: Voice processing module

[0072] 325: Voice recording module

[0073] 400: User Interface

[0074] 401~404: Play button

[0075] 410: First Area

[0076] 411A~411D: Name field

[0077] 420: Second Area

[0078] 421-425: Speech Materials

[0079] 425a: Thumbnail

[0080] 425b: Name

[0081] 425c: Text content

[0082] 425d: Receiving time

[0083] 425e: Play button

[0084] 430: Timeline

[0085] 431~435: Time segment

[0086] 440:Scrollbar

[0087] A~D:Image

[0088] S205-S220: Steps of the method for recording a video conference DETAILED DESCRIPTION

[0089] To facilitate understanding of the present application, the present application will be described more fully below with reference to the accompanying drawings. The accompanying drawings provide embodiments of the present application. However, the present application may be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to make the disclosure of the present application more thorough and comprehensive.

[0090] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application pertains. The terms used herein in the specification of this application are for the purpose of describing specific embodiments only and are not intended to limit this application.

[0091] It will be understood that the terms "first," "second," etc., used herein may be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish a first element from another element. For example, a first resistor may be referred to as a second resistor, and similarly, a second resistor may be referred to as a first resistor without departing from the scope of this application. The first resistor and the second resistor are both resistors, but they are not the same resistor.

[0092] It can be understood that the “connection” in the following embodiments should be understood as “electrical connection”, “communication connection”, etc. if there is transmission of electrical signals or data between the connected circuits, modules, units, etc.

[0093] It is understood that “at least one” refers to one or more, “a plurality” refers to two or more, and “at least a portion of an element” refers to a portion or all of an element.

[0094] As used herein, the singular forms "a," "an," and "said" may also include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the terms "include," "comprising," "having," and the like specify the presence of stated features, integers, steps, operations, components, parts, or combinations thereof, but do not preclude the presence or addition of one or more other features, integers, steps, operations, components, parts, or combinations thereof. Furthermore, the term "and / or" as used in this specification includes any and all combinations of the relevant listed items.

[0095] Figure 1 This is a schematic diagram of the structure of a video conferencing system according to an embodiment. Figure 1 The video conferencing system 100 includes a processor 110, a storage 120, an imaging device 130, an audio receiving device 140, and a display device 150. The processor 110 is coupled to the storage 120, the imaging device 130, the audio receiving device 140, and the display device 150.

[0096] The processor 110 is, for example, a central processing unit (CPU), a physical processing unit (PPU), a programmable microprocessor, an embedded control chip, a digital signal processor (DSP), an application specific integrated circuit (ASIC), or other similar devices.

[0097] Storage 120 may be, for example, any type of fixed or removable random access memory (RAM), read-only memory (ROM), flash memory, a hard disk, or other similar device, or a combination of these devices. Storage 120 stores one or more code segments, which, after being installed, are executed by processor 110. In this embodiment, storage 120 includes an application 121 for executing a video conference. When a video conference is initiated, processor 110 executes the following method for recording a video conference.

[0098] The imaging device 130 may be a video camera or a still camera that uses a charge coupled device (CCD) lens or a complementary metal oxide semiconductor transistor (CMOS) lens. For example, the imaging device 130 may be a wide-angle camera, a hemispherical camera, a full spherical camera, or the like.

[0099] The sound receiving device 140 is, for example, a microphone. In one embodiment, only one sound receiving device 140 may be provided. In other embodiments, multiple sound receiving devices 140 may also be provided.

[0100] The display device 150 is used to present a user interface and can be implemented, for example, by a liquid crystal display (LCD), a plasma display, a projection system, or the like.

[0101] Figure 2 This is a flowchart of a method for recording a video conference according to an embodiment. Figure 1 and Figure 2 In this embodiment, processor 110 executes application 121 to initiate a video conference. While application 121 is being initiated, processor 110 can also drive imaging device 130 and audio receiver 140 to obtain video and audio signals, respectively. Furthermore, when the video conference is initiated, processor 110 executes steps S205 to S220.

[0102] In step S205, a user interface is provided to display device 150. The user interface includes a first area, a second area, and a timeline. In step S210, in response to obtaining images corresponding to multiple participants from the video signal using a human face recognition algorithm, the images of each participant are displayed in the first area. The first area is used to present the speaker information in the video conference. For example, the human face recognition algorithm extracts the image of each participant from one or more frames of the video signal and displays it in the first area of ​​the user interface.

[0103] In step S215, in response to converting the audio clip of one of the participants obtained from the audio signal into text content via a voice processing algorithm, the text content is associated with the corresponding participant and displayed in the second area based on the speaking order. The voice processing algorithm is, for example, a speech-to-text algorithm. For example, processor 110 extracts audio clips of speeches made by the same speaker over a continuous period of time and performs speech-to-text processing. The second area is used to record the text content of speeches made during the video conference and display the multiple speeches according to the speaking order.

[0104] Furthermore, in step S220, the time length of the time axis is adjusted according to the recording time of the video conference. In addition, the processor 110 associates the recognized text content and its corresponding audio segment with the corresponding participant and stores them.

[0105] The following examples illustrate the detailed process of the application example of the video conferencing system 100.

[0106] Figure 3 It is a structural diagram of an application example of a video conferencing system according to an embodiment. Figure 3 Shown Figure 1 In the application example, the storage 120 further includes an image module 310 and an audio module 320. In this embodiment, the image module 310 and the audio module 320 are software modules provided independently of the application 121. However, in other embodiments, the image module 310 and the audio module 320 may also be integrated into the application 121, and this is not limited here.

[0107] The image module 310 includes a portrait recognition module 311, a screenshot module 313, and an image recording module 315. The portrait recognition module 311 is used to identify multiple participants in the video signal. The screenshot module 313 is used to capture images of each participant in the video signal. The image recording module 315 is used to record the entire video conference file and can also record the video clips corresponding to each participant's speech in real time.

[0108] The audio module 320 includes a voiceprint recognition module 321, a speech processing module 323, and a speech recording module 325. The voiceprint recognition module 321 is used to identify the speaker's location. The speech processing module 323 is used to convert audio clips into text. The speech recording module 325 is used to record the audio clips and text content corresponding to each participant's speech in real time.

[0109] Specifically, when the application 121 is started, the imaging device 130 and the sound receiving device 140 are driven to obtain video signals and audio signals respectively, so that the video signal obtained by the imaging device 130 is processed by the image module 310, and the audio signal obtained by the sound receiving device 140 is processed by the sound module 320.

[0110] After receiving the video and audio signals, the image recognition module 311 and voiceprint recognition module 321 can first match the participant's image with the corresponding voice. Specifically, the image recognition module 311 executes a human image recognition algorithm to identify each participant in the video signal and determine each participant's relative position within the conference space. For example, the image recognition module 311 can identify each video frame of the video conference and locate immovable objects in the conference space, such as furniture or furnishings, to determine the relative position of each participant within the conference space. Alternatively, the image recognition module 311 can determine the relative positions of movable objects to determine the relative positions of each participant within the conference space. Furthermore, the voiceprint recognition module 321 separates the voices of different individuals from the audio signal and records the corresponding voiceprint for each voice. The voiceprint recognition module 321 further executes a sound source localization algorithm to determine the source location of each voice within the conference space.

[0111] Based on each relative position obtained by the portrait recognition module 311 and each source position obtained by the voiceprint recognition module 321, the processor 110 matches each human voice with the image of one of the corresponding participants. Here, since not every participant will speak within the corresponding recording time, the number of human voices in the audio signal may be smaller than the number of participants in the video signal. Accordingly, the processor 110 matches the separated human voice with the image of the participant. In addition, since the voiceprint of the recognized human voice will be recorded, in the subsequently received audio signal, if the voiceprint of the recognized human voice has been recorded, the processor 110 can directly determine the participant corresponding to the human voice based on the previous match, without having to perform the matching work between the human voice and the participant again.

[0112] After matching the current speaker's voice with the corresponding participant, the voice processing module 323 executes a voice processing algorithm to convert the audio clip corresponding to the voice into text content, which is then displayed in the user interface.

[0113] Figure 4 This is a schematic diagram of a user interface of an embodiment. Figure 4 , the user interface 400 includes a first area 410 , a second area 420 and a timeline 430 .

[0114] The first area 410 is used to display images captured by the screenshot module 313. That is, after the portrait recognition module 311 identifies each participant in the video signal, the screenshot module 313 takes a screenshot of each participant in the video signal to obtain an image corresponding to each participant, and displays the image in the first area 410 of the user interface 400.

[0115] In this embodiment, first area 410 includes images A through D of four participants, each of which has a corresponding name field 411A through 411D. Name fields 411A through 411D are used to display the corresponding name of each participant. In one embodiment, after screenshot module 313 captures images A through D of each participant, default names can be directly displayed in name fields 411A through 411D. Furthermore, first area 410 may also provide an editing function to rename the corresponding names of images A through D. For example, name fields 411A through 411D each have an editing function, allowing the user to directly re-enter the desired name in name fields 411A through 411D.

[0116] In addition, images A-D in the first area 410 also have corresponding play buttons 401-404. In one embodiment, after the voiceprint recognition module 321 identifies one or more human voices and matches each voice with a corresponding participant, an audio clip that clearly identifies the voice of each participant is automatically extracted and associated with the corresponding play button 401-404. When one of the play buttons 401-404 is activated, the audio clip of the corresponding participant's voice is played.

[0117] The second area 420 is used to present multiple speech data 421 to 425. The second area 420 also includes a scroll bar 440, so that the content displayed in the second area 420 can scroll up and down in the vertical direction. Specifically, after the voice processing module 323 converts the audio segment corresponding to the human voice separated from the audio signal into text content through the voice processing algorithm, the processor 110 captures the image of one of the participants matching the human voice and the corresponding name from the first area 410. Then, the processor 110 displays the obtained image, name, text content and the reception time of the audio segment as a speech data in the second area 420. For the speech data 425, it is the speech data of the participant corresponding to image A. The speech data 425 includes the image corresponding to image A. Figure 4425a in the , the name 425b corresponding to the name field 411A, the text content of the speech 425c, the receiving time 425d of the audio clip (ie, the speech time).

[0118] Each speech data entry is also equipped with a corresponding play function. When the play function is enabled, the audio clip corresponding to the text content is played. For example, speech data 425 is provided with a corresponding play button 425e. Play button 425e is associated with the audio clip corresponding to text content 425c. When play button 425e is enabled, processor 110 plays the audio clip corresponding to text content 425c through the speaker. Similarly, speech data 421-424 have the same structure.

[0119] Timeline 430 includes time segments 431 to 435, which correspond to speech data 421 to 425, respectively. Specifically, after converting the audio segment corresponding to the human voice into text content, processor 110 can also associate the time segments 431 to 435 corresponding to the audio segment in timeline 430 with the text content. Accordingly, the corresponding speech data can be selected by selecting from the time segments 431 to 435 on timeline 430. Alternatively, the corresponding time segment in timeline 430 can be selected by selecting from the speech data 421 to 425. For speech data 425, text content 425c is associated with time segment 435. Similarly, speech data 421 to 424 also have the same structure.

[0120] Furthermore, after obtaining the images A-D corresponding to the participant, the processor 110 may further store the images A-D in a database and associate the subsequent video clips, audio clips, and text content of each participant's speech with their images. The video clips, audio clips, and text content may be stored in separate databases or in the same database.

[0121] In addition, the user interface 400 may further provide a marking function. Based on the time at which the marking function is enabled, the processor 110 highlights the text content in the second area 420 corresponding to the time point. For example, if a function button is provided in the user interface and the time at which this function button is enabled is 3:05 PM, the processor 110 highlights the text content 425c. For example, text content 425c may be marked in a color different from other text, or text content 425c may be visually displayed using a highlighting or fluorescent marking technique. Alternatively, the marking function may be enabled by setting a shortcut key on the remote control.

[0122] In summary, this method can identify the source of a call in real time, associate the text content converted from audio clips with the speaker, and provide a user interface to present the meeting transcript. The user interface summarizes the speaker's image and the corresponding audio clip. Furthermore, through this implementation, recording and generating meeting transcripts can be started in real time, eliminating the need to re-establish a participant database each time a video conference is initiated.

[0123] In the description of this specification, reference to the terms "some embodiments" or "other embodiments" means that a specific feature, structure, material, or characteristic described in conjunction with the embodiment or example is included in at least one embodiment or example of the present application. In this specification, the schematic descriptions of the above terms do not necessarily refer to the same embodiment or example.

[0124] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0125] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A method for recording a video conference, characterized in that: A processor is used to execute the following steps when starting a video conference, including: Providing a user interface to a display device, wherein the user interface includes a first area, a second area, and a timeline; In response to obtaining an image corresponding to each of a plurality of participants from a video signal through a portrait recognition algorithm, displaying the image of each participant in the first area; In response to converting an audio clip of one of the participants obtained from an audio signal into text content through a speech processing algorithm, associating the text content with the corresponding one of the participants, and displaying the text content in the second area based on a speaking order; and The time length of the time axis is adjusted according to the recording time of the video conference.

2. The method for recording a video conference according to claim 1, characterized in that: The first area provides an editing function, and after displaying the image of each participant in the first area, the method further includes: The image is renamed using the editing function.

3. The method for recording a video conference according to claim 1, characterized in that: Wherein, using the processor to start the video conference further includes executing the following steps, including: Using the portrait recognition algorithm, identify each of the participants included in the video signal, and obtain the relative position of each of the participants in a conference space; Separating a human voice from the audio signal through a voiceprint recognition module; Execute a sound source localization algorithm to determine a source position of the human voice in the conference space; Matching the human voice with one of the participants corresponding to the human voice based on the relative position and the source position; and The audio segment corresponding to the human voice is converted into the text content through the voice processing algorithm.

4. The method for recording a video conference according to claim 3, characterized in that: The step of displaying the text content in the second area includes: After converting the audio segment corresponding to the human voice into the text content through the voice processing algorithm, extracting the image of one of the participants matching the human voice and a corresponding name from the first area; and The image, the name, the text content and a receiving time of the audio clip are displayed in the second area.

5. The method for recording a video conference according to claim 4, characterized in that: After the audio segment corresponding to the human voice is converted into the text content through the voice processing algorithm, the method further includes: The time segment corresponding to the audio segment in the timeline is associated with the text content.

6. The method for recording a video conference according to claim 1, characterized in that: Wherein the user interface provides a marking function, the method further comprises: Based on a time point when the marking function is enabled, a highlight mark is performed on the text content corresponding to the time point in the second area.

7. The method for recording a video conference according to claim 1, characterized in that: The text content presented in the second area has a playback function, and the method further includes: When the play function is enabled, the audio segment corresponding to the text content is played.

8. A video conferencing system, characterized in that: include: a display device; a storage device including an application program; as well as A processor is coupled to the display device and the storage, and is configured to execute the application to start a video conference, and when starting the video conference, includes: Providing a user interface to the display device, wherein the user interface includes a first area, a second area, and a timeline; In response to obtaining an image corresponding to each of a plurality of participants from a video signal through a portrait recognition algorithm, displaying the image of each participant in the first area; In response to converting an audio clip of one of the participants obtained from an audio signal into text content through a speech processing algorithm, associating the text content with the corresponding one of the participants, and displaying the text content in the second area based on a speaking order; and The time length of the time axis is adjusted according to the recording time of the video conference.

9. The video conferencing system according to claim 8, characterized in that: The first area provides an editing function, and the processor is configured to: After displaying the image of each participant in the first area, a name corresponding to the image is renamed through the editing function.

10. The video conferencing system according to claim 8, characterized in that: The storage device further includes a portrait recognition module, a voiceprint recognition module and a speech processing module. The processor is configured to: Executing the portrait recognition algorithm through the portrait recognition module to identify each of the participants included in the video signal and obtain the relative position of each of the participants in a conference space; A human voice is separated from an audio signal through the voiceprint recognition module; Executing a sound source localization algorithm through the voiceprint recognition module to determine a source position of the human voice in the conference space; Matching the human voice with one of the participants corresponding to the human voice based on the relative position and the source position; and The voice processing algorithm is executed by the voice processing module to convert the audio segment corresponding to the human voice into the text content.

11. The video conferencing system according to claim 10, characterized in that: wherein the processor is configured to: After converting the audio segment corresponding to the human voice into the text content through the voice processing module, extracting a name corresponding to the image of one of the participants matching the human voice from the first area; and The name, the text content and a receiving time of the audio clip are displayed in the second area.

12. The video conferencing system according to claim 11, characterized in that: wherein the processor is configured to: After the audio segment corresponding to the human voice is converted into the text content through the voice processing module, the time segment corresponding to the audio segment in the time axis is associated with the text content.

13. The video conferencing system according to claim 8, characterized in that: Wherein the user interface provides a marking function, the processor is configured to: Based on a time point when the marking function is enabled, a highlight mark is performed on the text content corresponding to the time point in the second area.

14. The video conferencing system according to claim 8, characterized in that: The text content presented in the second area has a playback function, and the processor is configured to: When the play function is enabled, the audio segment corresponding to the text content is played.