Video conference processing method and system, electronic equipment and storage medium

By obtaining audio data in real time and generating structured meeting minutes using pre-trained artificial intelligence models, the problem of time-consuming and labor-consuming traditional video conference records is solved, and efficient and intelligent meeting minutes generation and distribution are achieved.

CN120302001APending Publication Date: 2025-07-11VISIONVERA INFORMATION TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510263780.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

Traditional video conference records rely on manual or simple automation tools, resulting in time-consuming and labor-intensive and lack of structured processing of conference content and intelligent refining of key points.

Method used

By obtaining audio data from each conference terminal in real time, using pre-trained artificial intelligence models to generate structured conference minutes data, including speech recognition and natural language processing technology, generating structured conference minutes data containing information such as conference topics, key points, to-do items, etc., and output them in electronic form.

Benefits of technology

It realizes the automated processing of meeting records, reduces manual participation, and generates efficient and structured meeting minutes data, which facilitates quick access and management of participants.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120302001A_ABST
    Figure CN120302001A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a video conference processing method and system, and the method comprises the steps: obtaining the audio data of each conference terminal in real time in a conference process; generating structured conference summary data according to the audio data and a pre-trained artificial intelligence large model; and outputting the conference summary data. According to the embodiment of the invention, the audio data of each conference terminal is acquired in real time, and the structured conference summary data is generated by using the pre-trained artificial intelligence large model, so that the automation of the conference recording process is realized. Compared with manual recording or automatic recording only generating an original text in a traditional method, the method has the advantages that the requirement of manual participation is reduced, the conference content can be intelligently processed, and the conference summary data can be generated and output.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video conferencing, and in particular, to a method for processing video conferencing, a system for processing video conferencing, an electronic device, and a computer-readable storage medium. Background Art

[0002] Traditional video conferencing systems usually generate a summary of the meeting content through manual recording or simple automation tools. However, these methods have significant deficiencies.

[0003] Manual recording is not only time-consuming and laborious, but also prone to missing key information. Moreover, existing automated recording technologies mostly rely on basic speech-to-text functions, and the generated meeting records are often just raw text, lacking structured processing of the meeting content and intelligent extraction of key points. Summary of the Invention

[0004] In view of the above problems, embodiments of the present invention are proposed to provide a method for processing video conferencing, a system for processing video conferencing, an electronic device, and a computer-readable storage medium that overcome the above problems or at least partially solve the above problems.

[0005] To solve the above problems, embodiments of the present invention disclose a method for processing video conferencing, the method comprising:

[0006] During the meeting, real-time obtain the audio data of each meeting terminal;

[0007] Generate structured meeting minutes data according to the audio data and a pre-trained large artificial intelligence model;

[0008] Output the meeting minutes data.

[0009] Optionally, the generating structured meeting minutes data according to the audio data and a pre-trained large artificial intelligence model includes:

[0010] Use speech recognition technology to convert the audio data into text data;

[0011] Input the text data into the large artificial intelligence model, and supplement the missing text content in the text data through the large artificial intelligence model, and output the meeting minutes data.

[0012] Optionally, the inputting the text data into the large artificial intelligence model, supplementing the missing text content in the text data through the large artificial intelligence model, and outputting the meeting minutes data includes:

[0013] Input the text data into the artificial intelligence large model, analyze the context relationship in the text data through the artificial intelligence large model, and identify and supplement the missing text content in the text data according to the context relationship;

[0014] Screen, classify, and label the supplemented text data to obtain the meeting minutes data;

[0015] Among them, the meeting minutes data includes at least one of the following: meeting theme information, meeting key point information, to-do item information, decision-making item information, content label information, theme duration information, meeting summary information.

[0016] Optionally, the output of the meeting minutes data includes:

[0017] Output the meeting minutes data in electronic form to each of the meeting terminals.

[0018] Optionally, the output of the meeting minutes data in electronic form to each of the meeting terminals includes:

[0019] Generate an access identifier pointing to the meeting minutes data and output the access identifier to each of the meeting terminals;

[0020] Among them, the access identifier includes: a QR code or a download link.

[0021] Optionally, after the output of the meeting minutes data, the method further includes:

[0022] Obtain access request information for the meeting minutes data;

[0023] Determine the access mode of the meeting minutes data according to the access request information, and the access mode includes a free mode or a paid mode;

[0024] Provide the meeting minutes data based on the access mode.

[0025] Optionally, the providing of the meeting minutes data based on the access mode includes:

[0026] When the access mode is the paid mode, receive payment information from the payment system, and when the payment information indicates successful payment, transmit and display the meeting minutes data to the device corresponding to the acquisition request information.

[0027] An embodiment of the present invention also discloses a processing system for a video conference, and the system includes:

[0028] An audio data acquisition module, configured to acquire the audio data of each meeting terminal in real time during the meeting;

[0029] A meeting minutes generation module, configured to generate structured meeting minutes data according to the audio data and a pre-trained large artificial intelligence model;

[0030] A meeting minutes output module, configured to output the meeting minutes data.

[0031] Optionally, the meeting minutes generation module includes:

[0032] An audio conversion module, configured to convert the audio data into text data by using speech recognition technology;

[0033] A text input module, configured to input the text data into the large artificial intelligence model, supplement the missing text content in the text data through the large artificial intelligence model, and output the meeting minutes data.

[0034] Optionally, the text input module is configured to input the text data into the large artificial intelligence model, analyze the context relationship in the text data through the large artificial intelligence model, identify and supplement the missing text content in the text data according to the context relationship; screen, classify and label the supplemented text data to obtain the meeting minutes data;

[0035] Wherein, the meeting minutes data includes at least one of the following: meeting theme information, meeting key point information, to-do item information, decision-making item information, content label information, theme duration information, meeting summary information.

[0036] Optionally, the meeting minutes output module is configured to output the meeting minutes data in an electronic form to each of the meeting terminals.

[0037] Optionally, the meeting minutes output module is configured to generate an access identifier pointing to the meeting minutes data, and output the access identifier to each of the meeting terminals;

[0038] Wherein, the access identifier includes: a QR code or a download link.

[0039] Optionally, the system further includes:

[0040] An access request acquisition module, configured to acquire access request information for the meeting minutes data after the meeting minutes output module outputs the meeting minutes data;

[0041] An access mode determination module, configured to determine the access mode of the meeting minutes data according to the access request information, where the access mode includes a free mode or a paid mode;

[0042] A meeting minutes providing module, configured to provide the meeting minutes data based on the access mode.

[0043] Optionally, the meeting minutes providing module is configured to, when the access mode is the paid mode, receive payment information from a payment system, and if the payment information indicates successful payment, transmit and display the meeting minutes data to a device corresponding to the acquisition request information.

[0044] An embodiment of the present invention also discloses an electronic device, including: one or more processors; and one or more machine-readable media storing instructions thereon, which, when executed by the one or more processors, cause the electronic device to execute the processing method of the video conference as described above.

[0045] An embodiment of the present invention also discloses a computer-readable storage medium, and a computer program stored thereon causes a processor to execute the processing method of the video conference as described above.

[0046] The embodiments of the present invention include the following advantages:

[0047] The processing solution of the video conference provided by the embodiments of the present invention, during the meeting process, obtains the audio data of each meeting terminal in real time; generates structured meeting minutes data according to the audio data and a pre-trained artificial intelligence large model; outputs the meeting minutes data.

[0048] Compared with the background art, the embodiments of the present invention have the following beneficial effects:

[0049] By obtaining the audio data of each meeting terminal in real time and using a pre-trained artificial intelligence large model to generate structured meeting minutes data, the embodiments of the present invention realize the automation of the meeting recording process. Compared with the traditional manual recording method or the automated recording that only generates raw text, this method reduces the need for manual participation and can perform intelligent processing on the meeting content, generating and outputting the meeting minutes data. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 is a flowchart of the steps of a processing method of a video conference according to an embodiment of the present invention;

[0051] Figure 2 is a block diagram of the structure of a processing system of a video conference according to an embodiment of the present invention. DETAILED DESCRIPTION

[0052] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0053] An embodiment of the present invention provides a processing solution for video conferencing. During the meeting, by real-time acquiring the audio data of each meeting terminal, using a pre-trained large artificial intelligence model to generate structured meeting minutes data, and outputting the meeting minutes data. Specifically, this solution collects audio data from the meeting terminals in real time, and converts the audio data into text data through speech recognition technology. Subsequently, the text data is input into the pre-trained large artificial intelligence model, and the artificial intelligence model supplements, filters, classifies, and marks the meeting content, etc., to generate structured meeting minutes data including information such as meeting topics, key points, and to-do items. Finally, the meeting minutes data is output to each meeting terminal in electronic form, and access identifiers such as QR codes or download links can be generated for users to obtain, and data provision in free or paid modes is supported, thereby realizing the automated processing and efficient distribution of meeting records.

[0054] Refer to Figure 1 , which shows the step flow chart of a processing method for a video conference according to an embodiment of the present invention. This processing method for a video conference can be applied to a video conferencing system or a video conference management system, simply referred to as a system. The processing method for a video conference specifically may include the following steps:

[0055] Step 101, during the meeting, real-time acquire the audio data of each meeting terminal.

[0056] Each meeting terminal is usually equipped with a voice collection device, such as a microphone, for capturing the speech content of the participants. The system will real-time monitor the audio signals generated during the meeting and convert them into a digital format that can be processed. The acquisition of audio data is not limited to a single meeting terminal, but covers all meeting terminals participating in the meeting, ensuring that the speeches of multiple parties can be completely recorded. For example, in a remote team meeting, participants located in different locations join through their respective laptops or mobile devices, and the system will synchronously collect the speech audio of each person. These audio data can be continuous speech streams, containing speech content with different volumes, speech rates, and languages, laying a foundation for subsequent processing. In the actual application process, the system may transmit the audio data through existing network protocols to ensure the real-time and integrity of data acquisition. The key to this step is to accurately capture all the sound information in the meeting, provide the original material for generating the meeting minutes data, and at the same time ensure that the data format adapts to the requirements of subsequent artificial intelligence processing.

[0057] Step 102, generate structured meeting minutes data according to the audio data and the pre-trained large artificial intelligence model.

[0058] In the actual application process, speech recognition technology can be used to convert audio data into text data. For example, the speech of the participants can be converted from sound waveforms into readable text records. Then, these text data are input into a pre-trained large artificial intelligence model. This large artificial intelligence model has been trained with a large amount of corpus and can understand the context and extract key information. The large artificial intelligence model will analyze the text data, including screening important content, classifying discussion topics, and marking action items, and finally generate structured meeting minutes data. The generated structured meeting minutes data organizes the meeting content in a systematic and orderly manner into data with a clear format and logical hierarchy, rather than simply a pile of texts, and may include information such as meeting topics, key decisions, and to-do items. For example, in a project meeting, the large artificial intelligence model may extract core contents such as "complete the report" and "person in charge" from the lengthy discussion. This process does not require manual intervention. The large artificial intelligence model automatically completes it through natural language processing technology. The generated meeting minutes data has the characteristics of being structured, which is convenient for reading and archiving. The pre-trained large artificial intelligence model relies on previous training data and algorithm design to ensure that the output meeting minutes data accurately reflects the meeting content.

[0059] Step 103, output the meeting minutes data.

[0060] The system outputs the meeting minutes data in electronic form to each meeting terminal to ensure that the participants can obtain the meeting results. Specifically, an access identifier pointing to the meeting minutes data, such as a QR code or a download link, may be generated and displayed on the interface of the meeting terminal. The participants can obtain it by scanning or clicking. For example, after the meeting ends, the system displays a QR code on the screen of each participant. After scanning, a document containing the meeting minutes can be downloaded. In addition, the output method may also include transmitting the meeting minutes data to the display interface of the meeting terminal in real time, or sending it to an external storage system through the network for subsequent access. The system also supports providing data according to the access mode, such as directly outputting in the free mode, or transmitting it to the user-specified device after payment in the paid mode.

[0061] The processing solution for video conferencing provided by the embodiments of the present invention obtains the audio data of each meeting terminal in real time during the meeting; generates structured meeting minutes data according to the audio data and the pre-trained large artificial intelligence model; outputs the meeting minutes data.

[0062] Compared with the background technology, the embodiments of the present invention have the following beneficial effects:

[0063] In an embodiment of the present invention, by obtaining the audio data of each conference terminal in real time and using a pre-trained large artificial intelligence model to generate structured meeting minutes data, the automation of the meeting recording process is achieved. Compared with the manual recording of traditional methods or the automated recording that only generates raw text, this method reduces the need for manual participation and can intelligently process the meeting content to generate and output the meeting minutes data.

[0064] In an exemplary embodiment of the present invention, an implementation manner of generating structured meeting minutes data according to the audio data and the pre-trained large artificial intelligence model is as follows: using speech recognition technology to convert the audio data into text data; inputting the text data into the large artificial intelligence model, and supplementing the missing text content in the text data through the large artificial intelligence model to output the meeting minutes data.

[0065] In the actual application process, the system uses speech recognition technology to convert the audio data obtained from each conference terminal into text data. Speech recognition technology usually relies on Automatic Speech Recognition (ASR) technology. By analyzing the acoustic features of the audio signal, the speech of the participants is converted into a readable text form. This process ensures that the audio data can be processed by subsequent artificial intelligence models. For example, in a multi-party video conference, the participants discuss the project progress. The system uses ASR to convert the meeting dialogue into text in real time, such as "submit a report". Then, the system inputs these text data into the pre-trained large artificial intelligence model. This large artificial intelligence model is based on Natural Language Processing (NLP) technology and analyzes and processes the text data. Among them, the large artificial intelligence model can supplement the missing text content in the text data. The large artificial intelligence model has accumulated language understanding ability through pre-training, can recognize the context, extract key information and generate structured meeting minutes data. For example, it may extract "Meeting topic: Project progress" "To-do item: Submit a report" and other content from the long discussion and output it in the form of items. The entire process does not require manual editing, and the model directly outputs the meeting minutes data, ensuring the integrity and consistency of the information.

[0066] This implementation manner solves the problem of converting audio data into text through ASR technology, and then realizes the transformation from the original text to structured meeting minutes data with the help of the NLP-driven large artificial intelligence model. Compared with traditional manual recording or simple transcription, it can automatically convert audio data into systematic meeting minutes data, reduce human intervention, and at the same time retain the key information of the meeting content. It is applicable to scenarios that require quickly sorting out the speeches of multiple parties, improving the efficiency and usability of meeting records.

[0067] In an exemplary embodiment of the present invention, when inputting text data into an artificial intelligence large model to supplement the missing text content in the text data, one implementation manner of outputting meeting minutes data is as follows: input the text data into the artificial intelligence large model, analyze the context relationship in the text data through the artificial intelligence large model, identify and supplement the missing text content in the text data according to the context relationship; screen, classify, and label the supplemented text data to obtain meeting minutes data; wherein, the meeting minutes data includes at least one of the following: meeting theme information, meeting key point information, to-do item information, decision item information, content label information, theme duration information, meeting summary information.

[0068] In the actual application process, the text data converted from audio data is input into a pre-trained artificial intelligence large model. This artificial intelligence large model performs multi-level processing on the input text data based on NLP technology. Specifically, due to speech recognition errors or network transmission problems, some key contents may be missing in the text data, such as complete speech sentences, decision items, important conclusions, etc. At this time, the artificial intelligence large model analyzes the context relationship and conducts reasoning to automatically complete the missing text content to ensure the integrity and readability of the meeting record. The artificial intelligence large model eliminates redundant or irrelevant information through the screening function and retains the core content in the meeting. Through the classification function, the text data is grouped by theme or type. Through the labeling function, identifiers are added to important information, and finally structured meeting minutes data is generated. For example, in a project discussion, the text data may contain long conversations. The artificial intelligence large model will screen out key sentences such as "The design needs to be completed before Friday", classify it as a "to-do item", and label the person in charge, and output it as "To-do item: Complete the design before Friday (Person in charge: Li)". The meeting minutes data includes at least one specific content, such as meeting theme information (meeting purpose), meeting key point information (main discussion points), to-do item information (action plan), decision item information (consensus reached), content label information (keywords), theme duration information (topic duration), or meeting summary information (brief summary). The generation of these contents depends on the semantic understanding and structuring ability of the artificial intelligence large model for the text data, ensuring that the output meeting minutes data is clear and well-organized.

[0069] This implementation manner utilizes NLP technology to endow the artificial intelligence large model with the ability to extract key information from the original text and organize it into meeting minutes. Compared with traditional manual collation or simple recording, through completion, screening, classification, and labeling, it is possible to generate meeting minutes data containing multi-dimensional information from complex text data, reducing the workload of manual processing and providing a more comprehensive meeting record for subsequent reference and implementation.

[0070] In an exemplary embodiment of the present invention, one implementation of outputting the meeting minutes data is to output the meeting minutes data in electronic form to each meeting terminal.

[0071] In the actual application process, the meeting terminal usually refers to the device used by the participants, such as a laptop, a smartphone or a tablet computer, and these devices are connected to the system through the network. Electronic form means that the meeting minutes data is output in the form of a digital file or information stream, such as a text document, a table or the content displayed on the interface, rather than the traditional paper print output. Specifically, the system can directly send the meeting minutes data to the display interface of each meeting terminal, or distribute it to the terminal storage module through a digital transmission protocol (such as the Hyper Text Transfer Protocol, abbreviated as HTTP). For example, after a multi-person video conference ends, the system may transmit the meeting minutes data containing "Meeting Theme: Quarterly Plan" and "To-do Item: Submit Report" to the screens of the participants' devices for real-time viewing. This output method relies on the network function of the system to ensure that the data can reach all relevant terminals immediately, regardless of geographical location. In addition, the electronic form of output provides a basis for subsequent operations (such as saving and forwarding). The participants can further process the meeting minutes data as needed, and the whole process does not require physical media or manual intervention.

[0072] This implementation outputs the meeting minutes data in electronic form, solving the problems in the distribution of traditional meeting records that rely on paper documents or manual transmission. The system can quickly transmit the meeting minutes data to each meeting terminal, achieving the immediacy and universality of distribution, avoiding the cumbersome steps of the traditional method, and facilitating the participants to manage and use the meeting information on electronic devices.

[0073] In an exemplary embodiment of the present invention, one implementation of outputting the meeting minutes data in electronic form to each meeting terminal is to generate an access identifier pointing to the meeting minutes data and output the access identifier to each meeting terminal; wherein, the access identifier includes: a QR code or a download link.

[0074] In the actual application process, for the meeting minutes data generated by a pre-trained large artificial intelligence model, an access identifier pointing to the meeting minutes data is generated. This access identifier can serve as an electronic medium, including the access path to the meeting minutes data, and the specific forms include QR codes or download links. A QR code is a two-dimensional barcode (Quick Response Code, abbreviated as QR Code), which can quickly parse out the embedded address information through scanning. A download link is a hyperlink that points to the network location where the meeting minutes data is stored. After generating the access identifier, the system outputs it to each meeting terminal, such as the computer or mobile phone screen of the participants. For example, after a video conference ends, the system may display a QR code on each meeting terminal, and after scanning it with a mobile phone, the participants can access the meeting minutes data containing "Key points of the meeting: Product release plan"; or display a link, and after clicking it, directly download the relevant file. This method utilizes digital technology to distribute the access identifier to users through the display function of the meeting terminal, and then users can obtain the meeting minutes data through the identifier. The entire process relies on the network transmission protocol to ensure the effectiveness and accessibility of the access identifier.

[0075] This implementation method converts the output of the meeting minutes data into a simple electronic distribution form by generating a QR code or a download link as the access identifier. Compared with directly transmitting the complete file, the access identifier occupies less resources and has a faster distribution speed. Users can flexibly obtain the meeting minutes data through the meeting terminal according to their needs, thus simplifying the distribution process and improving the autonomy of acquisition.

[0076] In an exemplary embodiment of the present invention, an implementation method after outputting the meeting minutes data is: obtaining access request information for the meeting minutes data; determining the access mode of the meeting minutes data according to the access request information, and the access mode includes a free mode or a paid mode; providing the meeting minutes data based on the access mode.

[0077] In the actual application process, access request information for meeting minutes data is obtained. This access request information is usually initiated by the user through a meeting terminal, such as clicking a download link or submitting an acquisition application, and includes the user's identity identifier or request intention. Based on the access request information, the access mode of the meeting minutes data is judged. The access mode is divided into two types: free mode and paid mode. In the free mode, the user can obtain the data without additional conditions. In the paid mode, the user needs to complete the payment process. For example, after a training meeting, ordinary participants may directly download the meeting minutes data through the free mode, while external users may need to obtain the file containing the "course summary" through the paid mode. After determining the access mode, the system provides the meeting minutes data based on this access mode, which may transmit the meeting minutes data to the meeting terminal specified by the user through the Hypertext Transfer Protocol, or present it to the user through a display interface. Relying on the permission management function of the system, the controllability of data distribution is ensured. On the basis of outputting the meeting minutes data, an access control link is added. The user triggers the system response through the access request information, and the acquisition method varies according to the mode.

[0078] This implementation method adds flexibility and controllability to the provision of meeting minutes data by introducing access request information and access mode judgment. It can provide meeting minutes data in free or paid modes according to user needs and permissions, which not only meets the requirements of open sharing but also supports access restrictions in specific scenarios, improving the adaptability and management ability of data distribution.

[0079] In an exemplary embodiment of the present invention, an implementation method for providing meeting minutes data based on the access mode is as follows: when the access mode is the paid mode, payment information from the payment system is received, and when the payment information indicates successful payment, the meeting minutes data is transmitted and displayed on the device corresponding to the acquisition request information.

[0080] In the actual application process, for the previously determined access mode, special focus is placed on operations under the paid mode. When the access mode is the paid mode, the system receives payment information from the payment system. The payment system is usually an external or built-in electronic payment platform, such as a third-party payment interface, which is responsible for handling users' payment requests. The payment information includes data such as the transaction amount and payment status, which are submitted by the user through the conference terminal, for example, by scanning a QR code and then jumping to the payment page to complete the operation. After receiving the payment information, the system checks its status. If the payment information indicates successful payment (such as receiving a confirmation message returned by the payment platform), the system then triggers the next action. Subsequently, the system transmits and displays the meeting minutes data to the device corresponding to the obtained request information. Here, the device is usually the conference terminal from which the user initiates the access request, such as a smartphone or a computer, and the transmission process may be achieved through the Hypertext Transfer Protocol. For example, after a business meeting, an external user pays the fee, and the system transmits the "Meeting Decision: Product Pricing Plan" to their mobile phone and displays it on the screen. This process ensures that only users who have completed the payment can obtain the meeting minutes data, reflecting the rigor of access control.

[0081] This implementation method, through the payment verification mechanism, limits the provision of meeting minutes data to users who have paid successfully, increasing the conditional nature of distribution. The system can confirm user permissions through payment information in the paid mode, accurately transmit the meeting minutes data to the specified device, achieving controllability and pertinence of data access, and supporting commercial requirements in specific scenarios.

[0082] Based on the above related description of the embodiment of a video conferencing processing method, the following introduces a solution for generating meeting minutes data based on speech recognition and artificial intelligence large models, which is applicable to video conferencing systems or video conferencing management systems (hereinafter referred to as "systems"), aiming to achieve the automatic generation and convenient distribution of meeting records. The specific implementation process is as follows: First, the user initiates a video conference through the Visual Networking Conference Platform. The system uses the microphones of each conference terminal to collect audio data in real time and activates the speech recognition module. The speech recognition module adopts deep learning algorithms and converts the audio data into text data through Real-Time Communication (WebRTC) and Google Speech-to-Text services. During the conversion process, the system processes the audio signal, suppresses background noise, and annotates the speaker's identity according to acoustic features, such as distinguishing the speeches of "Zhang" and "Li". The generated text data is displayed on the interfaces of each conference terminal in real time as the meeting record. Secondly, the system inputs the meeting record into a pre-trained artificial intelligence large model. The model analyzes the text data based on natural language processing technology and generates meeting minutes data through screening, classification, and tagging. For example, in a project meeting, the model extracts content such as "To-do: Submit a report on Friday (Person in charge: Zhang)" and "Key decision: Adopt Plan A" from the discussion, generates content tags for each sub-topic (such as "Progress discussion"), and simultaneously calculates the topic duration (such as "10 minutes"). In addition, the model generates a short meeting summary, such as "The meeting discussed the project progress, decided to adopt Plan A, and assigned tasks", for quick browsing. The meeting minutes data and summary are automatically saved. Subsequently, after the meeting ends, the system generates a QR code link pointing to the meeting minutes data through the built-in QR code generation module and displays it on the conference terminal interface. Participants use their mobile phones or tablets to scan the QR code to access the download page. The system provides a free mode (direct download) or a paid mode (download after completing payment through a third-party payment platform) according to the preset permissions. The download link is encrypted to ensure data security. For example, internal participants can obtain it for free, while external users need to pay to download. Finally, participants can save or print the meeting minutes data to complete the sharing of meeting information.

[0083] It should be noted that for the method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of the present invention are not limited by the described action sequences, because according to the embodiments of the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential for the embodiments of the present invention.

[0084] Refer to Figure 2, showing a structural block diagram of a video conferencing processing system according to an embodiment of the present invention. The video conferencing processing system may specifically include the following modules.

[0085] An audio data acquisition module 21, configured to acquire the audio data of each conference terminal in real time during the conference;

[0086] A meeting minutes generation module 22, configured to generate structured meeting minutes data according to the audio data and a pre-trained large artificial intelligence model;

[0087] A meeting minutes output module 23, configured to output the meeting minutes data.

[0088] In an exemplary embodiment of the present invention, the meeting minutes generation module 22 includes:

[0089] An audio conversion module, configured to convert the audio data into text data by using speech recognition technology;

[0090] A text input module, configured to input the text data into the large artificial intelligence model, supplement the missing text content in the text data through the large artificial intelligence model, and output the meeting minutes data.

[0091] In an exemplary embodiment of the present invention, the text input module is configured to input the text data into the large artificial intelligence model, analyze the context relationship in the text data through the large artificial intelligence model, identify and supplement the missing text content in the text data according to the context relationship; screen, classify, and label the supplemented text data to obtain the meeting minutes data;

[0092] Wherein, the meeting minutes data includes at least one of the following: meeting theme information, meeting key point information, to-do item information, decision-making item information, content label information, theme duration information, meeting summary information.

[0093] In an exemplary embodiment of the present invention, the meeting minutes output module 23 is configured to output the meeting minutes data in an electronic form to each of the conference terminals.

[0094] In an exemplary embodiment of the present invention, the meeting minutes output module 23 is configured to generate an access identifier pointing to the meeting minutes data and output the access identifier to each of the conference terminals;

[0095] Wherein, the access identifier includes: a two-dimensional code or a download link.

[0096] In an exemplary embodiment of the present invention, the system further includes:

[0097] An access request acquisition module, configured to acquire access request information for the minutes data after the minutes output module 23 outputs the minutes data;

[0098] An access mode determination module, configured to determine an access mode for the minutes data according to the access request information, where the access mode includes a free mode or a paid mode;

[0099] A minutes providing module, configured to provide the minutes data based on the access mode.

[0100] In an exemplary embodiment of the present invention, the minutes providing module is configured to, when the access mode is the paid mode, receive payment information from a payment system, and when the payment information indicates successful payment, transmit and display the minutes data to a device corresponding to the acquisition request information.

[0101] For the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For related parts, refer to the partial description of the method embodiment.

[0102] Each embodiment in this specification is described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the embodiments, refer to each other.

[0103] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a device, or a computer program product. Therefore, the embodiments of the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.

[0104] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the method, terminal device (system), and computer program product according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0105] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instruction means that implement the function specified in one or more of the processes and / or blocks Figure 1 one or more of the processes and / or blocks Figure 1 specified in one or more of the blocks or blocks.

[0106] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the function specified in one or more of the processes and / or blocks Figure 1 one or more of the processes and / or blocks Figure 1 specified in one or more of the blocks or blocks.

[0107] Although preferred embodiments of the present invention have been described, additional changes and modifications can be made by those skilled in the art once they learn of the basic inventive concept. Therefore, the appended claims are intended to be construed to cover the preferred embodiments as well as all changes and modifications that fall within the scope of the embodiments of the present invention.

[0108] Finally, it should also be noted that, in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover a non-exclusive inclusion, such that a process, method, article or terminal device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the phrase "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or terminal device comprising the element.

[0109] The above has introduced in detail a method for processing a video conference and a system for processing a video conference provided by the present invention. Specific examples are used in this article to illustrate the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A method for processing video conferencing, characterized in that, The method includes: During the meeting, obtaining the audio data of each meeting terminal in real time; Generating structured meeting minutes data based on the audio data and a pre-trained large artificial intelligence model; Outputting the meeting minutes data.

2. The method according to claim 1, characterized in that, The generating of the structured meeting minutes data based on the audio data and the pre-trained large artificial intelligence model includes: Using speech recognition technology to convert the audio data into text data; Inputting the text data into the artificial intelligence model, supplementing the missing text content in the text data through the artificial intelligence model, and outputting the meeting minutes data.

3. The method according to claim 2, wherein The inputting of the text data into the artificial intelligence model, supplementing the missing text content in the text data through the artificial intelligence model, and outputting the meeting minutes data includes: Inputting the text data into the artificial intelligence model, analyzing the context relationship in the text data through the artificial intelligence model, and identifying and supplementing the missing text content in the text data according to the context relationship; Screening, classifying, and tagging the supplemented text data to obtain the meeting minutes data; Wherein, the meeting minutes data includes at least one of the following: meeting theme information, meeting key point information, to-do item information, decision-making item information, content tag information, theme duration information, meeting summary information.

4. The method according to any one of claims 1 to 3, characterized in that, The outputting of the meeting minutes data includes: Outputting the meeting minutes data in electronic form to each of the meeting terminals.

5. The method according to claim 4, characterized in that The outputting of the meeting minutes data in electronic form to each of the meeting terminals includes: Generating an access identifier pointing to the meeting minutes data and outputting the access identifier to each of the meeting terminals; Wherein, the access identifier includes: a QR code or a download link.

6. The method according to claim 1, wherein After the outputting of the meeting minutes data, the method further includes: Obtaining access request information for the meeting minutes data; Determining an access mode for the meeting minutes data according to the access request information, the access mode including a free mode or a paid mode; Providing the meeting minutes data based on the access mode.

7. The method according to claim 6, wherein The providing of the meeting minutes data based on the access mode includes: When the access mode is the paid mode, receiving payment information from a payment system, and when the payment information indicates successful payment, transmitting and displaying the meeting minutes data to the device corresponding to the acquisition request information.

8. A processing system for video conferencing, characterized in that, The system includes: An audio data acquisition module for obtaining the audio data of each meeting terminal in real time during a meeting; A meeting minutes generation module for generating structured meeting minutes data based on the audio data and a pre-trained large artificial intelligence model; A meeting minutes output module for outputting the meeting minutes data.

9. An electronic device, characterized in that, Including: One or more processors; And One or more machine-readable media having instructions stored thereon, which when executed by the one or more processors, cause the electronic device to execute the video conferencing processing method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The stored computer program causes the processor to execute the processing method of the video conference according to any one of claims 1 to 7.

Citation Information

Cited By

  • Audio transfer and summary generation method and device for conference

    CN120636411A

  • Conference data management method, terminal, server and conference management system

    CN120786017A