Video conference multi-mode real-time abstract generation method

By synchronously collecting multimodal data during video conferencing and performing real-time processing and feature fusion to generate dynamically updated summaries, the problems of information omission and insufficient real-time performance in existing technologies are solved, thereby improving the communication efficiency and decision support capabilities of the meeting.

CN120980187APending Publication Date: 2025-11-18SHENZHEN ZHONG XUN WANG LIAN SCI & TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511116733.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing video conferencing technologies cannot effectively integrate multimodal data, resulting in high information density, untimely recording, and easy omission of key information, failing to meet the need for real-time tracking of meeting progress and adjustment of discussion direction.

Method used

During video conferences, audio streams, video streams, and text chat logs are simultaneously captured. A multimodal fusion model is used for real-time preprocessing and feature extraction to generate a fusion feature set containing content relationships. Core topics, key conclusions, and key action items are identified, and real-time summaries are dynamically updated.

Benefits of technology

It achieves comprehensive coverage of meeting information, ensures the accuracy and timeliness of summaries, improves meeting communication efficiency, and provides a clear information framework and a basis for post-meeting action follow-up.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120980187A_ABST
    Figure CN120980187A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of video conference data processing, and discloses a video conference multi-mode real-time abstract generation method, which comprises the following steps of: synchronously acquiring an audio stream, a video stream and a text chat record of a conference, and converting the audio stream, the video stream and the text chat record into a time-aligned text, a key frame sequence and effective chat content through preprocessing; then text semantic features, visual scene features and interactive intention features are extracted, cross-modal correlation analysis is carried out through a multi-modal fusion model, and a fusion feature set is generated; and based on the identified core issue, the key conclusion and the action item, performing structured organization according to the time sequence and the importance degree to form a real-time abstract and performing dynamic updating. According to the method, multi-dimensional information is integrated, the one-sidedness problem of a traditional single-mode abstract is solved, the integrity, accuracy and timeliness of the abstract are improved, participants are assisted in mastering key points of a conference in real time, and the conference efficiency and decision-making quality are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of video conference data processing, in particular to a video conference multi-modal real-time summary generation method. BACKGROUND

[0002] With the popularity of remote work and collaboration, video conferencing has become the core way of communication for enterprises, institutions and teams. However, traditional video conferences often face problems such as high information density, untimely recording, and easy omission of key content. For example, participants need to record the main points during the meeting, which leads to a decrease in concentration; when summarizing the minutes after the meeting, it is often difficult to completely restore the discussion process due to fuzzy memory or fragmented information. Especially in the context of multi-participant and long-duration meetings, relying solely on manual recording not only is inefficient, but also may result in the loss of key information due to subjective bias.

[0003] Existing conference summary generation techniques mostly rely on single modal data (such as only based on audio transcription text), which has obvious limitations: text information cannot reflect non-verbal cues in videos (such as presentation content, body movement hints); instant feedback and supplementary points in text chat records are often ignored, resulting in insufficient completeness of the summary. In addition, traditional methods are mostly offline after the meeting, which cannot meet the needs of real-time follow-up of the meeting process and timely adjustment of the discussion direction, reducing the efficiency of meeting collaboration.

[0004] With the development of multi-modal technology, although some solutions attempt to integrate audio and text information, there are still technical bottlenecks in cross-modal data alignment and feature correlation analysis. For example, the time synchronization accuracy of audio transcription text and video pictures is insufficient, resulting in mismatch between semantic content and visual scene; the interaction intent mining of text chat records is not deep enough, making it difficult to capture the potential needs or emotional tendencies of participants. These problems make the conference summaries generated by existing technologies often lack accuracy, comprehensiveness and timeliness, which cannot effectively support conference decision-making and follow-up action. SUMMARY

[0005] The purpose of the present application is to provide a video conference multi-modal real-time summary generation method to solve the problems raised in the background.

[0006] To achieve the above purpose, the present application provides the following technical solution: a video conference multi-modal real-time summary generation method, the method comprising: During the video conference, the audio stream, video stream and text chat record of the conference site are synchronously collected as multi-modal input data; The collected multi-modal input data is pre-processed in real time, wherein the audio stream is converted into time-aligned text content, the video stream extracts key frame sequences, and the text chat record filters invalid information and labels timestamps; Feature extraction is performed on the preprocessed audio text content, video keyframe sequence and text chat record to form text semantic features, visual scene features and interactive intent features; Text semantic features, visual scene features, and interactive intent features are analyzed across modalities using a multimodal fusion model to generate a fusion feature set that includes content relationships. Based on the fusion feature set, the core topics, key conclusions and key action items in the meeting discussion are identified to form a key information set; The key information is organized in a structured manner according to the time sequence and importance of the information set, generating a real-time summary text that covers the main content of the meeting, and the summary content is dynamically updated as the meeting progresses.

[0007] Preferably, the acquired multimodal input data is preprocessed in real time. The specific steps are as follows: The audio stream is converted into text content sentence by sentence by a speech recognition model, while recording the start and end timestamps of each text segment. The video stream extracts keyframes at fixed time intervals, performs scene recognition on the keyframes and filters out duplicate or blurry images, retaining a frame sequence with content differentiation. Text chat logs will be edited to remove advertising information, emoticons, and content unrelated to the meeting topic. The remaining content will be labeled with the sending time and the user's identifier.

[0008] Preferably, feature extraction is performed on the preprocessed audio text content, video keyframe sequence, and text chat history. The specific steps are as follows: The audio text content is processed by word embedding through a pre-trained language model, and the semantic vector of each sentence is extracted as the text semantic feature. The video keyframe sequence extracts image features through a convolutional neural network, and combines them with a scene classification model to output scene labels for each frame as visual scene features; Text chat logs are processed using sentiment analysis and topic classification models to extract the sentiment and topic categories of the messages as features of interaction intent.

[0009] Preferably, textual semantic features, visual scene features, and interactive intent features are analyzed across modalities using a multimodal fusion model. The specific steps are as follows: Establish a time alignment mechanism to align text semantic features, visual scene features, and interactive intent features to the same timeline according to timestamps; The association weights between different modal features are calculated through an attention mechanism, focusing on associating text content with corresponding video scenes and chat records within the same time period. The associated features are spliced ​​and fused to generate a fused feature set containing information from the time dimension, content dimension, and interaction dimension.

[0010] Preferably, the core topics, key conclusions, and key action items in the meeting discussion are identified based on the fused feature set. The specific steps are as follows: By using a topic detection model to cluster the fused feature set into topics, the main topics discussed at the meeting and the discussion time periods for each topic can be determined. During the discussion period for each topic, key terms and conclusive statements that frequently occur are identified as key conclusions through a keyword extraction model. By analyzing the descriptions involving task allocation and time nodes using action recognition models, key action items that require follow-up can be extracted.

[0011] Preferably, the key information is organized in a structured manner according to the time sequence and importance of the information set. The specific steps are as follows: The core topics are arranged chronologically according to the meeting's progress, with each topic listed under its corresponding discussion time period. Under each topic, conclusions that have been repeatedly confirmed or are directly related to the topic's objectives are prioritized based on the frequency of their key findings and the clarity of their expression. Key action items are listed separately, with corresponding responsible persons and time requirements marked, forming a well-structured information list with clear hierarchy.

[0012] Preferably, a real-time summary text covering the main content of the meeting is generated and dynamically updated. The specific steps are as follows: The structured information list is converted into natural language text, and the content is organized using the expression logic of "time-issue-conclusion-action item". When the meeting enters a new discussion phase, the steps of multimodal data acquisition, preprocessing, feature extraction, fusion analysis, and key information identification are repeated to obtain new key information; Insert the newly added key information into the corresponding topic or action item position, adjust the structure and content of the original summary, and keep the summary in sync with the meeting process.

[0013] Preferably, when simultaneously collecting audio streams, video streams, and text chat logs from the meeting, the specific collection steps are as follows: The audio stream is captured through the microphone array of the conference terminal, and noise reduction algorithms are used to suppress ambient noise. The video stream is captured from multiple angles by the conference camera group, simultaneously recording the speaker's image and the projected content. Text chat logs are captured in real time through the conferencing software interface, including the speaker's ID, speaking time, and text content, forming a multi-source heterogeneous multimodal input data set.

[0014] Preferably, when calculating the association weights between different modal features using the attention mechanism, the specific calculation steps are as follows: Based on timestamps, text semantic features, visual scene features, and interactive intent features within the same time period are grouped into feature sets. For each set of features, calculate the matching degree between the text content and the visual elements; Calculate the consistency between the sentiment tendency of chat logs and the semantic sentiment of the text; Assign association weights based on matching degree and consistency results; higher weights indicate a closer association between modalities.

[0015] Preferably, when the audio stream is converted into text content sentence by sentence by the speech recognition model, the specific conversion steps are as follows: The audio stream is divided into fixed-duration speech segments, and endpoint detection is performed on each segment to identify the start and end positions of valid speech. Acoustic features are extracted from effective speech segments, and the acoustic features are converted into text candidate sequences through a pre-trained speech recognition model. The candidate text sequences are corrected using a language model, and the most reasonable text content is selected as the final conversion result by combining the contextual semantics.

[0016] Compared with the prior art, the beneficial effects of the present invention are: Compared to traditional single-modal processing methods, this invention simultaneously collects and analyzes multi-dimensional information. It can capture speech content through audio transcription, extract visual cues such as demonstration scenes and body language through video keyframes, and uncover real-time interaction intentions through text chat records. This achieves comprehensive coverage of meeting information and avoids summary bias caused by missing information.

[0017] In terms of real-time performance, this method uses a dynamic update mechanism to generate and adjust summary content as the meeting progresses. Participants can view the core topics and conclusions of the current discussion at any time, grasping key information without waiting for the meeting to end. This helps to clarify questions and adjust the direction of the discussion in a timely manner, significantly improving the efficiency of meeting communication. At the same time, the time alignment mechanism ensures the precise correlation of various modal features on the timeline. For example, the synchronous analysis of audio text and corresponding video footage enables the summary to accurately reflect contextual information such as "the viewpoint raised by a speaker when presenting a specific PPT slide," enhancing the traceability of the content.

[0018] Through cross-modal correlation analysis using a multimodal fusion model, this method can deeply uncover the intrinsic connections between different types of data. For example, matching the sentiment of text chat with the semantic sentiment of audio text can identify participants' agreement or skepticism towards a conclusion, ensuring that summaries not only contain objective content but also reflect subjective feedback during the discussion, providing richer reference for meeting decisions. Furthermore, the structured organization presents core topics, key conclusions, and action items in chronological order and according to their importance, clearly distinguishing the priority of each part of the information, facilitating participants to quickly locate key points, and providing a clear basis for post-meeting follow-up actions.

[0019] The preprocessing steps of this method reduce the interference of noisy data on feature extraction by filtering invalid information (such as advertisements and blurry images) and labeling with timestamps and user identifiers, thereby improving the accuracy of subsequent analysis. Furthermore, the application of pre-trained language models and convolutional neural networks ensures the efficient extraction of textual semantics, visual scene features, and other characteristics, laying a high-quality data foundation for multimodal fusion. Attached Figure Description

[0020] Figure 1 This is a schematic diagram illustrating the working principle of the video conferencing multimodal real-time summary generation method described in this invention. Figure 2 A flowchart for real-time preprocessing of multimodal input data; Figure 3 This is a flowchart for multimodal feature extraction; Figure 4 This is a flowchart of cross-modal correlation analysis for a multimodal fusion model; Figure 5 A flowchart for identifying core issues, key conclusions, and action items. Detailed Implementation

[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0022] Please see Figures 1-5 The present invention relates to a method for real-time multimodal video conferencing summary generation, the specific implementation steps of which are as follows: During the video conference, audio streams, video streams, and text chat logs are simultaneously collected as multimodal input data. The audio stream is captured via the conference terminal's microphone array, employing noise reduction algorithms to suppress ambient noise. The video stream is captured from multiple angles by the conference camera group, simultaneously recording the speaker's image and the projected content. Text chat logs are captured in real-time through the conference software interface, including the speaker's ID, speaking time, and text content, forming a multi-source, heterogeneous multimodal input data set.

[0023] The acquired multimodal input data undergoes real-time preprocessing. The audio stream is converted sentence-by-sentence into text content using a speech recognition model, while simultaneously recording the start and end timestamps for each text segment. Specifically, the audio stream is segmented into fixed-length speech segments, endpoint detection is performed on each segment to identify the effective start and end positions of the speech, acoustic features are extracted from the effective speech segments, and the acoustic features are converted into text candidate sequences using a pre-trained speech recognition model. The text candidate sequences are then corrected using a language model, and the most reasonable text content is selected as the final conversion result based on contextual semantics. Keyframes are extracted from the video stream at fixed time intervals. Scene recognition is performed on the keyframes, and duplicate or blurry images are filtered out, retaining frame sequences with content differentiation. Text chat logs are deleting advertising information, emoticons, and content unrelated to the meeting topic, and the remaining content is labeled with the sending time and the speaker's identifier.

[0024] Feature extraction was performed on the preprocessed audio text content, video keyframe sequences, and text chat logs to form text semantic features, visual scene features, and interaction intent features. The audio text content underwent word embedding processing using a pre-trained language model to extract the semantic vector of each sentence as text semantic features. The video keyframe sequences were processed using a convolutional neural network to extract image features, and combined with a scene classification model to output scene labels for each frame as visual scene features. The text chat logs were processed using a sentiment analysis model and a topic classification model to extract the sentiment tendency and topic category of the speech content as interaction intent features.

[0025] This paper employs a multimodal fusion model to perform cross-modal correlation analysis on text semantic features, visual scene features, and interactive intent features, generating a fused feature set containing content association relationships. A time alignment mechanism is established to align text semantic features, visual scene features, and interactive intent features to the same timeline based on timestamps. An attention mechanism is used to calculate the association weights between different modal features. Specifically, based on timestamps, text semantic features, visual scene features, and interactive intent features within the same time period are grouped into feature sets. For each feature set, the matching degree between text content and visual elements is calculated, and the consistency between chat history sentiment and text semantic sentiment is calculated. Association weights are assigned based on the matching degree and consistency results; higher weights indicate a tighter information association between modalities. The associated features are then concatenated and fused to generate a fused feature set containing information from time, content, and interaction dimensions.

[0026] Based on the fusion feature set, the core topics, key conclusions, and key action items in the meeting discussion are identified, forming a key information set. The fusion feature set is then used to perform topic clustering through a topic detection model to determine the main topics of the meeting discussion and the discussion time periods for each topic. Within the discussion time period for each topic, a keyword extraction model identifies frequently occurring key terms and conclusive statements as key conclusions. An action recognition model analyzes the statements involving task allocation and time nodes to extract key action items that require follow-up.

[0027] The key information is structured according to its chronological order and importance, generating real-time summary text covering the main content of the meeting, and dynamically updating the summary content as the meeting progresses. Core topics are arranged chronologically according to the meeting's progress, with corresponding discussion time periods listed under each topic. Under each topic, conclusions are prioritized based on their frequency of occurrence and clarity of expression, with conclusions that have been repeatedly confirmed or directly related to the topic's objectives being given priority. Key action items are listed separately, with corresponding responsible persons and time requirements indicated, forming a hierarchical structured information list. This structured information list is converted into natural language text, using a "time-topic-conclusion-action item" format to organize the content. When the meeting enters a new discussion phase, the multimodal data collection, preprocessing, feature extraction, fusion analysis, and key information identification steps are repeated to obtain new key information. This new key information is inserted into the corresponding topic or action item positions, adjusting the structure and content of the original summary to maintain synchronization between the summary and the meeting's progress.

[0028] Example 1: In this example, the collected multimodal input data is preprocessed in real time, specifically involving the processing of three types of data: audio stream, video stream, and text chat records.

[0029] In the audio stream processing stage, the audio stream is segmented into fixed-duration speech segments. The fixed duration can be set according to actual needs, for example, 10 seconds. This facilitates subsequent block processing of the audio. After segmentation, endpoint detection is performed on each speech segment. By analyzing the energy changes and frequency characteristics of the speech signal, the start and end positions of valid speech are identified. The purpose of this step is to eliminate silent segments without speech, retaining only the parts containing actual speech content, thereby improving the efficiency and accuracy of subsequent processing.

[0030] Acoustic features are extracted from the detected valid speech segments. Common acoustic features, such as Mel-frequency cepstral coefficients (MFCCs), can well describe the spectral characteristics of speech. By extracting these features, the speech signal can be converted into a digital feature vector that can be processed by a computer.

[0031] The extracted acoustic features are input into a pre-trained speech recognition model. This model can be an end-to-end speech recognition model based on deep learning, such as a model using a recurrent neural network (RNN) or Transformer architecture. These models, trained on a large amount of speech data, are able to convert acoustic features into text candidate sequences. However, due to factors such as environmental noise and non-standard pronunciation during the speech recognition process, the generated text candidate sequences may contain errors.

[0032] Language model refinement is required for the candidate text sequences. Using pre-trained language models, such as BERT or GPT series models, and combining contextual semantics, multiple candidate texts are analyzed and evaluated, and the text content that best conforms to semantic logic is selected as the final conversion result. After completing the speech-to-text conversion, the start and end timestamps of each text segment also need to be recorded for subsequent time alignment with data from other modalities.

[0033] For video stream processing, keyframes are first extracted at fixed time intervals. For example, one frame per second can be extracted, which reduces the amount of data while ensuring that the main content of the video is captured. After extracting the keyframes, scene recognition and filtering are required for these frames.

[0034] Scene recognition employs a scene classification model based on convolutional neural networks, such as ResNet or VGG. This model determines the scene type of each frame, such as a speech scene, a discussion scene, or a projected content display scene. Based on scene recognition, duplicate frames and blurry frames caused by poor focus or insufficient lighting are filtered out. Duplicate frames increase data redundancy, while blurry frames fail to provide effective visual information; therefore, frame sequences with content differentiation are retained to ensure that subsequent processing is based on clear and meaningful visual information.

[0035] In processing text chat logs, each message needs to be examined individually. First, remove advertising information, as this is irrelevant to the meeting topic and will interfere with subsequent analysis. Also, remove emoticons, as they are primarily used to express emotions and do not significantly contribute to the semantic analysis of the text.

[0036] By using methods such as keyword matching and topic classification, content irrelevant to the meeting topic is identified and deleted. For example, if the meeting topic is a discussion of project progress, then casual conversation unrelated to project progress needs to be filtered out. After filtering out invalid information, the remaining valid text chat records are labeled with the sending time and the user's identifier. This clarifies the chronological order and source of each record, providing necessary information for subsequent interaction intent analysis.

[0037] Throughout the preprocessing process, the processing of audio streams, video streams, and text chat logs is both independent and interconnected. Audio streams are converted into timestamped text content, providing a record of audio information for the meeting; keyframes are extracted from the video stream and scene recognition is performed, preserving key visual information; and text chat logs are filtered for invalid information and annotated with time and user data, supplementing the interactive information from the meeting. After preprocessing, all three types of data are timestamped, laying the foundation for subsequent multimodal feature extraction and fusion analysis. This allows for alignment and correlation analysis of data from different modalities over time, leading to a more comprehensive and accurate understanding of the meeting content. Through these detailed preprocessing steps, the raw multimodal input data is transformed into a more clearly structured and informationally more effective data format, providing high-quality data support for subsequent feature extraction, fusion analysis, and other steps, ensuring the accurate and efficient operation of the entire video conferencing multimodal real-time summarization method.

[0038] Example 2: In this example, feature extraction is performed on the preprocessed audio text content, video keyframe sequence and text chat records. This process involves feature transformation and representation of different modal data, aiming to extract feature vectors that can reflect the core information of various types of data, so as to provide a foundation for subsequent multimodal fusion analysis.

[0039] For preprocessed audio-text content, feature extraction is primarily achieved through pre-trained language models. After converting the audio stream into text content and undergoing preprocessing, the resulting text sentences are timestamped. This text content is then input into a pre-trained language model, such as the BERT (Bidirectional Encoder Representations from Transformers) model. These pre-trained language models are trained on large-scale corpora and are capable of learning deep semantic representations of language.

[0040] The pre-trained language model first performs word segmentation on the input text, breaking it down into individual lexical units. Then, it generates a corresponding word embedding vector for each lexical unit, which contains the semantic information of the word and its contextual relationship within the sentence. By processing the word embedding vectors of the entire sentence, for example using a multi-layer Transformer encoder in the model for feature extraction, the semantic vector for each sentence is finally obtained.

[0041] This semantic vector is a high-dimensional numerical vector that can comprehensively represent the semantic information of a sentence, such as its topic, sentiment, and logical relationships. For example, for a sentence like "This project needs to be completed by the end of the month," the semantic vector can capture key information such as "project progress," "end of the month," and "complete," as well as the relationships between them, thereby converting the text content into semantic features that a computer can understand and process.

[0042] In feature extraction from video keyframe sequences, preprocessing yields a retained keyframe sequence with content-discriminative characteristics. These keyframes require the extraction of visual scene features using computer vision methods. Each keyframe is then fed into a Convolutional Neural Network (CNN), such as ResNet (Residual Network) or VGG (Visual Geometry Group) network.

[0043] Convolutional neural networks (CNNs) possess powerful image feature extraction capabilities. Through multiple convolutional and pooling layers, they can extract features from images, ranging from low-level edges and textures to high-level object shapes and scene structures. After forward propagation through the CNN, the image feature vector of the keyframe is obtained. These feature vectors contain low-level visual information such as the image's color, texture, and shape.

[0044] These image features are further processed using a scene classification model. The scene classification model can be a CNN-based classification network that has already learned feature representations for different scene types during training. By inputting the image features extracted by the CNN into the scene classification model, the model can output the scene label for that keyframe, such as "speaker close-up," "projected content," or "conference table discussion scene."

[0045] These scene labels, as visual scene features, can accurately describe the meeting scene corresponding to the keyframe, providing a basis for subsequent analysis of the transition and distribution of different scenes during the meeting process. For example, scene labels can be used to determine whether the current meeting is in the stage of a presentation, discussion, or display of projected content.

[0046] For feature extraction from text chat logs, the preprocessing results in text records that have been filtered out of invalid information and labeled with time and user information. These records need to be processed using sentiment analysis models and topic classification models to extract interaction intent features.

[0047] Sentiment analysis models are used to analyze the sentiment tendency of spoken content. Common sentiment analysis models are based on recurrent neural networks (RNNs) or Transformer architectures. When text chat logs are input into a sentiment analysis model, the model analyzes the vocabulary, sentence structure, and other aspects of the text to determine whether the sentiment tendency of the spoken content is positive, negative, or neutral.

[0048] For example, a statement like "This solution is excellent, I fully support it" would be interpreted as positive emotion, while "This problem may be difficult to solve" might be interpreted as neutral or slightly negative emotion. Emotional characteristics reflect a speaker's attitude and emotional state towards the meeting content, which is crucial for understanding the interactive atmosphere and emotional changes of participants in a meeting.

[0049] Topic classification models are used to categorize text chat logs into different topic categories. These models can employ text classification algorithms, such as those based on TF-IDF (Term Frequency-Inverse Document Frequency) combined with Support Vector Machines (SVM), or deep learning-based text classification models.

[0050] When training a topic classification model, a large number of meeting chat logs need to be labeled first to divide them into different topic categories, such as "meeting agenda," "task assignment," "problem discussion," "progress report," and "resource request." Then, the preprocessed text chat logs are input into the topic classification model, which classifies them into the corresponding topic categories based on the lexical distribution and semantic features of the text content.

[0051] Topic category features can clearly define the category to which the core content of each chat log belongs. For example, if a chat log is categorized under the topic of "task assignment," it means that the log involves the assignment and arrangement of tasks. By extracting features based on sentiment and topic categories, we can extract the interactive intent features of the speakers from text chat logs, providing key feature information for subsequent analysis of interaction patterns and information flow in meetings.

[0052] Throughout the feature extraction process, semantic features of audio text content, visual scene features of video keyframes, and interactive intent features of text chat logs are extracted from the perspectives of language semantics, visual scene, and interactive intent, respectively. Figure ThreeThe meeting content was characterized across several dimensions. These features not only preserved key information from the original data but also transformed it into a unified numerical vector form, facilitating subsequent cross-modal correlation analysis using a multimodal fusion model. By extracting targeted features from different modalities, valuable information within each modality can be fully explored, laying the foundation for generating comprehensive and accurate real-time meeting summaries. This allows subsequent fusion analysis to better capture the relationships between different modalities, thereby more accurately identifying core topics, key conclusions, and key actions in the meeting.

[0053] Example 3: Text semantic features, visual scene features, and interactive intent features are analyzed across modalities using a multimodal fusion model. This process aims to integrate feature information from different modalities, establish relationships between them, and generate a fusion feature set containing multidimensional information, providing more comprehensive feature support for subsequent identification of key meeting information.

[0054] When conducting cross-modal association analysis, a time alignment mechanism is needed. Since audio text content, video keyframe sequences, and text chat logs are all timestamped during acquisition and preprocessing, these timestamps can be used to align features from different modalities onto the same timeline. Specifically, for each time point or time period, the text semantic features, visual scene features, and interaction intent features within that time period are correlated, enabling features from different modalities to correspond in the time dimension. This prepares the groundwork for subsequent analysis of the correlations between information from different modalities within the same time period.

[0055] The association weights between features from different modalities are calculated using an attention mechanism. This process uses timestamps as a benchmark, grouping textual semantic features, visual scene features, and interactive intent features within the same time period into feature groups. Each feature group contains feature information from different modalities within that time period, and the degree of correlation between these features needs to be analyzed.

[0056] For each feature group, the matching degree between the text content and the visual elements is first calculated. For example, does the description of objects and scenes mentioned in the text semantic features match the content in the visual scene features? In practice, the text semantic features can be converted into semantic vectors, and the image features in the visual scene features can also be converted into vector form. Then, the matching degree is determined by calculating the similarity between the two vectors. The higher the similarity, the higher the matching degree between the text content and the visual elements, and the stronger the association between them.

[0057] Simultaneously, the consistency between the sentiment tendency of the chat log and the semantic sentiment of the text is calculated. The interaction intent features of the text chat log contain sentiment information, while the semantic features of the audio text content also implicitly contain the text's sentiment information. The degree of consistency between these two sentiment tendencies is calculated by comparing them. For example, if the sentiment tendency of the chat log is positive, and the semantic sentiment of the corresponding time period is also positive, then the consistency is high; if the two sentiment tendencies are opposite, then the consistency is low. The consistency calculation can be achieved through the similarity calculation of sentiment feature vectors.

[0058] Based on the matching degree between text content and visual elements, and the consistency between the sentiment tendency of chat logs and the semantic sentiment of text, association weights are assigned between different modal features. The higher the matching degree and consistency, the greater the assigned association weight, indicating a closer information connection between the modalities, which requires more attention in subsequent fusion analysis.

[0059] After calculating the association weights, the associated features are concatenated and fused. Concatenation and fusion involves linking feature vectors from different modalities together in a specific order to form a new fused feature vector. This fused feature vector contains information from the time, content, and interaction dimensions. The time dimension is represented by timestamps, the content dimension includes meeting content information represented by textual semantic features and visual scene features, and the interaction dimension includes interaction information reflected by the interaction intent features of the text chat records.

[0060] By splicing and fusing features from different modalities, the features are integrated into a unified feature space, enabling the model to consider information from multiple modalities simultaneously and capture the relationships between them. For example, within a certain time period, textual semantic features describe discussions about project progress, visual scene features show the scene of the projected content, and interactive intent features from text chat logs indicate that the discussion involves task allocation. By fusing these features, a more comprehensive understanding can be gained that the meeting content within that time period involved discussions about task allocation related to project progress under a projected display.

[0061] The generated fusion feature set contains fusion feature vectors from various time periods. These feature vectors integrate multimodal information and can more comprehensively represent the conference content. The fusion feature set not only retains the original feature information of each modality, but also introduces the correlation between modalities through correlation analysis, providing rich feature information for subsequent identification of core topics, key conclusions, and key action items in the conference discussion based on the fusion feature set.

[0062] Throughout the cross-modal association analysis process, the time alignment mechanism ensures the correspondence between different modal features in the time dimension, the attention mechanism highlights key intermodal association information by calculating association weights, and the splicing and fusion integrates multimodal features into a unified feature representation. These three steps work together to enable the multimodal fusion model to effectively process and integrate feature information from different sources, generating a fused feature set containing content associations. This approach fully utilizes multimodal information from the meeting, avoiding the limitations of single-modal analysis, thus providing strong support for a more accurate understanding of the meeting content and the generation of comprehensive real-time meeting summaries. The fused feature set more comprehensively reflects the actual situation of the meeting, allowing subsequent key information identification steps to be based on richer and more accurate feature information, improving the accuracy and effectiveness of the entire video conferencing multimodal real-time summarization method.

[0063] Example 4: Identifying core topics, key conclusions, and important action items in meeting discussions based on fused feature sets. This process requires the use of multiple models to analyze the fused features and extract the key content of the meeting from multi-dimensional information. The following is a detailed explanation with specific examples.

[0064] Suppose a video conference's fusion feature set contains fusion features from multiple time periods, with each feature group corresponding to textual semantics, visual scene, and interactive intent features at a specific time. First, a topic detection model is used to cluster the fusion feature set into topics. For example, when the meeting is 10-20 minutes in, semantic vectors related to "project progress," "task allocation," and "development stage" appear repeatedly in the fusion features; visual scene features mainly show projected content; and interactive intent features frequently involve task allocation topics and positive sentiment in chat logs. A topic detection model (such as LDA) will cluster these similar features into one category, determining the main topic for this time period as "project progress arrangement discussion," and labeling the discussion period as 10:00-20:00.

[0065] After identifying the core topics, key conclusions need to be determined within the discussion timeframe for each topic. Taking the "Project Schedule Discussion" topic as an example, in the fusion feature analysis from 10:00 to 20:00, the audio text repeatedly contained phrases such as "complete requirements analysis by the end of the month," "start coding in the second week," and "testing is scheduled for the fourth week." Keyword extraction models (such as TextRank-based algorithms) extract frequently occurring key terms based on word frequency and semantic relationships in these texts, such as "requirements analysis," "coding," "testing phase," "by the end of the month," and "second week," and identify conclusive statements. For example, "requirements analysis must be completed by the 30th of this month" is identified as a key conclusion due to its high frequency and clear wording; similarly, "coding work will start on Monday of the second week" is listed as a key conclusion under this topic because it directly relates to the topic's objectives.

[0066] To extract key action items, an action recognition model is needed to analyze statements involving task allocation and timelines. Continuing with the above example, the audio text contains the statement, "Zhang San is responsible for writing the requirements analysis document and submitting it before March 15th," while the text chat log includes the statement, "Li Si, you follow up on the development of the coding module and ensure it starts within the second week." The action recognition model will analyze the action keywords ("responsible," "follow up"), responsible persons (Zhang San, Li Si), timelines (before March 15th, within the second week), and task content (writing the requirements analysis document, developing the coding module) in these statements. Through semantic analysis, the model extracts this information as key action items, such as "Zhang San needs to complete the writing of the requirements analysis document before March 15th" and "Li Si should start the development of the coding module within the second week."

[0067] For example, when the meeting was 30-40 minutes in, the text semantic features repeatedly mentioned "technical difficulties," "solutions," and "compatibility issues." The visual scene features included close-ups of the speaker and whiteboard discussions. The chat log topic was categorized as "problem discussion," and the sentiment was mostly neutral. The topic detection model clustered this into the topic "Discussion on Solutions to Technical Difficulties," with the discussion time period marked as 30:00-40:00. Under this topic, the keyword extraction model identified high-frequency terms such as "compatibility patch," "interface adjustment," and "testing plan" from the audio text, as well as conclusive statements such as "using a compatibility patch to solve system adaptation issues" and "interface adjustments need to be tested by next week," identifying these as key conclusions. The action recognition model extracted corresponding action items from statements such as "Wang Wu is responsible for developing the compatibility patch and testing it by March 20th" and "Zhao Liu coordinates testing resources for interface adjustments," clarifying the responsible persons and time requirements.

[0068] In another scenario, when the meeting was 50-60 minutes in, the text semantics in the fused features involved "meeting summary" and "next steps," the visual scene was a view of all participants, and the chat log topic in the interaction intent features was "progress report," with a positive sentiment. The topic detection model identified it as the "meeting summary and action plan" topic. The keyword extraction model extracted key conclusions from statements such as "this meeting determined the time nodes for requirements analysis and coding" and "each module leader needs to proceed according to the plan," while the action recognition model extracted action items from content such as "each group submits a progress report before next Monday" and "progress will be updated at next week's meeting," labeling the responsible person as "each group leader" and the time requirements as "before next Monday" and "next week's meeting."

[0069] Throughout the identification process, the topic detection model determines the topics and time periods through clustering and feature fusion, the keyword extraction model extracts key conclusions based on word frequency and semantics, and the action recognition model focuses on task allocation-related statements to extract action items. These three components work together to extract the core structure of the meeting—topics, conclusions, and action items—from the fused features. For example, in the above example, by analyzing the fused features of different time periods, three core topics were identified sequentially: "Discussion on Project Schedule," "Discussion on Solutions to Technical Challenges," and "Meeting Summary and Action Plan." Each topic corresponds to specific key conclusions and key action items, forming a complete set of key information. This information is not isolated but forms a logically coherent framework of key meeting content through the time dimension and content association within the fused features, providing a precise information foundation for subsequent structured organization and summary generation. In this way, the core elements of the meeting can be systematically extracted from complex multimodal fusion information, ensuring the completeness and accuracy of key information and providing effective support for users to quickly grasp the key points of the meeting.

[0070] Example 5: The key information is organized chronologically and according to its importance, generating a real-time summary text covering the main content of the meeting. This summary is dynamically updated as the meeting progresses. This process requires arranging the identified core issues, key conclusions, and key actions according to a logical sequence to form an easily understandable text structure. New information is continuously incorporated as the meeting continues, maintaining the real-time nature and accuracy of the summary. A detailed explanation is provided below with specific examples.

[0071] Suppose a video conference's key information set contains three core topics: "Project Schedule Discussion," "Technical Challenges and Solutions Discussion," and "Meeting Summary and Action Plan." Each topic corresponds to different key conclusions and action items. When organizing the information in a structured manner, the core topics are first arranged chronologically according to the meeting's progress. For example, the "Project Schedule Discussion" might take place 10-20 minutes after the meeting begins, the "Technical Challenges and Solutions Discussion" 30-40 minutes, and the "Meeting Summary and Action Plan" 50-60 minutes. Therefore, in the structured information list, these three topics are arranged chronologically, with each topic listing its corresponding discussion time slot, such as "Project Schedule Discussion (10:00-20:00)," "Technical Challenges and Solutions Discussion (30:00-40:00)," and "Meeting Summary and Action Plan (50:00-60:00)."

[0072] Under each topic, key conclusions are ranked according to their frequency of occurrence and clarity of expression. Taking the "Project Schedule Discussion" topic as an example, key conclusions include "Requirements analysis must be completed by the 30th of this month," "Coding work will begin on Monday of the second week," and "Testing is scheduled for the fourth week." The statement "Requirements analysis must be completed by the 30th of this month" is mentioned multiple times in both audio and text chat logs, and is clearly stated, directly related to the core objective of the project schedule. "Coding work will begin on Monday of the second week" is also repeatedly confirmed, with clear timelines and task content. While "Testing is scheduled for the fourth week" is an important conclusion, it appears relatively infrequently. Therefore, in the structured organization, the first two conclusions are prioritized, followed by "Testing is scheduled for the fourth week," forming the order: "Requirements analysis must be completed by the 30th of this month; Coding work will begin on Monday of the second week; Testing is scheduled for the fourth week."

[0073] Key action items are listed separately, with corresponding responsible persons and time requirements indicated. For example, action items extracted from content such as "Zhang San is responsible for writing the requirements analysis document and submitting it before March 15th," "Li Si follows up on the development of the coding module and ensures it starts within the second week," and "Wang Wu is responsible for developing the compatibility patch and testing it before March 20th" will be categorized separately in the structured information list, forming the following structure: Zhang San needs to complete the requirements analysis document by March 15th; Li Si should start developing the coding module within the second week; Wang Wu needs to complete the development and testing of the compatibility patch by March 20th.

[0074] Once all key information is structured according to chronological order and importance, the structured information list needs to be converted into natural language text. The content should be organized using a "time-issue-conclusion-action" format, for example: "The meeting from 10:00 to 20:00 discussed the project schedule, clarifying that requirements analysis must be completed by the 30th of this month, coding will begin on the Monday of the second week, and testing will be scheduled for the fourth week. Corresponding action items include Zhang San needing to complete the requirements analysis document by March 15th, and Li Si needing to start coding module development within the second week. From 30:00 to 40:00, discussions focused on solutions to technical difficulties, concluding that compatibility patches should be used to resolve system adaptation issues, and interface adjustments need to be tested by next week. The corresponding action item is Wang Wu being responsible for developing the compatibility patch and testing it by March 20th. From 50:00 to 60:00, the meeting was summarized, confirming that each module leader must proceed with the work as planned, each group must submit a progress report by next Monday, and progress will be updated at next week's regular meeting." As the meeting progresses and new discussion phases begin, the steps of multimodal data acquisition, preprocessing, feature extraction, fusion analysis, and key information identification need to be repeated to obtain new key information and dynamically update the summary content. For example, assuming the meeting enters the "resource allocation discussion" phase after 60 minutes, multimodal data processing identifies the core topic of this phase as "resource allocation discussion (60:00-70:00)," with key conclusions including "development resources should be tilted towards the coding module" and "testing resources need to be in place by the third week." Key action items are "Zhao Liu coordinates the allocation of development resources, to be completed before March 12" and "Qian Qi secures testing resources, to be deployed starting in the third week."

[0075] At this point, the newly added key information needs to be inserted into the corresponding positions, and the structure and content of the original summary need to be adjusted. The last part of the original summary contains the relevant content of "Meeting Summary and Action Plan (50:00-60:00)," and the newly added topic "Resource Allocation Discussion (60:00-70:00)" should follow it in chronological order. At the same time, key conclusions and action items should be arranged under the new topic, and the expression logic of the summary text should be updated to include the latest meeting content: "...50:00-60:00 Meeting summary, confirming that the heads of each module need to advance the work according to the plan, each group should submit a progress report by next Monday, and the progress will be updated at next week's regular meeting. 60:00-70:00 Further discussion on resource allocation, clarifying that development resources should be tilted towards the coding module, and testing resources should be in place by the third week. The corresponding action items are: Zhao Liu coordinates the allocation of development resources, to be completed by March 12; Qian Qi secures testing resources, to be deployed starting in the third week." During dynamic updates, it's crucial to ensure logical consistency between new and existing information. The adjusted summary structure should maintain clear hierarchy, chronological order, and highlight key information. For instance, if new action items have a temporal or dependent relationship with previous action items, this relationship should be appropriately reflected in the summary text. However, no additional explanation of the logical relationship is needed; simply arrange them according to chronological order and importance. Furthermore, when new key conclusions supplement or revise existing conclusions, the original content must be replaced or adjusted to maintain the accuracy of the summary. For example, if the conclusion that "the testing phase is scheduled for the fourth week" is revised in subsequent discussions, and the testing phase is moved to the latter half of the third week, this conclusion needs to be updated in the summary, and the time requirements in the action items adjusted accordingly.

[0076] Through this structured organization and dynamic updating approach, the generated real-time summary text comprehensively covers the main content of the meeting, clearly presenting each topic, key conclusions of each topic, and key actions requiring follow-up. Furthermore, it continuously updates as the meeting progresses, remaining synchronized with the overall meeting schedule. By reading the summary text, users can quickly grasp the overall framework and key points of the meeting without needing to review the entire meeting, thus improving information acquisition efficiency. At the same time, the structured organization makes the summary content easy to understand and review, with key information readily apparent, meeting users' actual needs for real-time meeting summaries.

[0077] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0078] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for real-time summarization of multimodal video conferencing, characterized in that, Includes the following steps: During the video conference, audio streams, video streams, and text chat logs from the meeting are collected simultaneously as multimodal input data. The collected multimodal input data is preprocessed in real time, including converting audio streams into time-aligned text content, extracting keyframe sequences from video streams, and filtering invalid information and adding timestamps to text chat records. Feature extraction is performed on the preprocessed audio text content, video keyframe sequence and text chat record to form text semantic features, visual scene features and interactive intent features; Text semantic features, visual scene features, and interactive intent features are analyzed across modalities using a multimodal fusion model to generate a fusion feature set that includes content relationships. Based on the fusion feature set, the core topics, key conclusions and key action items in the meeting discussion are identified to form a key information set; The key information is organized in a structured manner according to the time sequence and importance of the information set, generating a real-time summary text that covers the main content of the meeting, and the summary content is dynamically updated as the meeting progresses.

2. The method for generating real-time multimodal video conferencing summaries according to claim 1, characterized in that, The collected multimodal input data undergoes real-time preprocessing, and the specific steps are as follows: The audio stream is converted into text content sentence by sentence by a speech recognition model, while recording the start and end timestamps of each text segment. The video stream extracts keyframes at fixed time intervals, performs scene recognition on the keyframes and filters out duplicate or blurry images, retaining a frame sequence with content differentiation. Text chat logs will be edited to remove advertising information, emoticons, and content unrelated to the meeting topic. The remaining content will be labeled with the sending time and the user's identifier.

3. The method for generating real-time multimodal video conferencing summaries according to claim 1, characterized in that, Feature extraction was performed on the preprocessed audio text content, video keyframe sequences, and text chat logs. The specific steps are as follows: The audio text content is processed by word embedding through a pre-trained language model, and the semantic vector of each sentence is extracted as the text semantic feature. The video keyframe sequence extracts image features through a convolutional neural network, and combines them with a scene classification model to output scene labels for each frame as visual scene features; Text chat logs are processed using sentiment analysis and topic classification models to extract the sentiment and topic categories of the messages as features of interaction intent.

4. The method for generating real-time multimodal video conferencing summaries according to claim 3, characterized in that, The text semantic features, visual scene features, and interaction intent features are analyzed across modalities using a multimodal fusion model. The specific steps are as follows: Establish a time alignment mechanism to align text semantic features, visual scene features, and interactive intent features to the same timeline according to timestamps; The association weights between different modal features are calculated through an attention mechanism, focusing on associating text content with corresponding video scenes and chat records within the same time period. The associated features are spliced ​​and fused to generate a fused feature set containing information from the time dimension, content dimension, and interaction dimension.

5. The method for generating real-time multimodal video conferencing summaries according to claim 4, characterized in that, The core topics, key conclusions, and important action items in the meeting discussions are identified based on a fusion feature set. The specific steps are as follows: By using a topic detection model to cluster the fused feature set into topics, the main topics discussed at the meeting and the discussion time periods for each topic can be determined. During the discussion period for each topic, key terms and conclusive statements that frequently occur are identified as key conclusions through a keyword extraction model. By analyzing the descriptions involving task allocation and time nodes using action recognition models, key action items that require follow-up can be extracted.

6. The method for generating real-time multimodal video conferencing summaries according to claim 5, characterized in that, The key information is organized in a structured manner according to its chronological order and importance. The specific steps are as follows: The core topics are arranged chronologically according to the meeting's progress, with each topic listed under its corresponding discussion time period. Under each topic, conclusions that have been repeatedly confirmed or are directly related to the topic's objectives are prioritized based on the frequency of their key findings and the clarity of their expression. Key action items are listed separately, with corresponding responsible persons and time requirements marked, forming a well-structured information list with clear hierarchy.

7. The method for generating real-time multimodal video conferencing summaries according to claim 6, characterized in that, The specific steps for generating and dynamically updating real-time summary text covering the main content of the meeting are as follows: The structured information list is converted into natural language text, and the content is organized using the "time-issue-conclusion-action item" expression logic; When the meeting enters a new discussion phase, the steps of multimodal data acquisition, preprocessing, feature extraction, fusion analysis, and key information identification are repeated to obtain new key information; Insert the newly added key information into the corresponding topic or action item position, adjust the structure and content of the original summary, and keep the summary in sync with the meeting process.

8. The method for generating real-time multimodal video conferencing summaries according to claim 1, characterized in that, When simultaneously capturing audio streams, video streams, and text chat logs from a meeting, the specific capture steps are as follows: The audio stream is captured through the microphone array of the conference terminal, and noise reduction algorithms are used to suppress ambient noise. The video stream is captured from multiple angles by the conference camera group, simultaneously recording the speaker's image and the projected content. Text chat logs are captured in real time through the conferencing software interface, including the speaker's ID, speaking time, and text content, forming a multi-source heterogeneous multimodal input data set.

9. A method for generating real-time multimodal video conferencing summaries according to claim 4, characterized in that, When calculating the association weights between features of different modalities using the attention mechanism, the specific calculation steps are as follows: Based on timestamps, text semantic features, visual scene features, and interactive intent features within the same time period are grouped into feature sets. For each set of features, calculate the matching degree between the text content and the visual elements; Calculate the consistency between the sentiment tendency of chat logs and the semantic sentiment of the text; Assign association weights based on matching degree and consistency results; higher weights indicate a closer association between modalities.

10. A method for generating real-time multimodal video conferencing summaries according to claim 2, characterized in that, When an audio stream is converted into text content sentence by sentence by a speech recognition model, the specific conversion steps are as follows: The audio stream is divided into fixed-duration speech segments, and endpoint detection is performed on each segment to identify the start and end positions of valid speech. Acoustic features are extracted from effective speech segments, and the acoustic features are converted into text candidate sequences through a pre-trained speech recognition model. The candidate text sequences are corrected using a language model, and the most reasonable text content is selected as the final conversion result by combining the contextual semantics.

Citation Information

Cited By

  • Conference information intelligent summary generation system

    CN121503500A