Real-time conference summary optimization system
By combining brain-computer interfaces and multimodal large language models, meeting minutes are optimized in real time, solving the problems of information omission and insufficient understanding in traditional recording technologies, and achieving efficient and accurate meeting recording and content optimization.
Patent Information
- Application Number
- CN202510873997.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-11-21
AI Technical Summary
Existing meeting minutes recording technologies rely on manual methods or speech-to-text conversion, resulting in information omissions and a lack of in-depth understanding, making it difficult to achieve real-time optimization and accurate extraction in meeting scenarios.
By combining brain-computer interface technology and multimodal large language models, the system monitors participants' focus and emotions through an EEG acquisition module, and analyzes meeting content using a speech recognition module and a multimodal large language model to optimize meeting minutes, including content structure and contextual analysis.
It enables real-time and accurate meeting minutes recording, improving the accuracy and efficiency of meeting minutes, and can handle multilingual and complex content, adapting to the different thinking states of participants.
Smart Images

Figure 20O65RPATJLZWXG26ULR1FDVCGMN0MYBDACWPPIE 
Figure 3Q0G1YMHV0PBIL4WT7MR8HXH7EJ5EZSK5H7BRJR9 
Figure 4JNFTC7ZVJR8YTMJNDCFOIZM5SLNRD9BG69AUSWS
Abstract
Description
Technical Field
[0001] This invention relates to a meeting minutes system, and more specifically to a real-time meeting minutes optimization system. Background Technology
[0002] With the accelerating pace of modern work and the growing demand for remote collaboration, efficient and accurate meeting minutes recording has become a key element in improving work efficiency. Traditional meeting minutes rely on manual recording, which suffers from problems such as information omissions and subjective interpretation biases. While existing speech-to-text technology can automate recording, it lacks a deep understanding of the meeting content and contextual analysis, making it difficult to accurately extract the core points of the meeting. Meanwhile, brain-computer interface technology has demonstrated its potential to perceive users' thought states in the field of human-computer interaction, and multimodal large language models also possess powerful semantic understanding and content generation capabilities. However, currently, no technical solution deeply integrates these two approaches, failing to meet the needs of real-time optimization of meeting minutes in meeting scenarios, achieving accurate recording, efficient extraction, and continuous iterative optimization. Summary of the Invention
[0003] To address the shortcomings of existing technologies, the present invention aims to provide a real-time meeting minutes optimization system and method based on brain-computer interface technology and multi-model large language model. This system can automatically record meeting content and optimize meeting minutes based on the user's EEG signals, thought state, and multi-language model analysis results, thereby improving the accuracy, contextual relevance, and efficiency of meeting recordings, and providing data and capabilities for continuous analysis and optimization.
[0004] To achieve the above objectives, the present invention provides the following technical solution: a real-time meeting minutes optimization system, comprising: The EEG acquisition module monitors the participants' EEG signals in real time through smart wearable devices, processes the acquired EEG signals, and analyzes the participants' focus, emotional fluctuations, and subconscious four-dimensional capture. The speech recognition module captures audio data from the meeting via a microphone or voice interface and performs real-time conversion using speech-to-text technology. A multi-model large language model performs content analysis, summary extraction, topic classification, and key information identification based on context to understand the discussion content of the participants, generate or optimize meeting minutes in real time, and ensure the contextual coherence and accuracy of the content. The content optimization module intelligently optimizes the content of meeting minutes by combining the analysis results of EEG signals and multi-model large language models. The cloud storage module uploads the real-time generated meeting minutes to the cloud in an encrypted manner, ensuring the security and accessibility of meeting data at any time.
[0005] As a further improvement of the present invention, the electroencephalogram (EEG) acquisition module includes: Multiple electrode sensors are attached to the external human body to detect brain electrical activity and output signals; The data processing unit is connected to multiple electrode sensors to receive brain activity signals, perform frequency band analysis, and divide the brain wave signals into different frequency bands. Common brain wave frequency bands include: δ, θ, α, β, and γ. By analyzing the changes in these different frequency bands, feature extraction and fusion processing of multidimensional cognitive data are performed.
[0006] As a further improvement of the present invention, the specific steps of the data processing unit performing frequency band analysis are as follows: Step 1: First, perform quantitative modeling of focus by constructing a focus index based on the ratio of the power spectral density of the theta wave to the beta wave, and then use a sliding window to calculate the AI curve in real time. Step 2: Multi-scale feature fusion is performed. The front end extracts gamma wave energy features, and the back end performs temporal emotion classification based on LSTM. Then, the emotion label is used as an additional feature vector and input into the Transformer encoder to build an emotion-content decoupling network for adversarial training. Step 3: Capture preconscious thought processes, establish a mapping model between N400 event-related potentials and concept associations, and detect unconventional semantic associations through a pre-trained concept network.
[0007] As a further improvement of the present invention, the power spectral densities of the θ wave and β wave in step one are calculated in the following manner: First, the spectrum of the EEG signal is calculated based on the Fast Fourier Transform (FFT): Where X(f) is the frequency domain signal, x(t) is the time domain signal, f is the frequency, and T is the number of sampling points; Secondly, the power spectral density of the EEG signal is calculated by segmenting the frequency band: Where P is the power in the α-wave band.
[0008] As a further improvement of the present invention, the specific steps of the speech recognition module in performing real-time conversion using speech-to-text technology are as follows: Step four: First, the speech signal is preprocessed, and then features are extracted using short-time Fourier transform or Mel frequency cepstral coefficients. Step 5: Use a deep learning-based acoustic model to map audio features to text.
[0009] As a further improvement of the present invention, the multi-model large language model adopts the Transformer model, which has the following self-attention mechanism: Where Q is the query matrix, K is the key matrix, V is the value matrix, and dk is the dimension of the key; The pre-training objectives are as follows: The BERT model is trained using masked language modeling and next-sentence prediction: Where wt is the current word, and w1, w2, ..., wT are the context words.
[0010] As a further improvement of the present invention, the content optimization module optimizes the content in the following way: Attention analysis: Based on EEG analysis results, automatically adjust the content structure and summary of meeting minutes to ensure that key information is highlighted; Context optimization: Based on the contextual understanding capabilities of the language model, the syntax and expression are optimized and rewritten to ensure the fluency and accuracy of the meeting minutes, as detailed below: Use a language model to calculate the confidence score of the sentence and determine whether correction is needed: Where wi is a word, n is the sentence length, and P(wi|w1,w2,…,w{i-1}) is the word probability based on context; Topic optimization: Based on topic models (such as LDA or BERT), extract the core topics of the meeting and optimize the meeting minutes: Where z represents the topic, w represents the word, and P(w|z) represents the word distribution under the topic.
[0011] The beneficial effects of this invention are that it can accurately record and optimize meeting content in real time by combining EEG technology with a multi-model large language model, avoiding the omission of important information and improving the quality of meeting minutes. Through the collaborative work of multi-language models, the system can handle complex content in different languages, dialects, and contexts, making the generation of meeting minutes more accurate and efficient. The system can also intelligently adjust the meeting recording strategy according to the participants' thinking state, enhancing the relevance, completeness, and operability of the meeting minutes. Detailed Implementation
[0012] The present invention will be further described in detail below with reference to the given embodiments.
[0013] This embodiment of a real-time meeting minutes optimization system includes the following modules: Brainwave Acquisition Module: This module monitors participants' brainwave signals in real time using smart wearable devices (such as brainwave headphones and smart glasses). Electroencephalography (EEG) signals reflect the electrical activity of neurons in the cerebral cortex. The system acquires these signals through sensors and processes them to analyze participants' focus, emotional fluctuations, and subconscious four-dimensional capture. In this embodiment, the brainwave acquisition module primarily performs the following actions: EEG signal acquisition: Multiple electrode sensors are used to detect brain electrical activity and transmit the signals to the data processing unit.
[0014] Frequency band analysis: EEG signals are mainly divided into different frequency bands. Common EEG frequency bands include: delta (deep sleep), theta (relaxation, meditation), alpha (mild relaxation), beta (active state), and gamma (high cognitive state). By analyzing the changes in these different frequency bands, feature extraction and fusion processing of multidimensional cognitive data can be performed.
[0015] Further: 1. Focus on metric modeling Dynamic weight allocation algorithm: Constructs an Attention Index (AI) based on the power spectral density ratio of theta waves (4-8Hz) to beta waves (12-30Hz), and calculates the AI curve in real time using a sliding window. Application scenarios: Real-time content priority tagging: When AI exceeds a threshold, the current speech is automatically marked as a "high-value segment," triggering a multimodal big data model to perform deep semantic analysis on that segment. Attention heatmap generation: Constructing a meeting attention distribution map by combining timestamps for intelligent summarization of subsequent minutes. 2. Emotional State Decoding Network Multi-scale feature fusion architecture: Front end: Extraction of energy characteristics of gamma waves (30-100Hz) (emotional intensity) Backend: LSTM-based temporal sentiment classification (calm / excitement / confusion / resistance) Emotion enhancement processing: Sentiment label embedding: Inputting sentiment labels as additional feature vectors into the Transformer encoder Adversarial training mechanism: Constructing an emotion-content decoupled network to ensure that semantic understanding is not affected by emotional fluctuations. 3. Preconscious thought capture Decoding slow cortical potentials (SCP): Establish a mapping model between N400 event-related potentials and concept correlation. Detecting unconventional semantic associations using a pre-trained concept network. Application method: Potential demand prediction: Generate relevant background information pop-ups based on neural characteristics before the user explicitly expresses their needs. Thought Leap Compensation: When a discrete concept is detected to be active, logical connection statements are automatically added. formula: Electroencephalogram (EEG) spectral analysis: Calculating the spectrum of brainwave signals based on Fast Fourier Transform (FFT): Where X(f) is the frequency domain signal, x(t) is the time domain signal, f is the frequency, and T is the number of sampling points.
[0016] Brainwave frequency band power calculation: By segmenting the brainwave signal into frequency bands and calculating the power spectral density, the intensity of a specific frequency band can be determined. For example: Where P is the power in the α-wave band (8-12Hz).
[0017] Speech Recognition Module: This module captures audio data from the meeting via a microphone or voice interface and performs real-time conversion using Speech-to-Text (STT) technology. This module typically relies on deep learning models such as Recurrent Neural Networks (RNNs), Long Short-Term Memory Networks (LSTMs), or Transformer architectures. This module can handle multiple languages and dialects, ensuring efficient transcription of speech data. In this embodiment, the speech recognition module performs the following actions: Feature Extraction: After preprocessing, the speech signal is used for feature extraction via Short-Time Fourier Transform (STFT) or Mel-frequency cepstral coefficients (MFCC). Acoustic Model: Deep learning-based acoustic models (such as LSTM and Transformer) are used to map audio features to text. Language Model: A context-based language model is used to post-process the transcription results, ensuring the accuracy and fluency of the text. formula: MFCC Feature Extraction: By performing a Fourier transform on the audio signal, converting it into a frequency domain signal, the Mel-frequency cepstral coefficients are calculated: Where (x(t)) is the time-domain audio signal, and (DFT) is the Discrete Fourier Transform.
[0018] LSTM network formula: The core calculation formula of LSTM: Where ft, it, and ot are gating vectors, ht is the output, xt is the input data, and c_t is the cell state.
[0019] Multi-model Large Language Model: This module includes multiple large language models (such as GPT, BERT, etc.). These models can perform content analysis, summary extraction, topic classification, and key information identification based on context. The large language model can understand the discussion content of participants, generate or optimize meeting minutes in real time, and ensure the contextual coherence and accuracy of the content. The specific details of this large language model are as follows: Transformer model: It adopts the Transformer structure and captures long-distance dependencies in the input sequence through a self-attention mechanism.
[0020] Multi-task learning: By combining different language tasks (such as summary generation, sentiment analysis, grammar correction, etc.), the system generates high-quality meeting minutes through the collaborative work of multiple models.
[0021] formula: Self-attention mechanism: In Transformer, the formula for calculating the self-attention mechanism is as follows: Where Q is the query matrix, K is the key matrix, V is the value matrix, and dk is the dimension of the key.
[0022] BERT's pre-training objective: The BERT model is trained using Masked Language Model (MLM) and Next Sentence Prediction (NSP): Where wt is the current word, and w1, w2, ..., wT are the context words.
[0023] Content optimization module: Combining the analysis results of EEG signals and multi-model large language models, the system intelligently optimizes the content of meeting minutes. For example, when the system detects that a certain topic has attracted high attention from participants, it will automatically highlight that topic and adjust the focus of the minutes. Furthermore, based on the reasoning ability of the large language model, the system can automatically correct grammar, sentence structure, and vocabulary usage, improving the readability and professionalism of the minutes. The content optimization process is as follows: Attention Analysis: Based on EEG analysis results, automatically adjust the content structure and summary of meeting minutes to ensure that key information is highlighted.
[0024] Context optimization: Based on the contextual understanding capabilities of the language model, the grammar and expressions are optimized and rewritten to ensure the fluency and accuracy of the language in the meeting minutes.
[0025] formula: Grammar correction: Use a language model to calculate the confidence score of the sentence and determine whether correction is needed: Where (wi) is a word, (n) is the sentence length, and (P(wi|w1,w2,…,w{i-1})) is the context-based word probability.
[0026] Theme optimization: Based on topic models (such as LDA or BERT), extract the core topics of the meeting and optimize the meeting minutes: Where z represents the topic, w represents the word, and P(w|z) represents the word distribution under the topic.
[0027] The optimization process was accomplished through the following engines and systems working together: Dynamic content optimization engine Attention-guided summary compression: Python defdynamic_compress(text,ai_score): ifai_score>0.8: returngraph_attention_summarize(text,compression_ratio=0.3) else: returnbert_extractive_summarize(text,compression_ratio=0.6) Optimization of the expression of emotion perception: During periods of high stress, softening language ("Further discussion is recommended...") is automatically inserted. Automatic annotation of knowledge points is triggered when a confused state is detected. Cognitive State Visualization System 3D Decision Dashboard: X-axis: Time dimension Y-axis: Cognitive load intensity Z-axis: Emotional polarity Interactive correction mechanism: Allows users to verify the accuracy of the minutes by replaying neural feature data from a specific time period.
[0028] Personalized cognitive profile Create user-specific: Attention baseline model Emotion Response Pattern Library Conceptual association feature space accomplish: Automatic adaptation of meeting style Analysis of long-term cognitive ability development trends.
[0029] Cloud storage module: Real-time generated meeting minutes are uploaded to the cloud in an encrypted manner, ensuring the security and accessibility of meeting data at any time. Cloud storage supports simultaneous access from multiple devices, facilitating viewing and sharing by participants.
[0030] Data encryption: Encryption algorithms (such as AES) are used to encrypt meeting minutes to ensure data security during storage and transmission.
[0031] Cloud storage service: Utilize the storage services provided by the cloud platform to store the minutes data in a distributed system, ensuring high availability and scalability.
[0032] Based on the modules mentioned above, the meeting minutes can be optimized.
[0033] Specifically, the following application examples are provided in this embodiment: Example 1: In an online meeting of a multinational corporation, participants used smart headsets and microphones with brain-computer interfaces. The system was able to collect participants' brainwave signals and speech data. Using a multilingual large language model, the system converted the speech signals into text and generated meeting minutes in real time. Simultaneously, the system analyzed attention changes based on participants' brainwave data, automatically highlighting key discussion points. By leveraging neural features to bypass language barriers and directly achieve conceptual understanding, it generated clearly structured meeting records.
[0034] Example 2: In a multi-party remote meeting, the system acquires users' attention and emotion data through a wearable brain-computer interface device, and combines this data with a multi-model large language model to transcribe the meeting content in real time and optimize the minutes. Based on the different thought states of the participants, the system adjusts the presentation of the meeting minutes to ensure that the core issues of concern to each participant are highlighted and improved, achieving a non-voice recording mode that allows for direct "think-write" input.
[0035] Example 3: Creative Brainstorming Sessions: Capturing Fleeting Preconscious Creative Ideas In summary, the real-time meeting minutes optimization system of this embodiment combines the powerful capabilities of brain-computer interface technology, speech recognition technology, and large language models to improve the quality of automatic generation and optimization of meeting minutes. By collecting the brainwave signals of participants in real time and combining them with multilingual models for speech transcription and content analysis, efficient meeting minutes are customized for each participant, ensuring the accuracy and adaptability of the content.
[0036] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should also be considered within the scope of protection of the present invention.
Claims
1. A real-time meeting minutes optimization system, characterized in that: Comprise: Electroencephalogram acquisition module, through intelligent wearable device real-time monitoring and the brain waves of the participants, and the collected brain electrical signals are signal processing, analysis and focus of the participants, emotional fluctuations and subconscious four-dimensional capture; Voice recognition module, through the microphone or voice interface to capture the audio data in the meeting, and use speech to text technology for real-time conversion; Multi-model large language model, the multi-model large language model according to the context of content analysis, abstract extraction, theme classification, key information recognition, to understand the discussion content of the participants, real-time generation or optimization of meeting minutes, ensure the context coherence and accuracy of the content; Content optimization module, combined with the analysis results of brain wave signals and multi-model large language model, intelligent optimization of the content of the meeting minutes; Cloud storage module, the real-time generated meeting minutes are uploaded to the cloud through encryption method, ensure the safety and accessibility of the meeting data.
2. The live meeting minutes optimization system of claim 1, wherein: The electroencephalogram acquisition module comprises: A plurality of electrode sensors, attached to the outside of the human body, for detecting brain electrical activity and outputting signals; Data processing unit, connected with the plurality of electrode sensors, to receive brain electrical activity signals, perform frequency band analysis, and divide the electroencephalogram signals into different frequency bands, common electroencephalogram frequency bands include: δ, θ, α, β, γ, by analyzing the changes of these different frequency bands, multi-dimensional cognitive data feature extraction and fusion processing.
3. The live meeting minutes optimization system of claim 2, wherein: The specific steps of frequency band analysis of the data processing unit are as follows: Step one, first, focus quantization modeling, by the power spectral density ratio of θ wave and β wave to construct the focus index, using sliding window real-time calculation AI curve; Step two, multi-scale feature fusion, front-end γ wave energy feature extraction, back-end LSTM-based time sequence emotion classification, then the emotion label is input as an additional feature vector into the Transformer encoder, to build an emotion-content decoupling network for adversarial training; Step three, preconscious thought capture, establish the mapping model of N400 event-related potential and concept correlation, through the pre-trained concept network to detect unconventional semantic association.
4. The real-time meeting minutes optimization system of claim 3, wherein: The power spectral density of θ wave and β wave in step one is calculated as follows: First, calculate the frequency spectrum of electroencephalogram signal based on fast Fourier transform (FFT): Where X(f) is the frequency domain signal, x(t) is the time domain signal, f is the frequency, and T is the number of sampling points; Second, the power spectral density is calculated by frequency band segmentation of electroencephalogram signal: where P is the power of the alpha band.
5. The live meeting minutes optimization system of any one of claims 1 to 4, wherein: The specific steps of real-time conversion of the voice recognition module using speech to text technology are as follows: Step four, first, pre-process the speech signal, then extract features through short-time Fourier transform or mel frequency cepstral coefficient; Step five, use deep learning-based acoustic model to map audio features to text.
6. The live meeting minutes optimization system of any one of claims 1 to 4, wherein: The multi-model large language model uses Transformer model, which has the following self-attention mechanism: Where Q is the query matrix, K is the key matrix, V is the value matrix, and dk is the dimension of the key; Has the following pre-training target: BERT model is trained using Masked Language Modeling and Next Sentence Prediction: where wt is the current word and w1, w2, …, wT are the context words.
7. The live meeting minutes optimization system of any one of claims 1 to 4, wherein: The content optimization module optimizes the content as follows: Attention analysis: Based on the electroencephalogram analysis results, automatically adjust the content structure and summary content of the meeting minutes to ensure that the key information is highlighted. Context optimization: Based on the context understanding ability of the language model, optimize and rewrite the syntax and expression to ensure the fluency and accuracy of the meeting minutes language, as follows: Use the language model to calculate the confidence of the sentence to determine whether it needs to be corrected: where wi is a word, n is the length of the sentence, and P(wi|w1, w2, …, w{i-1}) is the context-based word probability. Topic optimization: Based on the topic model (such as LDA or BERT), extract the core theme of the meeting and optimize the meeting minutes: where z represents the topic, w represents the word, and P(w|z) represents the word distribution under the topic.