A meeting recording generation system
By analyzing conference audio and video streams in real time and generating multimodal feedback, the problem of insufficient interaction methods in existing systems in complex conference scenarios is solved, intelligent conference recording and human-computer collaboration are realized, and recording efficiency and user experience are improved.
Patent Information
- Application Number
- CN202510615432.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-05-14
AI Technical Summary
Existing voice interaction systems find it difficult to intelligently adjust the interaction mode according to the discussion atmosphere and progress in complex meeting scenarios, resulting in limited recording quality and collaborative experience, and are unable to effectively convey the recording status, affecting the smoothness of human-computer collaboration.
The information acquisition and analysis module analyzes the conference audio and video streams in real time, calculates the atmosphere tension and semantic analysis results, generates multimodal feedback and dynamically adjusts the interaction strategy, including visual, sound and tactile feedback, optimizes display parameters and feedback frequency, and combines user historical data and personalized control instructions.
It realizes intelligent meeting recording and adaptively adjusts interaction strategies, improves recording efficiency and user experience, and ensures feedback quality and interaction efficiency in complex environments.
Smart Images

Figure CN120260571B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of human-computer interaction, and particularly relates to a conference record generation system. BACKGROUND
[0002] The conference record system is an important field of human-computer interaction, and is crucial to improving work efficiency and information management. With the popularization of intelligent technology, voice-controlled conference record systems have gradually become indispensable tools in enterprises, academia and government agencies, and the core is to achieve efficient and accurate recording and collaboration through natural interaction.
[0003] However, the limitations of existing methods are significant, and most systems rely on complex operation interfaces or fixed voice commands, lack of adaptability to conference dynamics, and cannot intelligently adjust the interaction mode according to the discussion atmosphere or process. This makes it difficult for users to focus on the content itself in complex conference scenarios, and the recording quality and collaboration experience are limited. The core challenge is first reflected in how to achieve flexible control over the recording process through simple voice commands. Existing voice interaction systems are often limited to preset instructions and are difficult to cope with diverse conference scenarios and user habits, which makes the system may interfere with the conference process at critical discussion moments due to command misjudgment. Due to the lack of intelligent recognition of conference atmosphere and process, the system cannot dynamically adjust the interaction mode according to the tension of the discussion or the rest period, such as reducing the prompt frequency during intense debate, or providing a concise summary review during rest. The lack of such recognition capability further makes it difficult for the system to effectively convey the recording status and quality through visual cues, sound signals or tactile feedback, and users have difficulty in real-time understanding of the system's working condition, thereby affecting the smoothness of human-computer collaboration. SUMMARY
[0004] The purpose of the present application is to provide a conference record generation system to solve the problems existing in the prior art.
[0005] To achieve the above-mentioned purpose, the present application provides the following technical scheme: a conference record generation system, the system comprising:
[0006] An information acquisition and analysis module for acquiring conference audio and video streams, extracting voice commands through natural language processing, calculating the tension of the conference atmosphere using real-time signal analysis, obtaining semantic analysis results and atmosphere intensity indicators;
[0007] An evaluation and feedback generation module for calculating a quality evaluation score based on the obtained semantic analysis results and atmosphere intensity indicators, combining keyword hit rate and noise interference degree, using a multi-modal feedback generation algorithm to generate visual cues, sound signals and tactile feedback signals, and obtaining a multi-modal feedback sequence;
[0008] a feedback optimization module, configured to, if the visual prompt clarity of the multi-modal feedback sequence is lower than a preset threshold, optimize display parameters through real-time signal analysis, combine timing alignment errors and grammatical normativity, and obtain an enhanced feedback signal;
[0009] a process dynamic adjustment module, configured to, for the enhanced feedback signal, obtain a conference process timestamp, combine a historical correction trajectory, update feedback frequency and content through a process dynamic adjustment algorithm, and obtain a final interactive output.
[0010] Preferably, the evaluation and feedback generation module calculates the quality evaluation score according to the obtained semantic analysis result and atmosphere intensity index includes matching the voice command using a preset command mapping table according to the semantic analysis result, combining the interactive time sequence and the command triggering frequency, and determining the recording control instruction, including starting, pausing or summarizing the recording.
[0011] Preferably, the evaluation and feedback generation module calculates the quality evaluation score according to the obtained semantic analysis result and atmosphere intensity index also includes, if the atmosphere intensity index exceeds a preset tension threshold, reducing the voice prompt frequency according to the user fatigue index, generating a simplified interactive mode through real-time signal analysis, and obtaining an adjusted interactive strategy.
[0012] Preferably, the evaluation and feedback generation module calculates the quality evaluation score according to the obtained semantic analysis result and atmosphere intensity index also includes, for the adjusted interactive strategy, extracting scene preference labels and command conflict records from user historical interactive data, updating a personalized weight matrix through a user habit adaptation algorithm, and obtaining a personalized control instruction set.
[0013] Preferably, the evaluation and feedback generation module calculates the quality evaluation score according to the obtained semantic analysis result and atmosphere intensity index also includes extracting text content from the personalized control instruction set and the conference audio stream, using a recording quality evaluation algorithm, combining text coverage, semantic consistency and sentence breaking accuracy, and calculating the quality evaluation score.
[0014] Preferably, the information acquisition and analysis module further includes a speech recognition unit and a context analysis unit, wherein the speech recognition unit is configured to perform real-time speech recognition on the conference audio stream, and distinguish different speakers through speaker separation technology in the presence of multiple speakers; the context analysis unit is configured to combine a natural language processing model to perform semantic understanding on the extracted voice command, and optimize the recognition accuracy of the voice command based on the conference context information.
[0015] Preferably, the evaluation and feedback generation module further comprises a quality scoring unit and a feedback algorithm adjustment unit, wherein the quality scoring unit is configured to compare the quality evaluation score calculated based on the semantic analysis result and the atmosphere intensity indicator with a preset conference type standard, and consider the conference content complexity, the number of participants, and the real-time noise level factors; and the feedback algorithm adjustment unit is configured to dynamically adjust the parameters of the multi-modal feedback generation algorithm according to the comparison result, including adjusting the style of the visual prompt, the volume and frequency of the sound signal, and the intensity of the tactile feedback, to adapt to the needs of different conference environments.
[0016] Preferably, the feedback optimization module further comprises a visual optimization unit, a display parameter adjustment unit, and an error correction unit, wherein the visual optimization unit is configured to automatically adjust the brightness, contrast, and font size of the visual prompt according to the ambient light intensity of the conference room, the screen display size, and the user's visual distance; the display parameter adjustment unit is configured to optimize the display parameters based on real-time signal analysis when it is detected that the visual prompt clarity is lower than a preset threshold; and the error correction unit is configured to dynamically correct the visual prompt content in combination with the timing alignment error and the grammatical normativity, to generate an enhanced feedback signal, and to ensure that the user can obtain clear and accurate visual prompt information in various environments.
[0017] Preferably, the process dynamic adjustment module further comprises a user behavior analysis unit, a historical trajectory analysis unit, and a dynamic adjustment unit, wherein the user behavior analysis unit is configured to collect and analyze the operation behavior, preference feedback, and interaction habits of the user during the conference; the historical trajectory analysis unit is configured to integrate the correction trajectory and feedback frequency change trend of past conferences, and identify the user's preference mode for feedback frequency and content; and the dynamic adjustment unit is configured to update the feedback frequency and content in real time according to the analysis result and in combination with the current conference process timestamp, using a process dynamic adjustment algorithm, to provide the final interaction output that meets the user's expectations, and to improve the adaptive ability and user experience of the system.
[0018] Preferably, the speech recognition unit further comprises a noise suppression subunit configured to perform noise reduction processing on the input audio signal using an adaptive filtering and deep learning algorithm in the presence of background noise or multiple sound sources in the conference environment, to improve the accuracy of speech recognition.
[0019] From the above technical solutions, the present application has the following beneficial effects:
[0020] The conference record generation system acquires a conference audio and video stream, analyzes a voice command and a conference atmosphere in real time, and dynamically adjusts an interaction strategy. The conference record generation system matches the voice command according to a semantic analysis result, determines a record control instruction in combination with an interaction time sequence and frequency, adjusts an interaction mode according to a user fatigue index when an atmosphere tension exceeds a threshold value, updates a personalized control instruction set based on user historical data, extracts text from the audio stream and evaluates a record quality, generates a multi-modal feedback, optimizes the feedback when the feedback signal is not clear enough, and dynamically adjusts a feedback frequency and content in combination with a conference process. The conference record generation system realizes intelligent conference recording and human-computer interaction, can adaptively adjust according to a conference condition and user characteristics, improves record efficiency and user experience, and provides a new technical solution for an intelligent conference system. BRIEF DESCRIPTION OF DRAWINGS
[0021] Figure 1 The figure is a system module connection diagram of the present application. DETAILED DESCRIPTION
[0022] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0023] As shown in the figure, Figure 1 The present application provides a technical solution: a conference record generation system, the system comprising:
[0024] An information acquisition and analysis module is configured to acquire a conference audio stream and a video stream, extract a voice command through natural language processing, calculate an atmosphere tension of a conference by real-time signal analysis, obtain a semantic analysis result and an atmosphere intensity index.
[0025] An evaluation and feedback generation module is configured to calculate a quality evaluation score according to the obtained semantic analysis result and atmosphere intensity index, generate a visual prompt, a sound signal and a tactile feedback signal by using a multi-modal feedback generation algorithm in combination with a keyword hit rate and a noise interference degree according to the quality evaluation score, and obtain a multi-modal feedback sequence.
[0026] A feedback optimization module is configured to optimize display parameters by real-time signal analysis if the visual prompt of the multi-modal feedback sequence is below a preset threshold, obtain an enhanced feedback signal in combination with a time sequence alignment error and grammatical normativity.
[0027] A process dynamic adjustment module is configured to obtain a conference process timestamp for the enhanced feedback signal, update a feedback frequency and content by a process dynamic adjustment algorithm in combination with a historical correction trajectory, and obtain a final interaction output.
[0028] The conference recording generation system first receives conference audio and video streams through an information acquisition and analysis module, and uses natural language processing (NLP) techniques to extract speech commands in real time, while applying signal processing algorithms to calculate tension during the conference, forming semantic analysis results and atmosphere intensity indicators. These data are then input into the evaluation and feedback generation module to calculate the quality evaluation score of the conference content, and based on the keyword hit rate and noise interference degree, a feedback sequence including visual cues, sound signals and tactile feedback signals is created through a multi-modal feedback generation algorithm. When the visual cue clarity in the feedback sequence is detected to be lower than the preset threshold, the feedback optimization module adjusts the display parameters based on real-time signal analysis, while considering timing alignment errors and grammatical norms to generate enhanced feedback signals. Finally, the process dynamic adjustment module optimizes the feedback frequency and content based on the enhanced feedback signals and conference process timestamps, combined with historical correction trajectories, using a dynamic adjustment algorithm to ensure the accuracy and adaptability of the final interactive output.
[0029] The system workflow includes the following steps and key formulas:
[0030] 1. Speech command extraction and semantic analysis
[0031] A deep learning acoustic model (such as a speech recognizer based on Transformer or LSTM) is used to extract speech commands from the audio stream.
[0032] 2. Tension index calculation
[0033] The atmosphere tension index (T s ) is calculated by the following formula:
[0034] T s = α1·Δf + β1·Δa;
[0035] Where Δf is the standard deviation of the speech frequency (Hz), Δa is the standard deviation of the speech volume (dB), and α1, β1 are empirical weight coefficients, fitted according to historical data, with example values of α1 = 0.65 and β1 = 0.35.
[0036] Δf and Δa are obtained using sliding window Fourier transform (window 0.5 seconds) and energy envelope analysis.
[0037] 3. Quality evaluation score calculation
[0038] The quality evaluation score (Q s ) is defined as:
[0039] Q s = γ1·M r + γ2·(1-N r );
[0040] where M r is semantic relevance score (range 0-1, computed by NLU module), N r is noise interference degree (0-1, inverse calculated by SNR), γ1, γ2 are evaluation weights, example γ1 = 0.7, γ2 = 0.3.
[0041] M r is scored by RoBERTa or BERT model, N r is estimated using the signal energy ratio before and after voice enhancement.
[0042] 4、Multi-modal feedback generation
[0043] Visual clarity (V s ) evaluation formula:
[0044] V s = δ1·C l + δ2·L c ;
[0045] where C l is visual contrast index (0-1), L c is brightness contrast index (0-1), δ1, δ2 are visual feedback weights (each 0.5).
[0046] Feedback generation: use conditional generative adversarial network (cGAN) or Transformer to generate visual, sound and tactile feedback sequences.
[0047] 5、Feedback optimization
[0048] When V s < V th (threshold 0.75), adjust the display parameters.
[0049] Timing error (E d ) is defined as:
[0050] E d = |t c -t m |;
[0051] where t c is the current predicted timestamp (seconds), t m is the actual matching timestamp (seconds).
[0052] According to E d and grammaticalness (language model score), adjust display resolution, color saturation and other parameters.
[0053] 6、Process dynamic adjustment
[0054] New feedback frequency (F n ) calculation:
[0055]
[0056] Where F0 is the base feedback frequency (Hz), η1 is the dynamic adjustment coefficient (suggested 0.2-0.5), is the historical average timing error (seconds).
[0057] Parameter determination: Cumulative calculation through historical meeting feedback data.
[0058] The system effectively improves the accuracy of conference semantic analysis and atmosphere perception through multi-source data fusion and dynamic feedback mechanism. The multi-modal feedback design (visual, auditory, and tactile) makes information transmission more intuitive. The feedback optimization and process dynamic adjustment functions ensure that the system can continuously maintain feedback quality and interaction efficiency in complex and changing conference environments. Compared with traditional single feedback systems, it has significant advantages in real-time performance, accuracy, and user adaptability.
[0059] In a quarterly performance conference of a certain enterprise, the conference recording generation system of the present application is used. During the conference, the information acquisition and analysis module captures the speech audio and video of the CEO and financial officer in real time. The system detects that the audience tension increases (T value reaches 0.85) during the Q&A session, and timely adjusts the feedback frequency and increases the brightness of the visual prompt to help the recorder accurately identify key information points. At the same time, the feedback generation module generates clear tactile signals to remind the technical support team to prepare specific data according to the semantic matching and noise suppression. During the entire conference, the system dynamically adjusts the feedback content, and the final interactive output meets the high-precision recording and real-time response requirements, greatly reducing the burden of manual operation and improving the efficiency and information availability of the conference.
[0060] In the present system, the visual prompt clarity in the multi-modal feedback sequence is an important indicator of the quality of visual feedback. The setting of the preset threshold directly affects whether the feedback optimization module triggers optimization operations. The determination of this threshold follows the following principles and methods:
[0061] 1. Experimental data statistical method
[0062] First, during the development stage, a large number of user experience tests are conducted to collect user feedback on the acceptability of visual prompts at different clarity levels. The system generates visual prompts under different resolutions, contrasts, brightness, and color settings. Through objective means such as user surveys or eye tracking, record the recognition success rate and subjective satisfaction of users under various settings. Statistical analysis of the clarity level where the recognition success rate reaches more than 90% and the user satisfaction is higher than the preset standard (for example, 80 points / 100 points) is used as a candidate threshold.
[0063] 2. Standard vs. Standard Method
[0064] Reference to human factors standards or display industry standards (such as ISO 9241-303 on visual display requirements) to determine the minimum acceptable visual clarity indicators for most users. For example: contrast ratio should not be less than 3:1. Resolution should not be less than 60 pixels per degree of viewing angle. Combine these industry standards with experimental data to filter out reasonable preset thresholds.
[0065] 3. Adaptive Adjustment Mechanism
[0066] Although the preset threshold is set at the factory or system deployment, the system design allows dynamic adjustment according to the actual use environment.
[0067] The initial threshold (e.g. visual clarity indicator set to 0.75) can be gradually optimized through user feedback data. The system analyzes the interaction effect of the user and the visual feedback in real time through a machine learning model (such as linear regression or decision tree), and adjusts the threshold to adapt to different users and use scenarios.
[0068] 4. User Personalization Settings
[0069] To improve user experience, the system allows advanced users to customize the threshold. Especially for users with color weakness, poor eyesight or specific preferences, you can set a higher or lower clarity threshold according to individual needs.
[0070] The evaluation and feedback generation module calculates the quality evaluation score based on the semantic analysis result and the atmosphere intensity indicator, including matching the voice command with the preset command mapping table based on the semantic analysis result, combining the interaction time sequence and the command trigger frequency to determine the recording control instruction, including starting, pausing or summarizing the recording.
[0071] In this embodiment, the evaluation and feedback generation module first receives the processing result from the semantic analysis module, which contains the voice commands recognized during the meeting and the tension indicator of the meeting atmosphere.
[0072] Step 1: Voice Command Matching
[0073] The system has a pre-defined command mapping table, which contains all commonly used recording control commands, such as "Start Recording", "Pause Recording" and "Summarize Recording", etc. When the semantic analysis module recognizes a new voice command, the system will compare the text of the command with the pre-defined instructions in the mapping table. If the recognized voice command has a corresponding item in the mapping table, it is confirmed that the command is a valid control instruction. For example, if the recognized voice text is "Start Recording", and the mapping table contains this expression or its synonyms, the system considers the command as "Start Recording".
[0074] Parameter Determination: The mapping table is set by the system administrator or user in advance according to actual needs, and can support multiple languages and expressions, and can be updated at any time.
[0075] Second Step: Interactive Time Series Analysis
[0076] The system records the timestamps of all recognized voice commands, forming a time series. These time points reflect when the user issued which commands during the meeting. Within a set time window, such as five minutes, the system counts the number of times each command was triggered during that period. This process helps the system understand the user's operation frequency at different times.
[0077] Parameter Determination: The length of the time window can be set by the user, usually according to the type and duration of the meeting. Short meetings may be set to one minute, and long meetings may be set to five to ten minutes.
[0078] Third Step: Calculate Command Trigger Frequency
[0079] Within the time window, the system calculates the trigger frequency of each control command based on the number of commands counted. For example, if the "pause recording" command was triggered three times in the past five minutes, the system will record the trigger frequency of this command as three times every five minutes.
[0080] Parameter Determination: The trigger frequency is automatically calculated by the system based on statistical data, without human intervention.
[0081] Fourth Step: Combine Atmosphere and Quality Assessment
[0082] The system not only considers the number and frequency of command triggers, but also refers to the previously calculated meeting atmosphere tension and meeting content quality assessment scores through natural language processing and signal analysis. Atmosphere tension reflects the current meeting tension or activity level, usually automatically evaluated according to voice volume, speech speed and audio signal changes. The quality assessment score measures the accuracy of semantic analysis and the degree of influence of background noise.
[0083] Parameter Determination: Atmosphere tension is obtained by statistical analysis of real-time audio signals, such as voice volume fluctuations and speech speed changes. Quality assessment scores include semantic matching degree, background noise level and keyword recognition effect, which are calculated by the system's natural language processing module and signal processing module based on pre-trained models and real-time data.
[0084] Fifth Step: Determine Recording Control Instructions
[0085] Finally, the system takes the matched commands, command trigger frequency, meeting atmosphere tension, and quality assessment score as inputs, and automatically determines the recording control instruction to be executed through a decision logic. For example, if the system detects multiple triggers of the "pause recording" command while the meeting atmosphere is tense and the semantic recognition is accurate, the system will prioritize the "pause recording" operation. If the system detects that the user has issued a "summarize recording" command and the current meeting stage is nearing its end, the system will execute "summarize recording" and begin organizing the meeting content.
[0086] Parameter determination: The decision logic can be based on fixed rules, such as prioritizing high-frequency commands, or using machine learning algorithms to dynamically adjust based on historical data and user preferences.
[0087] Through the above steps, the system can intelligently understand and respond to voice commands during the meeting process, achieving automated recording control. This design reduces manual operations, improves the efficiency and accuracy of meeting records, and is particularly suitable for complex or dynamic meeting scenarios.
[0088] This implementation enables the system to automatically recognize and execute recording control commands based on real-time semantic and interactive features of the meeting, reducing manual intervention and improving the intelligence and flexibility of operations. By combining the interactive time series and command trigger frequency, the system can accurately identify user intent, even in cases of diverse command expressions or complex contexts, maintaining a high recognition accuracy and ensuring the continuity and integrity of meeting records. In addition, by combining atmosphere tension and quality assessment scores, the system can dynamically adjust the trigger strategy of control instructions, ensuring optimal response in different meeting situations.
[0089] The evaluation and feedback generation module calculates the quality assessment score based on the semantic analysis results and atmosphere intensity indicators, and further includes the following steps: if the atmosphere intensity indicator exceeds the preset tension threshold, the voice prompt frequency is reduced according to the user fatigue index, a simplified interaction mode is generated through real-time signal analysis, and an adjusted interaction strategy is obtained.
[0090] In this implementation, the evaluation and feedback generation module not only completes semantic analysis, command matching, and interaction frequency analysis, but also intelligently adjusts the system's interaction strategy based on the emotional state and fatigue of meeting participants. The specific process is as follows:
[0091] First step: Atmosphere intensity analysis
[0092] The system analyzes the speech rate, volume fluctuation, tone change, and complexity of background noise in the meeting audio to calculate the tension of the current meeting atmosphere. If the detected tension exceeds the pre-set standard value, such as the "high stress threshold" determined during the experiment, the system determines that the current meeting is in a tense state.
[0093] Parameter determination: The tension threshold is determined during system design and testing phases by analyzing a large amount of real meeting data, based on user subjective evaluation and physiological indicators (such as heart rate, skin conductance). Generally, the tension level when users feel obvious stress is chosen as the threshold.
[0094] Second step, user fatigue index acquisition
[0095] The system estimates the user's fatigue index based on the following data: meeting duration: the longer the duration, the higher the fatigue index. User interaction frequency changes: if the user's response speed to system feedback gradually slows down, it indicates that fatigue increases. Optional physiological signals (if hardware supports): such as camera tracking blink frequency or changes in posture.
[0096] Parameter determination: The baseline value of the fatigue index is established based on the data of the user's first use, and is automatically updated according to the historical data of each meeting afterwards.
[0097] Third step, reduce the frequency of voice prompts
[0098] When the atmosphere tension exceeds the threshold and the fatigue index is high, the system automatically reduces the number and frequency of voice prompts to avoid interrupting the normal rhythm of the meeting and reducing the user's cognitive burden. For example, reduce the original number of voice prompts per minute by half, or only issue prompts when key events occur.
[0099] Parameter determination: The adjustment amplitude of the voice prompt frequency is set according to historical interaction data and user preferences. The initial setting is usually based on research or standard human engineering recommended values, and can be continuously optimized through user feedback.
[0100] Fourth step, generate simplified interaction mode
[0101] The system further analyzes the meeting content and user response habits, and automatically simplifies complex feedback or interaction prompts. For example, reduce long text prompts, only keep key words or use icons instead of language instructions, or combine multi-step operations into a single step to simplify the user's operation burden.
[0102] Parameter determination: Simplification rules are set by user preferences or learned by machine learning models based on user past behavior habits. Optional models include decision trees or sequence-to-sequence models.
[0103] Fifth step, form an adjusted interaction strategy
[0104] Finally, the system integrates tension, fatigue index, and simplified interaction mode to develop an interaction strategy that adapts to the current meeting situation. This strategy guides the system's subsequent feedback frequency, prompt content, and interaction mode, ensuring that even in a tense or fatigued environment, good user experience and meeting record effect can be maintained.
[0105] This implementation enables the conference recording system to have intelligent adaptability by introducing emotion perception and user fatigue assessment mechanisms. When detecting rising tension and intensifying user fatigue, the system can actively reduce disruptive prompts and simplify interaction content, thereby reducing user's operational burden, reducing cognitive pressure, and improving the smoothness and comfort of the conference process. In addition, the generation of dynamic interaction strategies ensures that the system can still provide effective support in complex or high-pressure scenarios, significantly improving user satisfaction and the practicality of the system.
[0106] The evaluation and feedback generation module calculates the quality evaluation score based on the semantic analysis results and atmosphere intensity indicators, and also includes extracting scene preference tags and command conflict records from user historical interaction data for the adjusted interaction strategy, updating the personalized weight matrix through the user habit adaptation algorithm, and obtaining the personalized control instruction set.
[0107] In this embodiment, the evaluation and feedback generation module further integrates the ability of user habit learning and personalized adaptation, enabling the system to adjust feedback and control strategies based on user historical behavior. The specific process is as follows:
[0108] Step 1: Collect user historical interaction data
[0109] The system will continuously record interaction data in each conference during the user's long-term use. These data include the user's feedback preferences in different conference types, such as the frequency of visual prompts, the use of voice prompts, and the execution of specific control commands (such as "pause recording" or "summarize recording") in certain situations. At the same time, the system also records the command conflict situations that have occurred, such as the user quickly and continuously issuing contradictory instructions (for example, saying "pause" first, and then saying "continue"), or responding differently to the same instruction in different scenarios.
[0110] Parameter determination: User interaction data is automatically collected by the system without manual input from the user. The data includes command type, trigger time, conference scene description, and user response records to feedback.
[0111] Step 2: Extract scene preference tags
[0112] The system analyzes the user's preferences in specific scenarios based on the collected data. For example, in a "project discussion" conference, the user may prefer higher frequency visual feedback, while in a "financial audit" conference, the user may prefer low frequency voice prompts. These preferences are summarized as "scene preference tags", each tag describing the user's interaction preferences in a specific conference situation.
[0113] Parameter determination: Tags are generated through statistical analysis, taking into account the meeting type, meeting duration, number of participants, and user feedback to adjust behavior.
[0114] Step 3: Identify command conflict records
[0115] The system automatically identifies and records historical command conflicts, including instances where users quickly change commands or undo previous commands. It also records the frequency and context of conflicts, as well as the command that the user ultimately confirmed.
[0116] Parameter determination: Conflict records are derived from system logs, combined with timestamps and final user behavior.
[0117] Step 4: Update the personalized weight matrix
[0118] Based on scenario preference tags and command conflict records, the system dynamically adjusts control logic using a user habit adaptation algorithm. This algorithm prioritizes different types of feedback and control commands based on the user's historical behavior, forming a personalized weight matrix. For example, if a user prefers visual feedback in most scenarios and frequently dismisses voice prompts, the system will lower the priority of voice prompts.
[0119] Parameter determination: The initial weight matrix is set by user survey data during the development phase and subsequently optimized through an adaptive algorithm. This algorithm can use a weighted moving average or a simple machine learning model such as K-nearest neighbor or logistic regression.
[0120] Step 5: Generate a personalized control instruction set
[0121] Guided by the updated weight matrix, the system prioritizes feedback types and control commands that match user preferences in real-time operation. Ultimately, a personalized set of control commands is formed for the current meeting context, ensuring that the system's responses meet the user's usage habits and interaction expectations.
[0122] This implementation introduces a user habit learning mechanism, enabling the meeting recording system to continuously adapt to individual user needs. Instead of a one-size-fits-all control strategy, the system adjusts based on historical user behavior and feedback, providing a personalized interactive experience. This adaptive mechanism significantly improves user satisfaction, reduces user discomfort with system feedback, and reduces the likelihood of misoperation. It also provides the system with greater flexibility and intelligence for diverse users and meeting scenarios.
[0123] The evaluation and feedback generation module calculates the quality evaluation score based on the semantic analysis results and atmosphere intensity indicators. It also extracts text content from the personalized control instruction set and conference audio stream, uses the recording quality evaluation algorithm, and combines text coverage, semantic consistency and sentence segmentation accuracy to calculate the quality evaluation score.
[0124] In this embodiment, the evaluation and feedback generation module further enhances the quality evaluation capability of the meeting record content itself, ensuring that the system-generated record content meets user expectations and has high accuracy. The specific workflow is as follows:
[0125] First, extract the text content from the audio stream
[0126] The system uses advanced automatic speech recognition technology to transcribe the meeting audio stream into text in real time. At the same time, according to the control rules in the personalized control instruction set, it filters out key speech segments and related text. For example, for control instructions such as "summary record" or "highlight marking", the audio content of these parts is preferentially processed and transcribed into text.
[0127] Parameter determination: The speech recognition model uses a pre-trained large-scale language model, and the recognition accuracy is calibrated regularly according to the user's usage scenario. Key words and control instructions are obtained by user preference settings or system learning.
[0128] Second, calculate the text coverage rate
[0129] The system analyzes the matching degree of the meeting content covered in the transcribed text with the actual meeting agenda or preset points. The coverage rate reflects whether the system has recorded the core content of the meeting completely. The coverage rate is calculated based on the proportion of the actual recorded meeting points to the total points discussed in the meeting.
[0130] Parameter determination: Meeting points are extracted by user preset, meeting agenda file, or system natural language understanding module.
[0131] Third, evaluate semantic consistency
[0132] The system further compares the semantic consistency of the transcribed text with the original content of the meeting discussion, detects whether there is distortion of meaning or deviation of understanding. For example, if the speech content is misrecognized or oversimplified, the system will reduce the semantic consistency score.
[0133] Parameter determination: Semantic consistency is evaluated by semantic similarity algorithm, and the semantic embedding model used can be based on Sentence-BERT or similar large pre-trained semantic models, optimized for different types of meetings.
[0134] Fourth, evaluate the accuracy of sentence segmentation
[0135] The system checks the quality of sentence segmentation in the transcribed text to determine whether the boundaries of the sentences are correct, especially focusing on the sentence segmentation effect in complex sentences and multi-speaker scenarios. If it finds that the sentence segmentation is wrong, for example, a sentence is incorrectly divided into multiple sentences or multiple sentences are mistakenly combined into one sentence, the sentence segmentation accuracy score will be reduced.
[0136] Parameter determination: The punctuation rule refers to the standard punctuation model in the field of natural language processing, while combining the user's personalized language habits adjustment, supporting Mandarin, English and multi-language environment.
[0137] Step 5, Calculate the record quality evaluation score
[0138] The system integrates text coverage, semantic consistency and punctuation accuracy, according to the weight proportion set in advance, to calculate the final record quality evaluation score. This score directly reflects the overall quality of the meeting record, and can be used as an important basis for subsequent optimization feedback strategy and adjustment of interactive mode.
[0139] Parameter determination: The weight proportion is determined by user research in the system design stage, for example, semantic consistency may be more important in legal meetings, while coverage may be more important in technical discussions. The system supports users to adjust the weight according to specific needs.
[0140] This embodiment realizes the quantitative control of the quality of the meeting record by including the comprehensiveness of the text content, the semantic accuracy and the rationality of the punctuation into the quality evaluation system. This multi-dimensional evaluation method ensures that the meeting records generated by the system meet the user's expectations in terms of accuracy, completeness and readability, and is particularly suitable for scenarios with high requirements for record quality, such as legal, medical and financial fields. In addition, the combination of personalized control instructions and real-time analysis makes the evaluation process have intelligent adaptability, which can be flexibly adjusted according to different users and meeting types.
[0141] The information acquisition and analysis module further includes a speech recognition unit and a context analysis unit, wherein the speech recognition unit is used for real-time speech recognition of the meeting audio stream, and different speakers are distinguished by speaker separation technology in the presence of multiple speakers; the context analysis unit is used to combine natural language processing model to understand the semantics of the extracted voice command, and to optimize the recognition accuracy of the voice command based on the meeting context information.
[0142] In this embodiment, the information acquisition and analysis module is refined into two key sub-modules, namely the speech recognition unit and the context analysis unit, which work together to ensure high-precision, context-related voice command recognition in a multi-speaker meeting environment:
[0143] Step 1, Real-time speech recognition processing
[0144] The speech recognition unit continuously monitors and transcribes the meeting audio stream. This unit uses an end-to-end deep learning speech recognition model to convert real-time speech signals into text information. For clear and single speaker's voice stream, the recognition process directly generates text; while in the context of multiple participants speaking at the same time or in turn, the system automatically enables the speaker separation mechanism.
[0145] Parameter determination: The speaker separation technology is based on voiceprint recognition and speech signal clustering algorithm, such as based on mel frequency cepstral coefficient and speaker vector model, through clustering to divide and attribute the audio to each speaker, and the recognition accuracy can reach more than 90%.
[0146] Second step, speaker distinction and label management
[0147] Each segment of the transcribed text is bound with a speaker label, and the system can uniformly number the speakers based on historical voice features, such as "Speaker 1", "Speaker 2", etc., to facilitate the generation of subsequent structured meeting records. If the system knows the identity of the speaker, the voiceprint recognition result can be further used to bind the label with the specific name.
[0148] Label matching can use the meeting participant roster and voiceprint pre-registration method to improve accuracy, and the system supports dynamic learning and correction.
[0149] Third step, semantic understanding and context optimization
[0150] The context analysis unit performs semantic analysis on the recognized text. This unit uses natural language understanding models, such as BERT or RoBERTa-based context modeling methods, to extract intent information and keywords from the text and identify whether it is an instructional sentence (such as "pause recording", "continue"). To improve accuracy, the system also references the current meeting context information, such as the current topic being discussed, the content of the previous sentence, the historical command sequence, etc., to determine the meaning of the sentence and avoid misidentification, such as recognizing "I want to pause for a moment" as an ordinary sentence rather than a control command.
[0151] Context information is achieved through the establishment of a "dynamic semantic cache". The cache records current topic keywords, recently recognized commands, user speaking habits, etc. The cache supports context window settings, such as the last 10 rounds of speaking or the semantic sequence within 30 seconds.
[0152] This embodiment significantly improves the system's voice processing capability in complex, multi-speaker meeting environments. By combining the voice recognition unit and the speaker separation technology, it ensures that the voice content of each speaker is accurately recognized and correctly classified, avoiding information confusion in a multi-speaker context. At the same time, the context analysis unit combines natural language processing and context modeling to effectively solve the problem of recognizing instructional sentences in natural language, making the system more accurate and robust in recognizing voice commands. This design greatly improves the accuracy, structure, and interactive intelligence of meeting records.
[0153] The evaluation and feedback generation module further comprises a quality scoring unit and a feedback algorithm adjustment unit. The quality scoring unit is configured to compare the quality evaluation score calculated based on the semantic analysis result and the atmosphere intensity indicator with a preset conference type standard, and consider the conference content complexity, the number of participants, and the real-time noise level factors. The feedback algorithm adjustment unit is configured to dynamically adjust the parameters of the multi-modal feedback generation algorithm according to the comparison result, including adjusting the style of visual cues, the volume and frequency of sound signals, and the intensity of tactile feedback, to adapt to the needs of different conference environments.
[0154] The present embodiment introduces two new sub-units in the evaluation and feedback generation module, namely the quality scoring unit and the feedback algorithm adjustment unit. These two units cooperate to achieve dynamic optimization of conference feedback, ensuring that the system can adaptively adjust the feedback strategy according to different conference scenarios and real-time environmental changes. The specific process is as follows:
[0155] First step, generate quality evaluation score
[0156] The system calculates the current quality evaluation score of the conference based on the semantic analysis result and the atmosphere intensity indicator. This score reflects the comprehensive quality of speech recognition, semantic understanding, atmosphere perception, and interaction effect.
[0157] The calculation method of the quality evaluation score considers the semantic accuracy, feedback response speed, and noise interference level as described in the previous embodiment.
[0158] Second step, compare with conference type standard
[0159] The quality scoring unit compares the calculated quality evaluation score with the preset standard value. The preset standard is set in advance according to the conference type, such as strategic meetings, project meetings, or training meetings, each type having different minimum quality requirements.
[0160] The conference type standard value is set through user demand investigation or historical data analysis. For example, high-level decision-making meetings require higher semantic consistency, and customer training meetings require more flexible feedback frequency.
[0161] Third step, consider conference content complexity, number of participants, and noise level
[0162] The quality scoring unit further adjusts the comparison benchmark according to the specific circumstances of the conference. For example, when the conference content complexity is high, the system appropriately relaxes the requirement for feedback response speed; when the number of participants is large, the fault tolerance rate of speaker recognition is increased; and when the real-time noise level rises, the requirement for speech prompt clarity is reduced to avoid false triggering of feedback optimization.
[0163] The complexity of the meeting is determined according to the number of keywords in the meeting agenda and the complexity of the sentences. The number of participants is automatically detected through the meeting management system or the number of voice channels. The noise level is measured by the background noise energy in the audio signal, using the same noise evaluation standard as the speech recognition module.
[0164] Step 4: Dynamic adjustment of feedback algorithm parameters
[0165] The feedback algorithm adjustment unit automatically adjusts the parameters of the multi-modal feedback generation algorithm based on the comparison results of the quality scoring unit.
[0166] If the current quality score is lower than the standard and the content is complex or the noise is high, the system will increase the contrast and font size of the visual cues to improve the recognition clarity. The volume and frequency of the sound signal can also be increased or decreased as needed to prevent information overload or omission. The intensity of the tactile feedback can be moderately increased within the range acceptable to the user to enhance the perception effect in a noisy environment.
[0167] The initial values of the feedback parameters are set by the user or determined based on industry standards. The adjustment range is automatically limited based on user preferences and device capabilities, such as the maximum font size of visual cues not exceeding a percentage of the pre-set screen size.
[0168] Step 5: Adaptive learning
[0169] The system continuously records the adjustment results and user feedback to optimize future feedback parameter settings, improving adaptability and user satisfaction.
[0170] This implementation realizes dynamic monitoring and real-time optimization of feedback effects through the introduction of the quality scoring unit and the feedback algorithm adjustment unit. The system can automatically adjust the feedback strategy according to different types of meetings and real-time environmental changes, significantly improving the effectiveness of feedback and reducing human intervention, while meeting the differentiated needs of different meeting environments. Especially in complex or high-noise environments, it can effectively guarantee the user's interactive experience and improve the accuracy and timeliness of the meeting record.
[0171] The feedback optimization module further includes a visual optimization unit, a display parameter adjustment unit, and an error correction unit. The visual optimization unit automatically adjusts the brightness, contrast, and font size of visual cues based on the ambient light intensity of the meeting room, the screen display size, and the user's viewing distance. The display parameter adjustment unit optimizes the display parameters based on real-time signal analysis when the visual cue clarity is detected to be below a pre-set threshold. The error correction unit dynamically corrects the visual cue content by combining timing alignment errors and grammatical correctness to generate enhanced feedback signals, ensuring that users can obtain clear and accurate visual cue information in various environments.
[0172] In this embodiment, the feedback optimization module is refined into three sub-modules that work together: the visual optimization unit, the display parameter adjustment unit, and the error correction unit. The goal is to continuously ensure the clarity and accuracy of visual cues in different environments. The specific process is as follows:
[0173] First step: Environment perception and preliminary optimization
[0174] The visual optimization unit is responsible for real-time collection of visual-related information of the conference room environment, including environmental light intensity (such as illuminance), physical size of the display screen, and the visual distance between the user and the screen. These information is usually obtained through environmental light sensors on the device, screen parameter interface, and user settings or camera visual ranging functions.
[0175] Based on the above information, the system automatically adjusts the brightness, contrast, and font size of the visual cues. For example, when the environmental light is strong, the brightness and font thickness are increased, and when the display screen size is small, higher contrast and larger font are used to improve information readability.
[0176] The environmental light intensity is measured in lux, with a common range from 100 lux (low light) to 1000 lux (bright conference room).
[0177] The screen size is read through internal device parameters, such as 13 inches, 27 inches, etc.
[0178] The visual distance is usually set by the user or obtained by the device's ranging function, generally ranging from 0.5 meters to 3 meters.
[0179] Second step: Clarity detection and parameter readjustment
[0180] The display parameter adjustment unit monitors whether the clarity of the current visual cues is below the system's preset threshold. Clarity evaluation can consider indicators such as brightness contrast, visual edge clarity, and font edge sharpness.
[0181] If the clarity is found to be insufficient, the system immediately triggers the parameter optimization process, resetting color, font edge sharpness, and screen refresh mode to improve visual effects.
[0182] The preset threshold is set by experimental testing and user feedback. For example, if the visual clarity score is below 0.75 (full score 1.0), optimization is triggered. The specific score is output in real time by the image processing module.
[0183] Third step: Content correction and semantic enhancement
[0184] The error correction unit is mainly responsible for optimizing the semantic level of visual cues, including correcting the alignment error between recognized text and actual speech time, and correcting common grammar errors such as missing punctuation and sentence breaks.
[0185] The system employs a language model to reparse the output text and align the display timing of the prompts with the speech timestamp information, ensuring that the information received by the user is accurate, logically coherent, and reasonably time-aligned.
[0186] The timing alignment error is calculated by comparing the speech recognition output timestamp with the actual speaking time. The grammatical correctness is determined by a natural language model to assess the completeness and structural reasonableness of the sentence.
[0187] This embodiment refines the visual adjustment and error correction capabilities in the feedback optimization module, enabling the system to have good environmental adaptability and semantic correction capabilities. By sensing the conference environment conditions in real time and automatically adjusting the prompt content style, users can obtain clear and comfortable visual feedback experience in bright, dim, close, or distant conditions. At the same time, the error correction mechanism ensures the language accuracy and time synchronization of the prompt content, especially suitable for information-intensive, fast-paced, or high-environmental interference conference scenarios, greatly improving the reliability and professionalism of conference recording interaction.
[0188] The process dynamic adjustment module further includes a user behavior analysis unit, a historical trajectory analysis unit, and a dynamic adjustment unit. The user behavior analysis unit is used to collect and analyze the user's operation behavior, preference feedback, and interaction habits during the conference. The historical trajectory analysis unit is used to integrate the correction trajectory and feedback frequency trend of past conferences, and identify the user's preference pattern for feedback frequency and content. The dynamic adjustment unit is used to update the feedback frequency and content in real time according to the analysis results and the current conference process timestamp, provide the final interaction output that meets the user's expectations, and improve the system's adaptive ability and user experience.
[0189] This embodiment further improves the intelligent level of the process dynamic adjustment module, enabling it to dynamically adjust the feedback strategy based on the user's operation habits and historical preferences, and achieve personalized interaction optimization. The specific process is as follows:
[0190] Step 1: User behavior analysis
[0191] The user behavior analysis unit continuously monitors the user's various operation behaviors during the conference, such as adjusting the frequency of visual prompts, pausing or starting recording, modifying feedback content, etc. At the same time, the system records the user's response to different types of feedback, such as accepting, ignoring, or manually adjusting the feedback, to identify the user's preference feedback and interaction habits.
[0192] Operation behaviors include clicking, voice commands, gesture inputs, etc. Preference feedback is determined based on whether the user actively responds or adjusts the feedback content, and interaction habits are identified by counting the user's repeated behaviors in different conference types.
[0193] Second Step, Historical Trajectory Analysis
[0194] The historical trajectory analysis unit integrates all past meeting data recorded by the system, including user records of adjusting feedback content and frequency, time points when feedback was accepted or ignored, and historical trajectories of automatic corrections by the system. These data are used to identify user preference patterns in specific situations, such as a tendency to reduce prompt frequency in long meetings and increase feedback detail during key topic discussion phases.
[0195] Data is derived from system logs, including timestamps, feedback types, user response records, etc. Feedback frequency change trends are calculated by counting feedback adjustment intervals and times, and correction trajectories are identified based on user modification history of system output.
[0196] Third Step, Dynamic Feedback Strategy Generation
[0197] The dynamic adjustment unit, based on the aforementioned analysis results, combines the current meeting's progress timestamp, applies a process dynamic adjustment algorithm, and updates feedback frequency and content in real time. For example, when the system detects that the meeting has entered the summary phase and user historical behavior shows a preference for brief feedback at this stage, it automatically reduces feedback frequency and simplifies prompt content; if it is in the problem discussion phase and the user prefers detailed feedback, it increases feedback frequency and content richness.
[0198] The initial rules of the dynamic adjustment algorithm are set based on user research and industry best practices, and are continuously optimized through machine learning, such as using decision trees or time series prediction models to adjust feedback strategies.
[0199] Fourth Step, Generation of Final Interaction Output
[0200] The system integrates all factors and outputs feedback content and frequency that meet the current meeting situation and user preferences in real time, ensuring that the interaction process meets information needs without causing user burden, significantly improving overall user experience and the system's adaptive ability.
[0201] This implementation gives the system strong personalized learning and dynamic adjustment capabilities. By deeply analyzing user behavior and historical data, the system can accurately identify feedback preferences of different users and different meeting situations, and adjust interaction strategies in real time. This mechanism significantly reduces user manual adjustment burden in the meeting process, improves the naturalness and comfort of interaction, and enables the system to actively adapt to changes in the meeting process, enhancing user experience and work efficiency.
[0202] The speech recognition unit also includes a noise suppression subunit for reducing noise in input audio signals using adaptive filtering and deep learning algorithms when there is background noise or multiple sound sources in the meeting environment, improving the accuracy of speech recognition.
[0203] In this embodiment, the speech recognition unit further integrates a noise suppression subunit, which is specifically designed to deal with the complex noise problems commonly encountered in conference environments, improving the accuracy and robustness of overall speech recognition. The workflow is as follows:
[0204] Step 1: Environmental audio acquisition
[0205] The system acquires audio signals in the conference environment through a microphone array or a single microphone. The audio signal usually contains the user's speech and background noise, such as air conditioner sound, keyboard tapping sound, walking sound, and even other non-target speech of speakers.
[0206] The audio sampling rate is usually set at 16 kHz or higher to ensure that the entire frequency range of human speech is covered, while ensuring sufficient signal details for subsequent processing.
[0207] Step 2: Adaptive filter noise reduction
[0208] The noise suppression subunit first applies adaptive filtering technology to preliminarily suppress the background noise in the audio signal. The filter dynamically adjusts the filtering parameters according to the real-time detected background noise characteristics, such as automatically selecting appropriate filtering methods such as linear prediction filtering or Kalman filtering according to the frequency components and trends of the noise.
[0209] The filtering parameters are calculated in real time based on the frequency distribution and energy level of the noise, and the adaptation speed of the filter is determined by the noise variation speed and the fidelity requirements of the speech content.
[0210] Step 3: Deep learning algorithm for further noise reduction
[0211] Based on the filtered audio, the system uses a deep learning algorithm to further identify and separate noise and target speech. The model used is usually a convolutional neural network, a recurrent neural network, or a specially designed deep reinforcement learning algorithm, which can accurately extract target speech in a complex, multi-sound source environment.
[0212] This algorithm not only identifies traditional background noise, but also identifies interference sources similar in frequency to the target speech, such as other speakers' speech, effectively reducing the occurrence of "cross-talking".
[0213] The deep learning model is trained with a large amount of conference speech and noise data, and the model complexity (number of layers and number of parameters) is determined by the device computing power and real-time requirements. The system can be periodically retrained or fine-tuned with new conference data to continuously optimize noise reduction performance.
[0214] Step 4: Clean audio input speech recognition module
[0215] The audio signal processed by both filtering and deep learning noise reduction is input to the speech recognition module for real-time transcription. Due to the significant reduction in noise interference, the recognition accuracy is significantly improved.
[0216] By integrating the noise suppression subunit, the system significantly improves the accuracy and stability of speech recognition in noisy or multi-source conference environments. The adaptive filtering can quickly respond to changes in environmental noise, and the deep learning algorithm can handle complex noise and sound overlap problems, ensuring accurate recognition of the speaker's voice commands and speech content even in high-noise or multi-person simultaneous speaking situations. This design greatly improves the environmental adaptability of the system, reduces false recognition and false touch, and improves user experience.
[0217] Although embodiments of the present application have been shown and described, it will be understood by those having ordinary skill in the art that various changes, modifications, substitutions and alterations can be made therein without departing from the principles and spirit of the application, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A meeting record generation system, characterized in that: The system comprises: The information acquisition and analysis module is used to obtain the conference audio and video streams, extract voice commands through natural language processing, and use real-time signal analysis to calculate the tension of the conference atmosphere to obtain semantic analysis results and atmosphere intensity indicators; The evaluation and feedback generation module is used to calculate the quality evaluation score based on the semantic analysis results and atmosphere intensity index. Based on the quality evaluation score, combined with the keyword hit rate and noise interference, a multimodal feedback generation algorithm is used to generate visual prompts, sound signals and tactile feedback signals to obtain a multimodal feedback sequence; The evaluation and feedback generation module calculates the quality evaluation score based on the semantic analysis results and the atmosphere intensity index, and further includes reducing the frequency of voice prompts according to the user fatigue index if the atmosphere intensity index exceeds a preset tension threshold, and generating a simplified interaction method through real-time signal analysis to obtain an adjusted interaction strategy; The evaluation and feedback generation module calculates the quality evaluation score based on the semantic analysis results and atmosphere intensity index, and further extracts scene preference labels and command conflict records from the user's historical interaction data for the adjusted interaction strategy, updates the personalized weight matrix through the user habit adaptation algorithm, and obtains a personalized control instruction set; The evaluation and feedback generation module calculates the quality evaluation score based on the semantic analysis results and the atmosphere intensity index, further comprising extracting text content from the personalized control instruction set and the conference audio stream, and using a recording quality evaluation algorithm to calculate the quality evaluation score by combining text coverage, semantic consistency, and sentence segmentation accuracy; A feedback optimization module is used to optimize display parameters through real-time signal analysis if the visual cue clarity of the multimodal feedback sequence is lower than a preset threshold, combining timing alignment errors and grammatical standardization to obtain an enhanced feedback signal; The process dynamic adjustment module is used to obtain the conference process timestamp based on the enhanced feedback signal, combine it with the historical correction trajectory, and update the feedback frequency and content through the process dynamic adjustment algorithm to obtain the final interactive output.
2. A meeting record generation system according to claim 1, characterized in that: The evaluation and feedback generation module calculates the quality evaluation score based on the semantic analysis results and atmosphere intensity index, including matching voice commands using a preset command mapping table based on the semantic analysis results, and determining the recording control instructions, including starting, pausing or summarizing the recording, in combination with the interaction time sequence and command trigger frequency.
3. A meeting record generation system according to claim 1, characterized in that: The information acquisition and analysis module also includes a speech recognition unit and a context analysis unit, wherein the speech recognition unit is used to perform real-time speech recognition on the conference audio stream and distinguish different speakers through speaker separation technology when there are multiple parties speaking; the context analysis unit is used to combine the natural language processing model to perform semantic understanding on the extracted voice commands and optimize the recognition accuracy of the voice commands based on the conference context information.
4. A meeting record generation system according to claim 1, characterized in that: The evaluation and feedback generation module also includes a quality scoring unit and a feedback algorithm adjustment unit, wherein the quality scoring unit is used to compare the quality assessment score calculated based on the semantic parsing results and the atmosphere intensity index with the preset meeting type standard, and take into account the complexity of the meeting content, the number of participants and the real-time noise level factors; the feedback algorithm adjustment unit is used to dynamically adjust the parameters of the multimodal feedback generation algorithm according to the comparison results, including adjusting the style of visual prompts, the volume and frequency of the sound signal and the intensity of tactile feedback to adapt to the needs of different meeting environments.
5. A meeting record generation system according to claim 1, characterized in that: The feedback optimization module also includes a visual optimization unit, a display parameter adjustment unit and an error correction unit, wherein the visual optimization unit is used to automatically adjust the brightness, contrast and font size of the visual prompt according to the ambient light intensity of the conference room, the screen display size and the user's viewing distance; the display parameter adjustment unit is used to optimize the display parameters based on real-time signal analysis when it detects that the clarity of the visual prompt is lower than a preset threshold; the error correction unit is used to dynamically correct the visual prompt content in combination with timing alignment errors and grammatical norms, generate an enhanced feedback signal, and ensure that users can obtain clear and accurate visual prompt information in various environments.
6. A meeting record generation system according to claim 1, characterized in that: The process dynamic adjustment module also includes a user behavior analysis unit, a historical trajectory analysis unit and a dynamic adjustment unit, wherein the user behavior analysis unit is used to collect and analyze the user's operating behavior, preference feedback and interaction habits during the meeting; the historical trajectory analysis unit is used to integrate the correction trajectory and feedback frequency change trend of past meetings, and identify the user's preference pattern for feedback frequency and content; the dynamic adjustment unit is used to update the feedback frequency and content in real time using a process dynamic adjustment algorithm based on the analysis results and the current meeting process timestamp, to provide a final interactive output that meets the user's expectations, thereby improving the system's adaptability and user experience.
7. A meeting record generation system according to claim 3, characterized in that: The speech recognition unit also includes a noise suppression subunit, which is used to use adaptive filtering and deep learning algorithms to reduce the noise of the input audio signal when there is background noise or multiple sound sources in the conference environment, thereby improving the accuracy of speech recognition.
Citation Information
Patent Citations
Conference process optimization method and device, electronic equipment and storage medium
CN117875904A
Multimedia conference room system
CN119376267A