Conference enhancement method, device, equipment and storage medium
Through a hierarchical module design conference enhancement system, multimodal data fusion, participant status analysis and decision-making process tracking are achieved, which solves the problem of insufficient efficiency and results of the existing conference system, and provides real-time intervention suggestions to improve the quality of the conference.
Patent Information
- Application Number
- CN202510536515.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2045-04-27
AI Technical Summary
The existing conference system lacks the comprehensive ability to process multimodal conference information, analyze participants' status and track decision-making process, resulting in inefficient meetings and limited results.
A conference enhancement system designed with hierarchical modules, including data preprocessing, participant analysis, conference content processing, decision analysis and intervention output components. Real-time intervention suggestions are generated through multimodal data fusion, participant identity identification and status analysis, structured representation of conference content and decision-making process analysis.
It improves the efficiency of conferences and the effectiveness of results, can identify implicit communication barriers and provide precise intervention, and improves the quality of decision-making.
Smart Images

Figure CN120075203B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a conference enhancement method, apparatus, device and storage medium. Background Art
[0002] As organizational complexity and decision-making environments become increasingly uncertain, meetings, as a core venue for team collaboration and decision-making, have a significant impact on business operations, both in terms of efficiency and quality. Currently, mid- and senior-level managers spend an average of a significant portion of their working time in meetings, many of which are perceived by participants as inefficient or of limited effectiveness. This inefficiency not only wastes time and resources but can also delay or reduce the quality of critical decisions.
[0003] Existing meeting enhancement solutions address this issue through voice transcription and recording systems, and basic collaboration platforms that provide document and screen sharing capabilities. However, these solutions often only support processing a single type of data. Their lack of comprehensive analysis of participants' emotions and other aspects of their state hinders their ability to identify hidden communication barriers. Furthermore, their use of fixed strategies for intervention prompts often results in untimely or inaccurate interventions. Summary of the Invention
[0004] In view of this, the present invention aims to provide a conference enhancement method, apparatus, device, and storage medium that effectively address the shortcomings of existing solutions in multimodal conference information processing, participant status analysis, and decision-making process tracking, thereby improving the efficiency of meetings and the effectiveness and quality of meeting outcomes. The specific solution is as follows:
[0005] In a first aspect, the present application provides a conference enhancement method, which is applied to a preset conference enhancement system, wherein the preset conference enhancement system includes a data preprocessing component, a participant analysis component, a conference content processing component, a decision analysis component, and an intervention output component; the method includes:
[0006] The data preprocessing component performs corresponding data preprocessing on the original information of each mode in the multimodal original information of the conference, and performs multimodal data fusion and time alignment processing based on the preprocessing results to determine the target data packet;
[0007] Identify the identities of the participants using the participant analysis component and the target data packet, and perform participant status analysis based on the identification results to determine a target model for analyzing the behavior and status of the participants, and determine the speaker context of the conference based on the target model;
[0008] Performing a structured representation of conference content through the conference content processing component, the speaker context, and the target data packet, and performing argument analysis based on the structured content to determine an argument analysis result;
[0009] By using the decision analysis component, the target model, the structured content, and the argumentation analysis results, and adopting a holistic perspective, the team interaction pattern and decision-making process of the meeting are analyzed to determine the analysis results;
[0010] Whether a meeting intervention is currently being performed is determined through the intervention output component, the parsing result, the structured content, the current meeting state information, and the historical intervention feedback information, and when so, a target intervention suggestion is determined and output.
[0011] Optionally, the data preprocessing component performs corresponding data preprocessing on the original information of each modality in the multimodal original information of the conference, and performs multimodal data fusion and time alignment processing based on the preprocessing results, including:
[0012] Receiving original signals sent by various hardware devices through the data preprocessing component; the hardware devices include microphones, cameras, and environmental sensors;
[0013] digitizing and normalizing the original signal to determine processed audio data, video data, and environmental data;
[0014] Performing noise filtering on the audio data based on a first preset algorithm to determine first preprocessed data;
[0015] Identifying a region of interest on the video data based on a second preset algorithm, and allocating resources based on the region identification result to determine second preprocessed data;
[0016] Performing an environmental state assessment on the environmental data based on a standardized environmental parameter vector and a multi-layer perceptron structure to determine third pre-processed data; the standardized environmental parameter vector includes a standardized humidity vector and a standardized temperature vector;
[0017] After adjusting the data acquisition configuration and resource allocation based on the hierarchical data acquisition protocol, the resource allocation algorithm and the currently determined first preprocessed data, the second preprocessed data and the third preprocessed data, data fusion is performed on the first preprocessed data, the second preprocessed data and the third preprocessed data, and the fused data is time-aligned through a timestamp alignment mechanism to determine the target data packet.
[0018] The method of identifying a participant through the participant analysis component and the target data packet, and performing participant status analysis based on the identification result to determine a target model for analyzing participant behavior and status, and determining the speaker context of the conference based on the target model, includes:
[0019] Simultaneously analyzing acoustic and visual features in the target data packet using the participant analysis component and the progressive recognition strategy to determine the participant identity and corresponding spatial position tracking results;
[0020] Based on the participant state analyzer in the participant analysis component, the participant identity, and the spatial position tracking result, the participant's current emotion and current cognitive state are evaluated using a voice-gesture-expression collaborative analysis framework to determine a participant evaluation result;
[0021] Analyze the historical speech content and interaction patterns of the participants based on the role and expertise modeler in the participant analysis component and the participant evaluation results to construct a knowledge model corresponding to the current identity of each participant;
[0022] Adjusting weights of different types of expertise and role attributes based on the knowledge model, the weight allocation mechanism, the identities of the participants, and the current agenda of the meeting to determine a participant identity mapping table;
[0023] The knowledge model is updated based on the interactive behavior adaptive modeling mechanism and the recorded behavioral characteristics of the participants, and the behavioral trends are captured and analyzed based on the updated knowledge model and the participant identity mapping table to determine the speaker context of the meeting; the speaker context includes the speaker identity information, the speaker status information and the speaker cognitive load information.
[0024] Optionally, the structured representation of the conference content using the conference content processing component, the speaker context, and the target data packet, and the argumentation analysis based on the structured content, includes:
[0025] Receiving the speaker context and the target data packet through the conference content processing component;
[0026] Performing text conversion on the target data packet to determine a conversion result;
[0027] Performing a timestamp and a speaker identity tag on the conversion result based on the speaker context to determine a tagged text;
[0028] Based on a content structured analyzer and a hierarchical attention network architecture, topic identification, semantic segmentation, and relationship extraction are performed on the labeled text to determine structured content;
[0029] The structured content is subjected to a structural argumentation analysis based on the speaker context, and the structured content is subjected to an information source analysis and an information reliability analysis based on preset analysis rules to determine an argumentation analysis result.
[0030] Optionally, analyzing the team interaction pattern and decision-making process of the meeting from a holistic perspective using the decision analysis component, the target model, the structured content, and the argumentation analysis results includes:
[0031] receiving, via the decision analysis component, the target model, the structured content, and the argumentation analysis result;
[0032] Constructing and analyzing a time sequence diagram based on the interactive network analyzer in the decision analysis component and the structured content, and analyzing the influence propagation path, information bottleneck points, and key connector roles in the determined dynamic interactive network to determine a first analysis result;
[0033] Identifying decision points on the structured content based on a decision process tracker in the decision analysis component, and tracking the evolution of the decision points based on the decision point identification results, analyzing the decision formation process based on the tracking results to determine a second analysis result;
[0034] Performing a multi-dimensional quantitative evaluation based on the evaluator in the decision analysis component, the meeting objectives of the meeting, the argumentation analysis results, the first analysis results, and the second analysis results to determine a corresponding quality report;
[0035] Hidden obstacle detection is performed based on the dynamic interactive network, and whether to trigger a cognitive gap marking operation is determined according to the obstacle monitoring result and the target model.
[0036] Optionally, determining whether to currently perform a meeting intervention through the intervention output component, the parsing result, the structured content, the current meeting status information, and the historical intervention feedback information, and determining and outputting a target intervention suggestion when the intervention is performed, includes:
[0037] Performing heterogeneous data integration on the received parsing results, the structured content, the speaker context, and the argument analysis results through the intervention output component, and prioritizing the data based on the integration results to determine a ranking result;
[0038] Determining whether to currently perform a meeting intervention based on the auxiliary generator in the intervention output component, the sorting result, the current meeting state information, and the historical intervention feedback information to obtain an intervention determination result;
[0039] When the intervention judgment result is yes, determining a target intervention suggestion and a corresponding intervention type based on the current meeting state information;
[0040] An intervention output mode corresponding to the target intervention suggestion is determined based on the intervention type, and an intervention output operation of the target intervention suggestion is triggered according to the intervention output mode; the target intervention suggestion includes intervention timing, intervention content, and intervention form.
[0041] Optionally, the method further includes:
[0042] After outputting the target intervention suggestion, collecting intervention feedback information corresponding to the target intervention suggestion through the intervention output component;
[0043] Performing an effectiveness evaluation of the intervention based on the intervention feedback information, and sending the intervention evaluation result to the decision analysis component so that the decision analysis component triggers an analysis optimization operation based on the intervention evaluation result;
[0044] After the meeting is over, the multimodal original information and the analysis results during the meeting are integrated through the intervention output component and based on the current target model, so as to determine and distribute post-meeting resources based on the integration results; the post-meeting resources include meeting summaries, decision records and learning resource recommendation information.
[0045] In a second aspect, the present application provides a conference enhancement device for use in a preset conference enhancement system, wherein the preset conference enhancement system includes a data preprocessing component, a participant analysis component, a conference content processing component, a decision analysis component, and an intervention output component; the device includes:
[0046] A data preprocessing module, configured to perform corresponding data preprocessing on the original information of each modality in the multimodal original information of the conference through the data preprocessing component, and perform multimodal data fusion and time alignment processing based on the preprocessing results to determine the target data packet;
[0047] a participant analysis module, configured to identify the identities of the participants using the participant analysis component and the target data packet, and perform participant status analysis based on the identification results to determine a target model for analyzing the behavior and status of the participants, and determine the speaker context of the conference based on the target model;
[0048] a content processing module, configured to perform a structured representation of the conference content using the conference content processing component, the speaker context, and the target data packet, and perform argument analysis based on the structured content to determine an argument analysis result;
[0049] A meeting analysis module is used to analyze the team interaction mode and decision-making process of the meeting from a holistic perspective through the decision analysis component, the target model, the structured content, and the argumentation analysis results to determine an analysis result;
[0050] The conference intervention module is used to determine whether conference intervention is currently being performed through the intervention output component, the parsing result, the structured content, the current conference status information and the historical intervention feedback information, and when so, determine and output a target intervention suggestion.
[0051] In a third aspect, the present application provides an electronic device, comprising:
[0052] Memory, used to store computer programs;
[0053] The processor is used to execute the computer program to implement the steps of the aforementioned conference enhancement method.
[0054] In a fourth aspect, the present application provides a computer-readable storage medium for storing a computer program, which implements the steps of the aforementioned conference enhancement method when executed by a processor.
[0055] It can be seen that the present application is applied to a preset conference enhancement system, and the preset conference enhancement system includes a data preprocessing component, a participant analysis component, a conference content processing component, a decision analysis component and an intervention output component; the method includes: performing corresponding data preprocessing on the original information of each mode in the multimodal original information of the conference through the data preprocessing component, and performing multimodal data fusion and time alignment processing based on the preprocessing result to determine the target data packet; performing participant identity recognition through the participant analysis component and the target data packet, and performing participant status analysis based on the recognition result to determine the target model for analyzing the participant behavior and status, and based on the target The model determines the speaker context of the meeting; the meeting content is structured through the meeting content processing component, the speaker context and the target data packet, and argument analysis is performed based on the structured content to determine the argument analysis results; the team interaction mode and decision-making process of the meeting are analyzed respectively from a holistic perspective through the decision analysis component, the target model, the structured content and the argument analysis results to determine the analysis results; the intervention output component, the analysis results, the structured content, the current meeting status information and the historical intervention feedback information are used to determine whether the meeting intervention is currently being performed, and when so, the target intervention suggestion is determined and output. That is, in this application, the meeting enhancement is achieved through a hierarchical, multi-module collaborative preset meeting enhancement system. Specifically, the data preprocessing component in the system is first used to preprocess the multimodal original information of the meeting, and multimodal data fusion and time alignment processing are performed to determine the target data packet. Then, the participant analysis component in the system and the target data packet are used to identify the identity of the participant and perform participant status analysis to determine the speaker context of the meeting using the target model obtained for analyzing the behavior and status of the participant. The system then uses the meeting content processing component and the speaker context to determine the structured content of the meeting and conduct argumentation analysis. The system then uses the decision analysis component, structured content, and argumentation analysis results to analyze the team interaction model and decision-making process of the meeting from a holistic perspective. The system then uses the intervention output component, analysis results, current meeting status information, and historical intervention feedback to determine whether to intervene in the meeting. If intervention is necessary, the target intervention recommendation is determined. This effectively addresses the shortcomings of existing solutions in multimodal meeting information processing, participant status analysis, and decision-making process tracking, improving the efficiency of meetings and the effectiveness and quality of meeting outcomes. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0057] Figure 1 A flow chart of a conference enhancement method provided for this application;
[0058] Figure 2 A schematic diagram of the interaction flow of various components in a preset conference enhancement system provided by this application;
[0059] Figure 3 A schematic diagram of the data processing flow of a data preprocessing component provided in this application;
[0060] Figure 4 A schematic diagram of the data processing flow of a participant analysis component provided for this application;
[0061] Figure 5 A data processing flow diagram of a conference content processing component provided in this application;
[0062] Figure 6 A schematic diagram of the data processing flow of a decision analysis component provided in this application;
[0063] Figure 7 A schematic diagram of the data processing flow of an intervention output component provided in this application;
[0064] Figure 8 A schematic diagram of the structure of a conference enhancement device provided in this application;
[0065] Figure 9 This is a structural diagram of an electronic device provided in this application. DETAILED DESCRIPTION
[0066] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0067] Existing conference enhancement solutions solve the problem through voice transcription and recording systems, and basic collaboration platforms that provide functions such as document sharing and screen sharing. However, these solutions often only support the processing of a single type of data. Due to the lack of comprehensive analysis of the emotional state of the participants, they lack the ability to identify implicit communication barriers. In addition, the use of fixed strategies for auxiliary intervention prompts often leads to untimely or inaccurate intervention. To this end, the present application provides a conference enhancement solution that can effectively solve the shortcomings of existing related solutions in multimodal conference information processing, participant status analysis, and decision-making process tracking, thereby improving the efficiency of meetings and the effectiveness and quality of meeting results.
[0068] See also Figure 1 As shown, an embodiment of the present invention discloses a conference enhancement method, which is applied to a preset conference enhancement system. The preset conference enhancement system includes a data preprocessing component, a participant analysis component, a conference content processing component, a decision analysis component, and an intervention output component. The method includes:
[0069] Step S11: perform corresponding data preprocessing on the original information of each mode in the multimodal original information of the conference through the data preprocessing component, and perform multimodal data fusion and time alignment processing based on the preprocessing results to determine the target data packet.
[0070] It is important to understand that, combined with Figure 2 The system architecture diagram of the preset conference enhancement system shown in the figure adopts a hierarchical module design and includes five core functional modules that work together: data acquisition and preprocessing module (i.e., data preprocessing component), participant analysis module (i.e., participant analysis component), content understanding and structuring module (i.e., conference content processing component), group dynamics and decision analysis module (i.e., decision analysis component) and intelligent assistance and output module (i.e., intervention output component). These modules form a complete intelligent conference management link through a carefully designed interaction mechanism.
[0071] Specifically, in the data processing flow of the system, the data preprocessing component first receives multimodal original information such as audio, video and environmental data in the meeting, and preprocesses the original information of each modality respectively. That is, the original signal sent by each hardware device is first received through the data preprocessing component; the hardware devices include microphones, cameras and environmental sensors; the original signal is digitized and standardized to determine the processed audio data, video data and environmental data; the audio data is noise filtered based on the first preset algorithm to determine the first preprocessed data; the video data is identified for the region of interest based on the second preset algorithm, and resources are allocated based on the region identification result to determine the second preprocessed data; the environmental data is evaluated for its environmental status based on the standardized environmental parameter vector and the multi-layer perceptron structure to determine the third preprocessed data; the standardized environmental parameter vector includes a standardized humidity vector and a standardized temperature vector; after adjusting the data acquisition configuration and resource allocation based on the hierarchical data acquisition protocol, the resource allocation algorithm and the currently determined first preprocessed data, second preprocessed data and third preprocessed data, the first preprocessed data, the second preprocessed data and the third preprocessed data are data fused, and the fused data are time-aligned through the timestamp alignment mechanism to determine the target data packet. It is important to understand that the data preprocessing component uses a multi-channel parallel processing architecture to simultaneously process multimodal raw information such as audio, video, and environmental data, and achieve preliminary synchronization and alignment across modalities.
[0072] Regarding the audio data processing flow, audio processing utilizes context-aware adaptive filtering technology, implemented through the following steps: first, using time-frequency domain analysis to decompose the audio signal into a time-frequency representation. A neural network structure with long-short-term memory (LSTM) is then applied to model the temporal evolution of noise. The network then receives the current frame's acoustic feature vector and historical state as input, and calculates the noise probability distribution through LSTM network units. This network is responsible for capturing temporal dependencies and optimally utilizing the generated noise probabilities. The system then enhances the original signal through a soft masking mechanism. This process not only considers current acoustic features but also incorporates historical information and knowledge of the meeting context, effectively suppressing various types of noise interference while preserving speech integrity.
[0073] Regarding the video data processing flow, visual processing uses an attention region priority mechanism: first, a low-complexity global scan is performed to identify potential regions of interest using a lightweight feature detector. Then, dynamic importance weights are assigned to each region based on the meeting's dynamic information and historical interaction patterns. A comprehensive importance assessment is then performed, considering three factors: the importance score based on the current speaking status, the influence factor of historical interaction patterns, and the complexity of the regional content. Finally, based on the calculated importance map, the system prioritizes computing resources for high-weighted regions. This differentiated processing strategy enables the system to capture key visual information in meetings while meeting real-time requirements.
[0074] Regarding the analysis process of environmental data, environmental processing establishes a correlation model between environmental parameters and cognitive performance: first receive the standardized environmental parameter vector (including temperature, humidity, concentration, etc.), then uses a multi-layer perceptron architecture to perform nonlinear transformations, mapping environmental parameters into scores of their impact on cognitive performance. Finally, based on the output scores, the system generates adjustment recommendations when environmental parameters deviate from their optimal ranges. The resulting association model, based on extensive empirical research data, establishes a mapping between environmental factors and meeting efficiency indicators.
[0075] Afterwards, the target data package is sent to the participant analysis component, enabling the latter to perform participant modeling based on high-quality data.
[0076] Step S12: identify the identity of the participant through the participant analysis component and the target data packet, and perform participant status analysis based on the identification result to determine a target model for analyzing the behavior and status of the participant, and determine the speaker context of the meeting based on the target model.
[0077] In this embodiment, the participant analysis component will use the data transmitted by the data preprocessing component to perform identity recognition and content modeling, converting the participants' explicit behaviors and implicit states into machine-understandable structured representations, providing important contextual support for subsequent meeting content analysis. That is, first, the acoustic features and visual features in the target data packet are analyzed simultaneously through the participant analysis component and the progressive recognition strategy to determine the participant identity and the corresponding spatial position tracking results; based on the participant state analyzer, participant identity and spatial position tracking results in the participant analysis component, the current emotion and current cognitive state of the participant are evaluated using the voice-gesture-expression collaborative analysis framework to determine the participant evaluation results; based on the role and expertise modeler and participant evaluation results in the participant analysis component, the historical speech content and interaction mode of the participant are analyzed to construct a knowledge model corresponding to the current identity of each participant; based on the knowledge model, weight distribution mechanism, participant identity and the current topic of the meeting, the weights of different types of expertise and role attributes are adjusted to determine the participant identity mapping table; based on the interactive behavior adaptive modeling mechanism and the recorded participant behavior characteristics, the knowledge model is updated, and the behavior trend is captured and analyzed based on the updated knowledge model and participant identity mapping table to determine the speaker context of the meeting; the speaker context includes speaker identity information, speaker state information and speaker cognitive load information. It can be understood that Figure 2 The participant model in is the target model, and the state data represents the state data related to the participant.
[0078] Regarding the identity recognition process, the multi-feature fusion identity recognition algorithm is implemented through the following steps:
[0079] 1) Construct a two-stream neural network architecture. The first stream processes short-term acoustic features (such as fundamental frequency distribution, harmonic structure, and formant characteristics), while the second stream analyzes long-term speaking style features (such as vocabulary usage habits, syntactic structure preferences, and expression patterns).
[0080] 2) Both streams produce identity embedding vectors;
[0081] 3) Perform weighted fusion through the attention mechanism to form the final identity representation;
[0082] 4) Adopting a progressive recognition strategy, a set of possible identity candidates is quickly generated and assigned an initial probability distribution at the beginning of the meeting. As more interaction data is acquired, the distribution is continuously updated until a high-confidence recognition is achieved.
[0083] Regarding the state analysis process, the participant state analyzer adopts a speech-posture-expression collaborative analysis framework: first, the original features are extracted from each modality. Speech features include pitch changes, speaking speed, and rhythmic features. Visual features include: facial expression units, eye focus, and head posture. Then, the complementary and consistent relationships between modalities are modeled through a graph neural network structure. Different modal features are represented as nodes in the graph, and the associations between modalities are represented as edges. Information fusion is achieved through the inter-node message passing mechanism to handle incomplete modal situations. When the participant is temporarily out of the camera's field of view, a reliable state assessment can still be given based on voice and historical status.
[0084] Regarding the knowledge modeling process, the role and expertise modeler adopts an automatic generation method of expertise graphs based on speech content: first, professional terms and key concepts are extracted from the speech text, and then the initial knowledge graph is constructed based on co-occurrence analysis and semantic association. The expression patterns of participants and domain experts are compared, and their professional depth and influence at each knowledge node are evaluated to form a dynamically updated personal knowledge model, provide a reference for content understanding and intervention strategies, implement a dynamic weight distribution mechanism, and automatically adjust the importance of different expertise and role attributes according to the current discussion topic.
[0085] Regarding interactive behavior modeling, an adaptive interactive behavior modeling mechanism is used to continuously learn and update the interaction patterns of participants: tracking the behavioral characteristics of participants such as speaking frequency, response patterns, and influence performance. Then, as the meeting progresses, the model is continuously updated to capture the evolution trend of participants' behavior (such as from initial conservative observation to later active participation), and analyze role changes under different topics (such as professional dominance of one topic to sideways listening on another topic).
[0086] Step S13: Structured representation of the conference content is performed through the conference content processing component, the speaker context, and the target data packet, and argument analysis is performed based on the structured content to determine an argument analysis result.
[0087] Specifically, in this embodiment, at the content processing layer, the meeting content processing component performs deep semantic analysis and structured representation of the meeting content based on information transmitted by the participant analysis component. Through deep semantic analysis and contextual reasoning, it captures the literal information and implicit intent, logical relationships, and knowledge elements of the meeting content, providing a foundation for subsequent decision analysis. Specifically, the meeting content processing component receives the speaker context and target data packet; performs text conversion on the target data packet to determine the conversion result; timestamps and identifies the speaker based on the speaker context to determine the tagged text; performs topic identification, semantic segmentation, and relationship extraction on the tagged text based on a content structured analyzer and a hierarchical attention network architecture to determine the structured content; performs structural argumentation analysis on the structured content based on the speaker context, and performs information source analysis and information reliability analysis on the structured content based on preset analysis rules to determine the argumentation analysis results. The structured meeting content, argumentation analysis results, and target model are then passed to the decision analysis component to support higher-level interaction and decision analysis.
[0088] Regarding the speech understanding process, the context-aware professional speech recognition system is implemented in the following way: first, a basic language model for domain adaptation is built using transfer learning methods, and then new terms are detected and learned in real time during the meeting: low-confidence areas are identified; whether the area contains important information (based on the context and the speaker's professional background) is determined, and the expression is learned from subsequent context, continuously improving domain adaptability as the meeting progresses.
[0089] Regarding the content structuring process, a dynamic adaptation mechanism based on a conference-specific language model employs a hierarchical attention network architecture, capturing semantic units of varying granularity through multi-level attention computation: word-level identification of key terms and sentiment words; sentence-level identification of key statements and discourse acts; and paragraph-level establishment of topic hierarchies. This process also integrates speaker roles, sentiment markers, and interaction context to achieve comprehensive semantic understanding. A progressive structuring strategy is employed, initially quickly identifying the overall meeting framework and key topics, then gradually refining the content structure as the discussion deepens.
[0090] Regarding the argumentation and quality assessment process, based on the identification of implicit premises within the meeting context, a complete argumentation diagram is first constructed, including clearly expressed premises, conclusions, and supporting relationships. Breakpoints in the logical chain are then detected and compared with the domain knowledge graph to identify possible bridging concepts. Finally, potential implicit premises are marked, prompting participants to confirm or clarify if necessary. When assessing information quality and reliability, a credibility score is first assigned based on indicators such as information source, supporting evidence, and internal consistency. Fact verification is then conducted using the enterprise knowledge base and trusted data sources. The value of the meeting content is then assessed based on relevance, novelty, and reliability. Finally, based on the overall value score, attention resources are allocated to highlight high-value information.
[0091] Step S14: using the decision analysis component, the target model, the structured content, and the argumentation analysis results, and adopting a holistic perspective to analyze the team interaction pattern and decision-making process of the meeting, respectively, to determine the analysis results.
[0092] Specifically, in this embodiment, the decision analysis component of the decision-making layer is responsible for analyzing team interaction patterns and the decision-making process. By focusing on the collaborative dynamics, influence mechanisms, and decision paths at the group level, it reveals the implicit social structure and cognitive processes in the meeting, providing a deep theoretical basis for meeting optimization. Specifically, the decision analysis component first receives the target model, structured content, and argument analysis results. Based on the interaction network analyzer in the decision analysis component and the structured content, a time sequence diagram is constructed and analyzed. The influence propagation paths, information bottlenecks, and key connector roles in the determined dynamic interaction network are analyzed to determine a first analysis result. Based on the decision process tracker in the decision analysis component, decision points are identified in the structured content. Based on the decision point identification results, the evolution of the decision points is tracked. Based on the tracking results, the decision-making process is analyzed to determine a second analysis result. A multi-dimensional quantitative assessment is performed based on the evaluator in the decision analysis component, the meeting goals, the argument analysis results, the first analysis results, and the second analysis results to determine a corresponding quality report. Hidden obstacles are detected based on the dynamic interaction network, and based on the obstacle monitoring results and the target model, a cognitive gap marking operation is determined.
[0093] Regarding the interactive network analysis process, dynamic modeling and prediction of time-varying interactive graphs:
[0094] 1) Representing meeting interactions as a dynamically evolving graph structure with participants as nodes and interactions as edges, capturing various interactive behaviors (direct conversations, responses, quotations, non-verbal feedback, etc.);
[0095] 2) Using a temporal graph neural network architecture, we construct an interaction graph at each time point, extract node representations and graph-level representations through a graph neural network layer, and capture temporal dependencies through a recurrent neural network layer;
[0096] 3) Identify key patterns, including: influencing communication paths, information bottlenecks, and key connector roles.
[0097] Regarding the decision analysis process, multi-level decision trees are constructed and evaluated in real time: first, the decision points and option branches are identified, then the basis and assumptions supporting each option are extracted, and then a dynamically growing decision tree is constructed to reflect the logical development of the decision under discussion and evaluate the quality of the decision process, including: sufficiency of evidence, degree of hypothesis verification, and breadth of exploration of alternative options. The Bayesian network model is used to represent decision dependencies and to evaluate the impact of changes in preconditions on decision consequences.
[0098] Regarding the efficiency evaluation process, it is used to quantitatively evaluate the meeting value generation process: first, consider multiple dimensions: time utilization (the ratio of effective discussion time to total meeting time), participation balance (uniformity of distribution of participation opportunities), information quality (the ratio of new information generation to redundant information), decision progress (decision point resolution rate), and then integrate multi-dimensional indicators through a hierarchical analysis process, and dynamically adjust the weight of each indicator according to the meeting type and goals. Finally, compare the current meeting with the historical performance benchmark to identify efficiency trends.
[0099] Regarding dynamic importance perception, it is used to automatically identify key moments and turning points in meetings: first analyze multiple signals (changes in interaction patterns, emotional markers, sudden changes in keyword frequency), then evaluate the importance of the discussion in real time, and then adjust the processing depth and attention resources accordingly. Finally, conduct in-depth analysis of key paragraphs and use lightweight processing for routine discussions.
[0100] Hidden barriers identification is used to uncover hidden factors that hinder effective communication and decision-making. First, three common hidden barriers are detected: differences in terminology (different meanings assigned to the same term), conflicts in implicit assumptions (reasoning based on different, unspoken premises), and expertise gaps (understanding barriers caused by a lack of shared background knowledge). Then, the system compares the terminology usage patterns, reasoning paths, and knowledge models of the participants, flagging potential cognitive gaps and providing bridging information when appropriate to promote effective communication.
[0101] Step S15: Determine whether conference intervention is currently being performed through the intervention output component, the parsing result, the structured content, the current conference status information, and the historical intervention feedback information, and if so, determine and output a target intervention suggestion.
[0102] In this embodiment, the intervention output component generates real-time intervention suggestions and post-meeting resources based on the results of the aforementioned analysis, and continuously collects feedback on the intervention effect to dynamically optimize the entire system. In this way, by converting the pre-process analysis into actual action, including real-time assistance during the meeting and output generation after the meeting, it constitutes a key link in realizing the value of the system. Specifically, regarding intervention, the intervention output component first integrates the received parsing results, structured content, speaker context, and argument analysis results, and prioritizes the data based on the integration results to determine the sorting result. Based on the auxiliary generator in the intervention output component, the sorting result, the current meeting status information, and historical intervention feedback information, it determines whether to intervene in the meeting at the moment to obtain an intervention judgment result. If the intervention judgment result is yes, the target intervention suggestion and the corresponding intervention type are determined based on the current meeting status information. Based on the intervention type, the intervention output method corresponding to the target intervention suggestion is determined, and the intervention output operation of the target intervention suggestion is triggered according to the intervention output method. The target intervention suggestion includes the intervention timing, intervention content, and intervention form.
[0103] Regarding the intervention decision-making process, a precise intervention algorithm based on contextual awareness: first, the intervention is regarded as a continuous decision-making problem, with the goal of maximizing the intervention utility while minimizing the intervention cost; then, based on the reinforcement learning framework, the current meeting status (discussion content, participant status, meeting progress, etc.) is observed, and intervention actions are selected (whether to intervene, intervention content, intervention form, etc.) to obtain immediate benefits (based on intervention effect evaluation) and move to a new state; then, the intervention decision takes multiple factors into consideration: the timing of intervention: based on the meeting rhythm, current importance, and cognitive load of the participants; the content of intervention: providing missing information, clarifying terms, prompting logical loopholes, and suggesting related knowledge; the form of intervention: visual prompts, text suggestions, and post-meeting annotations; focus on non-invasive intervention, inserting prompts at natural pause points; using terminology familiar to the participants; and adjusting the significance of prompts according to the degree of urgency.
[0104] Regarding the visualization process, complex data is transformed into visual representations. Data preprocessing begins with dimensionality reduction, clustering, and key feature extraction. Visualization templates (decision tree diagrams, interactive network diagrams, and topic evolution diagrams) are then selected based on the data type and purpose. The visualization complexity is adjusted based on the user's cognitive load and professional background, and the visualization content is integrated into the user interface. Next, information presentation is optimized based on cognitive load. Appropriate visualization strategies are selected based on the participants' professional backgrounds, attention spans, and task urgency. Simplified visual representations are employed in high-pressure decision-making situations to highlight core information. Detailed, multi-layered visualizations are then provided during the in-depth analysis phase.
[0105] The meeting resource generation process, based on a multi-level, multi-angle adaptive summarization algorithm, first constructs a hierarchical representation of the meeting content (topic layer, issue layer, viewpoint layer, detail layer). Then, a customized summary is generated based on user needs and roles: Decision makers are provided with a concise summary focusing on decision points and conclusions; Executors are provided with detailed action items and background information; and absentees are provided with a panoramic overview of key discussions and decisions. Next, action items are extracted and organized: commitments, task assignments, and decisions are automatically identified, followed by the extraction of key elements (responsible persons, deadlines, dependencies, and priorities). Action items are then structured and integrated with enterprise task management systems to establish an action item tracking mechanism and provide progress updates. Learning resource generation is also performed: based on the meeting content and the knowledge models of the participants, learning needs and knowledge gaps are identified, and targeted learning resources (term explanations, background knowledge links, and relevant document recommendations) are generated.
[0106] Furthermore, in this embodiment, after outputting the target intervention suggestion, the intervention feedback information corresponding to the target intervention suggestion is collected through the intervention output component; the effectiveness of the intervention is evaluated based on the intervention feedback information, and the intervention evaluation result is sent to the decision analysis component so that the decision analysis component triggers the analysis optimization operation based on the intervention evaluation result; after the meeting, the multimodal original information and analysis results during the meeting are integrated through the intervention output component and based on the current target model, so as to determine and distribute post-meeting resources based on the integration results; post-meeting resources include meeting summaries, decision records, and learning resource recommendation information. Combined with Figure 2 As can be seen, after the intervention output, the various layers of the system will form a closed-loop self-optimization mechanism through the established extensive feedback channels, including intervention effect evaluation, model update signals, and resource allocation instructions, to feed forward optimization layer by layer. Specifically, the intervention output component feeds the collected intervention feedback information back to the decision analysis component, which then analyzes the decision focus and feeds it back to the meeting content processing component. The meeting content processing component then performs semantic association and feeds back the optimized processing strategy to the participant analysis component. The participant analysis component then guides the focus and feeds it forward to the data preprocessing component.
[0107] In summary, the preset conference enhancement system in this embodiment simultaneously collects and processes audio, video, and environmental data through a multi-channel parallel processing architecture; utilizes a multi-feature fusion algorithm to identify the identities and analyze the status of participants; understands the content of the conference based on context-aware professional speech recognition and multi-level structured analysis; analyzes group dynamics and decision-making processes through time-varying interaction graphs and multi-level decision trees; and ultimately achieves precise intervention based on contextual awareness and multi-level adaptive resource generation. This invention effectively addresses the shortcomings of existing conference systems in terms of professional terminology recognition, participant status analysis, implicit communication barrier identification, and decision-making process tracking, thereby improving conference efficiency and decision-making quality. It is particularly suitable for high-value conference scenarios involving cross-domain expertise and complex decisions.
[0108] As can be seen, the present application implements conference enhancement through a hierarchical, multi-module collaborative pre-configured conference enhancement system. Specifically, the system's data preprocessing component first preprocesses the multimodal raw information of the conference, and performs multimodal data fusion and time alignment to determine the target data packet. The system's participant analysis component and target data packet are then used to identify the participants and perform participant status analysis. The resulting target model for analyzing participant behavior and status is then used to determine the speaker context of the meeting. The system's meeting content processing component and speaker context are then used to determine the structured content of the meeting and perform argumentation analysis. The system's decision analysis component, structured content, and argumentation analysis results are then used to analyze the meeting's team interaction patterns and decision-making process from a holistic perspective. The system's intervention output component, along with the analysis results, current meeting status information, and historical intervention feedback, determines whether to intervene in the meeting. If intervention is necessary, a target intervention recommendation is then determined. This effectively addresses the shortcomings of existing solutions in multimodal conference information processing, participant status analysis, and decision-making process tracking, improving meeting efficiency and the effectiveness and quality of meeting outcomes.
[0109] The following combination Figure 2-Figure 7 The schematic diagram disclosed in the figure specifically illustrates the technical solution of the embodiment of the present application.
[0110] Combine Figure 2 It can be seen that a hierarchical, multi-module collaborative conference cognition framework is deployed in the preset conference enhancement system in this embodiment.
[0111] Combine Figure 3As shown in the figure, the data preprocessing component in the system is the perception foundation of the entire system. It is responsible for obtaining raw information from multiple data sources and performing preliminary processing to provide high-quality input for subsequent analysis. In one specific embodiment, the component is internally organized into three collaborative sub-units: an audio acquisition and preprocessing unit, a visual data acquisition unit, and an environment and context perception unit. Each sub-unit adopts a specialized processing strategy for a specific data type while maintaining overall coordination. The data flow path within the component is as follows:
[0112] (1) Receive the original signals from each hardware device (microphone array sound waveform, camera video frame sequence, environmental sensor numerical reading);
[0113] (2) Conduct preliminary digitization and standardization;
[0114] (3) Divert to special processing channels:
[0115] 1) Audio preprocessing unit, which filters ambient noise from audio data using an adaptive noise suppression algorithm;
[0116] 2) Video data pre-processing unit, which performs preliminary analysis of video data through a regional priority processing mechanism;
[0117] 3) Environmental pre-processing unit, which integrates with preset meeting information to form an environmental status assessment;
[0118] (4) Establish cross-modal associations through the timestamp alignment mechanism, laying the foundation for subsequent fusion analysis.
[0119] It's important to understand that before modal fusion and time alignment, the data preprocessing component also utilizes a resource allocation and optimization mechanism for dynamic resource management. This management implementation includes: a hierarchical collection protocol that dynamically adjusts data collection configuration based on meeting type and importance; an intelligent resource allocation algorithm that employs a reinforcement learning framework and treats resource allocation as a sequential decision-making problem. The system uses intelligent agents to continuously observe the current state and select resource allocation actions to achieve immediate benefits. Through repeated training, the system learns the optimal resource allocation strategy for different meeting scenarios, ensuring the most valuable information is obtained within limited computing and storage resources. The final data package includes a clear audio stream with speaker location markers, a visual tracking data package with facial ROI (Region of Interest) and expression markers, and an environmental status report with meeting context.
[0120] In addition, regarding the interaction and feedback mechanism of the data preprocessing component, it realizes bidirectional information flow with other modules:
[0121] 1) Provide processed high-quality data to downstream modules;
[0122] 2) Receive feedback signals to adjust processing strategies (e.g., increase sampling rate or processing depth in specific areas);
[0123] 3) Dynamically adjust processing parameters according to actual needs to maximize overall system performance.
[0124] Combine Figure 4 As shown, regarding the participant analysis component in the system, it is the core perception unit of the system, responsible for identifying meeting participants and building their behavior and state models. This component occupies a key position in the system architecture, and is responsible for: receiving high-quality audio and video streams from the data acquisition and preprocessing module, performing in-depth feature extraction and multi-dimensional analysis; outputting a complete dynamic model of the participant, including information such as identity, location, emotional state, cognitive load, attention level, and professional knowledge structure; and providing a comprehensive human context for the understanding and analysis of meeting content. In addition, this component consists of three collaborative functional units: an identity recognition and tracker for determining the identity of meeting participants and tracking their physical locations, a participant state analyzer for real-time assessment of the participant's emotions, cognition, and attention state; and a role and expertise modeler for building the participant's professional background and meeting role model. These three units work together through information sharing and complementary verification mechanisms to build a complete and consistent participant representation. Regarding the data processing flow of this component:
[0125] 1) Receive pre-processed audio stream, video features, and pre-set meeting participant information (as well as historical participant data including organization and meeting records in the figure);
[0126] 2) The identity recognition and tracking device simultaneously analyzes acoustic and visual features to generate a participant identity map and spatial position tracking;
[0127] 3) The recognition results are passed to the participant state analyzer, which combines voice features, facial expressions and body language to assess the participant’s current emotional and cognitive state;
[0128] 4) At the same time, the role and expertise modeler analyzes the historical speech content and interaction patterns of the participants to build knowledge domain and meeting role models;
[0129] 5) The three pieces of information are finally integrated into a complete participant model and passed to subsequent components for use.
[0130] After obtaining these three pieces of information, the participant analysis component needs to adjust the importance of each topic based on a dynamic weighting mechanism and continuously update the participant model based on adaptive modeling of interactive behavior. Finally, the data is integrated to produce an output, which includes a participant identity mapping table with spatial location, a data stream of emotional and cognitive states related to real-time emotions and attention, and a dynamic participant model related to professional fields and interaction styles. Furthermore, this output can be used to provide optimization feedback to the aforementioned identity recognition and tracking components, participant status analyzer, and role and expertise modeler.
[0131] Combine Figure 5 As shown, regarding the meeting content processing component in the system, it is the core cognitive processing unit of the system, responsible for converting unstructured communication in the meeting into structured knowledge representation that can be understood by machines. The position of this component in the system architecture is responsible for: receiving the audio stream from the data acquisition preprocessing module and the speaker context from the participant analysis module; performing deep semantic understanding and structured processing; outputting a multi-level representation of the meeting content (text transcription, topic structure, argumentation framework, knowledge association and information value assessment); and providing a semantic basis for group dynamics and decision analysis. This component consists of three core units: a speech understanding and transcription unit for converting audio into text and annotating basic semantic features, a content structuring analyzer for identifying topic structure and semantic relationships, and an argumentation and quality assessment unit for analyzing the logical framework of the discussion and the quality of the information. The data processing flow in this component is as follows:
[0132] 1) Receive audio stream and speaker information;
[0133] 2) The speech understanding and transcription unit converts it into text with timestamps and speaker tags;
[0134] 3) The content structure analyzer performs topic identification, semantic segmentation, and relationship extraction to generate a hierarchical representation of the conference content;
[0135] 4) At the same time, the Argumentation and Quality Assessment Unit analyses the argument structure, information sources and reliability;
[0136] 5) Integrate the analysis results into a unified structured representation and pass it to subsequent modules.
[0137] Furthermore, the information output by the meeting content processing component provides forward learning feedback. This component establishes multi-directional interaction with other modules: it receives speaker context (identity, emotional state, and professional background) from the participant analysis module; provides structured meeting content and argument analysis to the group dynamics and decision analysis module; and receives interaction network and decision path analysis from the group dynamics module. This feedback is then used to optimize content structure and importance assessment. This multi-directional interaction ensures the coordinated evolution of the system's various modules, improves overall performance, and provides comprehensive support for efficient meeting decision-making.
[0138] Combine Figure 6 As shown, regarding the decision analysis component in the system, it is the system's advanced cognitive processing unit, responsible for analyzing team interaction patterns and decision-making processes from a holistic perspective. The position of this component in the system architecture: located at the top of the analysis chain, responsible for receiving the participant model of the participant analysis module and the structured meeting content of the content understanding module; responsible for fusion analysis and high-order reasoning; responsible for outputting interactive network analysis, decision path evaluation and efficiency quality reports; responsible for providing intervention basis for the intelligent assistance module and feedback to the preceding module to optimize the processing strategy. The component is composed of three functional units: an interactive network analyzer for modeling interaction patterns and influence propagation between participants, a decision process tracker for identifying and analyzing decision points and their evolution paths, and an efficiency and quality evaluator for evaluating meeting efficiency and outcome quality from multiple dimensions. About the data processing flow of this component:
[0139] 1) Receive participant status and interaction records;
[0140] 2) Interaction Network Analyzer constructs dynamic interaction networks through timing graph analysis;
[0141] 3) Decision Process Tracker identifies decision points based on structured meeting content and tracks their evolution;
[0142] 4) The efficiency and quality evaluator integrates the analysis results of the first two units and generates a multi-dimensional evaluation based on the meeting goals;
[0143] 5) Integrate analysis results into group dynamics and decision-making reports to guide intelligent intervention and meeting optimization.
[0144] Furthermore, the dynamic importance perception mechanism and output information within the decision analysis component feed back into the interactive network analyzer. Its interactions with other modules include receiving status and role information from the participant analysis module; receiving structured content and argument analysis from the content understanding module; providing interaction networks, decision paths, and efficiency assessments to the intelligent assistance and output module; and receiving intervention effectiveness assessments from the intelligent assistance module to optimize the analysis model. This multi-directional interaction enables this module to continuously optimize its analytical framework, providing accurate and valuable insights into group dynamics, supporting improved meeting efficiency and decision quality.
[0145] Combine Figure 7 As shown, regarding the intervention output component in the system, it is the terminal interactive interface of the entire system, responsible for converting in-depth analysis into practical interventions and valuable resources. Position in the system architecture: Located at the end of the execution chain, it is responsible for receiving the analysis results of all previous modules, especially the high-level insights of the group dynamics and decision analysis modules; responsible for comprehensive processing and strategic decision-making; responsible for generating user-oriented interventions and resources; responsible for providing feedback on intervention effects and supporting the overall optimization cycle of the system. It is composed of three functional units: a real-time auxiliary generator for providing timely and appropriate interventions during the meeting, a visualization engine for converting complex analysis results into intuitive and understandable visual expressions, and a conference resource generator for integrating meeting content and analysis results to generate post-meeting resources. Its data processing flow is as follows:
[0146] 1) Receive multiple result streams from the previous analysis (participant status data, structured meeting content, interaction network analysis, decision evaluation, etc.);
[0147] 2) Integrate and prioritize heterogeneous data;
[0148] 3) The real-time assistance generator decides whether to intervene, when to intervene, and the content and form of intervention based on the meeting status and historical intervention effects;
[0149] 4) Select appropriate presentation methods based on the intervention type: a visualization engine generates visual content, and a dedicated processing unit generates text and audio prompts;
[0150] 5) Deliver intervention content through the user interface and record intervention time and content;
[0151] 6) Accumulate and analyze meeting content in parallel and continuously update the resource model;
[0152] 7) After the meeting, generate customized meeting summaries, action lists, decision records, and learning resources;
[0153] 8) Distribute post-conference resources and collect feedback on their use.
[0154] Furthermore, the intervention output component will implement an optimized feedback loop, requiring monitoring of intervention effectiveness. This involves recording detailed information about each intervention (time, content, and target recipients), tracking changes in the meeting following the intervention (attention shifts, discussion adjustments, and decision-making progress), comparing these changes with expected outcomes, evaluating intervention effectiveness, and using the evaluation results to optimize subsequent interventions and provide feedback to the system model. A continuous learning mechanism is also required: tracking the long-term implementation results of meeting decisions, linking these results with meeting process analysis, identifying areas for improvement, and providing targeted support for similar discussions in the future. Furthermore, user experience adaptation involves recording and analyzing user response patterns to interventions and resources, building personalized user models, and adjusting intervention strategies and resource generation parameters to match user preferences. This comprehensive intelligent assistance and output design ensures the system provides maximum support with minimal disruption, achieving the goal of "augmenting" rather than "replacing" human collaboration, thereby comprehensively improving meeting efficiency and decision-making quality.
[0155] See also Figure 8 As shown, the embodiment of the present application also discloses a conference enhancement device, which is applied to a preset conference enhancement system. The preset conference enhancement system includes a data preprocessing component, a participant analysis component, a conference content processing component, a decision analysis component, and an intervention output component; the device includes:
[0156] The data preprocessing module 11 is configured to perform corresponding data preprocessing on the original information of each mode in the multimodal original information of the conference through the data preprocessing component, and perform multimodal data fusion and time alignment processing based on the preprocessing results to determine the target data packet;
[0157] a participant analysis module 12 configured to identify the participant using the participant analysis component and the target data packet, and perform participant status analysis based on the identification result to determine a target model for analyzing the participant's behavior and status, and determine the speaker context of the conference based on the target model;
[0158] A content processing module 13 is configured to perform a structured representation of the conference content using the conference content processing component, the speaker context, and the target data packet, and perform argument analysis based on the structured content to determine an argument analysis result;
[0159] The meeting analysis module 14 is configured to analyze the team interaction pattern and decision-making process of the meeting from a holistic perspective using the decision analysis component, the target model, the structured content, and the argumentation analysis results to determine an analysis result.
[0160] The conference intervention module 15 is used to determine whether conference intervention is currently being performed through the intervention output component, the parsing result, the structured content, the current conference status information and the historical intervention feedback information, and to determine and output a target intervention suggestion when it is.
[0161] Among them, for more specific working processes of the above modules, please refer to the corresponding contents disclosed in the aforementioned embodiments, which will not be repeated here.
[0162] As can be seen, the present application implements conference enhancement through a hierarchical, multi-module collaborative pre-configured conference enhancement system. Specifically, the system's data preprocessing component first preprocesses the multimodal raw information of the conference, and performs multimodal data fusion and time alignment to determine the target data packet. The system's participant analysis component and target data packet are then used to identify the participants and perform participant status analysis. The resulting target model for analyzing participant behavior and status is then used to determine the speaker context of the conference. The system's meeting content processing component and speaker context are then used to determine the structured content of the conference and perform argumentation analysis. The system's decision analysis component, along with the structured content and argumentation analysis results, then analyzes the team interaction patterns and decision-making process of the conference from a holistic perspective. The system's intervention output component, along with the analysis results, current meeting status information, and historical intervention feedback, then determines whether to intervene in the meeting. If intervention is necessary, a target intervention recommendation is determined. This effectively addresses the shortcomings of existing solutions in multimodal conference information processing, participant status analysis, and decision-making process tracking, improving the efficiency of meetings and the effectiveness and quality of meeting outcomes.
[0163] In some specific embodiments, the data preprocessing module 11 can be specifically used to: receive the original signal sent by each hardware device through the data preprocessing component; the hardware device includes a microphone, a camera and an environmental sensor; digitize and standardize the original signal to determine the processed audio data, video data and environmental data; perform noise filtering on the audio data based on a first preset algorithm to determine the first preprocessed data; identify the region of interest of the video data based on a second preset algorithm, and allocate resources based on the region identification result to determine the second preprocessed data; evaluate the environmental status of the environmental data based on a standardized environmental parameter vector and a multi-layer perceptron structure to determine the third preprocessed data; the standardized environmental parameter vector includes a standardized humidity vector and a standardized temperature vector; after adjusting the data acquisition configuration and resource allocation based on the hierarchical data acquisition protocol, the resource allocation algorithm and the currently determined first preprocessed data, the second preprocessed data and the third preprocessed data, perform data fusion on the first preprocessed data, the second preprocessed data and the third preprocessed data, and time align the fused data through a timestamp alignment mechanism to determine the target data packet.
[0164] In some specific embodiments, the participant analysis module 12 can be specifically used to: simultaneously analyze the acoustic features and visual features in the target data packet through the participant analysis component and the progressive recognition strategy to determine the participant identity and the corresponding spatial position tracking result; based on the participant state analyzer in the participant analysis component, the participant identity and the spatial position tracking result, and using the voice-gesture-expression collaborative analysis framework to evaluate the participant's current emotion and current cognitive state to determine the participant evaluation result; based on the role and expertise modeler in the participant analysis component and the participant evaluation result, evaluate the participant's historical speech The content and interaction patterns are analyzed to construct a knowledge model corresponding to the current identities of each participant; based on the knowledge model, the weight distribution mechanism, the identities of the participants and the current agenda of the meeting, the weights of different types of professional knowledge and role attributes are adjusted to determine the participant identity mapping table; based on the interactive behavior adaptive modeling mechanism and the recorded participant behavior characteristics, the knowledge model is updated, and the behavior trends are captured and analyzed based on the updated knowledge model and the participant identity mapping table to determine the speaker context of the meeting; the speaker context includes the speaker identity information, the speaker status information and the speaker cognitive load information.
[0165] In some specific embodiments, the content processing module 13 can be specifically used to: receive the speaker context and the target data packet through the conference content processing component; perform text conversion on the target data packet to determine the conversion result; timestamp and speaker identity mark the conversion result based on the speaker context to determine the marked text; perform topic identification, semantic segmentation and relationship extraction on the marked text based on the content structured analyzer and the hierarchical attention network architecture to determine the structured content; perform structural argumentation analysis on the structured content based on the speaker context, and perform information source analysis and information reliability analysis on the structured content based on preset analysis rules to determine the argumentation analysis result.
[0166] In some specific embodiments, the meeting parsing module 14 can be specifically used to: receive the target model, the structured content and the argumentation analysis result through the decision analysis component; construct and analyze a timing diagram based on the interactive network analyzer in the decision analysis component and the structured content, and analyze the influence propagation path, information bottleneck points and key connector roles in the determined dynamic interactive network to determine a first analysis result; identify decision points of the structured content based on the decision process tracker in the decision analysis component, and track the evolution of the decision points based on the decision point identification results, so as to analyze the decision formation process based on the tracking results to determine a second analysis result; perform a multi-dimensional quantitative evaluation based on the evaluator in the decision analysis component, the meeting goal of the meeting, the argumentation analysis result, the first analysis result and the second analysis result to determine a corresponding quality report; perform hidden obstacle detection based on the dynamic interactive network, and determine whether to trigger a cognitive gap marking operation based on the obstacle monitoring results and the target model.
[0167] In some specific embodiments, the conference intervention module 15 can be specifically used to: integrate heterogeneous data of the received parsing results, the structured content, the speaker context and the argument analysis results through the intervention output component, and prioritize the data based on the integration result to determine the sorting result; judge whether to conduct conference intervention at present based on the auxiliary generator in the intervention output component, the sorting result, the current conference status information and the historical intervention feedback information to obtain an intervention judgment result; when the intervention judgment result is yes, determine the target intervention suggestion and the corresponding intervention type based on the current conference status information; determine the intervention output mode corresponding to the target intervention suggestion based on the intervention type, and trigger the intervention output operation of the target intervention suggestion according to the intervention output mode; the target intervention suggestion includes the timing of intervention, the content of intervention and the form of intervention.
[0168] In some specific embodiments, the conference enhancement device can also be used to: after outputting the target intervention suggestion, collect intervention feedback information corresponding to the target intervention suggestion through the intervention output component; evaluate the effectiveness of the intervention based on the intervention feedback information, and send the intervention evaluation result to the decision analysis component, so that the decision analysis component triggers analysis optimization operations based on the intervention evaluation result; after the meeting ends, integrate the multimodal original information and the analysis results during the meeting through the intervention output component and based on the current target model, so as to determine and distribute post-meeting resources based on the integration results; the post-meeting resources include meeting summaries, decision records, and learning resource recommendation information.
[0169] Furthermore, the embodiment of the present application also discloses an electronic device, Figure 9 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content in the diagram should not be considered as any limitation to the scope of application of the present application.
[0170] Figure 9 This is a schematic diagram of the structure of an electronic device 20 provided in an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 is used to store a computer program, which is loaded and executed by the processor 21 to implement the relevant steps of the conference enhancement method disclosed in any of the aforementioned embodiments. Furthermore, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0171] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and the external device. The communication protocol it follows is any communication protocol that can be applied to the technical solution of this application and is not specifically limited here; the input and output interface 25 is used to obtain external input data or output data to the outside world. Its specific interface type can be selected according to specific application needs and is not specifically limited here.
[0172] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or CD, etc. The resources stored thereon can include an operating system 221, a computer program 222, etc., and the storage method can be temporary storage or permanent storage.
[0173] The operating system 221 is used to manage and control the hardware devices on the electronic device 20 and the computer program 222. The operating system 221 can be Windows Server, NetWare, Unix, Linux, etc. In addition to including computer programs capable of implementing the conference enhancement method performed by the electronic device 20 as disclosed in any of the aforementioned embodiments, the computer program 222 can further include computer programs capable of performing other specific tasks.
[0174] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the aforementioned conference enhancement method is implemented. The specific steps of this method can be referred to the corresponding contents disclosed in the aforementioned embodiments and will not be repeated here.
[0175] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from the other embodiments. Reference can be made to the descriptions of the identical or similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and the relevant parts can be referred to the descriptions of the methods.
[0176] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0177] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0178] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0179] The above is a detailed introduction to the technical solution provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for those skilled in the art, according to the ideas of the present application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A conference enhancement method, characterized in that: Applied to a preset conference enhancement system, the preset conference enhancement system includes a data preprocessing component, a participant analysis component, a conference content processing component, a decision analysis component, and an intervention output component; the method includes: The data preprocessing component performs corresponding data preprocessing on the original information of each mode in the multimodal original information of the conference, and performs multimodal data fusion and time alignment processing based on the preprocessing results to determine the target data packet; wherein the multimodal original information includes audio data, video data and environmental data; Participant identity recognition is performed using the participant analysis component and the target data packet, and based on the recognition result, participant status analysis is performed using a speech-gesture-expression collaborative analysis framework to determine a target model for analyzing participant behavior and status, and based on the target model, a speaker context of the conference is determined; the speaker context includes speaker identity information, speaker status information, and speaker cognitive load information; Performing a structured representation of conference content through the conference content processing component, the speaker context, and the target data packet, and performing argument analysis based on the structured content to determine an argument analysis result; By using the decision analysis component, the target model, the structured content, and the argumentation analysis results, and adopting a holistic perspective, the team interaction pattern and decision-making process of the meeting are analyzed to determine the analysis results; Determining whether to currently perform a meeting intervention through the intervention output component, the parsing result, the structured content, the current meeting state information, and the historical intervention feedback information, and determining and outputting a target intervention suggestion when the intervention is performed; The data preprocessing component performs corresponding data preprocessing on the original information of each mode in the multimodal original information of the conference, and performs multimodal data fusion and time alignment processing based on the preprocessing results, including: Receiving original signals sent by various hardware devices through the data preprocessing component; the hardware devices include microphones, cameras, and environmental sensors; digitizing and normalizing the original signal to determine processed audio data, video data, and environmental data; Noise filtering is performed on the audio data based on a context-aware adaptive filtering algorithm to determine first preprocessed data; wherein the implementation process of the context-aware adaptive filtering algorithm includes: first decomposing the signal in the audio data into a time-frequency representation using time-frequency domain analysis, then applying a neural network structure with long-short-term memory capabilities to model the temporal evolution characteristics of the noise, and then the network receives the current frame acoustic feature vector and historical state as input, calculates the noise probability distribution, and captures the temporal dependency, and uses the generated noise probability to enhance the original audio signal through a soft masking mechanism; Identifying regions of interest on the video data based on an attention region priority processing mechanism, determining the dynamic importance weight of each region in the region identification result based on conference dynamic information and historical interaction patterns, performing a comprehensive importance assessment based on the current speech status importance score, the historical interaction pattern influence factor, and the content complexity of the region of interest, and allocating resources based on the calculated importance map to determine the second preprocessed data; Performing an environmental state assessment on the environmental data based on a standardized environmental parameter vector and a multi-layer perceptron structure to determine third pre-processed data; the standardized environmental parameter vector includes a standardized humidity vector and a standardized temperature vector; After adjusting data acquisition configuration and resource allocation based on the hierarchical data acquisition protocol, the resource allocation algorithm, and the currently determined first pre-processed data, the second pre-processed data, and the third pre-processed data, data fusion is performed on the first pre-processed data, the second pre-processed data, and the third pre-processed data, and time alignment of the fused data is performed using a timestamp alignment mechanism to determine a target data packet; The hierarchical data collection protocol includes dynamically adjusting data collection configuration based on the type and importance of the meeting. The resource allocation algorithm uses a reinforcement learning framework to treat resource allocation as a sequential decision problem, continuously observing the current state to determine and execute the optimal resource allocation strategy for the current meeting scenario. The target data packet includes a clear audio stream with speaker location markers, a visual tracking data packet with facial ROI and expression markers, and an environmental status report with a meeting context data packet. The decision analysis component, the target model, the structured content, and the argumentation analysis results are used to analyze the team interaction mode and decision-making process of the meeting from a holistic perspective, including: receiving, via the decision analysis component, the target model, the structured content, and the argumentation analysis result; Based on the interactive network analyzer in the decision analysis component and the structured content, the meeting interaction is represented as a dynamically evolving graph structure, with the participants as nodes and the interactions as edges, to capture various interactive behaviors. A time-series graph neural network architecture is used to construct a dynamic interactive network at each time point. The graph neural network layer extracts node representations and graph-level representations, and the recurrent neural network layer captures time dimension dependencies to analyze the influence propagation path, information bottleneck points, and key connector roles in the dynamic interactive network to determine a first analysis result. Identifying decision points and solution branches in the structured content based on a decision process tracker in the decision analysis component, extracting the basis and assumptions supporting each solution based on the identification results, constructing a dynamically growing decision tree, representing decision dependencies based on the decision tree and Bayesian network model, and evaluating the impact of changes in preconditions on decision consequences, thereby completing the analysis of the decision-making process and determining a second analysis result; Performing a multi-dimensional quantitative evaluation based on the evaluator in the decision analysis component, the meeting objectives of the meeting, the argumentation analysis results, the first analysis results, and the second analysis results to determine a corresponding quality report; Performing hidden obstacle detection based on the dynamic interactive network, and determining whether to trigger a cognitive gap marking operation based on obstacle monitoring results and the target model; The method further comprises: After outputting the target intervention suggestion, intervention feedback information is collected, intervention effectiveness evaluation is performed based on the intervention feedback information, and the intervention evaluation result is sent to the decision analysis component to trigger the analysis and optimization operation of the target model.
2. The conference enhancement method according to claim 1, characterized in that: The method includes: identifying the identity of the participant through the participant analysis component and the target data packet, and analyzing the participant status using a voice-gesture-expression collaborative analysis framework based on the identification result to determine a target model for analyzing the behavior and status of the participant, and determining the speaker context of the conference based on the target model, including: Simultaneously analyzing acoustic and visual features in the target data packet using the participant analysis component and the progressive recognition strategy to determine the participant identity and corresponding spatial position tracking results; Based on the participant state analyzer in the participant analysis component, the participant identity, and the spatial position tracking result, the participant's current emotion and current cognitive state are evaluated using a voice-gesture-expression collaborative analysis framework to determine a participant evaluation result; Analyze the historical speech content and interaction patterns of the participants based on the role and expertise modeler in the participant analysis component and the participant evaluation results to construct a knowledge model corresponding to the current identity of each participant; Adjusting weights of different types of expertise and role attributes based on the knowledge model, the weight allocation mechanism, the identities of the participants, and the current agenda of the meeting to determine a participant identity mapping table; The knowledge model is updated based on the interactive behavior adaptive modeling mechanism and the recorded behavioral characteristics of the participants, and the behavioral trends are captured and analyzed based on the updated knowledge model and the participant identity mapping table to determine the speaker context of the meeting; the speaker context includes the speaker identity information, the speaker status information and the speaker cognitive load information.
3. The conference enhancement method according to claim 1, characterized in that: The structured representation of the conference content by using the conference content processing component, the speaker context, and the target data packet, and the argumentation analysis based on the structured content, includes: Receiving the speaker context and the target data packet through the conference content processing component; Performing text conversion on the target data packet to determine a conversion result; Performing a timestamp and a speaker identity tag on the conversion result based on the speaker context to determine a tagged text; Based on a content structured analyzer and a hierarchical attention network architecture, topic identification, semantic segmentation, and relationship extraction are performed on the labeled text to determine structured content; The structured content is subjected to a structural argumentation analysis based on the speaker context, and the structured content is subjected to an information source analysis and an information reliability analysis based on preset analysis rules to determine an argumentation analysis result.
4. The conference enhancement method according to claim 1, characterized in that: The determining whether to currently perform a meeting intervention through the intervention output component, the parsing result, the structured content, the current meeting status information, and the historical intervention feedback information, and determining and outputting a target intervention suggestion when the intervention is currently performed, includes: Performing heterogeneous data integration on the received parsing results, the structured content, the speaker context, and the argument analysis results through the intervention output component, and prioritizing the data based on the integration results to determine a ranking result; Determining whether to currently perform a meeting intervention based on the auxiliary generator in the intervention output component, the sorting result, the current meeting state information, and the historical intervention feedback information to obtain an intervention determination result; When the intervention judgment result is yes, determining a target intervention suggestion and a corresponding intervention type based on the current meeting state information; An intervention output mode corresponding to the target intervention suggestion is determined based on the intervention type, and an intervention output operation of the target intervention suggestion is triggered according to the intervention output mode; the target intervention suggestion includes intervention timing, intervention content, and intervention form.
5. The conference enhancement method according to claim 1, characterized in that: Also includes: After outputting the target intervention suggestion, collecting intervention feedback information corresponding to the target intervention suggestion through the intervention output component; Performing an effectiveness evaluation of the intervention based on the intervention feedback information, and sending the intervention evaluation result to the decision analysis component so that the decision analysis component triggers an analysis optimization operation based on the intervention evaluation result; After the meeting is over, the multimodal original information and the analysis results during the meeting are integrated through the intervention output component and based on the current target model, so as to determine and distribute post-meeting resources based on the integration results; the post-meeting resources include meeting summaries, decision records and learning resource recommendation information.
6. A conference enhancement device, characterized in that: Applied to a preset conference enhancement system, the preset conference enhancement system includes a data preprocessing component, a participant analysis component, a conference content processing component, a decision analysis component, and an intervention output component; the device includes: A data preprocessing module, configured to perform corresponding data preprocessing on the original information of each modality in the multimodal original information of the conference through the data preprocessing component, and perform multimodal data fusion and time alignment processing based on the preprocessing results to determine the target data packet; wherein the multimodal original information includes audio data, video data, and environmental data; a participant analysis module for identifying the identity of a participant using the participant analysis component and the target data packet, and performing participant status analysis based on the identification results using a speech-gesture-expression collaborative analysis framework to determine a target model for analyzing the behavior and status of the participant, and determining the speaker context of the conference based on the target model; the speaker context includes speaker identity information, speaker status information, and speaker cognitive load information; a content processing module, configured to perform a structured representation of the conference content using the conference content processing component, the speaker context, and the target data packet, and perform argument analysis based on the structured content to determine an argument analysis result; A meeting analysis module is used to analyze the team interaction mode and decision-making process of the meeting from a holistic perspective through the decision analysis component, the target model, the structured content, and the argumentation analysis results to determine an analysis result; A conference intervention module, configured to determine whether to conduct a conference intervention at present through the intervention output component, the parsing result, the structured content, the current conference status information, and the historical intervention feedback information, and to determine and output a target intervention suggestion when the intervention is conducted; The data preprocessing module is used to: receive the original signal sent by each hardware device through the data preprocessing component; the hardware device includes a microphone, a camera and an environmental sensor; digitize and standardize the original signal to determine the processed audio data, video data and environmental data; perform noise filtering on the audio data based on the context-aware adaptive filtering algorithm to determine the first preprocessed data; the implementation process of the context-aware adaptive filtering algorithm includes: firstly decomposing the signal in the audio data into a time-frequency representation by using time-frequency domain analysis, and then applying a neural network structure with long-term and short-term memory capabilities to model the noise time series evolution characteristics, and then the network receives the current The frame acoustic feature vector and historical state are used as input to calculate the noise probability distribution and capture the temporal dependency. The generated noise probability is used to enhance the original audio signal through a soft masking mechanism. The video data is used to identify the region of interest based on the attention region priority processing mechanism. The dynamic importance weight of each region in the region identification result is determined according to the conference dynamic information and historical interaction pattern. The importance evaluation is performed based on the current speech state importance score, the historical interaction pattern influencing factor and the content complexity of the interest region. Resources are allocated according to the calculated importance map to determine the second pre-processed data. The comprehensive result of the importance evaluation includes the current speech state Importance score, historical interaction mode influencing factor and content complexity of the area of interest are integrated; based on the standardized environmental parameter vector and the multi-layer perceptron structure, the environmental data is evaluated for environmental status to determine the third pre-processed data; the standardized environmental parameter vector includes a standardized humidity vector and a standardized temperature vector; after adjusting the data acquisition configuration and resource allocation based on the hierarchical data acquisition protocol, the resource allocation algorithm and the currently determined first pre-processed data, the second pre-processed data and the third pre-processed data, the first pre-processed data, the second pre-processed data and the third pre-processed data are fused and aligned by a timestamp alignment machine The fused data is time-aligned to determine the target data packet; the hierarchical data acquisition protocol includes dynamically adjusting the data acquisition configuration according to the type and importance of the meeting; the resource allocation algorithm includes adopting a reinforcement learning framework, treating resource allocation as a sequential decision problem, and determining and executing the optimal resource allocation strategy in the current meeting scenario by continuously observing the current state; the target data packet includes a clear audio stream with speaker position markers, a visual tracking data packet with facial ROI and expression markers, and an environmental status report with a meeting context data packet; the meeting parsing module is used to: receive the target model, the structured content, and the argumentation analysis results through the decision analysis component;Based on the interactive network analyzer in the decision analysis component and the structured content, the conference interaction is represented as a dynamically evolving graph structure, with the participants as nodes and the interactions as edges, to capture various interactive behaviors. A time-series graph neural network architecture is used to construct a dynamic interactive network at each time point. The node representation and graph-level representation are extracted through the graph neural network layer, and the time dimension dependency is captured through the recurrent neural network layer to analyze the influence propagation path, information bottleneck points and key connector roles in the dynamic interactive network to determine the first analysis result. Based on the decision process tracker in the decision analysis component, the decision points and solution branches of the structured content are identified, and the decision points and solution branches are identified based on the decision process tracker in the decision analysis component. Extract the basis and assumptions supporting each solution based on the identification results, construct a dynamically growing decision tree, express decision dependencies based on the decision tree and Bayesian network model, evaluate the impact of changes in preconditions on decision consequences, complete the analysis of the decision-making process, and determine a second analysis result; perform a multi-dimensional quantitative evaluation based on the evaluator in the decision analysis component, the meeting objectives of the meeting, the argumentation analysis results, the first analysis results, and the second analysis results to determine a corresponding quality report; perform hidden obstacle detection based on the dynamic interactive network, and determine whether to trigger a cognitive gap marking operation based on the obstacle monitoring results and the target model; The device is further configured to: after outputting the target intervention suggestion, collect intervention feedback information, perform intervention effectiveness evaluation based on the intervention feedback information, and send the intervention evaluation result to the decision analysis component to trigger analysis and optimization operations of the target model.
7. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the conference enhancement method according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that Used to store a computer program, which, when executed by a processor, implements the conference enhancement method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Multi-modal conference data structuring method and device and computer equipment
CN114298170A
Multifunctional video conference interaction method and device, equipment and storage medium
CN118612378A