Conference enhancement method and device, equipment and storage medium

Through a hierarchical and multi-module collaboration preset conference enhancement system, the problem of low meeting efficiency and results quality in the existing technology is solved, and the effectiveness of multi-modal conference information processing, participant status analysis and decision-making process tracking is achieved, and the overall efficiency and results quality of the conference are improved.

CN120075203AActive Publication Date: 2025-05-30SHANDONG INSPUR SCI RES INST CO LTD

Patent Information

Application Number
CN202510536515.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-05-30
Estimated Expiration
2045-04-27

AI Technical Summary

Technical Problem

The existing conference enhancement scheme has shortcomings in multimodal conference information processing, participant status analysis and decision-making process tracking, resulting in poor conference efficiency and quality of results.

Method used

A hierarchical and multi-module collaboration preset conference enhancement system is adopted, including data preprocessing components, participant analysis components, conference content processing components, decision analysis components and intervention output components. Real-time intervention suggestions are generated through multimodal data fusion, participant state analysis, structured content representation and decision-making process analysis.

Benefits of technology

It effectively improves the efficiency and effectiveness and quality of the meeting, and solves the shortcomings of multimodal conference information processing, participant status analysis and decision-making process tracking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120075203A_ABST
    Figure CN120075203A_ABST
Patent Text Reader

Abstract

The invention discloses a conference enhancement method and device, equipment and a storage medium, and relates to the technical field of computers, and the method comprises the steps: preprocessing the multi-modal original information of a conference through a data preprocessing component, carrying out the multi-modal data fusion and time alignment processing based on a preprocessing result, and determining a target data packet; analyzing the state of a participant through a participant analysis component and the target data packet, determining a target model for analyzing the behavior and state of the participant, and determining the context of a speaking party of the conference based on the target model; structural representation of the conference content is carried out through the conference content processing component and the speaking party context; analyzing a team interaction mode and a decision forming process of the conference by adopting an overall perspective through a decision analysis component, a target model, the structured content and a corresponding argument analysis result; and determining whether conference intervention is performed or not according to the analysis result through an intervention output component, and if so, determining and outputting a target intervention suggestion. The effectiveness and the quality of the conference result are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and particularly to a meeting enhancement method, device, equipment and storage medium. Background Art

[0002] With the increase in organizational complexity and the uncertainty of the decision-making environment, meetings, as the core scenarios for team collaboration and decision-making, have a significant impact on enterprise operations in terms of efficiency and quality. Currently, middle and senior managers spend a large amount of working time in meetings on average, and most of these meetings are considered inefficient or have limited results by the participants. This inefficiency not only causes waste of time and resources but may also lead to delays or a decline in the quality of key decisions.

[0003] To this end, existing meeting enhancement-related solutions solve problems through voice transcription and recording systems, and basic collaboration platforms for providing functions such as document sharing and screen sharing. However, these solutions often only support the processing of a single type of data, lack the ability to identify implicit communication barriers due to the lack of comprehensive analysis of the states of participants such as emotions, and the intervention prompts often lead to untimely or inaccurate interventions due to the use of fixed strategies for assistance. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to provide a meeting enhancement method, device, equipment and storage medium, which can effectively solve the deficiencies of existing related solutions in aspects such as multi-modal meeting information processing, participant state analysis, and decision-making process tracking, and improve the efficiency of meetings and the effectiveness and quality of meeting results. The specific solutions are as follows:

[0005] In a first aspect, the present application provides a meeting enhancement method, which is applied to a preset meeting enhancement system. The preset meeting enhancement system includes a data preprocessing component, a participant analysis component, a meeting content processing component, a decision analysis component, and an intervention output component; the method includes:

[0006] Perform corresponding data preprocessing on the original information of each modality in the multi-modal original information of the meeting through the data preprocessing component, and perform multi-modal data fusion and time alignment processing based on the preprocessing results to determine a target data packet;

[0007] Perform participant identity recognition through the participant analysis component and the target data packet, and perform participant state analysis based on the recognition result to determine a target model for analyzing the behavior and state of participants, and determine the speaker context of the meeting based on the target model;

[0008] Perform a structured representation of the meeting content through the meeting content processing component, the speaker context, and the target data packet, and conduct an argument analysis based on the structured content to determine the argument analysis result;

[0009] Through the decision analysis component, the target model, the structured content, and the argument analysis result, and adopt an overall perspective to respectively analyze the team interaction mode and the decision-making process of the meeting to determine the analysis result;

[0010] Determine whether to conduct a meeting intervention currently through the intervention output component, the analysis result, the structured content, the current meeting status information, and the historical intervention feedback information, and when it is, determine and output the target intervention suggestion.

[0011] Optionally, the data preprocessing component performs corresponding data preprocessing on the original information of each modality in the multi-modal original information of the meeting, and performs multi-modal data fusion and time alignment processing based on the preprocessing result, including:

[0012] The data preprocessing component receives the original signals sent by each hardware device; the hardware devices include microphones, cameras, and environmental sensors;

[0013] Perform digitization and standardization processing on the original signals to determine the processed audio data, video data, and environmental data;

[0014] Perform noise filtering on the audio data based on a first preset algorithm to determine the first preprocessed data;

[0015] Perform region of interest recognition on the video data based on a second preset algorithm, and perform resource allocation based on the region recognition result to determine the second preprocessed data;

[0016] Perform environmental state assessment on the environmental data based on the standardized environmental parameter vector and the multi-layer perceptron structure to determine the third preprocessed data; the standardized environmental parameter vector includes a standardized humidity vector and a standardized temperature vector;

[0017] After adjusting the data acquisition configuration and resource allocation based on the hierarchical data acquisition protocol, the resource allocation algorithm, and the currently determined first preprocessed data, second preprocessed data, and third preprocessed data, perform data fusion on the first preprocessed data, second preprocessed data, and third preprocessed data, and perform time alignment on the fused data through the timestamp alignment mechanism to determine the target data packet.

[0018] Performing participant identity recognition through the participant analysis component and the target data packet, and performing participant status analysis based on the recognition result to determine a target model for analyzing participant behavior and status, and determining the speaker context of the meeting based on the target model, including:

[0019] Simultaneously analyzing the acoustic features and visual features in the target data packet through the participant analysis component and the progressive recognition strategy to determine the participant identity and the corresponding spatial position tracking result;

[0020] Based on the participant status analyzer in the participant analysis component, the participant identity, and the spatial position tracking result, and using the speech-gesture-expression collaborative analysis framework to evaluate the current emotion and current cognitive state of the participant to determine the participant evaluation result;

[0021] Analyzing the historical speech content and interaction patterns of the participant based on the role and expertise modeler in the participant analysis component and the participant evaluation result to construct a knowledge model corresponding to each current participant identity;

[0022] Adjusting the weights of different types of professional knowledge and role attributes based on the knowledge model, the weight assignment mechanism, the participant identity, and the current topic of the meeting to determine the participant identity mapping table;

[0023] Updating the knowledge model based on the interactive behavior adaptive modeling mechanism and the recorded participant behavior characteristics, and capturing and analyzing the behavior trend according to the updated knowledge model and the participant identity mapping table to determine the speaker context of the meeting; the speaker context includes speaker identity information, speaker status information, and speaker cognitive load information.

[0024] Optionally, performing structured representation of the meeting content through the meeting content processing component, the speaker context, and the target data packet, and performing argumentation analysis based on the structured content, including:

[0025] Receiving the speaker context and the target data packet through the meeting content processing component;

[0026] Converting the target data packet into text to determine the conversion result;

[0027] Performing timestamp marking and speaker identity marking on the conversion result based on the speaker context to determine the marked text;

[0028] Performing topic recognition, semantic segmentation, and relationship extraction on the marked text respectively based on the content structured analyzer and the hierarchical attention network architecture to determine the structured content;

[0029] Perform structural argument analysis on the structured content based on the speaker context, and perform information source analysis and information reliability analysis on the structured content based on preset analysis rules to determine the argument analysis result.

[0030] Optionally, through the decision analysis component, the target model, the structured content, and the argument analysis result, and adopt an overall perspective to respectively analyze the team interaction mode and the decision-making formation process of the meeting, including:

[0031] Receive the target model, the structured content, and the argument analysis result through the decision analysis component;

[0032] Based on the interaction network analyzer in the decision analysis component and the structured content, construct and analyze a sequence diagram, and analyze the influence propagation path, information bottleneck points, and key connector roles in the determined dynamic interaction network to determine the first analysis result;

[0033] Based on the decision process tracker in the decision analysis component, identify decision points for the structured content, and perform evolutionary tracking of decision points according to the decision point identification results, so as to analyze the decision-making formation process based on the tracking results to determine the second analysis result;

[0034] Based on the evaluator in the decision analysis component, the meeting objective of the meeting, the argument analysis result, the first analysis result, and the second analysis result, conduct multi-dimensional quantitative evaluation to determine the corresponding quality report;

[0035] Detect hidden obstacles based on the dynamic interaction network, and determine whether to trigger the cognitive gap marking operation according to the obstacle monitoring result and the target model.

[0036] Optionally, determine whether to perform meeting intervention currently through the intervention output component, the parsing result, the structured content, the current meeting status information, and the historical intervention feedback information, and when it is, determine and output the target intervention suggestion, including:

[0037] Integrate heterogeneous data for the received parsing result, structured content, speaker context, and argument analysis result through the intervention output component, and perform priority sorting of the data based on the integration result to determine the sorting result;

[0038] Based on the auxiliary generator in the intervention output component, the sorting result, the current meeting status information, and the historical intervention feedback information, judge whether to perform meeting intervention currently to obtain the intervention judgment result;

[0039] When the intervention judgment result is yes, determine the target intervention suggestion and the corresponding intervention type based on the current meeting status information;

[0040] Determine the intervention output mode corresponding to the target intervention suggestion based on the intervention type, and trigger the intervention output operation of the target intervention suggestion according to the intervention output mode; the target intervention suggestion includes the intervention timing, intervention content, and intervention form.

[0041] Optionally, the method further includes:

[0042] After outputting the target intervention suggestion, collect the intervention feedback information corresponding to the target intervention suggestion through the intervention output component;

[0043] Conduct an effectiveness evaluation of the intervention based on the intervention feedback information, and send the intervention evaluation result to the decision analysis component, so that the decision analysis component triggers an analysis and optimization operation based on the intervention evaluation result;

[0044] After the meeting ends, integrate the multimodal raw information and the parsing result during the meeting through the intervention output component and based on the current target model, so as to determine and distribute the post-meeting resources based on the integration result; the post-meeting resources include meeting summaries, decision records, and learning resource recommendation information.

[0045] In a second aspect, the present application provides a meeting enhancement device, which is applied to a preset meeting enhancement system. The preset meeting enhancement system includes a data preprocessing component, a participant analysis component, a meeting content processing component, a decision analysis component, and an intervention output component; the device includes:

[0046] A data preprocessing module, configured to perform corresponding data preprocessing on the raw information of each modality in the multimodal raw information of the meeting through the data preprocessing component, and perform multimodal data fusion and time alignment processing based on the preprocessing result to determine a target data packet;

[0047] A participant analysis module, configured to identify the participant identity through the participant analysis component and the target data packet, and perform participant status analysis based on the identification result to determine a target model for analyzing the behavior and status of the participant, and determine the speaker context of the meeting based on the target model;

[0048] A content processing module, configured to perform a structured representation of the meeting content through the meeting content processing component, the speaker context, and the target data packet, and perform argument analysis based on the structured content to determine an argument analysis result;

[0049] A meeting analysis module, configured to respectively analyze the team interaction pattern and the decision-making formation process of the meeting from an overall perspective by using the decision analysis component, the target model, the structured content, and the argument analysis result, so as to determine an analysis result;

[0050] A meeting intervention module, configured to determine whether to perform a meeting intervention currently by using the intervention output component, the analysis result, the structured content, the current meeting status information, and the historical intervention feedback information, and when the determination result is yes, determine and output a target intervention suggestion.

[0051] In a third aspect, the present application provides an electronic device, including:

[0052] A memory, configured to store a computer program;

[0053] A processor, configured to execute the computer program to implement the steps of the foregoing meeting enhancement method.

[0054] In a fourth aspect, the present application provides a computer-readable storage medium, configured to store a computer program, where when the computer program is executed by a processor, the steps of the foregoing meeting enhancement method are implemented.

[0055] It can be seen that in this application, it is applied to a preset conference enhancement system, which includes a data preprocessing component, a participant analysis component, a conference content processing component, a decision analysis component, and an intervention output component; the method includes: performing corresponding data preprocessing on the original information of each modality in the multimodal original information of the conference through the data preprocessing component, and performing multimodal data fusion and time alignment processing based on the preprocessing results to determine a target data packet; performing participant identity recognition through the participant analysis component and the target data packet, and performing participant status analysis based on the recognition results to determine a target model for analyzing participant behavior and status, and determining the speaker context of the conference based on the target model; performing a structured representation of the conference content through the conference content processing component, the speaker context, and the target data packet, and performing argument analysis based on the structured content to determine an argument analysis result; parsing the team interaction mode and the decision-making formation process of the conference from an overall perspective through the decision analysis component, the target model, the structured content, and the argument analysis result to determine an analysis result; determining whether to perform a conference intervention currently through the intervention output component, the analysis result, the structured content, the current conference status information, and the historical intervention feedback information, and when it is necessary, determining and outputting a target intervention suggestion. That is, in this application, conference enhancement is achieved through a preset conference enhancement system with hierarchical and multi-module collaboration. Specifically, first, the data preprocessing component in the system is used to preprocess the multimodal original information of the conference, and perform multimodal data fusion and time alignment processing to determine a target data packet. Then, the participant analysis component in the system and the target data packet are used to identify the participant identity and perform participant status analysis, so as to determine the speaker context of the conference by using the obtained target model for analyzing participant behavior and status. Then, the conference content processing component in the system and the speaker context are used to determine the structured content of the conference and perform argument analysis. Then, the decision analysis component, the structured content, and the argument analysis result in the system are used to parse the team interaction mode and the decision-making formation process of the conference from an overall perspective. Then, the intervention output component in the system and the analysis result, the current conference status information, and the historical intervention feedback information are used to determine whether to intervene in the conference, and if so, determine the target intervention suggestion. In this way, it can effectively solve the deficiencies of existing related solutions in aspects such as multimodal conference information processing, participant status analysis, and decision-making process tracking, and improve the efficiency of the conference and the effectiveness and quality of the conference results. Description of the Drawings

[0056] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the provided drawings.

[0057] Figure 1 Flowchart of a conference enhancement method provided by this application;

[0058] Figure 2 Schematic diagram of the interaction process of each component in a preset conference enhancement system provided by this application;

[0059] Figure 3 Schematic diagram of the data processing flow of a data preprocessing component provided by this application;

[0060] Figure 4 Schematic diagram of the data processing flow of a participant analysis component provided by this application;

[0061] Figure 5 Schematic diagram of the data processing flow of a conference content processing component provided by this application;

[0062] Figure 6 Schematic diagram of the data processing flow of a decision analysis component provided by this application;

[0063] Figure 7 Schematic diagram of the data processing flow of an intervention output component provided by this application;

[0064] Figure 8 Schematic diagram of the structure of a conference enhancement device provided by this application;

[0065] Figure 9 Schematic diagram of the structure of an electronic device provided by this application. Detailed implementation manners

[0066] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0067] In existing conference enhancement-related solutions, problems are solved through a voice transcription and recording system and a basic collaboration platform for providing functions such as document sharing and screen sharing. However, these solutions often only support the processing of a single type of data. Due to the lack of comprehensive analysis of the states of participants such as emotions, the ability to identify implicit communication barriers is lacking. And because fixed strategies are often used for assisted intervention prompts, it often leads to untimely or inaccurate intervention. For this reason, this application provides a conference enhancement solution, which can effectively solve the deficiencies of existing related solutions in aspects such as multi-modal conference information processing, analysis of the states of participants, and tracking of the decision-making process, and improve the efficiency of the conference and the effectiveness and quality of the conference results.

[0068] See Figure 1 As shown, an embodiment of the present invention discloses a conference enhancement method, which is applied to a preset conference enhancement system. The preset conference enhancement system includes a data preprocessing component, a participant analysis component, a conference content processing component, a decision analysis component, and an intervention output component; the method includes:

[0069] Step S11, perform corresponding data preprocessing on the original information of each modality in the multi-modal original information of the conference through the data preprocessing component, and perform multi-modal data fusion and time alignment processing based on the preprocessing results to determine the target data packet.

[0070] It should be understood that, combined with Figure 2 As shown in the system architecture schematic diagram of the preset conference enhancement system, the system adopts a hierarchical module design and includes five core functional modules that cooperate with each other: a data collection and preprocessing module (i.e., the data preprocessing component), a participant analysis module (i.e., the participant analysis component), a content understanding and structuring module (i.e., the conference content processing component), a group dynamics and decision analysis module (i.e., the decision analysis component), and an intelligent assistance and output module (i.e., the intervention output component). These modules form a complete intelligent conference management link through a carefully designed interaction mechanism.

[0071] Specifically, in the data processing flow of the system, first, the data preprocessing component receives multi-modal raw information such as audio, video, and environmental data in the meeting, and preprocesses the raw information of each modality respectively. That is, first, the data preprocessing component receives the raw signals sent by each hardware device; the hardware devices include microphones, cameras, and environmental sensors; the raw signals are digitized and standardized to determine the processed audio data, video data, and environmental data; the audio data is filtered for noise based on the first preset algorithm to determine the first preprocessed data; the region of interest of the video data is identified based on the second preset algorithm, and resource allocation is performed based on the region recognition result to determine the second preprocessed data; the environmental state of the environmental data is evaluated based on the standardized environmental parameter vector and the multi-layer perceptron structure to determine the third preprocessed data; the standardized environmental parameter vector includes a standardized humidity vector and a standardized temperature vector; after adjusting the data acquisition configuration and resource allocation based on the hierarchical data acquisition protocol, the resource allocation algorithm, and the currently determined first preprocessed data, second preprocessed data, and third preprocessed data, the first preprocessed data, second preprocessed data, and third preprocessed data are fused, and the fused data is time-aligned through the timestamp alignment mechanism to determine the target data packet. It should be understood that the data preprocessing component adopts a multi-channel parallel processing architecture, processes multi-modal raw information such as audio, video, and environmental data simultaneously, and realizes preliminary cross-modal synchronization and alignment.

[0072] Regarding the processing flow of audio data, the audio processing adopts context-aware adaptive filtering technology and is realized through the following steps: first, the audio signal is decomposed into a time-frequency representation by time-frequency domain analysis, then a neural network structure with long short-term memory ability is applied to model the temporal evolution characteristics of the noise, and then the network receives the current frame acoustic feature vector and the historical state as inputs, and calculates the noise probability distribution through the long short-term memory network unit. This network is responsible for capturing temporal dependencies. Preferably, using the generated noise probability, the system enhances the raw signal through a soft mask mechanism. This process not only considers the current acoustic features but also incorporates historical information and meeting context knowledge, and can effectively suppress various noise interferences while preserving the integrity of the speech.

[0073] Regarding the processing flow of video data, visual processing adopts a mechanism of prioritizing the processing of regions of interest: first, a low-complexity global scan is performed to identify potential regions of interest through a lightweight feature detector, and then, based on the meeting dynamic information and historical interaction patterns, dynamic importance weights are assigned to each region. After that, importance evaluation synthesis is carried out, considering three evaluation factors: the importance score based on the current speaking state, the influence factor of historical interaction patterns, and the complexity of the region content. Finally, according to the calculated importance map, the system preferentially allocates computing resources to high-weight regions. This differential processing strategy enables the system to capture key visual information in the meeting while meeting the real-time requirements.

[0074] Regarding the analysis flow of environmental data, environmental processing establishes a correlation model between environmental parameters and cognitive performance: first, a standardized environmental parameter vector (including temperature, humidity, concentration, etc.) is received, and then a non-linear transformation is performed through a multi-layer perceptron structure to map the environmental parameters to an influence score on cognitive performance. Finally, based on the output score, the system generates adjustment suggestions when the environmental parameters deviate from the optimal range. The obtained correlation model is based on a large amount of empirical research data and establishes a mapping relationship between environmental factors and meeting efficiency indicators.

[0075] After that, the target data packet is sent to the participant analysis component, enabling the latter to perform participant modeling based on high-quality data.

[0076] Step S12: Identify the participant identity through the participant analysis component and the target data packet, and perform participant status analysis based on the identification result to determine the target model for analyzing the behavior and status of the participant, and determine the speaker context of the meeting based on the target model.

[0077] In this embodiment, the participant analysis component will use the data transmitted by the data preprocessing component for identity recognition and content modeling, converting the explicit behaviors and implicit states of the participants into a machine-understandable structured representation, providing important context support for subsequent meeting content analysis. That is, first, the acoustic features and visual features in the target data packet are analyzed simultaneously through the participant analysis component and the progressive recognition strategy to determine the participant identity and the corresponding spatial position tracking results; based on the participant state analyzer, participant identity, and spatial position tracking results in the participant analysis component, and using the speech-gesture-expression collaborative analysis framework to evaluate the current emotion and current cognitive state of the participant to determine the participant evaluation results; based on the role and expertise modeler and the participant evaluation results in the participant analysis component, analyze the historical speech content and interaction patterns of the participants to construct a knowledge model corresponding to the current identities of each participant; based on the knowledge model, weight assignment mechanism, participant identity, and the current topic of the meeting, adjust the weights of different types of professional knowledge and role attributes to determine the participant identity mapping table; update the knowledge model based on the interactive behavior adaptive modeling mechanism and the recorded participant behavior characteristics, and capture and analyze the behavior trends according to the updated knowledge model and the participant identity mapping table to determine the speaker context of the meeting; the speaker context includes speaker identity information, speaker status information, and speaker cognitive load information. It can be understood that Figure 2 the participant model in it is the target model, and the state data represents the state data related to the participant.

[0078] Regarding the identity recognition process, the multi-feature fusion identity recognition algorithm is implemented through the following steps:

[0079] 1) Construct a two-stream neural network architecture. The first stream processes short-term acoustic features (such as fundamental frequency distribution, harmonic structure, formant features, etc.), and the second stream analyzes long-term speaking style features (such as vocabulary usage habits, syntactic structure preferences, expression patterns, etc.);

[0080] 2) Each of the two streams generates an identity embedding vector;

[0081] 3) Perform weighted fusion through the attention mechanism to form the final identity representation;

[0082] 4) Adopt a progressive recognition strategy. At the initial stage of the meeting, quickly generate a possible identity candidate set and assign an initial probability distribution. As more interaction data is obtained, continuously update the distribution until high-confidence recognition is achieved.

[0083] Regarding the status analysis process, the participant status analyzer adopts a voice-gesture-expression collaborative analysis framework: First, it extracts raw features from each modality. The voice features include pitch changes, speech rate, and prosody features. The visual features include facial expression units, eye fixation points, and head postures. Then, it models the complementary and consistency relationships between modalities through a graph neural network structure, represents different modality features as nodes in the graph, represents the associations between modalities as edges, and realizes information fusion through the message passing mechanism between nodes to handle the situation of incomplete modalities. When the participant is temporarily out of the camera's view, it can still give a reliable status assessment based on voice and historical status.

[0084] Regarding the knowledge modeling process, the role and professional knowledge modeler adopts an automatic generation method of a professional knowledge graph based on the content of the speech: First, it extracts professional terms and key concepts from the speech text, then constructs an initial knowledge graph based on co-occurrence analysis and semantic associations, compares the expression patterns of the participants and domain experts, evaluates their professional depth and influence at each knowledge node, forms a dynamically updated personal knowledge model, provides a reference for content understanding and intervention strategies, and realizes a dynamic weight allocation mechanism to automatically adjust the importance of different professional knowledge and role attributes according to the current discussion topic.

[0085] Regarding the interactive behavior modeling, an interactive behavior adaptive modeling mechanism continuously learns and updates the interactive patterns of the participants: It tracks the behavioral characteristics of the participants such as speech frequency, response pattern, and influence performance, then continuously updates the model as the meeting progresses, and captures the evolution trend of the participants' behaviors (such as from conservative observation in the initial stage to active participation in the later stage), and analyzes the role changes under different topics (such as from being a professional leader in one topic to being a side listener in another topic).

[0086] Step S13: Perform a structured representation of the meeting content through the meeting content processing component, the speaker context, and the target data packet, and conduct an argumentation analysis based on the structured content to determine the argumentation analysis result.

[0087] Specifically, in this embodiment, at the content processing layer, the meeting content processing component performs in-depth semantic analysis and structured representation on the meeting content based on the information transmitted by the participant analysis component. Through in-depth semantic analysis and context reasoning, it captures the literal information, implicit intentions, logical relationships, and knowledge elements of the meeting content, providing a basis for subsequent decision-making analysis. That is, the meeting content processing component receives the speaker's context and the target data packet; performs text conversion on the target data packet to determine the conversion result; marks the conversion result with a timestamp and the speaker's identity based on the speaker's context to determine the marked text; performs topic recognition, semantic segmentation, and relationship extraction on the marked text respectively based on the content structure analyzer and the hierarchical attention network architecture to determine the structured content; performs structural argumentation analysis on the structured content based on the speaker's context, and performs information source analysis and information reliability analysis on the structured content based on preset analysis rules to determine the argumentation analysis result. Then, the structured meeting content, the argumentation analysis result, and the target model are jointly transmitted to the decision-making analysis component to support higher-level interaction and decision-making analysis.

[0088] Regarding the speech understanding process, the context-aware professional speech recognition system is implemented in the following way: First, a domain-adapted basic language model is constructed using transfer learning methods, and then new terms are detected and learned in real time during the meeting: Identify low-confidence regions; determine whether this region contains important information (based on context and the speaker's professional background), and learn this expression from the subsequent context to continuously improve domain adaptability as the meeting progresses.

[0089] Regarding the content structuring process, a dynamic adaptation mechanism based on the meeting-specific language model: Adopt a hierarchical attention network architecture to capture semantic units at different granularities through multi-level attention calculations: Word level: Identify key terms and sentiment words; Sentence level: Identify key statements and speech acts; Paragraph level: Establish a topic hierarchy. Integrate the speaker's role, emotion markers, and interaction context to achieve comprehensive semantic understanding. Adopt a progressive structuring processing strategy. First, quickly identify the overall framework and main topics of the meeting, and gradually refine the content structure as the discussion deepens.

[0090] Regarding the argumentation and quality assessment process, based on the identification of implicit premises in the meeting context, first construct a complete argumentation graph, including explicitly expressed premises, conclusions, and support relationships. Then, detect the breakpoints in the logical chain, compare with the domain knowledge graph, identify possible bridging concepts, and finally mark potential implicit premises, and prompt the participants to confirm or clarify if necessary. When evaluating information quality and reliability, first assign credibility scores according to indicators such as information sources, supporting evidence, and internal consistency. Then, verify facts through the enterprise knowledge base and trusted data sources. Then, evaluate the value of the meeting content from three dimensions: relevance, novelty, and reliability. Finally, allocate attention resources according to the comprehensive value score to highlight high-value information.

[0091] Step S14: Through the decision analysis component, the target model, the structured content, and the argument analysis result, and adopting an overall perspective, respectively analyze the team interaction pattern and the decision-making formation process of the meeting to determine the analysis result.

[0092] Specifically, in this embodiment, the decision analysis component of the analysis decision layer is responsible for analyzing the team interaction pattern and the decision-making formation process. By focusing on the collaborative dynamics, influence mechanisms, and decision-making paths at the group level, it reveals the implicit social structure and cognitive process in the meeting, providing a deep theoretical basis for meeting optimization. That is, first, the decision analysis component receives the target model, structured content, and argument analysis result; based on the interaction network analyzer and the structured content in the decision analysis component, construct and analyze the time sequence diagram, and analyze the influence propagation path, information bottleneck points, and key connector roles in the determined dynamic interaction network to determine the first analysis result; based on the decision process tracker in the decision analysis component, identify the decision points in the structured content, and perform evolutionary tracking of the decision points according to the decision point identification result, and analyze the decision-making formation process based on the tracking result to determine the second analysis result; based on the evaluator in the decision analysis component, the meeting objective of the meeting, the argument analysis result, the first analysis result, and the second analysis result, conduct multi-dimensional quantitative evaluation to determine the corresponding quality report; perform implicit obstacle detection based on the dynamic interaction network, and determine whether to trigger the cognitive gap marking operation according to the obstacle monitoring result and the target model.

[0093] Regarding the interactive network analysis process, dynamic modeling and prediction of time-varying interactive graphs:

[0094] 1) Represent the meeting interaction as a dynamically evolving graph structure, with participants as nodes and interactions as edges, capturing various interaction behaviors (direct dialogue, response, citation, non-verbal feedback, etc.);

[0095] 2) Adopt a time sequence graph neural network architecture to construct an interactive graph at each time point, extract node representations and graph-level representations through graph neural network layers, and capture time dimension dependencies through recurrent neural network layers;

[0096] 3) Identify key patterns, including: influence propagation paths, information bottleneck points, and the roles of key connectors.

[0097] Regarding the decision analysis process, real-time construction and evaluation of a multi-level decision tree: First, identify decision points and alternative branches, then extract the basis and assumptions supporting each alternative, then construct a dynamically growing decision tree to reflect the development of the decision logic under discussion, and evaluate the quality of the decision-making process, including: sufficiency of evidence, degree of hypothesis verification, and breadth of exploration of alternative solutions. Use a Bayesian network model to represent decision dependencies and evaluate the impact of changes in preconditions on decision consequences.

[0098] Regarding the efficiency evaluation process, for the quantitative evaluation of the meeting value generation process: First, consider multiple dimensions: time utilization rate (the ratio of effective discussion time to total meeting time), participation balance (uniformity of participation opportunities), information quality (the ratio of new information generation to redundant information), and decision progress (the rate of decision point resolution). Then, integrate multi-dimensional indicators through an analytic hierarchy process and dynamically adjust the weights of each indicator according to the meeting type and objectives. Finally, compare the current meeting with the historical performance benchmark to identify efficiency trends.

[0099] Regarding dynamic importance perception, for automatically identifying critical moments and turning points in a meeting: First, analyze multiple signals (changes in interaction patterns, sentiment markers, sudden changes in keyword frequencies), then evaluate the importance of the discussion in real time, then adjust the processing depth and attention resources accordingly, and finally conduct in-depth analysis of key paragraphs and adopt lightweight processing for routine discussions.

[0100] Regarding the identification of hidden obstacles, for discovering hidden factors that hinder effective communication and decision-making: First, detect three common types of hidden obstacles, including: differences in term understanding (the same term is given different meanings), conflicts in implicit assumptions (reasoning based on different unstated premises), and professional knowledge gaps (understanding obstacles caused by the lack of shared background knowledge). Then, compare the term usage patterns, reasoning paths, and knowledge models of the participants, mark potential cognitive gaps, and provide bridging information at appropriate times to facilitate effective communication.

[0101] Step S15: Determine whether to conduct a meeting intervention currently based on the intervention output component, the parsing result, the structured content, the current meeting status information, and the historical intervention feedback information, and when it is necessary, determine and output the target intervention suggestion.

[0102] In this embodiment, the intervention output component will generate real-time intervention suggestions and post-meeting resources based on the results of the foregoing analysis, continuously collect feedback on the intervention effect, and dynamically optimize the entire system. In this way, by converting the previous analysis into actual actions, including real-time assistance during the meeting and output generation after the meeting, it constitutes a key link in the realization of the system value. That is, regarding the intervention, first, the intervention output component integrates heterogeneous data such as the parsed results, structured content, speaker context, and argumentation analysis results received, and performs data priority sorting based on the integration results to determine the sorting result; based on the auxiliary generator, sorting result, current meeting status information, and historical intervention feedback information in the intervention output component, it determines whether to conduct a meeting intervention currently to obtain an intervention judgment result; when the intervention judgment result is yes, it determines the target intervention suggestion and the corresponding intervention type based on the current meeting status information; determines the intervention output method corresponding to the target intervention suggestion based on the intervention type, and triggers the intervention output operation of the target intervention suggestion according to the intervention output method; the target intervention suggestion includes the intervention timing, intervention content, and intervention form.

[0103] Regarding the intervention decision-making process, a precise intervention algorithm based on situation awareness: First, consider the intervention as a continuous decision-making problem, with the goal of maximizing the intervention utility while minimizing the intervention cost; then, based on the reinforcement learning framework, observe the current meeting status (discussion content, participant status, meeting progress, etc.), select intervention actions (whether to intervene, intervention content, intervention form, etc.), obtain immediate benefits (based on the intervention effect evaluation), and transfer to a new state; then the intervention decision considers multiple factors, intervention timing: based on the meeting rhythm, current importance, and participant cognitive load, intervention content: provide missing information, clarify terms, prompt logical loopholes, suggest related knowledge, intervention form: visual cues, text suggestions, post-meeting notes; pay attention to the non-intrusiveness of the intervention, insert prompts at natural pause points; use terms familiar to the participants to express; adjust the prompt salience according to the urgency.

[0104] Regarding the visualization processing process, convert complex data into visual expressions: First, perform data preprocessing: dimensionality reduction, clustering, and extraction of important features, then select a visualization template (decision tree diagram, interactive network diagram, theme evolution diagram) according to the data type and purpose, adjust the visualization complexity according to the user's cognitive load and professional background, and integrate the visualization content into the user interface. Then, optimize the information presentation with cognitive load awareness: First, select a suitable visualization strategy according to the participant's professional background, attention state, and task urgency, perform simplified visual expressions in high-pressure decision-making situations, highlight the core information, and provide detailed multi-level visualizations in the in-depth analysis stage.

[0105] Regarding the meeting resource generation process, based on a multi-level and multi-angle adaptive summarization algorithm: First, construct a hierarchical representation of the meeting content (theme layer, topic layer, opinion layer, detail layer), and then generate customized summaries according to user needs and roles: Provide a concise summary focusing on decision-making points and conclusions for decision-makers; provide detailed action items and background information for implementers; provide a panoramic overview of key discussions and decisions for absentees. Then, extract and organize action items: First, automatically identify commitments, task assignments, and decisions, then extract key elements (responsible person, deadline, dependencies, priority), organize action items in a structured manner, integrate with the enterprise task management system, establish an action item tracking mechanism, and provide progress updates. Also, generate learning resources: Based on the meeting content and the knowledge models of the participants, identify learning needs and knowledge gaps, and generate targeted learning resources (term explanations, background knowledge links, relevant document recommendations).

[0106] Furthermore, in this embodiment, after outputting the target intervention recommendation, the intervention output component collects intervention feedback information corresponding to the target intervention recommendation; evaluates the effectiveness of the intervention based on the intervention feedback information, and sends the intervention evaluation result to the decision analysis component, so that the decision analysis component triggers analysis and optimization operations based on the intervention evaluation result; after the meeting, the intervention output component integrates the multi-modal raw information and parsing results during the meeting based on the current target model to determine and distribute post-meeting resources based on the integration result; the post-meeting resources include meeting summaries, decision records, and learning resource recommendation information. Combining Figure 2 It can be seen that after the intervention output, a closed-loop self-optimization mechanism will be formed among the various layers of the system through the established extensive feedback channels, including intervention effect evaluation, model update signals, and resource allocation instructions, and feedback and optimize layer by layer forward. Among them, the intervention output component feeds back the collected intervention feedback information to the decision analysis component, and then the decision analysis component analyzes the decision-making focus and feeds it back to the meeting content processing component. Then, the meeting content processing component performs semantic association to feedback an optimized processing strategy to the participant analysis component. Then, the participant analysis component conducts focus guidance and feeds it back forward to the data preprocessing component.

[0107] In summary, the preset conference enhancement system in this embodiment simultaneously collects and processes audio, video, and environmental data through a multi-channel parallel processing architecture; uses a multi-feature fusion algorithm for participant identity recognition and status analysis; performs professional speech recognition based on context awareness and multi-level structured analysis to understand the conference content; analyzes group dynamics and decision-making processes through time-varying interaction graphs and multi-level decision trees; and finally realizes precise intervention with context awareness and multi-level adaptive resource generation. The present invention effectively solves the deficiencies of existing conference systems in aspects such as professional term recognition, participant status analysis, implicit communication barrier recognition, and decision-making process tracking, improves the conference efficiency and decision-making quality, and is particularly suitable for high-value conference scenarios involving cross-domain professional knowledge and complex decision-making.

[0108] Thus, in this application, a preset conference enhancement system with hierarchical and multi-module collaboration is used to achieve conference enhancement. Specifically, first, the data preprocessing component in the system preprocesses the multi-modal raw information of the conference, and performs multi-modal data fusion and time alignment processing to determine the target data packet. Then, the participant analysis component in the system and the target data packet are used to identify the participant identity and perform participant status analysis, so as to determine the speaker context of the conference using the obtained target model for analyzing participant behavior and status. After that, the conference content processing component in the system and the speaker context are used to determine the structured content of the conference and perform argumentation analysis. Then, the decision analysis component, the structured content, and the argumentation analysis results in the system are used to respectively analyze the team interaction mode and decision-making formation process of the conference from an overall perspective. Then, the intervention output component in the system, the analysis results, the current conference status information, and the historical intervention feedback information are used to determine whether to intervene in the conference. If an intervention is required, the target intervention suggestion is determined. In this way, the deficiencies of existing related solutions in aspects such as multi-modal conference information processing, participant status analysis, and decision-making process tracking can be effectively solved, and the efficiency of the conference and the effectiveness and quality of the conference results are improved.

[0109] The following combines Figures 2 - 7 with the schematic diagrams disclosed in

[0110] to specifically describe the technical solutions of the embodiments of this application. Figure 2 As can be seen from

[0111] Combined with Figure 3As shown, regarding the data preprocessing component in the system, it is the perception foundation of the entire system, responsible for obtaining raw information from multiple data sources and performing preliminary processing to provide high-quality input for subsequent analysis. In a specific implementation, the component is internally organized into three collaborative sub-units: the audio acquisition and preprocessing unit, the visual data acquisition unit, and the environment and context perception unit. Each sub-unit adopts a dedicated processing strategy for specific data types while maintaining overall coordination. The data flow path within the component is as follows:

[0112] (1) Receive raw signals from various hardware devices (microphone array sound waveforms, camera video frame sequences, environmental sensor numerical readings);

[0113] (2) Perform preliminary digitization and standardization processing;

[0114] (3) Split to dedicated processing channels:

[0115] 1) The audio preprocessing unit filters out environmental noise in the audio data through an adaptive noise suppression algorithm;

[0116] 2) The video data preprocessing unit performs preliminary analysis of the video data through a region priority processing mechanism;

[0117] 3) The environmental preprocessing unit integrates with the preset meeting information to form an environmental status assessment;

[0118] (4) Establish cross-modal associations through a timestamp alignment mechanism to lay the foundation for subsequent fusion analysis.

[0119] It should be understood that before performing modal fusion and time alignment, the data preprocessing component also performs dynamic resource management using a resource allocation and optimization mechanism. The specific management implementation process includes: a hierarchical acquisition protocol: dynamically adjusts the data acquisition configuration according to the meeting type and importance; a resource intelligent allocation algorithm: adopts a reinforcement learning framework and regards resource allocation as a sequential decision-making problem; the system continuously observes the current state through an intelligent agent, selects resource allocation actions, and obtains immediate benefits. In this way, through repeated training, the system can learn the optimal resource allocation strategy in different meeting scenarios to ensure obtaining the most valuable information under limited computing and storage resources. The finally determined data packet includes a clear audio stream with speaker position markers, a visual tracking data packet with facial ROI (Region Of Interest) and expression markers, and an environmental status report with meeting context data packets.

[0120] In addition, regarding the interaction and feedback mechanism of the data preprocessing component, it realizes two-way information flow with other modules:

[0121] 1) Provide processed high-quality data to downstream modules;

[0122] 2) Receive feedback signals to adjust the processing strategy (such as enhancing the sampling rate or processing depth of specific regions);

[0123] 3) Dynamically adjust the processing parameters according to actual needs to maximize the overall system performance.

[0124] Combined with Figure 4 As shown, regarding the participant analysis component in the system, it is the core perception unit of the system, responsible for identifying meeting participants and constructing models of their behaviors and states. This component is in a crucial position in the system architecture, connecting the upper and lower levels, and is responsible for: receiving high-quality audio and video streams from the data acquisition and preprocessing module, performing in-depth feature extraction and multi-dimensional analysis; outputting a complete dynamic model of the participants, including information such as identity, location, emotional state, cognitive load, attention level, and professional knowledge structure; providing a comprehensive human context for the understanding and analysis of meeting content. Moreover, this component consists of three collaborative functional units: an identity recognition and tracker for determining the identity of meeting participants and tracking their physical locations, a participant state analyzer for real-time evaluating the emotional, cognitive, and attention states of participants; a role and professional knowledge modeler for constructing the professional background and meeting role models of participants. These three units work together through information sharing and complementary verification mechanisms to construct a complete and consistent representation of the participants. Regarding the data processing flow of this component:

[0125] 1) Receive the preprocessed audio stream, video features, and preset participant information (and historical participant data including organizational structures and meeting records in the figure);

[0126] 2) The identity recognition and tracker simultaneously analyze acoustic features and visual features to generate a participant identity mapping and spatial location tracking;

[0127] 3) The recognition results are passed to the participant state analyzer, which combines voice features, facial expressions, and body language to evaluate the current emotional and cognitive states of the participants;

[0128] 4) At the same time, the role and professional knowledge modeler analyzes the historical speech content and interaction patterns of the participants to construct a knowledge domain and meeting role model;

[0129] 5) The information from the three parts is finally integrated into a complete participant model and passed to subsequent components for use.

[0130] Among them, after obtaining the three parts of information, the participant analysis component also needs to adjust the importance according to the topic based on the dynamic weight allocation mechanism, and continuously update the participant model based on the adaptive modeling of interaction behaviors. Finally, the data is integrated to obtain the output, which includes a participant identity mapping table containing spatial positions, a data stream of emotional and cognitive states related to real-time emotions and attention, and a dynamic participant model related to professional fields and interaction styles. In addition, based on the output, optimization feedback can be provided for the aforementioned identity recognition and tracker, participant status analyzer, and role and expertise modeling tool.

[0131] Combined with Figure 5 As shown, regarding the meeting content processing component in the system, it is the core cognitive processing unit of the system, responsible for converting the unstructured communication in the meeting into a machine-understandable structured knowledge representation. The position of this component in the system architecture is responsible for: receiving the audio stream from the data acquisition and preprocessing module and the speaker context from the participant analysis module; performing in-depth semantic understanding and structured processing; outputting a multi-level representation of the meeting content (text transcription, topic structure, argumentation framework, knowledge association, and information value assessment); providing a semantic basis for group dynamics and decision analysis. This component consists of three core units: a speech understanding and transcription unit for converting audio into text and annotating basic semantic features, a content structuring analyzer for identifying topic structures and semantic relationships, and an argumentation and quality assessment unit for analyzing the logical framework and information quality of the discussion. The data processing flow in this component is as follows:

[0132] 1) Receive the audio stream and speaker information;

[0133] 2) The speech understanding and transcription unit converts it into text with timestamps and speaker tags;

[0134] 3) The content structuring analyzer performs topic recognition, semantic segmentation, and relationship extraction to generate a hierarchical representation of the meeting content;

[0135] 4) At the same time, the argumentation and quality assessment unit analyzes the argumentation structure, information source, and reliability;

[0136] 5) Integrate the analysis results into a unified structured representation and pass it to the subsequent module.

[0137] In addition, the information output by the meeting content processing component will also be fed back for learning in the forward direction. This component establishes multi-directional interactions with other modules: receiving speaker context (identity, emotional state, professional background) from the participant analysis module; providing structured meeting content and argument analysis to the group dynamics and decision-making analysis module; receiving interaction network and decision path analysis from the group dynamics module; and using the feedback information to optimize the content structure and importance assessment. This multi-directional interaction ensures the co-evolution of each module of the system, improves the overall performance, and provides comprehensive support for efficient meeting decision-making.

[0138] Combined Figure 6 As shown, regarding the decision-making analysis component in the system, it is the high-level cognitive processing unit of the system, responsible for analyzing the team interaction pattern and the decision-making formation process from an overall perspective. The position of this component in the system architecture: located at the top of the analysis chain, responsible for receiving the participant model from the participant analysis module and the structured meeting content from the content understanding module; responsible for conducting fusion analysis and high-order reasoning; responsible for outputting interaction network analysis, decision path evaluation, and efficiency and quality reports; responsible for providing the basis for intervention to the intelligent assistance module and feeding back to the previous module to optimize the processing strategy. The component is internally composed of three functional units: an interaction network analyzer for modeling the interaction pattern and influence propagation between participants, a decision process tracker for identifying and analyzing decision points and their evolution paths, and an efficiency and quality evaluator for evaluating the meeting efficiency and outcome quality from multiple dimensions. Regarding the data processing flow of this component:

[0139] 1) Receive participant status and interaction records;

[0140] 2) The interaction network analyzer constructs a dynamic interaction network through time series graph analysis;

[0141] 3) The decision process tracker identifies decision points based on the structured meeting content and tracks their evolution;

[0142] 4) The efficiency and quality evaluator integrates the analysis results of the first two units and generates a multi-dimensional evaluation in combination with the meeting objectives;

[0143] 5) Integrate the analysis results into a group dynamics and decision report to guide intelligent intervention and meeting optimization.

[0144] In addition, the dynamic importance perception mechanism in the decision-making analysis component and the output information will be used to optimize and feedback to the previous interaction network analyzer. Its interaction with other modules: receiving the status and role information from the participant analysis module; receiving the structured content and argument analysis from the content understanding module; providing interaction network, decision path, and efficiency evaluation to the intelligent assistance and output module; receiving the intervention effect evaluation from the intelligent assistance module to optimize the analysis model. This multi-directional interaction enables this module to continuously optimize the analysis framework, provide accurate and valuable insights into group dynamics, and support the improvement of meeting efficiency and decision-making quality.

[0145] Combined Figure 7 As shown, regarding the intervention output component in the system, it is the terminal interaction interface of the entire system, responsible for transforming in-depth analysis into practical interventions and valuable resources. Location in the system architecture: at the end of the execution chain, responsible for receiving the analysis results of all previous modules, especially the advanced insights of the group dynamics and decision-making analysis module; responsible for comprehensive processing and strategic decision-making; responsible for generating user-oriented interventions and resources; responsible for providing feedback on the intervention effect to support the overall optimization cycle of the system. It is internally composed of three functional units: a real-time assistance generator for providing timely and appropriate interventions during the meeting, a visualization engine for transforming complex analysis results into intuitive and understandable visual expressions, and a meeting resource generator for integrating meeting content and analysis results to generate post-meeting resources. Its data processing flow is as follows:

[0146] 1) Receive multiple result streams from the previous analysis (participant status data, structured meeting content, interaction network analysis, decision evaluation, etc.);

[0147] 2) Integrate heterogeneous data and prioritize them;

[0148] 3) The real-time assistance generator decides whether to intervene, when to intervene, the content and form of the intervention based on the meeting status and historical intervention effects;

[0149] 4) Select an appropriate presentation method according to the intervention type: the visualization engine generates visual content, and a dedicated processing unit generates text and audio prompts;

[0150] 5) Transmit the intervention content through the user interface and record the intervention time and content;

[0151] 6) Parallel accumulate and analyze meeting content, continuously update the resource model;

[0152] 7) After the meeting, generate customized meeting summaries, action item lists, decision records, and learning resources;

[0153] 8) Distribute post-meeting resources and collect usage feedback.

[0154] In addition, the intervention output component will also perform an optimized feedback loop operation, and intervention effect monitoring is required: record the details of each intervention (time, content, target recipient), track the changes in the meeting after the intervention (attention shift, discussion adjustment, decision-making progress), compare the changes with the expected effects, evaluate the effectiveness of the intervention, and use the evaluation results to optimize subsequent interventions and feedback to the system model. It is also necessary to be based on a continuous learning mechanism: track the long-term implementation results of the meeting decisions, associate the results with the meeting process analysis, identify areas for improvement, and provide targeted support for future similar discussions. And regarding user experience adaptation: record and analyze the response patterns of users to interventions and resources, build a personalized user model, adjust the intervention strategies and resource generation parameters, and match user preferences. This comprehensive intelligent assistance and output design ensures that the system can provide the maximum support under the principle of minimum interference, achieve the goal of "enhancing" rather than "replacing" human collaboration, and provide an all-round improvement for meeting efficiency and decision-making quality.

[0155] See Figure 8 As shown, an embodiment of the present application also correspondingly discloses a meeting enhancement device, which is applied to a preset meeting enhancement system. The preset meeting enhancement system includes a data preprocessing component, a participant analysis component, a meeting content processing component, a decision analysis component, and an intervention output component; the device includes:

[0156] A data preprocessing module 11, configured to perform corresponding data preprocessing on the original information of each modality in the multi-modal original information of the meeting through the data preprocessing component, and perform multi-modal data fusion and time alignment processing based on the preprocessing results to determine a target data packet;

[0157] A participant analysis module 12, configured to identify the participant identities through the participant analysis component and the target data packet, and perform participant status analysis based on the identification results to determine a target model for analyzing the behaviors and states of the participants, and determine the speaker context of the meeting based on the target model;

[0158] A content processing module 13, configured to perform a structured representation of the meeting content through the meeting content processing component, the speaker context, and the target data packet, and perform argument analysis based on the structured content to determine an argument analysis result;

[0159] A meeting analysis module 14, configured to parse the team interaction mode and the decision-making process of the meeting from an overall perspective through the decision analysis component, the target model, the structured content, and the argument analysis result to determine an analysis result;

[0160] The meeting intervention module 15 is configured to determine whether to conduct a meeting intervention currently based on the intervention output component, the parsing result, the structured content, the current meeting status information, and the historical intervention feedback information, and when the determination is affirmative, to determine and output a target intervention suggestion.

[0161] Among them, the more specific working processes of the above-mentioned various modules can refer to the corresponding content disclosed in the foregoing embodiments, and will not be elaborated herein.

[0162] Thus, in this application, a preset meeting enhancement system with hierarchical and multi-module collaboration is used to achieve meeting enhancement. Specifically, first, the multi-modal raw information of the meeting is preprocessed by the data preprocessing component in the system, and multi-modal data fusion and time alignment processing are performed to determine the target data packet. Then, the participant analysis component and the target data packet in the system are used to identify the participant identities and perform participant status analysis, so as to determine the speaker context of the meeting by using the target model obtained for analyzing the behaviors and statuses of the participants. After that, the meeting content processing component and the speaker context in the system are used to determine the structured content of the meeting and conduct argumentation analysis. Then, the decision analysis component, the structured content, and the argumentation analysis results in the system are used to respectively analyze the team interaction mode and the decision-making formation process of the meeting from an overall perspective. Then, the intervention output component and the parsing result, the current meeting status information, and the historical intervention feedback information in the system are used to determine whether to intervene in the meeting, and if so, to determine the target intervention suggestion. In this way, the deficiencies in aspects such as multi-modal meeting information processing, participant status analysis, and decision-making process tracking in the existing related solutions can be effectively solved, and the efficiency of the meeting and the effectiveness and quality of the meeting results are improved.

[0163] In some specific embodiments, the data preprocessing module 11 may specifically be configured to: receive the original signals sent by each hardware device through the data preprocessing component; the hardware devices include microphones, cameras, and environmental sensors; perform digitization and standardization processing on the original signals to determine the processed audio data, video data, and environmental data; perform noise filtering on the audio data based on a first preset algorithm to determine the first preprocessed data; perform region of interest recognition on the video data based on a second preset algorithm and perform resource allocation based on the region recognition result to determine the second preprocessed data; perform environmental state evaluation on the environmental data based on the standardized environmental parameter vector and the multi-layer perceptron structure to determine the third preprocessed data; the standardized environmental parameter vector includes a standardized humidity vector and a standardized temperature vector; after adjusting the data acquisition configuration and resource allocation based on the hierarchical data acquisition protocol, the resource allocation algorithm, and the currently determined first preprocessed data, second preprocessed data, and third preprocessed data, perform data fusion on the first preprocessed data, second preprocessed data, and third preprocessed data, and perform time alignment on the fused data through the timestamp alignment mechanism to determine the target data packet.

[0164] In some specific embodiments, the participant analysis module 12 may specifically be configured to: simultaneously analyze the acoustic features and visual features in the target data packet through the participant analysis component and the progressive recognition strategy to determine the participant identity and the corresponding spatial position tracking result; based on the participant status analyzer in the participant analysis component, the participant identity, and the spatial position tracking result, and adopt the speech-gesture-expression collaborative analysis framework to evaluate the current emotion and current cognitive state of the participant to determine the participant evaluation result; based on the role and expertise modeler in the participant analysis component, the participant evaluation result, analyze the historical speech content and interaction pattern of the participant to construct the knowledge model corresponding to each current participant identity; based on the knowledge model, the weight allocation mechanism, the participant identity, and the current topic of the meeting, adjust the weights of different types of professional knowledge and role attributes to determine the participant identity mapping table; update the knowledge model based on the interactive behavior adaptive modeling mechanism and the recorded participant behavior characteristics, and capture and analyze the behavior trend according to the updated knowledge model and the participant identity mapping table to determine the speaker context of the meeting; the speaker context includes speaker identity information, speaker status information, and speaker cognitive load information.

[0165] In some specific embodiments, the content processing module 13 may specifically be configured to: receive the speaker context and the target data packet through the conference content processing component; perform text conversion on the target data packet to determine a conversion result; perform timestamp marking and speaker identity marking on the conversion result based on the speaker context to determine the marked text; perform topic recognition, semantic segmentation, and relationship extraction on the marked text respectively based on a content structure analyzer and a hierarchical attention network architecture to determine the structured content; perform structural argument analysis on the structured content based on the speaker context, and perform information source analysis and information reliability analysis on the structured content based on a preset analysis rule to determine an argument analysis result.

[0166] In some specific embodiments, the conference parsing module 14 may specifically be configured to: receive the target model, the structured content, and the argument analysis result through the decision analysis component; construct and analyze a timing diagram based on an interaction network analyzer in the decision analysis component and the structured content, and analyze the influence propagation path, information bottleneck points, and key connector roles in the determined dynamic interaction network to determine a first analysis result; identify decision points for the structured content based on a decision process tracker in the decision analysis component, and perform evolutionary tracking of the decision points according to the decision point identification result to analyze the decision formation process based on the tracking result to determine a second analysis result; perform multi-dimensional quantitative evaluation based on an evaluator in the decision analysis component, the conference objective of the conference, the argument analysis result, the first analysis result, and the second analysis result to determine a corresponding quality report; perform implicit obstacle detection based on the dynamic interaction network, and determine whether to trigger a cognitive gap marking operation according to the obstacle monitoring result and the target model.

[0167] In some specific embodiments, the conference intervention module 15 may specifically be configured to: perform heterogeneous data integration on the received parsing result, the structured content, the speaker context, and the argument analysis result through the intervention output component, and perform priority sorting on the data based on the integration result to determine a sorting result; determine whether to perform conference intervention currently based on an auxiliary generator in the intervention output component, the sorting result, the current conference status information, and the historical intervention feedback information to obtain an intervention judgment result; when the intervention judgment result is yes, determine a target intervention suggestion and a corresponding intervention type based on the current conference status information; determine an intervention output mode corresponding to the target intervention suggestion based on the intervention type, and trigger an intervention output operation of the target intervention suggestion according to the intervention output mode; the target intervention suggestion includes an intervention timing, intervention content, and intervention form.

[0168] In some specific embodiments, the meeting enhancement device may specifically further be used for: after outputting the target intervention suggestion, collecting intervention feedback information corresponding to the target intervention suggestion through the intervention output component; evaluating the effectiveness of the intervention based on the intervention feedback information, and sending the intervention evaluation result to the decision analysis component, so that the decision analysis component triggers an analysis and optimization operation based on the intervention evaluation result; after the meeting ends, integrating the multi-modal original information and the parsing result during the meeting through the intervention output component and based on the current target model, to determine and distribute post-meeting resources based on the integration result; the post-meeting resources include a meeting summary, a decision record, and learning resource recommendation information.

[0169] Furthermore, an embodiment of the present application also discloses an electronic device, Figure 9 which is a structural diagram of an electronic device 20 shown according to an exemplary embodiment. The content in the figure should not be considered as any limitation on the scope of use of the present application.

[0170] Figure 9 This is a schematic structural diagram of an electronic device 20 provided by an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the meeting enhancement method disclosed in any of the foregoing embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0171] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows is any communication protocol applicable to the technical solution of the present application, and no specific limitation is imposed on it here; the input / output interface 25 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application needs, and no specific limitation is made here.

[0172] In addition, the memory 22, as a carrier for resource storage, may be a read-only memory, a random access memory, a disk, or an optical disc, etc. The resources stored thereon may include an operating system 221, a computer program 222, etc., and the storage method may be short-term storage or permanent storage.

[0173] Among them, the operating system 221 is used to manage and control each hardware device and computer program 222 on the electronic device 20, and it can be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program capable of implementing the conference enhancement method executed by the electronic device 20 disclosed in any of the foregoing embodiments, the computer program 222 may further include a computer program capable of performing other specific tasks.

[0174] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the foregoing disclosed conference enhancement method is implemented. For the specific steps of this method, reference may be made to the corresponding content disclosed in the foregoing embodiments, and details will not be repeated here.

[0175] In this specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and reference can be made to the description in the method part for related parts.

[0176] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0177] The steps of the method or algorithm described in combination with the embodiments disclosed herein can be directly implemented by hardware, a software module executed by a processor, or a combination of the two. The software module can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.

[0178] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.

[0179] The technical solutions provided by the present application have been introduced in detail above. Specific examples are used in this text to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A conference enhancement method, characterized in that: Applied to a preset conference enhancement system, the preset conference enhancement system includes a data preprocessing component, a participant analysis component, a conference content processing component, a decision analysis component and an intervention output component; the method includes: The data preprocessing component performs corresponding data preprocessing on the original information of each mode in the multimodal original information of the conference, and performs multimodal data fusion and time alignment processing based on the preprocessing result to determine the target data packet; Performing participant identification through the participant analysis component and the target data packet, and performing participant status analysis based on the identification result to determine a target model for analyzing the behavior and status of the participant, and determining the speaker context of the conference based on the target model; Performing a structured representation of the conference content through the conference content processing component, the speaker context, and the target data packet, and performing argumentation analysis based on the structured content to determine an argumentation analysis result; By using the decision analysis component, the target model, the structured content and the argumentation analysis results, and adopting an overall perspective, the team interaction mode and the decision-making process of the meeting are analyzed respectively to determine the analysis results; Whether conference intervention is currently being performed is determined through the intervention output component, the parsing result, the structured content, the current conference status information, and the historical intervention feedback information, and when so, a target intervention suggestion is determined and output.

2. The conference enhancement method according to claim 1, characterized in that: The data preprocessing component performs corresponding data preprocessing on the original information of each mode in the multimodal original information of the conference, and performs multimodal data fusion and time alignment processing based on the preprocessing result, including: Receiving original signals sent by various hardware devices through the data preprocessing component; the hardware devices include microphones, cameras and environmental sensors; Performing digitization and standardization processing on the original signal to determine processed audio data, video data and environmental data; Performing noise filtering on the audio data based on a first preset algorithm to determine first preprocessed data; Identifying a region of interest on the video data based on a second preset algorithm, and allocating resources based on the region identification result to determine second preprocessed data; Performing an environmental state assessment on the environmental data based on a standardized environmental parameter vector and a multi-layer perceptron structure to determine third pre-processed data; the standardized environmental parameter vector includes a standardized humidity vector and a standardized temperature vector; After adjusting the data acquisition configuration and resource allocation based on the hierarchical data acquisition protocol, the resource allocation algorithm and the currently determined first preprocessed data, the second preprocessed data and the third preprocessed data, data fusion is performed on the first preprocessed data, the second preprocessed data and the third preprocessed data, and the fused data is time-aligned through a timestamp alignment mechanism to determine the target data packet.

3. The conference enhancement method according to claim 1, characterized in that: The identifying of the participant through the participant analysis component and the target data packet, and analyzing the participant status based on the identification result to determine a target model for analyzing the behavior and status of the participant, and determining the speaker context of the conference based on the target model, includes: The acoustic features and visual features in the target data packet are simultaneously analyzed by the participant analysis component and the progressive recognition strategy to determine the participant identity and the corresponding spatial position tracking result; Based on the participant state analyzer in the participant analysis component, the participant identity and the spatial position tracking result, and using a voice-gesture-expression collaborative analysis framework to evaluate the participant's current emotion and current cognitive state, so as to determine the participant evaluation result; Analyze the historical speech content and interaction mode of the participants based on the role and expertise modeler in the participant analysis component and the participant evaluation results to construct a knowledge model corresponding to the current identity of each of the participants; Based on the knowledge model, the weight allocation mechanism, the identities of the participants, and the current agenda of the meeting, weights of different types of expertise and role attributes are adjusted to determine a participant identity mapping table; The knowledge model is updated based on the interactive behavior adaptive modeling mechanism and the recorded behavior characteristics of the participants, and the behavior trends are captured and analyzed based on the updated knowledge model and the participant identity mapping table to determine the speaker context of the meeting; the speaker context includes the speaker identity information, the speaker status information and the speaker cognitive load information.

4. The conference enhancement method according to claim 1, characterized in that: The structured representation of the conference content is performed through the conference content processing component, the speaker context and the target data packet, and the argumentation analysis is performed based on the structured content, including: Receiving the speaker context and the target data packet through the conference content processing component; Performing text conversion on the target data packet to determine a conversion result; Performing a timestamp and a speaker identity tag on the conversion result based on the speaker context to determine a tagged text; Based on a content structured analyzer and a hierarchical attention network architecture, topic identification, semantic segmentation, and relationship extraction are performed on the labeled text to determine structured content; The structured content is subjected to a structural argumentation analysis based on the speaker context, and the structured content is subjected to an information source analysis and an information reliability analysis based on preset analysis rules to determine an argumentation analysis result.

5. The conference enhancement method according to claim 1, characterized in that: The decision analysis component, the target model, the structured content and the argumentation analysis result are used to analyze the team interaction mode and decision-making process of the meeting from an overall perspective, including: Receiving the target model, the structured content and the argumentation analysis result through the decision analysis component; Constructing and analyzing a time sequence diagram based on the interactive network analyzer in the decision analysis component and the structured content, and analyzing the influence propagation path, information bottleneck points, and key connector roles in the determined dynamic interactive network to determine a first analysis result; Identifying decision points of the structured content based on a decision process tracker in the decision analysis component, and tracking the evolution of the decision points according to the decision point identification results, so as to analyze the decision formation process based on the tracking results to determine a second analysis result; Performing a multi-dimensional quantitative evaluation based on the evaluator in the decision analysis component, the meeting goal of the meeting, the argumentation analysis result, the first analysis result, and the second analysis result to determine a corresponding quality report; Hidden obstacle detection is performed based on the dynamic interactive network, and whether to trigger a cognitive gap marking operation is determined according to the obstacle monitoring result and the target model.

6. The conference enhancement method according to claim 1, characterized in that: The determining whether to conduct a conference intervention at present through the intervention output component, the parsing result, the structured content, the current conference status information and the historical intervention feedback information, and determining and outputting a target intervention suggestion when it is conducted, includes: Performing heterogeneous data integration on the received parsing results, the structured content, the speaker context, and the argumentation analysis results through the intervention output component, and prioritizing the data based on the integration results to determine a ranking result; Determine whether to conduct a conference intervention at present based on the auxiliary generator in the intervention output component, the sorting result, the current conference status information and the historical intervention feedback information to obtain an intervention judgment result; When the intervention judgment result is yes, determining a target intervention suggestion and a corresponding intervention type based on the current meeting state information; An intervention output mode corresponding to the target intervention suggestion is determined based on the intervention type, and an intervention output operation of the target intervention suggestion is triggered according to the intervention output mode; the target intervention suggestion includes intervention timing, intervention content and intervention form.

7. The conference enhancement method according to claim 1, characterized in that: Also includes: After outputting the target intervention suggestion, collecting intervention feedback information corresponding to the target intervention suggestion through the intervention output component; Performing an effectiveness evaluation of the intervention based on the intervention feedback information, and sending the intervention evaluation result to the decision analysis component, so that the decision analysis component triggers an analysis optimization operation based on the intervention evaluation result; After the meeting is over, the multimodal original information and the analysis results during the meeting are integrated through the intervention output component and based on the current target model, so as to determine and distribute post-meeting resources based on the integration results; the post-meeting resources include meeting summaries, decision records, and learning resource recommendation information.

8. A conference enhancement device, characterized in that: Applied to a preset conference enhancement system, the preset conference enhancement system includes a data preprocessing component, a participant analysis component, a conference content processing component, a decision analysis component and an intervention output component; the device includes: A data preprocessing module, used to perform corresponding data preprocessing on the original information of each mode in the multimodal original information of the conference through the data preprocessing component, and perform multimodal data fusion and time alignment processing based on the preprocessing result to determine the target data packet; A participant analysis module, configured to identify the identity of the participant through the participant analysis component and the target data packet, and perform participant status analysis based on the identification result to determine a target model for analyzing the behavior and status of the participant, and determine the speaker context of the conference based on the target model; A content processing module, used for performing a structured representation of the conference content through the conference content processing component, the speaker context and the target data packet, and performing argumentation analysis based on the structured content to determine an argumentation analysis result; A meeting analysis module, for analyzing the team interaction mode and decision-making process of the meeting respectively from an overall perspective through the decision analysis component, the target model, the structured content and the argumentation analysis result, so as to determine the analysis result; The conference intervention module is used to determine whether conference intervention is currently being performed through the intervention output component, the parsing result, the structured content, the current conference status information and the historical intervention feedback information, and when so, determine and output a target intervention suggestion.

9. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the conference enhancement method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: Used to store a computer program, which, when executed by a processor, implements the conference enhancement method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Conference participant scoring method and device based on video conference, equipment and medium

    CN111970471A

  • Remote video conference intelligent management system based on big data and cloud computing and cloud conference management platform

    CN112801608A

  • Multi-modal conference data structuring method and device and computer equipment

    CN114298170A

  • Hierarchical multi-modal sentiment analysis method based on multi-task learning

    CN114973045A

  • Multifunctional video conference interaction method and device, equipment and storage medium

    CN118612378A

Cited By

  • Intelligent conference recording and recording method and system

    CN120568005A

  • Video conference multi-modal data alignment method and device based on causal mask, equipment and medium

    CN120763869A

  • A video conference multi-modal data alignment method and device based on a causal mask, equipment and medium

    CN120763869B

  • Multi-role configuration and effective judgment method based on semantic arbitration

    CN120809299A

  • Intelligent voice conference behavior analysis method and system based on multi-source perception

    CN121462328A