Multi-modal data-based multi-user interaction innovation ability evaluation system
By constructing a multi-person interactive innovation capability assessment system based on multimodal data, the problems of insufficient ecological validity, strong subjectivity, and single data dimension in existing technologies have been solved. This system enables multi-dimensional assessment in real interactive scenarios, improving the objectivity and efficiency of the assessment results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SOUTHWEST UNIV
- Filing Date
- 2025-12-18
- Publication Date
- 2026-05-08
AI Technical Summary
Existing technologies, when assessing individual innovation capabilities, lack ecological validity, are highly subjective, have limited data dimensions, and suffer from poor efficiency and scalability, failing to meet the need for comprehensive insights in real-world interactive scenarios.
A multi-person interactive innovation capability assessment system based on multimodal data is constructed. The system synchronously collects and analyzes the voice, text, and video data of participants during the discussion process through an online meeting platform. It integrates natural language processing, speech signal processing, and emotion recognition artificial intelligence technologies to construct a multi-dimensional assessment model.
It enables multi-dimensional assessment in real-world interactive scenarios, improving the objectivity and comprehensiveness of assessment results, increasing assessment efficiency, generating structured feedback reports, and reducing human and time costs.
Smart Images

Figure CN121998474A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of evaluation systems, and more specifically to a multi-person interactive innovation capability evaluation system based on multimodal data. Background Technology
[0002] Current assessments of innovation ability primarily rely on individual-level situational tasks, using manual scoring or unimodal data analysis to assess subjects' independent innovation capabilities. These methods have the following limitations:
[0003] Insufficient ecological validity: Traditional assessments are detached from real-world interactive environments and fail to reflect individuals' social behaviors such as viewpoint integration and conflict resolution in team collaboration;
[0004] High subjectivity: Human scoring is easily influenced by the cultural background and experience of the raters, which may lead to biased evaluation results; in addition, single-modal data (such as text answers) is difficult to capture dynamic features such as language communication and emotional changes.
[0005] The data dimension is too narrow: it only focuses on the novelty, variety and quantity of the output viewpoints, and ignores process behaviors such as language process, intra-group role collaboration and viewpoint selection.
[0006] Poor efficiency and scalability: Manual evaluation is difficult to adapt to large-scale scenarios and lacks intelligent prediction and feedback mechanisms.
[0007] With the upgrading of educational reforms and corporate innovation needs, competency assessment in real-world interactive scenarios has become a necessity. Enterprises need to identify employees with innovative capabilities, emotional stability, and leadership skills, while educational settings require evaluating students' performance in group collaborations. However, existing technologies cannot meet the demand for comprehensive insights in real-world situations.
[0008] To address the aforementioned issues, the applicant proposes a multi-person interactive innovation capability assessment system based on multimodal data. Summary of the Invention
[0009] The purpose of this invention is to provide a multi-person interactive innovation capability assessment system based on multimodal data to solve the problems in the prior art.
[0010] To achieve the above objectives, the present invention provides the following technical solution: a multi-person interactive innovation ability assessment system based on multimodal data, the system comprising:
[0011] System architecture: Built on a real online meeting platform, it collects and analyzes multimodal behavioral data of participants during discussions by designing standardized creative task scenarios;
[0012] Core functional modules:
[0013] Scene setup and equipment debugging module: used to guide participants to enter the online meeting room, adjust the camera angle and equipment position, and ensure the quality of data collection;
[0014] Instructions and Practice Module: Explains the task objectives and operating procedures to participants, and provides example tasks for practice;
[0015] Testing Phase and Data Acquisition Module: Discuss creative questions within a specified time and collect audio, text, and video data;
[0016] Data processing and feature extraction module: performs unified organization and multi-level feature extraction on the collected three-modal data, including text semantic analysis, speech feature extraction and facial expression recognition;
[0017] Analysis, modeling, and evaluation output module: Based on the extracted features, a multidimensional evaluation model of individual innovation ability and leadership is constructed, and the evaluation results and structured feedback reports are output;
[0018] Technical features: It integrates natural language processing, speech signal processing and emotion recognition artificial intelligence technologies to achieve multimodal data fusion and automated evaluation.
[0019] Optionally, the multimodal data acquisition method specifically includes:
[0020] Voice data acquisition: Collect participants' voice data using high-quality audio equipment for voice feature analysis;
[0021] Text data acquisition: Converting speech data into text data using speech-to-text technology for semantic analysis and time-series data analysis;
[0022] Video data acquisition: Facial video data of participants is acquired through cameras for facial expression and emotion recognition and analysis.
[0023] Optionally, the feature fusion model is specifically as follows:
[0024] A multimodal feature fusion model is constructed, which integrates text semantic features, speech acoustic features, and facial expression features. An assessment model for individual innovation ability and leadership is built using traditional machine learning algorithms or deep neural networks.
[0025] Optionally, the evaluation index system includes:
[0026] Textual semantic metrics: Based on the Word2Vec word vector model, participants' speech-to-text transcription is mapped to word vectors, and a semantic network is constructed on this basis. The originality of viewpoints is assessed by calculating semantic distance, hierarchical clustering is used to analyze viewpoint distribution to quantify thinking flexibility, and the fluency of expression is assessed by combining text quantity. Simultaneously, the degree centrality of nodes in the semantic network is calculated to predict an individual's leadership in the discussion.
[0027] Text temporal metrics: A temporal semantic network is constructed based on a combination of speech order and semantic similarity. Novelty is used to measure the difference between an individual's speech and previous content; jump rate is used to assess the leap between adjacent speeches; influence is used to characterize the sustained effect of preceding speeches on subsequent speeches; and centroid distance reflects the degree of deviation of an individual from the overall semantic center of the group. Simultaneously, directed centrality is calculated in the semantic network with influence as the weighted edge to capture an individual's position in the group's semantic flow, and its contribution level is measured in conjunction with the number of speeches.
[0028] Speech acoustic metrics: Multidimensional acoustic features are extracted from participants' speech signals, including speech rate (number of syllables per unit time), pitch (fundamental frequency F0), intonation (pitch variation pattern), pause duration and frequency, and energy intensity. Based on these features, emotional tension (through pitch and energy fluctuation amplitude), interactive initiative (through speech proportion and response delay time), and emotional stability (through pitch and energy variance) can be calculated, thereby quantifying an individual's emotional expression and interactive state in discussions.
[0029] Facial expression metrics: Utilizing facial action unit (AU) detection and expression recognition algorithms, this method extracts an individual's emotion category (e.g., pleasure, surprise, focus), amplitude of emotion changes (differences in facial expression features across different time periods), and emotional stability (variance or standard deviation of expression changes) during discussions. Based on these metrics, an individual's emotional expression ability, emotion regulation ability, and social adaptability and engagement in team interactions can be quantified. Optionally, it also includes application scenario extensions, which include:
[0030] Educational Scenario: Applicable to the assessment of student ability development in secondary education and university general education courses, it can comprehensively evaluate students' innovative thinking, language expression ability, teamwork level and emotional regulation ability; at the same time, it can provide a scientific basis for personalized learning guidance, ability enhancement plan development and academic potential prediction, and support teachers to conduct targeted intervention and feedback in classroom teaching and extracurricular activities.
[0031] Enterprise Scenario: Suitable for talent recruitment, team capability assessment and innovation potential evaluation for technology innovation enterprises. It can help enterprises identify core talents with creativity, collaboration ability and leadership potential, and provide data support for team building and job matching.
[0032] Other multi-person interactive scenarios: Applicable to any scenario requiring assessment of an individual's innovation ability, leadership potential, and social adaptability in group collaboration, such as team brainstorming, project collaboration, corporate training workshops, academic group discussions, and innovation competitions. It can quantitatively analyze participants' language expression, opinion integration ability, emotional regulation, and influence during interactions, providing a scientific basis for team building, talent selection, and skills development. Optionally, the voice data acquisition process also includes a preprocessing step for the voice data to improve the accuracy of voice feature extraction.
[0033] Optionally, the feature fusion model may also include normalization of multimodal features during its construction process to eliminate dimensional differences between different modal data and improve the stability of the evaluation model.
[0034] Optionally, the evaluation index system may also include weighting of each index during quantitative evaluation to reflect the relative importance of different indicators in the evaluation of innovation capability and leadership.
[0035] Optionally, the application scenario expansion also includes customized development for industries and fields to meet the specific needs of different industries for evaluating innovation capabilities and leadership.
[0036] Beneficial effects:
[0037] 1. Ecological effectiveness has been significantly improved.
[0038] Traditional methods are divorced from real-world interactive scenarios. This invention simulates team collaboration through an online meeting platform, simultaneously collecting language, voice, and emotion data from natural communication. This achieves a leap from static individual assessment to dynamic group interaction, making the assessment results more closely resemble actual educational / organizational environments.
[0039] 2. Objectivity and data dimensions have been comprehensively expanded.
[0040] By integrating multimodal artificial intelligence technology, it overcomes the limitations of single-modal data and manual scoring, and constructs a multidimensional assessment model covering cognitive ability, emotional stability and leadership through text semantic analysis, speech acoustic feature extraction and facial expression recognition.
[0041] 3. Advantages of automated evaluation efficiency
[0042] The system can complete data collection, feature extraction, and modeling analysis without human intervention. The automated evaluation efficiency is improved compared with traditional methods, and it generates structured feedback reports, which greatly reduces manpower and time costs.
[0043] 4. Multimodal fusion enhances the comprehensiveness of the evaluation.
[0044] Text data reflects the innovativeness of viewpoints, voice data reveals the tension of interaction, and facial expression data captures the ability to regulate emotions. The fusion of the three modalities improves the accuracy of prediction of innovation ability and leadership, and achieves a deeper assessment from surface indicators to deep psychological characteristics. Attached Figure Description
[0045] Figure 1 This is a flowchart illustrating the scenario deployment and equipment debugging process according to an embodiment of the present invention;
[0046] Figure 2 This is a flowchart illustrating the instructions and exercises for embodiments of the present invention.
[0047] Figure 3 This is a flowchart illustrating the data acquisition process in an embodiment of the present invention.
[0048] Figure 4 This is a flowchart of data processing and feature extraction in an embodiment of the present invention;
[0049] Figure 5 This is a flowchart illustrating the feature fusion and modeling evaluation process in an embodiment of the present invention.
[0050] Figure 6 This is a flowchart illustrating the application scenarios and feedback output of an embodiment of the present invention. Detailed Implementation
[0051] The preferred embodiments of the present invention are described below with reference to the accompanying drawings to make the technical content clearer and easier to understand. The present invention can be embodied in many different forms, and the scope of protection of the present invention is not limited to the embodiments mentioned herein.
[0052] This invention proposes a multi-person interactive innovation capability assessment system based on multimodal data. Its technical solution covers multiple aspects, including system architecture design, data acquisition and processing flow, multimodal feature extraction method, feature fusion model construction, assessment index system, automated assessment implementation steps, and implementation methods for specific application scenarios. The following is a detailed explanation of these technical details.
[0053] In terms of system architecture, this system is built on a commonly used video conferencing platform, such as Tencent Meeting, leveraging its stability and broad user base to achieve multi-person interactive evaluation in real online meeting scenarios. The system is divided into five main parts: data acquisition layer, data processing layer, feature extraction layer, model building layer, and evaluation output layer. The data acquisition layer is responsible for simultaneously collecting multimodal data from participants during discussions, including textual language, speech features, and facial emotional changes. The data processing layer cleans, preprocesses, and standardizes the format of the collected raw data. The feature extraction layer uses natural language processing, speech signal processing, and emotion recognition technologies to extract key features reflecting individual innovation capabilities and leadership from the processed data. The model building layer fuses the extracted multimodal features to construct a multidimensional evaluation model. The evaluation output layer outputs each individual's innovation capability score, leadership tendency index, and a summary of their performance in group collaboration based on the model, and generates a structured feedback report.
[0054] Data acquisition and processing are fundamental to system operation. During the scenario setup and equipment debugging phase, three participants entered the online meeting room via the selected video conferencing platform. Before the test began, the instructor instructed all participants to turn on their cameras and microphones and adjust the camera angles and device positions according to standardized requirements. Participants were ensured that their faces were directly in front of the camera, unobstructed, and approximately 60 centimeters away from the camera. They were encouraged to maintain a stable sitting posture and natural facial expressions, ensuring the head and facial features were fully visible in the video feed, avoiding obstructions and strong light interference. Testers pre-checked the video and audio signals to ensure complete and accurate data acquisition before proceeding to the next stage. In the instruction and practice phase, after the equipment was properly configured, testers explained the task objectives and operating procedures to the participants. To help them familiarize themselves with the process, the system provided a sample task for practice. No formal data was collected during the practice phase; it was only used to adapt to the discussion format and interaction rhythm. At the start of the formal testing phase, participants discussed creative problems within a specified time according to the rules. After the discussion, the system prompted each participant to select the most influential member based on the discussion process and recorded the voting results. If a member received at least two votes, they were marked as the leader of the group. The entire testing phase collected audio, text, and video data. After data collection, the data processing phase began. The system uniformly organized the three modalities of data. Text data was cleaned to remove noise and irrelevant information, followed by sentence segmentation and semantic representation processing. Audio and video data also underwent corresponding preprocessing to ensure the accuracy of subsequent feature extraction.
[0055] Multimodal feature extraction is one of the core components of the system. For text data, voiceprint recognition technology is first used to distinguish the speech information of different individuals, achieving participant-level text data annotation. Then, the text content is segmented into sentences and task-relevant filtering is performed, retaining language fragments highly relevant to the problem-solving task to construct an analysis corpus. Based on the word2vec word embedding model, the retained text is mapped to word vectors. Semantic distance calculation and hierarchical clustering methods are used to extract relevant indicators reflecting individual innovation capabilities. For example, to calculate the semantic distance between keywords in participants' speeches, we first convert the participants' speeches into word vector representations, assuming two word vectors are... and The semantic distance between word vectors is calculated using cosine distance: dist( , ) = 1- Where d represents the word vector dimension. A larger semantic distance indicates a greater difference between two words in the semantic space, representing higher novelty and divergence in the statements, and can be used as a basis for assessing the activity of innovative thinking. Furthermore, a semantic network is constructed based on each group of three dialogues, with each keyword or phrase as a node. Edges between nodes are established based on semantic similarity. If the cosine similarity between two words is greater than a certain threshold, an edge is established between them. Let the semantic network be G = (V, E), where V is the set of nodes and E is the set of edges. The degree centrality of a node v∈V is defined as: , where deg(v) represents the number of neighboring nodes of node v. For each participant, the degree centrality index of its corresponding node set reflects the importance of the participant's contributions in the discussion within the semantic network. It can be used to quantify the participant's linguistic influence or centrality in the dialogue, thereby aiding in the assessment of their leadership potential.
[0056] In addition, based on the text data characteristics generated by multi-person interaction, we introduced temporal feature indicators to more comprehensively reflect the dynamic characteristics of the discussion process. Specifically, we extracted six core indicators based on the semantic similarity and temporal relationship of the speeches: (1) Novelty, which measures the degree of difference between a new speech and all previous speeches; (2) Jump, which assesses the semantic distance between the current speech and its immediate predecessor, thus reflecting the coherence and leap of thought; (3) Influence, which characterizes the degree of response and continuation of the preceding speech to the subsequent speech by combining semantic similarity with the time decay function; (4) Centroid distance, which represents the degree of difference between the overall semantic vector of an individual speech and the average semantic vector of the group, reflecting the deviation and independence of the individual relative to the group; (5) Directed centrality, which uses network centrality algorithms (such as PageRank) to evaluate the core position of an individual speech in the group semantic flow in a semantic network with speeches as nodes and influence as weighted edges; (6) Number of speeches, which is the amount of contribution of an individual in the discussion. By calculating these temporal features, we can capture semantic content while quantifying the dynamic patterns of the interaction process, providing a richer and more refined representation for predicting individual innovation capabilities.
[0057] For speech data, the system performs frame-level segmentation and feature extraction on the raw speech data, focusing on extracting various acoustic features such as pitch, speaking rate, intonation fluctuations, and energy changes to quantify the emotional tension, expressive initiative, and interactive activity of an individual's speech. At the same time, an acoustic emotion recognition model is introduced to identify the speaker's emotional state at different time periods and calculate indicators such as the frequency, transition frequency, and stability of their emotional expression. These features can help identify individuals with higher participation, emotional infectivity, and proactive expression tendencies, thus providing important supplementary information for innovation ability and leadership inclination. For facial image data, the system uses deep convolutional neural networks to automatically encode and recognize participants' facial expressions based on facial key point information and action unit features extracted from video data. During the analysis, the system extracts the distribution of each participant's emotional state during the interaction and calculates the positiveness index, frequency of expression changes, and stability of expression. In addition, the system combines emotional time series analysis to evaluate the individual's facial response patterns to others' viewpoints during the discussion, thereby capturing their empathy, interaction sensitivity, and emotional regulation tendencies. These dimensions, as important supplementary features of the facial channel, help the multimodal fusion model to more comprehensively evaluate an individual's emotional expression ability, social interaction adaptability, and social influence in innovative discussions.
[0058] Feature fusion model construction is a crucial step in integrating extracted multimodal features. After feature extraction, the system fuses and models the three-modal data to construct a multidimensional assessment model of individual innovation ability and leadership. This model can employ traditional machine learning algorithms or deep neural networks, and improve prediction accuracy through methods such as cross-validation. For example, it can use traditional machine learning algorithms such as support vector machines and random forests, or deep learning models such as convolutional neural networks and recurrent neural networks. Features extracted from text, speech, and facial expressions are used as inputs. Through model training and optimization, model parameters that can accurately assess individual innovation ability and leadership are obtained.
[0059] An evaluation index system is an important basis for measuring an individual's innovation ability and leadership. This system constructs an evaluation index system from multiple dimensions, including the diversity and integration ability of language expression, the emotional tension and initiative of speech, and the positivity and stability of facial expressions. In terms of language expression, the system assesses the originality, flexibility, and fluency of an individual's viewpoints, as well as their coreness and dominance in semantic communication, through indicators such as semantic distance, number of clusters, and degree centrality. In terms of speech, the system measures the emotional tension, initiative, and interactive activity of an individual's speech using acoustic features such as pitch, speaking rate, intonation fluctuations, and energy changes, as well as relevant indicators of emotional state. In terms of facial expressions, the system assesses an individual's emotional expression ability, social interaction adaptability, and social influence in innovative discussions based on indicators such as emotional state distribution, facial expression positivity index, frequency of facial expression changes, and facial expression stability.
[0060] The automated assessment implementation steps outline the specific workflow of the system. First, scenario setup and equipment debugging are conducted to ensure the data collection environment and equipment meet requirements. Next, instructions and exercises are provided to familiarize participants with the task flow. Then, formal testing and data collection are carried out to acquire multimodal data. Following this, data processing and feature extraction are performed to extract key features from the raw data. Next, a feature fusion model is built to integrate the multimodal features. Finally, the model outputs assessment results, including each individual's innovation ability score, leadership inclination index, and a summary of their performance in group collaboration, generating a structured feedback report.
[0061] In terms of specific application scenarios, this system can be widely deployed in educational, enterprise, and other multi-person interactive scenarios. In education, the system is suitable for evaluating student development in secondary school and university general education courses. It can comprehensively assess students' innovative thinking, language expression skills, teamwork level, and emotional regulation abilities. It supports the generation of personalized learning guidance and skills enhancement plans based on assessment results and can predict academic potential, assisting teachers in implementing targeted interventions and feedback in classroom teaching and extracurricular activities. For example, in group discussion assignments, the system can analyze students' diverse viewpoints, language integration abilities, and emotional stability in real time, helping teachers identify strengths and weaknesses and adjust teaching strategies. In enterprise scenarios, this system is suitable for talent recruitment, team capability assessment, and innovation potential evaluation in technology innovation enterprises. It can help enterprises identify core talents with creativity, collaboration skills, and leadership potential, and provide data support for team building and job matching. For example, when recruiting team project leaders, applicants can be organized into group discussions. The system can analyze their innovation ability indicators and leadership tendency index, combined with other assessment information, to accurately select the most suitable candidates. In other multi-person interactive scenarios, this system is applicable to any situation requiring the assessment of an individual's innovation ability, leadership potential, and social adaptability in group collaboration, including team brainstorming, cross-departmental project collaboration, corporate training workshops, academic group discussions, and innovation competitions. The system can quantitatively analyze participants' language expression, ability to integrate viewpoints, emotional regulation, and influence during the interaction process, providing a scientific basis for team building, talent selection, and skills development. Optionally, a speech preprocessing step can be added during speech data acquisition to improve the accuracy of acoustic feature extraction and emotion recognition.
[0062] In summary, the multi-modal data-based multi-person interactive innovation ability assessment system proposed in this invention, through its technical details in system architecture design, data acquisition and processing flow, multimodal feature extraction method, feature fusion model construction, assessment index system, automated assessment implementation steps, and implementation methods for specific application scenarios, achieves accurate identification of individual innovation ability and leadership potential in real interactive scenarios. It provides a more intelligent, efficient, and reliable data-driven solution for educational assessment and talent selection, and has significant social value and promotion potential.
[0063] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. It will be apparent to those skilled in the art that the invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered illustrative and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the scope of the invention. No reference numerals in the claims should be construed as limiting the scope of the claims.
[0064] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A multi-person interactive innovation ability assessment system based on multimodal data, characterized in that, The system includes: System architecture: Built on a real online meeting platform, it collects and analyzes multimodal behavioral data of participants during discussions by designing standardized creative task scenarios; Core functional modules: Scene setup and equipment debugging module: used to guide participants to enter the online meeting room, adjust the camera angle and equipment position, and ensure the quality of data collection; Instructions and Practice Module: Explains the task objectives and operating procedures to participants, and provides example tasks for practice; Testing Phase and Data Acquisition Module: Discuss creative questions within a specified time and collect audio, text, and video data; Data processing and feature extraction module: performs unified organization and multi-level feature extraction on the collected three-modal data, including text semantic analysis, speech feature extraction and facial expression recognition; Analysis, modeling, and evaluation output module: Based on the extracted features, a multidimensional evaluation model of individual innovation ability and leadership is constructed, and the evaluation results and structured feedback reports are output; It integrates natural language processing, speech signal processing and emotion recognition artificial intelligence technologies to achieve multimodal data fusion and automated evaluation.
2. The system according to claim 1, characterized in that, The multimodal data acquisition methods specifically include: Voice data acquisition: Collect participants' voice data using high-quality audio equipment for voice feature analysis; Text data acquisition: Converting speech data into text data using speech-to-text technology for semantic analysis and time-series data analysis; Video data acquisition: Facial video data of participants is acquired through cameras for facial expression and emotion recognition and analysis.
3. The system according to claim 1, characterized in that, The feature fusion model is specifically as follows: A multimodal feature fusion model is constructed, which integrates text semantic and temporal features, speech acoustic features and facial expression features. An evaluation model for individual innovation ability and leadership is built through traditional machine learning algorithms or deep neural networks.
4. The system according to claim 1, characterized in that, The evaluation index system includes: Textual semantic metrics: Based on the Word2Vec word vector model, participants' speech-to-text transcription is mapped to word vectors, and a semantic network is constructed on this basis. The originality of viewpoints is assessed by calculating semantic distance, the distribution of viewpoints is analyzed using hierarchical clustering to quantify the flexibility of thinking, and the fluency of expression is assessed by combining the number of texts. At the same time, the degree centrality of nodes in the semantic network is calculated to predict an individual's leadership in the discussion. Textual temporal metrics: A temporal semantic network is constructed based on a combination of speaking order and semantic similarity. Novelty is used to measure the difference between an individual's speech and previous content; jump rate is used to assess the leap between adjacent speeches; influence is used to characterize the sustained effect of preceding speeches on subsequent speeches; and centroid distance reflects the degree of deviation of an individual from the overall semantic center of the group. Simultaneously, directed centrality is calculated in the semantic network with influence as the weighted edge to capture an individual's position in the group's semantic flow, and its contribution level is measured in conjunction with the number of times it speaks. Speech acoustic metrics: Multidimensional acoustic features are extracted from participants' speech signals, including speech rate (number of syllables per unit time), pitch (fundamental frequency F0), intonation (pitch variation pattern), pause duration and frequency, and energy intensity. Based on these features, emotional tension (through pitch and energy fluctuation amplitude), interactive initiative (through speech proportion and response delay time), and emotional stability (through pitch and energy variance) can be calculated, thereby quantifying an individual's emotional expression and interactive state in the discussion. Facial expression metrics: Utilizing facial action unit (AU) detection and expression recognition algorithms, this study extracts an individual's emotion category (e.g., pleasure, surprise, focus), amplitude of emotion changes (differences in facial expression characteristics across different time periods), and emotional stability (variance or standard deviation of expression changes) during discussions. Based on these metrics, an individual's emotional expression ability, emotion regulation ability, and social adaptability and engagement in team interactions can be quantified.
5. The system according to claim 1, characterized in that, It also includes application scenario extensions, which include: Educational Scenario: Applicable to the assessment of student ability development in secondary education and university general education courses, it can comprehensively evaluate students' innovative thinking, language expression ability, teamwork level and emotional regulation ability; at the same time, it can provide a scientific basis for personalized learning guidance, ability enhancement plan development and academic potential prediction, and support teachers to conduct targeted intervention and feedback in classroom teaching and extracurricular activities. Enterprise Scenario: Suitable for talent recruitment, team capability assessment and innovation potential evaluation for technology innovation enterprises. It can help enterprises identify core talents with creativity, collaboration ability and leadership potential, and provide data support for team building and job matching. Other multi-person interactive scenarios: Applicable to any scenario requiring assessment of an individual's innovation ability, leadership potential, and social adaptability in group collaboration, such as team brainstorming, project collaboration, corporate training workshops, academic group discussions, and innovation competitions. It can quantitatively analyze participants' verbal expression, opinion integration ability, emotional regulation, and influence during interactions, providing a scientific basis for team building, talent selection, and skills development.
6. The system according to claim 2, characterized in that, The voice data acquisition process also includes a preprocessing step for the voice data to improve the accuracy of voice feature extraction.
7. The system according to claim 3, characterized in that, The feature fusion model also includes normalization processing of multimodal features during its construction process to eliminate dimensional differences between different modal data and improve the stability of the evaluation model.
8. The system according to claim 4, characterized in that, The evaluation index system also includes weighting of each index during quantitative evaluation to reflect the relative importance of different indicators in the evaluation of innovation capability and leadership.
9. The system according to claim 5, characterized in that, The expanded application scenarios also include customized development for specific industries and fields to meet the specific needs of different industries for assessing innovation capabilities and leadership.