A method, system and terminal for quantitatively evaluating user cooperation and interaction capability
By designing team tasks and utilizing machine learning algorithms and prior rules to perform multi-stage mapping of linguistic and non-linguistic features, this approach addresses the lack of in-depth interpersonal interaction assessment in existing technologies, enabling accurate quantitative assessment of users' collaboration and interaction abilities.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-08
- Publication Date
- 2026-04-07
AI Technical Summary
Existing technologies lack quantitative assessment methods for deep interpersonal interactions and behavioral patterns among team users, resulting in inaccurate assessment results of user collaboration and interaction capabilities.
Team tasks are designed based on group dynamics. Through machine learning algorithms and prior rules, linguistic features, paralinguistic features and non-linguistic features are mapped in multiple stages to obtain key social behaviors. Based on prior knowledge and observation experience, cumulative scores are calculated to achieve a quantitative assessment of collaboration and interaction abilities.
It enables an objective and scientific quantitative assessment of users' collaboration and interaction capabilities, overcoming inherent limitations such as response bias, cultural prejudice, and impression management, and providing more accurate assessment results.
Smart Images

Figure CN121258339B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a method, system, and terminal for quantitatively evaluating user collaboration and interaction capabilities. Background Technology
[0002] In assessing interpersonal interaction and teamwork skills, traditional methods primarily rely on psychometric questionnaires. For example, interpersonal behavior scales assess traits in interpersonal relationships by having test takers answer a series of questions. These scales aim to help individuals understand their interpersonal behavior patterns.
[0003] However, such methods have significant limitations: First, their results are highly dependent on the test-taker's subjective perception and honesty, making it impossible to ensure the authenticity of the data. Second, the scale assessment is conducted at a static point in time, failing to capture, in real-time and dynamically, the test-taker's behavioral changes in real-world interactive situations, such as emotional fluctuations and adjustments in communication strategies. This assessment method is severely out of touch with the dynamic and complex team collaboration environment of the real world.
[0004] Currently, there are numerous digital tools and platforms on the market for team collaboration. These tools are designed to facilitate collaboration rather than evaluate it. They primarily focus on task and information flow management, lacking the ability to quantitatively assess the deep interpersonal interactions and behavioral patterns among team members. The essence of interpersonal interaction and team collaboration lies not only in task completion but also in complex and dynamic social behaviors such as nonverbal communication, emotional synchronization, role division, and conflict resolution.
[0005] Therefore, existing technologies still need to be improved and developed. Summary of the Invention
[0006] The main objective of this invention is to provide a method, system, terminal, and computer-readable storage medium for quantitatively evaluating users' collaboration and interaction capabilities. This aims to solve the problem that the lack of a quantitative evaluation method for deep interpersonal interactions and behavioral patterns among team users in the prior art leads to inaccurate evaluation results of users' collaboration and interaction capabilities.
[0007] To achieve the above objectives, the present invention provides a method for quantitatively evaluating user collaboration and interaction capabilities, the method comprising the following steps:
[0008] Team tasks were designed based on group dynamics, and multiple users were selected to participate in the team tasks. The linguistic features, paralinguistic features, and nonlinguistic features of the multiple users at different stages of the team tasks were obtained.
[0009] The key social behaviors are obtained by performing a one-stage mapping of the linguistic features, paralinguistic features, and nonlinguistic features through machine learning algorithms and prior rules.
[0010] Based on prior knowledge and observational experience, the key social behaviors are mapped in two stages to obtain mapping scores under multiple attribute windows. All the mapping scores are then accumulated to obtain a quantitative evaluation result of the collaboration and interaction capabilities of all the users.
[0011] Optionally, the quantitative evaluation method for user collaboration and interaction capabilities includes a team task comprising an observation phase, a personal description phase, a free discussion phase, and a representative summary phase.
[0012] The linguistic features include speech quality, participation, and number of interruptions; the paralinguistic features include the fundamental frequency and loudness of the audio data; and the nonlinguistic features include human posture features, gesture features, head posture features, and facial expression features.
[0013] Optionally, the quantitative evaluation method for user collaboration and interaction capabilities, wherein obtaining the linguistic features, paralinguistic features, and nonlinguistic features of multiple users at different stages of the team task specifically includes:
[0014] Video data of different stages of the team's task is acquired by a preset number of vertically placed cameras;
[0015] The video data is analyzed frame by frame using a human pose estimation framework to extract key skeletal points of the human body. Based on the key skeletal points, the limb movement trajectory and posture changes of each user are characterized. The limb movement trajectory and posture changes are modeled in combination with time series to obtain the human pose features and gesture features of each user.
[0016] The 6D pose estimation algorithm is used to estimate the 6-DOF head pose of the faces in the video data to obtain the attention direction, listening attitude and object orientation of each user. Based on the attention direction, listening attitude and object orientation, the head pose features of each user are analyzed.
[0017] By using cascaded multi-layer convolution and attention mechanisms, the correspondence between local facial muscle movements and overall facial expression structure is captured from the video data. Based on the correspondence, temporal modeling is performed in conjunction with facial key point features to obtain the facial expression features of each user.
[0018] The system collects facial features of each user from multiple perspectives and simultaneously acquires audio and video stream data from three cameras.
[0019] Based on the aforementioned frontal facial features, the frontal facial regions of each user are located and tracked in real time using face detection and facial key point detection algorithms to obtain the dynamic key point coordinates of each user's lips. The dynamic key point coordinates of the lips are used as the lip movement sequence to obtain the lip movement sequence of each user. The lip movement sequence is then combined with the speaker recognition model to obtain the speaker label for each frame in the audio and video stream data.
[0020] Speech activity detection is performed on the audio and video stream data to segment out effective audio segments. The effective audio segments are then matched with the speaker tags across modalities using an audio and video timestamp alignment strategy to obtain the mapping relationship between audio segments and speakers.
[0021] Based on the mapping relationship, open-source audio analysis tools and source speech-to-text models are used to obtain the speaking quality, participation, number of interruptions, fundamental frequency of audio data, and loudness of audio data for each user.
[0022] Optionally, the quantitative evaluation method for user collaboration and interaction capabilities includes key social behaviors such as: dominant speaking behavior, interrupting speaking behavior, summary statement level, gesture category, interaction behavior, and emotional state.
[0023] Optionally, the quantitative evaluation method for user collaboration and interaction capabilities, wherein the step of performing a one-stage mapping of the linguistic features, paralinguistic features, and nonlinguistic features through machine learning algorithms and prior rules to obtain key social behaviors specifically includes:
[0024] The dominant speaking behavior and interruption behavior of the corresponding user are obtained based on the participation level and the number of interruption behaviors;
[0025] Obtain the speaking quality, audio data fundamental frequency, and audio data loudness of each user during the representative summary phase, and obtain the summary statement level of each user based on the speaking quality, audio data fundamental frequency, and audio data loudness of the representative users;
[0026] Define multiple preset gesture categories, and perform similarity threshold matching between the gesture features and the multiple preset gesture categories to obtain the gesture category of the corresponding user;
[0027] The angle between the line connecting the head posture features and the coordinate points of each user's head is calculated, and the interaction behavior of each user is obtained based on the angle. The interaction behavior includes head-down behavior and eye contact interaction behavior.
[0028] Based on the facial expression features, an emotion change curve for each user is plotted, and the emotional state of each user is obtained based on the emotion change curve. The emotional state includes the percentage of emotion and the degree of emotional fluctuation.
[0029] Optionally, the quantitative evaluation method for user collaboration and interaction capabilities, wherein the key social behaviors are mapped in a two-stage manner based on prior knowledge and observational experience to obtain mapping scores under multiple attribute windows, specifically includes:
[0030] The system obtains the number of speaking frames for each user during personal summaries, calculates the speaking duration and number of summary segments based on the number of speaking frames, determines the initial leadership score for the corresponding user based on the speaking duration, the number of summary segments, the speaking quality, and the dominant speaking behavior, and supplements the initial leadership score based on the summary statement level to obtain the target leadership score for the corresponding user.
[0031] Based on the interruption behavior and the interaction behavior, a score is determined on the user's listening attitude when others are speaking.
[0032] By using preset rule thresholds, the shoulder tension, posture openness, and group convergence of the corresponding user are calculated based on the human posture characteristics. The overall posture score is obtained by fusing the shoulder tension, posture openness, and group convergence.
[0033] Obtain the human posture estimation coordinate vector of the corresponding user, calculate the gesture kinetic energy density of the corresponding user based on the human posture estimation coordinate vector, obtain positive and negative indicators from the gesture category, and add the positive indicators, the negative indicators and the gesture kinetic energy density to obtain the nonverbal leadership potential score.
[0034] Each user is assigned the same initial emotion score. If the degree of emotion fluctuation exceeds a first preset threshold or the proportion of neutral emotions in the emotion proportion exceeds a second preset threshold, the initial emotion score is modified to obtain the target emotion score for the corresponding user.
[0035] Optionally, the method for quantitatively evaluating user collaboration and interaction capabilities, wherein accumulating all the mapping scores to obtain a quantitative evaluation result of the collaboration and interaction capabilities of all users specifically includes:
[0036] The target leadership score, the listening attitude score, the overall posture score, the nonverbal leadership potential score, and the target emotional score are normalized.
[0037] The normalized target leadership score, listening attitude score, overall posture score, nonverbal leadership potential score, and target emotion score are weighted and accumulated to obtain a quantitative assessment result of the collaboration and interaction ability of each user.
[0038] Furthermore, to achieve the above objectives, the present invention also provides a quantitative evaluation system for user collaboration and interaction capabilities, wherein the quantitative evaluation system for user collaboration and interaction capabilities includes:
[0039] The feature acquisition module is used to design team tasks based on group dynamics, select multiple users to participate in the team tasks, and acquire the linguistic features, paralinguistic features, and nonlinguistic features of the multiple users at different stages of the team tasks.
[0040] The first mapping module is used to perform a one-stage mapping of the language features, paralinguistic features and nonlinguistic features through machine learning algorithms and prior rules to obtain key social behaviors.
[0041] The second mapping module is used to perform a two-stage mapping of the key social behaviors based on prior knowledge and observation experience, obtain mapping scores under multiple attribute windows, and accumulate all the mapping scores to obtain a quantitative evaluation result of the collaboration and interaction capabilities of all the users.
[0042] Furthermore, to achieve the above objectives, the present invention also provides a terminal, wherein the terminal includes: a memory, a processor, and a quantitative evaluation program for user collaboration and interaction capabilities stored in the memory and executable on the processor, wherein when the quantitative evaluation program for user collaboration and interaction capabilities is executed by the processor, it implements the steps of the quantitative evaluation method for user collaboration and interaction capabilities as described above.
[0043] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a quantitative evaluation program for user collaboration and interaction capabilities, and when the quantitative evaluation program for user collaboration and interaction capabilities is executed by a processor, it implements the steps of the quantitative evaluation method for user collaboration and interaction capabilities as described above.
[0044] In this invention, team tasks are designed based on group dynamics. Multiple users are selected to participate in the team tasks, and the linguistic, paralinguistic, and nonverbal characteristics of these users at different stages of the tasks are obtained. A first-stage mapping of these linguistic, paralinguistic, and nonverbal characteristics is performed using machine learning algorithms and prior rules to obtain key social behaviors. Based on prior knowledge and observational experience, these key social behaviors are then mapped in a second stage to obtain mapping scores under multiple attribute windows. All these mapping scores are accumulated to obtain a quantitative evaluation result of the collaboration and interaction abilities of all users. This invention overcomes inherent limitations such as response bias, cultural bias, and impression management, achieving an objective and scientific quantitative evaluation of users' collaboration and interaction abilities. Attached Figure Description
[0045] Figure 1 This is a flowchart of a preferred embodiment of the quantitative evaluation method for user collaboration and interaction capabilities of the present invention;
[0046] Figure 2 This is a schematic diagram illustrating the extraction of social behavior features from the quantitative evaluation method for user collaboration and interaction capabilities of this invention.
[0047] Figure 3 This is a schematic diagram illustrating the quantification of social behavior in the quantitative evaluation method for user collaboration and interaction capabilities of this invention.
[0048] Figure 4 This is a structural diagram of a preferred embodiment of the quantitative evaluation system for user collaboration and interaction capabilities of the present invention;
[0049] Figure 5 This is a structural diagram of a preferred embodiment of the terminal of the present invention. Detailed Implementation
[0050] This application provides a method, system, and terminal for quantitatively evaluating user collaboration and interaction capabilities. To make the purpose, technical solution, and effects of this application clearer and more explicit, the following detailed description is provided with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only for explaining this application and are not intended to limit this application.
[0051] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have the same meaning as in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless specifically defined as herein.
[0052] Furthermore, if the embodiments of this invention involve descriptions such as "first" or "second," these descriptions are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined with "first" or "second" may explicitly or implicitly include at least one of those features. Additionally, the technical solutions of the various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. If the combination of technical solutions is contradictory or impossible to implement, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed by this invention.
[0053] The preferred embodiment of the present invention describes a method for quantitatively evaluating user collaboration and interaction capabilities, such as... Figure 1 As shown, the quantitative evaluation method for user collaboration and interaction capabilities includes the following steps:
[0054] Step S10: Design a team task based on group dynamics, select multiple users to participate in the team task, and obtain the language features, paralinguistic features, and nonlinguistic features of the multiple users at different stages of the team task.
[0055] In this embodiment, based on the framework of group dynamics theory, the present invention first designs a team task under stress. This team task is designed with customized task specifications and revised rules. Based on the theme of cracking a secret puzzle, it is made suitable for teams of 6-8 people. A team activity is developed to simulate a scenario in a joint operation where a drone swarm conducts multi-angle reconnaissance of a target area behind enemy lines, and the image data link is fragmented due to electromagnetic interference. New recruits are required to integrate fragmented intelligence to reconstruct the complete dynamic chain of the battlefield, providing the command with accurate tactical decision-making support. The activity process is as follows: The first step is the intelligence reception stage, where multiple sets of pictures with temporal, spatial, or interpersonal logical relationships are prepared as task materials; the second step is the independent reconnaissance analysis and tactical briefing stage, where each participant randomly selects a single image from different picture sets and, under strict 30-second independent observation time limits and without the option to review or display the image, gives a personal presentation describing the content of the image. The third step is the battlefield situation simulation phase, where users can freely discuss and exchange information one-on-one or in small groups. Ultimately, through team consultation, they determine the correct order of all images and decide which piece of information is false and will be eliminated. The accuracy of the final decision and ordering is used as one of the core performance evaluation indicators. This activity, through information fragmentation, timed communication, and collective decision-making, systematically examines participants' information processing, verbal expression, and teamwork abilities.
[0056] The activity requires participants to function in a high-intensity communication environment, involving various cognitive abilities such as short-term memory, verbal description, proactive communication, listening and interpretation, reasoning and construction, and team decision-making, as well as leadership and organizational skills. Participants need to successfully complete various tasks within this setting. During the debriefing process, participants often discover that even small actions or decisions can lead to drastically different outcomes. Furthermore, this process helps to observe key personality traits such as initiative, responsibility, empathy, and adaptability. Microphones and cameras are used to capture multimodal data of the team activities, including audio, visual, and natural language, estimating participants' head posture, visual attention focus, nonverbal communication patterns, and discussion content, thereby enabling the analysis and evaluation of interpersonal interaction styles and collaborative abilities.
[0057] In the discussion scenario of cracking the secret puzzle, a camera is placed in each of the four vertical directions to ensure that the key facial areas of each person can be captured during the discussion; at the same time, the camera will capture the audio data in the scene in order to analyze the players' voice signals later.
[0058] Furthermore, the team task includes an observation phase, a personal description phase, a free discussion phase, and a representative summary phase.
[0059] Because the entire process is recorded with video and audio from four perspectives, the long video needs to be divided into multiple segments based on the task progress. Then, a group behavior analysis is performed on each segment based on its characteristics. Therefore, this analysis process divides the long video into multiple task segments according to the different stages of the team task. (1) Observation phase: All users listen to the rules and receive a card envelope. They open the envelope and observe the card. During this process, attention should be paid to each user's facial expressions, head posture, and questioning language. (2) Personal description phase: Each user takes turns describing the content of the picture on their card while other users listen. During this process, attention should be paid to each speaker's paralinguistic features, body posture and gestures, head posture, and the listener's head posture, body posture and gestures, as well as the tendency to interrupt. (3) Free discussion phase: All users freely discuss the order of the card pictures. During this process, attention should be paid to each user's enthusiasm for speaking, the speaker's paralinguistic features, body posture and gestures, and head posture, and the listener's body posture, gestures, and head posture. (4) Representative summary phase: Users select a representative to summarize the results of the discussion on the order of the pictures. During this process, attention should be paid to the representative's body posture and gestures, head posture, and paralinguistic features, as well as the response features of others' body posture and gestures, head posture, and supplementary language.
[0060] Furthermore, the language features described in this invention include speech quality, participation, and number of interruptions; the sub-language features include the fundamental frequency and loudness of audio data; and the non-language features include human posture features, gesture features, head posture features, and facial expression features.
[0061] like Figure 2 As shown, further, the acquisition of the linguistic features, paralinguistic features, and nonlinguistic features of multiple users at different stages of the team task specifically includes:
[0062] Video data of different stages of the team's task is acquired by a preset number of vertically placed cameras;
[0063] The video data is analyzed frame by frame using a human pose estimation framework to extract key skeletal points of the human body. Based on the key skeletal points, the limb movement trajectory and posture changes of each user are characterized. The limb movement trajectory and posture changes are modeled in combination with time series to obtain the human pose features and gesture features of each user.
[0064] The 6D pose estimation algorithm is used to estimate the 6-DOF head pose of the faces in the video data to obtain the attention direction, listening attitude and object orientation of each user. Based on the attention direction, listening attitude and object orientation, the head pose features of each user are analyzed.
[0065] By using cascaded multi-layer convolution and attention mechanisms, the correspondence between local facial muscle movements and overall facial expression structure is captured from the video data. Based on the correspondence, temporal modeling is performed in conjunction with facial key point features to obtain the facial expression features of each user.
[0066] The system collects facial features of each user from multiple perspectives and simultaneously acquires audio and video stream data from three cameras.
[0067] Based on the aforementioned frontal facial features, the frontal facial regions of each user are located and tracked in real time using face detection and facial key point detection algorithms to obtain the dynamic key point coordinates of each user's lips. The dynamic key point coordinates of the lips are used as the lip movement sequence to obtain the lip movement sequence of each user. The lip movement sequence is then combined with the speaker recognition model to obtain the speaker label for each frame in the audio and video stream data.
[0068] Speech activity detection is performed on the audio and video stream data to segment out effective audio segments. The effective audio segments are then matched with the speaker tags across modalities using an audio and video timestamp alignment strategy to obtain the mapping relationship between audio segments and speakers.
[0069] Based on the mapping relationship, open-source audio analysis tools and source speech-to-text models are used to obtain the speaking quality, participation, number of interruptions, fundamental frequency of audio data, and loudness of audio data for each user.
[0070] Understandably, in human posture and gesture recognition schemes, to capture user body movements and gesture changes during communication, this embodiment employs the MediaPipe Pose framework to analyze video data frame by frame, extracting 33 key skeletal points of the human body in real time. These points include major motion areas such as the shoulders, elbows, wrists, torso, and lower limbs, enabling precise depiction of an individual's limb movement trajectory and posture changes. Combined with time-series modeling, further analysis of non-verbal signals related to social interaction, such as pointing, outstretched gestures, and crossed arms, can support the assessment of task collaboration status and role participation levels.
[0071] In the head pose estimation scheme, a 6D pose estimation algorithm, namely 6DRepNet (6D Pose Regression Network), is introduced to perform 6-DOF head pose estimation on the face in each frame of the image. This approach has good robustness under different lighting and occlusion conditions and can be used to determine an individual's attention direction, listening attitude and target, thereby analyzing the attention distribution and potential interaction intentions among users in group interaction scenarios.
[0072] In this facial expression recognition scheme, to identify an individual's emotional expression during communication, this embodiment employs a Dual Attention Network (DAN) to classify facial expressions. DAN captures the relationship between local facial muscle movements and the overall facial structure through cascaded multi-layer convolutions and attention mechanisms, enabling the recognition of basic emotion categories (such as happiness, anger, surprise, and disgust). By combining this with temporal modeling, it is possible to further reveal the dynamic changes in emotions among users and their impact on the collaborative process.
[0073] In the language feature extraction scheme, the task is to extract the language features of each user in a multi-person group task. The challenge is to identify the speaker from a shared microphone array recording multiple users and obtain the speaker's audio features and further textual semantic features.
[0074] Specifically, to determine the user to whom the audio belongs, the speaker recognition scheme is set as follows: Based on a multimodal data fusion scheme, the system first simultaneously acquires audio and video stream data from eight users (with fixed positional labels) using three cameras. For video processing, the features closest to the frontal face of all users are selected from multiple perspectives. YOLO object detection algorithm, face detection, and MediaPipe Face algorithm are used to locate and track the frontal face region of each user in real time. The 21-dimensional dynamic keypoint coordinates of the lips are extracted as temporal features through mouth movement detection. Based on the aligned lip movement sequence and a matching threshold is set according to the temporal rules of mouth movement, the speaker is detected. By combining this with a temporal lip movement speaker recognition model, speaker identity is determined, and the speaker label and confidence score for each frame in the video stream are output.
[0075] To further extract the linguistic and paralinguistic features for each user, the audio processing scheme is as follows: The audio processing end uses Voice Activity Detection (VAD) to detect speech intervals in the original audio to locate valid segments containing speech. Using an audio-video timestamp alignment strategy, the audio segments segmented by VAD are matched cross-modally with the speaker labels output from the video end to establish an "audio segment-speaker" mapping relationship. For the attributed, clean single-person speech, open-source audio analysis tools are used to extract audio parameters such as fundamental frequency and loudness from each speech segment. The Whisper speech recognition model is then used to convert these parameters into a timestamped text feature sequence, further enriching the expression of linguistic information, ultimately yielding linguistic features (text) and paralinguistic features (fundamental frequency, loudness).
[0076] As can be seen, this invention divides social behavior characteristics into three parts: linguistic features, paralinguistic features, and non-linguistic features, and extracts social behavior representations based on this. At the non-linguistic feature level, it extracts members' body posture, gestures, facial expressions, and head posture to analyze the behavioral interaction patterns of each member with other members from a visual perspective; at the paralinguistic feature level, it extracts speech tone, energy, etc., to analyze the degree of speech participation of each member; at the linguistic feature level, it extracts the basic semantic attributes of the language text to analyze the language interaction patterns of each member.
[0077] Step S20: Perform a one-stage mapping of the language features, paralinguistic features, and nonlinguistic features using machine learning algorithms and prior rules to obtain key social behaviors.
[0078] like Figure 3 As shown, the key social behaviors include: dominant speaking behavior, interrupting speaking behavior, summary statement level, gesture category, interaction behavior, and emotional state.
[0079] The process of performing a one-stage mapping of the linguistic features, paralinguistic features, and nonlinguistic features using machine learning algorithms and prior rules to obtain key social behaviors specifically includes:
[0080] The dominant speaking behavior and interruption behavior of the corresponding user are obtained based on the participation level and the number of interruption behaviors;
[0081] Obtain the speaking quality, audio data fundamental frequency, and audio data loudness of each user during the representative summary phase, and obtain the summary statement level of each user based on the speaking quality, audio data fundamental frequency, and audio data loudness of the representative users;
[0082] Define multiple preset gesture categories, and perform similarity threshold matching between the gesture features and the multiple preset gesture categories to obtain the gesture category of the corresponding user;
[0083] The angle between the line connecting the head posture features and the coordinate points of each user's head is calculated, and the interaction behavior of each user is obtained based on the angle. The interaction behavior includes head-down behavior and eye contact interaction behavior.
[0084] Based on the facial expression features, an emotion change curve for each user is plotted, and the emotional state of each user is obtained based on the emotion change curve. The emotional state includes the percentage of emotion and the degree of emotional fluctuation.
[0085] Understandably, to improve the interpretability of the task, this invention designs a two-stage mapping scheme. First, social behavioral features are mapped to key social behaviors using machine learning methods. Then, these key social behaviors are mapped to interpersonal interaction, teamwork attributes, and scores. Specifically, the first mapping utilizes machine learning and prior rules to analyze user behavior responses based on social behavioral features (human posture, gestures, facial expressions, head posture, tone of voice, semantics, etc.). For example, posture estimation and gestures are used to determine whether the user is gesturing wildly, pointing, issuing commands, or refusing to communicate; facial expressions are used to determine their emotional state; voice features are used to determine whether the user interrupts others, dominates discussions for extended periods, or is successfully interrupted; and head posture is used to determine their attention span and interaction initiation and response. This comprehensive analysis identifies which key social behavior the user exhibits in each window, thereby establishing a feature-to-behavior mapping relationship.
[0086] Specifically, regarding the mapping of dominant speaking behavior, in multi-person discussion segments, the speaking time of each player is obtained by counting the number of speaking frames of each speaker. This indicator reflects the player's participation level, whether they "try to understand other people's ideas to solve problems together", whether they "understand the opinions of other teammates in a timely manner", and whether they "actively participate and give encouragement or agreement in a timely manner". The duration is sorted and players are scored from highest to lowest as 6, 5, ..., 1 points, which serve as the participation score for multi-person (taking 6 people as an example) discussions. The participation score is used as the standard for dominant speaking behavior.
[0087] The mapping of interrupting speech implies that listeners other than the speaker should remain silent while others are speaking. Speech recognition is used to detect overlapping speech; if this occurs, it indicates that some listeners have broken the established rules, and the user is deemed to have interrupted.
[0088] For the mapping of summary presentation levels, during each person's speech, speaker recognition and speech recognition are used to obtain the content and duration of each user's speech in turn. The speaker's volume can also be obtained by calculating the amplitude or decibel value of the audio signal. After obtaining the content and duration of the speech, indicators such as descriptive text length, duration, and speech rate (speech text length divided by speech duration) are extracted from language features, and the speaker's level of detail in the description is evaluated based on a weighted average of these indicators. The quality of a speaker's presentation is measured by combining the detail and volume of their speech. The quality of the presentation is measured by whether the speaker is "willing to provide the most accurate information to other teammates and exchange information and ideas with them," and whether they "understand the meaning of the information obtained." By evaluating the presentation sessions, the quality of the presentations can be ranked from highest to lowest, and the corresponding speakers can be scored 6, 5, ..., 1 points, which serve as the speaker's score.
[0089] For the mapping of gesture categories, this embodiment defines multiple categories of gestures. Gesture recognition is obtained by matching human postures with corresponding gesture examples using a similarity threshold, including positive categories. "Indicative" gestures: extending the index finger to point at others; "Emphasis" gestures: waving the hand rapidly up and down; "Control" gestures: pressing down or raising the palm; "Cooperation" gestures: spreading both hands towards others; "Confrontation" gestures: crossing arms or pushing with the palms; Negative categories. "Closed type": arms crossed or hands pushing against the surface; "Hidden type": hands under the table.
[0090] For mapping interactive behaviors, for speakers identified through speaker recognition, it's necessary to focus on the speaker's eye contact with other users. This is achieved by analyzing head posture to determine the proportion of non-interactive behaviors (such as looking down) versus interactive behaviors (such as eye contact with other users, where the angle between the speaker's head posture and that of other users is less than a threshold, indicating interaction). This yields a score representing the speaker's eye contact time percentage, which can reflect their social influence, dominance, and integration to some extent. In the four stages, for other non-speaking listeners, it's necessary to assess their level of attention to the speaker, reflecting their social interaction and participation. This is primarily done by detecting the angle between all listeners' head postures and the speaker's; angles less than a threshold indicate attention, and the proportion of attention time is calculated. Additionally, for listeners, nodding can be identified using a window threshold, reflecting their agreement with the speaker and indirectly indicating their social participation. A weighted sum of these nodding actions, along with the interaction levels during the speaking and listening processes, can reflect interpersonal interaction and collaborative attributes to some extent.
[0091] For mapping emotional states, the emotional changes of each individual can be assessed in all four stages, thus indirectly reflecting the participants' attitudes, participation, and seriousness towards the group task. Furthermore, by quantifying emotional change curves, emotional percentages, and emotional fluctuations, the characteristics of participants can be refined to a certain extent, which helps to map them to interpersonal interaction and collaborative attributes. Specifically, emotions are mainly realized through facial expression recognition. The YOLO object detection algorithm detects the human body, the Mediapipe library detects faces, and the DAN network extracts expression categories. Each person's facial expression is detected, resulting in an emotional change curve, which is then refined into various indicators.
[0092] The quantitative evaluation indicators are as follows: Emotion change curve: plot the emotion change curve of each person to show the trend of emotion change over time. Currently, emotions are divided into three categories, which are positive, neutral and negative emotions. An emotion curve evaluation is carried out at each stage.
[0093] Step S30: Based on prior knowledge and observation experience, perform a two-stage mapping of the key social behaviors to obtain mapping scores under multiple attribute windows, and accumulate all the mapping scores to obtain a quantitative evaluation result of the collaboration and interaction capabilities of all the users.
[0094] Based on prior knowledge and observational experience, the key social behaviors are mapped in a two-stage manner to obtain mapping scores under multiple attribute windows, specifically including:
[0095] The system obtains the number of speaking frames for each user during personal summaries, calculates the speaking duration and number of summary segments based on the number of speaking frames, determines the initial leadership score for the corresponding user based on the speaking duration, the number of summary segments, the speaking quality, and the dominant speaking behavior, and supplements the initial leadership score based on the summary statement level to obtain the target leadership score for the corresponding user.
[0096] Based on the interruption behavior and the interaction behavior, a score is determined on the user's listening attitude when others are speaking.
[0097] By using preset rule thresholds, the shoulder tension, posture openness, and group convergence of the corresponding user are calculated based on the human posture characteristics. The overall posture score is obtained by fusing the shoulder tension, posture openness, and group convergence.
[0098] Obtain the human posture estimation coordinate vector of the corresponding user, calculate the gesture kinetic energy density of the corresponding user based on the human posture estimation coordinate vector, obtain positive and negative indicators from the gesture category, and add the positive indicators, the negative indicators and the gesture kinetic energy density to obtain the nonverbal leadership potential score.
[0099] Each user is assigned the same initial emotion score. If the degree of emotion fluctuation exceeds a first preset threshold or the proportion of neutral emotions in the emotion proportion exceeds a second preset threshold, the initial emotion score is modified to obtain the target emotion score for the corresponding user.
[0100] It is understood that the mapping score includes target leadership score, listening attitude score, overall posture score, nonverbal leadership potential score, and target emotion score.
[0101] The target leadership score is obtained as follows: The speaking duration is determined by counting the number of frames spoken during each individual's summary. The number of summary segments completed by each person is also counted. This indicator reflects the player's leadership and information gathering abilities, whether they "consult and adopt suggestions from other teammates to make a summary" and "lead and motivate the team at appropriate times to achieve high-quality performance." The text of the summary is obtained through speech recognition, and the volume of the audio signal is calculated to comprehensively obtain the speaking quality during the individual summary (similar to the personal statement stage, evaluating the speaker's speaking quality). The player's individual summary ability is evaluated based on the summarization duration, number of summary segments, and speaking quality, and players are scored sequentially from highest to lowest as 6, 5, ..., 1 points, serving as their initial leadership score. Simultaneously, the initial leadership score is supplemented based on the summary statement level to obtain the corresponding user's target leadership score.
[0102] The listening attitude score is obtained as follows: A user's listening attitude score is determined based on interruption behavior and interaction behavior. First, lip movement detection is used to identify the speaker. After excluding speakers whose turn it is, listeners who interrupt are penalized 3 points, while other listeners are penalized 1 point for interruption. The reason for emphasizing interruption behavior in personal statements is that it seriously deviates from the excellent teamwork qualities of "maintaining attentive listening to others' opinions and views," "caring about other team members' feelings," and "respecting other teammates' opinions and views." Next, since the interaction behavior includes looking down and eye contact, looking down deducts points from the listening attitude score, while eye contact adds points.
[0103] The overall posture score acquisition scheme is as follows: This embodiment proposes a set of posture analysis indicators and conducts specific assessments and adjustments for members at relevant stages. Specifically, three body posture features are defined, and the results are calculated by extracting human postures and applying rule thresholds. Based on the human posture features, the shoulder tension of the corresponding user is calculated. Openness of posture and group convergence ,in, The vector angle of the shoulder and elbow (the value is negative when tense or inclined to protect oneself). The ratio of the widest distance between the upper limbs to the shoulder width, when ,but When it is 0, Before, then The value is 1. The overall posture score is obtained by integrating the shoulder tension, posture openness, and group convergence.
[0104] The scheme for obtaining the nonverbal leadership potential score is as follows: obtain the human posture estimation coordinate vector of the corresponding user, and calculate the gesture kinetic energy density of the corresponding user based on the human posture estimation coordinate vector. : ;in, , , , These represent the x, y, and vertices of the human pose estimation coordinate vector, respectively. For time span (number of frames).
[0105] Furthermore, positive (leadership) indicators and negative (leadership) indicators are obtained from the aforementioned gesture categories: ; ;in, Indicates a positive hand gesture In hand posture The proportion of positive indicators. Indicates a negative hand posture In hand posture The proportion of negative indicators. It represents the time span within a basic unit.
[0106] Finally, the positive indicators, the negative indicators, and the gesture kinetic energy density are added together to obtain the nonverbal leadership potential energy. The nonverbal leadership potential energy is then converted according to a preset scoring rule to obtain a nonverbal leadership potential energy score.
[0107] The method for obtaining the target emotion score is as follows: First, set the same initial emotion score for each user. Then, calculate the average value of the absolute value of the emotion change between the two frames before and after the whole stage to evaluate the degree of emotion fluctuation. If the degree of emotion fluctuation exceeds the first preset threshold or the proportion of neutral emotion in the emotion proportion exceeds the second preset threshold, then modify the initial emotion score to obtain the target emotion score for the corresponding user.
[0108] Furthermore, the step of accumulating all the mapping scores to obtain a quantitative evaluation result of the collaboration and interaction capabilities of all the users specifically includes:
[0109] The target leadership score, the listening attitude score, the overall posture score, the nonverbal leadership potential score, and the target emotional score are normalized.
[0110] The normalized target leadership score, listening attitude score, overall posture score, nonverbal leadership potential score, and target emotion score are weighted and accumulated to obtain a quantitative assessment result of the collaboration and interaction ability of each user.
[0111] In this embodiment, the quantitative evaluation results of the above four aspects—language features, human posture and gestures, head posture, and facial expressions—can be weighted according to the actual situation and mapped to the corresponding behavioral scores as an evaluation of interpersonal interaction and collaboration attributes, thus obtaining qualitative evaluation analysis results.
[0112] Furthermore, such as Figure 4 As shown, based on the above-mentioned quantitative evaluation method for user collaboration and interaction capabilities, the present invention also provides a quantitative evaluation system for user collaboration and interaction capabilities, wherein the quantitative evaluation system for user collaboration and interaction capabilities includes:
[0113] The feature acquisition module 51 is used to design team tasks based on group dynamics, select multiple users to participate in the team tasks, and acquire the language features, paralinguistic features, and nonlinguistic features of the multiple users at different stages of the team tasks.
[0114] The first mapping module 52 is used to perform a one-stage mapping of the language features, paralinguistic features and nonlinguistic features through machine learning algorithms and prior rules to obtain key social behaviors.
[0115] The second mapping module 53 is used to perform a two-stage mapping of the key social behaviors based on prior knowledge and observation experience, obtain mapping scores under multiple attribute windows, and accumulate all the mapping scores to obtain a quantitative evaluation result of the collaboration and interaction capabilities of all the users.
[0116] Furthermore, such as Figure 5 As shown, based on the above-mentioned quantitative evaluation method and system for user collaboration and interaction capabilities, the present invention also provides a terminal, which includes a processor 10, a memory 20 and a display 30. Figure 5 Only some of the terminal components are shown; however, it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.
[0117] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as a hard drive or memory. In other embodiments, the memory 20 may be an external storage device of the terminal, such as a plug-in hard drive, smart media card (SMC), secure digital card (SD), flash card, etc. Further, the memory 20 may include both internal and external storage devices. The memory 20 is used to store application software and various types of data installed on the terminal, such as program code installed on the terminal. The memory 20 can also be used to temporarily store data that has been output or will be output. In one embodiment, the memory 20 stores a quantitative evaluation program 40 for user collaboration and interaction capabilities, which can be executed by the processor 10 to implement the quantitative evaluation method for user collaboration and interaction capabilities in this application.
[0118] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor, or other data processing chip, used to run program code stored in the memory 20 or process data, such as executing the quantitative evaluation method for user collaboration and interaction capabilities.
[0119] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen. The display 30 is used to display information on the terminal and to display a visual user interface. The components of the terminal communicate with each other via a system bus.
[0120] In one embodiment, when the processor 10 executes the quantitative evaluation program 40 for user collaboration and interaction capabilities stored in the memory 20, the following steps are performed:
[0121] Team tasks were designed based on group dynamics, and multiple users were selected to participate in the team tasks. The linguistic features, paralinguistic features, and nonlinguistic features of the multiple users at different stages of the team tasks were obtained.
[0122] The key social behaviors are obtained by performing a one-stage mapping of the linguistic features, paralinguistic features, and nonlinguistic features through machine learning algorithms and prior rules.
[0123] Based on prior knowledge and observational experience, the key social behaviors are mapped in two stages to obtain mapping scores under multiple attribute windows. All the mapping scores are then accumulated to obtain a quantitative evaluation result of the collaboration and interaction capabilities of all the users.
[0124] The team tasks include an observation phase, a personal description phase, a free discussion phase, and a representative summary phase.
[0125] The linguistic features include speech quality, participation, and number of interruptions; the paralinguistic features include the fundamental frequency and loudness of the audio data; and the nonlinguistic features include human posture features, gesture features, head posture features, and facial expression features.
[0126] Specifically, obtaining the linguistic features, paralinguistic features, and nonlinguistic features of multiple users at different stages of the team task includes:
[0127] Video data of different stages of the team's task is acquired by a preset number of vertically placed cameras;
[0128] The video data is analyzed frame by frame using a human pose estimation framework to extract key skeletal points of the human body. Based on the key skeletal points, the limb movement trajectory and posture changes of each user are characterized. The limb movement trajectory and posture changes are modeled in combination with time series to obtain the human pose features and gesture features of each user.
[0129] The 6D pose estimation algorithm is used to estimate the 6-DOF head pose of the faces in the video data to obtain the attention direction, listening attitude and object orientation of each user. Based on the attention direction, listening attitude and object orientation, the head pose features of each user are analyzed.
[0130] By using cascaded multi-layer convolution and attention mechanisms, the correspondence between local facial muscle movements and overall facial expression structure is captured from the video data. Based on the correspondence, temporal modeling is performed in conjunction with facial key point features to obtain the facial expression features of each user.
[0131] The system collects facial features of each user from multiple perspectives and simultaneously acquires audio and video stream data from three cameras.
[0132] Based on the aforementioned frontal facial features, the frontal facial regions of each user are located and tracked in real time using face detection and facial key point detection algorithms to obtain the dynamic key point coordinates of each user's lips. The dynamic key point coordinates of the lips are used as the lip movement sequence to obtain the lip movement sequence of each user. The lip movement sequence is then combined with the speaker recognition model to obtain the speaker label for each frame in the audio and video stream data.
[0133] Speech activity detection is performed on the audio and video stream data to segment out effective audio segments. The effective audio segments are then matched with the speaker tags across modalities using an audio and video timestamp alignment strategy to obtain the mapping relationship between audio segments and speakers.
[0134] Based on the mapping relationship, open-source audio analysis tools and source speech-to-text models are used to obtain the speaking quality, participation, number of interruptions, fundamental frequency of audio data, and loudness of audio data for each user.
[0135] The key social behaviors include: dominating speech behavior, interrupting speech behavior, summarizing statement level, gesture category, interactive behavior, and emotional state.
[0136] Specifically, the process of performing a one-stage mapping of the linguistic features, paralinguistic features, and nonlinguistic features using machine learning algorithms and prior rules to obtain key social behaviors includes:
[0137] The dominant speaking behavior and interruption behavior of the corresponding user are obtained based on the participation level and the number of interruption behaviors;
[0138] Obtain the speaking quality, audio data fundamental frequency, and audio data loudness of each user during the representative summary phase, and obtain the summary statement level of each user based on the speaking quality, audio data fundamental frequency, and audio data loudness of the representative users;
[0139] Define multiple preset gesture categories, and perform similarity threshold matching between the gesture features and the multiple preset gesture categories to obtain the gesture category of the corresponding user;
[0140] The angle between the line connecting the head posture features and the coordinate points of each user's head is calculated, and the interaction behavior of each user is obtained based on the angle. The interaction behavior includes head-down behavior and eye contact interaction behavior.
[0141] Based on the facial expression features, an emotion change curve for each user is plotted, and the emotional state of each user is obtained based on the emotion change curve. The emotional state includes the percentage of emotion and the degree of emotional fluctuation.
[0142] Specifically, the key social behaviors are mapped in a two-stage manner based on prior knowledge and observational experience to obtain mapping scores under multiple attribute windows, including:
[0143] The system obtains the number of speaking frames for each user during personal summaries, calculates the speaking duration and number of summary segments based on the number of speaking frames, determines the initial leadership score for the corresponding user based on the speaking duration, the number of summary segments, the speaking quality, and the dominant speaking behavior, and supplements the initial leadership score based on the summary statement level to obtain the target leadership score for the corresponding user.
[0144] Based on the interruption behavior and the interaction behavior, a score is determined on the user's listening attitude when others are speaking.
[0145] By using preset rule thresholds, the shoulder tension, posture openness, and group convergence of the corresponding user are calculated based on the human posture characteristics. The overall posture score is obtained by fusing the shoulder tension, posture openness, and group convergence.
[0146] Obtain the human posture estimation coordinate vector of the corresponding user, calculate the gesture kinetic energy density of the corresponding user based on the human posture estimation coordinate vector, obtain positive and negative indicators from the gesture category, and add the positive indicators, the negative indicators and the gesture kinetic energy density to obtain the nonverbal leadership potential score.
[0147] Each user is assigned the same initial emotion score. If the degree of emotion fluctuation exceeds a first preset threshold or the proportion of neutral emotions in the emotion proportion exceeds a second preset threshold, the initial emotion score is modified to obtain the target emotion score for the corresponding user.
[0148] Specifically, the step of accumulating all the mapping scores to obtain a quantitative evaluation result of the collaboration and interaction capabilities of all the users includes:
[0149] The target leadership score, the listening attitude score, the overall posture score, the nonverbal leadership potential score, and the target emotional score are normalized.
[0150] The normalized target leadership score, listening attitude score, overall posture score, nonverbal leadership potential score, and target emotion score are weighted and accumulated to obtain a quantitative assessment result of the collaboration and interaction ability of each user.
[0151] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a quantitative evaluation program for user collaboration and interaction capabilities, and the quantitative evaluation program for user collaboration and interaction capabilities, when executed by a processor, implements the steps of the quantitative evaluation method for user collaboration and interaction capabilities as described above.
[0152] In summary, this invention provides a method, system, and terminal for quantitatively evaluating users' collaboration and interaction abilities. The method includes: designing team tasks based on group dynamics; selecting multiple users to participate in the team tasks; and acquiring the linguistic, paralinguistic, and nonlinguistic features of the multiple users at different stages of the team tasks; performing a one-stage mapping of the linguistic, paralinguistic, and nonlinguistic features using machine learning algorithms and prior rules to obtain key social behaviors; and performing a two-stage mapping of the key social behaviors based on prior knowledge and observational experience to obtain mapping scores under multiple attribute windows, and accumulating all the mapping scores to obtain a quantitative evaluation result of the collaboration and interaction abilities of all the users. This invention can overcome inherent limitations such as response bias, cultural bias, and impression management, and achieve an objective and scientific quantitative evaluation of users' collaboration and interaction abilities. It provides a more realistic and insightful basis for ability assessment, team building, and optimization, and its evaluation results have higher reliability and validity, significantly improving the accuracy and application value of the evaluation.
[0153] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal that includes that element.
[0154] Of course, those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided by this invention can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0155] It should be understood that the application of the present invention is not limited to the examples above. Those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.
Claims
1. A method for quantitatively evaluating users' collaboration and interaction capabilities, characterized in that, The quantitative evaluation methods for user collaboration and interaction capabilities include: Team tasks were designed based on group dynamics, and multiple users were selected to participate in the team tasks. The linguistic features, paralinguistic features, and nonlinguistic features of the multiple users at different stages of the team tasks were obtained. The team task includes an observation phase, a personal description phase, a free discussion phase, and a representative summary phase; the language features include speaking quality, participation, and the number of interruptions; the paralinguistic features include the fundamental frequency and loudness of the audio data; and the non-verbal features include human posture features, gesture features, head posture features, and facial expression features. A one-stage mapping is performed on the linguistic features, paralinguistic features, and nonlinguistic features using machine learning algorithms and prior rules to obtain key social behaviors, specifically including: The dominant speaking behavior and interruption behavior of the corresponding user are obtained based on the participation level and the number of interruption behaviors; Obtain the speaking quality, audio data fundamental frequency, and audio data loudness of each user during the representative summary phase, and obtain the summary statement level of each user based on the speaking quality, audio data fundamental frequency, and audio data loudness of the representative users; Define multiple preset gesture categories, and perform similarity threshold matching between the gesture features and the multiple preset gesture categories to obtain the gesture category of the corresponding user; The angle between the line connecting the head posture features and the coordinate points of each user's head is calculated, and the interaction behavior of each user is obtained based on the angle. The interaction behavior includes head-down behavior and eye contact interaction behavior. Based on the facial expression features, an emotion change curve for each user is plotted, and the emotional state of each user is obtained based on the emotion change curve. The emotional state includes the percentage of emotion and the degree of emotional fluctuation. Based on prior knowledge and observational experience, the key social behaviors are mapped in a two-stage manner to obtain mapping scores under multiple attribute windows, specifically including: The system obtains the number of speaking frames for each user during personal summaries, calculates the speaking duration and number of summary segments based on the number of speaking frames, determines the initial leadership score for the corresponding user based on the speaking duration, the number of summary segments, the speaking quality, and the dominant speaking behavior, and supplements the initial leadership score based on the summary statement level to obtain the target leadership score for the corresponding user. Based on the interruption behavior and the interaction behavior, a score is determined on the user's listening attitude when others are speaking. By using preset rule thresholds, the shoulder tension of the corresponding user is calculated based on the human posture characteristics. Openness of posture and group convergence fusion , and The overall posture score is obtained, among which, The vector angle of the shoulder and elbow. The ratio of the widest distance between the upper limbs to the shoulder width, when ,but When it is 0, Before, then =1; Obtain the human pose estimation coordinate vector of the corresponding user. Calculate the kinetic energy density of the corresponding user's gesture. : ,in, , For time span; Positive and negative indicators are obtained from the gesture categories. The positive indicator is the proportion of positive hand gestures in the total hand gestures, and the negative indicator is the proportion of negative hand gestures in the total hand gestures. The positive indicators, the negative indicators, and the gesture kinetic energy density are added together to obtain the nonverbal leadership potential score. Set the same initial emotion score for each user. If the degree of emotion fluctuation exceeds a first preset threshold or the proportion of neutral emotions in the emotion proportion exceeds a second preset threshold, then modify the initial emotion score to obtain the target emotion score for the corresponding user. By accumulating all the mapping scores, a quantitative assessment of the collaboration and interaction capabilities of all the users is obtained.
2. The method for quantitatively evaluating user collaboration and interaction capabilities according to claim 1, characterized in that, The acquisition of linguistic features, paralinguistic features, and nonlinguistic features of multiple users at different stages of the team task specifically includes: Video data of different stages of the team's task is acquired by a preset number of vertically placed cameras; The video data is analyzed frame by frame using a human pose estimation framework to extract key skeletal points of the human body. Based on the key skeletal points, the limb movement trajectory and posture changes of each user are characterized. The limb movement trajectory and posture changes are modeled in combination with time series to obtain the human pose features and gesture features of each user. The 6D pose estimation algorithm is used to estimate the 6-DOF head pose of the faces in the video data to obtain the attention direction, listening attitude and object orientation of each user. Based on the attention direction, listening attitude and object orientation, the head pose features of each user are analyzed. By using cascaded multi-layer convolution and attention mechanisms, the correspondence between local facial muscle movements and overall facial expression structure is captured from the video data. Based on the correspondence, temporal modeling is performed in conjunction with facial key point features to obtain the facial expression features of each user. The system collects facial features of each user from multiple perspectives and simultaneously acquires audio and video stream data from three cameras. Based on the aforementioned frontal facial features, the frontal facial regions of each user are located and tracked in real time using face detection and facial key point detection algorithms to obtain the dynamic key point coordinates of each user's lips. The dynamic key point coordinates of the lips are used as the lip movement sequence to obtain the lip movement sequence of each user. The lip movement sequence is then combined with the speaker recognition model to obtain the speaker label for each frame in the audio and video stream data. Speech activity detection is performed on the audio and video stream data to segment out effective audio segments. The effective audio segments are then matched with the speaker tags across modalities using an audio and video timestamp alignment strategy to obtain the mapping relationship between audio segments and speakers. Based on the mapping relationship, open-source audio analysis tools and source speech-to-text models are used to obtain the speaking quality, participation, number of interruptions, fundamental frequency of audio data, and loudness of audio data for each user.
3. The method for quantitatively evaluating user collaboration and interaction capabilities according to claim 1, characterized in that, The step of accumulating all the mapping scores to obtain a quantitative evaluation result of the collaboration and interaction capabilities of all the users specifically includes: The target leadership score, the listening attitude score, the overall posture score, the nonverbal leadership potential score, and the target emotional score are normalized. The normalized target leadership score, listening attitude score, overall posture score, nonverbal leadership potential score, and target emotion score are weighted and accumulated to obtain a quantitative assessment result of the collaboration and interaction ability of each user.
4. A quantitative evaluation system for user collaboration and interaction capabilities, characterized in that, The quantitative evaluation system for user collaboration and interaction capabilities is used to implement the quantitative evaluation method for user collaboration and interaction capabilities as described in any one of claims 1-3, and the quantitative evaluation system for user collaboration and interaction capabilities includes: The feature acquisition module is used to design team tasks based on group dynamics, select multiple users to participate in the team tasks, and acquire the linguistic features, paralinguistic features, and nonlinguistic features of the multiple users at different stages of the team tasks. The first mapping module is used to perform a one-stage mapping of the language features, paralinguistic features and nonlinguistic features through machine learning algorithms and prior rules to obtain key social behaviors. The second mapping module is used to perform a two-stage mapping of the key social behaviors based on prior knowledge and observation experience, obtain mapping scores under multiple attribute windows, and accumulate all the mapping scores to obtain a quantitative evaluation result of the collaboration and interaction capabilities of all the users.
5. A terminal, characterized in that, The terminal includes: a memory, a processor, and a quantitative evaluation program for user collaboration and interaction capabilities stored in the memory and executable on the processor. When the quantitative evaluation program for user collaboration and interaction capabilities is executed by the processor, it implements the steps of the quantitative evaluation method for user collaboration and interaction capabilities as described in any one of claims 1-3.
6. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a quantitative evaluation program for user collaboration and interaction capabilities, which, when executed by a processor, implements the steps of the quantitative evaluation method for user collaboration and interaction capabilities as described in any one of claims 1-3.
Citation Information
Patent Citations
Team cooperation capability evaluation method and system based on audio analysis
CN115430155A