Employee emotion index evaluation method and device based on multi-modal data

By using a multimodal data evaluation method that combines voice, text, video, and physiological signals, an employee emotion index is generated. This solves the problem that the emotion evaluation in existing technologies is not comprehensive and accurate enough, improves the real-time performance and accuracy of employee emotion evaluation, and promotes team collaboration and employee management.

CN120950901AActive Publication Date: 2025-11-14BEIJING NORTH LATITUDE 30 DEGREE NETWORK TECH CO LTD

Patent Information

Application Number
CN202511479051.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-16
Publication Date
2025-11-14
Estimated Expiration
2045-10-16

AI Technical Summary

Technical Problem

Existing technologies for employee emotion assessment suffer from problems such as high subjectivity, poor real-time performance, low data collection frequency, and insufficient accuracy. They are unable to fully capture the complexity of emotions, leading to low team collaboration efficiency, declining employee performance, and talent loss.

Method used

A multimodal data evaluation method is adopted, including collecting employees' voice data, text data from enterprise communication tools, video data captured by visual behavior, and physiological signals. Through feature extraction and cross-modal alignment, data quality coefficient and scene correlation coefficient are calculated to generate a comprehensive emotion index, including emotion tendency, emotion intensity, and stress index, and calibrated in combination with environmental parameters.

Benefits of technology

It enables real-time, comprehensive, and accurate assessment of employee morale, improves team collaboration efficiency, reduces the risk of talent loss, and optimizes work management strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120950901A_ABST
    Figure CN120950901A_ABST
Patent Text Reader

Abstract

The invention provides an employee emotion index evaluation method and device based on multi-modal data. The method comprises the following steps: acquiring multi-modal related data of an employee; the related data comprises voice data of the employees, text data in an enterprise communication tool, video data captured by visual behaviors and physiological signals of the employees; carrying out feature extraction on the related data and carrying out cross-modal alignment to obtain multi-modal technical features; calculating a data quality coefficient and a scene correlation coefficient of each mode so as to calculate a final dynamic weight of each mode; fusing the multi-modal technical features based on the final dynamic weight of each modal, and performing analysis to obtain a comprehensive emotion index of the employee; the comprehensive emotional index comprises emotional tendency, emotional intensity and a pressure index. According to the method and the device, the comprehensive emotion index of the employee is obtained by comprehensively analyzing the data of the multiple modes, so that the defects that the current emotion evaluation of the employee is not comprehensive enough and low in accuracy are overcome.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a method and apparatus for evaluating employee sentiment index based on multimodal data. Background Technology

[0002] Currently, most companies use methods such as questionnaires, interviews, and observations to understand employee emotions. However, traditional methods cannot accurately and in real time detect employee emotions, which creates numerous obstacles to team building. In terms of team collaboration, failure to promptly detect conflicts or misunderstandings arising from emotional issues among employees can lead to poor communication and low collaboration efficiency within the team. Employees who are chronically in a negative emotional state will experience a significant decline in work performance, thereby impacting the company's overall performance. Furthermore, if employee emotional problems spread within the company, it can lead to talent loss, increase the company's human resource costs, weaken its market competitiveness, and hinder the successful achievement of company goals.

[0003] Existing technologies suffer from the following shortcomings: they are highly subjective, lack real-time performance, and have a low data collection frequency, making it difficult to fully capture the complexity of emotions. Furthermore, relying solely on text analysis limits the accuracy of judgments. Summary of the Invention

[0004] The main objective of this invention is to provide a method and apparatus for evaluating employee sentiment index based on multimodal data, aiming to overcome the shortcomings of current employee sentiment evaluation methods, which are not comprehensive enough and have low accuracy.

[0005] To achieve the above objectives, this invention provides a method for evaluating employee sentiment index based on multimodal data, comprising the following steps: Collect relevant data on employees' multimodal behavior; the relevant data includes employees' voice data, text data from enterprise communication tools, video data captured by visual behavior, and employees' physiological signals; Feature extraction and cross-modal alignment are performed on the relevant data to obtain multimodal technology features; Calculate the data quality coefficient and scene correlation coefficient for each modality to determine the final dynamic weight of each modality; The multimodal technical features are fused and analyzed based on the final dynamic weights of each modality to obtain the comprehensive emotional index of the employee; the comprehensive emotional index includes emotional tendency, emotional intensity, and stress index.

[0006] Further, feature extraction is performed on the relevant data, including: Acoustic features such as speech rate, pitch, and pause frequency are extracted from real-time captured conference / call audio. Analyze the sentiment of text data in enterprise communication tools and identify semantic features in combination with context; Analyze the implicit emotional features of facial micro-expressions and body language in video data captured by cameras; Analyze the physiological characteristics of physiological signals collected by wearable devices.

[0007] Furthermore, the data quality coefficients for each modality are calculated, including: The data quality coefficient of speech data modality is calculated based on the signal-to-noise ratio of speech data and the integrity of speech activity detection. The data quality coefficient of the video data modality is calculated based on the area of ​​the camera obstruction and the proportion of the effective area. The text confidence score and special symbol processing status were used to calculate the data quality coefficient of the text data modality. The data quality coefficient of physiological signal modality is calculated based on the stability of physiological signals and the data missing rate.

[0008] Furthermore, the calculation of the scene relevance coefficient includes: Different correlation coefficients are preset for each modal data in different scenario types; the scenario types include at least remote meetings, independent office work, and team collaborative discussions. Based on the current scene type when collecting relevant data, determine the scene correlation coefficient for each modality.

[0009] Furthermore, the formula for calculating the final dynamic weights of each mode is as follows: ; in, , These are the data quality coefficients for the i-th and j-th modes, respectively. , These are the scene correlation coefficients for the i-th and j-th modalities, respectively. It is the final dynamic weight of the i-th mode.

[0010] Furthermore, after obtaining the employee's comprehensive emotional index, the following is included: Based on the comprehensive emotion index, the emotion type is determined. Based on the emotion type, combined with the employee's current overtime hours and daily work report content, a coping strategy is generated. The coping strategy includes automatically triggering workflows to adjust task priorities and recommending rest intervals. Establish a correlation model between the comprehensive sentiment index and performance indicators to identify employees with a high risk of turnover.

[0011] Furthermore, the method also includes: Real-time removal of identity information from voice and text data; The behavioral data in the relevant data is processed using k-anonymization technology.

[0012] Furthermore, after analyzing and obtaining the employee's comprehensive emotional index, the analysis also includes: Obtain environmental parameters, including office noise level, light intensity, temperature and humidity parameters, and team personnel density; Obtain the interference coefficients corresponding to each of the environmental parameters; wherein, a correlation analysis between the environmental parameters and the emotion recognition error is performed in advance to generate the interference coefficients corresponding to each environmental parameter. The comprehensive emotion index is calibrated based on the interference coefficients of various environmental parameters to obtain the final comprehensive emotion index.

[0013] Furthermore, after analyzing and obtaining the employee's comprehensive emotional index, the analysis also includes: Construct a team working relationship graph and calculate the emotion transmission coefficient between employees based on the frequency of collaborative work and communication density of employees in the team; Real-time monitoring of employees with abnormal comprehensive emotional index in the team, and calculation of the potential impact of abnormal employees on the overall team mood based on the emotional transmission coefficient among employees; When the potential impact value exceeds a preset threshold, a differentiated intervention plan is automatically generated.

[0014] The present invention also provides an employee sentiment index assessment device based on multimodal data, comprising: The data acquisition unit is used to collect relevant data on employees' multimodal behavior; the relevant data includes employees' voice data, text data from enterprise communication tools, video data captured by visual behavior, and employees' physiological signals. The alignment unit is used to extract features from the relevant data and perform cross-modal alignment to obtain multimodal technical features; The calculation unit is used to calculate the data quality coefficient and scene correlation coefficient of each modality in order to calculate the final dynamic weight of each modality. The analysis unit is used to fuse multimodal technical features based on the final dynamic weights of each modality and analyze them to obtain the comprehensive emotional index of the employee; the comprehensive emotional index includes emotional tendency, emotional intensity, and stress index.

[0015] The present invention provides a method and apparatus for evaluating employee emotional index based on multimodal data, comprising: collecting relevant multimodal data of employees; the relevant data including employee voice data, text data from enterprise communication tools, video data captured by visual behavior capture, and employee physiological signals; extracting features from the relevant data and performing cross-modal alignment to obtain multimodal technical features; calculating the data quality coefficient and scene correlation coefficient of each modality to calculate the final dynamic weight of each modality; fusing the multimodal technical features based on the final dynamic weight of each modality and analyzing them to obtain the comprehensive emotional index of the employee; the comprehensive emotional index includes emotional tendency, emotional intensity, and stress index. In this invention, the comprehensive emotional index of employees is obtained by comprehensively analyzing data from multiple modalities, overcoming the shortcomings of current employee emotional assessments being incomplete and inaccurate. Attached Figure Description

[0016] Figure 1 This is a schematic diagram of the steps of an employee sentiment index assessment method based on multimodal data in one embodiment of the present invention; Figure 2 This is a structural block diagram of an employee sentiment index assessment device based on multimodal data in one embodiment of the present invention; Figure 3 This is a schematic block diagram of the structure of a computer device according to an embodiment of the present invention.

[0017] The implementation, functional features, and advantages of this invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0019] Reference Figure 1 One embodiment of the present invention provides a method for evaluating employee sentiment index based on multimodal data, including the following steps: Step S1: Collect relevant data on employee multimodal behavior; the relevant data includes employee voice data, text data from enterprise communication tools, video data captured by visual behavior, and employee physiological signals. Step S2: Extract features from the relevant data and perform cross-modal alignment to obtain multimodal technology features; Step S3: Calculate the data quality coefficient and scene correlation coefficient for each modality to calculate the final dynamic weight of each modality; Step S4: Based on the final dynamic weights of each modality, the multimodal technical features are fused and analyzed to obtain the comprehensive emotional index of the employee; the comprehensive emotional index includes emotional tendency, emotional intensity, and stress index.

[0020] In this embodiment, as described in step S1 above, the aim is to acquire basic information that comprehensively reflects the emotional state of employees through diversified data collection methods. Specifically, voice data collection covers real-time voice streams from employee presentations and work calls, while also supporting retrospective acquisition of historical recordings to capture voice expression characteristics at different times; text data collection from enterprise communication tools covers emails, instant messaging messages, etc., including not only the text itself but also auxiliary expressive elements such as emoticons and special punctuation, while fully preserving contextual relationships; visual behavior data is recorded through compliantly deployed image acquisition devices, recording changes in employees' facial micro-expressions and body movement trajectories, focusing on capturing those outward expressions of emotions that are difficult to conceal; physiological signals are continuously monitored through wearable devices worn by employees, monitoring physiological indicators such as heart rate fluctuations and skin conductance. These physiological indicators can directly reflect the activity state of the autonomic nervous system, providing an objective basis for judging internal emotional changes. During the collection process, metadata such as the specific scenario and time node of data generation is recorded simultaneously, and all raw data can be encrypted to ensure data security.

[0021] As described in step S2 above, this is a crucial step in transforming raw data into analyzable features. For speech data, acoustic analysis techniques are used to extract features such as speech rate changes, pitch fluctuations, and pause patterns. These features reflect speech changes caused by emotional fluctuations. For text data, natural language processing techniques are used to analyze the emotional tendencies and semantic connotations in the text, while also identifying the emotional features implied in complex expressions such as irony and metaphor. For visual data, computer vision algorithms are used to extract facial key point motion trajectories (such as the curvature of the corners of the mouth and changes in eyebrow position) and body posture features (such as gesture amplitude and body tilt angle) from the video stream. These features can reflect unconscious emotional expressions. For physiological data, signal processing techniques are used to extract physiological features that reflect stress and emotional states, such as heart rate variability and skin conductance intensity. After feature extraction, standardization is used to map features from different modalities to a unified numerical range. At the same time, a time window segmentation strategy is used to solve the problem of temporal differences. Finally, a feature mapping model projects various features into a shared feature space, achieving effective alignment of cross-modal features and ensuring the comparability of different types of features.

[0022] As described in step S3 above, the weight of each modality in the comprehensive analysis is determined through quantitative evaluation. The calculation of the data quality coefficient is based on differentiated standards according to the characteristics of different modalities: speech data primarily considers clarity and the proportion of effective speech; text data focuses on the reliability of semantic analysis and the recognition effect of special symbols; visual data focuses on the degree of image occlusion and the visibility of key areas; and physiological data is based on signal stability and data integrity. The scene relevance coefficient is preset according to the emotional response capabilities of each modality in different work scenarios: for example, in remote meeting scenarios, speech data has a higher relevance; in independent office scenarios, physiological signals have greater reference value; and in team collaboration scenarios, the relevance of speech and visual data is relatively prominent. The final weights are calculated by combining the comprehensive data quality coefficient and the scene relevance coefficient, ensuring that high-quality data with a high degree of matching with the current scenario has a higher weight in the fusion analysis, thereby improving the accuracy of the evaluation.

[0023] As described in step S4 above, a quantitative indicator that comprehensively reflects the emotional state of employees is generated by fusing multimodal features. Based on the weights determined in step S3, the cross-modal aligned features are weighted and fused to ultimately form a comprehensive emotional index with three dimensions: emotional tendency, which distinguishes between positive and negative emotional attributes; emotional intensity, which reflects the intensity of emotional expression; and stress index, which measures the employee's current level of psychological stress. The generation of these indices is not only based on real-time data but also incorporates historical features for trend analysis and is linked to specific work scenarios and task information to trace the potential triggers for emotional changes. Through this multi-dimensional quantitative assessment, the emotional state of employees can be objectively presented.

[0024] In one embodiment, feature extraction of the relevant data includes: Acoustic features such as speech rate, pitch, and pause frequency are extracted from real-time captured conference / call audio. Analyze the sentiment of text data in enterprise communication tools and identify semantic features in combination with context; Analyze the implicit emotional features of facial micro-expressions and body language in video data captured by cameras; Analyze the physiological characteristics of physiological signals collected by wearable devices.

[0025] In this embodiment, acoustic analysis technology is used to extract core features reflecting emotional states from real-time captured conference or call audio. Specifically, speech rate is obtained by statistically analyzing the number of syllables and their rate of change per unit time, reflecting states of tension or relaxation; pitch features focus on the mean, fluctuation range, and trend of the fundamental frequency, showing significant differences for different emotions (e.g., pitch rises when angry, pitch falls when frustrated); and pause frequency is captured by calculating the number of speech interruptions and the percentage of pause duration per unit time, indicating changes in expression rhythm caused by hesitation, thought, or emotional fluctuations.

[0026] For text data in enterprise communication tools, natural language processing (NLP) techniques are used for in-depth analysis. First, sentiment analysis identifies the positive, negative, or neutral emotional tone inherent in the text, while simultaneously understanding semantic coherence within the context to avoid biases caused by isolated interpretations. Building upon this foundation, semantic features are further extracted, including the emotional intensity of words, sentence structure features (such as the strong emotions expressed by rhetorical questions), and the emotional orientation of special symbols (such as exclamation marks and emoticons). Particular attention is paid to identifying indirect expressions such as irony and metaphor to comprehensively capture the true emotions behind the text.

[0027] For video data captured by cameras, computer vision algorithms are used to extract emotional features. Facial key point detection technology is used to capture micro-expression changes, such as the raising or lowering of eyebrows, the stretching or contraction of the corners of the mouth, and the movement of eye muscles. These subtle movements often reflect genuine emotions that are difficult to conceal. At the same time, body language features are analyzed through human posture estimation, including the amplitude and frequency of gestures, leaning forward or backward, and changes in sitting posture. These non-verbal signals are converted into quantifiable emotional features to supplement the shortcomings of facial expression analysis.

[0028] For physiological signals collected by wearable devices, physiological feature analysis techniques are used to extract indicators closely related to emotional states. A key focus is on heart rate variability (the fluctuation of heart rate over time), whose changes are directly related to emotions such as stress and anxiety. Simultaneously, the intensity and frequency of skin conductance responses are analyzed; this indicator reflects the level of sympathetic nerve excitation and indirectly reflects the level of emotional arousal. Other auxiliary physiological features include changes in body temperature and respiratory rate. These indicators collectively constitute a set of physiological features reflecting internal emotional fluctuations.

[0029] In one embodiment, calculating the data quality coefficient for each modality includes: The data quality coefficient of speech data modality is calculated based on the signal-to-noise ratio of speech data and the integrity of speech activity detection. The data quality coefficient of the video data modality is calculated based on the area of ​​the camera obstruction and the proportion of the effective area. The text confidence score and special symbol processing status were used to calculate the data quality coefficient of the text data modality. The data quality coefficient of physiological signal modality is calculated based on the stability of physiological signals and the data missing rate.

[0030] In this embodiment, the data quality coefficient calculation for the speech data modality uses signal-to-noise ratio (SNR) and speech activity detection integrity as core evaluation indicators. The coefficient value is determined by comprehensively analyzing the performance of these two indicators. SNR reflects the ratio of the speech signal to background noise; a higher value indicates better speech clarity, less susceptibility to environmental interference, and stronger support for the reliability of emotion feature extraction. Conversely, a low SNR means the speech is overwhelmed by noise, potentially leading to distortion in acoustic feature extraction. Speech activity detection integrity measures the proportion of effective speech segments throughout the entire acquisition period. Higher integrity indicates a lower proportion of invalid silences, noise, and other non-speech segments, resulting in more sufficient effective speech information for emotion analysis. Insufficient integrity indicates a scarcity of effective speech data, potentially affecting the comprehensiveness of emotion feature extraction. In the specific calculation, signal processing techniques are used to obtain the specific values ​​of SNR and the integrity ratio of speech activity detection. Then, based on a preset mapping rule, the quantitative results of these two indicators are merged into a single quality coefficient, which objectively reflects the reliability of the speech data for emotion analysis.

[0031] The data quality coefficient calculation for video data modalities revolves around the proportion of occluded areas and effective areas from the camera viewpoint. Quality is quantified through a comprehensive evaluation of these two indicators. Occluded areas refer to regions in the captured image where key facial or limb features are not visible due to obstacles. A higher proportion of occluded areas means less effective visual information available for analysis, potentially leading to a loss of micro-expressions or body language features; conversely, a lower proportion indicates better visual data integrity. The effective area proportion focuses on the percentage of key areas related to emotional expression (such as the face and hands) within the entire captured image. A higher proportion indicates more thorough capture of details in these key areas, facilitating the extraction of accurate emotional features; a low effective area proportion may result in blurred or lost key emotional expression information. The calculation involves using computer vision algorithms to perform region analysis on video frames, obtaining the proportion of occluded areas and the effective area proportion separately. These are then weighted and fused to convert them into a unified quality coefficient, reflecting the quality of the video data's support for emotion analysis.

[0032] The data quality coefficient for the text data modality is calculated based on the text confidence score and special symbol processing status output by the BERT model. The final coefficient is determined by integrating the performance of these two indicators. The text confidence score calculated by the BERT model reflects the reliability of the model's understanding of text semantics and judgment of sentiment tendency. The higher the confidence score, the better the accuracy of text feature extraction and the greater the reference value for sentiment analysis. If the confidence score is low, it means that the text semantics may be misinterpreted, affecting the effectiveness of sentiment features. The special symbol processing status is used to evaluate the ability to identify and interpret non-textual emotional expression elements such as emoticons and special punctuation marks in the text. The better the processing status, the more complete these auxiliary emotional expression information are captured and transformed into effective features, which can supplement the emotional expression of the text itself. If the processing status is poor, it may lead to the omission of non-textual emotional information, affecting the comprehensiveness of text sentiment analysis. In the calculation process, the confidence score value of the BERT model and the completion index of special symbol processing need to be obtained first. Then, according to the preset fusion rules, the two are integrated into a single quality coefficient to quantify the quality level of text data used for sentiment analysis.

[0033] The data quality coefficient calculation for physiological signal modalities uses the stability of physiological signals and the data missing rate as core indicators. Quality is quantified through a comprehensive analysis of these two indicators. The stability of physiological signals reflects the degree of fluctuation during the acquisition process. Higher stability indicates less influence from factors such as equipment interference, and the more accurately the extracted physiological features (such as heart rate variability and skin conductance) reflect emotional states. Insufficient stability can lead to distortion of physiological features due to noise mixed in the signal, affecting the accuracy of emotion analysis. The data missing rate measures the proportion of complete physiological signals in the total acquisition period. A lower missing rate indicates better completeness of physiological data, providing a continuous and complete trajectory of emotional changes. A high missing rate means the loss of physiological information during key periods, potentially causing a break in the analysis of emotional features. During calculation, signal processing techniques are used to evaluate the stability level and data missing rate of physiological signals separately. Then, a pre-defined fusion algorithm converts these two factors into a unified quality coefficient to objectively reflect the reliability of physiological data for emotion analysis.

[0034] In one embodiment, the calculation of the scene correlation coefficient includes: Different correlation coefficients are preset for each modal data in different scenario types; the scenario types include at least remote meetings, independent office work, and team collaborative discussions. Based on the current scene type when collecting relevant data, determine the scene correlation coefficient for each modality.

[0035] In this embodiment, the preset scenario correlation coefficient is based on the difference in the actual contribution of each modality of data to emotional expression under different work scenarios, and establishes a standardized coefficient mapping rule. First, it is necessary to sort out the core work scenarios commonly encountered by enterprises, covering at least three typical scenarios: remote meetings, independent office work, and team collaborative discussions. These three scenarios have significant differences in communication methods, interaction frequency, and emotional expression carriers, and are key scenarios for changes in employees' emotional state in their daily work.

[0036] For each scenario, the correlation between each modality of data (speech, text, video, and physiological signals) and emotion recognition needs to be analyzed: In remote meeting scenarios, speech is the main communication medium for employees to convey their views and express their attitudes, and emotional fluctuations are often directly reflected in changes in speech rate and tone. Therefore, the correlation coefficient of the speech modality needs to be set at a high level. Text is mostly used as a supplementary record of meeting content or a brief interaction, and its reflection of emotions is weaker than that of speech, so the correlation coefficient is second. Video can capture facial micro-expressions (such as frowning and drooping corners of the mouth) to help judge emotional state, and the correlation coefficient needs to be higher than that of text but lower than that of speech. Physiological signals are mainly used to monitor the implicit stress brought about by long meetings, and their reflection of immediate emotional expression is relatively indirect, so the correlation coefficient is set at a low level.

[0037] In independent office settings, employees primarily work alone with minimal voice communication, significantly reducing the contribution of voice modality to emotion recognition, resulting in the lowest correlation coefficient. Text primarily consists of work-related messages or task records with colleagues, where emotional expression is more subtle, leading to a slightly higher correlation coefficient than voice. Video captures employees' natural states when alone (such as restlessness after prolonged sitting or relaxation after completing a task), providing a more realistic reflection of emotions, resulting in a higher correlation coefficient than text. Physiological signals can monitor psychological stress when alone in real time and are the core carrier of emotion recognition, thus the highest correlation coefficient is set.

[0038] In team collaboration scenarios, employees frequently exchange opinions via voice, and emotions often fluctuate with the clash of viewpoints, resulting in a high correlation coefficient for voice modality. Text is mostly used for temporary communication or document sharing during collaboration, and its emotional response is weaker, leading to a lower correlation coefficient than voice. Video can capture body language (such as gesture amplitude and body orientation) during team interactions, helping to determine whether emotions are positive, and its correlation coefficient is higher than that of text. Physiological signals are used to monitor collaborative stress (such as skin conductance response during disagreements), but their priority in reflecting emotions is lower than that of voice and video, and their correlation coefficient is set to a moderate level.

[0039] Through the correlation analysis of the above scenarios and modalities, fixed correlation coefficients are preset for each modal data in each scenario, forming a standardized coefficient library to ensure the consistency and rationality of subsequent coefficient calls.

[0040] When conducting employee sentiment index assessments, it is necessary to first clarify the specific scenario type during data collection. The scenario corresponding to the current data (remote meeting, independent work, or team collaboration discussion) can be determined through automatic system identification (such as using meeting software to determine whether it is a remote meeting or using office area cameras to determine whether it is an independent office) or manual annotation.

[0041] After determining the scenario type, retrieve the correlation coefficients corresponding to each modality of data in that scenario from the preset coefficient library: if the current scenario is a remote meeting, directly match the preset coefficients for voice, text, video, and physiological signals in the remote meeting scenario; if the current scenario is independent office work, retrieve the corresponding coefficients for the independent office scenario; if it is a team collaboration discussion scenario, retrieve the preset coefficients for that scenario.

[0042] In one embodiment, the formula for calculating the final dynamic weights of each mode is as follows: ; in, , These are the data quality coefficients for the i-th and j-th modes, respectively. , These are the scene correlation coefficients for the i-th and j-th modalities, respectively. It is the final dynamic weight of the i-th mode.

[0043] In this embodiment, data reliability and scene adaptability are considered simultaneously. The numerator of the formula multiplies the data quality coefficient of the i-th modality by the scene relevance coefficient. This reflects both the reliability of the modality data itself (e.g., the clarity of voice data, the accuracy of text analysis) and its matching degree with the current scene (e.g., the high relevance of voice data in a remote meeting). This multiplicative relationship ensures that only modality data with both high quality and high scene adaptability can occupy a dominant position in the final weight, avoiding bias caused by single-dimensional evaluation. For example, even if a modality data has extremely high quality, but low relevance to the current scene (e.g., voice data in an independent office scene), its weight will be reasonably suppressed.

[0044] Through the above calculation logic, the weights of each modality can objectively reflect its actual information value in the current scenario: high-quality and highly relevant modal data receive higher weights and play a dominant role in the calculation of the comprehensive sentiment index; low-quality or low-relevance modal data have reduced weights to minimize their interference with the overall results. This differentiated weight allocation mechanism makes the multimodal feature fusion process more consistent with the actual emotional expression patterns, ultimately improving the accuracy and reliability of employee sentiment index assessment and providing a more scientific quantitative basis for the formulation of subsequent sentiment intervention strategies.

[0045] In one embodiment, after obtaining the employee's overall emotional index, the process includes: Based on the comprehensive emotion index, the emotion type is determined. Based on the emotion type, combined with the employee's current overtime hours and daily work report content, a coping strategy is generated. The coping strategy includes automatically triggering workflows to adjust task priorities and recommending rest intervals. Establish a correlation model between the comprehensive sentiment index and performance indicators to identify employees with a high risk of turnover.

[0046] In this embodiment, after obtaining the employee's comprehensive emotional index, the current emotional type of the employee is accurately defined based on the combined characteristics of emotional tendency, emotional intensity and stress index. For example, when the emotional tendency shows obvious negative characteristics, the emotional intensity is at a high level and the stress index exceeds the normal threshold, it can be judged as high-stress negative emotion; if the emotional tendency is positive but the emotional intensity is low and the stress index is in the normal range, it can be judged as stable positive emotion. Through the collaborative analysis of multi-dimensional indices, the accuracy of the emotional type classification is ensured.

[0047] Based on the identification of emotion types, further in-depth correlation analysis is conducted by combining real-time work status data of employees: on the one hand, the current overtime hours of employees are retrieved to determine whether their emotional abnormalities are caused by long-term high-intensity work; for example, if high-pressure negative emotions are accompanied by several consecutive days of overtime, it usually points to emotional problems caused by workload overload; on the other hand, the key content in the employees' daily work reports is analyzed to extract information such as task progress and feedback on difficulties, and to identify whether the emotional abnormalities are related to specific work tasks; for example, if the daily work reports frequently mention that the task difficulty exceeds expectations and collaboration is hindered, and negative emotional characteristics appear at the same time, the cause of the emotional problem can be identified as being related to the specific work scenario.

[0048] Based on the correlation analysis results between the above-mentioned emotion types and work status, differentiated coping strategies are generated: For high-stress negative emotions caused by workload overload, a workflow adjustment mechanism is automatically triggered, prioritizing the reduction of the employee's current task priority, allocating urgent but non-core tasks to other collaborators, and recommending reasonable rest intervals based on overtime hours and stress index to prevent further deterioration of emotions; For negative emotions caused by task difficulty, in addition to adjusting task priority, relevant resource support (such as technical guidance documents and collaboration channels) will be simultaneously pushed, and tasks will be broken down in stages to help employees gradually alleviate stress; For employees with stable and positive emotions, it is recommended to appropriately increase challenging tasks or maintain the current work rhythm to prolong the good emotional state. The entire coping strategy generation process is data-driven at its core, ensuring the relevance and feasibility of the strategies, and achieving a deep integration of emotion intervention and work management.

[0049] Next, through data integration, comprehensive emotional index data (including temporal changes in emotional tendency, emotional intensity, and stress index) and corresponding performance indicators (such as task completion rate, work quality score, and project contribution) are collected from employees over historical periods to construct an analytical dataset containing multi-dimensional data. During the data preprocessing stage, the emotional index and performance indicators need to be aligned along the time dimension to ensure that each set of emotional data matches the performance of the same period, while outlier data (such as extreme emotional values ​​and performance statistical error data) is removed to ensure the reliability of the dataset.

[0050] Subsequently, data analysis techniques were used to construct a correlation model between the comprehensive sentiment index and performance indicators. Correlation analysis was used to identify the dimension of the sentiment index that has the most significant impact on performance. Further regression analysis or machine learning algorithms were employed to quantify the impact of changes in the sentiment index on performance indicators, forming a predictable correlation model. Based on this correlation model, and combined with the temporal variation characteristics of employee sentiment indices, employees with high-risk turnover tendencies were identified: If an employee's comprehensive sentiment index shows a long-term negative trend (e.g., persistently negative sentiment tendency, stress index consistently above the threshold), and this sentiment change has been reflected in a significant decline in performance indicators through the correlation model (e.g., decreased task completion rate, lower work quality score), while excluding non-continuous influencing factors such as short-term work task adjustments and temporary personal matters, then the employee can be determined to have a high-risk turnover tendency. Such employees are then marked as key monitoring targets, and warning information is sent to managers. Simultaneously, based on the triggers for their sentiment and performance changes, targeted intervention measures are implemented (e.g., one-on-one communication and guidance, re-matching of work tasks, career development path planning, etc.) to help the company intervene early and reduce the risk of losing core talent.

[0051] In one embodiment, after obtaining the employee's overall emotional index, the process includes: Generate individual employee mood dashboards, displaying real-time mood radar charts, historical trend curves, cross-team comparative analysis, and generating daily employee mood reports; Group sentiment distribution is displayed by department / project team; Hierarchical access control is set up so that employees can view their personal data independently, while managers can only see group statistics and risk warning data, and sensitive operation audit logs are set up.

[0052] In this embodiment, after obtaining the employee's comprehensive emotional index, a personal emotional dashboard is first constructed. This dashboard uses multi-dimensional data visualization to intuitively present the dynamic changes and comparisons of the employee's emotional state. The real-time emotional radar chart visually displays the specific numerical distribution of current emotional tendency, emotional intensity, and stress index using multi-axis coordinates. The length of different axes corresponds to the level of the index value, helping employees quickly locate the core characteristics of their current emotions. The historical trend curve, with time as the horizontal axis and the emotional index as the vertical axis, presents the trajectory of changes in the employee's emotional tendency, emotional intensity, and stress index over a recent period. The curve fluctuations clearly identify the stability of emotions or points of abnormal change. For example, if the stress index curve continuously rises over a certain period, it can prompt employees to pay attention to the impact of their work or life situation on their emotions during that period. Cross-team comparative analysis compares the employee's personal emotional index with the average emotional index of other employees in the same department or position, displaying the differences in the form of bar charts or line graphs. This helps employees objectively judge their emotional state's position within the group, avoiding cognitive biases caused by singular self-perception.

[0053] Building upon the personal mood dashboard, a daily employee mood report is generated. The report focuses on changes in the daily mood index, including not only a real-time mood radar chart and a summary of core data on the day's mood trends, but also connects mood change points to the day's work scenarios (such as meeting periods or busy task periods) to analyze potential triggers for mood fluctuations. Furthermore, it provides preliminary adjustment plans for abnormal mood situations (such as stress levels exceeding normal thresholds), enabling employees to adjust their state promptly based on the report's content and achieve self-management and regulation of their emotions.

[0054] To support organizational-level emotion management decisions, the overall emotion index of individual employees is categorized and integrated by department or project team to generate a visualization of group emotion distribution. Specifically, firstly, by department or project team, the average, median, and distribution range of the emotion tendency, emotion intensity, and stress index of all employees within the group are calculated, and the statistical data reflects the overall emotional characteristics of the group.

[0055] In terms of visualization, group emotion distribution can be presented using heatmaps, box plots, or pie charts. Heatmaps, organized by department / project team, use different color gradients to represent the average emotional tendency of the group (e.g., red for negative, green for positive) or the average stress index (e.g., dark for high stress, light for low stress), allowing managers to intuitively judge the differences in emotional atmosphere among different groups. Box plots, on the other hand, show the dispersion of the emotional index distribution among different groups, helping managers identify the balance of emotional states within the group. Through the visualization of group emotion distribution, managers can quickly grasp the overall emotional status of each department and project team at the organizational level, promptly identify groups with abnormal emotional atmospheres, and provide direction for optimizing team management and improving organizational effectiveness.

[0056] To ensure the security of employee emotional data, a strict hierarchical access control mechanism has been established, with different data access permissions defined based on user roles (employees and managers). For individual employees, access to their own emotional data is restricted; they can view their personal emotional dashboard, historical trend curves, and daily emotional reports, ensuring they can independently monitor their own emotional state while preventing their personal data from being accessed by others. For managers, access is limited to the group emotional statistics within their jurisdiction and early warning information for employees with high-risk emotional states within that group, but they cannot view detailed emotional data for specific employees. This satisfies managers' decision-making needs while preventing individual data leaks. Hierarchical access control is implemented through account role binding and data access interface restrictions. When data is accessed, user roles and access permissions are automatically verified, rejecting data access requests that exceed permissions, forming the first line of defense for data protection.

[0057] Meanwhile, a sensitive operation audit log is set up to record all critical operations involving sentiment data in real time. The audit log includes the operation time, user, type, target, and result, ensuring that every sensitive operation is traceable. When data access anomalies occur, the audit log provides crucial evidence for tracing the source, helping administrators promptly identify security risks. Furthermore, the audit log is archived regularly to meet data security compliance requirements, further strengthening the security protection and control of sentiment data. This ensures that the entire sentiment index assessment system is both compliant and secure during data use, enhancing employees' trust in data security.

[0058] In one embodiment, the method further includes: Real-time removal of identity information from voice and text data; The behavioral data in the relevant data is processed using k-anonymization technology.

[0059] In this embodiment, during the acquisition and processing of voice data, information that can be associated with employee identity, such as voiceprint features and names mentioned in the voice, is stripped away in real time. When processing text data, information that can directly identify an individual, such as email address, username, and employee number, is deleted simultaneously to prevent the leakage of identity information from the source and ensure that subsequent data processing only revolves around sentiment analysis and does not involve the association of the employee's personal identity.

[0060] The collected behavioral data related to employee visual behavior and physiological signals were processed using k-anonymization technology. By grouping the data and hiding individual unique characteristics, the processed data must be presented together with the data of at least k other individuals, making it impossible to identify the specific behavioral information of a single employee. This preserves the data's value for sentiment analysis while ensuring that employee behavioral data cannot be tracked or identified individually.

[0061] In one embodiment, after analyzing and obtaining the employee's comprehensive emotional index, the method further includes: Obtain environmental parameters, including office noise level, light intensity, temperature and humidity parameters, and team personnel density; Obtain the interference coefficients corresponding to each of the environmental parameters; wherein, a correlation analysis between the environmental parameters and the emotion recognition error is performed in advance to generate the interference coefficients corresponding to each environmental parameter. The comprehensive emotion index is calibrated based on the interference coefficients of various environmental parameters to obtain the final comprehensive emotion index.

[0062] In this embodiment, sensors deployed in the office area collect key environmental data in real time that affect employees' emotional perception and expression. These data include noise levels in the office area (such as equipment operating noise and the volume of conversations), light intensity (such as natural light intensity and indoor lighting brightness), temperature and humidity parameters (such as ambient temperature and air humidity), and team personnel density (such as the number of employees present in the office area at the same time and the degree of crowding at workstations). These environmental parameters can all indirectly affect employees' emotional expression or the emotion recognition process. For example, high-decibel noise may cause employees to become irritable and may also interfere with the accurate extraction of voice emotion features; therefore, they need to be included in the calibration scope.

[0063] Secondly, the environmental interference coefficient is determined by conducting a correlation analysis between environmental parameters and emotion recognition errors. Through extensive experimental data, the degree of deviation between the emotion index assessment results and the actual emotional state of employees under different environmental parameter values ​​is studied. For example, when the noise level exceeds a certain threshold, the error in recognizing emotional features in the speech modality increases significantly; when the light intensity is too weak, the accuracy of micro-expression recognition in the visual modality decreases. Based on these correlations, a corresponding interference coefficient is set for each environmental parameter. The magnitude of the coefficient is positively correlated with the emotion recognition error caused by that parameter; that is, the stronger the interference of the parameter on emotion recognition, the higher the corresponding interference coefficient. In practical applications, the preset interference coefficient is directly matched based on the environmental parameters acquired in real time.

[0064] Finally, the comprehensive emotion index is calibrated. This involves integrating the interference coefficients corresponding to each environmental parameter and considering their weighted impact on different modalities of emotional characteristics to correct the initially obtained comprehensive emotion index. For example, if the noise level in the real-time environment is high (corresponding to a high interference coefficient) and the light intensity is insufficient (corresponding to a medium interference coefficient), then the focus is on correcting the speech emotion characteristics affected by noise and the visual emotion characteristics affected by light. The comprehensive emotion index is then recalculated, ultimately yielding the final comprehensive emotion index after eliminating environmental interference, making the evaluation results more closely reflect the employees' true emotional state in the absence of environmental interference.

[0065] In one embodiment, after analyzing and obtaining the employee's comprehensive emotional index, the method further includes: Construct a team working relationship graph and calculate the emotion transmission coefficient between employees based on the frequency of collaborative work and communication density of employees in the team; Real-time monitoring of employees with abnormal comprehensive emotional index in the team, and calculation of the potential impact of abnormal employees on the overall team mood based on the emotional transmission coefficient among employees; When the potential impact value exceeds a preset threshold, a differentiated intervention plan is automatically generated.

[0066] In this embodiment, firstly, a team work relationship graph is constructed. This is achieved by connecting to data sources such as internal project management tools, collaborative office software, and communication systems to collect interaction information among team members. Collaborative work frequency is measured by the number of projects members participate in together, the number of tasks they handle together, and the duration of collaborative work, reflecting the closeness of their work overlap. Communication density is measured by the number of phone calls between members, the frequency of instant messaging, and the duration of joint meetings, reflecting the frequency of information exchange between members. This interaction data is then transformed into nodes and edges in the graph: each employee is an independent node in the graph, and the edges between nodes represent the work relationships between employees. The weight of each edge is determined by a comprehensive calculation of collaborative work frequency and communication density; a higher weight indicates a closer work relationship between the two employees.

[0067] Based on the team's working relationship graph, an emotion transmission coefficient is calculated between employees. This coefficient quantifies the likelihood and intensity of an employee's emotional state influencing another employee's emotions; its value is positively correlated with the tightness of the working relationship between the two employees. The tighter the working relationship (i.e., the higher the weight of the edges between nodes in the graph), the greater the probability of mutual emotion transmission and the stronger the impact, resulting in a higher emotion transmission coefficient. Conversely, the lower the emotion transmission coefficient, the less tightly related the employees are. The calculation process requires adjustments based on historical emotion transmission data. For example, if historical data shows that for two closely related employees, the probability of the other experiencing negative emotions shortly after one's negative emotions is significantly higher than with other relationship combinations, then the emotion transmission coefficient needs to be increased to ensure that it reflects both the working relationship and the actual pattern of emotion transmission. The resulting emotion transmission coefficient matrix provides a precise quantitative basis for subsequent analysis of the diffusion and impact of emotions within the team.

[0068] Next, monitor the overall emotional index of employees in the team in real time. Update the overall emotional index of each employee at a set frequency (such as hourly or half-day). Based on the preset abnormal judgment criteria (such as emotional tendency being in the negative range, stress index exceeding the normal threshold by a certain percentage, and emotional intensity fluctuating drastically in a short period of time), automatically mark employees with abnormal emotional indices. These employees are regarded as potential sources of emotional risk transmission.

[0069] After identifying employees with abnormal emotions, the potential impact on the team's overall mood is calculated based on the obtained emotion transmission coefficient. This calculation covers all employees associated with the abnormal employee in the team's work relationship graph. For each member with a work relationship with the abnormal employee, their emotion transmission coefficient is multiplied by the degree of the abnormal employee's emotional abnormality (e.g., the magnitude of emotional deviation from neutral values, the proportion of stress indices exceeding thresholds) to obtain the individual impact value of the abnormal employee on that member. Subsequently, the individual impact values ​​of all associated members are aggregated and weighted according to the total number of team members and the importance of the associated members' roles in the team (e.g., core business personnel, team managers) to finally obtain the potential impact value of the abnormal employee on the team's overall mood. For example, if the abnormal employee is a core team member with close work relationships with most members (high emotion transmission coefficient) and has a high degree of emotional abnormality, the calculated potential impact value will be significantly higher than the impact value when ordinary members have abnormal emotions, accurately reflecting the scope and severity of emotional risk within the team.

[0070] Once it is determined that the impact of emotional risks on the team exceeds a controllable range, targeted intervention strategies are developed to curb the spread of emotional risks and improve the team's emotional atmosphere. The core of differentiated intervention programs is to develop tiered strategies based on the different roles and degrees of impact of the intervention recipients. For the employee at the source of the emotional abnormality, the intervention plan focuses on individual emotional guidance and problem-solving. Based on the possible triggers for the emotional abnormality (such as excessive workload or overly difficult tasks), it automatically pushes one-on-one communication appointment reminders to managers and recommends suitable emotional regulation resources (such as stress relief courses and psychological counseling). If correlation data shows that the emotional abnormality is related to a specific work task, the employee's task priority will be adjusted or collaborative support will be assigned. For related members affected by the emotional transmission, the intervention plan emphasizes emotional prevention and atmosphere guidance. Positive work atmosphere-building content is pushed through team announcements, and a list of related members is sent to managers, prompting them to pay attention to their emotional changes. Small team communication meetings are held when necessary to promptly resolve potential negative emotional transmission. For the entire team, the intervention plan aims to optimize the emotional atmosphere. If the potential impact value is significantly high, non-urgent high-intensity work tasks are suspended, short team relaxation activities are organized, or office environment parameters are adjusted (such as reducing noise and optimizing lighting) to reduce the breeding ground for negative emotional spread from the environmental and task arrangement perspectives. The entire intervention plan generation process requires no manual intervention and seamlessly integrates with the company's existing workflow, ensuring that intervention measures can be implemented quickly and effectively reduce the impact of emotional risks on team performance.

[0071] In one embodiment, the method further includes: The team work relationship graph is mapped to an undirected work relationship graph. The variance of the overall emotional index of employees corresponding to each node in the team work relationship graph is mapped to the nodes of the undirected graph, and the nodes are assigned weights to obtain the undirected graph of fluctuation. The nodes of the fluctuating undirected graph are traversed, and the fluctuation variance of each node is concatenated with the node weight and converted into a hexadecimal encoded string. The hexadecimal encoded string is used as the initial input, and combined with the environmental parameter sequence composed of the environmental parameters of the day, to form the original seed pool. The original seed pool is iteratively processed, and the number of iterations is set to a preset multiple of the number of current employees in the team. In each iteration, the fluctuation variance of the nodes of the fluctuating undirected graph is used as the fine-tuning amount of the chaos parameter to obtain the chaotic confusion seed. Based on the team's average comprehensive sentiment index for the day, the chaotic confusion seed is segmented to obtain segmented confusion seeds; after inserting extended parameters into each segmented confusion seed, they are hashed separately and combined into a hash string. The overall sentiment index of each employee in the team is mapped to an integer, and the bits corresponding to each integer in the hash string are extracted and combined to form a key for encrypting and storing the relevant data.

[0072] In this embodiment, firstly, the team's working relationship graph is mapped into an undirected working relationship graph, visually representing the collaborative connections between employees. Based on this, the variance of the comprehensive emotional index fluctuation for each employee corresponding to each node is extracted. This variance reflects the stability of the employee's emotional state; greater fluctuation indicates greater emotional instability. This variance is used as a node attribute of the undirected graph, and each node is assigned a weight. The weight value is typically positively correlated with the employee's collaboration frequency and role importance within the team, ultimately forming a fluctuating undirected graph that integrates working relationship and emotional fluctuation characteristics. Through this transformation, the initial data for key generation is deeply bound to the team's actual business characteristics and employee emotional states.

[0073] Next, the nodes of the fluctuating undirected graph are traversed, following a descending order of node weights. During this traversal, the fluctuation variance and node weight of each node are concatenated in a fixed format to form a continuous feature data chain. This data chain is then converted into a hexadecimal encoded string. Encoding compresses the original data volume and transforms the numerical information into a character format suitable for subsequent encryption operations, providing standardized input for constructing the key seed.

[0074] Next, the randomness and complexity of the key seed are enhanced by fusing multi-dimensional dynamic data with chaotic processing. First, the generated hexadecimal encoded string is used as the core input and fused with the daily environmental parameter sequence. This sequence consists of real-time data on office noise levels, light intensity, temperature, humidity, and personnel density in a fixed order. By introducing dynamically changing environmental factors, the uncertainty of the seed pool is increased. The two are then concatenated in an alternating order of the feature encoded string and the environmental parameter sequence to form a variable-length original seed pool.

[0075] Subsequently, the original seed pool undergoes chaotic iterative processing: the number of iterations is set to a preset multiple (e.g., 3 times) of the current number of employees in the team, ensuring that the number of iterations dynamically adapts to the team size; each iteration uses a chaotic mapping algorithm, and the variance of the corresponding node in the fluctuating undirected graph is used as the fine-tuning amount of the chaotic parameters, so that the iteration process dynamically changes with the emotional fluctuation characteristics of employees; the greater the emotional fluctuation of a node, the greater the adjustment of the chaotic parameters. Through multiple rounds of chaotic iteration, the characteristics of the original seed pool are fully obfuscated, generating a chaotic obfuscated seed with high randomness, breaking the linear correlation between data, and providing a secure foundation for subsequent key generation.

[0076] Furthermore, by combining the team's overall emotional characteristics with segmented computation, the complexity of the key material is further enhanced. Based on the team's average daily comprehensive emotional index, the chaotic confusion seed is divided into several segments; the higher the average, the more segments there are, thus linking the segmentation logic with the team's overall emotional state. Each segment corresponds to a segmental confusion seed, and extended parameters are inserted into each segment. The extended parameters are derived from the average edge weights of the fluctuating undirected graph, reflecting the overall closeness of the team's collaboration.

[0077] Each segment of the obfuscation seed after inserting extended parameters is hashed separately. The hashing process uses a mutation hashing algorithm, which avoids the security risks associated with a fixed algorithm pattern by adjusting the initial constants and round function parameters of the standard algorithm. The hash results of all segments are then concatenated in sequence to form a complete hash string, so that the final key material simultaneously incorporates team emotional characteristics, collaborative relationships, and chaotic randomness.

[0078] Then, by mapping individual employee emotional characteristics, key information is extracted from the hash string to form the final key. The overall emotional index of each employee in the team is mapped to a specific integer. The mapping rule is based on the normalization and interval division of the index value, ensuring that each employee's emotional state corresponds to a unique integer identifier. Based on these integers, the corresponding bits in the hash string are located, and these bits are sorted and combined according to the closeness of the collaborative relationship between employees to form the final encryption key.

[0079] The key generation process described above incorporates multi-dimensional business characteristics such as team working relationships, employee emotional fluctuations, and environmental parameters. It also ensures encryption security through chaotic processing and hash operations, deeply coupling the key with business logic. This effectively improves the security of encrypted storage of relevant data while reducing the risk of being cracked.

[0080] In one embodiment, the method further includes: Add the characters representing the overall emotional index of each employee in the team to each node of the undirected graph in sequence. Based on the comprehensive emotional index of each employee in the team, the average emotional tendency and average emotional intensity of the team are calculated. Retrieve the preset encoding table; the preset encoding table includes an index character column and an encoded character column; Based on the average emotional tendency, the index character column of the preset encoding table is adjusted by character swapping; based on the average emotional intensity, the shift step is calculated, and the entire encoded character column of the preset encoding table is cyclically shifted by the corresponding step to obtain the first encoding table; Based on the number of edges in the team's working relationship graph, the specified characters in the coded character column of the first encoding table are replaced to obtain the second encoding table; The undirected graph and the second encoding table are superimposed on the same layer. The characters that overlap with the encoded character columns of the second encoding table in the undirected graph are concatenated and replaced in the corresponding encoded character columns to obtain the third encoding table. The relevant data is encoded and stored based on the third encoding table.

[0081] In this embodiment, firstly, employee emotional characteristics are associated with an undirected graph structure, providing a foundation for subsequent encoding table mutations. The characters contained in the comprehensive emotional index of each employee in the team are added sequentially to the corresponding nodes of the undirected graph, according to the order of the employees' nodes in the team's work relationship graph. Each node stores one character of its corresponding employee's comprehensive emotional index.

[0082] Next, by quantifying the overall emotional state of the team, a basis for adjusting the initial variation of the coding table is provided. Based on the comprehensive emotional index of each employee in the team, the mean of the team's emotional tendency and the mean of its emotional intensity are calculated. The mean of emotional tendency reflects the team's overall emotional bias (such as a bias towards positive or negative), while the mean of emotional intensity reflects the intensity of the team's overall emotional expression. These two, as core quantitative indicators of the team's emotional characteristics, will be used to drive the dynamic adjustment of the coding table, making the variation of the coding table closely related to the team's emotional state.

[0083] Next, a preset encoding table is obtained. This encoding table adopts a two-dimensional structure, containing an index character column and an encoded character column. The index character column consists of a series of base characters used for positioning, while the encoded character column contains encoded mapping characters that correspond one-to-one with the index characters. Together, they constitute the basic character mapping relationship, providing the original template for subsequent variation adjustments. The preset encoding table is usually designed based on common encoding rules (such as Base64) to ensure that the initial mapping relationship has basic encoding functions.

[0084] Then, by introducing team emotional characteristics, the preset coding table undergoes initial variation. Based on the calculated mean of emotional tendency, the index character column of the preset coding table is adjusted by character swapping: if the mean of emotional tendency is positive, the even-positioned character in the index character column is swapped with the character preceding it; if it is negative, the odd-positioned character is swapped with the character preceding it, so that the order of the index characters dynamically changes with the team's emotional tendency. Simultaneously, a shift step size is calculated based on the mean of emotional intensity, rounded down by a preset ratio (e.g., 1 / 10). The entire coding character column of the preset coding table is cyclically shifted according to this step size, so that the positional distribution of the coding characters is correlated with the team's emotional intensity. After the index character swapping and coding character shifting, the first coding table is formed, completing the first dynamic adjustment of the coding table.

[0085] Furthermore, based on the characteristics of team working relationships, a second mutation is performed on the first encoding table. According to the total number of edges in the team working relationship graph, the specified characters to be replaced in the encoded character column of the first encoding table are determined: the position of the specified character is an integer multiple of the total number of edges (e.g., if the number of edges is n, then the characters at positions n, 2n, 3n, etc., are the specified characters). The replacement rule is to add the ASCII code value of the specified character to the corresponding number of edges to form the second encoding table.

[0086] Next, the encoding table undergoes three mutations through spatial overlay interaction between the undirected graph and the encoding table. The constructed undirected graph incorporating the sentiment index characters is scaled proportionally and overlaid on the same layer as the generated second encoding table, causing spatial overlap between the node characters of the undirected graph and the encoded character columns of the encoding table. Overlapping characters between the node characters in the undirected graph and the encoded character columns of the second encoding table are identified, and these overlapping characters are concatenated and replaced in their corresponding encoded character columns, ensuring that the character distribution in the encoding table is simultaneously influenced by both the undirected graph topology and the sentiment characters. After character swapping, a third encoding table is formed. In other embodiments, deduplication and interpolation processing can also be performed on the encoded character columns of the third encoding table.

[0087] Finally, employee-related data is segmented into character sequences according to a preset format. Using the index character column of the third encoding table as a reference, the corresponding position of each character in the index column is found, and the character in the encoding character column corresponding to that position is used as the encoding result. After mapping all characters, they are concatenated to form complete encoded data, which is then stored in a designated medium. Because the third encoding table incorporates dynamic team characteristics and undergoes multiple rounds of mutation, the uniqueness and security of the encoding results are ensured. At the same time, the deep coupling between the encoding process and the team's business characteristics also enhances the data storage's resistance to hacking.

[0088] Reference Figure 2 In another embodiment of the present invention, an employee sentiment index assessment device based on multimodal data is also provided, comprising: The data acquisition unit is used to collect relevant data on employees' multimodal behavior; the relevant data includes employees' voice data, text data from enterprise communication tools, video data captured by visual behavior, and employees' physiological signals. The alignment unit is used to extract features from the relevant data and perform cross-modal alignment to obtain multimodal technical features; The calculation unit is used to calculate the data quality coefficient and scene correlation coefficient of each modality in order to calculate the final dynamic weight of each modality. The analysis unit is used to fuse multimodal technical features based on the final dynamic weights of each modality and analyze them to obtain the comprehensive emotional index of the employee; the comprehensive emotional index includes emotional tendency, emotional intensity, and stress index.

[0089] In this embodiment, the specific implementation of each unit in the above device embodiment is described in the above method embodiment, and will not be repeated here.

[0090] Reference Figure 3 This invention also provides a computer device, which can be a server, and its internal structure can be as follows: Figure 3 As shown, the computer device includes a processor, memory, display screen, input device, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database stores the data corresponding to this embodiment. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements the above-described method.

[0091] Those skilled in the art will understand that Figure 3The structures shown are merely block diagrams of some structures related to the present invention and do not constitute a limitation on the computer devices on which the present invention is applied.

[0092] An embodiment of the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described method. It is understood that the computer-readable storage medium in this embodiment can be a volatile readable storage medium or a non-volatile readable storage medium.

[0093] In summary, the employee emotional index assessment method and apparatus based on multimodal data provided in this embodiment of the invention includes: collecting relevant multimodal data of employees; the relevant data includes employees' voice data, text data from enterprise communication tools, video data captured by visual behavior, and employees' physiological signals; extracting features from the relevant data and performing cross-modal alignment to obtain multimodal technical features; calculating the data quality coefficient and scene correlation coefficient of each modality to calculate the final dynamic weight of each modality; fusing the multimodal technical features based on the final dynamic weight of each modality and analyzing them to obtain the comprehensive emotional index of the employee; the comprehensive emotional index includes emotional tendency, emotional intensity, and stress index. In this invention, by comprehensively analyzing data from multiple modalities to obtain the comprehensive emotional index of employees, the shortcomings of current employee emotional assessment methods, such as insufficient comprehensiveness and low accuracy, are overcome.

[0094] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the present invention and embodiments can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual-rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM, etc.

[0095] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, apparatus, article, or method that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, apparatus, article, or method. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, apparatus, article, or method that includes that element.

[0096] The above description is only a preferred embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.

Claims

1. A method for evaluating employee sentiment index based on multimodal data, characterized in that, Includes the following steps: Collect relevant data on employees' multimodal behavior; the relevant data includes employees' voice data, text data from enterprise communication tools, video data captured by visual behavior, and employees' physiological signals; Feature extraction and cross-modal alignment are performed on the relevant data to obtain multimodal technology features; Calculate the data quality coefficient and scene correlation coefficient for each modality to determine the final dynamic weight of each modality; The multimodal technical features are fused and analyzed based on the final dynamic weights of each modality to obtain the comprehensive emotional index of the employee; the comprehensive emotional index includes emotional tendency, emotional intensity, and stress index.

2. The employee sentiment index assessment method based on multimodal data according to claim 1, characterized in that, Feature extraction is performed on the relevant data, including: Acoustic features such as speech rate, pitch, and pause frequency are extracted from real-time captured conference / call audio. Analyze the sentiment of text data in enterprise communication tools and identify semantic features in combination with context; Analyze the implicit emotional features of facial micro-expressions and body language in video data captured by cameras; Analyze the physiological characteristics of physiological signals collected by wearable devices.

3. The employee sentiment index assessment method based on multimodal data according to claim 1, characterized in that, Calculate the data quality coefficients for each modality, including: The data quality coefficient of speech data modality is calculated based on the signal-to-noise ratio of speech data and the integrity of speech activity detection. The data quality coefficient of the video data modality is calculated based on the area of ​​the camera obstruction and the proportion of the effective area. The text confidence score and special symbol processing status were used to calculate the data quality coefficient of the text data modality. The data quality coefficient of physiological signal modality is calculated based on the stability of physiological signals and the data missing rate.

4. The employee sentiment index assessment method based on multimodal data according to claim 1, characterized in that, The calculation of the scene relevance coefficient includes: Different correlation coefficients are preset for each modal data in different scenario types; the scenario types include at least remote meetings, independent office work, and team collaborative discussions. Based on the current scene type when collecting relevant data, determine the scene correlation coefficient for each modality.

5. The employee sentiment index assessment method based on multimodal data according to claim 1, characterized in that, The formula for calculating the final dynamic weights of each modality is as follows: ; in, , These are the data quality coefficients for the i-th and j-th modes, respectively. , These are the scene correlation coefficients for the i-th and j-th modalities, respectively. It is the final dynamic weight of the i-th mode.

6. The employee sentiment index assessment method based on multimodal data according to claim 1, characterized in that, After obtaining the employee's overall emotional index, the following is included: Based on the comprehensive emotion index, the emotion type is determined. Based on the emotion type, combined with the employee's current overtime hours and daily work report content, a coping strategy is generated. The coping strategy includes automatically triggering workflows to adjust task priorities and recommending rest intervals. Establish a correlation model between the comprehensive sentiment index and performance indicators to identify employees with a high risk of turnover.

7. The employee sentiment index assessment method based on multimodal data according to claim 1, characterized in that, The method further includes: Real-time removal of identity information from voice and text data; The behavioral data in the relevant data is processed using k-anonymization technology.

8. The employee sentiment index assessment method based on multimodal data according to claim 1, characterized in that, After obtaining the employee's overall emotional index, the following is also included: Obtain environmental parameters, including office noise level, light intensity, temperature and humidity parameters, and team personnel density; Obtain the interference coefficients corresponding to each of the environmental parameters; wherein, a correlation analysis between the environmental parameters and the emotion recognition error is performed in advance to generate the interference coefficients corresponding to each environmental parameter. The comprehensive emotion index is calibrated based on the interference coefficients of various environmental parameters to obtain the final comprehensive emotion index.

9. The employee sentiment index assessment method based on multimodal data according to claim 1, characterized in that, After obtaining the employee's overall emotional index, the following is also included: Construct a team working relationship graph and calculate the emotion transmission coefficient between employees based on the frequency of collaborative work and communication density of employees in the team; Real-time monitoring of employees with abnormal comprehensive emotional index in the team, and calculation of the potential impact of abnormal employees on the overall team mood based on the emotional transmission coefficient among employees; When the potential impact value exceeds a preset threshold, a differentiated intervention plan is automatically generated.

10. A device for evaluating employee sentiment index based on multimodal data, characterized in that, include: The data acquisition unit is used to collect relevant multimodal data from employees. The relevant data includes employees' voice data, text data from enterprise communication tools, video data captured by visual behavior, and employees' physiological signals. The alignment unit is used to extract features from the relevant data and perform cross-modal alignment to obtain multimodal technical features; The calculation unit is used to calculate the data quality coefficient and scene correlation coefficient of each modality in order to calculate the final dynamic weight of each modality. The analysis unit is used to fuse multimodal technical features based on the final dynamic weights of each modality and analyze them to obtain the comprehensive emotional index of the employee; the comprehensive emotional index includes emotional tendency, emotional intensity, and stress index.

Citation Information

Patent Citations

  • Camera-based driver emotion recognition method assisted by 5G vehicle-mounted network cloud

    CN111444863A

  • Customer service staff emotion recognition method and system based on multi-modal data fusion

    CN118820840A

  • Dynamic calibration method and system of vehicle-mounted emotion recognition system

    CN120448917A

  • Emotion estimation server device, emotion estimation method, presentation device and emotion estimation system

    JP2018138155A

  • Computer-Based Systems and Methods for Sentiment Analysis

    US20230119405A1

Cited By

  • Method and system for identifying emotional state of operator in main control room of nuclear power plant

    CN121579952A