Emotion recognition model establishment method and system based on multi-modal data

Through multimodal data analysis, a multimodal data emotion recognition model is constructed, which solves the problem of low accuracy in traditional single-modal data recognition, and achieves more accurate emotion recognition, providing a reliable decision-making basis for enterprise management.

CN120145309APending Publication Date: 2025-06-13BEIJING TAIJI INFORMATION SYST TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510252189.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Traditional emotion recognition methods rely on a single data source, resulting in the accuracy of emotion recognition being affected by factors such as environmental noise, cultural differences, and individual physiological characteristics, making it difficult to accurately capture the true emotions of employees.

Method used

By obtaining multimodal data (such as speech, text, facial expression), constructing a sub-multimodal data model, calculating a sequence of modal data analysis values, and determining the modal data analysis factors through random allocation and analysis, finally establishing a multimodal data emotion recognition model.

Benefits of technology

It realizes more accurate emotional recognition, can comprehensively and accurately reflect the true emotional state of employees, and provides a reliable basis for corporate management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120145309A_ABST
    Figure CN120145309A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of model establishment, and discloses an emotion recognition model establishment method and system based on multi-modal data, and the method comprises the steps: obtaining a plurality of pieces of multi-modal data, outputting a modal data analysis value of sub-multi-modal data corresponding to each piece of multi-modal data based on a sub-multi-modal data model, and obtaining a plurality of modal data analysis value sequences; randomly distributing the modal data analysis value sequence to the modal data analysis value point diagram, and calculating modal data analysis factors of the modal data analysis value sequence; carrying out numerical value sorting on all modal data analysis factors, determining a plurality of candidate sub-multi-modal data, and taking the candidate sub-multi-modal data as a candidate sub-multi-modal data feature set; according to the method, the candidate sub-multi-modal data feature set is divided to obtain the model establishment set, the multi-modal data emotion recognition model is established, and the multi-modal data is utilized to construct the more accurate emotion recognition model, so that the real emotion state of the employee can be reflected more comprehensively and accurately, and a reliable decision basis is provided for enterprise management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of model establishment, and in particular, to a method and system for establishing an emotion recognition model based on multimodal data. Background Art

[0002] In today's enterprise management and human resources fields, accurately understanding the emotional state of employees is of crucial significance for improving work efficiency, optimizing team collaboration, and ensuring the mental health of employees.

[0003] Traditional emotion recognition methods usually rely on a single data source. For example, they infer the emotions of employees only through single-modal data such as voice analysis, text content processing, or facial expression recognition. For instance, the emotion recognition model based on voice mainly focuses on the features in the voice signal, such as pitch, speech rate, timbre, etc. However, the voice may be interfered by various factors such as environmental noise, speaker accent, and tone habits, resulting in certain limitations in the accuracy of emotion recognition. The emotion recognition model of text analysis identifies emotions by performing semantic analysis on the text content written by employees (such as emails, work reports, social media posts, etc.). However, there may be problems such as semantic ambiguity and implicit expression in the text information, and different people may have different understandings and emotional tendencies towards the same words or sentences, making it difficult to accurately capture the true emotions of employees solely relying on text analysis. Although facial expression recognition can directly obtain the facial expression information of employees, facial expressions may also be affected by factors such as disguise, cultural differences, and individual physiological characteristics. For example, some employees may conceal their true emotions for social courtesy, or the same expression may have different meanings under different cultural backgrounds, leading to misjudgment. Summary of the Invention

[0004] Embodiments of the present invention provide a method and system for establishing an emotion recognition model based on multimodal data. The present invention uses multimodal data to construct a more accurate emotion recognition model, which can more comprehensively and accurately reflect the true emotional state of employees and provide a reliable decision-making basis for enterprise management.

[0005] To achieve the above object, the present invention provides a method for establishing an emotion recognition model based on multimodal data, including: Obtaining a plurality of multimodal data, respectively outputting the modal data analysis values of the sub-multimodal data corresponding to each multimodal data based on the corresponding sub-multimodal data models, and processing all the modal data analysis values to obtain a plurality of modal data analysis value sequences; Randomly allocating the modal data analysis value sequences to a modal data analysis value dot plot, analyzing the modal data analysis value dot plot, and calculating the modal data analysis factors of the modal data analysis value sequences based on the analysis results; Sort the numerical values of all modal data analysis factors, determine multiple candidate sub-multi-modal data based on the sorting result, and use the candidate sub-multi-modal data as the candidate sub-multi-modal data feature set; Divide the candidate sub-multi-modal data feature set to obtain a model building set, and establish a multi-modal data emotion recognition model based on the model building set, where the model building set includes a training feature set and a validation feature set.

[0006] Further, when analyzing the modal data analysis value point graph and calculating the modal data analysis factors of the modal data analysis value sequence based on the analysis result, it includes: Determine the maximum modal data analysis value and the minimum modal data analysis value from the modal data analysis value sequence, and calculate the labeled modal data analysis value Q according to the maximum modal data analysis value and the minimum modal data analysis value, where Q = (q1 - q2) / 2, q1 is the maximum modal data analysis value, and q2 is the minimum modal data analysis value; Mark the labeled modal data analysis value on the modal data analysis value point graph to determine the data marking point, and divide the modal data analysis values between the initial modal data analysis value and the data marking point on the modal data analysis value point graph into the first modal data analysis value sequence; Divide the modal data analysis values between the end modal data analysis value and the data marking point on the modal data analysis value point graph into the second modal data analysis value sequence; Calculate the first modal data analysis factor of the modal data analysis value sequence based on the first modal data analysis value sequence; Calculate the second modal data analysis factor of the modal data analysis value sequence based on the second modal data analysis value sequence; Calculate the sum value of the first modal data analysis factor and the second modal data analysis factor, and use it as the modal data analysis factor of the modal data analysis value sequence.

[0007] Further, when calculating the first modal data analysis factor of the modal data analysis value sequence based on the first modal data analysis value sequence, it includes: Randomly match the modal data analysis values in the first modal data analysis value sequence in pairs to determine multiple modal data analysis value combinations; Judge whether the two modal data analysis values in the modal data analysis value combination are equal. If so, delete the corresponding modal data analysis value combination; Extract the remaining modal data analysis value combinations, calculate the modal data analysis value sum of two modal data analysis values ​​in the modal data analysis value combination, and select the maximum modal data analysis value sum and the minimum modal data analysis value sum; determining a modal data analysis value variance of all modal data analysis values ​​and values, and generating a first modal data analysis value code for all modal data analysis values ​​and values ​​that are greater than the modal data analysis value variance; generating a second modal data analysis value code for all modal data analysis values ​​and values ​​that are less than or equal to the modal data analysis value variance; A first modal data analysis factor of the modal data analysis value sequence is calculated based on the maximum modal data analysis value sum, the minimum modal data analysis value sum, the first modal data analysis value code, and the second modal data analysis value code.

[0008] Further, when calculating the first modal data analysis factor of the modal data analysis value sequence according to the maximum modal data analysis value and value, the minimum modal data analysis value and value, the first modal data analysis value code and the second modal data analysis value code, it includes: The first modal data analysis factor of the modal data analysis value sequence is calculated according to the following formula: ; Where w is the first modal data analysis factor of the modal data analysis value sequence, y is the number of modal data analysis value combinations, r1 is the first calculation coefficient, r2 is the second calculation coefficient, t u are the modal data analysis value and value corresponding to the u-th modal data analysis value combination, p1 is the maximum modal data analysis value and value, p2 is the minimum modal data analysis value and value, d1 is the number of first modal data analysis value codes, and d2 is the number of second modal data analysis value codes.

[0009] Further, when calculating the second modal data analysis factor of the modal data analysis value sequence based on the second modal data analysis value sequence, it includes: extracting the same modal data analysis value from the second modal data analysis value sequence, and obtaining a plurality of sub-modal data analysis value sequences; Counting the number of first modal data analysis values ​​of the submodal data analysis value sequence; Extracting a modal data analysis value from each of all sub-modal data analysis value sequences, and calculating a first modal data analysis value and a value; Obtaining a preset modal data analysis value, eliminating all sub-modal data analysis value sequences that are less than the preset modal data analysis value, and counting the number of second modal data analysis values ​​of the remaining sub-modal data analysis value sequences; Extract a modal data analysis value from the remaining sub-modal data analysis value sequences respectively, and calculate the second modal data analysis value and value; Calculate the second modal data analysis factor of the modal data analysis value sequence according to the number of the first modal data analysis values, the number of the second modal data analysis values, the sum value of the first modal data analysis values and the sum value of the second modal data analysis values.

[0010] Further, when calculating the second modal data analysis factor of the modal data analysis value sequence according to the number of the first modal data analysis values, the number of the second modal data analysis values, the sum value of the first modal data analysis values and the sum value of the second modal data analysis values, it includes: Calculate the second modal data analysis factor of the modal data analysis value sequence according to the following formula: ; where a is the second modal data analysis factor of the modal data analysis value sequence, n is the number of modal data analysis values in the modal data analysis value sequence, b j is the j-th modal data analysis value in the modal data analysis value sequence, v is the variance of the modal data analysis value sequence, g1 is the number of the first modal data analysis values, g2 is the number of the second modal data analysis values, k1 is the sum value of the first modal data analysis values, and k2 is the sum value of the second modal data analysis values.

[0011] Further, when sorting the numerical magnitudes of all the modal data analysis factors and determining multiple candidate sub-multi-modal data based on the sorting result, it includes: Select the modal data analysis value sequences corresponding to the first c1 modal data analysis factors as candidate modal data analysis value sequences; Select the first c2 sub-multi-modal data from each candidate modal data analysis value sequence as the candidate sub-multi-modal data.

[0012] Further, when partitioning the candidate sub-multi-modal data feature set to obtain a model building set and building a multi-modal data emotion recognition model based on the model building set, it includes: Construct an initial multi-modal data emotion recognition model based on the training feature set by using a preset machine learning algorithm; Use the validation feature set to validate the initial multi-modal data emotion recognition model, and when the validation result meets the preset validation result, obtain the multi-modal data emotion recognition model; When the validation result does not meet the preset validation result, adjust the model parameters according to the validation result to obtain the multi-modal data emotion recognition model.

[0013] Further, the preset machine learning algorithm includes one or more of a decision tree algorithm, a support vector machine algorithm, and a neural network algorithm; The model parameters include the depth of the decision tree, the penalty parameter of the support vector machine, the number of hidden layer nodes of the neural network, and the learning rate.

[0014] To achieve the above object, the present invention also provides a system for establishing an emotion recognition model based on multi-modal data, including: A data processing module, configured to obtain a plurality of multi-modal data, output the modal data analysis values of the sub-multi-modal data corresponding to each multi-modal data based on the corresponding sub-multi-modal data model, process all the modal data analysis values, and obtain a plurality of modal data analysis value sequences; A factor calculation module, configured to randomly assign the modal data analysis value sequences to a modal data analysis value dot plot, analyze the modal data analysis value dot plot, and calculate the modal data analysis factors of the modal data analysis value sequences based on the analysis results; A data determination module, configured to sort all the modal data analysis factors by numerical magnitude, determine a plurality of candidate sub-multi-modal data based on the magnitude sorting result, and use the candidate sub-multi-modal data as a candidate sub-multi-modal data feature set; A model establishment module, configured to divide the candidate sub-multi-modal data feature set to obtain a model establishment set, and establish a multi-modal data emotion recognition model based on the model establishment set, where the model establishment set includes a training feature set and a verification feature set.

[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention discloses a method and a system for establishing an emotion recognition model based on multi-modal data, which obtain a plurality of multi-modal data, output the modal data analysis values of the sub-multi-modal data corresponding to each multi-modal data based on the sub-multi-modal data model, and obtain a plurality of modal data analysis value sequences; randomly assign the modal data analysis value sequences to a modal data analysis value dot plot, calculate the modal data analysis factors of the modal data analysis value sequences; sort all the modal data analysis factors by numerical magnitude, determine a plurality of candidate sub-multi-modal data, and use them as a candidate sub-multi-modal data feature set; divide the candidate sub-multi-modal data feature set to obtain a model establishment set, and establish a multi-modal data emotion recognition model. Using multi-modal data to construct a more accurate emotion recognition model can more comprehensively and accurately reflect the true emotion state of employees, and provide a reliable decision-making basis for enterprise management. Description of the Drawings

[0016] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become apparent to those of ordinary skill in the art. The drawings are only for the purpose of illustrating the preferred embodiments and are not considered to be a limitation of the present invention. Also, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings: Figure 1 FIG. shows a schematic flow chart of a method for establishing an emotion recognition model based on multi-modal data in an embodiment of the present invention; Figure 2 FIG. shows a schematic structural diagram of a system for establishing an emotion recognition model based on multi-modal data in an embodiment of the present invention. Specific Embodiments

[0017] The following further describes in detail the specific embodiments of the present invention with reference to the drawings and embodiments. The following embodiments are used to illustrate the present invention but are not intended to limit the scope of the present invention.

[0018] In the description of the present application, it should be understood that the terms "center", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present application.

[0019] The terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present application, unless otherwise stated, the meaning of "a plurality" is two or more.

[0020] In the description of the present application, it should be noted that unless otherwise clearly defined and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific situations.

[0021] The following is a description of the preferred embodiments of the present invention with reference to the drawings.

[0022] As Figure 1 shown, the embodiments of the present invention disclose a method for establishing an emotion recognition model based on multi-modal data, including: S110: Obtain multiple multimodal data, respectively output the modal data analysis values of the sub-multimodal data corresponding to each multimodal data based on the corresponding sub-multimodal data models, process all the modal data analysis values, and obtain multiple sequences of modal data analysis values; In this embodiment, the multimodal data includes voice data, text data, facial expression data, etc. Among them, each multimodal data corresponds to different sub-multimodal data. As described above, the voice data corresponds to pitch, speech rate, timbre, etc., the text data corresponds to emails, work reports, social media posts, etc., and the facial expressions correspond to eyebrows, eyes, nose, mouth, etc., which are not shown one by one here.

[0023] In this embodiment, the sub-multimodal data model refers to the one trained according to each sub-multimodal data. For example, the training steps of the sub-multimodal data model corresponding to pitch are as follows: collect multiple historical pitch data, construct a data set according to the historical pitch data; sample the data set according to a preset ratio to obtain a training subset and a test subset; obtain a pre-selected neural network model, and perform iterative training on the neural network model according to the training subset, evaluate the iteratively trained neural network model according to the test subset, and obtain the sub-multimodal data model corresponding to pitch. The sub-multimodal data model corresponding to pitch can output the modal data analysis value corresponding to pitch, and the others are not shown one by one.

[0024] In this embodiment, divide the modal data analysis values of the sub-multimodal data corresponding to each multimodal data into a sequence, that is, obtain a sequence of modal data analysis values. Therefore, multiple sequences of modal data analysis values can be obtained.

[0025] S120: Randomly allocate the sequence of modal data analysis values to the modal data analysis dot plot, analyze the modal data analysis dot plot, and calculate the modal data analysis factor of the sequence of modal data analysis values based on the analysis result; In some embodiments of the present application, when analyzing the modal data analysis dot plot and calculating the modal data analysis factor of the sequence of modal data analysis values based on the analysis result, it includes: Determine the maximum modal data analysis value and the minimum modal data analysis value from the sequence of modal data analysis values, calculate the marked modal data analysis value Q according to the maximum modal data analysis value and the minimum modal data analysis value, where Q = (q1 - q2) / 2, q1 is the maximum modal data analysis value, and q2 is the minimum modal data analysis value; Annotate the modal data analysis values on the modal data analysis value dot plot, determine the data annotation points, and divide the modal data analysis values between the initial modal data analysis value and the data annotation points on the modal data analysis value dot plot into a first modal data analysis value sequence; Divide the modal data analysis values between the terminal modal data analysis value and the data annotation points on the modal data analysis value dot plot into a second modal data analysis value sequence; Calculate the first modal data analysis factor of the modal data analysis value sequence based on the first modal data analysis value sequence; Calculate the second modal data analysis factor of the modal data analysis value sequence based on the second modal data analysis value sequence; Calculate the sum value of the first modal data analysis factor and the second modal data analysis factor, and use it as the modal data analysis factor of the modal data analysis value sequence.

[0026] In this embodiment, neither the first modal data analysis value sequence nor the second modal data analysis value sequence includes the modal data analysis value corresponding to the data annotation point.

[0027] The beneficial effects of the above technical solution are as follows: The present invention calculates the first modal data analysis factor of the modal data analysis value sequence based on the first modal data analysis value sequence; calculates the second modal data analysis factor of the modal data analysis value sequence based on the second modal data analysis value sequence, laying a foundation for determining the modal data analysis factor and ensuring the calculation accuracy of the modal data analysis factor.

[0028] In some embodiments of the present application, when calculating the first modal data analysis factor of the modal data analysis value sequence based on the first modal data analysis value sequence, it includes: Randomly pair the modal data analysis values in the first modal data analysis value sequence to determine multiple modal data analysis value combinations; Judge whether the two modal data analysis values in the modal data analysis value combination are equal. If so, delete the corresponding modal data analysis value combination; Extract the remaining modal data analysis value combinations, calculate the modal data analysis value sum value of the two modal data analysis values in the modal data analysis value combination, and select the maximum modal data analysis value sum value and the minimum modal data analysis value sum value; Determine the modal data analysis value variance of all modal data analysis value sum values, and generate a first modal data analysis value code for all modal data analysis value sum values greater than the modal data analysis value variance; Generate a second modal data analysis value code for all modal data analysis value sum values less than or equal to the modal data analysis value variance; Calculate the first modal data analysis factor of the modal data analysis value sequence according to the maximum modal data analysis value and value, the minimum modal data analysis value and value, the first modal data analysis value coding, and the second modal data analysis value coding.

[0029] In this embodiment, if there is an unmatched modal data analysis value, delete this unmatched modal data analysis value.

[0030] In this embodiment, a modal data analysis value and value can be calculated according to each combination of modal data analysis values.

[0031] The beneficial effects of the above technical solution are as follows: The present invention calculates the first modal data analysis factor of the modal data analysis value sequence according to the maximum modal data analysis value and value, the minimum modal data analysis value and value, the first modal data analysis value coding, and the second modal data analysis value coding, ensuring the calculation accuracy and calculation efficiency of the first modal data analysis factor, and providing a basis for calculating the modal data analysis factor.

[0032] In some embodiments of the present application, when calculating the first modal data analysis factor of the modal data analysis value sequence according to the maximum modal data analysis value and value, the minimum modal data analysis value and value, the first modal data analysis value coding, and the second modal data analysis value coding, it includes: Calculate the first modal data analysis factor of the modal data analysis value sequence according to the following formula: ; where w is the first modal data analysis factor of the modal data analysis value sequence, y is the number of combinations of modal data analysis values, r1 is the first calculation coefficient, r2 is the second calculation coefficient, t u is the modal data analysis value and value corresponding to the u-th combination of modal data analysis values, p1 is the maximum modal data analysis value and value, p2 is the minimum modal data analysis value and value, d1 is the number of first modal data analysis value codings, and d2 is the number of second modal data analysis value codings.

[0033] In some embodiments of the present application, when calculating the second modal data analysis factor of the modal data analysis value sequence based on the second modal data analysis value sequence, it includes: Extract the same modal data analysis values from the second modal data analysis value sequence to obtain multiple sub-modal data analysis value sequences; Count the number of first modal data analysis values in the sub-modal data analysis value sequences; Extract one modal data analysis value from each of all the sub-modal data analysis value sequences and calculate the first modal data analysis value and value; Obtain the preset modal data analysis value, eliminate all sub-modal data analysis value sequences smaller than the preset modal data analysis value, and count the number of second modal data analysis values of the remaining sub-modal data analysis value sequences; Extract one modal data analysis value from each of the remaining sub-modal data analysis value sequences, and calculate the sum value of the second modal data analysis values; Calculate the second modal data analysis factor of the modal data analysis value sequence according to the number of first modal data analysis values, the number of second modal data analysis values, the sum value of the first modal data analysis values, and the sum value of the second modal data analysis values.

[0034] In this embodiment, the preset modal data analysis value set is the standard deviation of all modal data analysis values in the second modal data analysis value sequence.

[0035] The beneficial effects of the above technical solution are: The present invention calculates the second modal data analysis factor of the modal data analysis value sequence according to the number of first modal data analysis values, the number of second modal data analysis values, the sum value of the first modal data analysis values, and the sum value of the second modal data analysis values, which not only ensures the calculation accuracy of the second modal data analysis factor, but also provides another basis for calculating the modal data analysis factor.

[0036] In some embodiments of the present application, when calculating the second modal data analysis factor of the modal data analysis value sequence according to the number of first modal data analysis values, the number of second modal data analysis values, the sum value of the first modal data analysis values, and the sum value of the second modal data analysis values, it includes: Calculate the second modal data analysis factor of the modal data analysis value sequence according to the following formula: ; where a is the second modal data analysis factor of the modal data analysis value sequence, n is the number of modal data analysis values in the modal data analysis value sequence, b j is the j-th modal data analysis value in the modal data analysis value sequence, v is the variance of the modal data analysis value sequence, g1 is the number of first modal data analysis values, g2 is the number of second modal data analysis values, k1 is the sum value of the first modal data analysis values, and k2 is the sum value of the second modal data analysis values.

[0037] S130: Sort all modal data analysis factors by numerical magnitude, and determine multiple candidate sub-multi-modal data based on the sorting result, and use the candidate sub-multi-modal data as the candidate sub-multi-modal data feature set; In some embodiments of the present application, when sorting all modal data analysis factors by numerical magnitude and determining multiple candidate sub-multi-modal data based on the sorting result, it includes: Select the modal data analysis value sequences corresponding to the first c1 modal data analysis factors as the candidate modal data analysis value sequences; Select the first c2 sub-multi-modal data from each candidate modal data analysis value sequence as the candidate sub-multi-modal data.

[0038] In this embodiment, c1 is preferably 8 here, and c2 is preferably 6 here. Specifically, it can also be adaptively adjusted according to actual requirements.

[0039] The beneficial effects of the above technical solutions are as follows: The present invention sorts all modal data analysis factors by numerical size, determines multiple candidate sub-multi-modal data based on the sorting result, and uses the candidate sub-multi-modal data as the candidate sub-multi-modal data feature set, which can select some high-quality data. This means that not all data will be selected, but those data that are most likely to contribute to the accuracy of the emotion recognition model are selected, and some data that affect model establishment are excluded, which can reduce the amount of data that the model needs to process, while retaining the information that is most helpful for model establishment, thereby ensuring the establishment accuracy of the multi-modal data emotion recognition model and avoiding large errors.

[0040] S140: Divide the candidate sub-multi-modal data feature set to obtain a model establishment set, and establish a multi-modal data emotion recognition model based on the model establishment set, where the model establishment set includes a training feature set and a verification feature set.

[0041] In some embodiments of the present application, when dividing the candidate sub-multi-modal data feature set to obtain a model establishment set and establishing a multi-modal data emotion recognition model based on the model establishment set, it includes: Based on the training feature set, construct an initial multi-modal data emotion recognition model using a preset machine learning algorithm; Use the verification feature set to verify the initial multi-modal data emotion recognition model. When the verification result meets the preset verification result, obtain the multi-modal data emotion recognition model; When the verification result does not meet the preset verification result, adjust the model parameters according to the verification result to obtain the multi-modal data emotion recognition model.

[0042] In some embodiments of the present application, the preset machine learning algorithm includes one or more of a decision tree algorithm, a support vector machine algorithm, and a neural network algorithm; The model parameters include the depth of the decision tree, the penalty parameter of the support vector machine, the number of hidden layer nodes of the neural network, and the learning rate.

[0043] In this embodiment, according to the characteristics of multimodal data and the task requirements of emotion recognition, a suitable machine learning algorithm is selected, including one or more of decision tree algorithm, support vector machine algorithm, and neural network algorithm. Model training: The preprocessed training feature set is input into the selected machine learning algorithm for training to learn the parameters and structure of the model. During the training process, techniques such as cross-validation can be used to evaluate the performance of the model and select the optimal model parameters. For example, K-fold cross-validation can be adopted, where the dataset is divided into K parts. Each time, one part is used as the validation set, and the remaining parts are used as the training set. This is repeated multiple times, and the average value is taken as the performance evaluation index of the model. Use the validation feature set to verify the initial multimodal data emotion recognition model. Input the validation set into the trained initial multimodal data emotion recognition model, and calculate performance indicators such as accuracy, recall rate, and F1 value of the model on the validation set. These indicators can reflect the model's recognition ability and accuracy for different emotion categories. For example, the accuracy represents the proportion of samples correctly predicted by the model in the total samples; the recall rate represents the proportion of samples of a certain emotion correctly recognized by the model in the actual samples of that category; the F1 value is the harmonic mean of the accuracy and recall rate, comprehensively considering the precision and recall rate of the model. Compare the performance indicators of the model on the validation set with the preset threshold. If indicators such as the accuracy, recall rate, and F1 value of the model meet the preset requirements, it indicates that the generalization ability of the model is good, and it can be considered that the model has converged to a good state. Determine the current model parameters and structure as the final multimodal data emotion recognition model. This model can be used to perform emotion recognition and classification on new multimodal data. When the verification result does not meet the preset verification result, it is necessary to analyze the verification result to find out the problems existing in the model. Possible reasons include model overfitting, underfitting, unreasonable feature selection, improper model parameter settings, etc. For example, if the model performs well on the training set but poorly on the validation set, it may be that the model overfits the training data, and it is necessary to increase the diversity of the data or adopt techniques such as regularization to prevent overfitting. According to the analysis result, adjust the parameters of the model. The parameters that can be adjusted include the depth of the decision tree, the penalty parameter of the support vector machine, the number of hidden layer nodes and the learning rate of the neural network. The learning rate can be appropriately reduced or the regularization intensity can be increased; if the model is underfitting, the number of network layers or nodes can be increased. After adjusting the parameters, retrain the model and use the validation set for verification again until the verification result meets the preset requirements.

[0044] The beneficial effects of the above technical solution are as follows: By integrating data of multiple modalities such as voice, text, and facial expressions, the model of the present invention can obtain information about employees' emotions from multiple perspectives. Voice data can provide clues such as intonation and speech rate; text data can analyze the content and semantics expressed by employees; facial expressions directly reflect the external emotional expressions of employees. This multi-dimensional information fusion enables the model to more comprehensively understand the emotional state of employees, avoiding the one-sidedness and misleadingness that may be brought by single-modal data. For example, when an employee shows a calm intonation in voice, but expresses some negative emotion words in text, and at the same time shows a trace of anxiety in facial expression, the multi-modal emotion recognition model can integrate this information and more accurately judge that the true emotion of the employee may be a complex mixed emotion, rather than relying solely on single-modal judgment. A more accurate emotion recognition model can provide more detailed and accurate employee emotion information for enterprise managers, helping managers timely understand the working status and psychological needs of employees, so as to take corresponding management measures. For example, when it is found that an employee is in a negative emotional state, the manager can communicate with the employee in time and provide necessary help and support to avoid the negative impact of the employee's emotion problem on work efficiency and team atmosphere. In addition, through long-term analysis and monitoring of employees' emotion data, enterprises can also optimize work processes, improve work environments, increase employees' job satisfaction and loyalty, and thus enhance the overall performance and competitiveness of enterprises.

[0045] To further elaborate on the technical concept of the present invention, the technical solution of the present invention will be described in combination with specific application scenarios.

[0046] Correspondingly, as Figure 2 shown, the present application also provides a system for establishing an emotion recognition model based on multi-modal data, including: A data processing module, configured to obtain multiple multi-modal data, output the modal data analysis values of each multi-modal data corresponding to the sub-multi-modal data based on the corresponding sub-multi-modal data models respectively, and process all the modal data analysis values to obtain multiple modal data analysis value sequences; A factor calculation module, configured to randomly assign the modal data analysis value sequences to a modal data analysis value dot plot, analyze the modal data analysis value dot plot, and calculate the modal data analysis factors of the modal data analysis value sequences based on the analysis results; A data determination module, configured to sort the numerical magnitudes of all the modal data analysis factors, determine multiple candidate sub-multi-modal data based on the sorting result of the magnitudes, and use the candidate sub-multi-modal data as a candidate sub-multi-modal data feature set; A model establishment module is configured to divide the candidate sub-multimodal data feature set to obtain a model establishment set, and establish a multimodal data emotion recognition model based on the model establishment set, wherein the model establishment set includes a training feature set and a validation feature set.

[0047] In the description of the above embodiments, specific features, structures, materials, or characteristics may be combined in any one or more embodiments or examples in a suitable manner.

[0048] Although the present invention has been described above with reference to the embodiments, various improvements can be made to it and components therein can be replaced with equivalents without departing from the scope of the present invention. In particular, as long as there is no structural conflict, the various features in the embodiments disclosed in the present invention can be combined with each other in any way, and the situations of these combinations are not all described in this specification only for the sake of saving space and resources.

[0049] Those of ordinary skill in the art can understand that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for establishing an emotion recognition model based on multimodal data, characterized in that: include: Acquire multiple multimodal data, output modal data analysis values ​​of sub-multimodal data corresponding to each multimodal data based on the corresponding sub-multimodal data model, process all modal data analysis values, and obtain multiple modal data analysis value sequences; Randomly assigning the modal data analysis value sequence to a modal data analysis value point diagram, analyzing the modal data analysis value point diagram, and calculating a modal data analysis factor of the modal data analysis value sequence based on the analysis result; Sorting all modal data analysis factors by numerical value, and determining a plurality of candidate sub-multimodal data based on the sorting result, and using the candidate sub-multimodal data as a candidate sub-multimodal data feature set; The candidate sub-multimodal data feature sets are divided to obtain a model building set, and a multimodal data emotion recognition model is established based on the model building set, wherein the model building set includes a training feature set and a verification feature set.

2. The method for establishing an emotion recognition model based on multimodal data according to claim 1, characterized in that: When analyzing the modal data analysis value point diagram and calculating the modal data analysis factor of the modal data analysis value sequence based on the analysis result, it includes: Determine a maximum modal data analysis value and a minimum modal data analysis value from the modal data analysis value sequence, and calculate a labeled modal data analysis value Q according to the maximum modal data analysis value and the minimum modal data analysis value, wherein Q=(q1-q2) / 2, q1 is the maximum modal data analysis value, and q2 is the minimum modal data analysis value; Annotating the annotated modal data analysis values ​​on the modal data analysis value point graph, determining data annotated points, and dividing the initial modal data analysis values ​​on the modal data analysis value point graph and the modal data analysis values ​​between the data annotated points into a first modal data analysis value sequence; Dividing the end modal data analysis value on the modal data analysis value point graph and the modal data analysis value between the data annotation points into a second modal data analysis value sequence; Calculating a first modal data analysis factor of the modal data analysis value sequence based on the first modal data analysis value sequence; Calculating a second modal data analysis factor of the modal data analysis value sequence based on the second modal data analysis value sequence; A sum of the first modal data analysis factor and the second modal data analysis factor is calculated and used as the modal data analysis factor of the modal data analysis value sequence.

3. The method for establishing an emotion recognition model based on multimodal data according to claim 2, characterized in that: When calculating a first modal data analysis factor of the modal data analysis value sequence based on the first modal data analysis value sequence, the method comprises: Randomly matching modal data analysis values ​​in the first modal data analysis value sequence in pairs to determine a plurality of modal data analysis value combinations; Determining whether two modal data analysis values ​​in the modal data analysis value combination are equal, and if so, deleting the corresponding modal data analysis value combination; Extract the remaining modal data analysis value combinations, calculate the modal data analysis value sum of two modal data analysis values ​​in the modal data analysis value combination, and select the maximum modal data analysis value sum and the minimum modal data analysis value sum; determining a modal data analysis value variance of all modal data analysis values ​​and values, and generating a first modal data analysis value code for all modal data analysis values ​​and values ​​that are greater than the modal data analysis value variance; generating a second modal data analysis value code for all modal data analysis values ​​and values ​​that are less than or equal to the modal data analysis value variance; A first modal data analysis factor of the modal data analysis value sequence is calculated based on the maximum modal data analysis value sum, the minimum modal data analysis value sum, the first modal data analysis value code, and the second modal data analysis value code.

4. The method for establishing an emotion recognition model based on multimodal data according to claim 3, characterized in that: When calculating a first modal data analysis factor of the modal data analysis value sequence according to the maximum modal data analysis value and value, the minimum modal data analysis value and value, the first modal data analysis value code and the second modal data analysis value code, comprising: The first modal data analysis factor of the modal data analysis value sequence is calculated according to the following formula: ; Where w is the first modal data analysis factor of the modal data analysis value sequence, y is the number of modal data analysis value combinations, r1 is the first calculation coefficient, r2 is the second calculation coefficient, t u are the modal data analysis value and value corresponding to the u-th modal data analysis value combination, p1 is the maximum modal data analysis value and value, p2 is the minimum modal data analysis value and value, d1 is the number of first modal data analysis value codes, and d2 is the number of second modal data analysis value codes.

5. The method for establishing an emotion recognition model based on multimodal data according to claim 2, characterized in that: When calculating the second modal data analysis factor of the modal data analysis value sequence based on the second modal data analysis value sequence, the method comprises: extracting the same modal data analysis value from the second modal data analysis value sequence, and obtaining a plurality of sub-modal data analysis value sequences; Counting the number of first modal data analysis values ​​of the submodal data analysis value sequence; Extracting a modal data analysis value from each of all sub-modal data analysis value sequences, and calculating a first modal data analysis value and a value; Obtaining a preset modal data analysis value, eliminating all sub-modal data analysis value sequences that are less than the preset modal data analysis value, and counting the number of second modal data analysis values ​​of the remaining sub-modal data analysis value sequences; extracting a modal data analysis value from each of the remaining sub-modal data analysis value sequences, and calculating a second modal data analysis value and a value; A second modal data analysis factor of the modal data analysis value sequence is calculated based on the number of the first modal data analysis values, the number of the second modal data analysis values, the sum of the first modal data analysis values, and the sum of the second modal data analysis values.

6. The method for establishing an emotion recognition model based on multimodal data according to claim 5, characterized in that: When calculating a second modal data analysis factor of the modal data analysis value sequence according to the number of the first modal data analysis values, the number of the second modal data analysis values, the sum of the first modal data analysis values ​​and the sum of the second modal data analysis values, the method comprises: The second modal data analysis factor of the modal data analysis value sequence is calculated according to the following formula: ; Where a is the second modal data analysis factor of the modal data analysis value sequence, n is the number of modal data analysis values ​​in the modal data analysis value sequence, and b is j is the j-th modal data analysis value in the modal data analysis value sequence, v is the variance of the modal data analysis value sequence, g1 is the number of the first modal data analysis values, g2 is the number of the second modal data analysis values, k1 is the sum of the first modal data analysis values, and k2 is the sum of the second modal data analysis values.

7. The method for establishing an emotion recognition model based on multimodal data according to claim 1, characterized in that: When all modal data analysis factors are numerically sorted, and multiple candidate sub-multimodal data are determined based on the sorting results, including: Select the modal data analysis value sequences corresponding to the first c1 modal data analysis factors as candidate modal data analysis value sequences; The first c2 sub-multimodal data are selected from each candidate modal data analysis value sequence as the candidate sub-multimodal data.

8. The method for establishing an emotion recognition model based on multimodal data according to claim 1, characterized in that: When the candidate sub-multimodal data feature set is divided to obtain a model building set, and a multimodal data emotion recognition model is established based on the model building set, it includes: Based on the training feature set, an initial multimodal data emotion recognition model is constructed using a preset machine learning algorithm; Using the verification feature set to verify the initial multimodal data emotion recognition model, when the verification result meets the preset verification result, the multimodal data emotion recognition model is obtained; When the verification result does not meet the preset verification result, the model parameters are adjusted according to the verification result to obtain the multimodal data emotion recognition model.

9. The method for establishing an emotion recognition model based on multimodal data according to claim 8, characterized in that: The preset machine learning algorithm includes one or more of a decision tree algorithm, a support vector machine algorithm, and a neural network algorithm; The model parameters include the depth of the decision tree, the penalty parameter of the support vector machine, the number of hidden layer nodes of the neural network and the learning rate.

10. A system for establishing an emotion recognition model based on multimodal data, applied to the method for establishing an emotion recognition model based on multimodal data as claimed in any one of claims 1 to 9, characterized in that: include: A data processing module is used to obtain multiple multimodal data, output modal data analysis values ​​of sub-multimodal data corresponding to each multimodal data based on the corresponding sub-multimodal data model, process all modal data analysis values, and obtain multiple modal data analysis value sequences; A factor calculation module, used for randomly assigning the modal data analysis value sequence to a modal data analysis value point diagram, analyzing the modal data analysis value point diagram, and calculating the modal data analysis factor of the modal data analysis value sequence based on the analysis result; A data determination module is used to sort all modal data analysis factors by numerical value, and determine a plurality of candidate sub-multimodal data based on the sorting result, and use the candidate sub-multimodal data as a candidate sub-multimodal data feature set; The model building module is used to divide the candidate sub-multimodal data feature set to obtain a model building set, and build a multimodal data emotion recognition model based on the model building set, wherein the model building set includes a training feature set and a verification feature set.