Intelligent AI seat emotion recognition system and method

By combining differentiated noise separation and explicit emotion priority recognition with implicit emotion judgment and personalized calibration, the problems of noise interference and individual differences in intelligent AI agent emotion recognition are solved. This achieves high recognizability of voice features and accurate hierarchical recognition of emotions, improving the comprehensiveness and accuracy of user emotion recognition.

CN121997201APending Publication Date: 2026-05-08SHENZHEN YUNKEPAI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN YUNKEPAI TECH CO LTD
Filing Date
2026-01-09
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing intelligent AI agent emotion recognition technologies suffer from several drawbacks: voice features are susceptible to noise interference, resulting in insufficient recognition; the lack of explicit and implicit emotion stratification makes it difficult to balance recognition efficiency with the ability to capture implicit emotions; the use of fixed weights for multimodal features cannot adapt to the differences in scenarios with different emotion types and feature intensity levels; and the neglect of individual user emotion expression habits leads to misjudgments and omissions of implicit emotions, making it impossible to achieve comprehensive and accurate recognition of user emotions.

Method used

By differentiating noise separation, constructing a solidified knowledge base to enhance general emotion association features, combining a explicit emotion keyword library with a dynamic weight rule library and statistical analysis of multimodal feature contributions, explicit emotions are identified first and then implicit emotions are determined. User profiles are introduced for personalized calibration to improve the recognizability of voice features and the accuracy of implicit emotion recognition.

Benefits of technology

It improves the recognizability and reliability of voice features, achieves accurate hierarchical recognition of explicit and implicit emotions, avoids misjudgment and omission caused by individual expression differences, and ensures the comprehensiveness and accuracy of emotion recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121997201A_ABST
    Figure CN121997201A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent AI seat emotion recognition system and method, and the method comprises the steps: constructing a dominant emotion keyword library, carrying out the feature extraction of high-recognition voice features, text data and interactive behavior data based on a fixed knowledge library and the dominant emotion keyword library, generating multi-modal features, constructing a sample database, and carrying out the recognition of the emotion of a user. Extracting a dominant emotion labeling sample set from the sample database to perform dominant emotion multi-modal feature contribution degree statistical analysis, dynamic weight rule base construction and dominant emotion recognition operation, and judging the dominant emotion of the user according to a dominant emotion recognition result. Or extracting implicit micro-features, semantic contradiction degree and interactive behavior features from the high-recognition voice features, the text data and the interactive behavior data to perform implicit emotion recognition and personalized calibration, and judging the implicit emotion of the user or judging that the user has no implicit emotion. According to the invention, the efficiency of dominant emotion recognition is ensured, the problem of difficulty in recognizing the pain point of the recessive emotion in the prior art is solved, and the whole scene of user emotion expression is fully covered.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence and sentiment analysis technology, specifically to an intelligent AI agent emotion recognition system and method. Background Technology

[0002] In the wave of digital transformation, the customer service industry is undergoing profound changes. The traditional human agent service model faces numerous challenges, including high labor costs, limited service hours, and the impact of emotional fluctuations on service quality. Meanwhile, with the maturity of artificial intelligence technologies such as natural language processing, speech recognition, and machine learning, intelligent AI agents have emerged, bringing revolutionary changes to customer service. Intelligent AI agents are automated customer service systems based on artificial intelligence technology, utilizing natural language processing (NLP) and machine learning techniques to simulate the interaction between human customer service representatives and customers. (1) Compared to traditional human agents, intelligent AI agents offer significant advantages such as 24 / 7 uninterrupted service, standardized process execution, and millisecond-level response speed. However, existing intelligent AI agent emotion recognition technologies generally suffer from several shortcomings. These include: voice features being susceptible to noise interference leading to insufficient recognition accuracy; a lack of explicit and implicit emotion stratification, making it difficult to balance recognition efficiency with the ability to capture implicit emotions; the use of fixed weights for multimodal features, failing to adapt to different scenarios with varying emotion types and feature intensity levels; and ignoring individual differences in user emotion expression habits, resulting in misjudgments and missed detections of implicit emotions. These deficiencies collectively prevent current technologies from achieving comprehensive and accurate identification of user emotions, thereby impacting the targeting of AI agent service strategies and the improvement of user interaction experience. Summary of the Invention

[0003] To address the aforementioned technical problems, the present invention aims to provide an intelligent AI agent emotion recognition method, comprising the following steps: Step s1: Real-time acquisition of voice data, text data and interaction behavior data during the interaction between the user and the AI ​​agent; differential noise separation of the voice data to obtain clean voice data; construction of a solidified knowledge base; general emotion association feature enhancement of the clean voice data based on the solidified knowledge base to obtain highly recognizable voice features. Step s2: Construct an explicit emotion keyword library. Based on the fixed knowledge base and explicit emotion keyword library, extract features from high-recognition speech features, text data and interactive behavior data to generate multimodal features. Construct a sample database. Extract explicit emotion labeled sample sets from the sample database to perform statistical analysis of the contribution of explicit emotion multimodal features, construct a dynamic weight rule base and perform explicit emotion recognition operations. Determine the user's explicit emotion based on the explicit emotion recognition results or proceed to step s3. Step s3: Extract latent micro-features, semantic contradiction degree and interactive behavior features from high-recognition speech features, text data and interactive behavior data to construct multimodal features, perform latent emotion recognition and personalized calibration on multimodal features, determine the user's latent emotions or determine the user as having no latent emotions.

[0004] Furthermore, the process of differential noise separation for speech data includes: A noise spectrum feature library is constructed, which contains baseline features for several noise types. The collected speech data is segmented into frames, and the short-time Fourier transform spectrum of each frame is obtained. A spectrogram is constructed, and the spectral features in the spectrogram are extracted. The spectral features are matched with the baseline features in the noise spectrum feature library to obtain the noise type of the spectral features. Based on the noise type of the spectral features, differential noise separation is performed on the spectral features. Finally, the spectral features with noise separation are reconstructed into clean speech data.

[0005] Furthermore, the process of building a solidified knowledge base includes: A multi-emotion speech sample library is constructed, which contains speech samples of all target emotion types. All acoustic features of each speech sample in the sample library are extracted. Correlation strength analysis is performed on each acoustic feature and the target emotion type to obtain the correlation coefficient between each acoustic feature and any target emotion type. A strong correlation threshold and an irrelevant threshold are preset. The acoustic feature set with an absolute value of the correlation coefficient with any target emotion type greater than the strong correlation threshold is selected and marked as a general emotion association feature. The acoustic feature set with an absolute value of the correlation coefficient with any target emotion type less than the irrelevant threshold is selected and marked as an irrelevant feature. A solidified knowledge base is constructed based on general emotion-related and irrelevant features.

[0006] Furthermore, the process of enhancing the general emotion-related features of clean speech data based on a fixed knowledge base includes: All acoustic features are extracted from the clean speech data to construct an acoustic feature vector. Each acoustic feature in the acoustic feature vector is matched with general emotion-related features and irrelevant features in the knowledge base. If the acoustic feature is included in the general emotion-related features, a fixed enhancement coefficient is assigned to the acoustic feature; if the acoustic feature is included in the irrelevant features, a fixed attenuation coefficient is assigned to the acoustic feature. Weighted calculations are performed on each acoustic feature in the acoustic feature vector and its corresponding fixed enhancement coefficient or fixed attenuation coefficient to obtain enhanced features. PCA dimensionality reduction is performed on the enhanced features to generate highly recognizable speech features.

[0007] Furthermore, a explicit emotion keyword database is constructed. The process of extracting features from highly recognizable speech features, text data, and interactive behavior data based on the fixed knowledge base and the explicit emotion keyword database includes: Acoustic features with an absolute correlation coefficient greater than a strong correlation threshold for any explicit emotion are extracted from a fixed knowledge base. These acoustic features are then labeled as explicit speech features, and explicit speech features are extracted from the highly recognizable speech features. A keyword library for explicit emotions is constructed, which includes keywords associated with each type of explicit emotion and their corresponding weights. The text data is preprocessed, and statistical analysis is performed on the preprocessed text data based on the explicit emotion keyword library to obtain the frequency of occurrence of keywords associated with each type of explicit emotion in the text data. The score of each type of explicit emotion is obtained based on the frequency of occurrence of keywords associated with each type of explicit emotion and the weight of the keywords. At the same time, sentence structure analysis is performed on the preprocessed text data to obtain the proportion of each type of sentence structure. Text features are constructed based on the scores of each type of explicit emotion and the proportion of each type of sentence structure. The interaction behavior data is quantified to obtain standardized feature values ​​of various types, and interaction behavior features are constructed based on the standardized feature values ​​of various types. Multimodal features are constructed based on explicit speech features, text features, and interactive behavior features.

[0008] Furthermore, the process of constructing a sample database, extracting a set of explicitly labeled emotional samples from the database, performing statistical analysis of the contribution of explicitly labeled emotional multimodal features, constructing a dynamic weight rule base, and identifying explicitly labeled emotions includes: A sample database is constructed, which contains multimodal features of several explicit and implicit emotion scenarios. Several multimodal features of explicit emotion scenarios are extracted from the sample database as explicit emotion annotation sample sets. We conducted statistical analysis on the contribution of multimodal features of dominant emotions in the labeled sample set, obtained the contribution ratio of the three types of features under different intensity levels, and constructed a dynamic weight rule base based on the contribution ratio. A dominant emotion recognition model is constructed. The dominant emotion recognition model is trained using a dominant emotion labeled sample set to obtain the trained dominant emotion recognition model. The intensity level of the three types of features in the current multimodal features is obtained. The dynamic weights of the three types of features in the multimodal features are set according to the contribution ratio corresponding to the intensity level and the dynamic weight rule base. The multimodal features with the completed dynamic weight settings are input into the dominant emotion recognition model. The confidence of each type of dominant emotion is output according to the dominant emotion recognition model. A preset confidence threshold for explicit emotions is set. If there is an explicit emotion with a confidence level greater than the explicit emotion confidence threshold, the explicit emotion with the highest confidence level is selected from the explicit emotions and determined as the user's explicit emotion. If the confidence level of each type of explicit emotion is less than the explicit emotion confidence threshold, then implicit emotion identification is performed.

[0009] Furthermore, the process of conducting statistical analysis on the contribution of multimodal features of explicit emotions to the explicit emotion-labeled sample set and constructing a dynamic weight rule base includes: The explicit emotion labeled sample set is grouped according to explicit emotion type to obtain explicit emotion samples of each type. The three types of features in the multimodal features of each explicit emotion sample are normalized and mapped to a standardized feature value range. Threshold points are selected within the standardized feature value range to divide sub-intervals of different intensity levels. The intensity levels of the three types of features in the multimodal features of each explicit emotion sample are divided to generate sub-samples of each type of feature at different intensity levels. Emotion labels and modal contribution labels are applied to the sub-samples of each type of feature at different intensity levels in each explicit emotion sample. The number of core contribution samples and auxiliary contribution samples of the three types of features in each explicit emotion sample at different intensity levels are obtained. Set the core contribution weight coefficient and the auxiliary contribution weight coefficient, and obtain the contribution ratio of the three features under different intensity levels under each type of explicit emotion condition based on the number of core contribution samples, the number of auxiliary contribution samples, the core contribution weight coefficient, and the auxiliary contribution weight coefficient. Set contribution percentage thresholds and dynamic weight adjustment rules, compare the contribution percentage of various features at different intensity levels with the contribution percentage thresholds, and if the contribution percentage of a certain feature at an intensity level is greater than the contribution percentage threshold, add dynamic weight adjustment rules for that feature and its corresponding intensity level. Construct a dynamic weight rule library based on the various features with added dynamic weight adjustment rules and their corresponding intensity levels.

[0010] Furthermore, step s3 includes the following process: Acoustic features with absolute correlation coefficients greater than a strong correlation threshold for any latent emotion are extracted from a fixed knowledge base. These acoustic features are then labeled as latent micro-features. Latent micro-features are extracted from highly recognizable speech features. The text features of the user's current text data are compared with the text features of historical text data to obtain semantic contradiction. Interactive behavior features of the current interactive behavior data are obtained. Multimodal features are constructed based on latent micro-features, semantic contradiction, and interactive behavior features. To construct an implicit emotion recognition model, multimodal features in several implicit emotion scenarios are extracted from the sample database as an implicit emotion annotation sample set. The implicit emotion recognition model is trained using the implicit emotion annotation sample set to obtain the trained implicit emotion recognition model. Multimodal features are input into the latent emotion recognition model. The model outputs the confidence scores of various latent emotions. A default confidence score threshold for latent emotions is set. The confidence scores of various latent emotions are compared with the latent emotion confidence score threshold. If the confidence scores of all latent emotions are less than the latent emotion confidence score threshold, the user is determined to have no latent emotions. If the confidence scores of latent emotions are greater than the latent emotion confidence score threshold, the confidence scores of latent emotions are personalized and calibrated.

[0011] Furthermore, the process of performing personalized calibration includes: Build a user profile library to store static profile features of users. Extract static profile features of users from the user profile library, and extract dynamic features from the interaction between users and AI agents. Build user profiles based on static profile features and dynamic features. Construct a profile association reference table, which includes the association weights and association coefficients between different profile features and various types of implicit emotions. Obtain calibration coefficients based on the static and dynamic profile features in the user profile and the profile association reference table. The confidence level of implicit emotions is calibrated based on the calibration coefficient. If the confidence level of implicit emotions after calibration is greater than the confidence level threshold of implicit emotions, then the implicit emotions are determined to be the user's implicit emotions.

[0012] An intelligent AI-powered agent emotion recognition system includes a cloud platform, which is connected via communication to a data acquisition module, an explicit emotion recognition module, and a implicit emotion recognition module. The data acquisition module is used to acquire voice data, text data and interaction behavior data in real time during the interaction between the user and the AI ​​agent. It performs differential noise separation on the voice data to obtain clean voice data, builds a solidified knowledge base, and enhances the clean voice data with general emotion association features based on the solidified knowledge base to obtain highly recognizable voice features. The explicit emotion recognition module is used to build an explicit emotion keyword library. Based on the fixed knowledge base and explicit emotion keyword library, it extracts features from high-recognition speech features, text data and interactive behavior data to generate multimodal features, build a sample database, extract explicit emotion labeled sample sets from the sample database to perform explicit emotion multimodal feature contribution statistical analysis, dynamic weight rule base construction and explicit emotion recognition operation, and determine the user's explicit emotion or execute the implicit emotion recognition module based on the explicit emotion recognition results. The latent emotion recognition module is used to extract latent micro-features, semantic contradictions, and interactive behavior features from high-recognition speech features, text data, and interactive behavior data to construct multimodal features. It then performs latent emotion recognition and personalized calibration on the multimodal features to determine the user's latent emotions or classify the user as having no latent emotions.

[0013] Compared with the prior art, the beneficial effects of the present invention are: 1. Enhancing the Recognition and Reliability of Speech Features: This invention innovatively adopts a differentiated noise separation strategy. By constructing a noise spectrum feature library and performing similarity matching, different types of noise can be separated in a targeted manner, avoiding the speech feature distortion problem caused by the "one-size-fits-all" approach of traditional noise suppression methods. At the same time, based on a solidified knowledge base, general emotion-related feature enhancement is performed on clean speech data, strengthening acoustic features strongly correlated with emotions and weakening interference from irrelevant features, effectively improving the recognition of speech features and laying a high-quality data foundation for subsequent emotion recognition.

[0014] 2. Achieving precise hierarchical identification of explicit and implicit emotions: This invention employs a hierarchical architecture that prioritizes explicit emotion identification and supplements it with implicit emotion assessment. By constructing a keyword library and a dynamic weighted rule library for explicit emotions, combined with statistical analysis of multimodal feature contributions, it can quickly and accurately identify users' explicit emotions. For scenarios where explicit emotions are not detected, implicit micro-features, semantic contradictions, and interactive behavior features are further extracted to effectively capture implicit emotions. This hierarchical design ensures the efficiency of explicit emotion identification while addressing the pain point of existing technologies' difficulty in identifying implicit emotions, comprehensively covering all scenarios of user emotional expression.

[0015] 3. Personalized Calibration Further Enhances the Accuracy of Implicit Emotion Recognition. This invention addresses the characteristic of implicit emotions—the inconsistency between outward expression and inner emotion—by introducing user profiles for personalized calibration. Combining static user profile features with real-time dynamic features, a calibration coefficient is obtained through a profile association lookup table to correct the confidence level of implicit emotions. This design fully considers the differences in emotional expression habits among different users, effectively avoiding misjudgments and missed judgments caused by individual expression differences, significantly improving the accuracy of implicit emotion recognition, and making the recognition results more closely match the user's true emotional state. Attached Figure Description

[0016] Figure 1 This is a flowchart illustrating an intelligent AI agent emotion recognition method according to an embodiment of this application.

[0017] Figure 2 This is a flowchart of an intelligent AI agent emotion recognition system according to an embodiment of this application. Detailed Implementation

[0018] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0019] like Figure 1 As shown, an intelligent AI agent emotion recognition method includes the following steps: Step s1: Real-time acquisition of voice data, text data and interaction behavior data during the interaction between the user and the AI ​​agent; differential noise separation of the voice data to obtain clean voice data; construction of a solidified knowledge base; general emotion association feature enhancement of the clean voice data based on the solidified knowledge base to obtain highly recognizable voice features. Step s2: Construct an explicit emotion keyword library. Based on the fixed knowledge base and explicit emotion keyword library, extract features from high-recognition speech features, text data and interactive behavior data to generate multimodal features. Construct a sample database. Extract explicit emotion labeled sample sets from the sample database to perform statistical analysis of the contribution of explicit emotion multimodal features, construct a dynamic weight rule base and perform explicit emotion recognition operations. Determine the user's explicit emotion based on the explicit emotion recognition results or proceed to step s3. Step s3: Extract latent micro-features, semantic contradiction degree and interactive behavior features from high-recognition speech features, text data and interactive behavior data to construct multimodal features, perform latent emotion recognition and personalized calibration on multimodal features, determine the user's latent emotions or determine the user as having no latent emotions.

[0020] It should be further explained that, in the specific implementation process, the process of performing differential noise separation on the speech data to obtain clean speech data includes: Emotional features in user speech (such as fundamental frequency F0 fluctuations and tail decay rate) are low-energy details that are easily masked by environmental noise (such as traffic noise, keyboard noise) and equipment noise (such as electrical noise). For example, if a user's true emotion is "suppressed dissatisfaction," which is manifested by small fundamental frequency F0 fluctuations and fast tail decay, the steady-state noise of the background air conditioner will smooth out these details, causing the model to fail to recognize them. If a user is calm, but a sudden keyboard tapping sound (non-steady-state noise) will produce a volume pulse, which the model will misjudge as a sudden increase in volume characteristic of "anger."

[0021] Typical noise samples from various scenarios (home, outdoor, office, and traditional customer service centers) of AI agents are pre-collected. Spectral features are extracted from each noise sample as baseline features for noise identification, constructing a noise spectrum feature library. This library contains baseline features for several noise types: steady-state noise features: stable spectral distribution (e.g., 50Hz / 100Hz low-frequency peaks of air conditioner noise), small time-domain energy fluctuations (fluctuation coefficient < 0.1); non-steady-state noise features: randomly changing spectral peaks (e.g., 2kHz-5kHz random pulses of keyboard noise), large time-domain energy fluctuations (fluctuation coefficient > 0.5); link noise features: fixed-frequency interference (e.g., 50Hz / 60Hz current noise), intermittent features of lost speech frames. The collected speech data is then segmented into frames (20ms frame length, 10ms frame shift), with each segment divided into 128 frames. The short-time Fourier transform spectrum of each frame is obtained, and the short-time Fourier transform spectrum of each frame is... The frequency range in the spectrum (e.g., 0-8kHz, corresponding to the main frequency range of the speech signal) is divided into 128 equally spaced frequency points. Each point represents the signal energy intensity at that frequency. A 128×128 (128 speech frames × 128 frequency points) spectrum map is constructed. Spectral features in the spectrum map are extracted (the spectrum map is input into a pre-trained MobileNetV3 model, which extracts spectral features through 3 layers of deep separable convolution). The spectral features are matched with the baseline features in the noise spectrum feature library to obtain the similarity between the spectral features and each noise type. A preset similarity threshold (0.8) is used. When the similarity between the spectral features and the noise type is greater than the similarity threshold, it is determined to be the corresponding noise type. The noise type (steady-state / non-steady-state / link) of the spectral features is obtained. Differentiated noise separation is performed on the spectral features according to the noise type of the spectral features. Then, the spectral features with noise separation are reconstructed into clean speech data.

[0022] It should be further explained that, in the specific implementation process, the process of differentially separating spectral features based on the noise type of the spectral characteristics includes: Steady-state noise separation: Calculating the average power spectrum of the noise frame (Selecting the silent segment of the speech signal as the noise reference frame), calculate the power spectrum of the noisy speech frame. Obtain the power spectrum of pure speech The process is as follows: ; in, As an over-subtraction factor, it takes the value 1.3. As a residual factor, a value of 0.2 is set (to preserve low-power emotional detail features), and then the power spectrum of the pure speech is obtained by inverse STFT. Converted into a time-domain speech signal.

[0023] Non-steady-state noise separation: Non-steady-state noise has no fixed spectral characteristics, making it difficult to effectively separate using traditional methods. This method employs a time-series prediction model to achieve noise cancellation. Step 1: Convert the noisy speech signal into a temporal feature sequence and input it into the bidirectional LSTM model; Step 2: The LSTM model predicts the noise feature vector of the current frame based on historical temporal features. , Step 3: Noise Cancellation Calculation: ,in Features of noisy speech Features of pure speech The predicted weights (obtained from model training, ranging from 0.8 to 1.0); Step 4: Reconstruct the canceled feature sequence to obtain a clean speech signal in the time domain.

[0024] Link noise separation: For speech stuttering / frame loss caused by link noise, a frame interpolation repair algorithm is executed: Step 1: Identify lost / distorted speech frames through endpoint detection (frames with energy below the threshold are marked as lost frames); Step 2: Use linear interpolation to complete the acoustic features of the lost frame based on the effective frame features before and after the lost frame; Step 3: Smooth the repaired speech frame to avoid abrupt changes in features between frames and restore the continuity of emotional micro-features such as speech rate and pause intervals.

[0025] By accurately separating user voice signals from noise in complex noisy environments, restoring masked emotional features, and strengthening acoustic features strongly correlated with emotions, this method addresses the pain point of traditional emotion recognition methods experiencing a sharp drop in accuracy under noise interference.

[0026] It should be further explained that, in the specific implementation process, the process of building a solidified knowledge base includes: A multi-emotion speech sample library is constructed, which contains speech samples of all target emotion types (such as overt emotions like anger and depression, and latent emotions like suppressed dissatisfaction and pseudo-satisfaction). The samples cover users of different ages, genders, and dialects. All acoustic features (fundamental frequency F0 fluctuation amplitude, tail decay rate, volume change rate, etc.) of each speech sample in the sample library are extracted. The correlation strength analysis between each acoustic feature and the target emotion type is performed to obtain the correlation coefficient between each acoustic feature and any target emotion type. A strong correlation threshold (0.7) and an irrelevance threshold (0.3) are preset. The acoustic feature set with an absolute value of the correlation coefficient with any target emotion type greater than the strong correlation threshold is selected and marked as a general emotion association feature. The acoustic feature set with an absolute value of the correlation coefficient with any target emotion type less than the irrelevance threshold is selected and marked as an irrelevant feature.

[0027] A solidified knowledge base is constructed based on general emotion-related and irrelevant features.

[0028] The process of performing correlation strength analysis between each acoustic feature and the target emotion type is as follows: ; in, Let be a certain acoustic feature value of the i-th speech sample. Let be the average of a certain acoustic feature from n speech samples. Quantize the emotion label of the i-th voice sample (e.g., anger=1, neutral=0, depressed=-1). Quantize the average of the emotion labels for n voice samples. The correlation coefficient has a value range of [-1, 1]. The larger the value, the stronger the correlation between the characteristic and the emotion.

[0029] It should be further explained that, in the specific implementation process, the process of enhancing the general emotion association features of the clean speech data based on the fixed knowledge base to obtain highly recognizable speech features includes: Extract all acoustic features from clean speech data and construct an acoustic feature vector. n1 represents the number of acoustic features, and the acoustic feature vector is... Each acoustic feature in (j=1,2,...,n1) is matched with general emotion-related features and irrelevant features in the fixed knowledge base. If the acoustic feature is included in the general emotion-related features, a fixed enhancement coefficient (1.3) is assigned to the acoustic feature. If the acoustic feature is included in the irrelevant features, a fixed attenuation coefficient (0.7) is assigned to the acoustic feature. A weighted calculation is performed on each acoustic feature in the acoustic feature vector and its corresponding fixed enhancement coefficient or fixed attenuation coefficient. , This represents the fixed enhancement coefficient or fixed attenuation coefficient corresponding to the j-th acoustic feature, used to obtain the enhancement feature. PCA dimensionality reduction is performed on the enhanced features, retaining principal components with a cumulative contribution rate of ≥95%, and generating highly recognizable speech features.

[0030] The logical chain of the entire process can be summarized as: noisy speech Noise reduction compensation Enhanced general emotional characteristics Highly recognizable features Input recognition model By outputting specific emotions and improving the signal-to-noise ratio of all emotional features, the subsequent model can more easily "see" these features—regardless of whether the end user's emotion is anger or suppressed dissatisfaction, the relevant features have been amplified.

[0031] However, the feature enhancement stage of existing technologies first identifies emotions and then enhances the corresponding features in a targeted manner, which leads to a logical dead loop: it requires an emotional result. Talent targeted enhancement However, enhancement is a prerequisite for recognition. Without enhancement, the identification will be inaccurate; The advantages of this embodiment are: it breaks the cycle, eliminates the need to know the emotion in advance, directly improves the recognizability of all emotional features, and provides high-quality input for subsequent recognition; it is also compatible with all emotion types, avoiding the omission of certain emotions due to targeted enhancement (e.g., enhancing only the anger feature will cause the feature of suppressed dissatisfaction to be ignored); at the same time, it eliminates the intermediate step of "identifying the emotion first", reduces memory usage and computational complexity through PCA dimensionality reduction, and meets the latency requirements of real-time interaction of AI agents.

[0032] It should be further explained that, in the specific implementation process, the process of constructing an explicit emotion keyword library, and extracting features from high-discrimination speech features, text data, and interactive behavior data based on the fixed knowledge base and the explicit emotion keyword library to generate multimodal features includes: Acoustic features with an absolute correlation coefficient greater than a strong correlation threshold for any explicit emotion are extracted from a fixed knowledge base. These acoustic features are then labeled as explicit speech features, and explicit speech features are extracted from the highly recognizable speech features. A keyword library for explicit emotions is constructed, comprising keywords associated with each type of explicit emotion and their corresponding weights. Text data is preprocessed (format cleaning, word segmentation, and part-of-speech tagging). Statistical analysis is performed on the preprocessed text data based on the explicit emotion keyword library to obtain the frequency of occurrence of keywords associated with each type of explicit emotion. Scores for each type of explicit emotion are obtained based on the frequency of occurrence and the corresponding weights. Explicit emotions are often accompanied by specific sentence structures (such as exclamatory sentences and rhetorical questions), requiring the extraction of sentence structure features to aid in emotion determination: Sentence structure analysis is performed on the preprocessed text data to obtain the proportion of each sentence type, including: exclamatory sentence proportion: number of exclamatory sentences in the text / total number of sentences; rhetorical question proportion: number of rhetorical questions in the text / total number of sentences; short sentence proportion: number of sentences with a length of ≤5 characters / total number of sentences. Text features are constructed based on the scores of each type of explicit emotion and the proportion of each sentence type. The interactive behavior data is quantified to obtain standardized feature values ​​for each type, including: response duration: based on the median of normal response duration in the same scenario, it is standardized to [0,1], where 0 indicates extremely short (≤50% of the baseline) and 1 indicates normal; number of repeated questions: quantified as 0 (no repetition), 0.5 (1-2 repetitions), and 1 (≥3 repetitions); termination command trigger status: 0 (not triggered) and 1 (triggered); interactive behavior features are constructed based on the standardized feature values ​​for each type. Multimodal features are constructed based on explicit speech features, text features, and interactive behavior features.

[0033] The formula for multimodal features is as follows: ; in, For multimodal features, These are explicit speech features. For text features, As interactive behavior characteristics, , , These are the dynamic weights corresponding to high-recognition speech features, text features, and interaction behavior features, respectively. + + =1.

[0034] It should be further noted that, in the specific implementation process, the explicit emotion keyword database is shown in Table 1 below: Table 1

[0035] Assign differentiated weights to keywords (Core keyword weight 1.0, extended keyword weight 0.6), calculate the score corresponding to the explicit sentiment: ,in, This represents the frequency of the k-th keyword associated with overt emotions. is the weight of the k-th keyword, and n2 is the total number of k-th keywords associated with explicit emotions.

[0036] It should be further explained that, in the specific implementation process, the process of constructing a sample database, extracting a set of explicit emotion-labeled samples from the sample database for statistical analysis of the contribution of explicit emotion multimodal features, constructing a dynamic weight rule base, and identifying explicit emotions, and determining the user's explicit emotion or executing step s3 based on the explicit emotion identification results includes: Construct a sample database (collecting historical dialogue data from intelligent customer service, covering explicit emotional scenarios such as anger, happiness, and satisfaction, as well as implicit emotional scenarios such as suppressed dissatisfaction, pseudo-satisfaction, and implicit anxiety; each data point is manually labeled with the corresponding explicit emotion tag and multimodal features). The sample database contains multimodal features under several explicit and implicit emotional scenarios. Extract several multimodal features under explicit emotional scenarios from the sample database as an explicit emotion labeling sample set. The explicit emotion labeling sample set covers all target explicit emotion types: anger, happiness, depression, irritability, and satisfaction; the number of samples for each emotion type is no less than 1000 (to avoid statistical bias caused by small samples); each sample must simultaneously contain voice features, text features, and interaction behavior features, with no missing modal data; We conducted statistical analysis on the contribution of multimodal features of dominant emotions in the labeled sample set, obtained the contribution ratio of the three types of features under different intensity levels, and constructed a dynamic weight rule base based on the contribution ratio. A lightweight hybrid neural network model is used to construct an explicit emotion recognition model, which consists of convolutional layers (CNN), bidirectional long short-term memory layers (Bi-LSTM), and fully connected layers (FC). The convolutional layers employ three different sized convolutional kernels (3×3, 5×5, and 7×7) to extract local correlation features from the fused features (such as the correlation between emotion keywords and speech features). The Bi-LSTM layer captures the temporal dependencies of feature sequences (such as the temporal features of speech pitch changes and the contextual correlation of text semantics). The fully connected layer outputs probability scores for various explicit emotions through a softmax activation function. The loss function is cross-entropy, with the optimization objective being to minimize the error between the predicted probability and the true label. The explicit emotion recognition model is trained using a set of explicitly labeled emotion samples to obtain the trained model. The intensity levels of the three types of features in the current multimodal features are determined. Based on the contribution percentage corresponding to each intensity level and the dynamic weight rule base, the dynamic weights of the three types of features (explicit speech features, text features, and interactive behavior features) in the multimodal features are set. Taking the "voice volume extreme value" feature of anger as an example, the contribution percentage of the speech modality under different intensity levels is statistically shown in Table 2 below. Table 2

[0037] When the volume extreme value is ≥0.6 under the emotion of anger, the contribution of speech reaches more than 70%. Therefore, a dynamic weighting rule is applied: if the volume extreme value in the speech features is >0.6, the speech weight is increased. Up to 0.6, and maintain + + =1, input the multimodal features with dynamic weight settings into the explicit emotion recognition model, and output the confidence score of each type of explicit emotion based on the explicit emotion recognition model; A default confidence threshold of 0.7 is set. If there is a corresponding explicit emotion with a confidence level greater than the explicit emotion confidence threshold, the explicit emotion with the highest confidence level is selected from the explicit emotions and determined as the user's explicit emotion. If the confidence level of each type of explicit emotion is less than the explicit emotion confidence threshold, then implicit emotion identification is performed.

[0038] It should be further explained that, in the specific implementation process, the process of conducting statistical analysis on the contribution of dominant emotion multimodal features to the dominant emotion-labeled sample set and constructing a dynamic weight rule base includes: The explicit emotion labeled sample set is grouped according to explicit emotion type to obtain various explicit emotion samples (anger, happiness, depression, irritability). The three types of features in the multimodal features of each explicit emotion sample are normalized and mapped to a standardized feature value range [0,1]. Within the standardized feature value range, threshold points are selected to divide sub-intervals of different intensity levels (extremely low (0-0.2), low (0.2-0.4), medium (0.4-0.6), high (0.6-0.8), extremely high (0.8-1.0)). The intensity levels of the three types of features in the multimodal features of each explicit emotion sample are then divided, generating sub-samples for each feature at different intensity levels (sub-samples of a certain feature at different intensity levels require that the feature be within the corresponding intensity level; the other two types of features are not considered). The requirements are as follows: For each type of overt emotion sample, emotion labels and modal contribution labels are performed on subsamples of each feature at different intensity levels. The emotion labeling rules are as follows: Each subsample in the overt emotion labeling sample set is independently labeled with an overt emotion by two or more human labelers. Samples with a labeling consistency rate of ≥90% are then labeled with modal contribution. The modal contribution labeling rules are as follows: If removing a modal feature within a specified intensity level reduces the accuracy of human emotion judgment by ≥30%, then the modality is marked as a core contributing modality. If the accuracy decreases by 10%-30% after removal, then the modality is marked as an auxiliary contributing modality. If the accuracy decreases by <10% after removal, then the modality is marked as a non-contributing modality. The number of core contributing samples and auxiliary contributing samples of the three features at different intensity levels in each type of overt emotion sample are obtained. Set the core contribution weight coefficient and the auxiliary contribution weight coefficient, and obtain the contribution ratio of the three features under different intensity levels under each type of explicit emotion condition based on the number of core contribution samples, the number of auxiliary contribution samples, the core contribution weight coefficient, and the auxiliary contribution weight coefficient. The formula for calculating the contribution percentage is: ; in, The contribution percentage of the i-th feature (i=1, 2, 3). The core contribution weighting coefficient is 0.7. The auxiliary contribution weighting coefficient is 0.3. The number of core contributing samples for the i-th feature. The number of auxiliary contribution samples for the i-th feature. This represents the total number of subsamples within the explicit emotion sample.

[0039] Set contribution percentage thresholds and dynamic weight adjustment rules. Compare the contribution percentage of various features at different intensity levels with the contribution percentage thresholds. If the contribution percentage of a certain feature at an intensity level is greater than the contribution percentage threshold, add a dynamic weight adjustment rule (increase the weight to 0.6) for that feature and its corresponding intensity level. Construct a dynamic weight rule library based on the various features with added dynamic weight adjustment rules and their corresponding intensity levels.

[0040] It should be further explained that, in the specific implementation process, the process of extracting latent micro-features, semantic contradictions, and interactive behavior features from high-recognition speech features, text data, and interactive behavior data to construct multimodal features, and then performing latent emotion recognition and personalized calibration on these multimodal features to determine the user's latent emotions or classify the user as having no latent emotions includes: Latent emotions include three categories: suppressed dissatisfaction, pseudo-satisfaction, and hidden anxiety. Suppressed dissatisfaction manifests outwardly as a flat tone, brief replies (such as "okay" or "I understand"), and the absence of obvious negative keywords. The underlying emotion is dissatisfaction with the service but an unwillingness to confront it directly. For example, if the after-sales solution does not meet expectations, the user may not object but will subsequently refuse to cooperate. Pseudo-satisfaction manifests outwardly as neutral to positive words such as "okay" or "okay," but with a flat tone and a fast speaking speed. The underlying emotion is superficial compromise without actually addressing the core demand. For example, if the refund amount does not meet expectations, the user may say "okay" but sigh before hanging up. Hidden anxiety manifests outwardly as repeatedly asking the same question and slightly illogical statements, but without anxiety keywords. The underlying emotion is worry that the demand will not be met and inner tension. Latent emotions differ from overt emotions such as anger and joy, which are clearly expressed outwardly. Their core characteristic is that outward language / vocal behavior is inconsistent with internal emotional tendencies. Traditional emotion recognition methods, which rely on overt features (such as negative keywords or sudden changes in tone), are difficult to capture effectively.

[0041] Acoustic features with absolute correlation coefficients greater than a strong correlation threshold for any latent emotion are extracted from a fixed knowledge base. These acoustic features are labeled as latent micro-features. Latent micro-features are extracted from high-recognition speech features. The text features of the user's current text data are compared with the text features of historical text data (the historical text data of the user's request in the complete dialogue of the user's request) to obtain semantic contradiction (the scores of each type of explicit emotion and the proportion of each type of sentence structure in the text features of the current text data are compared with the scores of each type of explicit emotion and the proportion of each type of sentence structure in the text features of the historical text data to obtain cosine similarity, where semantic contradiction = 1 - cosine similarity). Interactive behavior features of the current interactive behavior data are obtained. Multimodal features are constructed based on latent micro-features, semantic contradiction, and interactive behavior features. A three-layer fully connected neural network (FCN) is used to construct an implicit emotion recognition model. The input layer has a multimodal feature dimension (e.g., 128 dimensions), the hidden layer has a 64-dimensional dimension, and the output layer has a 3-dimensional dimension (corresponding to the probabilities of three types of implicit emotions). The loss function adopts a combination of cross-entropy loss and focal loss to address the problem of few implicit emotion samples and their susceptibility to being overwhelmed by explicit emotion samples. The formula is as follows: ;in, =0.75 (balanced positive and negative samples). =2 (Focusing on difficult-to-distinguish samples) For real labels, To predict the probability, multimodal features in several implicit emotion scenarios are extracted from the sample database as an implicit emotion annotation sample set. The implicit emotion recognition model is trained using the implicit emotion annotation sample set to obtain the trained implicit emotion recognition model. Multimodal features are input into the latent emotion recognition model. The model outputs the confidence scores of various latent emotions. A default confidence score threshold (0.65) is set. The confidence scores of various latent emotions are compared with the latent emotion confidence score threshold. If the confidence scores of various latent emotions are all less than the latent emotion confidence score threshold, it is determined that the user has no latent emotions. If the confidence scores of latent emotions are greater than the latent emotion confidence score threshold, the confidence scores of latent emotions are personalized and calibrated.

[0042] It should be further explained that, in the specific implementation process, the personalized calibration process includes: A user profile database is built to store static profile features of users (personality tags, age range, and historical interaction emotion records (emotional types in the last 3 conversations)). The fields of the user profile database need to be obtained through user authorization and behavior analysis. That is, static profile features are generated through the initial interaction questionnaire and historical behavior clustering. The static profile features of users are extracted from the user profile database, and dynamic features (type of request, number of conversation rounds, and degree of request achievement) are extracted from the interaction between users and AI agents. The user profile is built based on the static profile features and dynamic features of users. Statistical analysis was performed on the sample data using statistical methods to construct a profile correlation table. This table includes the correlation weights and coefficients between different profile features (including static and dynamic features) and various types of latent emotions. The correlation weight represents the strength of the correlation between the profile feature and the non-overt emotion, with a value range of [0.1, 0.3]. A larger value indicates a more significant impact of the field on emotion assessment. The correlation coefficient represents the direction of the influence of the profile feature on the non-overt emotion, with values ​​of 1 (positive correlation), -1 (negative correlation), or 0 (no correlation). Calibration coefficients were obtained based on the static and dynamic profile features in the user profile and the profile correlation table. For example, if the latent emotion is suppressed dissatisfaction, the profile features are: Personality label = introverted (strongly correlated with repressed dissatisfaction): (correlation weight wf = 0.3, correlation coefficient kf = 1); (if personality label = extroverted, then the correlation coefficient kf = -1). Current complaint type = after-sales complaint (moderate correlation with suppressed dissatisfaction): (correlation weight wf=0.2, correlation coefficient kf=1); There is a historical record of suppressed discontent (strongly correlated with suppressed discontent): (correlation weight wf=0.3, correlation coefficient kf=1).

[0043] Calibration coefficient ;in, For calibration coefficients, The association weights of the portrait feature f are: Let f be the correlation coefficient of the portrait feature f. This represents the number of portrait features.

[0044] Then calibration coefficient =1+0.3+0.2+0.3=1.9.

[0045] The confidence level of implicit emotions is calibrated based on the calibration coefficient (confidence level × calibration coefficient). If the confidence level of implicit emotions after calibration is greater than the confidence level threshold of implicit emotions, then the implicit emotions are determined to be the user's implicit emotions.

[0046] like Figure 2As shown, an intelligent AI agent emotion recognition system includes a cloud platform, and the cloud platform is connected to a data acquisition module, an explicit emotion recognition module, and an implicit emotion recognition module. The data acquisition module is used to acquire voice data, text data and interaction behavior data in real time during the interaction between the user and the AI ​​agent. It performs differential noise separation on the voice data to obtain clean voice data, builds a solidified knowledge base, and enhances the clean voice data with general emotion association features based on the solidified knowledge base to obtain highly recognizable voice features. The explicit emotion recognition module is used to build an explicit emotion keyword library. Based on the fixed knowledge base and explicit emotion keyword library, it extracts features from high-recognition speech features, text data and interactive behavior data to generate multimodal features, build a sample database, extract explicit emotion labeled sample sets from the sample database to perform explicit emotion multimodal feature contribution statistical analysis, dynamic weight rule base construction and explicit emotion recognition operation, and determine the user's explicit emotion or execute the implicit emotion recognition module based on the explicit emotion recognition results.

[0047] The latent emotion recognition module is used to extract latent micro-features, semantic contradictions, and interactive behavior features from high-recognition speech features, text data, and interactive behavior data to construct multimodal features. It then performs latent emotion recognition and personalized calibration on the multimodal features to determine the user's latent emotions or classify the user as having no latent emotions.

[0048] The above embodiments are only used to illustrate the technical methods of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical methods of the present invention without departing from the spirit and scope of the technical methods of the present invention.

Claims

1. A method for intelligent AI agent emotion recognition, characterized in that, Includes the following steps: Step s1: Real-time acquisition of voice data, text data and interaction behavior data during the interaction between the user and the AI ​​agent; differential noise separation of the voice data to obtain clean voice data; construction of a solidified knowledge base; general emotion association feature enhancement of the clean voice data based on the solidified knowledge base to obtain highly recognizable voice features. Step s2: Construct an explicit emotion keyword library. Based on the fixed knowledge base and explicit emotion keyword library, extract features from high-recognition speech features, text data and interactive behavior data to generate multimodal features. Construct a sample database. Extract explicit emotion labeled sample sets from the sample database to perform statistical analysis of the contribution of explicit emotion multimodal features, construct a dynamic weight rule base and perform explicit emotion recognition operations. Determine the user's explicit emotion based on the explicit emotion recognition results or proceed to step s3. Step s3: Extract latent micro-features, semantic contradiction degree and interactive behavior features from high-recognition speech features, text data and interactive behavior data to construct multimodal features, perform latent emotion recognition and personalized calibration on multimodal features, determine the user's latent emotions or determine the user as having no latent emotions.

2. The intelligent AI agent emotion recognition method according to claim 1, characterized in that, The process of performing differential noise separation on speech data includes: A noise spectrum feature library is constructed, which contains baseline features for several noise types. The collected speech data is segmented into frames, and the short-time Fourier transform spectrum of each frame is obtained. A spectrogram is constructed, and the spectral features in the spectrogram are extracted. The spectral features are matched with the baseline features in the noise spectrum feature library to obtain the noise type of the spectral features. Based on the noise type of the spectral features, differential noise separation is performed on the spectral features. Finally, the spectral features with noise separation are reconstructed into clean speech data.

3. The intelligent AI agent emotion recognition method according to claim 2, characterized in that, The process of building a solidified knowledge base includes: A multi-emotion speech sample library is constructed, which contains speech samples of all target emotion types. All acoustic features of each speech sample in the sample library are extracted. Correlation strength analysis is performed on each acoustic feature and the target emotion type to obtain the correlation coefficient between each acoustic feature and any target emotion type. A strong correlation threshold and an irrelevant threshold are preset. The acoustic feature set with an absolute value of the correlation coefficient with any target emotion type greater than the strong correlation threshold is selected and marked as a general emotion association feature. The acoustic feature set with an absolute value of the correlation coefficient with any target emotion type less than the irrelevant threshold is selected and marked as an irrelevant feature. A solidified knowledge base is constructed based on general emotion-related and irrelevant features.

4. The intelligent AI agent emotion recognition method according to claim 3, characterized in that, The process of enhancing general emotion-related features in clean speech data based on a fixed knowledge base includes: All acoustic features are extracted from the clean speech data to construct an acoustic feature vector. Each acoustic feature in the acoustic feature vector is matched with general emotion-related features and irrelevant features in the knowledge base. If the acoustic feature is included in the general emotion-related features, a fixed enhancement coefficient is assigned to the acoustic feature; if the acoustic feature is included in the irrelevant features, a fixed attenuation coefficient is assigned to the acoustic feature. Weighted calculations are performed on each acoustic feature in the acoustic feature vector and its corresponding fixed enhancement coefficient or fixed attenuation coefficient to obtain enhanced features. PCA dimensionality reduction is performed on the enhanced features to generate highly recognizable speech features.

5. The intelligent AI agent emotion recognition method according to claim 4, characterized in that, The process of constructing an explicit emotion keyword database and extracting features from highly recognizable speech features, text data, and interactive behavior data based on a fixed knowledge base and the explicit emotion keyword database includes: Acoustic features with an absolute correlation coefficient greater than a strong correlation threshold for any explicit emotion are extracted from a fixed knowledge base. These acoustic features are then labeled as explicit speech features, and explicit speech features are extracted from the highly recognizable speech features. A keyword library for explicit emotions is constructed, which includes keywords associated with each type of explicit emotion and their corresponding weights. The text data is preprocessed, and statistical analysis is performed on the preprocessed text data based on the explicit emotion keyword library to obtain the frequency of occurrence of keywords associated with each type of explicit emotion in the text data. The score of each type of explicit emotion is obtained based on the frequency of occurrence of keywords associated with each type of explicit emotion and the weight of the keywords. At the same time, sentence structure analysis is performed on the preprocessed text data to obtain the proportion of each type of sentence structure. Text features are constructed based on the scores of each type of explicit emotion and the proportion of each type of sentence structure. The interaction behavior data is quantified to obtain standardized feature values ​​of various types, and interaction behavior features are constructed based on the standardized feature values ​​of various types. Multimodal features are constructed based on explicit speech features, text features, and interactive behavior features.

6. The intelligent AI agent emotion recognition method according to claim 5, characterized in that, The process of constructing a sample database, extracting a set of explicitly labeled emotional samples from the database, performing statistical analysis of the contribution of explicitly labeled emotional multimodal features, constructing a dynamic weight rule base, and identifying explicitly labeled emotions includes: A sample database is constructed, which contains multimodal features of several explicit and implicit emotion scenarios. Several multimodal features of explicit emotion scenarios are extracted from the sample database as explicit emotion annotation sample sets. We conducted statistical analysis on the contribution of multimodal features of dominant emotions in the labeled sample set, obtained the contribution ratio of the three types of features under different intensity levels, and constructed a dynamic weight rule base based on the contribution ratio. A dominant emotion recognition model is constructed. The dominant emotion recognition model is trained using a dominant emotion labeled sample set to obtain the trained dominant emotion recognition model. The intensity level of the three types of features in the current multimodal features is obtained. The dynamic weights of the three types of features in the multimodal features are set according to the contribution ratio corresponding to the intensity level and the dynamic weight rule base. The multimodal features with the completed dynamic weight settings are input into the dominant emotion recognition model. The confidence of each type of dominant emotion is output according to the dominant emotion recognition model. A preset confidence threshold for explicit emotions is set. If there is an explicit emotion with a confidence level greater than the explicit emotion confidence threshold, the explicit emotion with the highest confidence level is selected from the explicit emotions and determined as the user's explicit emotion. If the confidence level of each type of explicit emotion is less than the explicit emotion confidence threshold, then implicit emotion identification is performed.

7. The intelligent AI agent emotion recognition method according to claim 6, characterized in that, The process of performing statistical analysis on the contribution of dominant emotion multimodal features to the dominant emotion-labeled sample set and constructing a dynamic weight rule base includes: The explicit emotion labeled sample set is grouped according to explicit emotion type to obtain explicit emotion samples of each type. The three types of features in the multimodal features of each explicit emotion sample are normalized and mapped to a standardized feature value range. Threshold points are selected within the standardized feature value range to divide sub-intervals of different intensity levels. The intensity levels of the three types of features in the multimodal features of each explicit emotion sample are divided to generate sub-samples of each type of feature at different intensity levels. Emotion labels and modal contribution labels are applied to the sub-samples of each type of feature at different intensity levels in each explicit emotion sample. The number of core contribution samples and auxiliary contribution samples of the three types of features in each explicit emotion sample at different intensity levels are obtained. Set the core contribution weight coefficient and the auxiliary contribution weight coefficient, and obtain the contribution ratio of the three features under different intensity levels under each type of explicit emotion condition based on the number of core contribution samples, the number of auxiliary contribution samples, the core contribution weight coefficient, and the auxiliary contribution weight coefficient. Set contribution percentage thresholds and dynamic weight adjustment rules, compare the contribution percentage of various features at different intensity levels with the contribution percentage thresholds, and if the contribution percentage of a certain feature at an intensity level is greater than the contribution percentage threshold, add dynamic weight adjustment rules for that feature and its corresponding intensity level. Construct a dynamic weight rule library based on the various features with added dynamic weight adjustment rules and their corresponding intensity levels.

8. The intelligent AI agent emotion recognition method according to claim 7, characterized in that, Step s3 includes the following process: Acoustic features with absolute correlation coefficients greater than a strong correlation threshold for any latent emotion are extracted from a fixed knowledge base. These acoustic features are then labeled as latent micro-features. Latent micro-features are extracted from highly recognizable speech features. The text features of the user's current text data are compared with the text features of historical text data to obtain semantic contradiction. Interactive behavior features of the current interactive behavior data are obtained. Multimodal features are constructed based on latent micro-features, semantic contradiction, and interactive behavior features. To construct an implicit emotion recognition model, multimodal features in several implicit emotion scenarios are extracted from the sample database as an implicit emotion annotation sample set. The implicit emotion recognition model is trained using the implicit emotion annotation sample set to obtain the trained implicit emotion recognition model. Multimodal features are input into the latent emotion recognition model. The model outputs the confidence scores of various latent emotions. A default confidence score threshold for latent emotions is set. The confidence scores of various latent emotions are compared with the latent emotion confidence score threshold. If the confidence scores of all latent emotions are less than the latent emotion confidence score threshold, the user is determined to have no latent emotions. If the confidence scores of latent emotions are greater than the latent emotion confidence score threshold, the confidence scores of latent emotions are personalized and calibrated.

9. The intelligent AI agent emotion recognition method according to claim 8, characterized in that, The process of performing personalized calibration includes: Build a user profile library to store static profile features of users. Extract static profile features of users from the user profile library, and extract dynamic features from the interaction between users and AI agents. Build user profiles based on static profile features and dynamic features. Construct a profile association reference table, which includes the association weights and association coefficients between different profile features and various types of implicit emotions. Obtain calibration coefficients based on the static and dynamic profile features in the user profile and the profile association reference table. The confidence level of implicit emotions is calibrated based on the calibration coefficient. If the confidence level of implicit emotions after calibration is greater than the confidence level threshold of implicit emotions, then the implicit emotions are determined to be the user's implicit emotions.

10. An intelligent AI agent emotion recognition system, specifically applied to the intelligent AI agent emotion recognition method according to any one of claims 1 to 9, characterized in that, This includes the cloud, and the cloud communication connection includes a data acquisition module, an explicit emotion recognition module, and an implicit emotion recognition module; The data acquisition module is used to acquire voice data, text data and interaction behavior data in real time during the interaction between the user and the AI ​​agent. It performs differential noise separation on the voice data to obtain clean voice data, builds a solidified knowledge base, and enhances the clean voice data with general emotion association features based on the solidified knowledge base to obtain highly recognizable voice features. The explicit emotion recognition module is used to build an explicit emotion keyword library. Based on the fixed knowledge base and explicit emotion keyword library, it extracts features from high-recognition speech features, text data and interactive behavior data to generate multimodal features, build a sample database, extract explicit emotion labeled sample sets from the sample database to perform explicit emotion multimodal feature contribution statistical analysis, dynamic weight rule base construction and explicit emotion recognition operation, and determine the user's explicit emotion or execute the implicit emotion recognition module based on the explicit emotion recognition results. The latent emotion recognition module is used to extract latent micro-features, semantic contradictions, and interactive behavior features from high-recognition speech features, text data, and interactive behavior data to construct multimodal features. It then performs latent emotion recognition and personalized calibration on the multimodal features to determine the user's latent emotions or classify the user as having no latent emotions.