Customer emotion detection method in telephone customer service, medium and system
By using a combination of multi-head attention mechanism and LSTM model in telephone customer service, the problem that the existing technology is difficult to capture the dynamic changes in customer emotions is solved, and a more accurate and sensitive identification of customer emotions is achieved.
Patent Information
- Application Number
- CN202411900993.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art is difficult to accurately capture the dynamic changing characteristics of customer emotions over a long period of time in telephone customer service, especially in complex voice environments.
A method combining the multi-head attention mechanism (MHA) and LSTM model is adopted to extract and process features through residual network (ResNet). The multi-head attention mechanism divides the input features into multiple subspaces and performs self-attention calculations. The LSTM model decouples short-term mood fluctuations and long-term emotional trends.
It realizes more sensitive and accurate recognition of customer emotions, can effectively capture local and global emotional information in voice signals, and improves the accuracy and robustness of emotional recognition.
Smart Images

Figure CN119943096A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of customer emotions, and in particular, relates to a method, medium and system for monitoring customer emotions in telephone customer service. Background Art
[0002] In the telephone customer service of the power industry, customer mood fluctuations directly affect service quality and customer satisfaction. Effective customer emotion detection can not only help customer service personnel respond to customer needs in a timely manner, but also improve customer experience. As a way to detect customer emotions, speech emotion recognition has been widely used in customer service systems, but it still faces challenges such as subtle emotion changes, background noise interference, and complex emotion characteristics.
[0003] With the digital transformation of the customer service industry, speech emotion recognition technology is increasingly used in power company telephone customer service. Accurately identifying customer emotions can not only improve customer experience, but also help customer service staff to make appropriate responses in a timely manner. However, emotion recognition faces many challenges, especially in complex voice environments. How to accurately capture customer emotional changes is crucial.
[0004] Existing emotion recognition technologies mainly rely on short-term speech feature analysis, and usually classify speech signals through static feature extraction methods (such as Mel-frequency cepstral coefficients MFCC or linear prediction cepstral coefficients LPC). Such methods perform relatively well when processing single emotions or short-term speech signals, but their limitation is that they ignore the time dependence of emotion changes and it is difficult to effectively capture the dynamic changes of customer emotions over a long period of time. For example, in a telephone customer service scenario, the customer's emotions may continue to change as the conversation progresses (such as from anxiety to calmness or from neutral to anger), and it is difficult for existing technologies to capture these emotional trends through the features of a single moment.
[0005] In view of this, the present invention is proposed. Summary of the invention
[0006] The technical problem to be solved by the present invention is to overcome the deficiencies of the prior art and provide a method, medium and system for monitoring customer emotions in telephone customer service, which solves the problems raised in the above-mentioned background technology.
[0007] In order to solve the above technical problems, the basic concept of the technical solution adopted by the present invention is:
[0008] A method for monitoring customer emotions in telephone customer service comprises the following steps:
[0009] Receive the customer's voice signal, first remove the background noise through denoising technology to ensure the clarity of the voice signal. Then, extract features from the original voice signal, including but not limited to Mel-frequency cepstral coefficients (MFCC) or spectrogram; Receive the customer's voice signal, denoise and extract features, including Mel-frequency cepstral coefficients
[0010] The extracted features are input into the residual network (ResNet), where the input data is processed by the multi-head attention mechanism (MHA). This mechanism divides the input feature data into multiple subspaces (heads), each head undergoes a linear transformation through independent query, key, and value, and then performs self-attention calculation on each head to extract multiple emotion features; the extracted features are input into the residual network and processed by the multi-head attention mechanism, where each head independently performs a linear transformation of query, key, and value and calculates self-attention to extract emotion features, and the output results of all heads are concatenated and passed through the fully connected layer to obtain the final emotion feature representation;
[0011] Feed the output of the multi-head attention mechanism into the LSTM model to decouple short-term sentiment fluctuations from long-term sentiment trends. Long-term sentiment trends represent continuous changes in customer sentiment, while short-term fluctuations represent immediate sentiment changes.
[0012] Use the Softmax classifier to classify the emotional features and output the customer's emotional label. Emotional labels include but are not limited to anger, anxiety, and calmness.
[0013] Optionally, the steps of receiving the customer's voice signal, denoising and extracting features, wherein the features include Mel-frequency cepstral coefficients, are as follows:
[0014] Receive the customer's voice signal x(t), where t represents the time series. Then, use a bandpass filter to denoise the voice signal to obtain the signal x clean (t);
[0015] For the denoised signal x clean (t) Perform frame processing, divide the speech signal into multiple small segments, each segment is 20 to 30 milliseconds, and the overlapping part is half of each frame. The signal of each frame is processed by windowing function to reduce the influence of spectrum leakage. Its expression is: frame (n) = x clean (n)·w(n), where x frame (n) is the signal of each frame after windowing, w(n) is the window function (e.g., transparent window);
[0016] For each frame of signal x frame(n) Perform fast Fourier transform (FFT) to convert the time domain signal into a frequency domain signal, thereby obtaining the spectrum X(k), which is expressed as: Among them, X(k) is the spectrum coefficient, k is the frequency index, and N is the number of sampling points per frame;
[0017] Use a Mel-frequency filter bank to filter the spectrum |X(k)| 2 After processing, the Mel frequency energy is obtained. The logarithmic scale of the Mel scale is more in line with the auditory perception of the human ear. The output energy of each Mel filter is calculated as: Among them, E m is the output energy of the first m Mel filters, H m (k) is the frequency response of the mth filter;
[0018] Logarithmically compress the energy output by the Mel frequency filter bank to obtain the logarithmic Mel frequency energy To enhance the contribution of low-energy frequency bands to emotion recognition, the logarithmic Mel-frequency energy is then discrete cosine transformed and converted into the cepstral domain to obtain the Mel-frequency cepstral coefficients. Where MFCC[n] is the first n Mel-frequency cepstral coefficients, and M is the number of Mel filters;
[0019] The process is processed by a multi-head attention mechanism, where each head independently performs linear transformation of query, key, and value and calculates self-attention to extract sentiment features:
[0020] The extracted Mel frequency cepstral coefficients MFCC are input into the residual network (ResNet). The Mel frequency cepstral coefficients MFCC are mapped through each layer of the residual network to obtain the intermediate output And add it with the Mel frequency cepstral coefficient MFCC to get the final output ResNet(x), which is expressed as: in, is the mapping operation of the residual network, W i is the weight of the network layer, It is the result after network layer transformation.
[0021] The feature ResNet(x) output by the residual network is input into the multi-head attention mechanism (MHA) for further emotion feature extraction. The multi-head attention mechanism divides the input feature data x into multiple subspaces (i.e., multiple "heads"). Each head is linearly transformed through independent query, key, and value matrices to generate the self-attention of each head. Q = ResNet(x)W Q ,K=ResNet(x)W K , V = ResNet(x)W V, where: ResNet(x) is the input feature (the feature output from the residual network); W Q , W K , W V It is the weight matrix of query, key and value, which is used to map the input features to different spaces; Q, K, V are the matrices of query, key and value respectively.
[0022] For each head, self-attention is calculated, that is, the attention weight is calculated by the dot product of the query Q and the key K, and then the result is normalized using Softmax, and the value V is weighted summed according to the weight. The self-attention calculation formula is: Where: Q i , K i , V i are the query, key, and value matrices of the i-th head respectively; d k is the dimension of the key matrix, is the scaling factor used to scale the dot product result; softmax(·) is the Softmax function used to normalize the attention weights of the calculation; Attention i is the attention output of the i-th head. Finally, the extracted MFCC features are normalized to eliminate the feature scale differences between different speech samples.
[0023] Optionally, the steps of concatenating the output results of all heads and passing them through a fully connected layer to obtain the final emotion feature representation are:
[0024] The output results of all heads are concatenated to obtain a multi-dimensional emotion feature vector. This concatenation result contains emotion information of multiple subspaces, providing rich feature representation for subsequent emotion analysis;
[0025] The concatenated multi-head attention output passes through the fully connected layer and further optimizes the feature representation through linear transformation. The function of this layer is to map the multi-head output to the emotion feature space so that the emotion features meet the subsequent emotion classification tasks.
[0026] Optionally, the LSTM model includes a forget gate f t , input gate i t , output gate o t and cell status C t . Input is the current time further data x t and the hidden state h at the previous time step t-1 , the output is the further hidden state h at the current time t . Forget Gate f t Controls the retention of information about the previous cell state, input gate i t Controls whether the current data is added to the unit state, output gate o tControls the influence of cell state on hidden state. The weight range of the gated unit is [0,1], which is controlled by the activation function sigmoid. The cell state is nonlinearly mapped by the activation function tanh.
[0027] A multi-scale temporal convolution module is introduced before the input of the LSTM model. The module includes three parallel temporal convolution layers. The convolution kernel sizes of the three parallel temporal convolution layers are 3, 5, and 7 respectively. The speech emotion features are extracted at different time scales. The short time window convolution is used to capture short-term emotion fluctuations, and the long time window convolution is used to model long-term emotion trends. The output of the multi-scale temporal convolution module is consistent with the cell state C of the LSTM. t and the hidden state h t Combined, it is used to decouple short-term sentiment fluctuations from long-term sentiment trends.
[0028] Optionally, the output of the multi-head attention mechanism is fed into a multi-scale temporal convolution module, which consists of multiple parallel temporal convolutional layers, each of which uses convolutional kernels of different sizes (3, 5, 7) to capture short-term and long-term temporal features:
[0029] The convolution kernel size of the short time window is 3 to extract local short-term emotional fluctuations;
[0030] The convolution kernel size of the long time window is 7 to capture the overall long-term sentiment trend;
[0031] The output of the multi-scale time convolution module is used as input and fed into the LSTM model for time series modeling. The core function of the LSTM model is to further learn the temporal dependency of emotional changes, and to store and represent long-term emotional trends and short-term emotional fluctuations through its cell state and hidden state, respectively. Its expression is: b t =o t ·tanh(S t ) Among them, S t Represents the long-term sentiment trend, recorded by the unit state; b t Represents short-term emotional fluctuations and is output through hidden states; f t ,i t , o t They are the weights of the forget gate, input gate, and output gate respectively; is the further candidate cell state at the current time.
[0032] Optionally, use the Softmax classifier to classify the emotion features and output the customer's emotion label as follows:
[0033] The Softmax classifier first performs a linear transformation on the input emotion features to generate a score for each emotion category. The score is calculated as f c(x) = W c ·x+b c , where f c (x) is the score of emotion category c, W c is the weight matrix of the classifier, which represents the weight of category c on the input feature x, b c is the bias term of category c, and x is the input emotional feature;
[0034] The above emotion category scores are converted into probability distributions through the Softmax function, which represents the probability that the input feature belongs to each emotion category. The expression is: Where P(y=c|x) is the predicted probability that the input emotion feature x belongs to category c, and f c (x) is the score of category c, is the exponential sum of all emotion category scores, used for normalization;
[0035] According to the output probability of the Softmax function, select the emotion category c with the highest probability * As the final emotion label. The selected rules are in, c is the final emotional label.
[0036] A detection system includes: a processor, wherein the processor is used to execute the customer emotion monitoring method in telephone customer service.
[0037] A storage medium stores one or more computer instructions, wherein the one or more computer instructions are used to implement the customer emotion monitoring method in telephone customer service.
[0038] After adopting the above technical solution, the present invention has the following beneficial effects compared with the prior art. Of course, any product implementing the present invention does not necessarily need to achieve all the advantages described below at the same time:
[0039] 1. Through the multi-head attention mechanism, the extracted speech features are deeply processed, and the input features are divided into multiple attention heads. Each head independently calculates the linear transformation of the query, key, and value. This mechanism can adaptively capture the key emotional features in the speech signal. After the output results of each attention head are spliced, they are fused into a high-dimensional feature representation through a fully connected layer, thereby comprehensively reflecting the local and global emotional information in the speech. The output of the multi-head attention mechanism is then sent to the LSTM model for decoupling: it can capture short-term fluctuations in emotions and fully preserve the long-term trends of emotions. This decoupling mechanism makes the model more sensitive and accurate to changes in emotions at multiple time scales, providing rich and accurate emotional information for emotion recognition tasks.
[0040] 2. Extract the features of speech signals at different time scales in parallel through convolution kernels of different sizes. Convolution kernels with short time windows are suitable for capturing short-term emotional fluctuations in speech signals, such as sudden changes in tone or faster speaking speed; convolution kernels with long time windows can model long-term trends in emotions, such as the gradual calming or strengthening of emotions during a call. This multi-scale feature extraction method allows the model to focus on local and global emotional features at the same time, avoiding the limitations of a single scale. In addition, the multi-scale time convolution module improves the computational efficiency of the model through parallel processing, reduces the complexity of long time series feature modeling, and provides a richer and more comprehensive feature representation for subsequent emotion recognition tasks.
[0041] The specific implementation modes of the present invention are further described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] The drawings described below are only some embodiments. For those skilled in the art, other drawings can be obtained based on these drawings without creative work. In the drawings:
[0043] Figure 1 Flowchart of customer emotion detection method.
[0044] It should be noted that these drawings and textual descriptions are not intended to limit the conceptual scope of the present invention in any way, but are intended to illustrate the concept of the present invention for those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION
[0045] The present invention will now be described in further detail with reference to the accompanying drawings.
[0046] See also Figure 1 As shown, in this embodiment, a method for monitoring customer emotions in telephone customer service is provided, including the following steps:
[0047] Receive the customer's voice signal, first remove the background noise through denoising technology to ensure the clarity of the voice signal. Then, extract features from the original voice signal, including but not limited to Mel-frequency cepstral coefficients (MFCC) or spectrogram; Receive the customer's voice signal, denoise and extract features, including Mel-frequency cepstral coefficients
[0048] The extracted features are input into the residual network (ResNet), where the input data is processed by the multi-head attention mechanism (MHA). This mechanism divides the input feature data into multiple subspaces (heads), each head undergoes a linear transformation through independent query, key, and value, and then performs self-attention calculation on each head to extract multiple emotion features; the extracted features are input into the residual network and processed by the multi-head attention mechanism, where each head independently performs a linear transformation of query, key, and value and calculates self-attention to extract emotion features, and the output results of all heads are concatenated and passed through the fully connected layer to obtain the final emotion feature representation;
[0049] Feed the output of the multi-head attention mechanism into the LSTM model to decouple short-term sentiment fluctuations from long-term sentiment trends. Long-term sentiment trends represent continuous changes in customer sentiment, while short-term fluctuations represent immediate sentiment changes.
[0050] Use the Softmax classifier to classify the emotional features and output the customer's emotional label. Emotional labels include but are not limited to anger, anxiety, and calmness.
[0051] In this embodiment, the steps of receiving the voice signal of the customer, denoising and extracting features, wherein the features include Mel-frequency cepstral coefficients, are as follows:
[0052] Receive the customer's voice signal x(t), where t represents the time series. Then, use a bandpass filter to denoise the voice signal to obtain the signal x clean (t);
[0053] For the denoised signal x clean (t) Perform frame processing, divide the speech signal into multiple small segments, each segment is 20 to 30 milliseconds, and the overlapping part is half of each frame. The signal of each frame is processed by windowing function to reduce the influence of spectrum leakage. Its expression is: frame (n) = x clean (n)·w(n), where x frame (n) is the signal of each frame after windowing, w(n) is the window function (e.g., transparent window);
[0054] For each frame of signal x frame (n) Perform fast Fourier transform (FFT) to convert the time domain signal into a frequency domain signal, thereby obtaining the spectrum X(k), which is expressed as: Among them, X(k) is the spectrum coefficient, k is the frequency index, and N is the number of sampling points per frame;
[0055] Use a Mel-frequency filter bank to filter the spectrum |X(k)| 2After processing, the Mel frequency energy is obtained. The logarithmic scale of the Mel scale is more in line with the auditory perception of the human ear. The output energy of each Mel filter is calculated as: Among them, E m is the output energy of the first m Mel filters, H m (k) is the frequency response of the mth filter;
[0056] Logarithmically compress the energy output by the Mel frequency filter bank to obtain the logarithmic Mel frequency energy To enhance the contribution of low-energy frequency bands to emotion recognition, the logarithmic Mel-frequency energy is then discrete cosine transformed and converted into the cepstral domain to obtain the Mel-frequency cepstral coefficients. Where MFCC[n] is the first n Mel-frequency cepstral coefficients, and M is the number of Mel filters;
[0057] The process is processed by a multi-head attention mechanism, where each head independently performs linear transformation of query, key, and value and calculates self-attention to extract sentiment features:
[0058] The extracted Mel frequency cepstral coefficients MFCC are input into the residual network (ResNet). The Mel frequency cepstral coefficients MFCC are mapped through each layer of the residual network to obtain the intermediate output And add it with the Mel frequency cepstral coefficient MFCC to get the final output ResNet(x), which is expressed as: in, is the mapping operation of the residual network, W i is the weight of the network layer, It is the result after network layer transformation.
[0059] The feature ResNet(x) output by the residual network is input into the multi-head attention mechanism (MHA) for further emotion feature extraction. The multi-head attention mechanism divides the input feature data x into multiple subspaces (i.e., multiple "heads"). Each head is linearly transformed through independent query, key, and value matrices to generate the self-attention of each head. Q = ResNet(x)W Q ,K=ResNet(x)W K , V = ResNet(x)W V , where: ResNet(x) is the input feature (the feature output from the residual network); W Q , W K , W V It is the weight matrix of query, key and value, which is used to map the input features to different spaces; Q, K, V are the matrices of query, key and value respectively.
[0060] For each head, self-attention is calculated, that is, the attention weight is calculated by the dot product of the query Q and the key K, and then the result is normalized using Softmax, and the value V is weighted summed according to the weight. The self-attention calculation formula is: Where: Q i , K i , V i are the query, key, and value matrices of the i-th head respectively; d k is the dimension of the key matrix, is the scaling factor used to scale the dot product result; softmax(·) is the Softmax function used to normalize the attention weights of the calculation; Attention i is the attention output of the i-th head. Finally, the extracted MFCC features are normalized to eliminate the feature scale differences between different speech samples.
[0061] In this embodiment, the steps of concatenating the output results of all heads and passing through a fully connected layer to obtain the final emotion feature representation are:
[0062] The output results of all heads are concatenated to obtain a multi-dimensional emotion feature vector. This concatenation result contains emotion information of multiple subspaces, providing rich feature representation for subsequent emotion analysis;
[0063] The concatenated multi-head attention output passes through the fully connected layer and further optimizes the feature representation through linear transformation. The function of this layer is to map the multi-head output to the emotion feature space so that the emotion features meet the subsequent emotion classification tasks.
[0064] In this embodiment, the LSTM model includes a forget gate f t , input gate i t , output gate o t and cell status C t . Input is the current time further data x t and the hidden state h at the previous time step t-1 , the output is the further hidden state h at the current time t . Forget Gate f t Controls the retention of information about the previous cell state, input gate i t Controls whether the current data is added to the unit state, output gate o t Controls the influence of the cell state on the hidden state. The weight range of the gated unit is [0, 1], which is controlled by the activation function sigmoid. The cell state is nonlinearly mapped by the activation function tanh.
[0065] A multi-scale temporal convolution module is introduced before the input of the LSTM model. The module includes three parallel temporal convolution layers. The convolution kernel sizes of the three parallel temporal convolution layers are 3, 5, and 7 respectively. The speech emotion features are extracted at different time scales. The short time window convolution is used to capture short-term emotion fluctuations, and the long time window convolution is used to model long-term emotion trends. The output of the multi-scale temporal convolution module is consistent with the cell state C of the LSTM. t and the hidden state h t Combined, it is used to decouple short-term sentiment fluctuations from long-term sentiment trends.
[0066] In this embodiment, the output of the multi-head attention mechanism is fed into the LSTM model to decouple short-term sentiment fluctuations from long-term sentiment trends:
[0067] The output of the multi-head attention mechanism is fed into a multi-scale temporal convolution module, which consists of multiple parallel temporal convolutional layers, each of which uses convolutional kernels of different sizes (3, 5, 7) to capture short-term and long-term temporal features:
[0068] The convolution kernel size of the short time window is 3 to extract local short-term emotional fluctuations;
[0069] The convolution kernel size of the long time window is 7 to capture the overall long-term sentiment trend;
[0070] The output of the multi-scale time convolution module is used as input and fed into the LSTM model for time series modeling. The core function of the LSTM model is to further learn the temporal dependency of emotional changes, and to store and represent long-term emotional trends and short-term emotional fluctuations through its cell state and hidden state, respectively. Its expression is: b t =o t ·tanh(S t ) Among them, S t Represents the long-term sentiment trend, recorded by the unit state; b t Represents short-term emotional fluctuations and is output through hidden states; f t ,i t , o t They are the weights of the forget gate, input gate, and output gate respectively; is the further candidate cell state at the current time.
[0071] In this embodiment, the steps of using a Softmax classifier to classify the emotion features and outputting the customer's emotion label are as follows:
[0072] The Softmax classifier first performs a linear transformation on the input emotion features to generate a score for each emotion category. The score is calculated as f c (x) = W c ·x+bc , where f c (x) is the score of emotion category c, W c is the weight matrix of the classifier, which represents the weight of category c on the input feature x, b c is the bias term of category c, and x is the input emotional feature;
[0073] The above emotion category scores are converted into probability distributions through the Softmax function, which represents the probability that the input feature belongs to each emotion category. The expression is: Where P(y=c|x) is the predicted probability that the input emotion feature x belongs to category c, and f c (x) is the score of category c, is the exponential sum of all emotion category scores, used for normalization;
[0074] According to the output probability of the Softmax function, select the emotion category c with the highest probability * As the final emotion label. The selected rules are Among them, c * is the final emotional label.
[0075] A detection system includes: a processor, wherein the processor is used to execute the customer emotion monitoring method in telephone customer service.
[0076] A storage medium stores one or more computer instructions, wherein the one or more computer instructions are used to implement the customer emotion monitoring method in telephone customer service.
[0077] Comparative example: The solution based on convolutional neural network (CNN) extracts local emotional features of speech through convolutional layers and classifies them. It is suitable for processing short speech segments, but has limitations in capturing long-term dependencies. These two solutions represent typical methods of traditional machine learning and deep learning, respectively, as comparison benchmarks.
[0078] To ensure the fairness of the comparative experiment, all schemes were tested under the same dataset, hardware environment and hyperparameter settings. The speech datasets used in the experiment are CASIA and EmoDB. Data preprocessing includes silence removal, Gaussian noise addition and waveform displacement expansion to ensure consistent data quality and diversity. The hardware environment is Windows 10 system, RTX3060 graphics card and 32GB memory, and the code is based on Python 3.9. In the hyperparameter setting, the batch size is 64, the number of training times is 200, the learning rate is 0.002, the optimizer uses Adam, and the loss function is cross entropy loss. All models input the same feature data, and the division ratio is 80% for training set, 10% for validation set, and 10% for test set.
[0079] The experiment tested the classification performance of the three schemes under the same conditions and calculated the accuracy as the performance evaluation index. Accuracy is defined as the proportion of samples correctly classified by the model on the test set, and its expression is: Among them, TP is the number of positive samples correctly classified; TN is the number of negative samples correctly classified; FP is the number of negative samples misclassified as positive; FN is the number of positive samples misclassified as negative. The comparison results are shown in the table below.
[0080] plan CASIA accuracy (%) EmoDB accuracy (%) The present invention 91.5 89.8 CNN 83.7 81.5
[0081] Experimental results show that on the CASIA dataset, the accuracy of the present invention is 91.5%, which is significantly higher than CNN (83.7%); on the EmoDB dataset, the accuracy of the present invention is 89.8%, which is also significantly better than CNN (81.5%). The results show that the present invention has superior performance on multilingual and diversified speech emotion data.
[0082] The present invention introduces a multi-head attention mechanism (MHA) to assign different weights to different time steps of the input sequence data, emphasizing key features related to emotions. This mechanism allows the model to focus on specific emotional features from multiple subspaces in parallel, thereby enhancing the model's feature extraction capabilities in complex speech data. CNN only relies on local convolutional features and has limited ability to capture global dependencies, while the MHA in the present invention can better handle the potential global emotional correlation in sequence features, making the model perform better on long time series.
[0083] The multi-scale time convolution module extracts features of different time scales through parallel convolution layers. The convolution of short time windows captures short-term emotional fluctuations, and the convolution of long time windows models the overall emotional trend. This feature decomposition method makes the model highly adaptable to emotional changes at different time granularities. The fixed convolution kernel size of CNN limits its ability to model features of different time scales, and the present invention greatly improves the modeling ability of diversified emotional features through multi-scale feature extraction.
[0084] The present invention combines the LSTM model and effectively captures the dynamic characteristics of emotion evolution over time through its unique gating mechanism (forget gate, input gate, output gate). The unit state of LSTM stores long-term emotional trends, and the hidden state captures short-term emotional fluctuations, thereby achieving accurate decoupling of emotional features in time series. In contrast, CNN lacks sequence modeling capabilities and has difficulty in dealing with long-term dependency issues in emotion recognition tasks, and the introduction of LSTM makes up for this shortcoming.
[0085] A detection system includes: a processor, wherein the processor is used to execute the customer emotion monitoring method in telephone customer service.
[0086] A storage medium stores one or more computer instructions, wherein the one or more computer instructions are used to implement the customer emotion monitoring method in telephone customer service.
[0087] The present invention is not limited to the above-mentioned embodiments. Anyone should be aware that any structural changes made under the enlightenment of the present invention, and any technical solutions that are the same or similar to the present invention, fall within the protection scope of the present invention. The technology, shape, and structural parts not described in detail in the present invention are all well-known technologies.
Claims
1. A method for monitoring customer emotions in telephone customer service, characterized in that: The following steps are involved: Receive the customer's voice signal, denoise it and extract features; The extracted features are input into the residual network and processed by the multi-head attention mechanism, where each head independently performs linear transformation of query, key, and value and calculates self-attention to extract emotional features. The output results of all heads are concatenated and passed through the fully connected layer to obtain the final emotional feature representation; The output of the multi-head attention mechanism is fed into the LSTM model to decouple short-term sentiment fluctuations from long-term sentiment trends; Use the Softmax classifier to classify the sentiment features and output the customer’s sentiment label.
2. A method for monitoring customer emotions in telephone customer service according to claim 1, characterized in that: The steps of receiving the customer's voice signal, denoising and extracting features, including Mel-frequency cepstral coefficients, are as follows: Receive the customer's voice signal x(t), where t represents the time series. Then, use a bandpass filter to denoise the voice signal to obtain the signal x clean (t); For the denoised signal x clean (t) Perform frame processing, divide the speech signal into multiple small segments, each segment is 20 to 30 milliseconds, and the overlapping part is half of each frame. The signal of each frame is processed by windowing function to reduce the influence of spectrum leakage. Its expression is: frame (n) = x clean (n)·w(n), where x frame (n) is the signal of each frame after windowing, w(n) is the window function; For each frame of signal x frame (n) Perform fast Fourier transform to convert the time domain signal into a frequency domain signal, thereby obtaining the spectrum X(k), which is expressed as: Among them, X(k) is the spectrum coefficient, k is the frequency index, and N is the number of sampling points per frame; Use a Mel-frequency filter bank to filter the spectrum |X(k)| 2 After processing, the Mel frequency energy is obtained. The logarithmic scale of the Mel scale is more in line with the auditory perception of the human ear. The output energy of each Mel filter is calculated as: Among them, E m is the output energy of the first m Mel filters, H m (k) is the frequency response of the mth filter; Logarithmically compress the energy output by the Mel frequency filter bank to obtain the logarithmic Mel frequency energy To enhance the contribution of low-energy frequency bands to emotion recognition, the logarithmic Mel-frequency energy is then discrete cosine transformed to the cepstrum domain to obtain the Mel-frequency cepstrum coefficients. Where MFCC[n] is the first n Mel-frequency cepstral coefficients, and M is the number of Mel filters.
3. A method for monitoring customer emotions in telephone customer service according to claim 1, characterized in that: The process is processed by a multi-head attention mechanism, where each head independently performs linear transformation of query, key, and value and calculates self-attention to extract sentiment features: The extracted Mel frequency cepstral coefficients MFCC are input into the residual network, and the Mel frequency cepstral coefficients MFCC are mapped through each layer of the residual network to obtain the intermediate output And add it with the Mel frequency cepstral coefficient MFCC to get the final output ResNet(x), which is expressed as: in, is the mapping operation of the residual network, W i is the weight of the network layer, It is the result after transformation at the network layer; The feature ResNet(x) output by the residual network is input into the multi-head attention mechanism for further emotion feature extraction. The multi-head attention mechanism divides the input feature data x into multiple subspaces. Each head is linearly transformed through independent query, key and value matrices to generate the self-attention of each head, Q = ResNet(x)W Q ,K=ResNet(x)W K , V = ResNet(x)W V , where: ResNet(x) is the input feature; W Q , W K , W V is the weight matrix of query, key and value, which is used to map the input features to different spaces; Q, K, V are the matrices of query, key and value respectively; For each head, self-attention is calculated, that is, the attention weight is calculated by the dot product of the query Q and the key K, and then the result is normalized using Softmax, and the value V is weighted summed according to the weight. The self-attention calculation formula is: Where: Q i , K i , V i are the query, key, and value matrices of the i-th head respectively; d k is the dimension of the key matrix, is the scaling factor used to scale the dot product result; softmax(·) is the Softmax function used to normalize the attention weights of the calculation; Attention i is the attention output of the first i heads. Finally, the extracted MFCC features are normalized to eliminate the feature scale differences between different speech samples.
4. A method for monitoring customer emotions in telephone customer service according to claim 1, characterized in that: The steps of concatenating the output results of all heads and passing them through the fully connected layer to obtain the final emotion feature representation are as follows: The output results of all heads are concatenated to obtain a multi-dimensional emotion feature vector. This concatenation result contains emotion information of multiple subspaces, providing rich feature representation for subsequent emotion analysis. The concatenated multi-head attention output passes through the fully connected layer and further optimizes the feature representation through linear transformation. The function of this layer is to map the multi-head output to the emotion feature space so that the emotion features meet the subsequent emotion classification tasks.
5. The method for monitoring customer emotions in telephone customer service according to claim 1, characterized in that: The LSTM model includes a forget gate f t , input gate i t , output gate o t and cell status C t , input is the current time further data x t and the hidden state h at the previous time step t-1 , the output is the further hidden state h at the current time t , forget gate f t Controls the retention of information about the previous cell state, input gate i t Controls whether the current data is added to the unit state, output gate o t Controls the influence of cell state on hidden state. The weight range of gated unit is [0, 1], which is controlled by activation function sigmoid. The cell state is nonlinearly mapped by activation function tanh. A multi-scale time convolution module is introduced before the input of the LSTM model. The module includes three parallel time convolution layers. The convolution kernel sizes of the three parallel time convolution layers are 3, 5, and 7, respectively. Speech emotion features are extracted at different time scales, among which short-time window convolution is used to capture short-term emotional fluctuations, and long-time window convolution is used to model long-term emotional trends.
6. A method for monitoring customer emotions in telephone customer service according to claim 1, characterized in that: The output of the multi-head attention mechanism is fed into the LSTM model. The steps to decouple short-term sentiment fluctuations from long-term sentiment trends are: The output of the multi-head attention mechanism is fed into a multi-scale temporal convolution module, which consists of multiple parallel temporal convolutional layers, each of which uses convolutional kernels of different sizes to capture short-term and long-term temporal features. The output of the multi-scale time convolution module is used as input and sent to the LSTM model for time series modeling. The core function of the LSTM model is to further learn the time dependency of emotional changes and store and represent long-term emotional trends and short-term emotional fluctuations through its cell state and hidden state. The expression is: b t =o t ·tanh(S t ) Among them, S t Represents the long-term sentiment trend, recorded by the unit state; b t Represents short-term emotional fluctuations and is output through hidden states; f t ,i t , o t They are the weights of the forget gate, input gate, and output gate respectively; is the further candidate cell state at the current time.
7. The method for monitoring customer emotions in telephone customer service according to claim 1, characterized in that: Use the Softmax classifier to classify the emotional features and output the customer's emotional label as follows: The Softmax classifier first performs a linear transformation on the input emotion features to generate a score for each emotion category. The score is calculated as f c (x)=W c ·x+b c , where f c (x) is the score of emotion category c, W c is the weight matrix of the classifier, which represents the weight of category c on the input feature x, b c is the bias term of category c, and x is the input emotional feature; The above emotion category scores are converted into probability distributions through the Softmax function, which represents the probability that the input feature belongs to each emotion category. The expression is: Where P(y=c|x) is the predicted probability that the input emotion feature x belongs to category c, and f c (x) is the score of category c, is the exponential sum of all emotion category scores, used for normalization; According to the output probability of the Softmax function, select the emotion category c with the highest probability * As the final emotion label, the selected rule is Among them, c * is the final emotion label.
8. A detection system, characterized in that: include: A processor, wherein the processor is used to execute the customer emotion monitoring method in telephone customer service according to any one of claims 1 to 7.
9. A storage medium, characterized in that: The storage medium stores one or more computer instructions, and the one or more computer instructions are used to implement the customer emotion monitoring method in telephone customer service as described in any one of claims 1 to 7.