Customer service response queuing method and customer service response queuing device based on emotion recognition
By converting voice into text in the telephone customer service system and analyzing customer emotional states in combination with neural networks, predicting willing waiting time, the problem of insufficient customer emotional state recognition in the prior art is solved, and the efficiency of telephone consultation and customer satisfaction are improved.
Patent Information
- Application Number
- CN202510487587.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-04
AI Technical Summary
The prior art cannot effectively identify the emotional state of customers in telephone customer service, resulting in too long queue time, improper resource allocation and reduced customer satisfaction. The complexity bias of existing voice sentiment detection models leads to waste of computing resources and insufficient real-time performance.
By receiving voice statement audio converted into text, the response department is obtained using Bayesian model classification, and combined with dimensional affective sub-neural networks and discrete affective sub-neural networks to analyze customer emotional states, predict willing waiting time, and optimize queuing strategies.
It improves the efficiency of telephone consultation, improves customer experience, realizes accurate identification of emotional state and optimized allocation of resources, and reduces customer waiting time.
Smart Images

Figure CN120263902A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine models, and particularly to a customer service response queuing method and a customer service response queuing device based on emotion recognition. Background Art
[0002] Telephone customer service is an important channel for enterprises to communicate with customers. Customers express their needs, complaints or feedbacks through telephone calls. Traditional telephone customer service relies on digital selection after connection to judge the business needs of customers and conduct different business diverts, such as pre-sales business, after-sales business, complaint business, and extreme business scenarios. With the development of artificial intelligence (AI) and speech recognition technology, the existing technology has been able to achieve that after a customer service call is connected, the user can directly and simply describe their problems, and then the involved business scenarios can be matched and automatically transferred to the correct business processing department.
[0003] However, in the actual application process, there are problems such as the long waiting time of customers caused by a large number of queuing customers during peak hours, the inability to effectively distinguish key customers from general customers (for example, a long waiting time for pre-sales consultation may result in missed business opportunities), and the complaint customers may be assigned to general customers due to allocation problems (in fact, they should be assigned to more experienced customer service representatives). Traditional means can only identify keywords in the customer's voice for transfer and cannot directly identify the customer's needs and emotional states, which is not only inefficient but also affects customer satisfaction and even misses business opportunities. The emotional state of customers during a phone call is crucial for problem-solving and service quality improvement. If the customer's emotions can be recognized in a timely manner before determining the specific transfer channel, such as in the following scenarios:
[0004] For a customer with an obviously excited voice, the queuing time can be slightly shortened. If a customer shows an obvious pre-sales purchase demand, they can be quickly transferred and assigned to the pre-sales customer service; in addition, it can play a monitoring role in the communication dialogue during customer service reception, etc.
[0005] For example, it includes the following application scenarios:
[0006] 1. Customer emotion analysis: By analyzing the emotional state of customers during a phone call, such as being happy, angry or frustrated, enterprises can discover the needs and pain points of customers, so as to better improve products and services.
[0007] 2. Evaluation of customer service representatives' performance: Emotion recognition technology can be used to evaluate the performance of customer service representatives, understand whether they can effectively solve customer problems, and whether their service attitudes and tones are appropriate, providing an objective and quantitative evaluation standard for enterprises.
[0008] 3. Priority processing: According to the prediction of the customer's emotional state and waiting time, high-priority customers are processed first, optimizing resource allocation and reducing customer waiting time.
[0009] 4. Personalized service: Enterprises can segment customers according to their emotional states and provide personalized services for customers with different emotional states. For example, a quick response channel can be provided for angry customers, and more preferential information can be provided for satisfied customers.
[0010] 5. Adding an emotional dimension to service quality inspection can more accurately detect the level of service quality and facilitate post-event optimization and improvement. As a service reference indicator, detecting negative user emotions in a conversation can trigger the automatic synchronization of recordings to quality inspection personnel to complete quality inspection and improvement.
[0011] 6. By mining the value of complaint voice files and providing decision support for operations based on the identified customer emotions, semantic information, etc., potential dissatisfaction tendencies / disconnection tendencies of customers can be obtained in advance, continuously improving the service experience.
[0012] In related technologies, voice emotion detection refers to the process of automatically identifying and extracting emotional information from voice signals. Through artificial intelligence technology, computers can understand and interpret human emotional expressions. For this purpose, a series of voice emotion detection models have been invented in the industry. However, there is generally a complexity bias in the development process of current voice emotion detection models, that is, the technical bias that more parameters and more complex structures can bring better performance, and then blindly pursue using more huge and complex models to explain human emotional expressions. This has led to poor interpretability of the models and a great waste of computing resources. These defects are further amplified in scenarios with high requirements for real-time performance and decision-making rationality such as voice customer service, ultimately resulting in a decrease in customer satisfaction. Summary of the Invention
[0013] To solve at least one of the above problems, the first embodiment of the present invention provides a customer service response queuing method based on emotion recognition, including:
[0014] Receiving the voice statement audio of a caller making a customer service call and converting the voice statement audio into a statement text;
[0015] According to the statement text, using a Bayesian model for classification to obtain the response department;
[0016] According to the voice statement audio and the statement text, using a decoupled emotion calculation model to analyze the emotional state of the caller and predict the willing waiting time based on the emotional state;
[0017] Queuing the call of the caller in the response department according to the willing waiting time.
[0018] For example, in the customer service response queuing method provided in some embodiments of the present application, the queuing the call of the caller in the response department according to the willing waiting time further includes:
[0019] Calculate the queuing priority of the caller according to the willing waiting time and the consumption record of the caller;
[0020] Queue the call answering of the caller in the answering department according to the willing waiting time and the queuing priority.
[0021] For example, in the customer service answering queuing method provided by some embodiments of the present application, the decoupled emotion calculation model is deployed on a distributed cluster device, and the distributed cluster device includes an independent and electrically connected first distributed device, a second distributed device, and a third distributed device. The decoupled emotion model includes a dimensional emotion sub-neural network model deployed on the first distributed device, a discrete emotion sub-neural network model deployed on the second distributed device, and a waiting time prediction sub-neural network model deployed on the third distributed device;
[0022] The further steps of analyzing the emotional state of the caller using the decoupled emotion calculation model according to the voice statement audio and the statement text, and predicting the willing waiting time of the caller based on the emotional state include:
[0023] According to the statement text, use the dimensional emotion sub-neural network model to extract the first emotional feature of the caller and output a valence arousal vector;
[0024] According to the voice statement audio, use the discrete emotion sub-neural network model to extract the second emotional feature and the third emotional feature of the caller respectively and output an emotion label encoding;
[0025] According to the valence arousal vector and the emotion label encoding, use the waiting time prediction sub-neural network model to obtain the willing waiting time of the caller.
[0026] For example, in the customer service answering queuing method provided by some embodiments of the present application,
[0027] The dimensional emotion sub-neural network model includes an embedding layer, a first bidirectional LSTM layer, a second bidirectional LSTM layer, a first Highway layer, and a first fully connected layer connected in sequence. The further steps of using the dimensional emotion sub-neural network model to extract the first emotional feature of the caller according to the statement text and output a valence arousal vector include:
[0028] Encode the statement text and generate a statement text encoding, and input the statement text encoding into the dimensional emotion sub-neural network model to obtain the first emotional feature of the caller and output the valence arousal vector;
[0029] The discrete emotion sub-neural network model includes a first branch, a second branch, and a third branch. The first branch includes a speech excitement index pooling layer and a one-dimensional convolutional layer connected in sequence. The second branch includes a separable two-dimensional convolutional layer and a Flatten layer connected in sequence. The third branch includes a first splicing layer, a second Highway layer, and a second fully connected layer. The step of using the discrete emotion sub-neural network model to extract the second emotion feature and the third emotion feature of the caller respectively according to the speech statement audio and output an emotion label encoding further includes:
[0030] Generating a Mel spectrogram and a sound spectrogram according to the speech statement audio, superimposing the Mel spectrogram and the sound spectrogram to generate a multi-channel audio feature map, using the speech statement audio as the input of the first branch and outputting the second emotion feature, using the multi-channel audio feature map as the input of the second branch and outputting the third emotion feature, and using the second emotion feature and the third emotion feature as the input of the third branch and outputting the emotion label encoding;
[0031] The waiting time prediction sub-neural network model includes a second splicing layer, a third fully connected layer, and a fourth fully connected layer connected in sequence. The step of using the waiting time prediction sub-neural network model to obtain the willing waiting time according to the valence arousal vector and the emotion label encoding further includes:
[0032] Mapping the valence arousal vector and the emotion label encoding into the willing waiting time corresponding to the caller.
[0033] For example, in the customer service response queuing method provided in some embodiments of the present application, the speech excitement index pooling layer is used to calculate the speech excitement index and perform pooling.
[0034] The speech excitement index is:
[0035]
[0036] where E is the speech excitement index, c is the maximum length of the pooling window, n is the length of the pooling window, E n is the speech excitement index component with the length of the pooling window being n, P is the skewness within the pooling window, K is the kurtosis within the pooling window, is the mean within the pooling window, i is the serial number of the value within the pooling window, x i is the i-th value within the pooling window, and s is the standard deviation within the pooling window.
[0037] For example, in the customer service response queuing method provided in some embodiments of the present application, before classifying the statement text using the Bayesian model to obtain the response department, the customer service response queuing method further includes:
[0038] Establish a first training dataset and train the Bayesian model according to the first training dataset.
[0039] For example, in the customer service response queuing method provided in some embodiments of the present application, before analyzing the emotional state of the caller using the decoupled emotion calculation model based on the voice statement audio and statement text and predicting the willing waiting time based on the emotional state, the customer service response queuing method further includes:
[0040] Establish a second training dataset and train the decoupled emotion calculation model according to the second training dataset.
[0041] For example, in the customer service response queuing method provided in some embodiments of the present application, the step of establishing a first training dataset and training the Bayesian model according to the first training dataset further includes:
[0042] Extract all accepted sample records from the customer service call history;
[0043] Extract one accepted sample record and collect the record text and the responding department therein, convert the record text into a text vector, generate a department code based on the responding department, and the text vector of each accepted sample record and the corresponding department code form a training data in the first training dataset.
[0044] Determine whether there are still uncollected accepted sample records. If so, jump to the step of extracting one accepted sample record and collecting the record text and the responding department therein.
[0045] For example, in the customer service response queuing method provided in some embodiments of the present application, the step of establishing a second training dataset and training the decoupled emotion calculation model according to the second training dataset further includes:
[0046] Extract all unresponded accepted sample records from the customer service call history;
[0047] Extract one unresponded accepted sample record and collect the waiting time and the voice statement audio therein, convert the voice statement audio into a statement text, generate a Mel spectrogram and a sound spectrogram according to the voice statement audio, superimpose the Mel spectrogram and the sound spectrogram to generate a multi-channel audio feature map, and the waiting time, the voice statement audio, the statement text and the multi-channel audio feature map of each unresponded accepted sample record form a training data in the second training dataset.
[0048] Determine whether there are still uncollected unresponded accepted sample records. If so, jump to the step of extracting one unresponded accepted sample record and collecting the waiting time and the voice statement audio therein.
[0049] For example, in the customer service response queuing method provided in some embodiments of the present application, the calculating the queuing priority of the caller according to the willing waiting time and the consumption record of the caller further includes:
[0050] The queuing priority of the caller is:
[0051] R = N / T;
[0052] Q = S * R;
[0053] Wherein, R represents the risk of losing customers, T represents the willing waiting time, N represents the number of callers currently queuing, S represents the consumption record of the caller, and Q is the queuing priority.
[0054] The second embodiment of the present invention provides a customer service response queuing device applying the customer service response queuing method as described in the first embodiment, including a distributed cluster device and a controller. The distributed cluster device includes an independent and electrically connected first distributed device, a second distributed device, and a third distributed device. The decoupled emotion model includes a dimensional emotion sub-neural network model, a discrete emotion sub-neural network model, and a waiting time prediction sub-neural network model. The first distributed device is used to deploy the dimensional emotion sub-neural network model, the second distributed device is used to deploy the discrete emotion sub-neural network model, and the third distributed device is used to deploy the waiting time prediction sub-neural network model; the controller is configured to:
[0055] Receive the voice statement audio of the caller calling the customer service phone, and convert the voice statement audio into a statement text;
[0056] According to the statement text, use a Bayesian model for classification to obtain the response department;
[0057] According to the voice statement audio and the statement text, use the decoupled emotion calculation model to analyze the emotional state of the caller and predict the willing waiting time of the caller based on the emotional state;
[0058] Queue the phone response of the caller in the response department according to the willing waiting time.
[0059] The third embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the method as described in the first embodiment.
[0060] The fourth embodiment of the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the method as described in the first embodiment.
[0061] The beneficial effects of the present invention are as follows:
[0062] In view of the existing problems, the present invention provides a customer service response queuing method and a customer service response queuing device based on emotion recognition. By converting the voice statement audio of the caller in a telephone consultation into a statement text, classifying it through a Bayesian model to obtain the response department according to the statement text, analyzing the emotional state of the caller based on the voice statement audio and the statement text through a decoupled emotion calculation model, predicting the willing waiting time of the caller based on the emotional state, and performing effective queuing according to the willing waiting time. That is, the embodiments provided in the present application identify the emotional state of the caller during the telephone consultation statement, calculate the willing waiting time of the caller based on the emotional state, thus making up for the problems existing in the prior art, effectively improving the efficiency of telephone consultation, improving the user experience of telephone consultation callers, and having a wide application prospect. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0064] Figure 1 A flowchart showing the customer service response queuing method according to an embodiment of the present invention;
[0065] Figure 2 A schematic structural diagram showing the dimensional emotion sub-neural network model according to an embodiment of the present invention;
[0066] Figure 3 A schematic structural diagram showing the discrete emotion sub-neural network model according to an embodiment of the present invention;
[0067] Figure 4 A schematic structural diagram showing the waiting time prediction sub-neural network model according to an embodiment of the present invention;
[0068] Figure 5 A schematic block diagram showing the customer service response queuing device according to an embodiment of the present invention;
[0069] Figure 6 A schematic structural diagram showing a computer device according to another embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0070] To more clearly illustrate the present invention, the present invention will be further described below in conjunction with preferred embodiments and the accompanying drawings. Similar components in the drawings are denoted by the same reference numerals. Those skilled in the art should understand that the content specifically described below is illustrative rather than restrictive, and should not be used to limit the protection scope of the present invention.
[0071] According to the problems existing in the related art, such as Figure 1 As shown, an embodiment of the present invention provides a customer service response queuing method based on emotion recognition, including:
[0072] Receiving the voice statement audio of the caller calling the customer service phone, and converting the voice statement audio into a statement text;
[0073] According to the statement text, using a Bayesian model for classification to obtain the response department;
[0074] According to the voice statement audio and the statement text, using a decoupled emotion calculation model to analyze the emotional state of the caller, and predicting the willing waiting time based on the emotional state;
[0075] Queuing the phone response of the caller in the response department according to the willing waiting time.
[0076] Considering the emotional state of the caller during the phone consultation statement, in this embodiment, the statement text is obtained by converting the voice statement audio of the caller in the phone consultation into text. According to the statement text, classification is performed through a Bayesian model to obtain the response department. According to the voice statement audio and the statement text, the emotional state of the caller is analyzed through a decoupled emotion calculation model, and the willing waiting time of the caller is predicted based on the emotional state, and effective queuing is performed according to the willing waiting time. That is, the embodiment provided in the present application effectively improves the efficiency of phone consultation and improves the user experience of the callers in phone consultation by identifying the emotional state of the caller during the phone consultation statement and calculating the willing waiting time of the caller based on the emotional state, and has broad application prospects.
[0077] To further illustrate the specific implementation manner of this embodiment, taking a specific caller calling the customer service phone for consultation as an example, as Figure 1 shown, the following steps are included:
[0078] The first step is to receive the voice statement audio of the caller calling the customer service phone, and convert the voice statement audio into a statement text.
[0079] In this embodiment, when a caller dials the customer service phone for consultation, the caller briefly describes the problem to be consulted according to the voice prompt, saves the voice described by the user as a voice statement audio, and converts the voice statement audio into a statement text according to the classification requirements of the answering department. Specifically, CMU Sphinx is used to implement the conversion of the voice statement audio into a statement text.
[0080] In the second step, according to the statement text, a Bayesian model is used for classification to obtain the answering department.
[0081] In this embodiment, based on the statement text obtained in the above steps, the statement text is further converted into a TF-IDF matrix, and a multinomial naive Bayesian model is used for classification to obtain the answering department. For example, if it is determined from the statement text that the purpose of the caller's phone consultation is to purchase a trial, the caller is assigned to the "pre-sales department" for answering, that is, the caller's consultation is classified into the corresponding answering department through the Bayesian model. In this embodiment, the multinomial naive Bayesian model analyzes the caller's needs based on the TF-IDF matrix of the text and transfers the caller to the corresponding answering department. The TF-IDF matrix can accurately reflect the relationship between the word frequency and the needs in the statement text, realizing intelligent automatic customer service transfer, which has the characteristic of lower cost compared with the manual transfer method, and can effectively improve the caller's usage experience compared with the caller's self-selection of the answering department according to the voice prompt.
[0082] In the third step, according to the voice statement audio and the statement text, a decoupled emotion calculation model is used to analyze the emotional state of the caller and predict the willing waiting time of the caller based on the emotional state.
[0083] In this embodiment, a decoupled emotion calculation model is used to parse the speech statement audio and the statement text to obtain the emotional state of the caller, and based on full consideration of the situation state of the caller, the willing waiting time of the caller is predicted. Specifically, the decoupled emotion calculation model is a calculation model that can recognize and understand human emotions, that is, it recognizes and understands the emotional state of the caller in terms of dimensional emotion and discrete emotion based on the caller's speech audio and the statement text converted from the speech audio, and predicts the willing waiting time by separately calculating the valence arousal vector based on dimensional emotion and the emotion label encoding based on discrete emotion. The decoupled emotion calculation model includes a dimensional emotion sub-neural network model, a discrete emotion sub-neural network model, and a waiting time prediction sub-neural network model. Among them, dimensional emotion originates from dimensional emotion theory. Dimensional emotion theory is a theory used to describe and explain emotions. It regards emotions as a psychological phenomenon that varies on multiple dimensions, which helps to more precisely understand the essence and differences of emotions. Among them, arousal and valence are two key dimensions. Arousal refers to the degree of physiological and psychological activation accompanied by emotions, and is used to reflect the physical and mental excitement level of an individual when experiencing a certain emotion. Valence refers to the degree to which the nature of the emotion is positive or negative, and mainly focuses on the position of the emotional experience on the pleasant and unpleasant dimensions. Discrete emotion originates from discrete emotion theory. Discrete emotion theory mainly emphasizes that emotions are composed of a series of independent and discrete emotion categories. It assumes that emotions can be clearly divided into different types, such as basic emotions: happiness, sadness, anger, fear, disgust, and surprise. The discrete emotion model has the advantages of simplicity and intuitiveness and has been widely used in the field of emotion calculation. The decoupled emotion calculation model of this embodiment uses the dimensional emotion sub-neural network model of the two dimensions of arousal and valence to describe the emotional state, representing the intensity and positive or negative nature of the emotion; at the same time, the emotional state is labeled as discrete adjective labels through the discrete emotion sub-neural network model to represent a limited number of single and clear emotion types; and then the waiting time prediction sub-neural network model makes a prediction according to the outputs of the dimensional emotion sub-neural network model and the discrete emotion sub-neural network model to obtain the willing waiting time.
[0084] In a specific embodiment, the decoupled emotion calculation model is deployed on a distributed cluster device. The distributed cluster device includes a first distributed device, a second distributed device, and a third distributed device that are independent and electrically connected. The decoupled emotion model includes a dimensional emotion sub-neural network model deployed on the first distributed device, a discrete emotion sub-neural network model deployed on the second distributed device, and a waiting time prediction sub-neural network model deployed on the third distributed device;
[0085] Said analyzing the emotional state of the caller according to the voice statement audio and statement text using a decoupled emotion calculation model, and predicting the willing waiting time of the caller based on the emotional state further includes:
[0086] According to the statement text, using a dimensional emotion sub-neural network model to extract the first emotional feature of the caller and output a valence-arousal vector;
[0087] According to the voice statement audio and statement text, using the discrete emotion sub-neural network model to extract the second emotional feature and the third emotional feature of the caller respectively and output an emotion label encoding;
[0088] According to the valence-arousal vector and the emotion label encoding, using the waiting time prediction sub-neural network model to obtain the willing waiting time of the caller.
[0089] In this embodiment, a decoupled emotion calculation model is used to predict the waiting time by calculating emotion labels and valence-arousal vectors. Compared with the voice emotion detection model in the related art, the decoupled emotion calculation neural network model in this embodiment is decoupled into three sub-neural network models according to the ideas of the discrete emotion theory and the dimensional emotion theory, namely, a dimensional emotion sub-neural network model, a discrete emotion sub-neural network model, and a waiting time prediction sub-neural network model; and in actual application, the three sub-neural network models are distributedly deployed on three independent and electrically connected devices. Specifically, the discrete emotion sub-neural network model is deployed on the first distributed device, the dimensional emotion sub-neural network model is deployed on the second distributed device, and the waiting time prediction sub-neural network model is deployed on the second distributed device. Among them, the discrete emotion sub-neural network model and the dimensional emotion sub-neural network model operate in parallel to save processing time, and the waiting time prediction sub-neural network model receives the valence-arousal vector output by the dimensional emotion sub-neural network model and the emotion label encoding output by the discrete emotion sub-neural network model for prediction, effectively saving the communication cost and computing power cost of distributed deployment, that is, this embodiment realizes multi-modal, highly interpretable, and highly real-time voice emotion detection while taking into account the discrete emotion theory and the dimensional emotion theory, and obtains the willing waiting time of the caller accordingly.
[0090] In a specific embodiment, such as Figure 2As shown, the dimensional sentiment sub-neural network model includes an embedding layer, a first bidirectional LSTM layer, a second bidirectional LSTM layer, a first Highway layer, and a first fully connected layer connected in sequence. Further, according to the statement text, using the dimensional sentiment sub-neural network model to extract the first sentiment feature of the caller and output a valence arousal vector includes: encoding the statement text to generate a statement text encoding, and inputting the statement text encoding into the dimensional sentiment sub-neural network model to obtain the first sentiment feature of the caller and output the valence arousal vector.
[0091] In this embodiment, the tokenizer of Keras is used to encode the statement text to obtain a statement text encoding. The statement text encoding is used as the input of the dimensional sentiment sub-neural network model. Through the dimensional sentiment sub-neural network model, the dimensional sentiment features of the caller including activation degree and valence degree are obtained, and a valence arousal vector representing the dimensional sentiment features is output. In this embodiment, by independently deploying the dimensional sentiment sub-neural network model on the first distributed device, the computing power cost is effectively saved.
[0092] In a specific embodiment, as Figure 3 shown, the discrete sentiment sub-neural network model includes a first branch, a second branch, and a third branch. The first branch includes a voice excitement index pooling layer and a one-dimensional convolutional layer connected in sequence. The second branch includes a separable two-dimensional convolutional layer and a Flatten layer connected in sequence. The third branch includes a first splicing layer, a second Highway layer, and a second fully connected layer. Further, according to the voice statement audio and the statement text, using the discrete sentiment sub-neural network model to extract the second sentiment feature and the third sentiment feature of the caller respectively and output an emotion label encoding includes:
[0093] Generating a Mel spectrogram and a sound spectrogram according to the voice statement audio, superimposing the Mel spectrogram and the sound spectrogram to generate a multi-channel audio feature map, using the voice statement audio as the input of the first branch and outputting the second sentiment feature, using the multi-channel audio feature map as the input of the second branch and outputting the third sentiment feature, and using the second sentiment feature and the third sentiment feature as the input of the third branch and outputting the emotion label encoding.
[0094] In this embodiment, the voice excitement index pooling layer of the first branch pools the voice excitement index according to the input voice statement audio. Specifically, the voice excitement index is:
[0095]
[0096]
[0097] Among them, E is the voice excitation index, c is the maximum length of the pooling window, n is the length of the pooling window, and E n is the voice excitation index component with the pooling window length of n, P is the skewness within the pooling window, K is the kurtosis within the pooling window, is the mean within the pooling window, i is the serial number of the value within the pooling window, and x i is the i-th value within the pooling window, and s is the standard deviation within the pooling window. In this embodiment, the maximum length of the pooling window is 7, and the voice excitation index pooling layer is implemented using layers.core.Lambda of Keras. Connect the voice excitation index pooling layer to a one-dimensional convolutional layer and output the second emotional feature characterizing the emotional discrete characteristics. The voice excitation index pooling layer in this embodiment is used to capture the relationship between features such as the average pitch, skewness, and kurtosis in the audio and the intensity of the emotion revealed by the caller when simply describing the consultation problem, and define the voice excitation index as an index for pooling. Since the pooling layer does not contain any learnable parameters, the computing power resources saved by this pooling operation are still greater than the computing power resources it consumes, realizing multi-feature and multi-scale feature extraction while shrinking the input information, extracting both overall and local features at the same time, effectively improving the accuracy of the model, and saving computing power resources.
[0098] At the same time, the separable two-dimensional convolutional layer of the second branch outputs the third emotional feature characterizing the emotional discrete characteristics according to the input multi-channel audio feature map. Specifically, in this embodiment, a Mel spectrogram and a sound spectrogram are generated from the speech statement audio, and the Mel spectrogram and the sound spectrogram are superimposed to generate a multi-channel audio feature map, that is, the physical characteristics of the sound and the non-linear perception of the sound frequency by the human ear are superimposed through the constructed multi-channel audio feature map, so as to construct a targeted feature engineering. Then, the separable two-dimensional convolutional layer performs depthwise convolution on the information of different channels and then pointwise convolution. Compared with directly recognizing the audio, the separable two-dimensional convolutional layer of the second branch in this embodiment can capture more detailed and diverse local features, and then perform flattening processing through the Flatten layer, that is, convert the multi-dimensional data output by the separable two-dimensional convolutional layer into one-dimensional data for subsequent data processing by the third branch.
[0099] The first branch and the second branch of this embodiment perform parallel operations, and respectively input the output second emotional feature and third emotional feature into the third branch. The first splicing layer of the third branch splices the second emotional feature and the third emotional feature, and then outputs the emotional encoding characterizing the emotional discrete characteristics after passing through the second Highway layer and the second fully connected layer.
[0100] In this embodiment, by independently deploying the discrete emotion sub-neural network model on the second distributed device, the computing power expenditure is effectively saved.
[0101] In a specific embodiment, Figure 4 As shown, the waiting time prediction sub-neural network model includes a second splicing layer, a third fully connected layer and a fourth fully connected layer connected in sequence, and the method of obtaining the caller's willingness to wait time using the waiting time prediction sub-neural network model according to the valence arousal vector and the emotion label encoding further includes:
[0102] The valence arousal vector and the emotion label encoding are mapped into the corresponding willingness to wait time of the caller.
[0103] In this embodiment, the dimensional emotion sub-neural network model and the discrete emotion sub-neural network model are respectively connected to the waiting time prediction sub-neural network model. The waiting time prediction sub-neural network model is independently deployed on a third distributed device. The valence arousal vector and the emotion label encoding output by the dimensional emotion sub-neural network model are mapped into the corresponding willingness to wait time of the caller. Since the waiting time prediction sub-neural network model is independently deployed on the third distributed device, the communication expenses and computing power expenses of the distributed deployment are effectively saved. This embodiment realizes multimodal, highly interpretable and highly real-time voice emotion detection while taking into account discrete emotion theory and dimensional emotion theory, and obtains the caller's willingness to wait time accordingly.
[0104] The fourth step is to queue the callers for telephone answering in the answering department according to the willing waiting time.
[0105] In this embodiment, based on the voice statement audio of the caller briefly describing the consultation problem, on the one hand, the voice statement audio is converted into statement text, and the Bayesian model is used to classify the statement text to determine the answering department for the caller's consultation. On the other hand, the emotional state of the caller is analyzed using a decoupled emotional computing model based on the voice statement audio, and the caller's willingness to wait time is predicted based on the emotional state, so that in the queue corresponding to the answering department of the telephone customer service system, the caller queues according to the obtained willingness to wait time, effectively improving the efficiency of telephone consultation and improving the user experience of telephone consultation callers. That is, the caller's willingness to wait time predicted by this embodiment based on the emotional state of the telephone consultation caller can more effectively capture the caller's emotions, explore the caller's needs and pain points, and thus improve products and services.
[0106] In order to further improve the answering efficiency, in an optional embodiment, queuing the callers for the telephone answering of the answering department according to the willing waiting time further comprises:
[0107] Calculate the queuing priority of the caller according to the willing waiting time and the consumption record of the caller;
[0108] Queue the call answering of the caller in the answering department according to the willing waiting time and the queuing priority.
[0109] In this embodiment, by further considering the historical consumption situation of the callers of telephone consultations, generate the queuing priority of the callers considering both the willing waiting time and the historical consumption, and queue in the telephone customer service system according to the willing waiting time and the queuing priority.
[0110] In a specific embodiment, the calculating the queuing priority of the caller according to the willing waiting time and the consumption record of the caller further includes:
[0111] The queuing priority of the caller is:
[0112] R = N / T;
[0113] Q = S * R;
[0114] Wherein, R represents the risk of losing customers, T represents the willing waiting time, N represents the number of callers currently queuing, S represents the consumption record of the caller, and Q is the queuing priority.
[0115] In this embodiment, first calculate the risk of losing customers according to the willing waiting time of the caller and the number of callers currently queuing, and then calculate the queuing priority according to the historical consumption amount of the caller and the risk of losing customers, so as to comprehensively consider the importance of the caller and the relative tolerance for queuing to form the queuing priority of the caller, and queue according to the queuing priority and the willing waiting time, which can effectively improve the comprehensive satisfaction of the caller, that is, improve the user experience.
[0116] In a specific example, for example, the willing waiting time of the caller is 50 seconds, the number of callers currently queuing is 10, and the historical consumption amount of the caller is 1000 yuan. It is calculated that the risk of losing customers of this user is 0.2, and further calculate the queuing priority of this caller as 200 according to the historical consumption amount. Then queue according to the willing waiting time and the queuing priority of the caller. In this embodiment, by associating the queuing priority of the caller with the historical consumption record, further distinguish high-value callers, ensure that high-value callers can obtain a higher priority and get a response from the customer service within the willing waiting time.
[0117] This embodiment classifies and processes the specific queuing situation:
[0118] Classification situation 1: The queuing priority has arrived but the willing waiting time has not arrived. Arrange customer service staff to answer the caller's consultation according to the queuing priority and answer in time, so as to ensure that the consultations of high-value callers can be processed in time and avoid missing business opportunities;
[0119] Classification scenario 2: When willing to wait until the time arrives but the queuing priority has not arrived and there is no customer service staff available to answer the call, the queuing priority is dynamically adjusted. For example, the queuing priority of the caller is increased to ensure that the customer service staff answers the caller's consultation as soon as possible and responds in a timely manner, so as to ensure that the callers within the willing waiting time will not be ignored due to priority issues.
[0120] It should be noted that the present application does not specifically limit the queuing processing method. Those skilled in the art can select an appropriate queuing processing method according to the above customer service response queuing method according to actual application requirements, with the design criterion of improving the efficiency of telephone consultation and improving the user experience of telephone consultation, which will not be elaborated here.
[0121] So far, the specific process of queuing and answering calls from callers who call the customer service for consultation is completed. By converting the voice statement audio of the caller in the telephone consultation into a statement text, classifying it through a Bayesian model according to the statement text to obtain the answering department, analyzing the emotional state of the caller through a decoupled emotion calculation model based on the voice statement audio and the statement text, predicting the willing waiting time of the caller and the queuing priority of the caller, and performing effective queuing according to the willing waiting time and the queuing priority. That is, the embodiment provided by the present application effectively improves the efficiency of telephone consultation and improves the user experience of callers in telephone consultation by identifying the emotional state of the caller during the telephone consultation statement and calculating the willing waiting time of the caller based on the emotional state, and has a wide range of application prospects.
[0122] Considering that the Bayesian model needs to be pre-trained before it can be used, in an optional embodiment, before classifying according to the statement text using the Bayesian model to obtain the answering department, the customer service response queuing method further includes: establishing a first training data set and training the Bayesian model according to the first training data set.
[0123] In this embodiment, the Bayesian model is trained by presetting a training data set for training the Bayesian model, so as to facilitate obtaining the answering department according to the statement text converted from the voice statement audio of the caller.
[0124] In a specific embodiment, the establishing a first training data set and training the Bayesian model according to the first training data set further includes;
[0125] Extract all acceptance sample records from the customer service call history;
[0126] Extract an accepted sample record and collect the record text and the responding department therein, convert the record text into a text vector, generate a department code based on the responding department, and a text vector of each accepted sample record and the corresponding department code form a training data in the first training dataset;
[0127] Determine whether there are still uncollected accepted sample records. If so, jump to the step of extracting an accepted sample record and collecting the record text and the responding department therein.
[0128] In this embodiment, MultinomialNB is used to construct a multinomial naive Bayes model as the Bayes model of this embodiment. In the process of constructing the training dataset, taking the customer service call history record as the data source, all accepted sample records are extracted, that is, all records of simple descriptions and responding department classifications during the caller's phone consultation. Specifically, traverse all accepted sample records, extract the statement text and the responding department from each accepted sample record; convert the statement text into a TF-IDF matrix, independently encode the responding department to generate a department code, and a TF-IDF matrix of an accepted sample record and the department code form a training data. In this embodiment, TfidfTransformer of Scikit-learn is used to convert the text into a TF-IDF matrix, and OneHotEncode of Scikit-learn is used to independently encode the responding department to generate a department code; data processing is sequentially performed on all accepted sample records to form a first training dataset for training the Bayes model. Finally, the first training dataset is divided into a Bayes training set and a Bayes test set in a ratio of 8:2, specifying the TF-IDF matrix in the first training dataset as the input and the department code as the target output to train the Bayes model.
[0129] Considering that the decoupled emotion calculation model also needs to be pre-trained before use, in an optional embodiment, before analyzing the caller's emotional state using the decoupled emotion calculation model based on the voice statement audio and the statement text and predicting the caller's willing waiting time based on the emotional state, the customer service response queuing method further includes: establishing a second training dataset and training the decoupled emotion calculation model according to the second training dataset.
[0130] In this embodiment, the decoupled emotion calculation model is trained by presetting a training dataset for training the decoupled emotion calculation model, so as to obtain the caller's willing waiting time according to the caller's voice statement audio.
[0131] In a specific embodiment, the step of establishing a second training dataset and training the decoupled emotion calculation model according to the second training dataset further includes:
[0132] Extract all the received unanswered sample records from the customer service call history;
[0133] Extract one received unanswered sample record and collect the waiting time and the voice statement audio therein. Convert the voice statement audio into a statement text, generate a Mel spectrogram and a sound spectrogram according to the voice statement audio, and superimpose the Mel spectrogram and the sound spectrogram to generate a multi-channel audio feature map. The waiting time, the voice statement audio, the statement text, and the multi-channel audio feature map of each received unanswered sample record form a training data in the second training dataset;
[0134] Determine whether there are still received unanswered sample records that have not been collected. If so, jump to the step of extracting one received unanswered sample record and collecting the waiting time and the voice statement audio therein.
[0135] In this embodiment, taking the customer service call history as the data source, extract the unanswered sample records in the received sample records, that is, the records where the caller calls to simply describe the consultation problem but hangs up the phone before the queue ends, that is, the situation of making a call, simply describing but not waiting for a response. Specifically, traverse all the received unanswered sample records, extract the waiting time and the voice statement audio from each received unanswered sample record, convert the voice statement audio into a statement text, and at the same time process the voice statement audio to generate a Mel spectrogram and a sound spectrogram. Superimpose the Mel spectrogram and the sound spectrogram to generate a multi-channel audio feature map. The waiting time, the voice statement audio, the statement text, and the multi-channel audio feature map in one received unanswered sample record form a training data, and sequentially perform data processing on all the received unanswered sample records to form a second training dataset for training the decoupled emotion calculation model.
[0136] The decoupled emotion calculation model includes a dimensional emotion sub-neural network model deployed on the first distributed device, a discrete emotion sub-neural network model deployed on the second distributed device, and a waiting time prediction sub-neural network model deployed on the third distributed device.
[0137] Specifically, the training of the dimensional emotion sub-neural network model includes: importing Chinese EmoBank, encoding the texts in Chinese EmoBank, and then dividing Chinese EmoBank into a Chinese EmoBank training set and a Chinese EmoBank test set at a ratio of 8:2. After encoding the texts in Chinese EmoBank, they are specified as the input, and the valence arousal vector in Chinese EmoBank is the target output. The dimensional emotion sub-neural network model is trained on the first distributed device. Among them, the loss function of the first fully connected layer is MSE, and the optimizer is adam. In this embodiment, the tokenizer of Keras is used to encode the texts in Chinese EmoBank.
[0138] The training of the discrete emotion sub-neural network model includes: importing the CASIA Chinese Emotion Audio Corpus, extracting the audio in the CASIA Chinese Emotion Audio Corpus to generate corresponding multi-channel audio feature maps and supplementing them into the CASIA Chinese Emotion Audio Corpus, encoding the emotion labels in the CASIA Chinese Emotion Audio Corpus to generate emotion label encodings, dividing the CASIA Chinese Emotion Audio Corpus into a CASIA training set and a CASIA test set at a ratio of 8:2, specifying the audio in the CASIA Chinese Emotion Audio Corpus as the input of the speech arousal index pooling layer, and the corresponding multi-channel audio feature maps of the audio in the CASIA Chinese Emotion Audio Corpus as the input of the separable two-dimensional convolutional layer, specifying the emotion label encoding as the target output, and training the discrete emotion sub-neural network model on the second distributed device. Among them, the loss function of the second fully connected layer is categorical_crossentropy, and the optimizer is adam. In this embodiment, OneHotEncode of Scikit-learn is used to encode the emotion labels in the CASIA Chinese Emotion Audio Corpus to generate emotion label encodings.
[0139] The training of the waiting time prediction sub-neural network model includes: using the dimensional emotion sub-neural network model and the discrete emotion sub-neural network model to process the second training data set, generating valence arousal vectors and emotion label encodings for each training data in the second training data set and supplementing them into the second training data set, specifying the valence arousal vectors and emotion label encodings in the second training data set as the input, and the waiting time as the target output, and training the waiting time prediction sub-neural network model on the third distributed device. Among them, the loss function of the fourth fully connected layer is MSE, and the optimizer is adam.
[0140] Based on the customer service response queuing method of the above embodiment, another embodiment of the present invention further provides a customer service response queuing device applying the above customer service response queuing method, such asFigure 5 As shown, it includes a distributed cluster device and a controller. The distributed cluster device includes a first distributed device, a second distributed device, and a third distributed device that are independent and electrically connected. The decoupled emotion model includes a dimensional emotion sub-neural network model, a discrete emotion sub-neural network model, and a waiting time prediction sub-neural network model. The first distributed device is used to deploy the dimensional emotion sub-neural network model, the second distributed device is used to deploy the discrete emotion sub-neural network model, and the third distributed device is used to deploy the waiting time prediction sub-neural network model. The controller is configured to:
[0141] Receive the voice statement audio of the caller making a call to the customer service, and convert the voice statement audio into a statement text;
[0142] According to the statement text, use the Bayesian model for classification to obtain the responding department;
[0143] According to the voice statement audio and the statement text, use the decoupled emotion calculation model to analyze the emotional state of the caller, and predict the willing waiting time of the caller based on the emotional state;
[0144] Queue the call answering of the caller in the responding department according to the willing waiting time.
[0145] As Figure 5 shown, the embodiments are only for facilitating the description of the specific implementation manners of the present application. Those skilled in the art should understand that the controller can be set independently or in one of the distributed cluster devices, and the design criterion is to meet the actual application requirements, which will not be elaborated here.
[0146] In this embodiment, by converting the voice statement audio of the caller in the telephone consultation into a statement text, classifying according to the statement text through the Bayesian model to obtain the responding department, analyzing the emotional state of the caller according to the voice statement audio and the statement text through the decoupled emotion calculation model, predicting the willing waiting time of the caller based on the emotional state, and performing effective queuing according to the willing waiting time. That is, the embodiments provided by the present application identify the emotional state of the caller during the telephone consultation statement, and calculate the willing waiting time of the caller based on the emotional state, thereby making up for the problems existing in the prior art, effectively improving the efficiency of telephone consultation, and improving the user experience of the callers in telephone consultation. For the specific implementation manners of this embodiment, refer to the foregoing embodiments, which will not be elaborated here.
[0147] It should be noted that the location for deploying the Bayesian model in this embodiment is not limited and can be any one of the distributed devices in the distributed cluster device. Those skilled in the art should select an appropriate distributed device according to actual application requirements, with the design criteria of improving the efficiency of telephone consultations and enhancing the user experience of telephone consultations, which will not be elaborated herein.
[0148] Another embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it realizes: receiving the voice statement audio of a caller making a call to the customer service, converting the voice statement audio into a statement text; classifying using the Bayesian model according to the statement text to obtain the answering department; analyzing the emotional state of the caller using the decoupled emotion calculation model based on the voice statement audio and the statement text, and predicting the willing waiting time of the caller based on the emotional state; queuing the call answering of the caller in the answering department according to the willing waiting time.
[0149] In practical applications, the computer-readable storage medium can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device.
[0150] The computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device.
[0151] The program code contained on a computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wire, optical fiber cable, RF, etc., or any suitable combination of the above.
[0152] The computer program code for performing the operations of the present invention can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0153] As Figure 6 shown, a schematic structural diagram of a computer device provided by another embodiment of the present invention. Figure 6 The displayed computer device T12 is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present invention.
[0154] As Figure 6 shown, the computer device T12 is presented in the form of a general-purpose computing device. The components of the computer device T12 may include but are not limited to: one or more processors or processing units T16, a system memory T28, and a bus T18 connecting different system components (including the system memory T28 and the processing unit T16).
[0155] The bus T18 represents one or more of several types of bus structures, including a memory bus or a memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any bus structure in a variety of bus structures. For example, these architectures include but are not limited to Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MAC) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.
[0156] The computer device T12 typically includes a variety of computer system-readable media. These media can be any available media that can be accessed by the computer device T12, including volatile and non-volatile media, removable and non-removable media.
[0157] System memory T28 can include computer system readable media in the form of volatile memory, such as random access memory (RAM) T30 and / or cache memory T32. The computer device T12 can further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, a storage system T34 can be used for reading and writing on non-removable, non-volatile magnetic media ( Figure 6 not shown, commonly referred to as a "hard disk drive"). Although Figure 6 not shown in, a disk drive for reading and writing on removable non-volatile disks (such as "floppy disks") and an optical disk drive for reading and writing on removable non-volatile optical disks (such as CD-ROM, DVD-ROM or other optical media) can be provided. In these cases, each drive can be connected to the bus T18 through one or more data media interfaces. The memory T28 can include at least one program product having a set (such as at least one) of program modules configured to perform the functions of the embodiments of the present invention.
[0158] A program / utility T40 having a set (at least one) of program modules T42 can be stored, for example, in the memory T28. Such program modules T42 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment. The program modules T42 generally perform the functions and / or methods in the embodiments described in the present invention.
[0159] The computer device T12 can also communicate with one or more external devices T14 (such as a keyboard, a pointing device, a display T24, etc.), and can also communicate with one or more devices that enable a user to interact with the computer device T12, and / or communicate with any device that enables the computer device T12 to communicate with one or more other computing devices (such as a network card, a modem, etc.). Such communication can be carried out through an input / output (I / O) interface T22. And, the computer device T12 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN) and / or a public network, such as the Internet) through a network adapter T20. As Figure 6 shown, the network adapter T20 communicates with other modules of the computer device T12 through the bus T18. It should be understood that although Figure 6 not shown in, other hardware and / or software modules can be used in combination with the computer device T12, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0160] The processor unit T16 executes various functional applications and data processing by running the programs stored in the system memory T28, for example, implementing a customer service response queuing method based on emotion recognition provided by the embodiments of the present invention.
[0161] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is impossible to enumerate all the implementation manners here. Any obvious changes or modifications derived from the technical solutions of the present invention still fall within the protection scope of the present invention.
Claims
1. A customer service response queuing method based on emotion recognition, characterized in that, Including: Receiving the audio of the voice statement of the caller making a call to the customer service, and converting the audio of the voice statement into a statement text; Classifying according to the statement text using a Bayesian model to obtain a response department; Analyzing the emotional state of the caller using a decoupled emotion calculation model based on the audio of the voice statement and the statement text, and predicting the willing waiting time based on the emotional state; Queuing the call response of the caller in the response department according to the willing waiting time.
2. The customer service response queuing method according to claim 1, wherein The queuing of the call response of the caller in the response department according to the willing waiting time further includes: Calculating the queuing priority of the caller according to the willing waiting time and the consumption record of the caller; Queuing the call response of the caller in the response department according to the willing waiting time and the queuing priority.
3. The customer service response queuing method according to claim 2, wherein The calculating the queuing priority of the caller according to the willing waiting time and the consumption record of the caller further includes: The queuing priority of the caller is: R = N / T; Q = S * R; wherein, R represents the risk of losing customers, T represents the willing waiting time, N represents the number of callers currently queuing, S represents the consumption record of the caller, and Q is the queuing priority.
4. The customer service response queuing method according to claim 2, wherein The decoupled emotion calculation model is deployed on a distributed cluster device, and the distributed cluster device includes an independent and electrically connected first distributed device, a second distributed device, and a third distributed device. The decoupled emotion model includes a dimensional emotion sub-neural network model deployed on the first distributed device, a discrete emotion sub-neural network model deployed on the second distributed device, and a waiting time prediction sub-neural network model deployed on the third distributed device; The analyzing the emotional state of the caller using a decoupled emotion calculation model based on the audio of the voice statement and the statement text, and predicting the willing waiting time based on the emotional state further includes: Extracting the first emotional feature of the caller using the dimensional emotion sub-neural network model according to the statement text and outputting a valence arousal vector; Extracting the second emotional feature and the third emotional feature of the caller respectively using the discrete emotion sub-neural network model according to the audio of the voice statement and outputting an emotion label encoding; Obtaining the willing waiting time using the waiting time prediction sub-neural network model according to the valence arousal vector and the emotion label encoding.
5. The customer service response queuing method according to claim 4, wherein The dimensional emotion sub-neural network model includes an embedding layer, a first bidirectional LSTM layer, a second bidirectional LSTM layer, a first Highway layer, and a first fully connected layer connected in sequence. The extracting the first emotional feature of the caller using the dimensional emotion sub-neural network model according to the statement text and outputting a valence arousal vector further includes: Encoding the statement text to generate a statement text encoding, inputting the statement text encoding into the dimensional emotion sub-neural network model to obtain the first emotional feature of the caller, and outputting the valence arousal vector; The discrete emotion sub-neural network model includes a first branch, a second branch, and a third branch. The first branch includes a speech excitement index pooling layer and a one-dimensional convolutional layer connected in sequence. The second branch includes a separable two-dimensional convolutional layer and a Flatten layer connected in sequence. The third branch includes a first splicing layer, a second Highway layer, and a second fully connected layer. The step of using the discrete emotion sub-neural network model to extract the second emotion feature and the third emotion feature of the caller from the speech statement audio and output an emotion label encoding further includes: Generating a Mel spectrogram and a sound spectrogram from the speech statement audio, superimposing the Mel spectrogram and the sound spectrogram to generate a multi-channel audio feature map, using the speech statement audio as the input of the first branch and outputting the second emotion feature, using the multi-channel audio feature map as the input of the second branch and outputting the third emotion feature, and using the second emotion feature and the third emotion feature as the input of the third branch and outputting the emotion label encoding; The waiting time prediction sub-neural network model includes a second splicing layer, a third fully connected layer, and a fourth fully connected layer connected in sequence. The step of using the waiting time prediction sub-neural network model to obtain the willing waiting time based on the valence arousal vector and the emotion label encoding further includes: Mapping the valence arousal vector and the emotion label encoding to the willing waiting time corresponding to the caller.
6. The customer service response queuing method according to claim 5, wherein, The speech excitement index pooling layer is used to calculate the speech excitement index and perform pooling. The speech excitement index is: Among them, E is the voice excitation index, c is the maximum length of the pooling window, n is the length of the pooling window, and E n is the voice excitation index component with the length of the pooling window being n, P is the skewness within the pooling window, K is the kurtosis within the pooling window, is the mean within the pooling window, i is the serial number of the values within the pooling window, and x i is the i-th value within the pooling window, and s is the standard deviation within the pooling window.
7. The customer service response queuing method according to claim 5, wherein Before using the Bayesian model to classify according to the statement text to obtain the response department, the customer service response queuing method further includes: Establishing a first training data set and training the Bayesian model according to the first training data set; Before using the decoupled emotion calculation model to analyze the emotion state of the caller based on the speech statement audio and the statement text and predicting the willing waiting time based on the emotion state, the customer service response queuing method further includes: Establishing a second training data set and training the decoupled emotion calculation model according to the second training data set.
8. The customer service response queuing method according to claim 7, wherein The step of establishing the first training data set and training the Bayesian model according to the first training data set further includes: Extracting all acceptance sample records from the customer service call history; Extracting one acceptance sample record and collecting the record text and the response department therein, converting the record text into a text vector, generating a department encoding based on the response department, and the text vector and the corresponding department encoding of each acceptance sample record form a training data in the first training data set; Judging whether there are still uncollected acceptance sample records. If so, jump to the step of extracting one acceptance sample record and collecting the record text and the response department therein; The establishment of the second training dataset and the further training of the decoupled sentiment calculation model based on the second training dataset further include: Extract all the unresponded acceptance sample records from the customer service call history; Extract one unresponded acceptance sample record and collect the waiting time and the voice statement audio therein. Convert the voice statement audio into a statement text, generate a Mel spectrogram and a sound spectrogram according to the voice statement audio, and superimpose the Mel spectrogram and the sound spectrogram to generate a multi-channel audio feature map. The waiting time, the voice statement audio, the statement text, and the multi-channel audio feature map of each unresponded acceptance sample record form a training data in the second training dataset; Determine whether there are still uncollected unresponded acceptance sample records. If so, jump to the step of extracting one unresponded acceptance sample record and collecting the waiting time and the voice statement audio therein.
9. A customer service response queuing device applying the customer service response queuing method according to any one of claims 1-8, characterized in that, It includes a distributed cluster device and a controller. The distributed cluster device includes a first distributed device, a second distributed device, and a third distributed device that are independent and electrically connected. The decoupled sentiment model includes a dimensional sentiment sub-neural network model, a discrete sentiment sub-neural network model, and a waiting time prediction sub-neural network model. The first distributed device is used to deploy the dimensional sentiment sub-neural network model, the second distributed device is used to deploy the discrete sentiment sub-neural network model, and the third distributed device is used to deploy the waiting time prediction sub-neural network model; The controller is configured to: Receive the voice statement audio of the caller making a customer service call and convert the voice statement audio into a statement text; Classify according to the statement text using a Bayesian model to obtain the responding department; Analyze the emotional state of the caller using the decoupled sentiment calculation model based on the voice statement audio and the statement text, and predict the willing waiting time of the caller based on the emotional state; Queue the call answering of the caller in the responding department according to the willing waiting time.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the method according to any one of claims 1-8.
11. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method according to any one of claims 1-8.