User Emotion Analysis Model Training Method, Device, Electronic Device, and Storage Medium
By obtaining training corpus from historical corpus and performing vocalprint clustering, and using speech psychology and convolutional neural network models to train user sentiment analysis models, the problem of insufficient accuracy and flexibility of emotion recognition models in the existing technology is solved, and higher prediction accuracy and adaptability are achieved.
Patent Information
- Application Number
- CN202210312084.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-28
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-03-28
AI Technical Summary
The existing emotion recognition model has poor accuracy and low flexibility when predicting user emotional states, and cannot adapt to user needs in different business scenarios.
By obtaining the training corpus from the preset historical corpus, determining the user's intentions and clustering them according to the voiceprints, dividing them into the same training set, using a convolutional neural network model based on speech psychology for training, extracting feature vectors and performing model training, adapting to sentiment analysis under different user's intentions.
It improves the prediction accuracy and flexibility of the model in different business scenarios, can adapt to emotional changes of different users, reduce negative emotions when users communicate with customer service personnel, and improve user experience.
Smart Images

Figure CN114925159B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technologies, and in particular, to a method, apparatus, electronic device, and storage medium for training a user sentiment analysis model. Background Art
[0002] With the development of communication technologies, users can seek help by contacting customer service online. Currently, after a call center customer service system assigns a customer service agent to serve an incoming user, during the process of serving the user, the customer service agent mainly relies on experience to manually judge the user's mental state, and based on their own experience, combined with the conversation scripts from the onboarding training, completes the user call service.
[0003] In the prior art, it is possible to collect the user's voice information and input it into an emotion recognition model to predict the user's emotional state, find the corresponding conversation script based on the predicted emotional state, and recommend the conversation script to the customer service agent who is having a conversation with the user, reducing the negative emotions during the communication between the user and the customer service agent. Among them, the emotion recognition model is a model trained based on mutually related semantic information samples, voice information samples, and emotion samples.
[0004] However, the above-mentioned emotion recognition model still has problems of poor prediction accuracy and low flexibility. Summary of the Invention
[0005] This application provides a method, apparatus, electronic device, and storage medium for training a user sentiment analysis model, which can train the model based on training sets of different users under different user intents. The input data is targeted, improving the prediction accuracy and the flexibility of model use.
[0006] In a first aspect, this application provides a method for training a user sentiment analysis model, the method including:
[0007] Obtain a plurality of training corpora from a preset historical corpus, and determine the user intent corresponding to each training corpus;
[0008] Perform clustering operations on the training corpora of each user intent according to voiceprints, to obtain multiple user voiceprint types;
[0009] Divide the training corpora with the same intent and the same user voiceprint type into the same training set, and train the corresponding user sentiment analysis model according to the training set.
[0010] Optionally, dividing the training corpora with the same intent and the same user voiceprint type into the same training set includes:
[0011] Extract features from the training corpora with the same intent and the same user voiceprint type, to obtain feature vectors of at least one session type;
[0012] Summarize the feature vectors of the at least one session type into the same training set.
[0013] Optionally, perform feature extraction on the training corpus with the same intent and the same user voiceprint type to obtain feature vectors of at least one session type, including:
[0014] Perform feature extraction on the training corpus with the same intent and the same user voiceprint type based on a neural network model to obtain the following feature vectors of at least one session type: the feature vector of the dialogue rhythm type, the feature vector of the tone type, the feature vector of the intonation type, and the feature vector of the sensitive word type.
[0015] Optionally, train a corresponding user sentiment analysis model according to the training set, including:
[0016] Train a user sentiment analysis model according to the feature vectors in the training set and the labels corresponding to the feature vectors; wherein, the labels are negative emotion labels or positive emotion labels obtained by performing semantic recognition on the training corpus using a preset dictionary library;
[0017] The user sentiment analysis model is a convolutional neural network model established based on speech psychology.
[0018] Optionally, the labels include emotion labels of multiple levels. Train a user sentiment analysis model according to the feature vectors in the training set and the labels corresponding to the feature vectors, including:
[0019] Find the emotion labels of the corresponding levels preset for each training set;
[0020] Train a user sentiment analysis model according to the feature vectors in each training set and the found emotion labels of the corresponding levels.
[0021] Optionally, the method further includes:
[0022] Obtain the personality characteristics of the user and the corresponding recommended conversation phrase library. The personality characteristics of the user are marked by feature vectors of at least one session type; the recommended conversation phrase library is used to recommend conversation maintenance phrases for the corresponding customer service staff to refer to during communication;
[0023] Divide the personality characteristics and the corresponding recommended conversation phrase library into a training sample data set and a test sample data set according to a random ratio, and input the training sample data set and the test sample data set into the user sentiment analysis model for model training and testing.
[0024] Optionally, the method further includes:
[0025] Obtain the voice information of the user, and parse the voice information into conversation corpus through natural language processing technology;
[0026] Extract the feature vectors of at least one session type from the session corpus;
[0027] Input the feature vectors of the at least one session type into the trained user sentiment analysis model, and determine whether to send a warning prompt message to the customer service based on the model output result.
[0028] Optionally, the method further includes:
[0029] If it is determined that a warning prompt message needs to be sent to the customer service, then look up the corresponding customer service emergency handling method in the lookup table based on the model output result; the customer service emergency handling method includes the adjustment methods for the speech rate, intonation, and conversation rhythm of the customer service staff, as well as the recommended retention scripts.
[0030] Optionally, the method further includes:
[0031] If it is determined not to send a warning prompt message to the customer service, then obtain the session corpus of the customer service staff to retrain the trained user sentiment analysis model based on the session corpus.
[0032] In a second aspect, the present application further provides a user sentiment analysis model training device, and the device includes:
[0033] An acquisition module, configured to acquire a plurality of training corpora from a preset historical corpus and determine the user intentions corresponding to each training corpus;
[0034] A clustering module, configured to cluster the training corpora of each user intention according to the voiceprint to obtain a plurality of user voiceprint types;
[0035] A training module, configured to divide the training corpora with the same intention and the same user voiceprint type into the same training set, and train the corresponding user sentiment analysis model according to the training set.
[0036] In a third aspect, the present application further provides an electronic device, including: a processor, a memory, and a computer program; wherein, the computer program is stored in the memory and is configured to be executed by the processor, and the computer program includes instructions for executing the user sentiment analysis model training method according to any one of the first aspect.
[0037] In a fourth aspect, the present application further provides a computer-readable storage medium, and the computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, they are used to implement the user sentiment analysis model training method according to any one of the first aspect.
[0038] In summary, the present application provides a method, an apparatus, an electronic device, and a storage medium for training a user sentiment analysis model. The method can obtain multiple training corpora from a preset historical corpus, determine the user intents corresponding to each training corpus, and further cluster the training corpora of each user intent according to voiceprints to obtain multiple user voiceprint types. Then, the training corpora with the same intent and the same user voiceprint type are divided into the same training set. Furthermore, the corresponding user sentiment analysis model can be trained based on the training set. In this way, the model is trained based on the training sets of different users under different user intents. Since the input data is targeted, the accuracy of model prediction can be improved, and the trained model can adapt to different users in different business scenarios, improving the flexibility of model application. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0040] Figure 1 FIG. is a schematic diagram of an application scenario provided by an embodiment of the present application;
[0041] Figure 2 FIG. is a schematic flowchart of a method for training a user sentiment analysis model provided by an embodiment of the present application;
[0042] Figure 3 FIG. is a schematic flowchart of a specific method for training a user sentiment analysis model provided by an embodiment of the present application;
[0043] Figure 4 FIG. is a schematic flowchart of a process of working using a user sentiment analysis model provided by an embodiment of the present application;
[0044] Figure 5 FIG. is a schematic structural diagram of a device for training a user sentiment analysis model provided by an embodiment of the present application;
[0045] Figure 6 FIG. is a schematic structural diagram of an electronic device provided by an embodiment of the present application.
[0046] Through the above-mentioned accompanying drawings, the specific embodiments of the present application have been shown, and there will be more detailed descriptions hereinafter. These drawings and the textual descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0048] For the convenience of clearly describing the technical solutions of the embodiments of the present application, in the embodiments of the present application, terms such as "first" and "second" are used to distinguish identical or similar items with basically the same functions and roles. For example, the first device and the second device are only used to distinguish different devices, and do not limit their order. Those skilled in the art can understand that terms such as "first" and "second" do not limit the quantity and execution order, and "first", "second", etc. do not necessarily mean different.
[0049] It should be noted that in the present application, words such as "exemplary" or "for example" are used to mean examples, illustrations, or explanations. Any embodiment or design described as "exemplary" or "for example" in the present application should not be construed as being more preferred or having more advantages than other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific manner.
[0050] In the present application, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one (item)" or its similar expression refers to any combination of these items, including any combination of single item (s) or plural item (s). For example, at least one (item) of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, c can be single or multiple.
[0051] In the communications field, the customer service industry has a high turnover of operators. Due to the lack of experience of new employees, simple pre-job training cannot guarantee that operators can judge the negative emotional changes of users in advance, which affects the user experience and may even cause user complaints. Therefore, there is an urgent need for intelligent auxiliary tools to help operators complete business services. Intelligent auxiliary tools can predict user emotions by inputting user conversation data into the user sentiment analysis model. When the user's emotions change, timely warnings are issued, and the model output is returned to help operators perform maintenance and rescue operations.
[0052] Among them, in order to obtain the user sentiment analysis model to process the user's conversation data in real time, it is necessary to train the user sentiment analysis model in advance so that the trained user sentiment analysis model can recognize and predict the results of user conversations in different business scenarios.
[0053] The embodiments of the present application are introduced below with reference to the accompanying drawings. Figure 1 This is a schematic diagram of an application scenario provided in an embodiment of the present application. The user sentiment analysis model training method provided in the present application can be applied to the following examples: Figure 1 In the application scenario shown. The application scenario includes: a historical corpus 101 and a server 102; the server 102 can obtain training corpus from the historical corpus 101 to train the model. Specifically, the server 102 can classify the training corpus according to the different business scenarios in the training corpus. For example, it can be divided into two categories based on business scenario 1 and business scenario 2. Further, for business scenario 1, clustering operations can be performed according to the voiceprints of different users to obtain the conversation corpus of user 1, the conversation corpus of user 2, and the conversation corpus of user 3; similarly, for business scenario 2, after clustering operations are performed according to the voiceprints of different users, the conversation corpus of user 1, the conversation corpus of user 4, and the conversation corpus of user 5 are obtained. Further, the above-mentioned corpus is respectively input into the user sentiment analysis model to train the model, so that the user sentiment analysis model can adapt to different users in different business scenarios when used, thereby improving the flexibility of model application and the accuracy of prediction.
[0054] It is understandable that the same user can consult customer service personnel regarding different businesses and obtain corresponding conversation data, such as the conversation data about user 1 in business scenario 1 and the conversation data about user 1 in business scenario 2 in the above scenario. The corresponding users and number in each business scenario are not specifically limited in the embodiments of the present application.
[0055] It should be noted that according to the classification of different business scenarios in the training corpus, the number and types of business scenarios are not specifically limited in the embodiments of the present application. The number of business scenarios in the above application scenarios and the number of users in each business scenario are only for illustrative purposes.
[0056] Optionally, the above application scenario can be a user emotion analysis model for user-customer service conversations trained to reduce user complaints in industries such as mobile communication, online shopping, and various commercial after-sales services.
[0057] In the prior art, the voice information of a user can be collected and input into an emotion recognition model to predict the emotion state of the user. Based on the predicted emotion state, the corresponding conversation words are found and recommended to the customer service staff who is having a conversation with the user, so as to reduce the negative emotions during the communication between the user and the customer service staff. Among them, the emotion recognition model is a model trained based on mutually related semantic information samples, sound information samples, and emotion samples.
[0058] However, different business scenarios may have different emotions. The above emotion recognition model uniformly inputs the samples into the model, resulting in poor accuracy of prediction using this emotion recognition model and a problem of low flexibility.
[0059] Therefore, the present application provides a method for training a user emotion analysis model. By processing historical normal conversation corpora and complaint conversation corpora and classifying them according to conversation intents, further, the corpora of user demands with the same user intent and the same voiceprint are clustered to obtain a training set, and the training set is input into a user emotion analysis model established based on speech psychology and convolutional neural network for model training. In this way, using the training sets of different users under different user intents to train the model, due to the pertinence of the input data, the accuracy of model prediction can be improved, and the trained model can adapt to different users in different business scenarios, improving the flexibility of model application.
[0060] The technical solution of the present application will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below with reference to the accompanying drawings.
[0061] Exemplarily, Figure 2 is a schematic flow chart of a method for training a user emotion analysis model provided by an embodiment of the present application; as Figure 2 shown, the method may include:
[0062] S201. Obtain a plurality of training corpora from a preset historical corpus and determine the user intent corresponding to each training corpus.
[0063] In the embodiments of the present application, the preset historical corpus includes historical complaint conversation information, historical normal conversation information, and psychological data. Among them, the historical complaint conversation information refers to the conversation information between the user and the customer service within a certain period, and this conversation information is mainly the conversation information of user complaints, and the corresponding user emotion is negative; the historical normal conversation information refers to the conversation information of normal conversations between the user and the customer service within a certain period, and the corresponding user emotion is positive; the psychological data may refer to the conversation quotations of the user and the customer service based on psychology, and each conversation quotation has an analysis of the corresponding user emotion.
[0064] In this step, the user intention may refer to the intention of the user to consult the customer service staff, that is, the corresponding different business scenarios. For example, the business scenarios of consulting and handling packages, conducting after-sales services, purchasing corresponding products, etc. The embodiments of the present application do not specifically limit the types and contents of the user intentions.
[0065] Exemplarily, in Figure 1 the application scenario, the server 102 can obtain multiple training corpora from the historical corpus 101. For example, the complaint recordings of the user and the customer service within a week, the normal recordings of the user and the customer service within a week, and the quotations of the user and the customer service obtained from the psychological database. Further, using natural speech processing technology to process the above-mentioned recording files and quotations, and determining that the user intentions corresponding to the above-mentioned recording files and quotations are business scenario 1 and business scenario 2.
[0066] S202. Cluster the training corpora of each user intention according to the voiceprint to obtain multiple user voiceprint types.
[0067] In this step, clustering the training corpora of each user intention according to the voiceprint means that for the training corpora of each user intention, classify them according to whether they belong to the voiceprint of the same user (that is, the sound wave spectrum of the speech information). Further, cluster the training corpora corresponding to different user voiceprints to obtain multiple user voiceprint types, and the voiceprint types corresponding to different users are different.
[0068] Exemplarily, in Figure 1 the application scenario, taking business scenario 1 as an example, the server 102 clusters the training corpora of business scenario 1 according to the voiceprint to obtain the conversation corpora of user 1, the conversation corpora of user 2, and the conversation corpora of user 3.
[0069] S203. Divide the training corpora with the same intention and the same user voiceprint type into the same training set, and train the corresponding user emotion analysis model according to the training set.
[0070] In the embodiments of the present application, the user emotion analysis model may refer to a convolutional neural network model established based on speech psychology. Among them, the convolutional neural network model is a type of feedforward neural network model that contains convolutional calculations and has a deep structure, and is one of the representative algorithms of deep learning. It can be used for speech synthesis and language modeling, and can well complete the input sentence without additional feature engineering requirements for data.
[0071] It should be noted that the user emotion analysis model may also be other deep learning models, or a combination of multiple models. For example, a hybrid model of a convolutional neural network and a Hidden Markov Model (HMM). The embodiments of the present application do not make specific limitations on this. The user emotion analysis model can process natural language to achieve the purpose of the present application.
[0072] Exemplarily, in Figure 1 the application scenario of, taking business scenario 1 as an example, the server 102 divides the conversation corpus of user 1 in business scenario 1 into a training sample, the conversation corpus of user 2 into a training sample, and the conversation corpus of user 3 into a training sample. Further, the above training samples are aggregated into the same training set, and the training set is input into the corresponding user emotion analysis model for model training.
[0073] It should be noted that each training sample in the training set corresponds to attribute information, and the attribute information is used to identify which type of training corpus the training sample belongs to, that is, which intention and user voiceprint type it belongs to. Further, each training sample and the corresponding attribute information are input into the model to train the user emotion analysis model in a classified manner. Correspondingly, when using the trained user emotion analysis model, it can be determined which intention and user voiceprint type the user's voice information belongs to, and the corresponding attribute information can be found, and then input into the user emotion analysis model to obtain the output result. Among them, the number of trained user emotion analysis models is one.
[0074] It can be understood that if the conversation corpus of the user's calls includes multiple times in the same intention and the same user voiceprint type, and the emotions corresponding to the conversation corpus of each user's call are different, then model training is performed for the conversation corpus of each user's call.
[0075] Optionally, training corpora with the same intention and the same user voiceprint type are divided into the same training set. Each training set contains multiple different training samples, that is, the training set corresponds to the same intention and the same user voiceprint type, but different emotions. Therefore, multiple corresponding user sentiment analysis models are trained for different training sets. Correspondingly, when using the trained user sentiment analysis models, the user's voice information can be judged to belong to which intention and user voiceprint type, and then input into the corresponding user sentiment analysis model to obtain the output result. Among them, the number of trained user sentiment analysis models is multiple.
[0076] Therefore, the present application provides a method for training a user sentiment analysis model. Multiple training corpora can be obtained from a preset historical corpus, and the user intentions corresponding to each training corpus can be determined. Further, the training corpora of each user intention are clustered according to the voiceprint to obtain multiple user voiceprint types, and the training corpora with the same intention and the same user voiceprint type are divided into the same training set. Then, the training set is input into the user sentiment analysis model for model training, that is, the model is trained based on the training sets of different users under different user intentions. Because the input data is targeted, the accuracy of model prediction can be improved, and the trained model can adapt to different users in different business scenarios, improving the flexibility of model application.
[0077] Optionally, dividing the training corpora with the same intention and the same user voiceprint type into the same training set includes:
[0078] Feature extraction is performed on the training corpora with the same intention and the same user voiceprint type to obtain feature vectors of at least one conversation type;
[0079] The feature vectors of the at least one conversation type are summarized into the same training set.
[0080] In this step, the conversation type may refer to the conversation type used to characterize the change in the user's emotion when the user communicates with the customer service staff. For example, if the conversation type is intonation type, that is, the tone corresponding to the conversation is any one of high, medium, and low.
[0081] Exemplarily, in Figure 1 the application scenario of, taking business scenario 1 as an example, the server 102 divides the conversation corpus of user 1 in business scenario 1 into a training sample. Further, the feature vectors in the training sample need to be extracted, such as the speed of the conversation rhythm, the height of the intonation, whether there are sensitive words, and what the corresponding sensitive words are. Similarly, the server 102 can extract the feature vectors in the conversation corpora of user 2 and user 3 in business scenario 1, and further summarize them into the same training set.
[0082] It should be noted that the extraction of feature vectors from the conversation corpus in this application is to meet the input requirements of the user sentiment analysis model and improve the processing efficiency of the model.
[0083] Therefore, in the embodiments of this application, by extracting features from the training corpus, at least one type of conversation feature vector that meets the input requirements of the user sentiment analysis model is obtained, and the model does not need to perform additional processing on the training corpus. While meeting the requirements, the processing efficiency of the model can be improved.
[0084] Optionally, extracting features from the training corpus with the same intention and the same user voiceprint type to obtain at least one type of conversation feature vector, including:
[0085] Based on the neural network model, extracting features from the training corpus with the same intention and the same user voiceprint type to obtain at least one of the following types of conversation feature vectors: the feature vector of the conversation rhythm type, the feature vector of the tone type, the feature vector of the intonation type, and the feature vector of the sensitive word type.
[0086] In the embodiments of this application, the conversation rhythm type may refer to the speed of the user's communication with the customer service staff, such as fast conversation rhythm or slow conversation rhythm, etc.; the tone type may refer to the degree of relaxation of the user's communication tone with the customer service staff, such as calm and comfortable communication, little fluctuation, gentle tone or impatient tone, etc.; the intonation type may refer to the high or low pitch of the user's communication tone with the customer service staff, such as high, medium, low, etc. of the voice; the sensitive word type may refer to the inclusion of sensitive words in the user's communication with the customer service staff, such as complaints, are you sick, etc. The server can extract the corresponding feature vectors based on the above conversation types.
[0087] In this step, when extracting features from the training corpus with the same intention and the same user voiceprint type based on the neural network model, the features of the training corpus extracted by each convolutional layer in the neural network model can be used as feature vectors, or the features output by the last convolutional layer or the last residual module containing the convolutional layer can be used as feature vectors. The embodiments of this application do not make specific limitations on this.
[0088] It should be noted that the embodiments of this application can also extract feature vectors from the training corpus through other algorithms, such as convolutional neural network models, etc. The embodiments of this application do not make specific limitations on this.
[0089] Exemplarily, in Figure 1In the application scenario, taking business scenario 2 as an example, server 102 can respectively extract features from the training corpora of user 1, user 4, and user 5 in business scenario 2 based on a neural network model. Further, for the training corpus of user 1, the feature vectors of the conversation type obtained can be that the conversation rhythm is fast, the tone is impatient, the intonation is high, and the sensitive word "complaint"; for the training corpus of user 2, the feature vectors of the conversation type obtained can be that the conversation rhythm is slow, the tone is gentle, and the intonation is low; for the training corpus of user 3, the feature vectors of the conversation type obtained can be that the conversation rhythm is slow, the tone is impatient, and the intonation is medium.
[0090] Therefore, by using a neural network model to extract features from the training corpus, manual participation in calculations is not required, which can improve the accuracy and speed of feature extraction.
[0091] Optionally, training a corresponding user sentiment analysis model according to the training set includes:
[0092] Training a user sentiment analysis model according to the feature vectors in the training set and the labels corresponding to the feature vectors; wherein, the label is a negative emotion label or a positive emotion label obtained by performing semantic recognition on the training corpus using a preset dictionary library;
[0093] The user sentiment analysis model is a convolutional neural network model established based on speech psychology.
[0094] Specifically, inputting the feature vectors in the training set into the corresponding user sentiment analysis model to obtain the model output result;
[0095] Calculating the error loss of the user sentiment analysis model according to the model output result and the label;
[0096] Updating the network parameters of the user sentiment analysis model according to the error loss, and repeating the training until the error loss meets the preset conditions and then stopping the training to obtain the trained user sentiment analysis model.
[0097] In the embodiments of the present application, the label corresponding to the feature vector can refer to knowing in advance whether a certain training corpus corresponds to a negative emotion label or a positive emotion label. The error loss can refer to the error loss between the predicted value and the true value obtained by the user sentiment analysis model. The network parameters can refer to the weights corresponding to each convolutional layer in the user sentiment analysis model, that is, the weights corresponding to each feature vector during calculation in the convolutional layer. The preset condition can refer to a preset threshold set by the system. When the error loss is less than this preset threshold, the training is stopped. This preset threshold can also be modified manually. The embodiments of the present application do not make specific limitations on the numerical value and setting form of the preset threshold.
[0098] Among them, during the process of training the user sentiment analysis model using the training set, it is necessary to continuously learn the error loss between the predicted value and the true value to adjust the network parameters of the user sentiment analysis model itself.
[0099] In this step, the present application pre-establishes a dictionary library, which contains all the words in the training corpus. Each word corresponds to a unique identification number. Using one-hot text representation, the semantic recognition result of the training corpus is obtained through text mapping, that is, negative sentiment labels and positive sentiment labels can be obtained.
[0100] Exemplarily, in Figure 1 In the application scenario, the server 102 can input the conversation corpus of user 1, user 2, and user 3 in business scenario 1 into the corresponding user sentiment analysis model to obtain 3 model output results; among them, the label corresponding to the conversation corpus of user 1 is a negative sentiment, the label corresponding to the conversation corpus of user 2 is a positive sentiment, and the label corresponding to the conversation corpus of user 3 is a positive sentiment. Further, according to the 3 model output results and the labels corresponding to the conversation corpus of each user above, the error loss of the user sentiment analysis model is calculated respectively. Furthermore, the network parameters of the user sentiment analysis model can be updated according to each calculated error loss, and the training is repeated until the above error loss is less than the preset threshold to stop training, and the trained user sentiment analysis model is obtained. The process of training the user sentiment analysis model with the conversation corpus of each user in business scenario 2 is similar to the above and will not be elaborated here.
[0101] Therefore, continuously training the user sentiment analysis model using multiple training sets until the error loss of the user sentiment analysis model meets the preset conditions and then stopping training can improve the calculation accuracy of the user sentiment analysis model.
[0102] Optionally, calculating the error loss of the user sentiment analysis model according to the model output result and the label includes:
[0103] Calculating the error loss of the user sentiment analysis model using a predefined algorithm according to the model output result and the label.
[0104] In the embodiments of the present application, the predefined algorithm can be methods such as gradient descent method, backpropagation method, focal loss function, regularization, etc. Through the predefined algorithm, the error loss of the user sentiment analysis model can be calculated. The embodiments of the present application do not limit the specific algorithm of the predefined algorithm. It can refer to the algorithms for calculating the error loss of the model in the prior art or can be a redefined algorithm.
[0105] Exemplarily, in Figure 1In the application scenario, taking business scenario 1 as an example, server 102 uses a preset dictionary library to perform semantic recognition on the conversation corpus of user 1, the conversation corpus of user 2, and the conversation corpus of user 3, and obtains that the label corresponding to the conversation corpus of user 1 is a negative emotion label, the label corresponding to the conversation corpus of user 2 is a positive emotion label, and the label corresponding to the conversation corpus of user 3 is a positive emotion label. Further, according to the three output results obtained by inputting the conversation corpus of user 1, the conversation corpus of user 2, and the conversation corpus of user 3 into the convolutional neural network model established based on speech psychology and the labels corresponding to the above user conversation corpus, the error loss of the user emotion analysis model is calculated respectively using the gradient descent method.
[0106] Therefore, the embodiment of the present application can obtain the label corresponding to the feature vector, that is, extract the features of the semantics. Further, the user emotion analysis model is trained using the label and the feature vector in the training set, which improves the training rate and accuracy of the model.
[0107] It should be noted that to establish a convolutional neural network model based on speech psychology, specifically, a large amount of data in the psychology database can be obtained. The psychology database includes the quotes of users talking to customer service and the user emotion labels corresponding to the quotes. Further, the quotes of users talking to customer service and the user emotion labels corresponding to the quotes are input into the convolutional neural network model for training to obtain an initial user emotion analysis model for subsequent training processes.
[0108] Among them, in the subsequent training process, multiple training corpora in the historical corpus can be input into the model for training.
[0109] Optionally, the label includes multiple levels of emotion labels. Training the user emotion analysis model according to the feature vector in the training set and the label corresponding to the feature vector includes:
[0110] Search for the preset corresponding level of emotion label based on each training set;
[0111] Train the user emotion analysis model according to the feature vector in each training set and the found corresponding level of emotion label.
[0112] Among them, the label is at least one level of negative emotion label and / or positive emotion label obtained by performing semantic recognition on the training corpus using a preset dictionary library. The emotion label of this level is a label corresponding to what type of corpus set in advance. For example, if there is the word "complaint" in a certain corpus, the set corresponding level is negative level two; specifically, if a certain training set includes 5 training samples, the labels corresponding to the 5 training samples can be positive level one, positive level two, negative level one, negative level two, and negative level three respectively.
[0113] It should be noted that the embodiments of the present application do not limit the number and specific levels of emotion tags including levels, and can be modified manually.
[0114] Specifically, training a user sentiment analysis model according to the feature vectors in each training set and the found emotion tags corresponding to the levels includes: inputting the feature vectors of each training set into the user sentiment analysis model to obtain an output result;
[0115] Calculating the error loss of the user sentiment analysis model according to the output result of each training set and the found emotion tags corresponding to the levels to obtain multiple error losses;
[0116] Updating the network parameters of the user sentiment analysis model according to the multiple error losses, repeating the training until the multiple error losses meet the preset conditions and then stopping the training to obtain a trained user sentiment analysis model.
[0117] In this step, each training set includes multiple training samples. Inputting each training sample into the user sentiment analysis model to obtain an output result. Since the tags include emotion tags of multiple levels, such as positive level 1 to positive level 4, negative level 1 to negative level 4, etc., correspondingly, the output result of the model also corresponds to emotion tags of multiple levels. For example, the output result of the model can be emotion tags with levels such as positive level 1, positive level 2, negative level 1, negative level 2, etc. The emotion tags can also be represented by numerical identifiers. For example, 0-4 corresponds to negative level 1 to negative level 5, and 5-8 corresponds to positive level 1 to positive level 4. The embodiments of the present application do not make specific limitations on this.
[0118] Exemplarily, in Figure 1 the application scenario of, taking business scenario 1 as an example, the server 102 can respectively find the preset emotion tags corresponding to the levels based on the training samples in each training set; for example, the preset tag corresponding to the conversation corpus of user 1 is negative level 2, the preset tag corresponding to the conversation corpus of user 2 is positive level 1, and the preset tag corresponding to the conversation corpus of user 3 is positive level 2. Further, obtaining the model output results corresponding to the conversation corpus of user 1, the conversation corpus of user 2, and the conversation corpus of user 3, and then calculating the error loss of the user sentiment analysis model according to the above model output results and the found emotion tags corresponding to the levels to obtain 3 error losses.
[0119] Therefore, in the embodiments of the present application, different emotion level tags can also be set for the output result of the model, and the classification result is more accurate. In this way, the user's emotion can be determined more precisely, the change of the user's emotion can be sensed, the adaptability is wider, and the user experience is improved.
[0120] Optionally, the method further includes:
[0121] Obtaining the personality characteristics of the user and the corresponding recommended speech library, wherein the personality characteristics of the user are marked by a feature vector of at least one conversation type; the recommended speech library is used to recommend maintenance speech to corresponding customer service personnel for reference during communication;
[0122] The personality traits and the corresponding recommendation speech library are divided into a training sample data set and a test sample data set according to a random ratio, and the training sample data set and the test sample data set are input into a user sentiment analysis model for model training and testing.
[0123] In the embodiment of the present application, personality traits may refer to the rhythm and social skills used to characterize the user's behavior, such as the ability to deal with people, which can be divided into eagle type, peacock type, dove type, owl type, etc., among which, eagle type refers to the personality traits of being quick to do things, decisive in decision-making, centered on facts and tasks, and not good at dealing with people; peacock type refers to the personality traits of being quick to do things, decisive in decision-making, and particularly strong in the ability to communicate with people, usually centered on people; dove type refers to the personality traits of being friendly, calm, and not impatient when doing things; owl type refers to the personality traits of not being easy to show friendliness to the other party and not being very talkative at ordinary times. The above personality traits can be marked by the feature vector of at least one conversation type, such as the eagle type can be marked by feature vectors such as fast conversation rhythm, medium tone, and decisive tone.
[0124] It is understandable that, in actual applications, the personality characteristics of the corresponding user can also be searched based on the feature vector of at least one conversation type, and then input into the model to obtain the corresponding recommendation speech library.
[0125] Among them, the recommended script library includes maintenance scripts for communicating with customers in various business scenarios. The maintenance scripts refer to scripts for telephone communication used to retain users, so that based on the customer's psychology, language can be spoken that is more acceptable to them.
[0126] In this step, the personality traits and the corresponding recommendation speech library are divided into a training sample data set and a test sample data set in a random ratio. The training sample data set and the test sample data set can be divided in a ratio of 10:1, or in other ratios. The embodiment of the present application does not make specific limitations on this, but usually the proportion occupied by the training sample data set is higher than the proportion occupied by the test sample data set.
[0127] For example, in Figure 1In the application scenario, the server 102 can obtain the personality characteristics of users marked by the feature vectors of at least one session type and the corresponding recommended speech database. For example, in business scenario 1, the feature vector corresponding to the conversation corpus of user 1 is that the conversation rhythm is fast, the intonation is medium, and the tone is decisive. The above feature vector can be marked as the personality characteristic of the eagle type. Similarly, obtain the personality characteristics corresponding to other users in business scenario 1 and business scenario 2. Further, obtain the recommended speech database corresponding to each personality characteristic, and divide the above personality characteristics and the corresponding recommended speech database into a training sample data set and a test sample data set according to a ratio of 5:1. Further, input the training sample data set and the test sample data set into the user sentiment analysis model for model training and testing.
[0128] It should be noted that in actual use, a large number of personality characteristics and the corresponding recommended speech database need to be obtained for model training and testing. The above is only an example for illustration.
[0129] Therefore, training the model based on the user's personality characteristics and the corresponding recommended speech database can enable the model to generate the corresponding recommended speech database in the same scenario based on the user's personality characteristics when in use, for the customer service staff to refer to during communication, providing convenience for the customer service staff, and improving the flexibility of model input and output.
[0130] Combined with the above embodiments, Figure 3 is a schematic flowchart of a specific method for training a user sentiment analysis model provided by an embodiment of the present application. As Figure 3 shown, the implementation method steps of the embodiment of the present application include:
[0131] Step A: Input the historical coordinate complaint session recording, the historical coordinate normal session recording, and the psychological data into the early warning and maintenance model (user sentiment analysis model) for training. Specifically, classify the user intentions of the recording files through natural language processing (NLP) technology. Further, cluster the corpus of user demands with the same user intention or the same voiceprint, and execute step B.
[0132] Step B: In the same session intention business scenario, extract the session feature vectors, that is, extract vectors for semantics to obtain positive or negative emotions; extract vectors for the conversation rhythm to obtain a fast or slow conversation rhythm; extract vectors for the tone to obtain an impatient or peaceful tone; extract vectors for the intonation to obtain a high or low intonation; extract vectors for sensitive words to obtain "complaint" or "are you sick", etc. After extracting the feature vectors, execute step C.
[0133] Step C: Input the feature vector into a model established based on a convolutional neural network to obtain an output result, where the output result includes the user's emotion label and forms a high-quality language corpus under the same intention. Among them, the model established based on the convolutional neural network is based on psychological data, and through mining positive corpus and complaint corpus under the same intention, a user emotion change model is formed. The user's emotion label can be approval, resistance, disgust, anger (i.e., different levels of user emotion labels). Further, based on psychological analysis theory, calculate the current mental state of the current dialogue user and make a judgment. If the mental state of the user is negative at the current moment, remind the customer service staff to pay attention and return the maintenance measures calculated by the model (such as maintenance words, customer service speech rate, customer service tone).
[0134] Optionally, the method further includes:
[0135] Obtain the user's voice information and parse the voice information into conversation corpus through natural language processing technology;
[0136] Extract the feature vectors of at least one conversation type from the conversation corpus;
[0137] Input the feature vectors of the at least one conversation type into a trained user emotion analysis model, and based on the model output result, judge whether it is necessary to send a warning prompt message to the customer service.
[0138] In the embodiments of the present application, natural language processing technology may refer to the technology of using natural language used by humans for communication to interact with machines. Through artificial processing of natural language, the computer can read and understand it.
[0139] In this step, the warning prompt message may refer to the message sent to the customer service to remind the user that the emotion has changed. The embodiments of the present application do not specifically limit the content and sending form of the warning prompt message. It can directly display a message prompt box on the corresponding terminal device of the customer service to remind the customer service, or it can be in the form of sending a vibration prompt tone, etc.
[0140] Exemplarily, the server can collect the user's conversation voice data in real time. Further, the collected corpus is transmitted to the background, and the natural language processing NLP technology is used to process the conversation corpus data in real time, extract the feature vectors of at least one conversation type, and then transmit them to the user emotion analysis model. When the model warning threshold is reached (i.e., the emotion changes to negative), a warning prompt message is sent to the customer service, and at the same time, the words output by the model can also be returned to help the operator carry out maintenance and rescue operations.
[0141] Therefore, when the user emotion analysis model trained in the embodiments of the present application is used, the voice information of the user is input to obtain a prediction result, and a model warning is performed based on the prediction result, which provides convenience for customer service personnel and timely discovers the emotional changes of the user.
[0142] Optionally, the method further includes:
[0143] If it is determined that a warning prompt message needs to be sent to the customer service, the corresponding customer service emergency handling method is searched in the look-up table based on the model output result; the customer service emergency handling method includes the adjustment methods for the speech rate, intonation, and conversation rhythm of the customer service personnel, as well as the recommended words for maintaining the relationship.
[0144] In this step, the look-up table refers to a table for storing the corresponding relationship between different emotion labels and customer service emergency handling methods. Among them, each emotion label has a corresponding customer service emergency handling method. The customer service emergency handling method includes how to adjust the speech rate, intonation, and conversation rhythm of the customer service personnel in a certain scenario, as well as the recommended words for maintaining the relationship to the customer service personnel. For example, in a certain scenario, the customer service emergency handling method can be to slow down the speech rate, lower the intonation, slow down the conversation rhythm, and recommend the words for maintaining the relationship 1 to the customer service personnel.
[0145] Exemplarily, in Figure 1 the application scenario, if the server 102 determines that a warning prompt message needs to be sent to the customer service, the corresponding customer service emergency handling method is searched in the look-up table based on the model output result, and the customer service emergency handling method is sent to the terminal device of the customer service personnel for the customer service personnel to refer to.
[0146] Therefore, the embodiments of the present application can find the corresponding strategy based on the output result. For example, if a warning prompt is performed after determining that a warning prompt message needs to be sent to the customer service, the customer service emergency handling method can be searched based on the output result to solve the problems encountered in communication, thereby reducing user complaints, retaining users, and improving the user experience.
[0147] Optionally, the method further includes:
[0148] If it is determined not to send a warning prompt message to the customer service, the conversation corpus of the customer service personnel is obtained to retrain the trained user emotion analysis model based on the conversation corpus.
[0149] Exemplarily, in Figure 1 the application scenario, if the server 102 determines that it is not necessary to send a warning prompt message to the customer service, the conversation corpus of the customer service personnel is automatically obtained and input into the user emotion analysis model again for training, so that the model is optimized and the prediction accuracy of the model is improved.
[0150] It can be understood that the customer service staff can modify the words in the recommended retention script, and then the server automatically obtains the words spoken by the customer service staff and inputs them into the model for model optimization.
[0151] Therefore, the embodiments of the present application can find corresponding strategies based on the output results. For example, after determining not to send a warning prompt message to the customer service, model optimization can be performed, thereby improving the accuracy of model prediction.
[0152] Optionally, the user sentiment analysis model provided by the embodiments of the present application can be put into use after training. During the use of the model, when it is monitored that the user's emotion changes to a negative emotion, the operator can be notified in a timely manner, and the retention script can be returned in real time to help reduce complaints. Exemplarily, Figure 4 FIG. is a schematic flowchart of a process of working using the user sentiment analysis model provided by the embodiments of the present application. As Figure 4 shown, the process of the user sentiment analysis model working includes the following steps:
[0153] Step A: Real-time monitor the customer service conversation voice and analyze the conversation corpus between the user and the customer service. Process the conversation corpus through natural language processing (NLP) technology, that is, extract feature vectors such as semantics, conversation rhythm, tone, intonation, and sensitive words for each round of conversation, and execute Step B.
[0154] Step B: Input the feature vectors into the warning and retention analysis model. The model analyzes the psychological changes of the conversation customer in real time according to the feature vectors and judges whether it is necessary to warn the customer service. If the user's emotion is negative, it is judged that it is necessary to warn the customer service, then warn the customer service and return the emergency handling for the customer service (emergency script, adjust the speech rate, adjust the tone), and the process ends; if the user's emotion is positive, it is judged that it is not necessary to warn the customer service, and the process ends.
[0155] Among them, the model will monitor the conversation corpus generated by the user's communication with the customer service and the provided retention script. After being confirmed by the operator as no problem or optimized, it can be fed back to the model again, that is, input into the model regularly to perform enhanced training on the model for optimizing and upgrading the model.
[0156] In the foregoing embodiments, the training method of the user sentiment analysis model provided by the embodiments of the present application is introduced. In order to implement each function in the method provided by the embodiments of the present application, the electronic device as the execution subject may include a hardware structure and / or a software module, and implement each of the above functions in the form of a hardware structure, a software module, or a combination of a hardware structure and a software module. Whether a certain function among the above functions is executed in the form of a hardware structure, a software module, or a combination of a hardware structure and a software module depends on the specific application and design constraint conditions of the technical solution.
[0157] For example, Figure 5The following is a schematic structural diagram of a user sentiment analysis model training device provided by an embodiment of the present application, as shown in Figure 5 As shown, the device includes: an acquisition module 510, a clustering module 520, and a training module 530; wherein, the acquisition module 510 is configured to obtain a plurality of training corpora from a preset historical corpus and determine the user intentions corresponding to each training corpus;
[0158] The clustering module 520 is configured to cluster the training corpora of each user intention according to voiceprints to obtain a plurality of user voiceprint types;
[0159] The training module 530 is configured to divide the training corpora with the same intention and the same user voiceprint type into the same training set, and train the corresponding user sentiment analysis model according to the training set.
[0160] Optionally, the training module 530 includes a division module and a training module;
[0161] Wherein, the division module includes an extraction unit and a summary unit;
[0162] Specifically, the extraction unit is configured to extract features from the training corpora with the same intention and the same user voiceprint type to obtain feature vectors of at least one session type;
[0163] The summary unit is configured to summarize the feature vectors of the at least one session type into the same training set.
[0164] Optionally, the extraction unit is specifically configured to:
[0165] Based on a neural network model, extract features from the training corpora with the same intention and the same user voiceprint type to obtain the following feature vectors of at least one session type: feature vectors of dialogue rhythm type, tone type, intonation type, and sensitive word type.
[0166] Optionally, the training module is specifically configured to:
[0167] Train a user sentiment analysis model according to the feature vectors in the training set and the labels corresponding to the feature vectors; wherein, the labels are negative emotion labels or positive emotion labels obtained by performing semantic recognition on the training corpora using a preset dictionary library;
[0168] The user sentiment analysis model is a convolutional neural network model established based on speech psychology.
[0169] Optionally, the labels include emotion labels of multiple levels, and the training module is specifically configured to:
[0170] Search for the corresponding level of emotion labels preset for each training set;
[0171] Train a user sentiment analysis model based on the feature vectors in each training set and the found emotion labels corresponding to the levels.
[0172] Optionally, the device further includes a training and testing module, and the training and testing module is used for:
[0173] Obtain the personality characteristics of the user and the corresponding recommended speech library, where the personality characteristics of the user are marked by feature vectors of at least one conversation type; the recommended speech library is used to recommend recovery speech to the corresponding customer service staff for reference during communication;
[0174] Divide the personality characteristics and the corresponding recommended speech library into a training sample data set and a testing sample data set according to a random ratio, and input the training sample data set and the testing sample data set into the user sentiment analysis model for model training and testing.
[0175] Optionally, the device further includes an application module, and the application module is used for:
[0176] Obtain the voice information of the user, and parse the voice information into conversation corpus through natural language processing technology;
[0177] Extract the feature vectors of at least one conversation type from the conversation corpus;
[0178] Input the feature vectors of the at least one conversation type into the trained user sentiment analysis model, and judge whether it is necessary to send a warning prompt message to the customer service based on the model output result.
[0179] Optionally, the device further includes a warning module, and the warning module is used for:
[0180] When it is determined that it is necessary to send a warning prompt message to the customer service, look up the corresponding customer service emergency handling method in the look-up table based on the model output result; the customer service emergency handling method includes the adjustment methods for the speaking speed, intonation, and conversation rhythm of the customer service staff, as well as the recommended recovery speech.
[0181] Optionally, the device further includes a retraining module, and the retraining module is used for:
[0182] When it is determined not to send a warning prompt message to the customer service, obtain the conversation corpus of the customer service staff to retrain the trained user sentiment analysis model based on the conversation corpus.
[0183] For the specific implementation principle and effect of the user sentiment analysis model training device provided in the embodiments of the present application, reference can be made to the corresponding descriptions and effects in the above embodiments, and details are not described here.
[0184] The embodiments of the present application also provide a schematic structural diagram of an electronic device. Figure 6 It is a schematic structural diagram of an electronic device provided by the embodiments of the present application. As Figure 6 shown, the electronic device may include: a processor 601 and a memory 602 communicatively connected to the processor; the memory 602 stores a computer program; the processor 601 executes the computer program stored in the memory 602, so that the processor 601 executes the method described in any of the foregoing embodiments.
[0185] Among them, the memory 602 and the processor 601 may be connected through a bus 603.
[0186] The embodiments of the present application also provide a computer-readable storage medium storing computer program execution instructions, which are used to implement the method described in any of the foregoing embodiments of the present application when executed by a processor.
[0187] The embodiments of the present application also provide a chip for running instructions, and the chip is used to execute the method described in any of the foregoing embodiments executed by an electronic device.
[0188] The embodiments of the present application also provide a computer program product, which includes a computer program that can implement the method described in any of the foregoing embodiments executed by an electronic device when executed by a processor.
[0189] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of devices or modules can be in an electrical, mechanical or other form.
[0190] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to implement the solution of this embodiment.
[0191] In addition, in each embodiment of the present application, the functional modules can be integrated into a processing unit, or each module can exist physically alone, or two or more modules can be integrated into one unit. The unit formed by the above modules can be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.
[0192] The integrated module implemented in the form of a software functional module can be stored in a computer-readable storage medium. The above software functional module is stored in a storage medium, including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute some steps of the methods described in each embodiment of the present application.
[0193] It should be understood that the above processor can be a Central Processing Unit (CPU for short), and can also be other general-purpose processors, Digital Signal Processors (DSP for short), Application Specific Integrated Circuits (ASIC for short), etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the application can be directly implemented by the execution of the hardware processor, or can be implemented by the combination of the hardware and software modules in the processor.
[0194] The memory may include a high-speed random access memory (Random Access memory, RAM for short), and may also include a non-volatile memory (Non-volatile Memory, NVM for short), such as at least one disk memory, and can also be a USB flash drive, a mobile hard disk, a read-only memory, a magnetic disk, or an optical disc, etc.
[0195] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, the bus in the drawings of the present application is not limited to only one bus or one type of bus.
[0196] The above storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random-Access Memory (SRAM), Electrically Erasable Programmable Read Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk. The storage medium can be any available medium accessible by a general-purpose or special-purpose computer.
[0197] An exemplary storage medium is coupled to the processor, enabling the processor to read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an Application Specific Integrated Circuit (ASIC). Of course, the processor and the storage medium can also exist as discrete components in an electronic device or a master control device.
[0198] As described above, the above is only the specific implementation manner of the embodiments of the present application, but the protection scope of the embodiments of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the embodiments of the present application should be covered by the protection scope of the embodiments of the present application. Therefore, the protection scope of the embodiments of the present application shall be subject to the protection scope of the claims.
Claims
1. A method for training a user sentiment analysis model, characterized in that, The method includes: Obtaining a plurality of training corpora from a preset historical corpus, and determining the user intentions corresponding to each training corpus; Performing clustering operations on the training corpora of each user intention according to voiceprints to obtain multiple user voiceprint types; Dividing the training corpora with the same intention and the same user voiceprint type into the same training set, and training the corresponding user sentiment analysis model according to the training set; Dividing the training corpora with the same intention and the same user voiceprint type into the same training set, including: Performing feature extraction on the training corpora with the same intention and the same user voiceprint type based on a neural network model to obtain feature vectors of at least one of the following conversation types: feature vectors of dialogue rhythm type, feature vectors of tone type, feature vectors of intonation type, and feature vectors of sensitive word type; Summarizing the feature vectors of at least one of the conversation types into the same training set.
2. The method according to claim 1, wherein Training the corresponding user sentiment analysis model according to the training set, including: Training the user sentiment analysis model according to the feature vectors in the training set and the labels corresponding to the feature vectors; wherein, the labels are negative emotion labels or positive emotion labels obtained by performing semantic recognition on the training corpora using a preset dictionary library; The user sentiment analysis model is a convolutional neural network model established based on speech psychology.
3. The method according to claim 2, wherein The labels include emotion labels of multiple levels. Training the user sentiment analysis model according to the feature vectors in the training set and the labels corresponding to the feature vectors, including: Searching for the corresponding level of emotion labels preset for each training set; Training the user sentiment analysis model according to the feature vectors in each training set and the found corresponding level of emotion labels.
4. The method according to claim 1, characterized in that The method further includes: Obtaining the personality characteristics of the user and the corresponding recommended conversation strategy library, where the personality characteristics of the user are marked by feature vectors of at least one conversation type; the recommended conversation strategy library is used to recommend conversation strategies for reference by corresponding customer service staff during communication; Dividing the personality characteristics and the corresponding recommended conversation strategy library into a training sample data set and a test sample data set according to a random ratio, and inputting the training sample data set and the test sample data set into the user sentiment analysis model for training and testing of the model.
5. The method according to any one of claims 1-4, characterized in that, The method further includes: Obtaining the voice information of the user, and parsing the voice information into conversation corpora through natural language processing technology; Extracting feature vectors of at least one conversation type from the conversation corpora; Inputting the feature vectors of at least one conversation type into the trained user sentiment analysis model, and determining whether to send a warning prompt message to the customer service based on the model output result.
6. The method according to claim 5, characterized in that, The method further includes: If it is determined that a warning prompt message needs to be sent to the customer service, then searching for the corresponding customer service emergency handling method in the lookup table based on the model output result; the customer service emergency handling method includes methods for adjusting the speech rate, intonation, and conversation rhythm of the customer service staff, as well as recommended conversation strategies.
7. The method according to claim 5, wherein The method further includes: If it is determined not to send a warning prompt message to the customer service, then obtaining the conversation corpora of the customer service staff to retrain the trained user sentiment analysis model based on the conversation corpora.
8. A user emotion analysis model training device, characterized in that The device includes: An acquisition module, configured to acquire a plurality of training corpora from a preset historical corpus and determine the user intents corresponding to the respective training corpora; A clustering module, configured to cluster the training corpora of each user intent according to voiceprints to obtain a plurality of user voiceprint types; A training module, configured to divide the training corpora with the same intent and the same user voiceprint type into the same training set and train a corresponding user sentiment analysis model according to the training set; The division module includes an extraction unit and a summarization unit; The extraction unit is configured to extract features from the training corpora with the same intent and the same user voiceprint type based on a neural network model to obtain feature vectors of at least one of the following conversation types: a feature vector of a conversation rhythm type, a feature vector of a tone type, a feature vector of an intonation type, and a feature vector of a sensitive word type; The summarization unit is configured to summarize the feature vectors of the at least one conversation type into the same training set.
9. An electronic device, characterized in that, Comprising: A processor, a memory, and a computer program; wherein, the computer program is stored in the memory and is configured to be executed by the processor, and the computer program includes instructions for executing the user sentiment analysis model training method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, they are used to implement the user sentiment analysis model training method according to any one of claims 1-7.
Citation Information
Patent Citations
A speech recommendation method and device, a computer device and a storage medium
CN109033257A
Method for synthesized speech generation using emotion information correction and apparatus
US20210074261A1