Telephone service data processing method and device, equipment and storage medium

By employing emotion recognition and tailored communication strategies based on user data, the method addresses low conversion rates and poor user experience in telephone sales by optimizing call timing and dialogue, improving answer rates and user satisfaction.

CN120321337APending Publication Date: 2025-07-15CHINA PING AN PROPERTY INSURANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510572146.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

In telesales, sales personnel are unable to accurately utilize target user information, resulting in low conversion rates and unstable user experience, especially when users are busy or resting, making calls can easily be considered harassment.

Method used

By obtaining user attribute data and historical behavior data, the time prediction model is used to determine the optimal call time, and use the emotion recognition model to adjust the speech when calling, avoid making calls when inappropriate, and adjust the speech in combination with the emotion recognition model to improve the user experience.

Benefits of technology

It improves phone answering rates, reduces the possibility of being considered harassment, and improves the quality and user experience of telephone services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120321337A_ABST
    Figure CN120321337A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of deep learning and voice processing, and provides a telephone service data processing method and device, equipment and a storage medium, and the method comprises the steps: obtaining the attribute data and historical behavior data of a user; based on a preset time prediction model, according to the attribute data and the historical behavior data, determining target power-off time of contacting the target user; acquiring an audio feature corresponding to the audio data of the target user when the target user is in a call within the target call time; based on a preset emotion recognition model, determining an emotion recognition result of the target user according to the audio features; and according to the emotion recognition result of the target user, adjusting the verbal skill of conversation with the target user. According to the invention, the target calling time of contacting the target user is predicted through the time prediction model, so that the call answering rate is improved and the possibility of being regarded as harassment is reduced; and the emotion of the target user is recognized through the emotion recognition model, so that the conversation verbal skill is adjusted according to the emotion of the target user, and the quality of the telephone service is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical fields of deep learning and speech processing, and particularly to a method, apparatus, device, and storage medium for processing telephone service data. Background Art

[0002] Telemarketing has become one of the important channels for enterprises' pre-sales services. Since salespersons cannot accurately utilize the information of target users, the overall sales conversion rate through telemarketing is very low; and when the user status is unstable, salespersons cannot adjust the service direction in a timely manner, resulting in the inability to provide a good user experience for users.

[0003] Taking the financial field as an example, with the intensifying competition in the insurance industry, it has become a common sales service method for the staff of insurance companies to communicate with users by phone. However, frequent telemarketing services are easily regarded as harassment by users, thus affecting the user experience. Summary of the Invention

[0004] The main purpose of this application is to provide a method, apparatus, device, and storage medium for processing telephone service data, which can avoid making calls when users are busy or resting, so as to improve the call answering rate and reduce the possibility of being regarded as harassment, thereby improving the quality of telephone services and the user experience.

[0005] In a first aspect, this application provides a method for processing telephone service data, including:

[0006] Obtain the attribute data and historical behavior data of the user;

[0007] Based on a preset time prediction model, determine the target call time for contacting the target user according to the attribute data and the historical behavior data;

[0008] Obtain the audio features corresponding to the audio data of the target user when making a call to the target user within the target call time;

[0009] Based on a preset emotion recognition model, determine the emotion recognition result of the target user according to the audio features;

[0010] Adjust the conversation script for communicating with the target user according to the emotion recognition result of the target user.

[0011] In a second aspect, this application also provides a device for processing telephone service data, including:

[0012] A first acquisition module, configured to obtain the attribute data and historical behavior data of the user;

[0013] A prediction module, configured to determine a target call time for contacting a target user based on a preset time prediction model, according to the attribute data and the historical behavior data;

[0014] A second acquisition module, configured to acquire audio features corresponding to audio data of the target user when making a call to the target user within the target call time;

[0015] An identification module, configured to determine an emotion recognition result of the target user based on a preset emotion recognition model according to the audio features;

[0016] An adjustment module, configured to adjust the conversation script for communicating with the target user according to the emotion recognition result of the target user.

[0017] In a third aspect, the present application further provides a computer device, which includes a memory and a processor;

[0018] The memory is used for storing a computer program;

[0019] The processor is configured to execute the computer program and implement the telephone service data processing method as described above when executing the computer program.

[0020] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the telephone service data processing method as described above are implemented.

[0021] The present application provides a telephone service data processing method, apparatus, device and storage medium. The method includes: acquiring attribute data and historical behavior data of a user; determining a target call time for contacting a target user based on a preset time prediction model according to the attribute data and the historical behavior data; acquiring audio features corresponding to audio data of the target user when making a call to the target user within the target call time; determining an emotion recognition result of the target user based on a preset emotion recognition model according to the audio features; and adjusting the conversation script for communicating with the target user according to the emotion recognition result of the target user. After acquiring the attribute data and historical behavior data of the user, the present application predicts a target call time suitable for contacting the target user through a preset time prediction model, avoiding making calls when the user is busy or resting, so as to improve the call answer rate and reduce the possibility of being regarded as harassment; and during the call with the target user, acquiring audio features corresponding to the audio data of the target user, so as to identify the emotion of the target user through a preset emotion recognition model, facilitating the adjustment of the conversation script according to the emotion of the target user, reducing the dissatisfaction of the user, and thus improving the quality of telephone service and the experience of the user. Description of the Drawings

[0022] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0023] Figure 1 It is a schematic flowchart of a method for processing telephone service data provided by an embodiment of the present application;

[0024] Figure 2 It is a schematic connection diagram of a server and a terminal device provided by an embodiment of the present application;

[0025] Figure 3 It is a schematic block diagram of a device for processing telephone service data provided by an embodiment of the present application;

[0026] Figure 4 It is a schematic block diagram of the structure of a computer device provided by an embodiment of the present application. Specific Embodiments

[0027] The following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, rather than all, of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present application.

[0028] The flowchart shown in the accompanying drawings is only an example illustration, and does not necessarily include all contents and operations / steps, nor does it necessarily execute in the described order. For example, some operations / steps can also be decomposed, combined, or partially merged. Therefore, the actual execution order may be changed according to the actual situation.

[0029] The embodiments of the present application provide a method, device, equipment, and storage medium for processing telephone service data. Among them, the method for processing telephone service data can be applied to a terminal device, and the terminal device can be a device such as a mobile phone, a tablet computer, a notebook computer, or a desktop computer. It can also be applied to a server, and the server can be a single server or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.

[0030] The following will describe in detail some embodiments of the present application with reference to the accompanying drawings. Without conflict, the following embodiments and the features in the embodiments may be combined with each other.

[0031] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a method for processing telephone service data provided by an embodiment of the present application. It should be noted that the method for processing telephone service data provided by the embodiment of the present application can be used in a terminal device, and of course, it can also be used in a server.

[0032] As Figure 2 shown, the method for processing telephone service data is applied to a server. The server is communicatively connected to the terminal device, and through the server, the script adjusted by the method for processing telephone service data can be sent to the terminal device. Of course, it is not limited to this, and no limitation is made here.

[0033] In specific implementation, the terminal device includes but is not limited to: any one of a mobile phone, a tablet computer, a notebook computer, and a desktop computer; the server can be a single server, a server cluster, or a cloud server providing cloud computing services.

[0034] As Figure 1 shown, the method for processing telephone service data includes steps S101 to S105.

[0035] Step S101: Obtain the attribute data and historical behavior data of the user.

[0036] It can be understood that the attribute data of the user may include the user's name, gender, age, occupation, address, model of the vehicle under the user's name, vehicle price, etc. The historical behavior data of the user may include the user's online browsing of auto insurance products, consultation of auto insurance products, underwriting records, and call records with the staff of the insurance company.

[0037] According to the obtained attribute data and historical behavior data of the user, the embodiment of the present application can classify the intention of the user to purchase products. For example, users who have a vehicle, browse auto insurance products, and consult auto insurance products are classified as high-intention users, that is, target users. Users who do not have a vehicle or have clearly refused similar services in the call records are classified as low-intention customers. Based on this, users with a relatively high intention to purchase auto insurance products can be screened out, thereby improving the call answering rate and the success rate of telephone services.

[0038] Step S102: Based on a preset time prediction model, determine the target call time for contacting the target user according to the attribute data and historical behavior data.

[0039] The time prediction model of the embodiment of the present application can predict the value of a future time point according to the attribute data and historical behavior data, that is, the target call time.

[0040] For example, the time prediction model may include XGBoost, LightGBM, and random forest. Among them, XGBoost iteratively trains multiple weak models (i.e., decision trees), and each new model focuses on correcting the error of the prediction time of the previous model, and finally generates a strong prediction model. LightGBM iteratively trains multiple weak models (i.e., decision trees), and after the iteration ends, the prediction results of the prediction time of multiple weak models are weighted and combined to generate a strong prediction model. Random forest improves the stability and prediction performance of the model by constructing multiple decision trees and integrating their prediction results of the prediction time (e.g., voting or averaging).

[0041] Based on the time prediction model, the target call time suitable for contacting the target user can be predicted according to the user's attribute data and historical behavior data. For example, the time in the area where the user is located can be learned based on the user's address, so as to predict the time suitable for calling the users in this area. For example, if the user's address is in Xinjiang, the call time cannot be predicted according to the working hours corresponding to Beijing time. The target call time suitable for contacting the user can be predicted according to the user's occupation and age. For example, if the user is a restaurant waiter, it can be predicted that 08:30-09:30 in the morning and 16:00-17:00 in the afternoon are suitable for making calls. The target call time can be predicted according to the call record of the user's call with the insurance company staff. For example, if the user indicates in the call record that he does not want to be called in the morning, it can be predicted that the suitable call time is generally in the afternoon or evening. Generally, the target call time for full-time workers is 18:00-20:00 (after work); the time of freelancers / self-employed individuals is more flexible, generally busy in the morning, and the corresponding target call time can be 15:00-17:00 (relatively free in the afternoon); the target call time for retirees can be 10:00-17:00.

[0042] Step S103, obtain the audio features corresponding to the audio data of the target user when communicating with the target user during the target call time.

[0043] It can be understood that the audio feature refers to the key information related to the emotion of the target user. The audio features may include speech rate, intonation, pause duration, and volume, etc. Different audio features can reflect different emotions of the target user. For example, a sudden increase or decrease in speech rate generally reflects the anxiety or impatience of the target user. A sudden rise or fall in intonation generally reflects the anger or dissatisfaction of the target user. A long pause generally reflects the indifference or boredom of the target user. A sudden increase in volume generally reflects the anger of the target user.

[0044] During a call with a target user, audio features corresponding to the audio data of the target user can be obtained to analyze the emotions of the user during the call based on these audio features, thereby improving the success rate of telephone services.

[0045] Step S104: Based on a preset emotion recognition model, determine the emotion recognition result of the target user according to the audio features.

[0046] The emotion recognition model in the embodiments of the present application may include at least one of a convolutional neural network, a recurrent neural network, and a long short-term memory network. For example, the emotion recognition model in the embodiments of the present application may be a hybrid model based on a convolutional neural network and a long short-term memory network, or may be a hybrid model based on a convolutional neural network and a recurrent neural network. Specifically, high-dimensional feature vectors of audio features can be extracted based on a convolutional neural network, and then the time series corresponding to the high-dimensional feature vectors can be modeled based on a recurrent neural network or a long short-term memory network to capture the temporal dependence relationship of the high-dimensional feature vectors. In the embodiments of the present application, audio features such as the speaking speed, intonation, pause duration, and volume of the user are analyzed through the emotion recognition model, and finally the emotion recognition result of the target user is output, improving the accuracy of emotion recognition.

[0047] Step S105: Adjust the conversation script for the call with the target user according to the emotion recognition result of the target user.

[0048] For example, after the emotion recognition result of the target user is output through the emotion recognition model, the conversation script for the call with the target user can be adjusted according to this emotion recognition result. For example, during a call with the target user, when talking about purchasing insurance products (such as critical illness insurance, accident insurance, property insurance, vehicle insurance, etc.), if the target user continuously asks relevant questions about vehicle insurance products and the emotion recognition result output by the emotion recognition model shows a positive emotion, the conversation script can be adjusted, for example, providing a conversation script for introducing vehicle insurance products that meet the needs of the target user. By identifying the emotions of the target user and adjusting the relevant conversation scripts in telephone services, dissatisfaction of the target user can be avoided, thereby improving the success rate of telephone services.

[0049] Exemplarily, a staff member can make a call to the target user. During the call with the target user, the staff member can adjust the conversation strategy according to the emotion recognition result of the target user. Specifically, during the process of the staff member introducing the property insurance product to the target user, when the emotion recognition result output by the emotion recognition model is a negative emotion, a reminder light can be emitted or a pop-up window including the adjusted conversation content can be displayed on the terminal used by the staff member to timely remind the staff member to adjust the relevant conversation strategy, so that the staff member can timely adjust the conversation strategy for introducing the property insurance product to the conversation termination strategy, for example, "Sorry to disturb you. Wish you a happy life." It can be understood that reminder lights of different colors correspond to different emotion recognition results. For example, a red reminder light corresponds to a negative emotion recognition result, a yellow reminder light corresponds to a neutral emotion recognition result, and a green reminder light corresponds to a positive emotion recognition result.

[0050] In another embodiment, a robot can make a call to the target user. During the call with the target user, the robot can automatically adjust the conversation strategy according to the emotion recognition result of the target user, further improving the intelligence of the telephone service. Specifically, during the process of the robot introducing the bank financial product to the target user, when the emotion recognition result output by the emotion recognition model deployed in the robot is a neutral emotion, the conversation strategy can be automatically adjusted, for example, adjusted to conversation strategies such as "Are you not interested in the financial product? Or do you want to learn about the insurance product?"

[0051] The telephone service data processing method provided by the above embodiment includes: obtaining the attribute data and historical behavior data of the user; based on a preset time prediction model, determining the target call time for contacting the target user according to the attribute data and historical behavior data; obtaining the audio features corresponding to the audio data of the target user during the call with the target user at the target call time; based on a preset emotion recognition model, determining the emotion recognition result of the target user according to the audio features; and adjusting the conversation strategy for the call with the target user according to the emotion recognition result of the target user. After obtaining the attribute data and historical behavior data of the user in the embodiment of the present application, the target call time suitable for contacting the target user is predicted through a preset time prediction model, avoiding making calls when the user is busy or resting, so as to improve the call answer rate and reduce the possibility of being regarded as harassment; and during the call with the target user, the audio features corresponding to the audio data of the target user are obtained to identify the emotion of the target user through a preset emotion recognition model, so as to adjust the conversation strategy according to the emotion of the target user, reduce the dissatisfaction of the user, and thus improve the quality of the telephone service and the experience of the user.

[0052] In an exemplary embodiment, the attribute data includes the user's occupation, age, and address; the historical behavior data includes historical call answering results and historical call durations; step S102 may include: inputting the user's occupation, age, address, historical call answering results, and historical call durations into a preset time prediction model to determine the target call time for contacting the target user through the time prediction model.

[0053] It can be understood that the attribute data may include the user's occupation, age, and address. Based on the user's attribute data, users can be classified. For example, users can be classified into categories such as full-time workers, freelancers or self-employed individuals, retirees, etc. The call times corresponding to different user categories are different. Based on the time prediction model in the embodiments of the present application, the target call time suitable for contacting the target user can be predicted according to the user's occupation, age, address, historical call answering results, and historical call durations.

[0054] Specifically, the time prediction model can predict the call time corresponding to the user category according to the user's occupation, age, and address, and then, based on the historical call answering results (such as answered and unanswered) and historical call durations (such as call time, call start time, and call end time, etc.), further predict the most suitable target call time for contacting the target user on the basis of the call time corresponding to the user category, so as to improve the call answering rate and reduce the possibility of being regarded as harassment.

[0055] In an exemplary embodiment, the method further includes steps S201 to S203.

[0056] Step S201: Obtain training data, where the training data includes the user's attribute data and historical behavior data.

[0057] Step S202: Through sampling with replacement, randomly sample in the training data to obtain a plurality of sub-datasets.

[0058] Step S203: Train corresponding decision trees according to each sub-dataset to obtain a time prediction model.

[0059] Taking the financial field as an example, the user's attribute data and historical behavior data can be obtained from the customer management system of an insurance company, and the user's attribute data and historical behavior data are used as the training data for training the time prediction model. Among them, the training data may include attribute data such as the user's age, occupation, and address, as well as historical behavior data such as answered, unanswered, and call time, as shown in Table 1.

[0060] Table 1

[0061]

[0062]

[0063] It is understandable that after obtaining the training data, data cleaning can be performed on the training data. Specifically, handle the missing values of the call time; unify all time data into a unified format, for example, the format of YYYY-MM-DD HH:MM; unify the occupations of users (such as insurance agents, clerks, finance, administration, etc.) into full-time workers; eliminate call durations that exceed the reasonable time range; delete duplicate data, etc.

[0064] The embodiments of the present application can adopt the sampling method with replacement to randomly sample in the training data to obtain multiple sub-datasets. For example, randomly sample in the training data. The first sub-dataset includes features such as full-time workers, 30 years old, and call time of 2025-03-04 18:00; the second sub-dataset includes features such as retirees, 68 years old, and call time of 2025-03-04 19:00. Furthermore, train a decision tree on each sub-dataset. When the decision tree splits nodes each time, randomly select some features for evaluation, and select the best split feature and threshold based on the Gini index.

[0065] Among them, the Gini index is:

[0066] P i is the proportion of the category in the current node. For example, the proportion of answered or not answered.

[0067] For example, a certain node in the decision tree has 10 samples, 6 answered, and 4 not answered. Then substituting into the formula of the Gini index, we can get:

[0068] Gini = 1 - (0.6 2 + 0.4 2 ) = 1 - (0.36 + 0.16) = 0.48.

[0069] The decision trees obtained by training each sub-training set are combined to obtain a time prediction model. When predicting the answering rate based on the time prediction model, each decision tree outputs a classification result (for example, answered and not answered, that is, 1 or 0), and then take the average probability of all decision trees.

[0070]

[0071] Among them, N is the number of decision trees, x is the data including time, and T i(x) is the prediction of the i-th decision tree for the input x (1 means answered, 0 means not answered). P(answered) is the answering probability predicted by the time prediction model for this time.

[0072] In an exemplary embodiment, step S103 may include steps S1031 to S1033.

[0073] Step S1031, obtain the audio data of the target user when communicating with the target user within the target call time.

[0074] Step S1032, preprocess the audio data of the target user to obtain the target audio corresponding to the target user.

[0075] Step S1033, extract features from the target audio corresponding to the target user to obtain the audio features corresponding to the audio data of the target user.

[0076] It can be understood that audio recording can be performed during the call with the target user to obtain the audio data of the target user. In the embodiments of the present application, preprocessing such as voice segmentation and noise reduction processing can be performed on the audio data of the target user to obtain multiple segments of target audio. Then, feature extraction is performed on the target audio to obtain audio features related to the emotion of the target user, thereby improving the recognition efficiency and accuracy of the model.

[0077] For example, it can be set that the number of times of calling the same user within a week shall not exceed three times, and the interval shall not be less than 48 hours. If the user clearly refuses the renewal or transfer service during the call, the user will not be contacted within a short period of time (within one month or half a year) to reduce the disturbance to the user. Of course, if the user clearly indicates that they do not want to answer similar call services anymore, the user will no longer be contacted.

[0078] In an exemplary embodiment, the emotion recognition model includes a first network, a second network, and a fully connected layer; step S104 may include steps S1041 to S1043.

[0079] Step S1041, input the audio features corresponding to the audio data of the target user into the first network to extract the local frequency domain features of the audio features through the first network.

[0080] Step S1042, input the time series corresponding to the local frequency domain features into the second network to capture the temporal dependence relationship of the time series corresponding to the local frequency domain features through the second network to obtain temporal features.

[0081] Step S1043, input the local frequency domain features and the temporal features into the fully connected layer to fuse the local frequency domain features and the temporal features through the fully connected layer to obtain the emotion recognition result of the target user.

[0082] Exemplarily, embodiments of the present application may first obtain the recording data of different users, and use the recording data of different users as the training data for training the emotion recognition model. Taking the financial field as an example, the historical recording data saved in the telephone service system of an insurance company may be used as the training data, or the recording data may be obtained from a publicly available speech emotion dataset, such as the RAVDESS dataset and the CREMA-D dataset. Among them, the RAVDESS dataset includes emotional speech and emotional singing samples, aiming to provide high-quality audio-visual resources for emotional expression analysis. The CREMA-D dataset is an emotional multi-modal actor dataset. In addition, the newly added recording data in the telephone service system of the insurance company may be used as the training data for training the emotion recognition model to continuously optimize the emotion recognition model.

[0083] The emotion recognition model in the embodiments of the present application may include a first network, a second network, and a fully connected layer. The first network may be a convolutional neural network. The second network may be a recurrent neural network and a long short-term memory network. Specifically, the audio features corresponding to the audio data of the target user may be converted into the format of a spectrogram and then input into the convolutional neural network. The local frequency domain features of the spectrogram, such as formants, pitch mutations, etc., are extracted through the convolutional layer of the convolutional neural network, and then the dimension of the local frequency domain features is reduced through the pooling layer of the convolutional neural network. Then, the time series corresponding to the local frequency domain features in the spectrogram is input into the long short-term memory network or the recurrent neural network, and the time series is modeled through the long short-term memory network or the recurrent neural network to capture the temporal dependence relationship and obtain the temporal features. Finally, the feature vector after splicing the local frequency domain features and the temporal features is mapped to the probability space of emotion categories through the fully connected layer to achieve emotion classification, so as to obtain the emotion recognition result of the target user. The emotion recognition model composed of a convolutional neural network and a long short-term memory network, or a convolutional neural network and a recurrent neural network in the embodiments of the present application can more comprehensively capture the spatio-temporal characteristics of emotions, thereby improving the accuracy of emotion recognition.

[0084] In an exemplary implementation manner, communicating with the target user includes: providing a product introduction service; step S105 may include steps S1051 to S1053.

[0085] Step S1051, when the emotion recognition result of the target user is a positive emotion, prompt the words for continuing to provide the product introduction service.

[0086] Step S1052, when the emotion recognition result of the target user is a neutral emotion, prompt the words for stopping to provide the product introduction service and asking about the needs of the target user.

[0087] Step S1053: When the emotion recognition result of the target user is a negative emotion, prompt the words for stopping providing the product introduction service or ending the call with the target user.

[0088] It can be understood that providing the product introduction service may be included during the call with the target user. Taking the financial field as an example, during the call between the staff of an insurance company and the target user, the insurance product introduction service can be provided. During the call between the staff of a bank and the target user, the interest rate introduction service corresponding to different savings periods, the financial product introduction service, or the loan service, etc. can be provided. According to the different emotion recognition results of the target user, prompt the words during the call with the target user, which can meet the needs of the target user in a targeted manner, reduce unnecessary call time, and thus avoid causing dissatisfaction among users.

[0089] In an exemplary embodiment, the method further includes step S100.

[0090] Step S100: Display the emotion recognition result of the target user and the corresponding word content on the terminal used by the staff.

[0091] It can be understood that in the embodiment of the present application, displaying the emotion recognition result of the target user and the corresponding word content on the terminal used by the staff can facilitate the staff to timely understand the emotion of the target user and timely adjust the words, so as to improve the success rate of the telephone service. Specifically, the current emotion recognition result and the previous emotion recognition results of the user can be displayed in the form of a list, as well as the word content corresponding to different emotion recognition results. For example, during the process of the staff of an insurance company recommending insurance products to the target user by phone, if the emotion recognition result of the target user is a negative emotion, then display "Sorry to trouble you. You can call us at any time if you need. Wish you a happy life"; if the emotion recognition result of the target user is a positive emotion, then display "Our company's insurance products include accident insurance, property insurance, vehicle insurance..., among which, there is a promotion for the vehicle insurance product recently..."; if the emotion recognition result of the target user is a neutral emotion, then display "Excuse me, does the insurance product introduced now not meet your needs? You can tell me what insurance products you are interested in, so that I can introduce them to you in a targeted manner." By displaying the emotion of the user and the corresponding word content on the terminal used by the staff, it can not only enable the staff to quickly establish a topic with the target user, but also save the time of the target user and improve the user experience of the telephone service.

[0092] Please refer to Figure 3 , Figure 3 which is a schematic block diagram of a telephone service data processing device provided by the embodiment of the present application. This telephone service data processing device can be configured in a server or a terminal device and is used to execute the foregoing telephone service data processing method.

[0093] As Figure 3 shown, the telephone service data processing device includes: a first acquisition module 110, a prediction module 120, a second acquisition module 130, an identification module 140, and an adjustment module 150.

[0094] The first acquisition module 110 is configured to acquire the attribute data and historical behavior data of the user.

[0095] The prediction module 120 is configured to determine the target call time for contacting the target user based on a preset time prediction model according to the attribute data and historical behavior data.

[0096] The second acquisition module 130 is configured to acquire the audio features corresponding to the audio data of the target user when communicating with the target user at the target call time.

[0097] The identification module 140 is configured to determine the emotion recognition result of the target user based on a preset emotion recognition model according to the audio features.

[0098] The adjustment module 150 is configured to adjust the conversation script for communicating with the target user according to the emotion recognition result of the target user.

[0099] In an exemplary embodiment, the attribute data includes the user's occupation, age, and address; the historical behavior data includes historical answering results and historical call times; the prediction module 120 may specifically be configured to input the user's occupation, age, address, historical answering results, and historical call times into a preset time prediction model to determine the target call time for contacting the target user through the time prediction model.

[0100] In an exemplary embodiment, the device further includes a first acquisition sub-module, a sampling sub-module, and a training sub-module.

[0101] The first acquisition sub-module is configured to acquire training data, and the training data includes the attribute data and historical behavior data of the user.

[0102] The sampling sub-module is configured to randomly sample in the training data by means of sampling with replacement to obtain a plurality of sub-datasets.

[0103] The training sub-module is configured to train corresponding decision trees according to each sub-dataset to obtain a time prediction model.

[0104] In an exemplary embodiment, the second acquisition module 130 may include a second acquisition sub-module, a preprocessing sub-module, and a feature extraction sub-module.

[0105] The second acquisition sub-module is configured to acquire the audio data of the target user when communicating with the target user at the target call time.

[0106] A preprocessing sub-module for preprocessing the audio data of the target user to obtain the target audio corresponding to the target user.

[0107] A feature extraction sub-module for extracting features from the target audio corresponding to the target user to obtain the audio features corresponding to the audio data of the target user.

[0108] In an exemplary embodiment, the emotion recognition model includes a first network, a second network, and a fully connected layer; the recognition module 140 may include a first extraction sub-module, a second extraction sub-module, and a fusion sub-module.

[0109] The first extraction sub-module is configured to input the audio features corresponding to the audio data of the target user into the first network to extract the local frequency domain features of the audio features through the first network.

[0110] The second extraction sub-module is configured to input the time series corresponding to the local frequency domain features into the second network to capture the temporal dependence relationship of the time series corresponding to the local frequency domain features through the second network, and obtain temporal features.

[0111] The fusion sub-module is configured to input the local frequency domain features and the temporal features into the fully connected layer to fuse the local frequency domain features and the temporal features through the fully connected layer to obtain the emotion recognition result of the target user.

[0112] In an exemplary embodiment, making a call with the target user includes: providing a product introduction service; the adjustment module 150 may specifically be used for:

[0113] When the emotion recognition result of the target user is a positive emotion, prompting the words for continuing to provide the product introduction service; when the emotion recognition result of the target user is a neutral emotion, prompting the words for stopping providing the product introduction service and asking about the needs of the target user; when the emotion recognition result of the target user is a negative emotion, prompting the words for stopping providing the product introduction service or ending the call with the target user.

[0114] In an exemplary embodiment, the device further includes a display module.

[0115] The display module is configured to display the emotion recognition result of the target user and the corresponding words on the terminal used by the staff.

[0116] It should be noted that those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described device and each module and unit can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0117] The method of the present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld devices or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0118] Exemplarily, the above-mentioned method and device can be implemented in the form of a computer program, and the computer program can run on a computer device.

[0119] Please refer to Figure 4 , Figure 4 which is a schematic block diagram of the structure of a computer device provided by an embodiment of the present application. The computer device can be a server or a terminal device.

[0120] As Figure 4 shown, the computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the memory can include a storage medium and an internal memory.

[0121] The storage medium can store an operating system and a computer program. The computer program includes program instructions, and when the program instructions are executed, the processor can be made to execute the steps of any one of the telephone service data processing methods.

[0122] The processor is used to provide computing and control capabilities to support the operation of the entire computer device.

[0123] The internal memory provides an environment for the operation of the computer program in the storage medium. When the computer program is executed by the processor, the processor can be made to execute the steps of any one of the telephone service data processing methods.

[0124] The network interface is used for network communication, such as sending assigned tasks, etc.

[0125] Those skilled in the art can understand that Figure 4The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.

[0126] It should be understood that the processor may be a Central Processing Unit (CPU), and the processor may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0127] Among them, in one embodiment, the processor is used to execute a computer program and can implement the following steps when executing the computer program:

[0128] Obtain the attribute data and historical behavior data of the user;

[0129] Based on a preset time prediction model, determine the target call time for contacting the target user according to the attribute data and historical behavior data;

[0130] Obtain the audio features corresponding to the audio data of the target user when making a call to the target user within the target call time;

[0131] Based on a preset emotion recognition model, determine the emotion recognition result of the target user according to the audio features;

[0132] Adjust the conversation script for communicating with the target user according to the emotion recognition result of the target user.

[0133] It should be noted that those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working process of the above-described telephone service data processing can refer to the corresponding process in the embodiment of the foregoing telephone service data processing method, and will not be elaborated herein.

[0134] The embodiment of this application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the following steps can be implemented:

[0135] Obtain the attribute data and historical behavior data of the user;

[0136] Based on a preset time prediction model, determine the target call time for contacting the target user according to the attribute data and historical behavior data;

[0137] Obtain the audio features corresponding to the audio data of the target user when communicating with the target user within the target call time;

[0138] Based on a preset emotion recognition model, determine the emotion recognition result of the target user according to the audio features;

[0139] Adjust the conversation script for communicating with the target user according to the emotion recognition result of the target user.

[0140] Among them, the computer-readable storage medium may be an internal storage unit of the computer device in the foregoing embodiment, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk equipped on the computer device, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc.

[0141] It should be noted that for the functions or steps that can be achieved by the above computer-readable storage medium, reference can be made to the embodiments of the foregoing telephone service data processing method.

[0142] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0143] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.

[0144] The above-described embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.

Claims

1. A method for processing telephone service data, characterized in that Including: Obtain the user's attribute data and historical behavior data; Based on a preset time prediction model, determine the target call time for contacting the target user according to the attribute data and the historical behavior data; Obtain the audio features corresponding to the audio data of the target user when making a call to the target user within the target call time; Based on a preset emotion recognition model, determine the emotion recognition result of the target user according to the audio features; Adjust the conversation script for the call with the target user according to the emotion recognition result of the target user.

2. The method for processing telephone service data according to claim 1, characterized in that, The attribute data includes the user's occupation, age, and address; the historical behavior data includes the historical call answer result and historical call time; the step of determining the target call time for contacting the target user based on a preset time prediction model according to the attribute data and the historical behavior data includes: Input the user's occupation, age, address, historical call answer result, and historical call time into a preset time prediction model to determine the target call time for contacting the target user through the time prediction model.

3. The method for processing telephone service data according to claim 1, wherein The method further includes: Obtain training data, where the training data includes the user's attribute data and historical behavior data; Through sampling with replacement, randomly sample in the training data to obtain multiple sub-datasets; Train corresponding decision trees according to each sub-dataset to obtain a time prediction model.

4. The telephone service data processing method according to claim 1, wherein The step of obtaining the audio features corresponding to the audio data of the target user when making a call to the target user within the target call time includes: Obtain the audio data of the target user when making a call to the target user within the target call time; Preprocess the audio data of the target user to obtain the target audio corresponding to the target user; Extract features from the target audio corresponding to the target user to obtain the audio features corresponding to the audio data of the target user.

5. The method for processing telephone service data according to claim 1, wherein The emotion recognition model includes a first network, a second network, and a fully connected layer; the step of determining the emotion recognition result of the target user based on a preset emotion recognition model according to the audio features includes: Input the audio features corresponding to the audio data of the target user into the first network to extract the local frequency domain features of the audio features through the first network; Input the time series corresponding to the local frequency domain features into the second network to capture the temporal dependence relationship of the time series corresponding to the local frequency domain features through the second network to obtain temporal features; Input the local frequency domain features and the temporal features into the fully connected layer to fuse the local frequency domain features and the temporal features through the fully connected layer to obtain the emotion recognition result of the target user.

6. The method for processing telephone service data according to claim 1, wherein The call with the target user includes: providing a product introduction service; the step of adjusting the conversation script for the call with the target user according to the emotion recognition result of the target user includes: When the emotion recognition result of the target user is a positive emotion, prompt the conversation script for continuing to provide the product introduction service; When the emotion recognition result of the target user is a neutral emotion, prompt the words for stopping the product introduction service and asking about the needs of the target user; When the emotion recognition result of the target user is a negative emotion, prompt the words for stopping the product introduction service or ending the call with the target user.

7. The method for processing telephone service data according to any one of claims 1 to 6, characterized in that, The method further includes: Displaying the emotion recognition result of the target user and the corresponding word content on the terminal used by the staff.

8. A telephone service data processing device, characterized in that, Including: A first acquisition module, configured to acquire the attribute data and historical behavior data of the user; A prediction module, configured to determine the target call time for contacting the target user based on the preset time prediction model according to the attribute data and the historical behavior data; A second acquisition module, configured to acquire the audio features corresponding to the audio data of the target user when calling the target user within the target call time; An identification module, configured to determine the emotion recognition result of the target user based on the preset emotion recognition model according to the audio features; An adjustment module, configured to adjust the words for calling the target user according to the emotion recognition result of the target user.

9. A computer device, characterized in that, The computer device includes a memory and a processor; The memory is used to store a computer program; The processor is configured to execute the computer program and implement the telephone service data processing method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the telephone service data processing method according to any one of claims 1 to 7 are implemented.