Intelligent outbound call method and device, computer device and storage medium
By performing speech recognition and text preprocessing on customer voice data, and combining it with a multi-intent recognition model, the problem of low intent recognition accuracy in intelligent outbound calling systems has been solved, achieving more efficient intent recognition and response.
Patent Information
- Application Number
- CN202210920853.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-02
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2042-08-02
AI Technical Summary
Existing intelligent outbound calling systems have low accuracy in intent recognition, resulting in an inability to respond effectively and promptly, thus impacting the service experience.
The system acquires customer voice data for speech recognition and semantic parsing, acquires customer text data for text preprocessing, uses a multi-intent recognition model corresponding to the target domain for intent recognition, and performs outbound calls based on the recognition results.
It improved the accuracy and efficiency of intent recognition, enhanced the response quality and efficiency of intelligent outbound calls, and improved the customer experience.
Smart Images

Figure CN115240676B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of artificial intelligence, and in particular to an intelligent outbound call method and device, computer equipment and a storage medium. BACKGROUND
[0002] Outbound call refers to automatically dialing a user's phone number through a computer and playing a recorded voice to the user through the computer. It is an indispensable part of the computer telephone integrated system of the customer service center. The intelligent outbound call system is a common method in the industry. Some repetitive notifications can be reported by machines to reduce human waste. In the intelligent outbound call system, it is very important to accurately identify the intent of the voice data returned by the customer in a timely manner. Only by accurately identifying the customer's intent can targeted responses be given, such as transferring to a human phone, information questions, or modifying personal information. When the intelligent outbound call system identifies the intent of the voice data, the low accuracy of intent recognition may result in a failure to timely and effectively respond, thereby affecting the service experience. SUMMARY
[0003] The embodiments of the present application provide an intelligent outbound call method, device, computer equipment and storage medium to solve the problem of low accuracy of intent recognition in the intelligent outbound call system.
[0004] An intelligent outbound call method comprises:
[0005] obtaining customer voice data in a target field;
[0006] performing voice recognition and semantic analysis on the customer voice data to obtain customer text data;
[0007] performing text preprocessing on the customer text data to obtain target key data;
[0008] using a multi-intent recognition model corresponding to the target field to identify the target key data to determine at least one target intent;
[0009] An intelligent outbound call device comprises:
[0010] a customer voice data obtaining module configured to obtain customer voice data in a target field;
[0011] a customer text data obtaining module configured to perform voice recognition and semantic analysis on the customer voice data to obtain customer text data;
[0012] a target key data obtaining module configured to perform text preprocessing on the customer text data to obtain target key data;
[0013] a target intention determination module configured to determine at least one target intention by using a multi-intention recognition model corresponding to the target field to recognize the target key data;
[0014] an outbound call module configured to obtain voice reply data corresponding to the at least one target intention, and perform outbound call based on the voice reply data.
[0015] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the intelligent outbound call method when executing the computer program.
[0016] A computer readable storage medium stores a computer program, and the computer program is executable by a processor to implement the intelligent outbound call method.
[0017] The intelligent outbound call method, device, computer device, and storage medium can perform voice recognition and semantic analysis on customer voice data of a target field, obtain customer text data, and ensure the feasibility of customer intention recognition. The target key data is obtained by preprocessing the customer text data, which helps to ensure the accuracy and efficiency of target intention recognition. The target key data is recognized by using a multi-intention recognition model corresponding to the target field, at least one target intention is determined, and the accuracy and efficiency of multi-intention recognition are improved. The outbound call is performed based on the voice reply data corresponding to the target intention, the corresponding outbound call operation is made according to the target intention, and the response quality and efficiency of intelligent outbound call are improved. BRIEF DESCRIPTION OF DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the description of the embodiments of the present application. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0019] Figure 1 is an application environment diagram of the intelligent outbound call method in an embodiment of the present application;
[0020] Figure 2 is a flowchart of the intelligent outbound call method in an embodiment of the present application;
[0021] Figure 3 is another flowchart of the intelligent outbound call method in an embodiment of the present application;
[0022] Figure 4 is another flowchart of the intelligent outbound call method in an embodiment of the present application;
[0023] Figure 5 is another flowchart of the intelligent outbound call method in an embodiment of the present application;
[0024] Figure 6 is another flowchart of the intelligent outbound call method in an embodiment of the present application;
[0025] Figure 7 is another flowchart of the intelligent outbound call method in an embodiment of the present application;
[0026] Figure 8 is another flowchart of the intelligent outbound call method in an embodiment of the present application;
[0027] Figure 9 is a schematic diagram of the intelligent outbound call device in an embodiment of the present application;
[0028] Figure 10 is a schematic diagram of the computer device in an embodiment of the present application. DETAILED DESCRIPTION
[0029] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of the present application.
[0030] The intelligent outbound call method provided by the embodiments of the present application can be applied in an application environment as shown in Figure 1 . Specifically, the intelligent outbound call method is applied in an intelligent outbound call system, which includes a client and a server as shown in Figure 1 . The client and the server communicate through a network to realize multi-intent recognition on customer voice data and intelligent outbound call based on at least one target intent recognized, which helps to improve customer experience. The client, also known as the user end, is a program that provides local services for customers corresponding to the server. The client can be installed on, but not limited to, various personal computers, notebook computers, smart phones, tablet computers and portable wearable devices. The server can be realized by an independent server or a server cluster composed of multiple servers.
[0031] In an embodiment, as shown in Figure 2 , an intelligent outbound call method is provided. Taking the server in Figure 1 as an example, the method includes the following steps:
[0032] S201: Obtain customer voice data in a target field;
[0033] S202: Perform voice recognition and semantic analysis on the customer voice data to obtain customer text data;
[0034] S203: Perform text preprocessing on the customer text data to obtain target key data;
[0035] S204: Use a multi-intent recognition model corresponding to the target domain to recognize the target key data and determine at least one target intent;
[0036] S205: Obtain voice reply data corresponding to the at least one target intent, and perform outbound call based on the voice reply data.
[0037] The target domain refers to the domain to which the content of the voice communication between the intelligent outbound call system and the customer belongs, for example, when the intelligent outbound call system is used to recommend insurance business, the target domain thereof is the insurance domain. The customer voice data refers to the voice data collected during the conversation between the intelligent outbound call system and the customer. Generally, during the conversation between the intelligent outbound call system and the customer, the customer will be prompted in advance that the conversation recording exists, and after the customer's permission, the customer voice data collected during the conversation can be obtained.
[0038] As an example, in step S201, during the conversation between the intelligent outbound call system and the customer, the server obtains the customer voice data corresponding to the target domain to which the conversation content belongs after the customer's permission, so as to perform intent recognition and response based on the customer voice data.
[0039] The customer text data is the text data obtained by performing voice recognition and semantic analysis on the customer voice data.
[0040] As an example, in step S202, the server performs voice recognition and semantic analysis on the customer voice data, which can specifically use a voice recognition tool to recognize the customer voice data to convert the voice content into text content, and then use a semantic analysis tool to analyze the text content recognized by the customer voice data to obtain the customer text data, thereby preparing for the subsequent customer intent recognition.
[0041] The target key data is the data obtained by performing text preprocessing on the customer text data. The text processing method uses word segmentation, and then removes stop words and punctuation marks for each sentence. The stop word refers to some words or characters that are automatically filtered out before or after processing natural language data (or text) in order to save storage space and improve search efficiency.
[0042] As an example, in step S203, the server pre-processes the customer text data, specifically, adopts a pre-set screening filtering logic to screen and filter the customer text data, so as to automatically filter out some words or phrases irrelevant to the intent recognition, so as to obtain the target key data, which can save the memory space required in the subsequent intent recognition process and improve the search efficiency, avoid the interference of irrelevant information, and help improve the processing efficiency and recognition accuracy of the intent recognition. The screening filtering logic is a pre-set processing logic for screening and filtering words or phrases irrelevant to the intent recognition.
[0043] The target intent is an intent identified from the target key data. The multi-intent recognition model corresponding to the target domain is a model trained from the training data corresponding to the target domain and capable of recognizing multiple intents.
[0044] As an example, in step S204, the server identifies the target key data in the same target domain by using the multi-intent recognition model corresponding to the target domain, which can recognize at least one target intent at the same time compared with the single-intent recognition model, and helps improve the accuracy of the target intent recognition. Since the multi-intent recognition model and the target key data correspond to the same target domain, the accuracy of the target intent recognition can be further guaranteed. In this example, the multi-intent recognition model is a multi-classification model constructed by using multiple single-intent recognition models. When the multi-intent recognition model is used to identify the target key data, each single-intent recognition model can output a target intent with a higher probability, so as to obtain at least one target intent at the same time and improve the recognition accuracy and efficiency of the at least one target intent.
[0045] The voice reply data is data obtained by replying to the customer according to the target intent
[0046] As an example, in step S205, after obtaining the at least one target intent, the server can determine whether there is an intent reply text corresponding to the at least one target intent according to the at least one target intent. The intent reply text is text data for replying to the at least one target intent. If there is an intent reply text corresponding to the at least one target intent, a text-to-speech tool is adopted to convert the intent reply text into voice to obtain voice reply data, and then an outbound call is performed based on the voice reply data, so as to realize intelligent voice reply according to the at least one target intent identified, which helps improve the voice reply efficiency and improve the customer experience. If there is no intent reply text corresponding to the at least one target intent, the current node jumps to a different outbound call processing mode, for example, can be transferred to a human, can be called later, or can be explained and pacified, so as to make corresponding outbound call operation according to the at least one target intent and improve the response efficiency of the intelligent outbound call.
[0047] In the example, the server can query the system database based on the at least one target intent, determine whether there is an intent response text corresponding to each target intent, and if there is an intent response file for each target intent, splice the intent response text corresponding to the at least one target intent to obtain an intent reply text, so as to realize quick response according to the at least one target intent, and help improve the response efficiency of intelligent outbound calls.
[0048] In the intelligent outbound call method provided in the embodiment, the customer voice data of the target field is subjected to speech recognition and semantic analysis to obtain customer text data, which guarantees the feasibility of customer intent recognition; the customer text data is preprocessed to obtain target key data, which helps guarantee the recognition accuracy and efficiency of target intent; the target key data is identified by using a multi-intent recognition model corresponding to the target field to determine at least one target intent, which improves the accuracy and efficiency of multi-intent recognition; and the outbound call is made based on the voice reply data corresponding to the target intent, and the corresponding outbound operation is made according to the target intent, which improves the response quality and efficiency of intelligent outbound calls.
[0049] In an embodiment, as shown in Figure 3 In step S203, the customer text data is subjected to text preprocessing to obtain target key data.
[0050] S301: Tokenizing the customer text data to obtain original key data;
[0051] S302: Removing stop words and punctuation marks from the original key data to obtain target key data.
[0052] The original key data is obtained by tokenizing the customer text data. A tokenizer is used for tokenizing the customer text.
[0053] As an example, in step S301, the server inputs the customer text data into the tokenizer for tokenization to obtain the original key data. In the example, the tokenizer can be, but is not limited to, jieba tokenization. Jieba tokenization is a Python Chinese tokenization component that can perform functions such as Chinese text tokenization, part-of-speech tagging, and keyword extraction, and supports custom dictionaries.
[0054] The target key data is obtained by removing stop words and punctuation marks from the original key data.
[0055] As an example, in step S302, the server can remove stop words and punctuation marks from the original key data to obtain the target key data. Specifically, the removal of stop words and punctuation marks is based on a dedicated list pre-selected by business logic. This dedicated list records the stop words and punctuation marks that need to be removed. The server can remove stop words and punctuation marks from the original key data according to the business logic to obtain the target key data, avoiding interference from stop words and punctuation marks on subsequent intent recognition and helping to ensure the accuracy of intent recognition.
[0056] In the intelligent outbound calling method provided in this embodiment, a word segmenter, including but not limited to jieba, is first used to segment the customer text data to obtain the original key data; then, the original key data is processed to remove stop words and punctuation to obtain the target key data. By segmenting the words first and then removing stop words and punctuation, we can avoid the situation where some words, when viewed alone, are unnecessary interjections, but when combined with other words, they become meaningful phrases. Therefore, the accuracy of intent recognition can be improved.
[0057] In one embodiment, such as Figure 4 As shown, in step S204, the multi-intent recognition model corresponding to the target domain includes multiple single-intent recognition models corresponding to the target domain.
[0058] A multi-intent recognition model corresponding to the target domain is used to identify key target data and determine at least one target intent, including:
[0059] S401: Employ multiple single-intent recognition models corresponding to the target domain to analyze the target key data and obtain the original intent and the recognition probability corresponding to the original intent output by the multiple single-intent recognition models.
[0060] S402: Identify at least one original intent with a probability greater than a preset probability as at least one target intent.
[0061] Here, the original intent is the intent predicted by the single intent recognition model based on the analysis of key target data. The recognition probability corresponding to the original intent is the probability of obtaining a certain original intent based on the analysis and evaluation of key target data by the single intent recognition model. Each original intent corresponds to a recognition probability, which represents the confidence level in predicting the original intent.
[0062] As an example, in step S401, the server adopts a plurality of single-intent recognition models corresponding to the target field to respectively analyze the target key data, and obtains an original intent predicted by each single-intent recognition model and a recognition probability corresponding to the original intent. In this example, the same target key data is recognized by using a plurality of single-intent recognition models corresponding to the target field, multi-intent recognition at the same time can be recognized, the recognition range is improved, and the multi-intent recognition efficiency is improved.
[0063] The preset probability is a probability preset for evaluating whether the recognition probability meets a large standard. The target intent is an original intent with a recognition probability greater than the preset probability.
[0064] As an example, in step S402, the server compares the recognition probability corresponding to the original intent with the preset probability set in advance. When the recognition probability is greater than the preset probability, it is determined that the recognition probability of this original intent meets a large standard, and the original intent with a recognition probability greater than the preset probability is regarded as a target intent. In this example, at least one target intent is determined according to the comparison result of the recognition probability corresponding to at least one original intent and the preset probability, and the accuracy and recognition efficiency of multi-intent recognition are improved.
[0065] In the intelligent outbound method provided in this embodiment, the same target key data is recognized by using a plurality of single-intent recognition models corresponding to the target field, multi-intent recognition at the same time can be recognized, the recognition range is improved, and the multi-intent recognition efficiency is improved; the recognition probability corresponding to the original intent is compared with the preset probability set in advance. When the recognition probability is greater than the preset probability, it is determined that the recognition probability of this original intent meets a large standard, and the original intent with a recognition probability greater than the preset probability is regarded as a target intent, and the accuracy and recognition efficiency of multi-intent recognition are improved.
[0066] In an embodiment, as shown in Figure 5 Before obtaining the customer voice data, the intelligent outbound method further includes:
[0067] S501: Obtain model training data corresponding to the same target field, and the model training data includes training text data and multi-intent labels;
[0068] S502: Obtain a plurality of single-intent recognition models, construct a multi-classification model based on the plurality of single-intent recognition models, set the last layer of the multi-classification model without an activation function, or set the last layer of the multi-classification model to use a sigmoid function or a logits function;
[0069] S503: Train the multi-classification model by using the model training data, and obtain a multi-intent recognition model.
[0070] The model training data corresponding to the same target field is data used for model training. The training text data is text data used for model training. The multi-intent label is a multi-label classification required by the training model. In order to maintain the unity and integrity of the input data and output data during training, the multi-intent label needs to be packaged into one-hot encoding form before training. One-hot encoding is also known as one-bit effective encoding. In this example, an N-bit status register can be used to encode N states, each state has its own register bit, and at any time, only one bit is valid.
[0071] As an example, in step S501, the server obtains training text data and multi-intent labels, and packages the multi-intent labels into one-hot encoding form. The obtained training text data and the multi-intent labels packaged into one-hot encoding form are used as model training data to prepare for multi-intent recognition model training. For example, the training text data is labeled with a multi-intent label 00101000, indicating that it has the third intent and the fifth intent, and no other intent. This allows the target field corresponding training text data to be distinguished by intent, so that the trained multi-intent recognition model can recognize multiple intents at the same time.
[0072] The single-intent recognition model refers to a model that can recognize a single intent. The single-intent recognition model can be selected from a variety of models such as SVM (support vector machines) binary classification model, logistic regression, and deep neural network. The multi-classification model is a model constructed from multiple single-classification models. The activation function is a function running on the neuron of the artificial neural network, responsible for mapping the input of the neuron to the output. The sigmoid function is the most widely used type of activation function, which has an exponential function shape. It is most similar to biological neurons in physical meaning and is a commonly used S-shaped function in biology, also known as the S-shaped growth curve. The logits function line is a common S-shaped function, and the logits function and the sigmoid function are inverse functions of each other.
[0073] As an example, in step S502, the server obtains a plurality of single-intent recognition models, constructs the plurality of single-intent recognition models into a multi-classification model, sets the last layer of the multi-classification model without an activation function, or sets the last layer of the multi-classification model to use a sigmoid function or a logits function. In this example, the server sets the last layer of the multi-classification model without an activation function, which can avoid the last layer performing weighting or other processing on the output result of the multi-classification model, so that each single-classification model outputs a label, and the multi-intent recognition model can perform multi-intent recognition. Generally, the activation function of the last layer of the single-intent recognition model often uses a softmax function. Softmax is an activation function that can normalize a numerical vector into a probability distribution vector, and the sum of each probability is 1. Therefore, the single label corresponding to the maximum probability cannot be selected. In this example, the sigmoid activation function and the logits activation function are used as the activation function of the last layer of the multi-intent recognition model. Since each classification is independent, it will not be equal to 1 like the softmax function. In this way, if the recognition probability of multiple classifications is higher than a preset probability, the data corresponds to multiple classifications, so that the multi-intent recognition model can perform multi-intent recognition.
[0074] The multi-intent recognition model is obtained by training the multi-classification model using model training data.
[0075] As an example, in step S503, the server trains the multi-classification model using model training data. The model training data includes training text data and multi-intent labels, where the multi-intent labels are in one-hot encoding form. Since the form of the labels is different, the loss function of the multi-classification model needs to be modified to binary cross entropy (binary cross entropy loss function), circle loss, or focal loss, etc. In this example, the server needs to modify the label code of the multi-classification model. The label code of the single-classification model is to calculate the coincidence rate of the classification with the maximum probability and the real classification. The multi-classification model needs to be modified to multi-label classification, so that it becomes to calculate the coincidence rate of the classification greater than the threshold value and the real classification. The server inputs the modified multi-classification model into the model training data to train and obtain a multi-intent recognition model, so that the multi-intent recognition model can perform multi-intent recognition.
[0076] The intelligent outbound method provided by the embodiment trains the model training data corresponding to the same target field, and prepares for multi-intent recognition model training; a multi-classification model is constructed based on multiple single-intent recognition models, the last layer of the multi-classification model has no activation function or uses a sigmoid function or a logits function, and the feasibility of multi-intent recognition model training is improved; and the modified multi-classification model is input into model training data for training to obtain a multi-intent recognition model, so that the multi-intent recognition model can perform multi-intent recognition.
[0077] In an embodiment, as shown in Figure 6 Before the target key data is recognized by the multi-intent recognition model and at least one target intent is determined, the intelligent outbound method further includes:
[0078] S601: The intent rule engine is used to perform matching processing on the target key data, and it is determined whether the intent rule engine can match the first intent corresponding to the target key data;
[0079] S602: If the first intent can be matched, the first intent is determined as the target intent;
[0080] S603: If the first intent cannot be matched, the target key data is recognized by the multi-intent recognition model corresponding to the target field, and at least one target intent is determined.
[0081] The rule engine refers to a component that can reduce the complexity of a complex business logic component, reduce the maintenance and scalability cost of an application program. The intent rule engine is a rule engine used to implement intent recognition. The first intent is the intent matched by the rule engine to the target key data.
[0082] As an example, in step S601, the server uses the intent rule engine to perform matching processing on the target key data, and it is determined whether the intent rule engine can match the first intent corresponding to the target key data, so that according to the matching result, the target intent corresponding to the target key data can be quickly determined,
[0083] As an example, in step S602, the server uses the intent rule engine to match the first intent corresponding to the target key data, and the server takes the matched first intent as the target intent. For example, the intent rule engine identifies that the target key data includes "I confirm" and other words reflecting the explicit intent of the customer, and then determines the first intent in the intent rule engine as the target intent. In this example, the intent rule engine is used to perform intent matching on the target key data, which helps to improve the recognition efficiency of the target intent.
[0084] As an example, in step S603, the server adopts the first intent corresponding to the target key data that cannot be matched by the intent rule engine, and the server performs the multi-intent recognition model corresponding to the target field to recognize the target key data and determine at least one target intent.
[0085] In the intelligent outbound method provided in the embodiment, the target key data is matched by using the intent rule engine to determine whether the first intent corresponding to the target key data can be matched, if the first intent can be matched, the first intent is taken as the target intent, if the first intent cannot be matched, the target key data is matched by using the retrieval analysis model to determine at least one target intent. Before the multi-intent recognition model is used to recognize the target key data, the intent rule engine is used for recognition, which reduces the complexity of the logic component of the intelligent outbound system, improves the accuracy and efficiency of intent recognition, and improves the efficiency of the intelligent outbound method.
[0086] In an embodiment, as shown in Figure 7 Before the multi-intent recognition model is used to recognize the target key data and determine at least one target intent, the intelligent outbound method further includes:
[0087] S701: The retrieval analysis model is used to match the target key data to determine whether the second intent corresponding to the target key data can be matched by the retrieval analysis model;
[0088] S702: If the second intent can be matched, the second intent is determined as the target intent;
[0089] S703: If the second intent cannot be matched, the multi-intent recognition model corresponding to the target field is used to recognize the target key data to determine at least one target intent.
[0090] The retrieval analysis model includes an ES retrieval algorithm and a semantic matching model. The ES (ElasticSearch) retrieval is an open source search engine based on Apache Lucene (TM), which is a distributed and highly scalable full-text retrieval search engine, and also provides near real-time indexing, analysis, and search functions. The second intent is an intent obtained by matching the target key data by the retrieval analysis model.
[0091] As an example, in step S701, the server matches the target key data by using the retrieval analysis model to determine whether the second intent corresponding to the target key data can be matched by the retrieval analysis model, so that the target intent corresponding to the target key data can be quickly determined according to the matching result.
[0092] As an example, in step S702, the server matches the target key data by using the retrieval analysis model, and when the retrieval analysis model matches the second intent corresponding to the target key data, the second intent is taken as the target intent. For example, the ES database connected to the server has pre-stored a plurality of ES configuration data and a second intent corresponding to each ES configuration data, and the server can match the target key data and the ES configuration data corresponding thereto, and if the matching is successful, the second intent corresponding to the ES configuration data is determined as the target intent.
[0093] As an example, in step S703, the server matches the target key data by using the retrieval analysis model, and when the retrieval analysis model does not match the second intent corresponding to the target key data, the multi-intent recognition model corresponding to the target domain is used to recognize the target key data to determine at least one target intent.
[0094] In the intelligent outbound method provided in the embodiment, if the retrieval analysis model can match the second intent corresponding to the target key data, the second intent is determined as the target intent; if the retrieval analysis model cannot match the second intent, the multi-intent recognition model is used to recognize the target key data to determine at least one target intent. Before the multi-intent recognition model is used to recognize the intent of the target key data, the retrieval analysis model is used to analyze the target key data, which can further improve the accuracy of intent recognition.
[0095] In an embodiment, as shown in FIG. 7, Figure 8 In step S701, the retrieval analysis model is used to match the target key data, and it is determined whether the retrieval analysis model can match the second intent corresponding to the target key data, including:
[0096] S801: The ES retrieval algorithm is used to preliminarily retrieve the target key data to obtain a plurality of ES configuration data, and each ES configuration data corresponds to a configuration intent;
[0097] S802: The semantic matching model is used to analyze the similarity of the plurality of ES configuration data and the target key data to obtain a target similarity corresponding to each ES configuration data;
[0098] S803: If there is a target similarity greater than a preset similarity, the configuration intent corresponding to the ES configuration data corresponding to the target similarity is determined as the second intent corresponding to the target key data, and it is determined that the retrieval analysis model can match the second intent.
[0099] S804: If all target similarities are not greater than the preset similarity, it is determined that the retrieval analysis model cannot match the second intent.
[0100] The ES configuration data is obtained by using an ES retrieval algorithm to preliminarily retrieve the target key data. The configuration intent is an intent obtained by manual annotation and log cleaning, and the configuration intent is stored in an ES retrieval library.
[0101] As an example, in step S801, the server uses an ES retrieval algorithm to preliminarily retrieve the target key data, obtains a plurality of ES configuration data, and each ES configuration data has a corresponding configuration intent. The ES configuration data herein can be understood as data containing all or part of the target keywords pre-stored in the ES database. The configuration intent refers to an intent pre-stored in the ES database and matched with the ES configuration data.
[0102] The semantic matching model is a model for judging whether two sentences express the same or similar meaning. The target similarity is a similarity obtained by using the semantic matching model to analyze the similarity between the ES configuration data and the target key data. In this example, since the ES retrieval data retrieves a plurality of ES configuration data, each ES configuration data will have a corresponding target similarity.
[0103] As an example, in step S802, the server uses a semantic matching model to score the similarity between the plurality of ES configuration data and the target key data, and obtains the target similarity between each ES configuration data and the target key data. In this example, the semantic matching model used is ESIM (Enhanced LSTM for Natural Language Inference), which is a text similarity calculation model.
[0104] The preset similarity is a preset similarity threshold.
[0105] As an example, in step S803, the server compares the target similarity corresponding to at least one ES configuration data with the preset similarity. If there is a target similarity greater than the preset similarity, the configuration intent corresponding to the ES configuration data corresponding to the target similarity is taken as the second intent.
[0106] As an example, in step S804, the server compares the target similarity corresponding to at least one ES configuration data with the preset similarity. If all target similarities are not greater than the preset similarity, there is no second intent in the configuration intent corresponding to the ES configuration data.
[0107] In the intelligent outbound method provided in the embodiment, the ES retrieval algorithm is used to perform matching processing on the target key data, obtain a plurality of ES configuration data, each ES configuration data corresponds to a configuration intention, and preparation is made for obtaining the second intention; the semantic matching model is used to perform similarity scoring on the ES configuration data and the target key data, obtain a target similarity, and when the target similarity is greater than a preset similarity, the configuration intention corresponding to the ES configuration data can be used as the second intention; if the target similarity is less than the preset similarity, there is no second intention, so that the accuracy and efficiency of intention recognition are improved.
[0108] It should be understood that the size of the serial number of each step in the above embodiment does not mean the order of execution, and the execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiment of the application.
[0109] In an embodiment, an intelligent outbound device is provided, which corresponds to the intelligent outbound method in the above embodiment. As shown in the figure, the intelligent outbound device includes a customer voice data acquisition module 901, a customer text data acquisition module 902, a target key data acquisition module 903, a target intention determination module 904, and an outbound module 905. The functions of each functional module are described in detail as follows: Figure 9
[0110] The customer voice data acquisition module 901 is configured to acquire customer voice data in a target field.
[0111] The customer text data acquisition module 902 is configured to perform voice recognition and semantic analysis on the customer voice data, and acquire customer text data.
[0112] The target key data acquisition module 903 is configured to perform text preprocessing on the customer text data, and acquire target key data.
[0113] The target intention determination module 904 is configured to use a multi-intention recognition model corresponding to the target field to recognize the target key data, and determine at least one target intention.
[0114] The outbound module 905 is configured to acquire voice reply data corresponding to the at least one target intention, and perform outbound based on the voice reply data.
[0115] In an embodiment, the target key data acquisition module 903 includes:
[0116] The original key data acquisition unit is configured to perform word segmentation on the customer text data, and acquire original key data.
[0117] The target key data acquisition unit is configured to perform stop word and punctuation symbol removal processing on the original key data, and acquire target key data.
[0118] In an embodiment, the target intent determination module 904 comprises:
[0119] An original intent recognition unit is configured to use a plurality of single-intent recognition models corresponding to the target domain to analyze the target key data respectively, and obtain original intents output by the plurality of single-intent recognition models and recognition probabilities corresponding to the original intents.
[0120] A target intent determination unit is configured to determine at least one original intent with a recognition probability greater than a preset probability as at least one target intent.
[0121] In an embodiment, the multi-intelligent outbound call device further comprises:
[0122] A model training data acquisition unit is configured to acquire model training data corresponding to the same target domain, the model training data comprising training text data and multi-intent labels.
[0123] A multi-classification model construction unit is configured to acquire a plurality of single-intent recognition models, construct a multi-classification model based on the plurality of single-intent recognition models, set the last layer of the multi-classification model without an activation function, or set the last layer of the multi-classification model to use a sigmoid function or a logits function.
[0124] A multi-intent recognition model acquisition unit is configured to train the multi-classification model using the model training data, and acquire a multi-intent recognition model.
[0125] In an embodiment, the multi-intelligent outbound call device further comprises:
[0126] A rule engine matching unit is configured to use an intent rule engine to perform matching processing on the target key data, and determine whether the intent rule engine can match a first intent corresponding to the target key data.
[0127] A first rule matching processing unit is configured to determine the first intent as a target intent if the first intent can be matched.
[0128] A second rule matching processing unit is configured to use a multi-intent recognition model corresponding to the target domain to recognize the target key data and determine at least one target intent if the first intent cannot be matched.
[0129] In an embodiment, the multi-intelligent outbound call device further comprises:
[0130] A retrieval analysis matching unit is configured to use a retrieval analysis model to perform matching processing on the target key data, and determine whether the retrieval analysis model can match a second intent corresponding to the target key data.
[0131] The first search matching processing unit is configured to determine the second intention as the target intention if the second intention can be matched.
[0132] The second search matching processing unit is configured to use a multi-intention recognition model corresponding to the target field to recognize the target key data and determine at least one target intention if the second intention cannot be matched.
[0133] In an embodiment, the search analysis matching unit comprises:
[0134] The ES configuration data acquisition subunit is configured to use an ES search algorithm to preliminarily search the target key data and acquire a plurality of ES configuration data, each of which corresponds to a configuration intention;
[0135] The target similarity acquisition subunit is configured to use a semantic matching model to analyze the similarity between the plurality of ES configuration data and the target key data and acquire a target similarity corresponding to each of the ES configuration data;
[0136] The first similarity processing subunit is configured to determine, if there is a target similarity greater than a preset similarity, a configuration intention corresponding to an ES configuration data corresponding to the target similarity as a second intention corresponding to the target key data and determine that the search analysis model can match the second intention.
[0137] The second similarity processing subunit is configured to determine that the search analysis model cannot match the second intention if all the target similarities are not greater than the preset similarity.
[0138] The specific limitations of the intelligent outbound device can refer to the limitations of the intelligent outbound method in the foregoing, which will not be repeated here. Each module in the intelligent outbound device described above can be realized by software, hardware, and a combination thereof in whole or in part. The modules described above can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory in the computer device in software form, so as to be called and executed by the processor to perform the operations corresponding to each module.
[0139] In one embodiment, a computer device is provided, which can be a server, and an internal structure diagram of the computer device can be as shown in Figure 10As shown in the figure. The computer device includes a processor, a memory, a network interface and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium, an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The database of the computer device is used to store the data used or generated in the execution of the intelligent outbound process. The network interface of the computer device is used to communicate with the terminal outside through the network connection. The computer program is executed by the processor to implement an intelligent outbound method.
[0140] In an embodiment, a computer device is provided, including a memory, a processor and a computer program stored on the memory and executable on the processor, and the processor executes the computer program to implement the intelligent outbound method in the above-mentioned embodiments, for example Figure 2 As shown in the figure, or Figures 3 to 8 As shown in the figure, for the sake of brevity, the functions of the customer voice data acquisition module 901, the customer text data acquisition module 902, the target key data acquisition module 903, the target intent determination module 904 and the outbound module 905 are not repeated here. Or, the processor executes the computer program to implement the functions of each module / unit in this embodiment of the intelligent outbound device, for example Figure 9 As shown in the figure, for the sake of brevity, the functions of the customer voice data acquisition module 901, the customer text data acquisition module 902, the target key data acquisition module 903, the target intent determination module 904 and the outbound module 905 are not repeated here. Or, the processor executes the computer program to implement the functions of each module / unit in this embodiment of the intelligent outbound device, for example
[0141] In an embodiment, a computer readable storage medium is provided, and the computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the intelligent outbound method in the above-mentioned embodiments, for example Figure 2 As shown in the figure, or Figures 3 to 8 As shown in the figure, for the sake of brevity, the functions of the customer voice data acquisition module 901, the customer text data acquisition module 902, the target key data acquisition module 903, the target intent determination module 904 and the outbound module 905 are not repeated here. Or, the processor executes the computer program to implement the functions of each module / unit in this embodiment of the intelligent outbound device, for example Figure 9 As shown in the figure, for the sake of brevity, the functions of the customer voice data acquisition module 901, the customer text data acquisition module 902, the target key data acquisition module 903, the target intent determination module 904 and the outbound module 905 are not repeated here. Or, the processor executes the computer program to implement the functions of each module / unit in this embodiment of the intelligent outbound device, for example
[0142] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments of each method. Any reference to memory, storage, database or other medium used in each embodiment provided by the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0143] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is exemplified, and in actual application, the above-mentioned functions can be completed by different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.
[0144] The above-mentioned embodiments are only used to illustrate the technical solutions of the present application, but not to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that the technical solutions recorded in the foregoing embodiments can be modified, or some technical features can be replaced by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.
Claims
1. A smart outbound calling method, characterized in that, include: Acquire customer voice data in the target area; The customer's voice data is subjected to speech recognition and semantic parsing to obtain the customer's text data; The customer text data is preprocessed to obtain key target data; A retrieval analysis model is used to match the target key data, and it is determined whether the retrieval analysis model can match the second intent corresponding to the target key data. Specifically, the matching process includes: first, using an ES retrieval algorithm to perform a preliminary retrieval of the target key data to obtain multiple ES configuration data, each ES configuration data corresponding to a configuration intent; the ES configuration data is data pre-stored in an ES database containing all or part of the keywords in the target keywords; the configuration intent refers to the intent pre-stored in the ES database that matches the ES configuration data; then, a semantic matching model is used to perform similarity analysis on the multiple ES configuration data and the target key data to obtain the target similarity corresponding to each ES configuration data; if there is a target similarity greater than a preset similarity, the configuration intent corresponding to the ES configuration data with the target similarity is determined as the second intent corresponding to the target key data, and it is determined that the retrieval analysis model can match the second intent; if all target similarities are not greater than the preset similarity, it is determined that the retrieval analysis model cannot match the second intent. If the second intent can be matched, then the second intent is determined as the target intent; If the second intent cannot be matched, then the multi-intent recognition model corresponding to the target domain is used to identify the target key data and determine at least one target intent; After obtaining the target intent, it is determined whether there is a corresponding intent response text based on the target intent. If there is a corresponding intent response text, a text-to-speech tool is used to convert the intent response text into speech, obtain the speech response data, and then make an outbound call based on the speech response data. If there is no corresponding intent response text, the call is switched to different outbound call processing methods according to the current node. The outbound call processing methods include at least one of the following: transfer to human agent, please call back later, and explanation and reassurance.
2. The intelligent outbound calling method as described in claim 1, characterized in that, The step of preprocessing the customer text data to obtain target key data includes: The customer text data is segmented into words to obtain the original key data; The original key data is processed by removing stop words and punctuation marks to obtain the target key data.
3. The intelligent outbound calling method as described in claim 1, characterized in that, The multi-intent recognition model corresponding to the target domain includes multiple single-intent recognition models corresponding to the target domain; The step of using a multi-intent recognition model corresponding to the target domain to identify the target key data and determine at least one target intent includes: Multiple single-intent recognition models corresponding to the target domain are used to analyze the target key data to obtain the original intent output by the multiple single-intent recognition models and the recognition probability corresponding to the original intent. At least one original intent whose recognition probability is greater than a preset probability is identified as at least one target intent.
4. The intelligent outbound calling method as described in claim 1, further comprising, before acquiring customer voice data in the target domain: Acquire model training data corresponding to the same target domain, wherein the model training data includes training text data and multi-intent labels; Obtain multiple single-intent recognition models, construct a multi-classification model based on the multiple single-intent recognition models, and set the last layer of the multi-classification model to use no activation function, sigmoid activation function, or logits activation function. The multi-classification model is trained using the model training data to obtain a multi-intent recognition model.
5. The intelligent outbound calling method as described in claim 1, before identifying the target key data and determining at least one target intent using a multi-intent recognition model corresponding to the target domain, the intelligent outbound calling method further includes: An intent rule engine is used to match the target key data, and it is determined whether the intent rule engine can match the first intent corresponding to the target key data. If the first intent can be matched, then the first intent is determined as the target intent; If the first intent cannot be matched, then the multi-intent recognition model corresponding to the target domain is used to identify the target key data and determine at least one target intent.
6. An intelligent outbound calling device, characterized in that, include: The customer voice data acquisition module is used to acquire customer voice data in the target area. The customer text data acquisition module is used to perform speech recognition and semantic parsing on the customer voice data to acquire customer text data; The target key data acquisition module is used to perform text preprocessing on the customer text data to acquire target key data; The target intent determination module is used to perform matching processing on the target key data using a retrieval analysis model, and determine whether the retrieval analysis model can match the second intent corresponding to the target key data. Specifically, the process of using a retrieval analysis model to match the target key data and determine whether the retrieval analysis model can match the second intent corresponding to the target key data includes: first, using an ES retrieval algorithm to perform a preliminary retrieval on the target key data to obtain multiple ES configuration data, each ES configuration data corresponding to a configuration intent; the ES configuration data is data pre-stored in an ES database containing all or part of the keywords in the target keywords; the configuration intent refers to the intent pre-stored in the ES database that matches the ES configuration data; then, using... A semantic matching model performs similarity analysis on multiple ES configuration data and the target key data to obtain the target similarity corresponding to each ES configuration data. If there is a target similarity greater than a preset similarity, the configuration intent corresponding to the ES configuration data with the target similarity is determined as a second intent corresponding to the target key data, and the retrieval analysis model is determined to be able to match the second intent. If all target similarities are not greater than the preset similarity, the retrieval analysis model is determined to be unable to match the second intent. If the second intent can be matched, the second intent is determined as the target intent. If the second intent cannot be matched, a multi-intent recognition model corresponding to the target domain is used to identify the target key data and determine at least one target intent. The outbound call module is used to determine whether there is an intent response text corresponding to the target intent after obtaining the target intent. If there is an intent response text corresponding to the target intent, a text-to-speech tool is used to convert the intent response text into speech, obtain speech response data, and then make an outbound call based on the speech response data. If there is no intent response text corresponding to the target intent, the module jumps to different outbound call processing methods according to the current node. The outbound call processing methods include at least one of the following: transfer to human agent, ask to call back later, and provide explanation and reassurance.
7. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the intelligent outbound calling method as described in any one of claims 1 to 5.
8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the intelligent outbound calling method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Outbound method and device based on intention recognition
CN111949784A
Intention recognition method and device, equipment and storage medium
CN114117037A