Call request answering method and device, computer equipment and readable storage medium
By combining the intention recognition model with robot and agent response, the automatic response system cannot accurately respond to user questions and the limited processing volume of agent response requests is solved, and efficient and accurate response to call requests is achieved.
Patent Information
- Application Number
- CN202510316816.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-07-04
AI Technical Summary
The existing automatic response system cannot accurately respond to user questions when processing call requests, and the limited number of requests to be answered by agents makes it impossible to reach a large number of call requests.
Through the trained intention recognition model, the user intention of the target call request is identified, and combined with other user intentions, the response method is determined from the preset response method, and robot response or agent response is preferred to improve the accuracy and reach efficiency of the response content.
The accuracy and reach efficiency of call requests are improved, and the illusion generation caused by low robot response accuracy is reduced. The problem of disconnecting from actual response speech caused by the low robot response accuracy is optimized.
Smart Images

Figure CN120263904A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a method, device, computer device, and computer-readable storage medium for answering call requests. Background Art
[0002] With the development of society and the progress of technology, people's demand for product or service consultation is increasing day by day. For example, in the financial business scenario, users often need to make call consultations about insurance products such as car insurance, and in the digital medical scenario, users often need to make call consultations about medical service processes. To meet the processing requirements of a large number of call requests, an automatic answering system (i.e., robot answering) is set up in the related technology to handle call requests.
[0003] However, the automatic answering system mainly includes rule-based intelligent answering methods, template matching-based intelligent answering methods, deep learning-based intelligent answering methods, and large language model-based intelligent answering methods. Although the automatic answering system can quickly complete the processing of call requests, sometimes the response content of call requests cannot accurately answer users' questions. If agent answering is used, the accuracy of the response content of call requests can be improved, but due to the limited processing capacity of agent answering, a large number of call requests cannot be reached. Summary of the Invention
[0004] This application provides a method, device, computer device, and computer-readable storage medium for answering call requests, belonging to the field of artificial intelligence technology, which can improve the reach efficiency of call requests and the accuracy of the response content of call requests.
[0005] In a first aspect, this application provides a method for answering call requests, the method comprising:
[0006] Based on the call content of the target call request in the business answering system, perform intent recognition through a trained intent recognition model to obtain the target user intent of the target call request;
[0007] Obtain other user intents of other call requests in the business answering system;
[0008] Based on the target user intent and the other user intents, determine the answering mode of the target call request from preset answering modes, where the preset answering modes include agent answering and robot answering;
[0009] Answer the target call request according to the answering mode of the target call request.
[0010] In a second aspect, this application provides a device for answering call requests, the device for answering call requests comprising:
[0011] An identification unit, configured to perform intent identification on the call content of a target call request in a service response system through a trained intent recognition model, so as to obtain the target user intent of the target call request;
[0012] An acquisition unit, configured to acquire the other user intents of other call requests in the service response system;
[0013] A determination unit, configured to determine the response mode of the target call request from preset response modes based on the target user intent and the other user intents, where the preset response modes include agent response and robot response;
[0014] A response unit, configured to respond to the target call request according to the response mode of the target call request.
[0015] In a third aspect, the present application further provides a computer device, where the computer device includes a memory and a processor; the memory is used to store a computer program; the processor is used to execute the computer program and implement the call request response method when executing the computer program.
[0016] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored, and the computer program is loaded by a processor to execute the call request response method.
[0017] In the present application, through a trained intent recognition model, intent recognition is performed on the call content of a target call request in a service response system to obtain the target user intent of the target call request. Based on the target user intent and other user intents, the response mode of the target call request is determined from preset response modes. In this way, when the accuracy of the robot response for the target user intent is relatively high, the robot can be preferentially used to respond to the target call request. When the accuracy of the robot response for the target user intent is relatively low, the agent can be preferentially used to respond to the target call request. Thus, the robot response and the agent response can be combined to process call requests, which can improve the reach efficiency of call requests and the accuracy of response content to a certain extent. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0019] Figure 1 It is a schematic flowchart of a call request response method provided by an embodiment of the present application;
[0020] Figure 2 It is a schematic diagram of the principle structure of the trained intent recognition model provided in the embodiments of the present application;
[0021] Figure 3 It is another schematic diagram of the principle structure of the trained intent recognition model provided in the embodiments of the present application;
[0022] Figure 4 It is a schematic diagram for explaining the overall process of call request response in the embodiments of the present application;
[0023] Figure 5 It is a schematic diagram of the training process of the intent recognition model provided in the embodiments of the present application;
[0024] Figure 6 It is a schematic diagram of the structure of an embodiment of the call request response device provided in the embodiments of the present application;
[0025] Figure 7 It is a schematic block diagram of the structure of a computer device provided in the embodiments of the present application. Detailed implementation manners
[0026] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.
[0027] The flowcharts shown in the accompanying drawings are only illustrative examples, and do not necessarily include all the contents and operations / steps, nor do they necessarily need to be executed in the described order. For example, some operations / steps can be decomposed, combined, or partially merged, so the actual execution order may change according to the actual situation.
[0028] In the description of the embodiments of the present application, it should be understood that the terms "first" and "second" are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of the described features. In the description of the embodiments of the present application, "a plurality of" means two or more, unless otherwise specifically defined.
[0029] The following description is provided to enable any person skilled in the art to implement and use this application. In the following description, details are set forth for purposes of explanation. It should be understood that those of ordinary skill in the art can recognize that this application can be implemented without using these specific details. In other instances, well-known processes are not elaborated in detail to avoid obscuring the description of the embodiments of this application with unnecessary details. Therefore, this application is not intended to be limited to the embodiments shown, but is to be accorded the widest scope consistent with the principles and features disclosed in the embodiments of this application.
[0030] Embodiments of this application provide a call request response method, apparatus, computer device, and computer-readable storage medium.
[0031] The execution subject of the call request response method in the embodiments of this application can be the call request response apparatus provided in the embodiments of this application or the computer device provided in the embodiments of this application. Among them, the call request response apparatus can be implemented in a hardware or software manner.
[0032] The following describes some embodiments of this application in detail with reference to the accompanying drawings. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0033] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a call request response method provided by the embodiments of this application. The call request response method includes steps 101 to 104, where:
[0034] 101. Based on the call content of the target call request in the service response system, perform intent recognition through a trained intent recognition model to obtain the target user intent of the target call request.
[0035] Among them, the service response system is a service system that can provide call response services. For example, in the insurance business scenario, the service response system can be an insurance product trading system. Another example is that in the wealth management business scenario, the service response system can be a wealth management product trading system. Still another example is that in the digital medical scenario, the service response system can be a digital medical service platform.
[0036] Among them, the target call request refers to the call request for which the response method needs to be determined.
[0037] Exemplarily, if the call content is in audio format, first, the format of the call content in the target call request can be converted. The call content in audio format (such as call voice including multiple statements) is converted into call content in text format (such as converted into call text including multiple statements). Then, the trained intent recognition model first extracts text features from the call content in text format to obtain the target text features of the target call request, and then uses the target text features for intent recognition to obtain the target user intent of the target call request.
[0038] To better understand this embodiment, the trained intent recognition model in this embodiment will be introduced first. Please refer to Figure 2 and Figure 3 , Figure 2 which is a schematic diagram of the principle structure of the trained intent recognition model provided in the embodiment of the present application. Figure 3 which is another schematic diagram of the principle structure of the trained intent recognition model provided in the embodiment of the present application. The trained intent recognition model may include an input layer, a text feature extraction layer, and an intent recognition layer. Further, the trained intent recognition model may also include an audio feature extraction layer. The trained intent recognition model can be obtained by training based on an initial intent recognition model.
[0039] As Figure 2 shown, taking the trained intent recognition model may include a text feature extraction layer, an audio feature extraction layer, and an intent recognition layer as an example, the working principles of each module are as follows:
[0040] 1. Input layer: On the one hand, it is used to perform embedding representation on the call content in text format to convert the call content in text format into a format that can be processed by the model, so that the subsequent text feature extraction layer can process the call content in text format. The call content in text format (such as call text including multiple statements) undergoes preprocessing by the input layer, and each statement is preprocessed to construct a standardized representation: S = R x*K , where x represents the length of the standardized statement, and K represents the dimension of the pre-trained word vector. On the other hand, it is used to represent the call content in audio format in a specified format, such as converting the call content in audio format into a spectrogram (such as a spectrogram or Mel Frequency Cepstral Coefficients MFCC)), so as to convert the call content in audio format into a format that can be processed by the model, facilitating the subsequent audio feature extraction layer to use the call content spectrogram for further feature extraction processing.
[0041] 2. The text feature extraction layer is used to extract features from the text-formatted call content after embedding representation to obtain the text features of the call request. Exemplarily, the text feature extraction layer can adopt a convolutional neural network structure. Using the convolutional neural network, feature extraction operations such as convolution, pooling, and activation are performed on the text-formatted call content (such as a call text including multiple sentences), and finally the text features of the text-formatted call content (such as a call text including multiple sentences) are obtained as the text features of the call request. Among them, each sentence is first preprocessed by the input layer to construct a standardized representation, and then passed through the text feature extraction layer. The text features of the text-formatted call content are obtained by extracting features from each sentence of the call text through filters of different dimensions (i = [3, 4, 5]). In this way, the text feature extraction layer can learn rich feature representations of the call request from the text-formatted call content. The intent recognition model using the text feature extraction layer can use richer text information for intent prediction, improving the accuracy of intent recognition of the call request.
[0042] 3. The audio feature extraction layer is used to extract features from the call content spectrogram output by the input layer to obtain the audio features of the call request. Exemplarily, the audio feature extraction layer can adopt a recurrent neural network (RNN) structure. Using the recurrent neural network, feature extraction operations are performed on the call content spectrogram, and finally the audio features of the audio-formatted call content are obtained as the audio features of the call request. In this way, the audio feature extraction layer can learn rich feature representations of the call request from the audio-formatted call content, effectively capturing the temporal dependence relationship of the audio signals in the audio-formatted call content. The intent recognition model using the audio feature extraction layer can use richer audio information for intent prediction, improving the accuracy of intent recognition of the call request.
[0043] 4. The intent recognition layer is used to splice and fuse the text features output by the text feature extraction layer and the audio features output by the audio feature extraction layer to obtain the fused features of the call request; intent prediction is performed according to the fused features of the call request to obtain the user intent of the call request. Through a series of processing steps such as the text feature extraction layer and the audio feature extraction layer, the intent recognition model fully mines and integrates the purpose of the call request, laying a solid foundation for the intent recognition of the call request; then, through the intent recognition layer for intent recognition, the accuracy of intent recognition can be improved.
[0044] As Figure 3 shown, taking the trained intent recognition model that can include an input layer, a text feature extraction layer, and an intent recognition layer as an example, the working principles of each module are as follows:
[0045] 1. Input layer, which is used to perform embedding representation on the text - formatted call content to convert the text - formatted call content into a format that can be processed by the model, so that the subsequent text feature extraction layer can process the text - formatted call content. The text - formatted call content (such as a call text including multiple sentences) is pre - processed by the input layer, and each sentence is pre - processed to construct a standardized representation: S = R x*K , where x represents the length of the standardized sentence, and K represents the dimension of the pre - trained word vector.
[0046] 2. Text feature extraction layer, which is used to extract features from the text - formatted call content output by the input layer to obtain the text features of the call request. Exemplarily, the text feature extraction layer can adopt a convolutional neural network structure. Using the convolutional neural network, feature extraction operations such as convolution, pooling, and activation are performed on the text - formatted call content (such as a call text including multiple sentences), and finally the text features of the text - formatted call content (such as a call text including multiple sentences) are obtained as the text features of the call request. Among them, each sentence is first pre - processed by the input layer to construct a standardized representation, and then passes through the text feature extraction layer. Filters with different dimensions (i = [3, 4, 5]) are used to extract features from each sentence of the call text to obtain the text features of the text - formatted call content. In this way, this text feature extraction layer can learn rich feature representations of the call request from the text - formatted call content. The intent recognition model using the text feature extraction layer can use richer text information for intent prediction, improving the accuracy of intent recognition of the call request.
[0047] 3. Intent recognition layer, which is used to perform intent prediction based on the text features of the call request to obtain the user intent of the call request.
[0048] There are multiple implementation methods for step 101. Exemplarily, they include:
[0049] (1) In some embodiments, the trained intent recognition model includes an input layer, a text feature extraction layer, and an intent recognition layer. At this time, step 101 can specifically include: through the input layer of the trained intent recognition model, perform embedding representation on the text - formatted call content of the target call request to obtain the embedded text - formatted call content; through the text feature extraction layer of the trained intent recognition model, extract features from the embedded text - formatted call content of the target call request to obtain the target text features of the target call request, where the text feature extraction layer is learned based on the text feature loss between the first call request samples of the first service type and the second call request samples of the second service type, and the first service type is the service type of the target call request; through the intent recognition layer of the trained intent recognition model, perform intent prediction based on the target text features to obtain the target user intent.
[0050] (2) In some embodiments, the trained intent recognition model includes a text feature extraction layer, an audio feature extraction layer, and an intent recognition layer. In this case, step 101 may specifically include: embedding the text format call content of the target call request through the input layer of the trained intent recognition model to obtain the text format call content after the embedded representation; converting the audio format call content of the target call request through the input layer of the trained intent recognition model to obtain a call content spectrum of the target call request; extracting features from the text format call content after the embedded representation of the target call request through the text feature extraction layer of the trained intent recognition model to obtain the target text feature of the target call request; extracting features from the call content spectrum of the target call request through the audio feature extraction layer of the trained intent recognition model to obtain the target audio feature of the target call request; splicing and fusing the call content spectrum of the target call request and the target audio feature through the intent recognition layer of the trained intent recognition model to obtain the fused feature of the target call request; and predicting intent based on the fused feature of the target call request to obtain the target user intent.
[0051] The text feature extraction layer is obtained based on the text feature loss learning between the first call request sample of the first business type and the second call request sample of the second business type, and the first business type is the business type of the target call request. In this way, on the first hand, during the training phase of the trained intent recognition model, the call request samples of the first business type and the call request samples of the second business type can be used for training at the same time, thereby improving the richness of the training data and avoiding the problem of low recognition accuracy of the trained intent recognition model due to less sample data of the first business type. Secondly, since the text feature extraction layer is learned based on the text feature loss between the first call request sample of the first business type and the second call request sample of the second business type, it is possible to learn the feature differences between the samples of the target domain (i.e., the first business type) and the samples of the source domain (i.e., the second business type), so that the text feature extraction layer can learn more consistent feature expressions in the sample data of two different business types, thereby reducing the problem of text feature extraction mismatch of call requests of the first business type caused by learning with the second call request samples of the second business type, thereby reducing the problem of inapplicability of intent recognition of the first business type caused by learning with the second call request samples of the second business type, thereby improving the accuracy of intent recognition, thereby improving the matching degree of the target call request's response speech when using the target user's intention to generate the response speech, and reducing the problem of generating response speech that is out of touch with reality due to hallucinations.
[0052] Among them, the text feature extraction layer is learned based on the text feature loss between the first call request sample of the first service type and the second call request sample of the second service type, and the first service type is the service type of the target call request. In this way, on the one hand, during the training stage of the trained intent recognition model, call request samples of the first service type and call request samples of the second service type can be used for training at the same time, improving the richness of training data and avoiding the problem of low recognition accuracy of the trained intent recognition model due to the small number of sample data of the first service type. On the other hand, since the audio feature extraction layer is learned based on the audio feature loss between the first call request sample of the first service type and the second call request sample of the second service type, therefore, the feature differences between the samples in the target domain (i.e., the first service type) and the samples in the source domain (i.e., the second service type) can be learned, enabling the audio feature extraction layer to learn more consistent feature expressions from the sample data of two different service types, reducing the problem of non-fitting in the audio feature extraction of call requests of the first service type caused by learning using the second call request sample of the second service type, and further reducing the problem of inapplicability in the intent recognition of the first service type caused by learning using the second call request sample of the second service type, thereby improving the intent recognition accuracy, and further improving the matching degree of the response speech for the target call request when generating the response speech using the target user intent, and reducing the problem of generating hallucinated and unrealistic response speech.
[0053] 102. Obtain other user intents of other call requests in the service response system.
[0054] Among them, other call requests refer to call requests in all call requests to be responded in the service response system except the target call request.
[0055] The obtaining method of "other user intents of other call requests" is similar to the obtaining method of "target user intent of the target call request", and specific reference can be made to the relevant description in step 101, which will not be elaborated here.
[0056] 103. Based on the target user intent and the other user intents, determine the response method for the target call request from the preset response methods.
[0057] Among them, the preset response methods include agent response and robot response.
[0058] In some embodiments, step 103 may specifically include the following steps 1031 to 1033:
[0059] 1031. Based on the target user intent and the other user intents, determine the processing order of the target call request among all call requests to be processed in the service response system.
[0060] Among them, the processing order of the target call request is negatively correlated with the accuracy of the robot response to the target user intention, that is, the higher the accuracy of the robot response to the target user intention, the later the processing order of the target call request; on the contrary, the lower the accuracy of the robot response to the target user intention, the earlier the processing order of the target call request. In this way, when the accuracy of the robot response to the target user intention is relatively high, the robot is preferentially used to respond to the target call request, and when the accuracy of the robot response to the target user intention is relatively low, the agent is preferentially used to respond to the target call request. Thus, the robot response and the agent response can be combined to process the call request, which can improve the reach efficiency of the call request and the accuracy of the response content of the call request to a certain extent.
[0061] Exemplarily, the accuracy of the robot response to the target user intention and the accuracy of the robot response to other user intentions can be obtained; based on the accuracy of the robot response to the target user intention and the accuracy of the robot response to other user intentions, the initial sorting of the target call request among all the call requests to be processed in the response system is determined; the user call activity of the target call request is obtained; according to the user call activity of the target call request, the initial sorting among all the call requests to be processed in the response system is adjusted to obtain the processing order of the target call request among all the call requests to be processed in the service response system.
[0062] 1032. Obtain the agent load capacity of the service response system.
[0063] Exemplarily, the number of artificial customer service agents currently assigned by the service response system to the first service type (i.e., the service type of the target call request) can be used as the agent load capacity of the service response system.
[0064] 1033. Determine the response method of the target call request according to the agent load capacity and the processing order.
[0065] Exemplarily, if the processing order is less than or equal to the agent load capacity, the agent response is used as the response method of the target call request; or, if the processing order is greater than the agent load capacity, the robot response is used as the response method of the target call request. In this way, the call requests can be sorted according to the user intention, and the call requests with a higher processing order can preferentially use the agent response, which can filter out the list of invalid call requests or the list of low-quality call requests, improve the contact efficiency of the agent with the call request list, and reduce the problems of time-consuming, laborious and low efficiency in the agent's contact with the call request list.
[0066] 104. Respond to the target call request according to the response method of the target call request.
[0067] For example, as Figure 4 shown, after a user initiates a target call request in a business response system, the call is first answered by a robot, and the call content of the target call request is collected by the robot (such as call voice, that is, call content in audio format); then, the audio format call content is converted into text format call content through speech recognition technology (ASR); then, through an error correction algorithm based on a deep model (FASPell), the possible errors in the text format call content are corrected; then, through a trained intent recognition model, intent recognition is performed based on the call content of the target call request to obtain the target user intent of the target call request; then, according to the other user intents of other call requests in the business response system and the target user intent of the target call request, the processing order of the target call request among all the call requests to be processed in the business response system is determined, and the response method for the target call request is determined from the preset response methods. For example, if the processing order of the target call request is greater than the seat capacity, the robot response is used as the response method for the target call request, and the robot customer service generates a response script to respond to the target call request; if the processing order of the target call request is less than or equal to the seat capacity, the seat response is used as the response method for the target call request, the target call request is connected to the seat customer service, and the seat customer service uses the corresponding response script to respond to the target call request.
[0068] Furthermore, when the robot response is used as the response method for the target call request, the response script can be generated based on the identified target user intent. Since intent recognition through the trained intent recognition model can improve the accuracy of intent recognition, using the target user intent to generate the response script can also improve the matching degree of the response script for the target call request to a certain extent and reduce the problem of generating unrealistic response scripts due to hallucinations.
[0069] Please refer to Figure 5 , Figure 5 which is a schematic diagram of a training process of the intent recognition model provided in an embodiment of the present application. Exemplarily, the trained intent recognition model can be obtained through the following steps 501 to 505:
[0070] 501. Obtain an initial intent recognition model trained with call request samples based on a second service type.
[0071] In an embodiment of the present application, considering that the call request samples of the first service type (i.e., the service type of the target call request) are relatively few, in order to enrich the sample data and improve the accuracy of intent recognition of the trained intent recognition model, the call request samples of the first service type and the call request samples of the second service type are used for training at the same time. The training of the intent recognition model includes two stages:
[0072] In the first stage, using the call request samples of the second service type, with the goal of "minimizing the error between the predicted intention and the actual intention of the call request samples of the second service type", train the preset intention recognition model to obtain a preliminary intention recognition model.
[0073] In the second stage, using the first call request samples (i.e., call request samples of the first service type) and the second call request samples (i.e., call request samples of the second service type) with the same intention, with the goal of "minimizing the text feature loss between the first call request samples and the second call request samples", train the preliminary intention recognition model to obtain a trained intention recognition model.
[0074] 502. Obtain the first call request samples of the first service type and the second call request samples of the second service type.
[0075] Among them, the labeled intention of the first call request samples is the same as the labeled intention of the second call request samples.
[0076] 503. Through the text feature extraction layer of the initial intention recognition model, extract text features from the first call request samples and the second call request samples respectively to obtain the first text features of the first call request samples and the second text features of the second call request samples.
[0077] 504. Obtain the loss between the first text features and the second text features as the text feature loss between the first call request samples and the second call request samples.
[0078] 505. Train the initial intention recognition model based on the text feature loss to obtain the trained intention recognition model.
[0079] There are various implementation manners for step 505 (for example, the initial intention recognition model can be trained using at least one of the text feature loss, audio feature loss, first intention loss, and second intention loss). Exemplarily, it includes the following cases ① to ④:
[0080] (1) In some embodiments, the intention recognition model includes an input layer, a text feature extraction layer, and an intention recognition layer.
[0081] ① Train using the text feature loss and the first intention loss.
[0082] At this time, the method further includes: through the intent recognition layer of the initial intent recognition model, performing intent prediction based on the first text feature to obtain the predicted intent of the first call request sample; obtaining the loss between the predicted intent of the first call request sample and the labeled intent of the first call request sample as the first intent loss of the initial intent recognition model. Step 505 may specifically include: training the initial intent recognition model based on the first intent loss and the text feature loss to obtain the trained intent recognition model. Specifically, the model parameters of the text feature extraction layer of the initial intent recognition model can be adjusted according to the text feature loss, and the model parameters of the prediction layer of the initial intent recognition model can be adjusted according to the first intent loss until the stop training condition is met, and the initial intent recognition model with adjusted parameters is used as the trained intent recognition model.
[0083] ② Train using the text feature loss, the first intent loss, and the second intent loss.
[0084] At this time, the method further includes: through the intent recognition layer of the initial intent recognition model, performing intent prediction based on the first text feature to obtain the predicted intent of the first call request sample; obtaining the loss between the predicted intent of the first call request sample and the labeled intent of the first call request sample as the first intent loss of the initial intent recognition model; through the intent recognition layer of the initial intent recognition model, performing intent prediction based on the second text feature to obtain the predicted intent of the second call request sample; obtaining the loss between the predicted intent of the second call request sample and the labeled intent of the second call request sample as the second intent loss of the initial intent recognition model. Step 505 may specifically include: training the initial intent recognition model based on the first intent loss, the second intent loss, and the text feature loss to obtain the trained intent recognition model. Specifically, the model parameters of the text feature extraction layer of the initial intent recognition model can be adjusted according to the text feature loss, the model parameters of the prediction layer of the initial intent recognition model can be adjusted according to the first intent loss, and the model parameters of the prediction layer of the initial intent recognition model can be adjusted according to the second intent loss until the stop training condition is met, and the initial intent recognition model with adjusted parameters is used as the trained intent recognition model.
[0085] (2) In some embodiments, the intent recognition model includes a text feature extraction layer, an audio feature extraction layer, and an intent recognition layer.
[0086] ③ Train using the text feature loss, the audio feature loss, and the first intent loss.
[0087] At this time, the method further includes: respectively extracting audio features of the first call request sample and the second call request sample through the audio feature extraction layer of the initial intent recognition model to obtain a first audio feature of the first call request sample and a second audio feature of the second call request sample; performing intent prediction based on the first text feature and the first audio feature through the intent recognition layer of the initial intent recognition model to obtain a predicted intent of the first call request sample; obtaining a loss between the predicted intent of the first call request sample and the labeled intent of the first call request sample as a first intent loss of the initial intent recognition model; and obtaining a loss between the first audio feature and the second audio feature as an audio feature loss between the first call request sample and the second call request sample. Step 505 may specifically include: training the initial intent recognition model based on the first intent loss, the text feature loss, and the audio feature loss to obtain the trained intent recognition model. Specifically, the model parameters of the text feature extraction layer of the initial intent recognition model may be adjusted according to the text feature loss, the model parameters of the audio feature extraction layer of the initial intent recognition model may be adjusted according to the audio feature loss, and the model parameters of the prediction layer of the initial intent recognition model may be adjusted according to the first intent loss. Until the stop training condition is met, the initial intent recognition model with adjusted parameters is used as the trained intent recognition model.
[0088] ④ Train using text feature loss, audio feature loss, first intent loss, and second intent loss.
[0089] At this time, the method further includes: respectively extracting audio features of the first call request sample and the second call request sample through the audio feature extraction layer of the initial intent recognition model to obtain a first audio feature of the first call request sample and a second audio feature of the second call request sample; performing intent prediction on the basis of the first text feature and the first audio feature through the intent recognition layer of the initial intent recognition model to obtain a predicted intent of the first call request sample; performing intent prediction on the basis of the second text feature and the second audio feature through the intent recognition layer of the initial intent recognition model to obtain a predicted intent of the second call request sample; obtaining a loss between the predicted intent of the first call request sample and the labeled intent of the first call request sample as a first intent loss of the initial intent recognition model; obtaining a loss between the predicted intent of the second call request sample and the labeled intent of the second call request sample as a second intent loss of the initial intent recognition model; obtaining a loss between the first audio feature and the second audio feature as an audio feature loss between the first call request sample and the second call request sample. Step 505 may specifically include: training the initial intent recognition model based on the first intent loss, the second intent loss, the text feature loss, and the audio feature loss to obtain the trained intent recognition model. Specifically, the model parameters of the text feature extraction layer of the initial intent recognition model may be adjusted according to the text feature loss, the model parameters of the audio feature extraction layer of the initial intent recognition model may be adjusted according to the audio feature loss, the model parameters of the prediction layer of the initial intent recognition model may be adjusted according to the first intent loss, and the model parameters of the prediction layer of the initial intent recognition model may be adjusted according to the second intent loss. Until the stop training condition is met, the initial intent recognition model with adjusted parameters is used as the trained intent recognition model.
[0090] From the above content, it can be seen that, firstly, through the trained intent recognition model, the intent recognition is performed based on the call content of the target call request in the business answering system to obtain the target user intent of the target call request. Based on the target user intent and other user intents, the answering method of the target call request is determined from the preset answering methods. In this way, when the robot answering accuracy of the target user intent of the target call request is relatively high, the robot can be used to answer the target call request first. When the robot answering accuracy of the target user intent of the target call request is relatively low, the agent can be used to answer the target call request first. In this way, the robot answer and the agent answer can be combined to process the call request, which can improve the contact efficiency of the call request and the accuracy of the answer content of the call request to a certain extent. Secondly, through the trained intent recognition model, the user intent recognition accuracy of the target call request can be improved, and then the reliability of the processing order of the target call request can be improved, so that the call request with a higher processing order can be answered by the agent first, so that the invalid call request list or the low-quality call request list can be filtered, the contact efficiency of the agent's call request list can be improved, and the time-consuming, labor-intensive and inefficient contact of the agent with the call request list can be reduced. Thirdly, by using a trained intent recognition model to identify intent, the accuracy of identifying the user intent of the target call request can be improved. Therefore, using the target user intent to generate response words can also improve the matching degree of the response words of the target call request to a certain extent, and reduce the problem of generating response words that are out of touch with reality due to hallucinations.
[0091] In addition, in order to better implement the call request answering method in the embodiment of the present application, based on the call request answering method, the embodiment of the present application also provides a call request answering device, such as Figure 6 FIG. 1 is a schematic diagram of a structure of an embodiment of a call request answering device provided in an embodiment of the present application. The call request answering device 600 includes:
[0092] The recognition unit 601 is used to perform intent recognition based on the call content of the target call request in the service answering system through the trained intent recognition model to obtain the target user intent of the target call request;
[0093] An acquisition unit 602 is used to acquire other user intentions of other call requests in the service answering system;
[0094] A determining unit 603 is used to determine an answering method for the target call request from preset answering methods based on the target user intention and the other user intentions, wherein the preset answering methods include agent answering and robot answering;
[0095] The answering unit 604 is used to answer the target call request according to the answering mode of the target call request.
[0096] In some embodiments, the recognition unit 601 is specifically configured to:
[0097] Extract features from the text format call content of the target call request through the text feature extraction layer of the trained intent recognition model, to obtain the target text features of the target call request, where the text feature extraction layer is learned based on the text feature loss between the first call request samples of the first service type and the second call request samples of the second service type, and the first service type is the service type of the target call request;
[0098] Perform intent prediction based on the target text features through the intent recognition layer of the trained intent recognition model, to obtain the target user intent.
[0099] In some embodiments, the call request response device further includes a training unit (not shown in the figure), and the training unit is specifically configured to:
[0100] Obtain an initial intent recognition model trained based on call request samples of the second service type;
[0101] Obtain the first call request samples of the first service type and the second call request samples of the second service type, where the labeled intent of the first call request samples is the same as the labeled intent of the second call request samples;
[0102] Extract text features from the first call request samples and the second call request samples respectively through the text feature extraction layer of the initial intent recognition model, to obtain the first text features of the first call request samples and the second text features of the second call request samples;
[0103] Obtain the loss between the first text features and the second text features as the text feature loss between the first call request samples and the second call request samples;
[0104] Train the initial intent recognition model based on the text feature loss to obtain the trained intent recognition model.
[0105] In some embodiments, the training unit is specifically configured to:
[0106] Perform intent prediction based on the first text features through the intent recognition layer of the initial intent recognition model, to obtain the predicted intent of the first call request samples;
[0107] Obtain the loss between the predicted intent of the first call request samples and the labeled intent of the first call request samples as the first intent loss of the initial intent recognition model;
[0108] In some embodiments, the training unit is specifically configured to:
[0109] Train the initial intent recognition model based on the first intent loss and the text feature loss to obtain the trained intent recognition model.
[0110] In some embodiments, the training unit is specifically configured to:
[0111] Extract audio features of the first call request sample respectively through the audio feature extraction layer of the initial intent recognition model to obtain the first audio features of the first call request sample;
[0112] The intent prediction of the first call request sample based on the first text feature through the intent recognition layer of the initial intent recognition model to obtain the predicted intent of the first call request sample includes:
[0113] Perform intent prediction on the first call request sample based on the first text feature and the first audio feature through the intent recognition layer of the initial intent recognition model to obtain the predicted intent of the first call request sample.
[0114] In some embodiments, the training unit is specifically configured to:
[0115] Extract audio features of the first call request sample and the second call request sample respectively through the audio feature extraction layer of the initial intent recognition model to obtain the first audio features of the first call request sample and the second audio features of the second call request sample;
[0116] Obtain the loss between the first audio feature and the second audio feature as the audio feature loss between the first call request sample and the second call request sample;
[0117] In some embodiments, the training unit is specifically configured to:
[0118] Train the initial intent recognition model based on the audio feature loss and the text feature loss to obtain the trained intent recognition model.
[0119] In some embodiments, the determining unit 603 is specifically configured to:
[0120] Determine the processing order of the target call request among all call requests to be processed in the service response system based on the target user intent and the other user intents, where the processing order is negatively correlated with the robot response accuracy of the target user intent;
[0121] Obtain the seat capacity of the service response system;
[0122] Determine the response mode of the target call request according to the seat capacity and the processing order.
[0123] In specific implementation, each of the above units can be implemented as an independent entity, or can be combined arbitrarily to be implemented as the same or several entities. For the specific implementation of each of the above units, reference can be made to the embodiments of the call request response method described above, which will not be elaborated here.
[0124] Please refer to Figure 7 , Figure 7 which is a schematic block diagram of the structure of a computer device provided by an embodiment of the present application. The computer device can be a server.
[0125] As Figure 7 shown, the computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the memory can include a non-volatile storage medium and an internal memory.
[0126] The non-volatile storage medium can store an operating system and a computer program. The computer program includes program instructions, and when the program instructions are executed, the processor can be made to execute any call request response method.
[0127] The processor is used to provide computing and control capabilities to support the operation of the entire computer device.
[0128] The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium. When the computer program is executed by the processor, the processor can be made to execute any call request response method.
[0129] The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art can understand that Figure 7 the structure shown in
[0130] It should be understood that the processor may be a Central Processing Unit (CPU), and the processor may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0131] Among them, in one embodiment, the processor is used to run a computer program stored in a memory to implement the following steps:
[0132] Through a trained intent recognition model, perform intent recognition based on the call content of the target call request in the service response system to obtain the target user intent of the target call request; obtain the other user intents of other call requests in the service response system; based on the target user intent and the other user intents, determine the response method for the target call request from preset response methods, where the preset response methods include agent response and robot response; and respond to the target call request according to the response method of the target call request.
[0133] Those of ordinary skill in the art can understand that all or part of the steps in the above call request response method can be completed by instructions, or by controlling relevant hardware through instructions. The instructions can be stored in a computer-readable storage medium and loaded and executed by the processor.
[0134] Therefore, an embodiment of the present application provides a computer-readable storage medium, which stores multiple computer programs that can be loaded by a processor to execute any call request response method provided by the embodiment of the present application.
[0135] Among them, the computer-readable storage medium may include: Read Only Memory (ROM), Random Access Memory (RAM), magnetic disk or optical disc, etc.
[0136] In the above embodiments of the call request response device and the computer-readable storage medium, the descriptions of the respective embodiments have their own focuses. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments. Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes and the beneficial effects brought by the above-described call request response device, computer-readable storage medium and their corresponding units can refer to the description of the call request response method in the above embodiments, and will not be elaborated here specifically.
[0137] The above has introduced in detail a call request response method, device, computer device and computer-readable storage medium provided by the embodiments of the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application. The non-company software tools or components appearing in the embodiments of the present application are only for illustrative introduction and do not represent actual use.
Claims
1. A call request response method, characterized in that, The method includes: Based on the call content of the target call request in the business response system, perform intent recognition through a trained intent recognition model to obtain the target user intent of the target call request; Obtain the other user intents of other call requests in the business response system; Based on the target user intent and the other user intents, determine the response method for the target call request from preset response methods, where the preset response methods include agent response and robot response; Respond to the target call request according to the response method of the target call request.
2. The call request response method according to claim 1, wherein The step of performing intent recognition through a trained intent recognition model based on the call content of the target call request in the business response system to obtain the target user intent of the target call request includes: Through the text feature extraction layer of the trained intent recognition model, perform feature extraction on the text format call content of the target call request to obtain the target text features of the target call request, where the text feature extraction layer is learned based on the text feature loss between the first call request samples of the first service type and the second call request samples of the second service type, and the first service type is the service type of the target call request; Through the intent recognition layer of the trained intent recognition model, perform intent prediction based on the target text features to obtain the target user intent.
3. The call request response method according to claim 2, wherein The trained intent recognition model is trained through the following method: Obtain an initial intent recognition model trained based on call request samples of the second service type; Obtain the first call request samples of the first service type and the second call request samples of the second service type, where the labeled intent of the first call request samples is the same as the labeled intent of the second call request samples; Through the text feature extraction layer of the initial intent recognition model, perform text feature extraction on the first call request samples and the second call request samples respectively to obtain the first text features of the first call request samples and the second text features of the second call request samples; Obtain the loss between the first text features and the second text features as the text feature loss between the first call request samples and the second call request samples; Based on the text feature loss, train the initial intent recognition model to obtain the trained intent recognition model.
4. The call request response method according to claim 3, wherein The method further includes: Through the intent recognition layer of the initial intent recognition model, perform intent prediction based on the first text features to obtain the predicted intent of the first call request samples; Obtain the loss between the predicted intent of the first call request samples and the labeled intent of the first call request samples as the first intent loss of the initial intent recognition model; The step of training the initial intent recognition model based on the text feature loss to obtain the trained intent recognition model includes: Based on the first intent loss and the text feature loss, train the initial intent recognition model to obtain the trained intent recognition model.
5. The call request response method according to claim 4, wherein The method further includes: Through the audio feature extraction layer of the initial intent recognition model, audio features of the first call request sample are extracted respectively to obtain first audio features of the first call request sample. The predicting the intent of the first call request sample based on the first text feature through the intent recognition layer of the initial intent recognition model includes: Predicting the intent of the first call request sample through the intent recognition layer of the initial intent recognition model based on the first text feature and the first audio features to obtain the predicted intent of the first call request sample.
6. The call request response method according to claim 3, wherein The method further includes: Through the audio feature extraction layer of the initial intent recognition model, audio features of the first call request sample and the second call request sample are extracted respectively to obtain first audio features of the first call request sample and second audio features of the second call request sample. Obtaining a loss between the first audio features and the second audio features as the audio feature loss between the first call request sample and the second call request sample. The training the initial intent recognition model based on the text feature loss to obtain the trained intent recognition model includes: Training the initial intent recognition model based on the audio feature loss and the text feature loss to obtain the trained intent recognition model.
7. The call request response method according to claim 1, wherein The determining the response mode of the target call request based on the target user intent and the other user intent includes: Determining the processing order of the target call request among all the call requests to be processed in the service response system based on the target user intent and the other user intent, where the processing order is negatively correlated with the robot response accuracy of the target user intent. Obtaining the seat capacity of the service response system. Determining the response mode of the target call request according to the seat capacity and the processing order.
8. A call request response device, characterized in that, The call request response device includes: An identification unit configured to identify the target user intent of the target call request in the service response system based on the call content of the target call request through the trained intent recognition model. An acquisition unit configured to acquire the other user intents of other call requests in the service response system. A determination unit configured to determine the response mode of the target call request from preset response modes based on the target user intent and the other user intents, where the preset response modes include agent response and robot response. A response unit configured to respond to the target call request according to the response mode of the target call request.
9. A computer device, characterized in that, The computer device includes a memory and a processor. The memory is used for storing a computer program. The processor is configured to execute the computer program and, when executing the computer program, implement the call request response method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and the computer program is loaded by the processor to execute the call request response method according to any one of claims 1 to 7.