Method, apparatus and electronic device for constructing FAQ knowledge base in insurance field

By building a FAQ knowledge base in the insurance field, using sentence recognition and objection models to extract questions and objection answer pairs from the conversation text, and combining the multi-source answer quality sorting model to screen high-quality conversation pairs, it solves the problem that insurance consultants can find it difficult for them to quickly solve customer problems, improves the quality and accuracy of answers, and assists in the company's operations and marketing strategy formulation.

CN114064873BActive Publication Date: 2025-07-29HUI ZE (CHENGDU) NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111354977.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-16
Publication Date
2025-07-29
Estimated Expiration
2041-11-16

AI Technical Summary

Technical Problem

When facing different customers, insurance consultants find it difficult to quickly and accurately solve customers' problems and doubts. The existing knowledge base fails to effectively analyze customer problems and objections, resulting in a long and inefficient solution, which makes it impossible to assist the company in formulating effective marketing strategies.

Method used

By building a FAQ knowledge base in the insurance field, obtain the conversation text of customers and consultants, extract questions and objection answer pairs, use sentence recognition and objection models to predict, combine the multi-source answer quality sorting model, filter high-quality conversation pairs, build a FAQ knowledge base, and integrate the context characteristics of similar conversation pairs to improve answer quality sorting.

Benefits of technology

It shortens the time for insurance consultants to solve customer questions, improves the quality and accuracy of answers, provides support for analyzing customer questions and objections, assists company operations and formulates better marketing strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114064873B_ABST
    Figure CN114064873B_ABST
Patent Text Reader

Abstract

The present application discloses a method, apparatus and electronic device for constructing an FAQ knowledge base in the insurance field. By extracting question-answer pairs and / or objection-answer pairs from the conversation texts between customers and consultants in the insurance field, conversation pairs are obtained, and quality ranking processing is performed on the answers in the conversation pairs and conversation pair filtering processing based on preset quality conditions is performed. Finally, an FAQ knowledge base in the insurance field that meets the quality conditions is constructed. When performing quality ranking processing on the answers in the extracted conversation pairs, for similar conversation pairs in the conversation pairs, the quality ranking result of the answers is controlled by the questions or objections in the similar conversation pairs. By constructing an FAQ knowledge base and controlling the quality ranking result of the answers by the questions or objections in the similar conversation pairs, the present application provides support for analyzing customer questions and objections, can effectively shorten the time for insurance consultants to solve customer doubts and improve the accuracy of solving customer doubts, and correspondingly can achieve the purpose of assisting the company's operation and formulating better marketing strategies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the fields of insurance and artificial intelligence, and particularly relates to a method, device, and electronic device for constructing a Frequently Asked Questions (FAQ) knowledge base in the insurance field. Background Art

[0002] As a means of risk aversion, insurance has gradually become well-known to the public in recent years. There are a wide variety of insurance products on the market. Due to the strong professionalism of insurance products, especially for long-term insurance or annuity insurance and other types of insurance, ordinary consumers will more often understand specific products through corresponding insurance advisors. Insurance advisors thus play an increasingly important role. To effectively promote sales conversion, it is required that insurance advisors can accurately solve customers' problems and doubts.

[0003] However, insurance advisors face different customers every day, and different customers have problems or doubts in different aspects. Coupled with the strong professionalism characteristics of insurance products and the relatively slow growth rate of insurance advisors, all these factors pose challenges for insurance advisors to accurately solve customers' problems and doubts. Therefore, shortening the time for insurance advisors to solve customers' questions and improving the accuracy of solving customers' questions through technical means, so as to achieve the purpose of assisting company operations and formulating better marketing strategies, has become a technical problem that urgently needs to be solved in this field. Summary of the Invention

[0004] In view of this, this application provides a method, device, and electronic device for constructing a FAQ knowledge base in the insurance field. By constructing a knowledge base for customer questions or objections in the insurance field, it provides support for analyzing customer questions and objections, so as to shorten the time for insurance advisors to solve customers' questions and improve the accuracy, and correspondingly achieve the purpose of assisting company operations and formulating better marketing strategies.

[0005] The specific technical solutions are as follows:

[0006] A method for constructing a FAQ knowledge base in the insurance field includes:

[0007] Obtain the conversation text between customers and advisors in the insurance field;

[0008] Extract questions and answers matching the questions, and / or extract objections and answers matching the objections from the conversation text to obtain conversation pairs including question-answer pairs and / or objection-answer pairs;

[0009] Perform quality ranking processing on the answers in the extracted conversation pairs to obtain an answer quality ranking result; among them, for similar conversation pairs in the conversation pairs, control the answer quality ranking result through the questions or objections in the similar conversation pairs;

[0010] Based on the sorted result of the answer quality, filter out the conversation pairs in which the answer quality does not meet the preset quality conditions, and construct an FAQ knowledge base for the insurance field including the conversation pairs that are not filtered out.

[0011] Optionally, extracting questions and answers matching the questions from the conversation text includes:

[0012] Use a pre-constructed sentence pattern recognition model to predict the sentence pattern of the customer statements in the conversation text, and obtain a sentence pattern prediction result indicating whether the customer statement is an interrogative sentence or a non-interrogative sentence;

[0013] Screen out the customer statements with the sentence pattern prediction result being an interrogative sentence as questions;

[0014] Take the first consecutive conversation content of the consultant corresponding to the question in the conversation text as the answer to the question, and obtain a question-answer pair;

[0015] The extracting objections and answers matching the objections includes:

[0016] Use a pre-constructed objection model to predict the content of the customer statements in the conversation text, and obtain a content prediction result indicating whether there is an objection or no objection to the content of the customer statement;

[0017] Extract the customer objection statements with the content prediction result indicating the existence of objections;

[0018] Take the consultant conversation content corresponding to the customer objection statement in the conversation text as the answer to the customer objection statement, and obtain an objection-answer pair.

[0019] Optionally, the processing of sorting the quality of the answers in the extracted conversation pairs to obtain the sorted result of the answer quality includes:

[0020] Perform clustering processing on the extracted conversation pairs based on a preset clustering algorithm to obtain multiple groups of different similar conversation pairs;

[0021] Determine the influence degree of different questions or objections in the similar conversation pairs on different answers;

[0022] Based on the influence degree of different questions or objections in the similar conversation pairs on different answers, perform feature interaction on the answers and their matching questions or objections in the similar conversation pairs to obtain the interaction features corresponding to the answers;

[0023] Perform fusion processing on the interaction features corresponding to the answers and the answer features of other answers in the similar conversation pairs to which the answers belong to obtain the fusion features corresponding to the answers in the similar conversation pairs.

[0024] Based on the fusion features corresponding to each answer in the similar conversation pairs, perform quality sorting processing on each answer in the similar conversation pairs.

[0025] Optionally, determining the influence degree of different questions or objections in the similar conversation pairs on different answers, and based on the influence degree of different questions or objections in the similar conversation pairs on different answers, performing feature interaction on the answers and their matching questions or objections in the similar conversation pairs to obtain interaction features corresponding to the answers, including:

[0026] Performing encoding processing on different questions or objections in the similar conversation pairs to obtain question vectors or objection vectors; performing encoding processing on different answers in the similar conversation pairs to obtain answer vectors;

[0027] Using the following calculation formula to calculate the influence degree of different questions or objections in the similar conversation pairs on different answers and the fused question vector corresponding to the answer respectively:

[0028] g(QE) = softmax(W * QE)

[0029]

[0030] where, W ∈ R n*n represents a weight mapping matrix, n * n represents the matrix dimension of the matrix, R represents the dimension symbol, n represents the number of questions and objections in the similar conversation pairs, softmax is a probability normalization function, and g(QE) ∈ R n*1 represents the weight distribution of different questions or objections in the similar conversation pairs on the answers, and is used to measure the influence degree of different questions or objections in the similar conversation pairs on different answers; att_QEi represents the fused question vector corresponding to the i-th answer;

[0031] Performing interaction processing on the answer vectors and the fused question vectors corresponding to different answers in the similar conversation pairs to obtain interaction features corresponding to different answers respectively.

[0032] Optionally, performing fusion processing on the interaction features corresponding to the answers and the answer features of other answers in the similar conversation pairs to which the answers belong to obtain the fused features corresponding to the answers in the similar conversation pairs, including:

[0033] Performing mean calculation processing on the interaction features corresponding to different answers in the similar conversation pairs to obtain a sorted fusion vector;

[0034] Concatenating the sorted fusion vector to the interaction features of each answer in the similar conversation pairs to obtain the fused features corresponding to each answer in the similar conversation pairs.

[0035] Optionally, performing quality ranking processing on the answers through a pre-trained multi-source answer quality ranking model;

[0036] Among them, the multi-source answer quality ranking model is: a model obtained by training with the pseudo-answer quality label as the model training label; the pseudo-answer quality label is a quality label generated by evaluating the quality of answers in each group of similar conversation pairs based on predetermined business rules.

[0037] Optionally, before clustering the extracted conversation pairs based on a preset clustering algorithm, the method further includes:

[0038] Based on a preset filtering rule, filtering out low-quality conversation pairs in the extracted conversation pairs, so as to perform clustering processing on the conversation pairs obtained after filtering out the low-quality conversation pairs.

[0039] Optionally, before performing quality ranking processing on each answer in the similar conversation pairs, the method further includes:

[0040] Extracting predetermined statistical features corresponding to the similar conversation pairs in the insurance field, so as to perform quality ranking processing on each answer in the similar conversation pairs in combination with the predetermined statistical features.

[0041] An insurance field FAQ knowledge base construction device includes:

[0042] A text acquisition unit, configured to acquire conversation texts between customers and consultants in the insurance field;

[0043] A conversation pair extraction unit, configured to extract questions and answers matching the questions, and / or extract objections and answers matching the objections from the conversation texts, to obtain conversation pairs including question-answer pairs and / or objection-answer pairs;

[0044] A quality ranking processing unit, configured to perform quality ranking processing on the answers in the extracted conversation pairs to obtain an answer quality ranking result; among them, for the similar conversation pairs in the conversation pairs, the answer quality ranking result is controlled by the questions or objections in the similar conversation pairs;

[0045] A knowledge base construction unit, configured to filter out conversation pairs in the conversation pairs whose answer quality does not meet the preset quality conditions based on the answer quality ranking result, and construct an insurance field FAQ knowledge base including the unfiltered conversation pairs.

[0046] An electronic device includes:

[0047] A memory, configured to store a computer instruction set;

[0048] A processor, configured to implement the insurance field FAQ knowledge base construction method as described in any one of the above by executing the instruction set stored in the memory.

[0049] Compared with the traditional technology, the present application has the following beneficial effects:

[0050] The method, device, and electronic device for constructing an FAQ knowledge base in the insurance field provided by this application obtain the conversation texts between customers and consultants in the insurance field, extract questions and answers matching the questions and / or extract objections and answers matching the objections from the conversation texts to obtain conversation pairs, and perform quality ranking processing on the answers in the conversation pairs and conversation pair filtering processing based on preset quality conditions, and finally construct an FAQ knowledge base in the insurance field including the unfiltered conversation pairs meeting the quality conditions. Among them, when performing quality ranking processing on the answers in the extracted conversation pairs, for similar conversation pairs in the conversation pairs, the quality ranking result of the answers is controlled by the questions or objections in the similar conversation pairs.

[0051] Therefore, this application proposes and implements an FAQ knowledge base construction solution that uses question-answer pairs and / or objection-answer pairs in the insurance field as knowledge pairs, provides support for analyzing customer questions and objections by constructing an FAQ knowledge base in the insurance field, and when constructing the FAQ knowledge base, controls the quality ranking result of the answers in the similar conversation pairs through the questions or objections in the similar conversation pairs, so that the context features reflected by the questions or objections in the similar conversation pairs are incorporated into the answer quality ranking, further improving the answer quality ranking performance, facilitating the further construction of an FAQ knowledge base including high-quality conversation pairs, providing a basis for shortening the time for insurance consultants to solve customer questions and improving the solution of customer questions, and correspondingly achieving the purpose of assisting company operations and formulating better marketing strategies. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] In order to more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following will briefly introduce the drawings required to be used in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of this application, and for those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0053] Figure 1 It is a flowchart of the method for constructing an FAQ knowledge base in the insurance field provided by this application;

[0054] Figure 2 It is a training framework of the sentence pattern recognition model provided by this application;

[0055] Figure 3 It is a model structure diagram of the multi-source answer quality ranking model provided by this application;

[0056] Figure 4 It is a flowchart of performing quality ranking processing on the answers provided by this application;

[0057] Figure 5 It is a schematic structural diagram of the device for constructing an FAQ knowledge base in the insurance field provided by this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0058] Next, in combination with the accompanying drawings in the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0059] The applicant's research found that there are already some types of knowledge bases in the current insurance industry. These knowledge bases are generally constructed insurance product knowledge graphs, mainly concentrated on the insurance product side. Insurance advisors can search for basic information of specific insurance products through this knowledge base, such as the maximum insurance amount, the insurance company to which it belongs, etc. However, these types of knowledge bases are only constructed for insurance products and do not involve the construction of knowledge bases for customer questions or objections, and cannot be used to analyze customer questions and objections. And these knowledge bases are small-scale manually constructed knowledge bases, and it is difficult to complete the construction of large-scale knowledge bases in the insurance field.

[0060] Therefore, the present application discloses a method, device and electronic device for constructing an FAQ knowledge base in the insurance field to solve the above technical problems.

[0061] The processing process of the method for constructing an FAQ knowledge base disclosed in the present application is as Figure 1 shown, and specifically includes:

[0062] Step 101, obtain the conversation text between customers and advisors in the insurance field.

[0063] First, obtain the dialogue text during the communication between the insurance advisor and the customer as the data basis.

[0064] Specifically, a part of the obtained dialogue text is used as a training sample for training related models (such as, the sentence pattern recognition model, objection model, multi-source answer quality ranking model, etc. involved later), and the other part is used as a prediction sample of the model to finally construct the required FAQ knowledge base.

[0065] The following provides a specific example of the obtained dialogue text through Table 1:

[0066] Table 1

[0067]

[0068] Step 102, extract questions and answers matching the questions, and / or extract objections and answers matching the objections from the conversation text to obtain conversation pairs including question-answer pairs and / or objection-answer pairs.

[0069] The embodiment of this application pre - constructs a question - answer pair extractor and an objection pair extractor, which are respectively used to extract question - answer pairs and objection - answer pairs in conversation texts. Among them, customer objections are different from customer questions. Specifically, customer objections refer to the situation where customers do not trust the products of insurance companies or sales platforms during the communication process, which affects the conclusion of a deal. Insurance consultants need to answer customer objections accordingly to promote the conclusion of a deal.

[0070] The question - answer pair extractor includes a sentence - pattern recognition model, which is specifically a binary - classification sentence - pattern recognition model used to predict sentence - pattern labels for customer statements in conversation texts. The predicted labels are: interrogative sentence or non - interrogative sentence.

[0071] Combined with Figure 2 As shown in the training framework of the sentence - pattern recognition model, the training process of the binary - classification sentence - pattern recognition model includes:

[0072] 11) Manually annotate the customer's speaking content in the conversation text, mark it as an interrogative sentence or a non - interrogative sentence. For non - customer content, mark it as empty, as shown in the following table:

[0073] Table 2

[0074] role content label customer Hello non-interrogative sentence customer I would like to ask what is the cash value? interrogative sentence advisor Hello! The so-called cash value of the insurance policy is blank customer Oh, I see non-interrogative sentence

[0075] 12) Perform custom word - segmentation on the customer - role content in step 11) and remove stop words;

[0076] Among them, a custom dictionary can be added specifically to effectively perform word - segmentation on vocabulary in a specific field. The custom dictionary can select business hot words in a specific field. The word - segmentation examples are as follows:

[0077] Table 3

[0078] original sentence after word segmentation I would like to ask what is the cash value? I, would like to, ask, what, is, the cash value,?

[0079] 13) Use the labeled samples in step 12) to train the binary - classification sentence - pattern recognition model.

[0080] Among them, frameworks such as FastText and TextCNN can be selected, but are not limited to, to train the binary - classification sentence - pattern recognition model.

[0081] The objection pair extractor includes an objection model. Similarly, the objection model is specifically a binary - classification objection model used to predict whether there are objections in customer statements in conversation texts. The predicted labels are: objection or non - objection.

[0082] The training process of the binary - classification objection model includes:

[0083] 21) Manually annotate the content labels of the customer's speaking content in the conversation text, including objections and non - objections;

[0084] Specific examples are shown in the following table:

[0085] Table 4

[0086] content label I want to take a look at this product non-objection But I think small insurance companies are not very reliable objection I trust offline insurance companies more objection

[0087] 22) Use the labeled samples to train a binary classification dissent model.

[0088] For the training of the binary classification dissent model, frameworks such as FastText and TextCNN can be used, but are not limited to these.

[0089] Based on the pre-constructed question-answer pair extractor and dissent pair extractor, for the dialogue text obtained when the insurance advisor communicates with the customer (such as the prediction samples therein), the question-answer pair extractor can be further used to extract the question-answer pairs of the customer from it, and the dissent pair extractor can be used to extract the dissent-answer pairs of the customer from it, obtaining the conversation pairs corresponding to the conversation text. That is, in the embodiments of the present application, the question-answer pairs and dissent-answer pairs are collectively referred to as conversation pairs.

[0090] Specifically, the process of using the question-answer pair extractor to extract question-answer pairs from the conversation text can be further implemented as follows:

[0091] 31) Use the sentence pattern recognition model in the question-answer pair extractor to predict the sentence pattern of the customer's statement in the conversation text, obtaining a sentence pattern prediction result indicating whether the customer's statement is an interrogative sentence or a non-interrogative sentence;

[0092] 32) Screen the customer's statements with the sentence pattern prediction result of an interrogative sentence as questions;

[0093] 33) Use the first consecutive conversation content of the advisor corresponding to the question in the conversation text as the answer to the question, obtaining a question-answer pair.

[0094] In the example of Table 1, specifically, the advisor conversation content indicated by index 4 and index 5 can be merged and used as the answer to the customer question indicated by index 3, correspondingly obtaining a question-answer pair. The example is as follows:

[0095] Table 5

[0096] question answer What does the cash value refer to? corresponding answer What is the difference between medical insurance and critical illness insurance? corresponding answer

[0097] The process of using the dissent pair extractor to extract dissent-answer pairs from the conversation text can be further implemented as follows:

[0098] 41) Use the pre-constructed dissent model to predict the content of the customer's statement in the conversation text, obtaining a content prediction result indicating whether there is a dissent or no dissent in the content of the customer's statement;

[0099] 42) Extract the customer dissent statements with the content prediction result indicating a dissent;

[0100] 43) Use the advisor's conversation content corresponding to the customer objection statement in the conversation text as the answer to the customer objection statement, and obtain the objection answer pair.

[0101] The following provides specific examples of objection answer pairs through Table 6:

[0102] Table 6

[0103]

[0104] Step 103: Perform quality ranking processing on the answers in the extracted conversation pairs to obtain the answer quality ranking result; among them, for the similar conversation pairs in the conversation pairs, control the answer quality ranking result through the questions or objections in the similar conversation pairs.

[0105] The applicant's research found that since the communication between insurance advisors and customers tends to be colloquial, simply using the advisor's speech content under the customer's questions or objections as the answer will result in a large number of low-quality answers, which in turn leads to redundancy and low quality in the knowledge pairs of the constructed knowledge base, and cannot achieve the actual use effect. In response to this situation, this embodiment further proposes to perform quality ranking processing on the answers in the extracted conversation pairs in order to screen out high-quality conversation pairs as knowledge pairs when constructing the knowledge base.

[0106] Specifically, this application embodiment proposes a multi-source answer quality ranking model, and performs quality ranking processing on the answers in the extracted conversation pairs based on this model.

[0107] The multi-source answer quality ranking model can be constructed through the following processing process:

[0108] I. Filter conversation pairs by rules

[0109] This step of filtering conversation pairs by rules is an optional step.

[0110] Since the FAQ knowledge pairs all come from the conversation text during the communication between insurance advisors and customers, and there is a large amount of meaningless text, therefore, preferably, this embodiment first filters the extracted conversation pairs using a preset rule logic. The specific steps are as follows:

[0111] 51) Filter conversation pairs where the question / objection and the answer have no repeated words;

[0112] First, use an open-source word segmentation tool (such as jieba) to segment the conversation pairs in the knowledge base. If there are no repeated words between the question or objection and the answer, then eliminate this conversation pair.

[0113] 52) Filter answers that still have doubts in the answer;

[0114] Use the sentence pattern recognition model or objection model provided above to predict the sentence pattern / content of the answer corresponding to the question or objection. If the predicted label corresponding to the answer is an interrogative sentence or an objection sentence, then eliminate this conversation pair.

[0115] 53) Filter out the obviously meaningless text in the answer, such as: um, ah.

[0116] II. Aggregation of Similar Conversation Pairs

[0117] There are still a large number of similar questions in the conversation pairs obtained after filtering by rules. To further reduce redundancy and improve the later use efficiency, for the insurance vertical domain, this application clusters the conversation pairs into multiple groups of different similar conversation pairs based on a predetermined clustering algorithm such as singlepass, as shown below;

[0118] Table 7

[0119]

[0120]

[0121] III. Pseudo-Labels for Answer Quality

[0122] The current judgment of answer quality strictly depends on manually annotating the quality of conversation pairs, and different people have different judgment criteria for answer quality. The answers automatically extracted from the communication text have not been strictly manually corrected, and the quality is uneven. To minimize manual intervention as much as possible, this application proposes a method for generating pseudo-labels for answer quality from a business perspective, specifically as follows:

[0123] 61) Distinguish each conversation pair in each group of similar conversation pairs extracted and clustered from the conversation text as coming from a completed order or a non-completed order;

[0124] 62) Since the conversation pairs extracted from the communication in the completed orders can indirectly reflect that the consultant's answer promotes the deal, in this embodiment, the answer quality category of the consultant or salesperson in the completed orders in each group of similar conversation pairs with a monthly conversion rate > R1 and a consultant / salesperson level > S1 is defined as: "good";

[0125] 63) The answer quality category of the consultant or salesperson whose monthly conversion rate < R2 and whose consultant / salesperson level < S2 in the communication of the non-completed orders within N months is defined as "ordinary";

[0126] The generated pseudo-labels are shown in the following table:

[0127] Table 8

[0128]

[0129] It should be noted that the rules adopted for generating the pseudo-labels of the answer quality from a business perspective above are only an example of this application. In implementation, the rules for generating the pseudo-labels of the answer quality can be flexibly set according to requirements.

[0130] IV. Training of Multi-source Answer Quality Ranking Model

[0131] The current quality assessment method simply trains the answer quality model by manually annotating the answer quality categories. As described above, this method severely relies on the quality labels manually annotated and uniformly models the answer quality of different session pairs, without fully considering the quality order among different answers in similar session pairs. Based on this, this application proposes a conditional control answer quality ranking model based on similar session pairs, that is, a multi-source answer quality ranking model.

[0132] Although the pseudo-labels generated for each session pair in each group of similar session pairs in Step 3 cannot be used as accurate labels, the generated pseudo-labels contain both "good" category labels and may also contain "ordinary" category labels. However, after the session pairs extracted from the transaction orders reach a certain number, a certain quality label such as the "good" label will dominate and can be used as the training label of the answer quality ranking model for model training.

[0133] The model structure of the multi-source answer quality ranking model constructed through training is as Figure 3 shown, specifically including: an answer encoding module (encoder), a conditional control module (Gate), an interaction module (interaction), and an answer ranking (context-ranking) module. The functions of each module will be described below through the process of performing quality ranking processing on the answers in the extracted session pairs using this multi-source answer quality ranking model.

[0134] Based on the multi-source answer quality ranking model proposed in the embodiment of this application, see Figure 4 , this step 103 (performing quality ranking processing on the answers in the extracted session pairs) can be further implemented as:

[0135] Step 401: Perform clustering processing on the extracted session pairs based on a preset clustering algorithm to obtain multiple groups of different similar session pairs.

[0136] Specifically, it can be but is not limited to clustering the session pairs used as prediction samples into multiple groups of different similar session pairs based on clustering algorithms such as singlepass. Optionally, before the clustering processing, the session pairs used as prediction samples can also be filtered based on the rule filtering method provided above.

[0137] Step 402: Determine the influence degree of different questions or objections in similar conversation pairs on different answers; based on the influence degree of different questions or objections in similar conversation pairs on different answers, perform feature interaction on the answers and their matching questions or objections in the similar conversation pairs to obtain the interaction features corresponding to the answers.

[0138] The applicant's research finds that information mining can be carried out from different question information in similar conversation pairs as context features to assist in improving the accuracy of answer quality ranking. Moreover, the applicant's research finds that different answers in similar conversation pairs have different emphases in answering. Based on this, in order to evaluate the answer quality more comprehensively and accurately, this embodiment proposes to fully model the question / objection text under similar conversation pairs and control the quality ranking results of different answers in similar conversation pairs through the questions / objections in the similar conversation pairs.

[0139] First, in the condition control module, different questions in each group of similar conversation pairs are encoded into corresponding question vectors QE1, QE2,..., QEn, and the encoding is as follows:

[0140] QEi = encoder(Qi) (1)

[0141] In formula (1), encoder can specifically be, but is not limited to, pre-trained model structures such as BERT and XLNet. The obtained group of similar question vectors (QE1, QE2,..., QEn) is used as a gating device to control the quality ranking results of different answers in similar conversation pairs.

[0142] After that, a weight mapping matrix W ∈ R n*n is introduced to calculate the influence degree of different questions in similar conversation pairs on different answers and calculate the fused question vectors corresponding to different answers in similar conversation pairs. The specific calculation is as follows:

[0143] g(QE) = softmax(W * QE) (2)

[0144]

[0145] In formulas (2)-(3), W ∈ R n*n represents the weight mapping matrix, n*n represents the matrix dimension of the matrix, R represents the dimension symbol, n represents the number of questions and objections in the similar question-and-answer pairs, softmax is the probability normalization function, and g(QE) ∈ R n*1 represents the weight distribution of different questions or objections in similar conversation pairs on answers, which is used to measure the influence degree of different questions or objections in similar conversation pairs on different answers; att_QEi represents the fused question vector corresponding to the i-th answer.

[0146] On this basis, the answer vectors corresponding to different answers in the similar conversation pairs and the fused question vectors are interacted to obtain the interaction features corresponding to different answers respectively.

[0147] Specifically, the interaction module is mainly used to perform feature interaction on the questions and answers in the similar conversation pairs. Accordingly, the interaction module can be used to interact the answer vector with its corresponding fused question vector, and add the fused matching question vector to the answer encoding vector to obtain the interaction features F_AQEi corresponding to different answers in the similar conversation pairs, which can be expressed as [fi1,...,fin]. The interaction process is as follows:

[0148]

[0149] In formula (4), AEi represents the i-th answer encoding vector in the similar conversation pair, and att_QEi represents the fused question vector corresponding to the i-th answer. denotes the multiplication of corresponding vector elements.

[0150] Step 403: Fuse the interaction features corresponding to the answers with the answer features of other answers in the similar conversation pairs to which the answers belong to obtain the fused features corresponding to the answers in the similar conversation pairs.

[0151] Step 404: Based on the fused features corresponding to each answer in the similar conversation pairs, perform quality ranking processing on each answer in the similar conversation pairs.

[0152] The current ranking methods mainly include methods based on poist-wise, pair-wise, and list-wise, etc. The point-wise and pair-wise based ranking methods cannot fully incorporate all document information, and the list-wise method has computational complexity problems. In order to further fully incorporate the potential context features provided by the similar conversation pairs, this application proposes a ranking feature fusion method, which further introduces the mutual interaction between answers when ranking different answers in the similar conversation pairs. Among them, the mutual interaction between answers further introduced when ranking different answers in the similar conversation pairs refers to the fusion process in step 403 of fusing the interaction features corresponding to the answers with the answer features of other answers in the similar conversation pairs to which the answers belong.

[0153] Among them, the fusion process of the interaction features corresponding to the answers with the answer features of other answers in the similar conversation pairs where the answers are located can be further implemented as: calculating the mean value of the interaction features corresponding to different answers in the similar conversation pairs to obtain a ranking fusion vector, and concatenating the ranking fusion vector to the interaction features of each answer in the similar conversation pairs to obtain the fused features corresponding to each answer in the similar conversation pairs.

[0154] Specifically, first, the feature vectors F_AQE1, ..., F_AQEn (i.e., the interaction features corresponding to different answers in similar conversation pairs) are averaged to obtain the sorted fusion vector FM_AQE, and the fusion vector is represented as [fm1, ..., fmn]. This vector can represent the different features of similar conversation pairs, and then the vector is concatenated to each vector F_AQE1, ..., F_AQEn. The concatenated answer vector incorporates the features of other answers in the similar conversation pairs in the group to which it belongs, and is represented as [f1, ..., fn, fm1, ..., fmn]. Correspondingly, the fusion features corresponding to each answer in the similar conversation pairs are obtained.

[0155] The following is a detailed description of the concatenation process between feature vectors:

[0156] Suppose one vector is [1, 2, 3] and the other vector is [3, 4, 5], then concatenate the two to obtain the concatenated vector [1, 2, 3, 3, 4, 5].

[0157] On this basis, the answer ranking module is used to perform quality ranking of each answer based on the fused features corresponding to each answer in the similar conversation pairs. Since the context vector of the group to which the answer belongs is integrated into the ranking stage, the ranking method of this application can further improve the quality ranking performance of each answer.

[0158] Step 104: Based on the answer quality ranking result, filter out conversation pairs whose answer quality does not meet the preset quality conditions, and build an insurance field FAQ knowledge base including the conversation pairs that have not been filtered out.

[0159] Based on the quality ranking of the different answers in each group of similar conversation pairs, conversation pairs that do not meet preset quality conditions can be further filtered out. Preset quality conditions can be, but are not limited to, setting the quality prediction confidence level to a set confidence threshold or ranking within the top k range (k is an integer greater than 0 and can be set as needed).

[0160] The insurance industry stipulates that online insurance sales must meet compliance requirements. Consultants' statements and responses must also pass compliance testing. Therefore, optionally, compliance testing can also be performed on conversation pairs. If a violation is detected, the offending conversation pair will be removed.

[0161] Finally, each conversation pair that meets the quality condition (or meets the quality condition and is compliant) is imported into the corresponding knowledge base as a knowledge base knowledge pair to realize knowledge base construction. In particular, different databases can be used as knowledge carriers to construct the knowledge base.

[0162] Optionally, in addition to including conversation pairs such as question-answer pairs / dissent-answer pairs that meet the requirements, the constructed knowledge base may further include quality assessment information (such as confidence values and / or quality rankings, etc.) of different answers in similar conversation pairs, so as to provide a richer reference for insurance consultants when solving customer questions.

[0163] In summary, the method of the embodiment of the present application. Thus, the present application proposes and implements a construction scheme for an FAQ knowledge base that uses question-answer pairs and / or dissent-answer pairs in the insurance field as knowledge pairs. By constructing an FAQ knowledge base in the insurance field, it provides support for analyzing customer questions and dissents. And when constructing the FAQ knowledge base, the quality ranking result of the answers in the similar conversation pairs is controlled by the questions or dissents in the similar conversation pairs, so that the context features reflected by the questions or dissents in the similar conversation pairs are incorporated into the answer quality ranking, further improving the answer quality ranking performance, facilitating the further construction of an FAQ knowledge base including high-quality conversation pairs, providing a basis for shortening the time for insurance consultants to solve customer questions and improving the solution of customer questions, and correspondingly achieving the purpose of assisting the company's operation and formulating better marketing strategies.

[0164] Optionally, in an embodiment, the method for constructing an FAQ knowledge base in the insurance field of the present application may further include before performing quality ranking processing on each answer in the similar conversation pairs:

[0165] Extracting predetermined statistical features corresponding to the similar conversation pairs in the insurance field, and performing quality ranking processing on each answer in the similar conversation pairs in combination with the extracted predetermined statistical features.

[0166] To further improve the accuracy of answer quality judgment, in addition to using answer features (such as the fused features corresponding to the answers) as the features for judging answer quality, this embodiment also proposes to participate in quality judgment by adding specific statistical features based on the business in the insurance sales field.

[0167] For the insurance field, the introduced statistical features include but are not limited to:

[0168] a. Insurance advisor level: It can be defined in combination with the advisor level under specific business;

[0169] b. Conversation gender: Male or female;

[0170] c. Conversation duration t: Usually, the longer the communication duration in a conversation, the higher the customer acceptance. In this embodiment, is taken as the duration feature.

[0171] It should be noted that although it is theoretically difficult to set different credibility / weights for different genders of male and female, which may have different impacts on the judgment of answer quality based on gender characteristics. However, in specific fields such as the insurance field, there may still be potential subtle differences between different genders (such as affinity and subjective credibility perception), which may have a certain impact on the judgment of answer quality. In view of this, this embodiment introduces the dialogue gender as an auxiliary feature (not the main distinguishing feature) to participate in model training and the quality ranking process of different answers in similar conversations based on the trained model.

[0172] In this embodiment, based on the business in the insurance sales field, specific statistical features are introduced to participate in the quality ranking process of different answers in similar conversation pairs, and potential factors affecting the evaluation of answer quality are mined from multiple aspects and dimensions as much as possible, which can further improve the accuracy of the final answer quality ranking result.

[0173] Corresponding to the above method for constructing the FAQ knowledge base in the insurance field, this application embodiment also discloses an apparatus for constructing the FAQ knowledge base in the insurance field, as Figure 5 shown, the apparatus includes:

[0174] A text acquisition unit 501, configured to acquire the conversation text between the customer and the advisor in the insurance field;

[0175] A conversation pair extraction unit 502, configured to extract questions and answers matching the questions, and / or extract objections and answers matching the objections from the conversation text, to obtain conversation pairs including question-answer pairs and / or objection-answer pairs;

[0176] A quality ranking processing unit 503, configured to perform quality ranking processing on the answers in the extracted conversation pairs to obtain an answer quality ranking result; wherein, for similar conversation pairs in the conversation pairs, the answer quality ranking result is controlled by the questions or objections in the similar conversation pairs;

[0177] A knowledge base construction unit 504, configured to filter out the conversation pairs in the conversation pairs whose answer quality does not meet the preset quality conditions based on the answer quality ranking result, and construct an FAQ knowledge base in the insurance field including the conversation pairs not filtered out.

[0178] In an implementation manner, when the conversation pair extraction unit 502 extracts questions and answers matching the questions from the conversation text, it is specifically configured to:

[0179] Use a pre-constructed sentence pattern recognition model to predict the sentence pattern of the customer's statement in the conversation text, and obtain a sentence pattern prediction result indicating whether the customer's statement is an interrogative sentence or a non-interrogative sentence;

[0180] Screen the customer's statements with the sentence pattern prediction result of an interrogative sentence as questions;

[0181] Use the first consecutive session content of the consultant corresponding to the question in the session text as the answer to the question, and obtain a question-answer pair.

[0182] The session pair extraction unit 502 is specifically configured to, when extracting objections and answers matching the objections from the session text:

[0183] Use a pre-constructed objection model to perform content prediction on the customer statements in the session text, and obtain a content prediction result indicating whether there are objections or no objections in the content of the customer statements;

[0184] Extract the customer objection statements whose content prediction results indicate objections;

[0185] Use the consultant session content corresponding to the customer objection statement in the session text as the answer to the customer objection statement, and obtain an objection-answer pair.

[0186] In one embodiment, the quality ranking processing unit 503 is specifically configured to:

[0187] Perform clustering processing on the extracted session pairs based on a preset clustering algorithm to obtain multiple groups of different similar session pairs;

[0188] Determine the influence degree of different questions or objections in the similar session pairs on different answers;

[0189] Based on the influence degree of different questions or objections in the similar session pairs on different answers, perform feature interaction on the answers and their matching questions or objections in the similar session pairs to obtain the interaction features corresponding to the answers;

[0190] Perform fusion processing on the interaction features corresponding to the answers and the answer features of other answers in the similar session pairs to which the answers belong, to obtain the fusion features corresponding to the answers in the similar session pairs.

[0191] Based on the fusion features respectively corresponding to each answer in the similar session pairs, perform quality ranking processing on each answer in the similar session pairs.

[0192] In one embodiment, when the quality ranking processing unit 503 determines the influence degree of different questions or objections in the similar session pairs on different answers, and based on the influence degree of different questions or objections in the similar session pairs on different answers, performs feature interaction on the answers and their matching questions or objections in the similar session pairs to obtain the interaction features corresponding to the answers, it is specifically configured to:

[0193] Perform encoding processing on the different questions or objections in the similar session pairs to obtain question vectors or objection vectors; perform encoding processing on the different answers in the similar session pairs to obtain answer vectors;

[0194] Use the following calculation formulas to calculate the influence degree of different questions or objections on different answers in similar conversation pairs and the fused question vectors corresponding to the answers respectively:

[0195] g(QE) = softmax(W * QE)

[0196]

[0197] where, W ∈ R n*n represents the weight mapping matrix, n * n represents the matrix dimension of the matrix, R represents the dimension symbol, n represents the number of questions and objections in the similar conversation pair, softmax is the probability normalization function, and g(QE) ∈ R n*1 represents the weight distribution of different questions or objections on the answers in the similar conversation pair, and is used to measure the influence degree of different questions or objections on different answers in the similar conversation pair; att_QEi represents the fused question vector corresponding to the i-th answer;

[0198] Perform an interaction process on the answer vectors and the fused question vectors respectively corresponding to different answers in the similar conversation pair to obtain the interaction features respectively corresponding to different answers.

[0199] In one embodiment, the quality ranking processing unit 503, when fusing the interaction features corresponding to the answers with the answer features of other answers in the similar conversation pairs to which the answers belong to obtain the fused features corresponding to the answers in the similar conversation pairs, is specifically used for:

[0200] Perform a mean calculation process on the interaction features respectively corresponding to different answers in the similar conversation pair to obtain a ranking fusion vector;

[0201] Concatenate the ranking fusion vector to the interaction features of each answer in the similar conversation pair to obtain the fused features corresponding to each answer in the similar conversation pair.

[0202] In one embodiment, the above device performs a quality ranking process on the answers in the extracted conversation pairs through a pre-trained multi-source answer quality ranking model;

[0203] where, the multi-source answer quality ranking model is: a model obtained by training with the pseudo-answer quality label as the model training label; the pseudo-answer quality label is a quality label generated by evaluating the quality of the answers in each group of similar conversation pairs based on a predetermined service rule.

[0204] In one embodiment, the quality ranking processing unit 503, before performing a clustering process on the extracted conversation pairs based on a preset clustering algorithm, is further used for:

[0205] Based on a preset filtering rule, filter out the low-quality conversation pairs in the extracted conversation pairs, so as to perform a clustering process on the conversation pairs obtained after filtering out the low-quality conversation pairs.

[0206] In one embodiment, before performing quality ranking processing on each answer in the similar conversation pairs, the quality ranking processing unit 503 is further configured to:

[0207] Extract the predetermined statistical features corresponding to the similar conversation pairs in the insurance field, and perform quality ranking processing on each answer in the similar conversation pairs by combining the predetermined statistical features.

[0208] For the insurance field FAQ knowledge base construction device disclosed in the embodiments of the present application, since it corresponds to the insurance field FAQ knowledge base construction method disclosed in the above method embodiments, the description is relatively simple. For relevant similarities, please refer to the description of the corresponding method embodiments above, and details are not described here again.

[0209] The embodiments of the present application also disclose an electronic device, which specifically includes:

[0210] A memory for storing a computer instruction set;

[0211] The computer instruction set can be implemented in the form of a computer program.

[0212] A processor for implementing the insurance field FAQ knowledge base construction method disclosed in any of the above method embodiments by executing the computer instruction set.

[0213] The processor can be a central processing unit (CPU), an application-specific integrated circuit (ASIC), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, etc.

[0214] In addition, the electronic device may further include components such as a communication interface and a communication bus. The memory, the processor, and the communication interface complete communication with each other through the communication bus.

[0215] The communication interface is used for communication between the electronic device and other devices. The communication bus can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc.

[0216] In addition, an embodiment of the present application also discloses a storage medium, which stores a computer instruction set. When the stored computer instruction set runs, it can be used to implement the method for constructing an FAQ knowledge base in the insurance field disclosed in any of the above method embodiments.

[0217] It should be noted that the embodiments in this specification are all described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other.

[0218] For the convenience of description, when describing the above system or device, it is divided into various modules or units according to functions for separate description. Of course, when implementing the present application, the functions of each unit can be realized in the same or multiple software and / or hardware.

[0219] From the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments of the present application.

[0220] Finally, it should also be noted that in this article, relational terms such as first, second, third, and fourth are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0221] The above are only the preferred embodiments of the present application. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A method for constructing an FAQ knowledge base in the insurance field, characterized in that, Including: Obtain the conversation text between customers and consultants in the insurance field; Extract questions and answers matching the questions, and / or extract objections and answers matching the objections from the conversation text to obtain conversation pairs including question-answer pairs and / or objection-answer pairs; Perform clustering processing on the extracted conversation pairs based on a preset clustering algorithm to obtain multiple groups of different similar conversation pairs; Determine the influence degree of different questions or objections in the similar conversation pairs on different answers; Based on the influence degree of different questions or objections in the similar conversation pairs on different answers, perform feature interaction on the answers and their matching questions or objections in the similar conversation pairs to obtain the interaction features corresponding to the answers; Perform mean calculation processing on the interaction features corresponding to different answers in the similar conversation pairs to obtain a sorted fusion vector; Concatenate the sorted fusion vector to the interaction features of each answer in the similar conversation pairs to obtain the fusion features corresponding to each answer in the similar conversation pairs; Based on the fusion features corresponding to each answer in the similar conversation pairs, perform quality ranking processing on each answer in the similar conversation pairs to obtain an answer quality ranking result; Based on the answer quality ranking result, filter out the conversation pairs in which the answer quality does not meet the preset quality conditions, and construct an insurance field FAQ knowledge base including the unfiltered conversation pairs.

2. The method according to claim 1, wherein The extracting questions and answers matching the questions from the conversation text includes: Use a pre-constructed sentence pattern recognition model to predict the sentence pattern of the customer statements in the conversation text to obtain a sentence pattern prediction result indicating whether the customer statement is an interrogative sentence or a non-interrogative sentence; Screen the customer statements with the sentence pattern prediction result of interrogative sentences as questions; Take the first consecutive conversation content of the consultant corresponding to the question in the conversation text as the answer to the question to obtain a question-answer pair; The extracting objections and answers matching the objections includes: Use a pre-constructed objection model to predict the content of the customer statements in the conversation text to obtain a content prediction result indicating whether there is an objection or no objection to the content of the customer statement; Extract the customer objection statements with the content prediction result indicating an objection; Take the consultant conversation content corresponding to the customer objection statement in the conversation text as the answer to the customer objection statement to obtain an objection-answer pair.

3. The method according to claim 1, wherein The determining the influence degree of different questions or objections in the similar conversation pairs on different answers, and performing feature interaction on the answers and their matching questions or objections in the similar conversation pairs based on the influence degree of different questions or objections in the similar conversation pairs on different answers to obtain the interaction features corresponding to the answers includes: Perform encoding processing on different questions or objections in the similar conversation pairs to obtain question vectors or objection vectors; perform encoding processing on different answers in the similar conversation pairs to obtain answer vectors; Use the following calculation formula to calculate the influence degree of different questions or objections in the similar conversation pairs on different answers and the fused question vectors corresponding to the answers respectively: g(QE) = softmax(W * QE) where \(W\in R\) n*n represents the weight mapping matrix, \(n\times n\) represents the matrix dimension of the matrix, \(R\) represents the dimension symbol, \(n\) represents the number of questions and objections in the similar conversation pair, softmax is the probability normalization function, and \(g(QE)\in R\) n*1 represents the weight distribution of different questions or objections to the answers in the similar conversation pair, and is used to measure the influence degree of different questions or objections in the similar conversation pair on different answers; \(att\_QE_i\) represents the fused question vector corresponding to the \(i\)-th answer Perform interaction processing on the answer vectors and the fused question vectors corresponding to different answers in the similar conversation pairs to obtain the interaction features corresponding to different answers.

4. The method according to claim 1, characterized in that, Perform quality ranking processing on the answers through a pre-trained multi-source answer quality ranking model; Among them, the multi-source answer quality ranking model is: a model obtained by training with pseudo-answer quality labels as model training labels; the pseudo-answer quality labels are quality labels generated by evaluating the quality of answers in each group of similar conversation pairs based on predetermined business rules.

5. The method according to claim 1, characterized in that, Before clustering the extracted conversation pairs based on a preset clustering algorithm, it further includes: Based on a preset filtering rule, filter out low-quality conversation pairs in the extracted conversation pairs, so as to cluster the conversation pairs obtained after filtering out low-quality conversation pairs.

6. The method according to claim 1, wherein Before performing quality ranking processing on each answer in the similar conversation pairs, it further includes: Extract the predetermined statistical features corresponding to the similar conversation pairs in the insurance field, so as to perform quality ranking processing on each answer in the similar conversation pairs in combination with the predetermined statistical features.

7. An apparatus for constructing an FAQ knowledge base in the insurance field, characterized in that, It includes: A text acquisition unit, configured to acquire the conversation text between customers and consultants in the insurance field; A conversation pair extraction unit, configured to extract questions and answers matching the questions, and / or extract objections and answers matching the objections from the conversation text, to obtain conversation pairs including question-answer pairs and / or objection-answer pairs; A quality ranking processing unit, configured to cluster the extracted conversation pairs based on a preset clustering algorithm to obtain multiple groups of different similar conversation pairs; Determine the influence degree of different questions or objections in the similar conversation pairs on different answers; Based on the influence degree of different questions or objections in the similar conversation pairs on different answers, perform feature interaction on the answers and their matching questions or objections in the similar conversation pairs to obtain the interaction features corresponding to the answers; Perform mean calculation processing on the interaction features corresponding to different answers in the similar conversation pairs to obtain a ranking fusion vector; concatenate the ranking fusion vector to the interaction features of each answer in the similar conversation pairs to obtain the fusion features corresponding to each answer in the similar conversation pairs; Based on the fusion features corresponding to each answer in the similar conversation pairs, perform quality ranking processing on each answer in the similar conversation pairs to obtain an answer quality ranking result; A knowledge base construction unit, configured to filter out conversation pairs with answer quality not meeting the preset quality conditions in the conversation pairs based on the answer quality ranking result, and construct an insurance field FAQ knowledge base including the unfiltered conversation pairs.

8. An electronic device, characterized in that, It includes: A memory, configured to store a computer instruction set; A processor, configured to implement the insurance field FAQ knowledge base construction method according to any one of claims 1-6 by executing the instruction set stored on the memory.

Citation Information

Patent Citations

  • Method for automatically extracting question and answer corpus, on-line intelligent customer service system and electronic device

    CN109508367A

  • Customer service knowledge base establishing method, device and equipment

    CN110019149A