Intelligent customer service FAQ matching method and system based on deep semantic understanding

The deep semantic understanding method preprocesses user queries to align them with financial domain terminology and perform parallel matching, addressing query diversity and complexity issues in financial smart customer service systems, enhancing accuracy and reliability.

CN120316237AActive Publication Date: 2025-07-15ZHONGQI LIANXIN (BEIJING) TECH CO LTD

Patent Information

Application Number
CN202510810974.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-07-15
Estimated Expiration
2045-06-17

AI Technical Summary

Technical Problem

In the existing financial intelligent customer service system, FAQ matching technology is difficult to cope with the diversity and complexity of user questions, and there is a risk of unreliability when generating answers from large language models.

Method used

Through a method based on deep semantic understanding, user questions are rewritten questions that conform to professional terms in the financial field, and combined with intention classification and entity recognition, semantic vectors are generated for matching to avoid directly generating answers.

Benefits of technology

It significantly improves the accuracy and reliability of FAQ matching, reduces the risk of untrusted answers, and enhances semantic alignment of query and knowledge base items.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120316237A_ABST
    Figure CN120316237A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent customer service FAQ matching method and system based on deep semantic understanding. The method comprises the steps that an original question submitted by a user is rewritten into a rewritten question conforming to financial field terminologies and an FAQ knowledge base format through a large language model; the intention classification model and the entity recognition model are used for recognizing the intention and the entity of the original question, the recognized intention and the recognized entity are fused with the original question and the rewritten question respectively, and an original spliced text and a rewritten spliced text are generated; and encoding the original spliced text and the rewritten spliced text into an original vector and a rewritten vector through a semantic matching model, and respectively calculating the similarity between the original vector and the semantic vector of the question item in the pre-processing FAQ knowledge base and the similarity between the rewritten vector and the semantic vector of the question item in the pre-processing FAQ knowledge base so as to screen an optimal answer. According to the method, the FAQ matching accuracy can be remarkably improved, the illusion risk of generating answers by a large model is avoided, and the reliability is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of text recognition, and in particular, to an intelligent customer service FAQ matching method and system based on deep semantic understanding. Background Art

[0002] With the wide application of artificial intelligence technology in the financial industry, intelligent customer service systems have become an important tool for financial institutions to improve service efficiency and customer experience.

[0003] Currently, mainstream financial intelligent customer service systems mainly rely on the FAQ (Frequently Asked Questions) Q&A mode and answer users' questions through a predefined Q&A pair library. However, traditional FAQ matching technologies usually rely on keyword or shallow semantic matching and are difficult to handle the diversity and complexity of users' questions. For example, users may express the same question using different synonyms, abbreviations, or word orders, resulting in a relatively high matching failure rate.

[0004] In recent years, the rise of large pre-trained language models (such as GPT, BERT, etc.) has provided new solutions for Q&A systems, improving the coverage rate by directly generating answers. However, such models are prone to the "hallucination" problem when generating answers, that is, generating inaccurate or false information. In the financial field, such errors may lead to serious compliance risks or customer losses. For example, the model may generate incorrect investment advice or account information, which is unacceptable risk for financial users. Therefore, relying solely on large models to directly generate answers is difficult to meet the strict requirements of the financial industry for accuracy and reliability. Summary of the Invention

[0005] In view of this, embodiments of the present invention provide an intelligent customer service FAQ matching method and system based on deep semantic understanding to eliminate or improve one or more defects existing in the prior art, and solve the problem that the flexibility of existing FAQ matching based on rules or small models is insufficient, while it is difficult to ensure the reliability of directly generating answers by large language models.

[0006] On the one hand, the present invention provides an intelligent customer service FAQ matching method based on deep semantic understanding, and the method includes: Input the original question submitted by the user into a preset large language model for rewriting to generate a rewritten question; wherein, the semantic of the rewritten question is the same as that of the original question, and the expression form of the rewritten question conforms to the professional terms in the financial field and the preset format of the FAQ knowledge base; Input the original question into a pre-trained intent classification model to generate an intent label of the original question; input the original question into a pre-trained entity recognition model to identify the entity and its category of the original question, and obtain an entity sequence; Fuse the original question sentence with the intent label and the entity sequence to obtain an original spliced text; fuse the rewritten question sentence with the intent label and the entity sequence to obtain a rewritten spliced text; Input the original spliced text and the rewritten spliced text into a pre-trained semantic matching model respectively to generate an original vector and a rewritten vector; calculate the text similarity between the original vector and the rewritten vector and the semantic vectors of each question item in the preprocessed FAQ knowledge base respectively and sort them, and select to obtain the final answer; Among them, the preprocessing of the FAQ knowledge base includes inputting each question item in the FAQ knowledge base into the intent classification model and the entity recognition model to generate the intent label and entity sequence of each question item; fusing the intent label and entity sequence of each question item with the corresponding question to obtain a question spliced text; inputting the question spliced text into the semantic matching model to generate the semantic vectors of each question item and store them in the FAQ knowledge base.

[0007] In some embodiments of the present invention, before inputting the original question sentence submitted by the user into a preset large language model for rewriting, the method further includes: Clean and standardize the original question sentence, including at least noise character removal, synonym replacement, and term normalization.

[0008] In some embodiments of the present invention, inputting the original question sentence submitted by the user into a preset large language model for rewriting to generate a rewritten question sentence includes: Rewrite the original question sentence into the rewritten question sentence through a preset prompt template or a sequence-to-sequence generation model.

[0009] In some embodiments of the present invention, the intent classification model uses a bidirectional encoder representation model based on the Transformer architecture for classification; the entity recognition model uses a joint model of a bidirectional encoder representation model and a conditional random field for predicting the entity sequence.

[0010] In some embodiments of the present invention, fusing the original question sentence with the intent label and the entity sequence to obtain an original spliced text; fusing the rewritten question sentence with the intent label and the entity sequence to obtain a rewritten spliced text includes: Add the intent label and the entity sequence to the front and back of the original question sentence with a preset marker to generate the original spliced text; add the intent label and the entity sequence to the front and back of the rewritten question sentence with a preset marker to generate the rewritten spliced text.

[0011] In some embodiments of the present invention, calculating the text similarity between the original vector and the rewritten vector and the semantic vectors of each question item in the preprocessed FAQ knowledge base and sorting them includes: Calculating the cosine similarity between the original vector and the semantic vectors of each question item in the preprocessed FAQ knowledge base to obtain matching scores and constructing a first matching score list; Calculating the cosine similarity between the rewritten vector and the semantic vectors of each question item in the preprocessed FAQ knowledge base to obtain matching scores and constructing a second matching score list; Merging the same question items in the first matching score list and the second matching score list and retaining the highest matching score; sorting the matching scores of the merged question items in descending order, selecting the top pre-set number of question items, and using the FAQ answers of the selected question items as the final answer.

[0012] In some embodiments of the present invention, the method further includes: Selecting the top pre-set number of question items as candidates; Calculating the consistency score between the intent label of the candidate and the intent label of the original question. If the consistency score is lower than the pre-set score, filtering the candidate; Counting the proportion of entities in the candidate that cover the original question. If the proportion is lower than the pre-set proportion, filtering the candidate; Using the FAQ answers of the remaining candidates after filtering as the final answer.

[0013] In some embodiments of the present invention, inputting the question splicing text into the semantic matching model to generate semantic vectors of each question item and storing them in the FAQ knowledge base includes: The semantic matching model adopts a sentence bidirectional encoder representation model; Encoding the question splicing text into semantic vectors through the sentence bidirectional encoder representation model; associating the generated semantic vectors with the corresponding question items and storing them in a structured vector retrieval library to construct the FAQ knowledge base; Among them, the structured vector retrieval library realizes vector indexing and fast retrieval through an efficient similarity search algorithm.

[0014] On the other hand, the present invention also provides an intelligent customer service FAQ matching system based on deep semantic understanding. The system includes: A question rewriting module for rewriting the original question mentioned by the user into a rewritten question that conforms to the professional terms in the financial field and the preset format of the FAQ knowledge base through a preset large language model; An intent classification module for identifying the intent category of the original question through a pre-trained intent classification model to obtain an intent label; An entity recognition module, configured to recognize entities and their categories in the original question sentence through a pre-trained entity recognition model, and obtain an entity sequence; A structured query representation construction module, configured to fuse the original question sentence with the intent label and the entity sequence to obtain an original concatenated text; fuse the rewritten question sentence with the intent label and the entity sequence to obtain a rewritten concatenated text; An FAQ knowledge base preprocessing module, configured to generate intent labels and entity sequences of each question item in the FAQ knowledge base through the intent classification model and the entity recognition model; fuse the intent labels and entity sequences of each question item with the corresponding question to obtain a question concatenated text; encode the question concatenated text through a pre-trained semantic matching model to obtain semantic vectors of each question item; A semantic matching module, configured to encode the original concatenated text and the rewritten concatenated text respectively through the semantic matching model to obtain an original vector and a rewritten vector; calculate the text similarity between the original vector and the rewritten vector and the semantic vectors of each question item in the preprocessed FAQ knowledge base respectively and sort them, and select to obtain a final answer.

[0015] On the other hand, the present invention also provides a computer-readable storage medium, on which a computer program / instructions are stored, and when the computer program / instructions are executed by a processor, the steps of any one of the methods mentioned above are implemented.

[0016] The present invention provides an intelligent customer service FAQ matching method and system based on deep semantic understanding. Compared with the prior art, the present invention only uses a large language model to rewrite the question sentence, rather than directly generating an answer, fundamentally avoiding the hallucination risk of the model generating untrustworthy answers, and significantly improving the reliability. At the same time, the original question sentence of the user is rewritten, making the expression form of the user's question more in line with the habits of the domain FAQ, narrowing the semantic gap; adding intent and entity information to the retrieval process enhances the semantic alignment between the query and the knowledge base items; parallel matching of the original question sentence, the rewritten question sentence and the question items in the FAQ knowledge base can better find semantically similar FAQ questions, significantly improving the FAQ matching accuracy.

[0017] Further, the intelligent customer service FAQ matching system based on deep semantic understanding adopts a modular design, and the functions of each module are independent of each other, and can be flexibly replaced and upgraded according to actual situations and requirements.

[0018] Additional advantages, objects, and features of the present invention will be partly set forth in the description which follows and, in part, will be obvious to those having ordinary skill in the art upon examination of the following, or may be learned from practice of the present invention. The objects and other advantages of the present invention may be realized and attained by the structure particularly pointed out in the specification and the drawings.

[0019] Those skilled in the art will understand that the objects and advantages that can be achieved by the present invention are not limited to those specifically described above, and the above and other objects that the present invention can achieve will be more clearly understood from the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The drawings described herein are for further understanding of the present invention, form a part of this application, and do not limit the present invention. In the drawings: Figure 1 It is a schematic diagram of the steps of the intelligent customer service FAQ matching method based on deep semantic understanding in an embodiment of the present invention.

[0021] Figure 2 It is a schematic flowchart of the intelligent customer service FAQ matching method based on deep semantic understanding in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0022] To make the objects, technical solutions, and advantages of the present invention more clearly understood, the present invention will be further described in detail below in conjunction with the embodiments and the drawings. Herein, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but do not limit the present invention.

[0023] Herein, it should also be noted that in order to avoid obscuring the present invention with unnecessary details, only the structures and / or processing steps closely related to the solution of the present invention are shown in the drawings, and other details less related to the present invention are omitted.

[0024] It should be emphasized that the term "comprising / including" when used herein refers to the presence of features, elements, steps, or components, but does not exclude the presence or addition of one or more other features, elements, steps, or components.

[0025] Herein, it should also be noted that if not otherwise specified, the term "connection" in this article can refer not only to direct connection but also to indirect connection with an intermediate.

[0026] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In the drawings, the same reference numerals represent the same or similar components, or the same or similar steps.

[0027] It should be emphasized here that the step marks mentioned below do not limit the order of the steps. Instead, it should be understood that the steps can be executed in the order mentioned in the embodiments, or different from the order in the embodiments, or several steps can be executed simultaneously.

[0028] To solve the problem that the existing FAQ matching based on rules or small models lacks flexibility, while it is difficult to ensure the reliability of directly generating answers by large language models, the present invention provides an intelligent customer service FAQ matching method based on deep semantic understanding. As Figure 1 shown, this method includes the following steps S101 to S104: Step S101: Input the original question submitted by the user into a preset large language model for rewriting to generate a rewritten question. Among them, the rewritten question has the same semantics as the original question, and the expression form of the rewritten question conforms to the professional terms in the financial field and the preset format of the FAQ knowledge base.

[0029] Step S102: Input the original question into a pre-trained intent classification model to generate an intent label for the original question. Input the original question into a pre-trained entity recognition model to identify the entity and its category in the original question, and obtain an entity sequence.

[0030] Step S103: Integrate the original question with the intent label and the entity sequence to obtain an original spliced text. Integrate the rewritten question with the intent label and the entity sequence to obtain a rewritten spliced text.

[0031] Step S104: Input the original spliced text and the rewritten spliced text into a pre-trained semantic matching model respectively to generate an original vector and a rewritten vector; calculate the text similarity between the original vector and the rewritten vector and the semantic vector of each question item in the pre-processed FAQ knowledge base respectively and sort them, and select to obtain the final answer.

[0032] Among them, the pre-processing of the FAQ knowledge base includes inputting each question item in the FAQ knowledge base into an intent classification model and an entity recognition model to generate an intent label and an entity sequence for each question item; integrating the intent label and the entity sequence of each question item with the corresponding question to obtain a question spliced text; inputting the question spliced text into a semantic matching model to generate the semantic vector of each question item and store it in the FAQ knowledge base.

[0033] As Figure 2 shown, it is a flow schematic diagram of an intelligent customer service FAQ matching method based on deep semantic understanding.

[0034] In step S101, considering that the user's question-asking methods vary greatly, the same question may have multiple expression forms, and the expression forms are usually colloquial, concise, or ambiguous. Therefore, the original question submitted by the user is input into a preset large language model for rewriting to generate a rewritten question. Among them, the generated rewritten question has the same semantics as the original question, but the expression form of the rewritten question conforms to the professional terms in the financial field and the preset format of the FAQ knowledge base.

[0035] In some embodiments, the preset large language model is a model fine-tuned on vertical domain corpus. Among them, the vertical domain corpus refers to a collection of text data specifically collected and annotated for a specific industry or application scenario (such as finance, medical, legal, etc.). The fine-tuned large language model (such as ChatGLM) can make the generated rewritten question more in line with the expression habits and professional requirements of the financial industry, and avoid colloquial or unprofessional expressions generated by general models.

[0036] In some embodiments, before inputting the original question submitted by the user into the preset large language model for rewriting, the original question is first cleaned and standardized, such as operations like noise character removal, synonym replacement, and term standardization, to purify the input text, unify semantic expressions, adapt to domain knowledge, and improve the quality of input data for subsequent question rewriting, intent classification, entity recognition, etc.

[0037] In some embodiments, the original question is rewritten into a rewritten question through a preset prompt template (Prompt) or a sequence-to-sequence generation model.

[0038] Specifically, by designing a specific prompt template (Prompt), the model is guided to generate a rewritten question that meets the requirements. Exemplarily, the template is set as: "Please rewrite the following user question into a professional question that conforms to the financial customer service FAQ format, keeping the original meaning unchanged: Original question: [user input] Rewritten question: " Optionally, domain constraints are added, such as: "Requirement: Start with 'Excuse me', include the full product name (such as 'CSI 300 Index Fund'), and avoid colloquial words." Thus, when rewriting, the original question is embedded in the Prompt template and input into the model, and the model outputs the generated rewritten question.

[0039] It can also be achieved based on sequence-to-sequence generation. The question rewriting task is modeled as an end-to-end generation task of "original question → rewritten question" and implemented through fine-tuning the model.

[0040] In step S102, the original question is input into a pre-trained intent classification model to generate an intent label for the original question, such as account opening consultation, financial product price inquiry, trading process, etc. The original question is input into a pre-trained entity recognition model to identify the entities (such as bank name, security code, product name, time, amount, etc.) in the original question and their categories, obtaining a structured entity sequence.

[0041] In some embodiments, the intent classification model employs a Bidirectional Encoder Representations from Transformers (BERT) model based on the Transformer architecture.

[0042] The training method of the intent classification model includes: Construct an intent classification training set, which includes a large number of user questions in the financial field, and label the true intent label for each user question.

[0043] Construct a model, which adopts the BERT model and includes an input layer, a BERT encoding layer, and a classification layer.

[0044] Use the intent classification training set to train the BERT model. Input the user questions with true intent labels into the BERT model in batches to generate the predicted intent labels for each user question. Construct a loss function for the true intent label and the predicted intent label, and optimize the model with the goal of minimizing the loss function to obtain the intent classification model.

[0045] In some embodiments, the entity recognition model adopts a joint model of a Bidirectional Encoder Representations from Transformers (BERT) model and a Conditional Random Field (CRF).

[0046] The training method of the entity recognition model includes: Construct an entity recognition training set, which includes a large corpus of entities annotated in the financial field, covering key entity types.

[0047] Construct a model, which adopts a joint model of BERT and CRF and includes an input layer, a BERT encoding layer, and a CRF correction layer.

[0048] Use the entity recognition training set to train the joint model. Input the corpus with true entity annotations into the joint model in batches to generate the predicted entities and their types. Construct a negative log-likelihood loss to optimize the difference between the true label and the predicted value to obtain the entity recognition model.

[0049] In step S103, the original question is fused with the intent label and the entity sequence to obtain the original concatenated text. The rewritten question is fused with the intent label and the entity sequence to obtain the rewritten concatenated text.

[0050] In some embodiments, intent tags and entity sequences are added before and after the original question or the rewritten question with preset tags to obtain the original spliced text and the rewritten spliced text.

[0051] Exemplarily, the intent tag is placed before the original question or the rewritten question, and the entity sequence is placed after the original question or the rewritten question: [INTENT: Inquiry about financial products] How to buy Fund A [ENTITY: Fund A].

[0052] In step S104, the original spliced text and the rewritten spliced text are respectively input into a pre-trained semantic matching model to generate an original vector and a rewritten vector; the text similarity between the original vector and the rewritten vector and the semantic vector of each question item in the preprocessed FAQ knowledge base is calculated and sorted respectively, and the final answer is selected.

[0053] In some embodiments, the semantic matching model adopts the Sentence-BERT model (Sentence Bidirectional Encoder Representations from Transformers model). The training method of the semantic matching model includes: Construct a semantic matching training set, which includes a large number of FAQ question-answer pairs in the financial field. Similar sentence pairs are labeled as 1 (semantically the same), and dissimilar sentence pairs are labeled as 0.

[0054] Construct a model, which adopts a Siamese network based on BERT, including an input layer, a feature extraction layer, and a similarity calculation layer. Exemplarily, the similarity calculation layer is a classification task, that is, the vectors of two sentences are concatenated and then input into a fully connected layer to output the similarity probability; or the similarity calculation layer is a regression task, that is, the cosine similarity is directly calculated and optimized with the mean squared error (MSE).

[0055] Use the semantic matching training set to train the model, input the labeled FAQ question-answer pairs into the model to generate predicted similarity values. If it is a classification task, construct a cross-entropy loss; if it is a regression task, construct a mean squared error loss or a cosine embedding loss to optimize the model and obtain the semantic matching model.

[0056] In some embodiments, calculate the cosine similarity between the original vector and the semantic vector of each question item in the preprocessed FAQ knowledge base to obtain a matching score and construct a first matching score list; calculate the cosine similarity between the rewritten vector and the semantic vector of each question item in the preprocessed FAQ knowledge base to obtain a matching score and construct a second matching score list.

[0057] Merge the same question items in the first matching score list and the second matching score list, and retain the highest matching score. For example, if the matching score of question item A in the first matching score list is 0.8 and the matching score of question item A in the second matching score list is 0.9, then merge question item A in the two matching score lists and take the matching score of 0.9. Sort the matching scores of the merged question items in descending order, and select the top pre-set number (TopK) of question items. For example, select the top three question items with the highest matching scores, and use the FAQ answers of the selected question items as the final answers.

[0058] In some embodiments, after selecting the top pre-set number of question items, it further includes: using the selected top pre-set number of question items as candidate items. Calculate the consistency score between the intent label of the candidate item and the intent label of the original question sentence. If the consistency score is lower than the pre-set score, then filter out the candidate item. At the same time, count the proportion of the entities in the candidate item that cover the original question sentence. If the proportion is lower than the pre-set proportion, then filter out the candidate item. Use the FAQ answers of the remaining candidate items after filtering as the final answers.

[0059] In some embodiments, for better matching, the FAQ knowledge base is preprocessed in advance, including inputting each question item in the FAQ knowledge base into the above-mentioned trained intent classification model and entity recognition model to generate the intent label and entity sequence of each question item; fusing the intent label and entity sequence of each question item with the corresponding question to obtain a question splicing text; inputting the question splicing text into the above-mentioned trained semantic matching model to generate semantic vectors of each question item and store them in the FAQ knowledge base.

[0060] In some embodiments, the semantic matching model uses a sentence bidirectional encoder representation model. Through the sentence bidirectional encoder representation model, the question splicing text is encoded into a semantic vector, and then the generated semantic vector is associated with the corresponding question item and stored in a structured vector retrieval library to obtain a preprocessed FAQ knowledge base.

[0061] In some embodiments, the structured vector retrieval library realizes vector indexing and fast retrieval through an efficient similarity search algorithm, such as the Facebook AI Similarity Search (Faiss), Milvus vector database (Milvus), etc.

[0062] Corresponding to the intelligent customer service FAQ matching method based on deep semantic understanding, the present invention also provides an intelligent customer service FAQ matching system based on deep semantic understanding, and the system includes: A question sentence rewriting module, configured to rewrite the original question sentence mentioned by the user into a rewritten question sentence that conforms to the professional terms in the financial field and the preset format of the FAQ knowledge base through a preset large language model.

[0063] An intent classification module, configured to identify the intent category of the original question sentence through a pre-trained intent classification model to obtain an intent label.

[0064] An entity recognition module, configured to identify the entity and its category of the original question sentence through a pre-trained entity recognition model to obtain an entity sequence.

[0065] A structured query representation construction module, configured to fuse the original question sentence with the intent label and the entity sequence to obtain an original concatenated text; and fuse the rewritten question sentence with the intent label and the entity sequence to obtain a rewritten concatenated text.

[0066] An FAQ knowledge base preprocessing module, configured to generate the intent label and the entity sequence of each question item in the FAQ knowledge base through the intent classification model and the entity recognition model; fuse the intent label and the entity sequence of each question item with the corresponding question to obtain a question concatenated text; and encode the question concatenated text through a pre-trained semantic matching model to obtain the semantic vector of each question item.

[0067] A semantic matching module, configured to encode the original concatenated text and the rewritten concatenated text respectively through a semantic matching model to obtain an original vector and a rewritten vector; calculate the text similarity between the original vector and the rewritten vector and the semantic vector of each question item in the preprocessed FAQ knowledge base respectively and sort them, and select to obtain the final answer.

[0068] Corresponding to the above method, the present invention further provides an electronic device, which includes a computer device. The computer device includes a processor and a memory. The memory stores computer instructions, and the processor is configured to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the electronic device implements the steps of the method described above.

[0069] The embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the method described above. The computer-readable storage medium may be a tangible storage medium, such as a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a floppy disk, a hard disk, a removable storage disk, a CD-ROM, or any other form of storage medium known in the technical field.

[0070] Those of ordinary skill in the art should understand that the various exemplary components, systems, and methods described in connection with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Specifically, whether to implement in hardware or software depends on the specific application and design constraints of the technical solution. A person skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present invention. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, and so on. When implemented in software, the elements of the present invention are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted via a data signal carried in a carrier wave on a transmission medium or a communication link.

[0071] It should be clear that the present invention is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present invention.

[0072] In the present invention, the features described and / or illustrated for one embodiment can be used in the same or a similar manner in one or more other embodiments, and / or combined with the features of other embodiments or replace the features of other embodiments.

[0073] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, various changes and variations can be made to the embodiments of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An intelligent customer service FAQ matching method based on deep semantic understanding, characterized in that, The method includes: Input the original question submitted by the user to the intelligent customer service into a preset large language model for rewriting to generate a rewritten question; wherein, the rewritten question has the same semantics as the original question, and the expression form of the rewritten question conforms to the professional terms in the financial field and the preset format of the FAQ knowledge base; Input the original question into a pre-trained intent classification model to generate an intent label for the original question; input the original question into a pre-trained entity recognition model to identify the entities and their categories in the original question, obtaining an entity sequence; Fuse the original question with the intent label and the entity sequence to obtain an original spliced text; fuse the rewritten question with the intent label and the entity sequence to obtain a rewritten spliced text; Input the original spliced text and the rewritten spliced text into a pre-trained semantic matching model respectively to generate an original vector and a rewritten vector; calculate the text similarity between the original vector and the rewritten vector and the semantic vectors of each question item in the preprocessed FAQ knowledge base respectively and sort them, and select to obtain the final answer; Among them, the preprocessing of the FAQ knowledge base includes inputting each question item in the FAQ knowledge base into the intent classification model and the entity recognition model to generate an intent label and an entity sequence for each question item; fusing the intent label and the entity sequence of each question item with the corresponding question to obtain a question spliced text; inputting the question spliced text into the semantic matching model to generate semantic vectors of each question item and store them in the FAQ knowledge base.

2. The intelligent customer service FAQ matching method based on deep semantic understanding according to claim 1, wherein Before inputting the original question submitted by the user into a preset large language model for rewriting, the method further includes: Clean and standardize the original question, including at least removing noise characters, replacing synonyms, and normalizing terms.

3. The intelligent customer service FAQ matching method based on deep semantic understanding according to claim 1, characterized in that Inputting the original question submitted by the user into a preset large language model for rewriting to generate a rewritten question includes: Rewriting the original question into the rewritten question through a preset prompt template or a sequence-to-sequence generation model.

4. The intelligent customer service FAQ matching method based on deep semantic understanding according to claim 1, wherein, The intent classification model uses a bidirectional encoder representation model based on the Transformer architecture for classification; the entity recognition model uses a joint model of a bidirectional encoder representation model and a conditional random field for predicting the entity sequence.

5. The intelligent customer service FAQ matching method based on deep semantic understanding according to claim 1, characterized in that Fuse the original question with the intent label and the entity sequence to obtain an original spliced text; Fusing the rewritten question with the intent label and the entity sequence to obtain a rewritten spliced text includes: Adding the intent label and the entity sequence to the front and back of the original question with a preset marker to generate the original spliced text; Adding the intent label and the entity sequence to the front and back of the rewritten question with a preset marker to generate the rewritten spliced text.

6. The intelligent customer service FAQ matching method based on deep semantic understanding according to claim 1, wherein Calculating the text similarity between the original vector and the rewritten vector and the semantic vectors of each question item in the preprocessed FAQ knowledge base respectively and sorting them includes: Calculating the cosine similarity between the original vector and the semantic vectors of each question item in the preprocessed FAQ knowledge base to obtain a matching score and constructing a first matching score list; Calculate the cosine similarity between the rewritten vector and the semantic vectors of each question item in the preprocessed FAQ knowledge base, obtain the matching scores, and construct a second matching score list; Merge the same question items in the first matching score list and the second matching score list, and retain the highest matching score; sort the matching scores of the merged question items in descending order, select the top pre-set number of question items, and use the FAQ answers of the selected question items as the final answer.

7. The intelligent customer service FAQ matching method based on deep semantic understanding according to claim 6, characterized in that, The method further includes: Select the top pre-set number of question items as candidates; Calculate the consistency score between the intent labels of the candidates and the intent label of the original question. If the consistency score is lower than the pre-set score, filter the candidates; Count the proportion of entities in the original question covered by the candidates. If the proportion is lower than the pre-set proportion, filter the candidates; Use the FAQ answers of the remaining candidates after filtering as the final answer.

8. The intelligent customer service FAQ matching method based on deep semantic understanding according to claim 1, characterized in that Input the question splicing text into the semantic matching model to generate the semantic vectors of each question item and store them in the FAQ knowledge base, including: The semantic matching model uses a sentence bidirectional encoder representation model; Encode the question splicing text into semantic vectors through the sentence bidirectional encoder representation model; associate the generated semantic vectors with the corresponding question items and store them in the structured vector retrieval library to construct the FAQ knowledge base; Among them, the structured vector retrieval library realizes vector indexing and fast retrieval through an efficient similarity search algorithm.

9. An intelligent customer service FAQ matching system based on deep semantic understanding, characterized in that, The system includes: A question rewriting module for rewriting the original question mentioned by the user into a rewritten question that conforms to the professional terms in the financial field and the preset format of the FAQ knowledge base through a preset large language model; An intent classification module for identifying the intent category of the original question through a pre-trained intent classification model to obtain an intent label; An entity recognition module for identifying the entity and its category of the original question through a pre-trained entity recognition model to obtain an entity sequence; A structured query representation construction module for fusing the original question with the intent label and the entity sequence to obtain an original splicing text; fusing the rewritten question with the intent label and the entity sequence to obtain a rewritten splicing text; An FAQ knowledge base preprocessing module for generating the intent labels and entity sequences of each question item in the FAQ knowledge base through the intent classification model and the entity recognition model; fusing the intent labels and entity sequences of each question item with the corresponding question to obtain a question splicing text; encoding the question splicing text through a pre-trained semantic matching model to obtain the semantic vectors of each question item; A semantic matching module for encoding the original splicing text and the rewritten splicing text respectively through the semantic matching model to obtain an original vector and a rewritten vector; calculating the text similarity between the original vector and the rewritten vector and the semantic vectors of each question item in the preprocessed FAQ knowledge base respectively and sorting, and selecting to obtain the final answer.

10. A computer-readable storage medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Query feedback method and device, computer equipment and storage medium

    CN111538894A

  • FAQ question similarity calculation method and system

    CN111581354A

  • Statement processing method and device, equipment, medium and computer program product

    CN114281959A

  • Intelligent question and answer method and device, storage medium and electronic equipment

    CN115878790A

  • Diversity-enhanced text retrieval enhancement generation method and system

    CN118260406A

Cited By

  • Large model-based intention recognition method, system and equipment

    CN120724981A