Question answering method based on FAQ system, computer device, and storage medium
By constructing a collection of candidate questions, including standard and fuzzy candidate questions, and using the search model to filter target candidate questions, the accuracy problem of the FAQ system when matching user questions is solved, and a more efficient and accurate question-and-answer service is achieved.
Patent Information
- Application Number
- PCT/CN2024/136972
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-21
- Filing Date
- 2024-12-05
- Publication Date
- 2025-06-26
AI Technical Summary
When matching user questions and setting topics, it is difficult for existing FAQ systems to accurately match the best question-and-answer pairs in the case of low literal similarity, high semantic similarity, or high literal similarity and low semantic similarity, resulting in low matching accuracy.
By obtaining user query statements, a collection of candidate questions is constructed, including standard candidate questions and fuzzy candidate questions. Standard candidate problems are obtained through similarity matching, and fuzzy candidate problems are the generalized expressions of standard candidate problems. Then, through preset search models, such as keyword search models and vector search models, the target candidate questions that best match the user query statements are selected, and the corresponding target answers are feedback.
It improves the accuracy and efficiency of the FAQ system when matching user problems, and can provide users with the answers they need faster and more accurately, solving the problems of stiff matching, low accuracy and recall of existing FAQ systems.
Smart Images

Figure CN2024136972_26062025_PF_FP_ABST
Abstract
Description
Question-answering method, computer device, and storage medium based on FAQ system
[0001] This application claims priority to Chinese patent application No. 202311776061.2, filed on December 21, 2023, with the invention name “Question and answer method, computer device and storage medium based on FAQ system”. The entire contents of the above Chinese patent application are incorporated into this application by reference. Technical Field
[0002] The present application relates to the field of question-answering technology, and specifically provides a question-answering method, computer device, and storage medium based on an FAQ system. Background Art
[0003] Frequently Asked Question (FAQ) is a relatively mature question-and-answer solution. When a user asks a question, the FAQ system uses semantic matching to find the most similar question-and-answer pair in a Q&A database and returns the answer to the user. This approach offers significant advantages, including controllable question-and-answer scope, ease of use, and a good user experience. However, its difficulty lies in determining the similarity between the user's question and the designated topic. Generally, a common approach involves manually defining a representative question and a response for a topic, then training a semantic matching model using a large amount of data. However, while using a semantic matching model to determine the optimal question-and-answer pair can ensure the FAQ system's response speed, it often fails to meet matching accuracy requirements. Specifically, when literal similarity is low but semantic similarity is high, or when literal similarity is high but semantic similarity is low, the question-and-answer pair matched by the semantic matching model may not be the optimal one. Consequently, current FAQ dialogue solutions suffer from low matching accuracy.
[0004] Accordingly, the art needs a new FAQ system question-answering solution to solve the above problems. Summary of the Invention
[0005] In order to overcome the above-mentioned defects, the present application is proposed to provide a solution or at least partially solve the above-mentioned technical problems.
[0006] In a first aspect, the present application provides a question-answering method based on a FAQ system, the method comprising:
[0007] Get user query statements;
[0008] Obtaining a set of candidate questions based on the user query, the set of candidate questions including standard candidate questions and fuzzy candidate questions, wherein the standard candidate questions are candidate questions obtained by similarity matching with the user query, and the fuzzy candidate questions are candidate questions that are generalized versions of the standard candidate questions;
[0009] determining a target candidate question based on the candidate question set;
[0010] Based on the target candidate questions, a target answer is determined and fed back to the user.
[0011] In one technical solution of the above-mentioned question-answering method based on the FAQ system, obtaining a set of candidate questions based on the user query statement includes:
[0012] Recalling the user query statement through a preset retrieval model to obtain an initial question set;
[0013] The data in the initial question set are sorted from high to low according to the similarity, and based on the sorting result, a preset number of candidate questions with high similarity rankings in the initial question set are selected to obtain a candidate question set.
[0014] In one technical solution of the above-mentioned FAQ system-based question-answering method, the preset retrieval model includes a keyword retrieval model and a vector retrieval model. Recalling the user query statement using the preset retrieval model to obtain an initial question set includes:
[0015] Based on the user query statement, obtaining a first initial question set through a keyword retrieval model;
[0016] Based on the user query statement, obtaining a second initial question set through a vector retrieval model;
[0017] An initial question set is obtained based on the first initial question set and the second initial question set.
[0018] In one technical solution of the above-mentioned question-answering method based on the FAQ system, the keyword retrieval model is constructed based on standard questions, fuzzy questions, and corresponding answers; wherein the standard questions and the corresponding answers are collected in advance, and the fuzzy questions are generalized by generalizing the standard questions based on a preset text generation model to obtain multiple generalized questions corresponding to the standard questions, and the generalized questions and the standard questions are checked for similarity, and the generalized questions are obtained based on the similarity with the standard questions that meets the preset conditions; and / or,
[0019] The vector retrieval model is constructed based on the vectorized standard questions, the fuzzy questions and the corresponding answers.
[0020] In one technical solution of the above-mentioned FAQ system-based question-answering method, in response to the target answer fed back to the user being no answer or a preset fallback answer, the method further includes:
[0021] Determine whether the similarity between the user query statement and the corresponding candidate question is greater than or equal to a preset similarity threshold; if so, add the user query statement to the keyword retrieval model and the vector retrieval model.
[0022] In one technical solution of the above-mentioned question-answering method based on the FAQ system, before sorting the data in the initial question set from high to low according to similarity, the method further includes:
[0023] Normalize the similarities between the user query statements and the candidate questions in the initial question set.
[0024] In one technical solution of the above-mentioned FAQ system-based question-answering method, before obtaining a set of candidate questions based on the user query, the method further includes:
[0025] The user query statement is preprocessed, and the preprocessing includes at least one of correcting format errors, removing non-keywords, replacing standard words, processing wake-up words, and vectorization processing.
[0026] In one technical solution of the above-mentioned FAQ system-based question-answering method, determining a target candidate question based on the candidate question set includes:
[0027] Combining the user query statement with the standard candidate question and the fuzzy candidate question in the candidate question set into statement pairs;
[0028] determining a target sentence pair based on the sentence pair;
[0029] A target candidate question is determined based on the similarity between the user query sentence and the corresponding candidate question in the target sentence pair.
[0030] In one technical solution of the above-mentioned question-answering method based on the FAQ system, determining the target sentence pair based on the sentence pair includes:
[0031] Determine the semantic matching score of each sentence pair based on the preset semantic recognition model;
[0032] The sentence pair with the highest semantic matching score is taken as the target sentence pair.
[0033] In one technical solution of the above-mentioned FAQ system-based question answering method, determining the target candidate question based on the similarity between the user query sentence and the corresponding candidate question in the target sentence pair includes:
[0034] Performing weighted calculation on the similarity and semantic matching score of the user query sentence of the target sentence pair and the corresponding candidate question to determine the final score of the target sentence pair;
[0035] Determining whether the final score is greater than or equal to a preset threshold;
[0036] When the final score is greater than or equal to a preset threshold, the candidate question corresponding to the target sentence pair is determined to be the target candidate question.
[0037] In one technical solution of the above-mentioned FAQ system-based question-answering method, determining a target answer based on the target candidate question and feeding it back to the user includes:
[0038] Obtaining a target answer corresponding to the target candidate question;
[0039] Feedback the target answer to the user in at least one of image, text, voice, and video; or
[0040] The method further comprises:
[0041] When the final score is less than a preset threshold, the user is fed back that there is no answer in at least one of image, text, voice, and video, or a preset fallback answer is fed back to the user.
[0042] In a second aspect, a computer device is provided, which includes a processor and a memory, wherein the memory is suitable for storing multiple program codes, and the program codes are suitable for being loaded and run by the processor to execute the question-answering method based on the FAQ system described in any one of the technical solutions of the above-mentioned question-answering method based on the FAQ system.
[0043] In a third aspect, a computer-readable storage medium is provided, which stores a plurality of program codes, wherein the program codes are suitable for being loaded and run by a processor to execute the question-answering method based on the FAQ system described in any one of the technical solutions of the above-mentioned question-answering method based on the FAQ system.
[0044] The above one or more technical solutions of this application have at least one or more of the following beneficial effects:
[0045] The present application provides a question-answering method based on an FAQ system, comprising: obtaining a user query statement; obtaining a candidate question set based on the user query statement, the candidate question set including standard candidate questions and fuzzy candidate questions, wherein the standard candidate question is a candidate question obtained by similarity matching with the user query statement, and the fuzzy candidate question is a candidate question generalized from the standard candidate question; determining a target candidate question based on the candidate question set; and determining a target answer based on the target candidate question and feeding it back to the user. The present application obtains a candidate question set through a user query statement, the set including two types of candidate questions, one being a standard candidate question obtained by similarity matching with the user query statement, and the other being a fuzzy candidate question generalized from the standard candidate question. By screening these candidate questions, the target candidate question that best matches the user query statement can be found, and the target answer corresponding to the target candidate question can be fed back to the user. It can help users obtain the answers they need for their queries more quickly and accurately, provide users with more convenient search services, and solve the problems of rigid matching, low accuracy and low recall rates in existing FAQ systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] The disclosure of this application will be more easily understood with reference to the accompanying drawings. Those skilled in the art will readily appreciate that these drawings are for illustrative purposes only and are not intended to limit the scope of protection of this application. Furthermore, similar numbers in the figures represent similar components, where:
[0047] FIG1 is a flow chart showing the main steps of a question-answering method based on a FAQ system according to an embodiment of the present application;
[0048] FIG2 is a schematic diagram of a flow chart of steps for collecting candidate questions according to an embodiment of the present application;
[0049] FIG3 is a schematic diagram of a complete step flow chart of a question-answering method based on a FAQ system according to an embodiment of the present application;
[0050] FIG4 is a schematic diagram of a main structural block diagram of a computer device according to an embodiment of the present application. DETAILED DESCRIPTION
[0051] Some embodiments of the present application are described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principles of the present application and are not intended to limit the scope of protection of the present application.
[0052] In the description of this application, "module" and "processor" may include hardware, software, or a combination of both. A module may include hardware circuitry, various suitable sensors, communication ports, and memory. It may also include software components, such as program code, or a combination of software and hardware. A processor may be a central processing unit, a microprocessor, an image processor, a digital signal processor, or any other suitable processor. A processor has data and / or signal processing capabilities. A processor may be implemented in software, hardware, or a combination of both. Non-transitory computer-readable storage media include any suitable medium capable of storing program code, such as magnetic disks, hard disks, optical disks, flash memory, read-only memory, random access memory, etc. The term "A and / or B" refers to all possible combinations of A and B, such as only A, only B, or both A and B. The terms "at least one of A or B" or "at least one of A and B" have similar meanings to "A and / or B" and may include only A, only B, or both A and B. The singular forms "a" and "the" may also include the plural forms.
[0053] Traditional FAQ systems currently use a method of manually defining a representative question and an answer for a specific topic, then using a large amount of data to train a semantic matching model. However, while using a semantic matching model to determine the optimal question-answer pair can ensure the FAQ system's response speed, it often fails to meet matching accuracy requirements. Specifically, when literal similarity is low but semantic similarity is high, or when literal similarity is high but semantic similarity is low, the question-answer pair matched by the semantic matching model may not be the optimal one. Therefore, current FAQ dialogue solutions suffer from low matching accuracy.
[0054] The present application provides a question-answering method based on an FAQ system, comprising: obtaining a user query statement; obtaining a candidate question set based on the user query statement, the candidate question set including standard candidate questions and fuzzy candidate questions, wherein the standard candidate question is a candidate question obtained by similarity matching with the user query statement, and the fuzzy candidate question is a candidate question generalized from the standard candidate question; determining a target candidate question based on the candidate question set; and determining a target answer based on the target candidate question and feeding it back to the user. The present application obtains a candidate question set through a user query statement, the set including two types of candidate questions, one being a standard candidate question obtained by similarity matching with the user query statement, and the other being a fuzzy candidate question after generalization of the standard candidate question. By screening these candidate questions, the target candidate question that best matches the user query statement can be found, and the target answer corresponding to the target candidate question can be fed back to the user. This can help users obtain the answers they need for their queries more quickly and accurately, provide users with more convenient search services, and solve the problems of rigid matching, low accuracy and low recall rates in existing FAQ systems.
[0055] Please refer to Figure 1, which is a flowchart of the main steps of a question-answering method based on a FAQ system according to an embodiment of the present application.
[0056] As shown in FIG1 , the question-answering method based on the FAQ system in the embodiment of the present application mainly includes the following steps S101 to S104 .
[0057] Step S101: Obtain user query statements.
[0058] In this embodiment, the user query statement may be a query statement that is input into a search engine by a user in various ways and then converted into text form, such as keyboard input, voice input, and the like.
[0059] Step S102: obtaining a set of candidate questions based on the user query statement, wherein the set of candidate questions includes standard candidate questions and fuzzy candidate questions, wherein the standard candidate questions are candidate questions obtained by similarity matching with the user query statement, and the fuzzy candidate questions are candidate questions that are generalized versions of the standard candidate questions.
[0060] In this embodiment, a set of candidate questions including standard candidate questions and fuzzy candidate questions is obtained through user query statements. The standard candidate questions are obtained by similarity matching between the user query statements and multiple preset candidate questions, and the fuzzy candidate questions are candidate questions that are generalized expressions of the standard candidate questions.
[0061] Step S103: determining a target candidate question based on the candidate question set.
[0062] In this embodiment, the candidate questions in the candidate question set are screened to determine the target candidate question that best matches the user query statement.
[0063] Step S104: Based on the target candidate question, a target answer is determined and fed back to the user.
[0064] Based on the above steps S101 to S104, the present application obtains a set of candidate questions through user query statements. The set includes two types of candidate questions, one is a standard candidate question obtained by similarity matching with the user query statement, and the other is a fuzzy candidate question after generalization of the standard candidate question. By screening these candidate questions, the target candidate question that best matches the user query statement can be found, and the target answer corresponding to the target candidate question can be fed back to the user. It can help users obtain the answers they need for queries more quickly and accurately, provide users with more convenient search services, and solve the problems of rigid matching, low accuracy and low recall rates in existing FAQ systems.
[0065] The above steps S101 to S104 are further explained below.
[0066] With respect to step S101 , a user query statement is obtained.
[0067] In this embodiment, the user query statement can be a query statement that the user enters into the search engine in various ways and converted into text form, such as keyboard input, voice input, etc., and the query statement may be "What is the weather like today?" or "How to turn on the air conditioner."
[0068] With respect to step S102, in some embodiments, obtaining a set of candidate questions based on the user query statement includes: performing a recall on the user query statement through a preset retrieval model to obtain an initial question set; sorting the data in the initial question set from high to low according to similarity, and based on the sorting result, selecting a preset number of candidate questions with high similarity ranking in the initial question set to obtain a set of candidate questions.
[0069] Specifically, based on the user query statement, a preset retrieval model is used to obtain an initial question set, which includes standard candidate questions and fuzzy candidate questions, as well as the similarity between each candidate question and the user query statement; then the candidate questions in the initial question set are sorted from high to low by similarity, and a candidate question set is obtained based on multiple candidate questions ranked high in the sequence.
[0070] In some embodiments, the preset retrieval model includes a keyword retrieval model and a vector retrieval model, and the recall of the user query statement through the preset retrieval model to obtain an initial question set includes: based on the user query statement, obtaining a first initial question set through the keyword retrieval model; based on the user query statement, obtaining a second initial question set through the vector retrieval model; and obtaining an initial question set based on the first initial question set and the second initial question set.
[0071] Specifically, the preset retrieval model includes two retrieval models. The keyword retrieval model can be the Elasticsearch retrieval engine, and the vector retrieval model can be the Faiss vector retrieval library. Through the similarity algorithm built into the keyword retrieval model, the user query statement is matched with the candidate questions in the keyword retrieval model to obtain a first initial question set. The first initial question set includes a first standard candidate question, a first fuzzy candidate question, a first standard similarity, and a first fuzzy similarity. The first standard candidate question is a candidate question obtained by similarity matching with the user query statement, and the first fuzzy candidate question is a candidate question that is a generalized expression of the first standard candidate question. The first standard similarity and the first fuzzy similarity are determined based on the similarity algorithm built into the keyword retrieval model. The similarity algorithm can adopt the edit distance judgment algorithm or the BM25 algorithm, which is not limited here.
[0072] The user query statement is matched with the candidate questions in the vector retrieval model through the similarity algorithm built into the vector retrieval model to obtain a second initial question set, which includes a second standard candidate question, a second fuzzy candidate question, a second standard similarity, and a second fuzzy similarity; the second standard candidate question is a candidate question obtained by similarity matching with the user query statement, and the second fuzzy candidate question is a candidate question that is a generalized expression of the second standard candidate question; the second standard similarity and the second fuzzy similarity are determined based on the similarity algorithm built into the vector retrieval model, and the similarity algorithm can be a cosine similarity algorithm, a Manhattan / Euclidean distance algorithm, which is not limited here.
[0073] In some specific embodiments, the user query statement is first vectorized using the Sentence-BERT model. The Sentence-BERT model is a twin network based on the pre-trained BERT and can generate a vector for each input statement. The vectorized user query statement is then matched with the candidate questions in the vector retrieval model to obtain a second initial question set. It is understandable that when constructing the vector retrieval model, the Sentence-BERT model will also be used to vectorize the candidate questions, and then the vector retrieval model will be constructed based on the vectorized candidate questions.
[0074] In some embodiments, the keyword retrieval model is constructed based on standard questions, fuzzy questions and corresponding answers; wherein the standard questions and corresponding answers are collected in advance, and the fuzzy questions are generalized by the standard questions based on a preset text generation model to obtain multiple generalized questions corresponding to the standard questions, and the generalized questions and the standard questions are checked for similarity, and the generalized questions are obtained based on the similarity with the standard questions that meets preset conditions; and / or, the vector retrieval model is constructed based on the vectorized standard questions, the fuzzy questions and the corresponding answers.
[0075] Specifically, we first collect standard questions and answers corresponding to the question topic. There is no limit to the number of standard questions that can be collected for each question topic. Then, based on a preset text generation model, we generate text for the standard questions to obtain multiple generalized questions that are generalized expressions of the standard questions. We then check the similarity between the generalized questions and the standard questions, and classify the generalized questions whose similarity to the standard questions meets the preset criteria as fuzzy questions. It is understandable that the standard questions and the fuzzy questions that are generalized from the standard questions have the same answers. Finally, we construct a keyword retrieval model based on the standard questions, fuzzy questions, and their corresponding answers.
[0076] When building a vector retrieval model, standard questions and fuzzy questions are vectorized using the Sentence-BERT model, and a vector retrieval model is built based on the vectorized standard questions, fuzzy questions, and the corresponding answers.
[0077] In addition, text clustering can be used to utilize online data to obtain question topics that can be added to the retrieval model and candidate questions that can be added to the already constructed question topics, and possible question topics and candidate questions can be added to the retrieval model.
[0078] Please refer to Figure 2, which is a schematic flow chart of the steps of collecting candidate questions according to an embodiment of the present application.
[0079] As shown in Figure 2, collecting candidate questions and building a retrieval model includes the following steps:
[0080] Step S201: Collect standard questions and answers corresponding to the question topic. In this embodiment, each question topic is expected to obtain about 100 standard questions, which greatly expands the semantic coverage of the question topic.
[0081] Step S202: Generate similar text based on the standard question to obtain a generalized question of the standard question. In this embodiment, a preset text generation model is used to generalize the standard question to obtain various generalized statements of the standard question.
[0082] Step S203: A similarity check is performed between the generalized questions and the standard questions, and generalized questions whose similarity to the standard questions meets a preset condition are classified as fuzzy questions. In this embodiment, considering that the generalized statements generated by the model may be incoherent or have semantic changes, a similarity check is used to select generalized questions whose similarity to the standard questions meets a preset condition as fuzzy questions.
[0083] Step S204: Construct a keyword retrieval model based on the standard questions, fuzzy questions and corresponding answers.
[0084] Step S205: performing vectorization processing on the standard question and the fuzzy question;
[0085] Step S206: constructing a vector retrieval model based on the vectorized standard questions, fuzzy questions, and corresponding answers.
[0086] In some embodiments, in response to the target answer fed back to the user being no answer or a preset fallback answer, the method further includes: determining whether the similarity between the user query statement and the corresponding candidate question is greater than or equal to a preset similarity threshold; if so, adding the user query statement to the keyword retrieval model and the vector retrieval model.
[0087] Specifically, for some user query statements, there may be a situation where no answer or a preset fallback answer is fed back to the user. The highest similarity between the user query statement and the candidate question can be obtained, and the similarity of the candidate question can be determined to be greater than or equal to the preset similarity threshold. The preset similarity threshold is a conditional threshold for determining whether a candidate question can be added. When the similarity of the candidate question is greater than or equal to the preset similarity threshold, the user query statement is added as a new candidate question to the question topic to which the candidate question with the highest similarity belongs. For missed call issues where there is no answer or a preset fallback answer, the user query statement is added to the corresponding question topic through similarity judgment, achieving a quick fix for missed calls.
[0088] In some implementations, before sorting the data in the initial question set from high to low according to similarity, the method further includes: normalizing the similarity between the user query statements and the candidate questions in the initial question set.
[0089] Specifically, the similarity between user query statements and candidate questions recalled using keyword retrieval models and vector retrieval models may be calculated differently. Therefore, it is necessary to normalize the similarity between user query statements and candidate questions in the initial question set to facilitate subsequent sorting of the data in the initial question set.
[0090] For example, the similarity between user query statements recalled by the keyword retrieval model and candidate questions is generally calculated using the similarity algorithm built into the Elasticsearch retrieval engine, with a score range of 0 to 50. The similarity between user query statements recalled by the vector retrieval model and candidate questions is generally calculated using the cosine similarity algorithm, with a score range of 0 to 1. The normalization method can be adjusted based on the similarity ratio of different candidate questions. For example, an Elasticsearch similarity of 0 to 5 points corresponds to a cosine similarity algorithm similarity of 0 to 0.2. Statistics are used to determine an appropriate normalization method and normalize the similarity.
[0091] In some embodiments, before obtaining a set of candidate questions based on the user query statement, the method further includes: preprocessing the user query statement, wherein the preprocessing includes at least one of correcting format errors, removing non-keywords, replacing standard words, wake-up word processing, and vectorization processing.
[0092] Specifically, user query sentences may be relatively colloquial texts and may contain some greetings or other noise characters, so user query sentences need to be preprocessed. The preprocessing can be at least one of correcting format errors, removing non-keywords, standard word replacement, wake-up word processing, and vectorization processing. Among them, correcting format errors is to clean up punctuation, capitalization, spaces, and simple question corrections; removing non-keywords is to clean up the beginning words, cleaning up the ending words, and processing modal particles, such as: please ask, hi, yeah, wow, etc.; standard word replacement is to replace keywords and synonyms with standard expressions, such as: wiper water is replaced with glass water, roof light is replaced with reading light, etc. Wake-up word processing is the wake-up word processing after the first wake-up and the replacement of custom wake-up words. For wake-up word processing, the wake-up word resources can be automatically and regularly updated from the wake-up word database. Vectorization processing is to vectorize user query sentences through a preset vectorization processing model, such as the Sentence-BERT model.
[0093] With respect to step S103, in some embodiments, determining the target candidate question based on the candidate question set includes: forming a sentence pair with the user query sentence and the standard candidate question and the fuzzy candidate question in the candidate question set respectively; determining a target sentence pair based on the sentence pairs; and determining the target candidate question based on the similarity between the user query sentence and the corresponding candidate question in the target sentence pair.
[0094] Specifically, the user query statement is combined with each candidate question in the candidate question set to form a statement pair, including a statement pair consisting of a user query statement and a standard candidate question, and a statement pair consisting of a user query statement and a fuzzy candidate question; then a target statement pair is determined based on all statement pairs, and finally the matching target candidate question is determined based on the similarity between the user query statement and the corresponding candidate question in the target statement pair.
[0095] In some embodiments, determining the target sentence pair based on the sentence pair includes: determining the semantic matching score of each sentence pair based on a preset semantic recognition model; and taking the sentence pair with the highest semantic matching score as the target sentence pair.
[0096] Specifically, the preset semantic recognition model includes but is not limited to the BERT model, the RoBerta model, or the TinyBERT model. The semantic matching score predicted by the preset semantic recognition model is a decimal between 0 and 1. The preset semantic recognition model is used to predict the semantic matching score between the user query and the candidate question in each statement pair. All statement pairs are sorted by the semantic matching score, and the statement pair with the highest semantic matching score is selected as the target statement pair.
[0097] In some embodiments, determining the target candidate question based on the similarity between the user query statement in the target statement pair and the corresponding candidate question includes: performing weighted calculation on the similarity and semantic matching score between the user query statement of the target statement pair and the corresponding candidate question to determine the final score of the target statement pair; judging whether the final score is greater than or equal to a preset threshold; and determining that the candidate question corresponding to the target statement pair is the target candidate question when the final score is greater than or equal to the preset threshold.
[0098] Specifically, weights of similarity and semantic matching scores are set respectively. Based on the weights of similarity and semantic matching scores, the similarity and semantic matching scores of the user query statement in the target statement pair and the corresponding candidate question are weightedly calculated to obtain the final score of the target statement pair; then, it is determined whether the final score is greater than or equal to the preset threshold. If so, the candidate question in the target statement pair is used as the target candidate question.
[0099] With respect to step S104, in some embodiments, determining the target answer based on the target candidate question and feeding it back to the user includes: obtaining the target answer corresponding to the target candidate question; feeding back the target answer to the user in at least one of image, text, voice, and video; or, the method further includes: when the final score is less than a preset threshold, feeding back to the user that there is no answer in at least one of image, text, voice, and video, or feeding back a preset fallback answer to the user.
[0100] Specifically, the target answer corresponding to the target candidate question in the retrieval model is obtained; and the target answer is fed back to the user in at least one of the following ways: image, text, voice, and video.
[0101] If the final score is less than the preset threshold, it is considered that no answer has been obtained for the user's query statement, and the user is fed back that there is no answer through at least one of images, text, voice, and video, or the preset fallback answer is fed back to the user.
[0102] Please refer to FIG3 , which is a schematic diagram of a complete step flow chart of a question-answering method based on a FAQ system according to an embodiment of the present application.
[0103] As shown in FIG3 , in this embodiment, the question-answering method based on the FAQ system includes the following steps:
[0104] Step S301: Start.
[0105] Step S302: Obtain user query statements.
[0106] Step S303: Preprocess the user query. In this embodiment, the preprocessing includes at least one of correcting format errors, removing non-keywords, replacing standard words, processing wake-up words, and vectorizing.
[0107] Step S304: Recall the pre-processed user query statement through the keyword retrieval model to obtain a first initial question set.
[0108] Step S305: Perform recall on the pre-processed user query statement through the vector retrieval model to obtain a second initial question set.
[0109] Step S306: Obtain an initial question set based on the first initial question set and the second initial question set.
[0110] Step S307: Based on a preset number of candidate questions ranked high in similarity in the initial question set, a candidate question set is obtained. In this embodiment, the similarity between the user query and the candidate questions in the initial question set is first normalized. Then, the data in the initial question set is sorted from high to low based on similarity. Based on the sorting results, a preset number of candidate questions ranked high in similarity in the initial question set are selected to obtain a candidate question set. The candidate question set includes standard candidate questions, fuzzy candidate questions, and the similarity between the user query and the candidate questions.
[0111] Step S308: forming a statement pair based on the user query and each candidate question in the candidate question set. In this embodiment, the user query is formed into at least one statement pair with each of the standard candidate question and the fuzzy candidate question in the candidate question set.
[0112] Step S309: Determine the semantic matching score of each sentence pair based on the preset semantic recognition model.
[0113] Step S310: Determine a target sentence pair based on the semantic matching score of each sentence pair. In this embodiment, the sentence pair with the highest semantic matching score is used as the target sentence pair.
[0114] Step S311: Calculate the final score of the target sentence pair. In this embodiment, the similarity and semantic matching score of the user query sentence and the corresponding candidate question in the target sentence pair are weighted to determine the final score of the target sentence pair.
[0115] Step S312: Determine whether the final score is greater than or equal to a preset threshold. If so, execute step S313; if not, execute step S314.
[0116] Step S313: Feedback the target answer corresponding to the candidate question in the target sentence pair to the user.
[0117] Step S314: Feedback the user with no answer or a preset fallback answer.
[0118] Step S315: Determine whether the similarity corresponding to the candidate question in the target sentence pair is greater than or equal to a preset similarity threshold. If so, execute step S316; if not, execute step S317.
[0119] Step S316: Perform missed call repair. In this embodiment, missed call repair refers to determining whether the similarity between the user query and the corresponding candidate question in the target sentence pair is greater than or equal to a preset similarity threshold when no answer or a preset fallback answer is fed back to the user. If so, the user query is added to the keyword search model and the vector search model.
[0120] Step S317: End.
[0121] It should be pointed out that although the various steps in the above embodiments are described in a specific order, those skilled in the art will understand that in order to achieve the effect of the present application, different steps do not have to be performed in such an order. They can be performed simultaneously (in parallel) or in other orders. These changes are within the scope of protection of the present application.
[0122] Furthermore, the present application also provides a question-answering device based on the FAQ system.
[0123] The question-and-answer device based on the FAQ system in the embodiment of the present application mainly includes a first acquisition module 10, a second acquisition module 20, a determination module 30 and a feedback module 40. In some embodiments, one or more of the first acquisition module 10, the second acquisition module 20, the determination module 30 and the feedback module 40 can be combined into one module. In some embodiments, the first acquisition module 10 can be configured to obtain a user query statement. The second acquisition module 20 can be configured to obtain a set of candidate questions based on the user query statement. The determination module 30 can be configured to determine a target candidate question based on the candidate question set. The feedback module 40 can be configured to determine a target answer based on the target candidate question and feed it back to the user. In one embodiment, the description of the specific implementation function can be found in steps S101-S104.
[0124] The above-mentioned question-and-answer device based on the FAQ system is used to execute the embodiment of the question-and-answer method based on the FAQ system shown in Figure 1. The technical principles, technical problems solved and technical effects produced by the two are similar. Technical personnel in this technical field can clearly understand that for the convenience and conciseness of description, the specific working process and related instructions of the question-and-answer device based on the FAQ system can refer to the contents described in the embodiment of the question-and-answer method based on the FAQ system, and will not be repeated here.
[0125] It will be understood by those skilled in the art that all or part of the processes in the method for implementing the above embodiment of the present application can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, it can implement the steps of each of the above method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file or some intermediate form. The computer-readable storage medium can include: any entity or device, medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory, random access memory, electric carrier signal, telecommunication signal and software distribution medium that can carry the computer program code. It should be noted that the content contained in the computer-readable storage medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable storage media do not include electric carrier signals and telecommunication signals.
[0126] Furthermore, the present application also provides a computer device.
[0127] Please refer to FIG4 , which is a schematic diagram of the main structure block diagram of a computer device according to an embodiment of the present application.
[0128] As shown in FIG4 , the computer device in the embodiment of the present invention primarily includes a memory 11 and a processor 12. Memory 11 can be configured to store a program that executes the FAQ system-based question-answering method of the above-described method embodiment, and processor 12 can be configured to execute the program in the memory, including but not limited to a program that executes the FAQ system-based question-answering method of the above-described method embodiment. For ease of illustration, only the portion relevant to the embodiment of the present invention is shown. For specific technical details not disclosed, please refer to the method section of the embodiment of the present invention.
[0129] In a computer device embodiment according to the present application, the computer device includes a processor and a memory. The memory can be configured to store a program for executing the FAQ system-based question-answering method of the above-mentioned method embodiment, and the processor can be configured to execute the program in the memory, which includes but is not limited to a program for executing the FAQ system-based question-answering method of the above-mentioned method embodiment. For ease of explanation, only the parts related to the embodiment of the present application are shown. For specific technical details not disclosed, please refer to the method section of the embodiment of the present application. The computer device can be a computer device formed by various electronic devices.
[0130] In an embodiment of the present invention, the computer device may be a computer device formed by various electronic devices. In some possible implementations, the computer device may include multiple memories 11 and multiple processors 12. The program for executing the question-answering method based on the FAQ system of the above-mentioned method embodiment can be divided into multiple subroutines, and each subroutine can be loaded and run by the processor to execute different steps of the question-answering method based on the FAQ system of the above-mentioned method embodiment. Specifically, each subroutine can be stored in different memories 11 respectively, and each processor 12 can be configured to execute the program in one or more memories 11 to jointly implement the question-answering method based on the FAQ system of the above-mentioned method embodiment, that is, each processor 12 executes different steps of the question-answering method based on the FAQ system of the above-mentioned method embodiment respectively to jointly implement the question-answering method based on the FAQ system of the above-mentioned method embodiment.
[0131] The multiple processors 12 may be processors deployed on the same device. For example, the computer device may be a high-performance device composed of multiple processors, and the multiple processors 12 may be processors configured on the high-performance device. Furthermore, the multiple processors 12 may be processors deployed on different devices. For example, the computer device may be a server cluster, and the multiple processors 12 may be processors on different servers in the server cluster.
[0132] Furthermore, the present application also provides a computer-readable storage medium. In a computer-readable storage medium embodiment according to the present application, the computer-readable storage medium can be configured to store a program for executing the FAQ system-based question-and-answer method of the above-mentioned method embodiment, and the program can be loaded and run by the processor to implement the above-mentioned FAQ system-based question-and-answer method. For ease of explanation, only the parts related to the embodiment of the present application are shown. For specific technical details not disclosed, please refer to the method part of the embodiment of the present application. The computer-readable storage medium can be a storage device formed by various electronic devices. Optionally, the computer-readable storage medium in the embodiment of the present application is a non-temporary computer-readable storage medium.
[0133] Furthermore, it should be understood that since the configuration of each module is merely for the purpose of illustrating the functional units of the apparatus of the present application, the physical devices corresponding to these modules may be the processor itself, or a portion of the software in the processor, a portion of the hardware, or a combination of software and hardware. Therefore, the number of modules in the figure is merely illustrative.
[0134] Those skilled in the art will appreciate that the various modules in the device can be adaptively split or merged. Such splitting or merging of specific modules will not cause the technical solution to deviate from the principles of this application. Therefore, the technical solutions after splitting or merging will fall within the scope of protection of this application.
[0135] Thus far, the technical solutions of the present application have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is readily understood by those skilled in the art that the scope of protection of the present application is obviously not limited to these specific embodiments. Without departing from the principles of the present application, those skilled in the art may make equivalent changes or substitutions to the relevant technical features, and the technical solutions after such changes or substitutions will fall within the scope of protection of the present application.
Claims
1. A question-answering method based on a FAQ system, characterized in that: The method comprises: Get the user query statement; Acquire a set of candidate questions based on the user query, the set of candidate questions including standard candidate questions and fuzzy candidate questions, wherein the standard candidate questions are candidate questions obtained by similarity matching with the user query, and the fuzzy candidate questions are candidate questions that are generalized from the standard candidate questions; Determine a target candidate question based on the candidate question set; Based on the target candidate questions, a target answer is determined and fed back to the user.
2. The question-answering method based on the FAQ system according to claim 1, characterized in that: The obtaining of a candidate question set based on the user query statement includes: Recalling the user query statement through a preset retrieval model to obtain an initial question set; The data in the initial question set are sorted from high to low according to the similarity, and based on the sorting result, a preset number of candidate questions with high similarity rankings in the initial question set are selected to obtain a candidate question set.
3. The question-answering method based on the FAQ system according to claim 2, characterized in that: The preset retrieval model includes a keyword retrieval model and a vector retrieval model; the recalling of the user query statement through the preset retrieval model to obtain an initial question set includes: Based on the user query statement, obtaining a first initial question set through a keyword retrieval model; Based on the user query statement, obtaining a second initial question set through a vector retrieval model; An initial question set is obtained based on the first initial question set and the second initial question set.
4. The question-answering method based on the FAQ system according to claim 3, characterized in that: The keyword search model is constructed based on standard questions, fuzzy questions and corresponding answers; The standard questions and the corresponding answers are collected in advance, the fuzzy questions are obtained by generalizing the standard questions based on a preset text generation model to obtain a plurality of generalized questions corresponding to the standard questions, and the generalized questions and the standard questions are checked for similarity, and the generalized questions are obtained based on the similarity with the standard questions that meets the preset conditions; and / or, The vector retrieval model is constructed based on the vectorized standard questions, the fuzzy questions and the corresponding answers.
5. The question-answering method based on the FAQ system according to claim 3, characterized in that: In response to the target answer fed back to the user being no answer or a preset fallback answer, the method further includes: Determine whether the similarity between the user query statement and the corresponding candidate question is greater than or equal to a preset similarity threshold; if so, add the user query statement to the keyword retrieval model and the vector retrieval model.
6. The question-answering method based on the FAQ system according to claim 2, characterized in that: Before sorting the data in the initial question set from high to low according to similarity, the method further includes: The similarities between the user query statements and the candidate questions in the initial question set are normalized.
7. The question-answering method based on the FAQ system according to claim 1, characterized in that: Before obtaining a candidate question set based on the user query statement, the method further includes: The user query statement is preprocessed, and the preprocessing includes at least one of correcting format errors, removing non-keywords, replacing standard words, processing wake-up words, and vectorization processing.
8. The question-answering method based on the FAQ system according to claim 1, characterized in that: The determining a target candidate question based on the candidate question set includes: The user query sentence is combined with the standard candidate question and the fuzzy candidate question in the candidate question set to form a sentence pair; determining a target sentence pair based on the sentence pair; The target candidate question is determined based on the similarity between the user query sentence and the corresponding candidate question in the target sentence pair.
9. The question-answering method based on the FAQ system according to claim 8, characterized in that: The determining a target sentence pair based on the sentence pair comprises: Determine the semantic matching score of each sentence pair based on a preset semantic recognition model; The sentence pair with the highest semantic matching score is taken as the target sentence pair.
10. The question-answering method based on the FAQ system according to claim 9, characterized in that: The determining the target candidate question based on the similarity between the user query sentence and the corresponding candidate question in the target sentence pair includes: Performing weighted calculation on the similarity and semantic matching score between the user query sentence of the target sentence pair and the corresponding candidate question to determine the final score of the target sentence pair; Determine whether the final score is greater than or equal to a preset threshold; When the final score is greater than or equal to a preset threshold, the candidate question corresponding to the target sentence pair is determined to be the target candidate question.
11. The question-answering method based on the FAQ system according to claim 10, characterized in that: The step of determining a target answer based on the target candidate question and feeding it back to the user includes: Obtaining a target answer corresponding to the target candidate question; Feedback the target answer to the user in at least one of image, text, voice, and video; or, The method further comprises: When the final score is less than a preset threshold, the user is fed back that there is no answer in at least one of image, text, voice, and video, or a preset fallback answer is fed back to the user.
12. A computer device comprising a processor and a memory, wherein the memory is suitable for storing a plurality of program codes, wherein: The program code is suitable for being loaded and run by the processor to execute the question-answering method based on the FAQ system according to any one of claims 1 to 11.
13. A computer-readable storage medium storing a plurality of program codes, characterized in that: The program code is suitable for being loaded and run by a processor to execute the question-answering method based on a FAQ system according to any one of claims 1 to 11.
Citation Information
Patent Citations
Intelligent question and answer method, device and equipment and storage medium
CN114416927A
Artificial intelligence question and answer model generation method, question and answer method and device and storage medium
CN114896382A
Multi-scene intelligent question answering method and system based on multi-path recall
CN115470338A
Text matching method and device and electronic equipment
CN116127005A
Question and answer method based on FAQ system, computer equipment and storage medium
CN117763108A
Cited By
Dynamic knowledge retrieval enhancement method based on large language model
CN120407570A
User input statement processing method and device and medium
CN120596505A
Text2SQL (Structured Query Language) medical data processing method based on large language model and electronic equipment
CN120723807A
Large model data scheduling method and system based on natural language and data view
CN121116961A
Cross-language retrieval method
CN121117156A