Query method, device, equipment and storage medium for question-and-answer data
By determining the positive and negative Q&A examples of the question statements in the corpus, and selecting the answer corpus using the similarity threshold, the problem of low accuracy of Q&A data query is solved, and more accurate query results are achieved.
Patent Information
- Application Number
- CN202111200482.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-14
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2041-10-14
AI Technical Summary
The accuracy of Q&A data query is low. The existing Q&A model outputs incorrect query results due to limited training corpus.
Determine the positive Q&A example and negative Q&A example corresponding to the question statement in the corpus, and judge through the similarity threshold, select the answer corpus of the positive Q&A example as the query result, and expand the corpus to store and associate the positive and negative Q&A examples.
It improves the accuracy of Q&A data query, avoids the output of wrong results of the Q&A model, and enhances the effectiveness of the query.
Smart Images

Figure CN114003688B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to a method, device, equipment, and computer-readable storage medium for querying question-and-answer data. Background Art
[0002] Intelligent customer service systems have been widely used by Internet companies and have achieved good application results. The main technical solution of the intelligent customer service system is to train a question-and-answer model based on question-and-answer data and return the answer corpus determined by the question-and-answer model to the user. However, due to limited training corpus, the question-and-answer model may output incorrect query results, resulting in users being unable to obtain effective answers. Summary of the Invention
[0003] The main purpose of the present invention is to provide a method, device, equipment, and computer-readable storage medium for querying question-and-answer data, aiming to solve the problem of low accuracy in querying question-and-answer data.
[0004] To achieve the above object, a method for querying question-and-answer data provided by the present invention includes the following steps:
[0005] Obtain a question statement, and determine a positive question-and-answer example and a negative question-and-answer example corresponding to the question statement in a corpus;
[0006] If the first similarity between the question statement and the query corpus in the positive question-and-answer example is greater than a preset first threshold, and the second similarity between the question statement and the query corpus in the negative question-and-answer example is less than a preset second threshold, then use the answer corpus of the positive question-and-answer example as the query result corresponding to the question statement.
[0007] In an embodiment, before the step of determining the positive question-and-answer example and the negative question-and-answer example corresponding to the question statement in the corpus, the method further includes:
[0008] Obtain the question statements corresponding to the incorrect query results output by a preset question-and-answer model within a preset time interval;
[0009] Input the question statement into a preset first corpus model to obtain a positive question-and-answer example;
[0010] Input the question statement into a preset second corpus model to obtain a negative question-and-answer example;
[0011] Associate and store the question statement, the positive question-and-answer example, and the negative question-and-answer example in the corpus.
[0012] In an embodiment, after the step of obtaining the question statements corresponding to the incorrect query results output by a preset question-and-answer model within a preset time interval, the method further includes:
[0013] Determine the relevance between the question statement and a preset standard Q&A pair;
[0014] If the relevance is greater than a preset third threshold, perform the step of inputting the question statement into a preset first corpus model to obtain the positive Q&A example.
[0015] In one embodiment, the step of determining the relevance between the question statement and a preset standard Q&A pair includes:
[0016] Determine a first similarity value between the semantic vector corresponding to the question statement and the semantic vector of the query corpus in the standard Q&A pair;
[0017] Generate an answer corpus corresponding to the question statement according to a preset semantic model;
[0018] Determine a second similarity value between the semantic vector corresponding to the answer corpus of the question statement and the semantic vector of the answer corpus in the standard Q&A pair;
[0019] Determine the relevance according to a preset weight value, the first similarity value, and the second similarity value.
[0020] In one embodiment, the step of associatively storing the question statement, the positive Q&A example, and the negative Q&A example in the corpus includes:
[0021] Determine the cosine value between the semantic vector of the positive Q&A example and a random vector, and determine a first storage location for the positive Q&A example according to the cosine value;
[0022] Determine a second storage location for the negative Q&A example according to the first storage location;
[0023] Determine the association relationship between the positive Q&A example and the negative Q&A example according to the first storage location and the second storage location;
[0024] Associatively store the positive Q&A example, the negative Q&A example, and the question statement in the corpus according to the association relationship.
[0025] In one embodiment, before the step of determining the positive Q&A example and the negative Q&A example corresponding to the question statement in the corpus, it further includes:
[0026] Input the question statement into a preset Q&A model to obtain a query result, and obtain a score corresponding to the query result;
[0027] If the score is less than a preset threshold, perform the step of determining the positive Q&A example and the negative Q&A example corresponding to the question statement in the corpus.
[0028] In one embodiment, after the step of using the answer corpus of the positive Q&A example as the query result corresponding to the question statement, the following steps are further included:
[0029] Using the answer corpus of the positive Q&A example and the question statement as training samples;
[0030] Retraining a preset Q&A model according to the training samples;
[0031] Saving the retrained preset Q&A model.
[0032] To achieve the above object, the present invention further provides a query device for Q&A data, and the query device for Q&A data includes:
[0033] An acquisition module, configured to acquire a question statement, and determine a positive Q&A example and a negative Q&A example corresponding to the question statement in a corpus;
[0034] A query module, configured to, if a first similarity between the question statement and the query corpus in the positive Q&A example is greater than a preset first threshold, and a second similarity between the question statement and the query corpus in the negative Q&A example is less than a preset second threshold, use the answer corpus of the positive Q&A example as the query result corresponding to the question statement.
[0035] To achieve the above object, the present invention further provides a query device for Q&A data, and the query device for Q&A data includes a memory, a processor, and a Q&A data query program stored in the memory and executable on the processor. When the Q&A data query program is executed by the processor, each step of the above-mentioned Q&A data query method is implemented.
[0036] To achieve the above object, the present invention further provides a computer-readable storage medium, which stores a Q&A data query program. When the Q&A data query program is executed by a processor, each step of the above-mentioned Q&A data query method is implemented.
[0037] A query method, device, equipment and computer-readable storage medium for question-and-answer data provided by the present invention obtain a question statement, and determine a positive question-and-answer example and a negative question-and-answer example corresponding to the question statement in a corpus; if the first similarity between the question statement and the query corpus in the positive question-and-answer example is greater than a preset first threshold, and the second similarity between the question statement and the query corpus in the negative question-and-answer example is less than a preset second threshold, then the answer corpus of the positive question-and-answer example is used as the query result corresponding to the question statement. By determining the positive question-and-answer example and the negative question-and-answer example corresponding to the question statement in the corpus, the query result corresponding to the question statement is obtained, and the accuracy of querying question-and-answer data is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 It is a schematic hardware structure diagram of a query device for question-and-answer data according to an embodiment of the present invention;
[0039] Figure 2 It is a schematic flowchart of the first embodiment of the query method for question-and-answer data of the present invention;
[0040] Figure 3 It is a schematic flowchart of the second embodiment of the query method for question-and-answer data of the present invention;
[0041] Figure 4 It is a schematic flowchart of the third embodiment of the query method for question-and-answer data of the present invention;
[0042] Figure 5 It is a schematic flowchart of step S60 in the fourth embodiment of the query method for question-and-answer data of the present invention;
[0043] Figure 6 It is a schematic diagram of the storage locations of positive and negative question-and-answer examples in the corpus of the present invention;
[0044] Figure 7 It is a schematic flowchart of the fifth embodiment of the query method for question-and-answer data of the present invention;
[0045] Figure 8 It is a schematic logical structure diagram of the query device for question-and-answer data of the present invention.
[0046] The realization, functional features and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0048] The main solution of the embodiment of the present invention is: obtain a question statement, and determine the positive Q&A example and the negative Q&A example corresponding to the question statement in the corpus; if the first similarity between the question statement and the query corpus in the positive Q&A example is greater than a preset first threshold, and the second similarity between the question statement and the query corpus in the negative Q&A example is less than a preset second threshold, then use the answer corpus of the positive Q&A example as the query result corresponding to the question statement.
[0049] By determining the positive Q&A example and the negative Q&A example corresponding to the question statement in the corpus, the query result corresponding to the question statement is obtained, improving the accuracy of Q&A data query.
[0050] As an implementation solution, the query device for Q&A data can be as Figure 1 shown.
[0051] The solution of the embodiment of the present invention relates to a query device for Q&A data. The query device for Q&A data includes: a processor 101, such as a CPU, a memory 102, and a communication bus 103. Among them, the communication bus 103 is used to realize the connection and communication between these components.
[0052] The memory 102 can be a high-speed RAM memory or a stable memory (non-volatile memory), such as a disk memory. As Figure 1 shown, the memory 102, as a computer-readable storage medium, may include a query program for Q&A data; and the processor 101 can be used to call the query program for Q&A data stored in the memory 102 and perform the following operations:
[0053] Obtain a question statement, and determine the positive Q&A example and the negative Q&A example corresponding to the question statement in the corpus;
[0054] If the first similarity between the question statement and the query corpus in the positive Q&A example is greater than a preset first threshold, and the second similarity between the question statement and the query corpus in the negative Q&A example is less than a preset second threshold, then use the answer corpus of the positive Q&A example as the query result corresponding to the question statement.
[0055] In an embodiment, the processor 101 can be used to call the query program for Q&A data stored in the memory 102 and perform the following operations:
[0056] Obtain the question statements corresponding to the wrong query results output by a preset Q&A model within a preset time interval;
[0057] Input the question statements into a preset first corpus model to obtain positive Q&A examples;
[0058] Input the question statement into a preset second corpus model to obtain a negative Q&A example;
[0059] Associate and store the question statement, the positive Q&A example, and the negative Q&A example in the corpus.
[0060] In one embodiment, the processor 101 can be used to call the query program of the Q&A data stored in the memory 102 and perform the following operations:
[0061] Determine the relevance between the question statement and a preset standard Q&A pair;
[0062] If the relevance is greater than a preset third threshold, execute the step of inputting the question statement into a preset first corpus model to obtain the positive Q&A example.
[0063] In one embodiment, the processor 101 can be used to call the query program of the Q&A data stored in the memory 102 and perform the following operations:
[0064] Determine the first similarity value between the semantic vector corresponding to the question statement and the semantic vector of the query corpus in the standard Q&A pair;
[0065] Generate the answer corpus corresponding to the question statement according to a preset semantic model;
[0066] Determine the second similarity value between the semantic vector corresponding to the answer corpus of the question statement and the semantic vector of the answer corpus in the standard Q&A pair;
[0067] Determine the relevance according to a preset weight value, the first similarity value, and the second similarity value.
[0068] In one embodiment, the processor 101 can be used to call the query program of the Q&A data stored in the memory 102 and perform the following operations:
[0069] Determine the cosine value between the semantic vector of the positive Q&A example and a random vector, and determine the first storage location of the positive Q&A example according to the cosine value;
[0070] Determine the second storage location of the negative Q&A example according to the first storage location;
[0071] Determine the association relationship between the positive Q&A example and the negative Q&A example according to the first storage location and the second storage location;
[0072] Associate and store the positive Q&A example, the negative Q&A example, and the question statement in the corpus according to the association relationship.
[0073] In one embodiment, the processor 101 may be used to call the query program of the Q&A data stored in the memory 102 and perform the following operations:
[0074] Input the question statement into a preset Q&A model to obtain a query result, and obtain the score corresponding to the query result;
[0075] If the score is less than a preset threshold, then perform the step of determining the positive Q&A example and the negative Q&A example corresponding to the question statement in the corpus.
[0076] In one embodiment, the processor 101 may be used to call the query program of the Q&A data stored in the memory 102 and perform the following operations:
[0077] Use the answer corpus of the positive Q&A example and the question statement as training samples;
[0078] Retrain the preset Q&A model according to the training samples;
[0079] Save the retrained preset Q&A model.
[0080] Based on the hardware architecture of the above Q&A data query device, an embodiment of the Q&A data query method of the present invention is proposed.
[0081] Refer to Figure 2 , Figure 2 This is the first embodiment of the Q&A data query method of the present invention. The Q&A data query method includes the following steps:
[0082] Step S10, obtain a question statement, and determine the positive Q&A example and the negative Q&A example corresponding to the question statement in the corpus.
[0083] Specifically, the question statement is the text information when the user makes a query or asks a question. The corpus contains the positive Q&A example and the negative Q&A example corresponding to the question statement. Among them, the positive Q&A example refers to the associated data between the question statement and the correct answer corpus. For example, the question statement of the positive Q&A example is "How to unsubscribe from the membership", and the correct answer corpus is "Click on my membership and then click on the unsubscribe button". The negative Q&A example is the associated data between the question statement and the wrong answer corpus. For example, the question statement of the negative Q&A example is "How to unsubscribe from the membership", and the wrong answer corpus is "The membership has the following rights and interests".
[0084] Step S20, if the first similarity between the question statement and the query corpus in the positive Q&A example is greater than a preset first threshold, and the second similarity between the question statement and the query corpus in the negative Q&A example is less than a preset second threshold, then use the answer corpus of the positive Q&A example as the query result corresponding to the question statement.
[0085] Specifically, when the first similarity between the query corpus of the question statement and the positive Q&A example is greater than a preset first threshold, it indicates that the similarity between the question statement and the positive Q&A example is relatively high. When the second similarity between the query corpus of the question statement and the negative Q&A example is less than a preset second threshold, it indicates that the similarity between the question statement and the negative Q&A example is relatively low. After meeting the above conditions, the answer corpus of the positive Q&A example is used as the query result corresponding to the question statement. Among them, to determine the first similarity between the query corpus of the question statement and the positive Q&A example, it can be to calculate the first cosine distance between the semantic vector of the question statement and the semantic vector of the query corpus in the positive Q&A example, and determine the first similarity according to the first cosine distance. To determine the second similarity between the query corpus of the question statement and the negative Q&A example, it can be to calculate the second cosine distance between the semantic vector of the question statement and the semantic vector of the query corpus in the negative Q&A example, and determine the second similarity according to the second cosine distance.
[0086] In the technical solution of this embodiment, a question statement is obtained, and a positive Q&A example and a negative Q&A example corresponding to the question statement are determined in the corpus; if the first similarity between the query corpus of the question statement and the positive Q&A example is greater than a preset first threshold, and the second similarity between the query corpus of the question statement and the negative Q&A example is less than a preset second threshold, then the answer corpus of the positive Q&A example is used as the query result corresponding to the question statement. By determining the positive Q&A example and the negative Q&A example corresponding to the question statement in the corpus, the query result corresponding to the question statement is obtained, improving the accuracy of the Q&A data query.
[0087] Refer to Figure 3 , Figure 3 This is the second embodiment of the query method for Q&A data of the present invention. Based on the first embodiment, before step S10, it further includes:
[0088] Step S30, obtaining the question statements corresponding to the incorrect query results output by a preset Q&A model within a preset time interval;
[0089] Step S40, inputting the question statement into a preset first corpus model to obtain a positive Q&A example;
[0090] Step S50, inputting the question statement into a preset second corpus model to obtain a negative Q&A example;
[0091] Step S60, associatively storing the question statement, the positive Q&A example, and the negative Q&A example in the corpus.
[0092] Specifically, obtain multiple question statements corresponding to the incorrect query results output by a preset Q&A model within a preset time interval. The preset Q&A model is used to output corresponding query results based on the user's question statements. For example, when the input question statement is "How to switch the camera", the output query result can be "Click the camera switching button". However, there are cases where the preset Q&A model outputs incorrect query results. For example, when the input question statement is "How to switch the camera", the output query result is "Click the service cancellation button". The determination of incorrect query results can obtain the user's rating of the query result. When the rating is less than the preset threshold, the query result is regarded as an incorrect query result.
[0093] Input the question statement into a preset first corpus model to obtain a positive Q&A example. The first corpus model is used to output positive Q&A examples based on the question statement, and the first corpus model is trained by a standard Q&A pair training set and an existing extended corpus.
[0094] Before inputting the question statement into the preset first corpus model to obtain a positive Q&A example, the Bert model can be used to transform the question statement into a preliminary semantic vector with a unified length, and determine the key points in the question statement according to the preset positive Attention function, that is, transform the preliminary semantic vector into a positive semantic vector, as shown in the following formula:
[0095] V r = attn(V t , V i )·V i ;
[0096] Among them, V r is the transformed positive semantic vector, V t is a randomly generated vector, and V i is the preliminary semantic vector to be transformed.
[0097] Input the transformed positive semantic vector into the first corpus model to obtain the positive Q&A example output by the first corpus model.
[0098] Input the question statement into a preset second corpus model to obtain a negative Q&A example. The second corpus model is used to output negative Q&A examples based on the question statement. The second corpus model is trained by a standard Q&A pair training set and an existing extended corpus. The different questions in the existing corpus can be shuffled and recombined as training data for training.
[0099] Before inputting the question statement into the preset second corpus model to obtain a negative Q&A example, the Bert model can be used to transform the question statement into a preliminary semantic vector with a unified length, and determine the non-key points in the question statement according to the preset negative Attention function, that is, transform the preliminary semantic vector into a negative semantic vector.
[0100]
[0101] Among them, V w is the transformed negative semantic vector, V t is a randomly generated vector, and V i is the preliminary semantic vector to be transformed.
[0102] Input the transformed negative semantic vector into the second corpus model to obtain the negative Q&A examples output by the second corpus model.
[0103] After obtaining the positive Q&A examples and negative Q&A examples corresponding to the question statement, associate and store the question statement, positive Q&A examples, and negative Q&A examples in the corpus. After obtaining the question statement input by the user, the positive Q&A examples and negative Q&A examples corresponding to the question statement can be obtained according to the association relationship.
[0104] In the technical solution of this embodiment, obtain the question statement corresponding to the wrong query result output by the preset Q&A model within a preset time interval; input the question statement into the preset first corpus model to obtain positive Q&A examples; input the question statement into the preset second corpus model to obtain negative Q&A examples; associate and store the question statement, positive Q&A examples, and negative Q&A examples in the corpus. By expanding the corpus of the question statement of the wrong query result, determine the positive Q&A examples and negative Q&A examples corresponding to the query result, so as to determine the query result corresponding to the question statement in the corpus, avoid the situation where the preset Q&A model cannot output the correct query result, and improve the accuracy of Q&A data query.
[0105] Refer to Figure 4 , Figure 4 This is the third embodiment of the query method for the Q&A data of the present invention. Based on the second embodiment, after the step S30, it further includes:
[0106] Step S70, determine the relevance between the question statement and the preset standard Q&A pair;
[0107] Step S80, if the relevance is greater than a preset third threshold, then execute the step of inputting the question statement into the preset first corpus model to obtain the positive Q&A examples.
[0108] Specifically, the standard Q&A pair is a set of question statements and answer corpora. Determining the relevance between the question statement and the query corpus in the preset standard Q&A pair can be to only determine the relevance between the question statement and the query corpus in the preset standard Q&A pair, and use the Word2Vec model to encode the question statement or the standard Q&A pair to obtain semantic vectors. As shown in the following formula:
[0109]
[0110] Among them, Sim is the relevance, Q t is the semantic vector of the question statement, Q i is the semantic vector of the question statement of the i-th standard Q&A pair, m is the total number of question statements of the standard Q&A pair, and attn() is the Attention function.
[0111] To determine the relevance between the question statement and the preset standard Q&A pair, the first similarity value between the semantic vector corresponding to the question statement and the semantic vector of the query corpus in the standard Q&A pair can also be determined first, as shown in the following formula:
[0112]
[0113] Among them, S1 is the first similarity value, Q t is the semantic vector of the question statement, Q i is the semantic vector of the question statement of the i-th standard Q&A pair, m is the total number of question statements of the standard Q&A pair, and attn() is the Attention function.
[0114] Generate the answer corpus corresponding to the question statement according to the preset semantic model. Among them, the semantic model is used to generate answers according to the user's question statement. Different from the Q&A model, the semantic model is trained with a general dataset of non-Q&A data, such as 27 million text datasets crawled from the web. The preset semantic model is the MGPT2 model. Different from the classic GPT2 model, the MGPT2 model is trained with multiple question statements and one answer, and an LSTM model is added to each layer on the basis of the GPT2 model to fuse the generated vectors of multiple question statements. Therefore, the MGPT2 model can not only learn the association between the question statement and the answer, but also learn the association between the question statements, so that the MGPT2 model can recognize different input question statements and generate similar content. For example, when the user inputs "Don't want the membership anymore" and "The membership is not practical", the MGPT2 model can output similar answer expressions such as "Just press the unsubscribe button" and "Click to unsubscribe".
[0115] Encode the answer corpus corresponding to the question statement using the Word2Vec model to obtain the semantic vector. Determine the second similarity value between the semantic vector corresponding to the answer corpus corresponding to the question statement and the semantic vector of the answer corpus in the standard Q&A pair;
[0116]
[0117] Among them, S2 is the second similarity value, A t is the semantic vector of the answer corpus corresponding to the question statement, A jIt is the semantic vector of the answer corpus for the j-th standard Q&A pair, n is the total number of answer corpora of the standard Q&A pairs, and attn() is the Attention function.
[0118] Determine the relevance according to the preset weight value, the first similarity value and the second similarity value.
[0119] Sim = α×S1+(1 - α)×S2;
[0120] Where, Sim is the relevance, S1 is the first similarity value, S2 is the second similarity value, and α is the preset weight value.
[0121] When the relevance is greater than the preset third threshold, it indicates that the question statement is relevant to the standard Q&A pair, and perform the step of inputting the question statement into the preset first corpus model to obtain a positive Q&A example. When the relevance is less than or equal to the preset third threshold, it indicates that the question statement is irrelevant to the existing standard Q&A pairs, and then it is necessary to expand the standard Q&A pairs corresponding to the preset Q&A model.
[0122] In the technical solution of this embodiment, determine the relevance between the question statement and the preset standard Q&A pair; if the relevance is greater than the preset third threshold, then perform the step of inputting the question statement into the preset first corpus model to obtain a positive Q&A example. Determine whether the question statement is relevant to the standard Q&A pair according to the relevance. When the question statement is relevant to the standard Q&A pair and the preset Q&A model outputs an incorrect query result, it indicates that the question statement needs to be expanded to avoid the situation of incorrect query results caused by the small number of standard Q&A pairs.
[0123] Refer to Figure 5 , Figure 5 This is the fourth embodiment of the query method for the Q&A data of the present invention. Based on the second embodiment, the step S60 includes:
[0124] Step S61, determine the cosine value between the semantic vector of the positive Q&A example and the random vector, and determine the first storage position of the positive Q&A example according to the cosine value;
[0125] Step S62, determine the second storage position of the negative Q&A example according to the first storage position;
[0126] Step S63, determine the association relationship between the positive Q&A example and the negative Q&A example according to the first storage position and the second storage position;
[0127] Step S64, store the positive Q&A example, the negative Q&A example and the question statement in the corpus in an associated manner according to the association relationship.
[0128] Specifically, the random vector is a randomly generated vector. Determine the cosine value of the semantic vector of the positive Q&A example and the random vector, and determine the first storage location of the positive Q&A example according to the cosine value.
[0129] Loc = crc64(cos(V0, V)) % 8192;
[0130] Where Loc represents the first storage location, crc() is a function that converts the cosine value to an integer using the CRC64 algorithm, and % is the modulo operation.
[0131] After determining the first storage location, determine the second storage location of the negative Q&A example according to the first storage location. You can add the number of slots maintained by the preset node to the slot number of the first storage location and store the negative Q&A example in the next node of the first storage location. Determine the association relationship between the positive Q&A example and the negative Q&A example according to the first storage location and the second storage location. Exemplarily, store the slot number of the first storage location of the positive Q&A example and the slot number of the negative Q&A example as mapping data into the node. As Figure 6 shown, store the positive Q&A example, the negative Q&A example, and the question statement in the corpus in an associated manner according to the association relationship.
[0132] In the technical solution of this embodiment, determine the cosine value of the semantic vector of the positive Q&A example and the random vector, and determine the first storage location of the positive Q&A example according to the cosine value; determine the second storage location of the negative Q&A example according to the first storage location; determine the association relationship between the positive Q&A example and the negative Q&A example according to the first storage location and the second storage location; store the positive Q&A example, the negative Q&A example, and the question statement in the corpus in an associated manner according to the association relationship. Through the associated storage of the positive Q&A example, the negative Q&A example, and the question statement, the positive Q&A example and the negative Q&A example are evenly distributed in the corpus according to their semantic information to achieve load balancing of storage and access.
[0133] Refer to Figure 7 , Figure 7 This is the fifth embodiment of the query method for the Q&A data of the present invention. Based on any one of the first to fourth embodiments, before step S10, it further includes:
[0134] Step S90, input the question statement into a preset Q&A model to obtain a query result, and obtain the score corresponding to the query result;
[0135] Step S100, if the score is less than a preset threshold, then execute the step of determining the positive Q&A example and the negative Q&A example corresponding to the question statement in the corpus.
[0136] Specifically, before determining the positive Q&A examples and negative Q&A examples corresponding to the question statement in the corpus, obtain the question statement input by the user, input the question statement into the preset Q&A model to obtain the query result, and obtain the score corresponding to the query result. If the score is less than the preset threshold, it means that the obtained query result is an incorrect query result or the user is not satisfied with the current query result, and it is necessary to determine the positive Q&A examples and negative Q&A examples corresponding to the question statement in the corpus.
[0137] After using the answer corpus of the positive Q&A example as the query result corresponding to the question statement, the answer corpus of the positive Q&A example and the question statement can also be used as training samples to retrain the preset Q&A model according to the training samples, and save the retrained preset Q&A model.
[0138] In the technical solution of this embodiment, input the question statement into the preset Q&A model to obtain the query result, and obtain the score corresponding to the query result; if the score is less than the preset threshold, execute the step of determining the positive Q&A examples and negative Q&A examples corresponding to the question statement in the corpus. By judging whether the query result corresponding to the question statement output by the preset Q&A model is correct before determining the positive Q&A examples and negative Q&A examples corresponding to the question statement in the corpus, it is avoided to query the question statement in the corpus when the Q&A model outputs the correct query result, which improves the efficiency of Q&A data query.
[0139] Refer to Figure 8 , the present invention also provides a query device for Q&A data, and the query device for Q&A data includes:
[0140] An acquisition module 100, configured to acquire a question statement, and determine positive Q&A examples and negative Q&A examples corresponding to the question statement in a corpus;
[0141] A query module 200, configured to, if a first similarity between the question statement and query corpus in the positive Q&A example is greater than a preset first threshold, and a second similarity between the question statement and query corpus in the negative Q&A example is less than a preset second threshold, use the answer corpus of the positive Q&A example as the query result corresponding to the question statement.
[0142] In one embodiment, before determining the positive Q&A examples and negative Q&A examples corresponding to the question statement in the corpus, the acquisition module 100 is specifically configured to:
[0143] Acquire question statements corresponding to incorrect query results output by the preset Q&A model within a preset time interval;
[0144] Input the question statement into a preset first corpus model to obtain positive Q&A examples;
[0145] Input the question statement into a preset second corpus model to obtain a negative Q&A example;
[0146] Associate and store the question statement, the positive Q&A example, and the negative Q&A example in the corpus.
[0147] In one embodiment, after obtaining the question statement corresponding to the error query result output by the preset Q&A model within a preset time interval, the obtaining module 100 is specifically configured to:
[0148] Determine the relevance between the question statement and a preset standard Q&A pair;
[0149] If the relevance is greater than a preset third threshold, execute the step of inputting the question statement into a preset first corpus model to obtain the positive Q&A example.
[0150] In one embodiment, in terms of determining the relevance between the question statement and a preset standard Q&A pair, the obtaining module 100 is specifically configured to:
[0151] Determine a first similarity value between the semantic vector corresponding to the question statement and the semantic vector of the query corpus in the standard Q&A pair;
[0152] Generate an answer corpus corresponding to the question statement according to a preset semantic model;
[0153] Determine a second similarity value between the semantic vector corresponding to the answer corpus of the question statement and the semantic vector of the answer corpus in the standard Q&A pair;
[0154] Determine the relevance according to a preset weight value, the first similarity value, and the second similarity value.
[0155] In one embodiment, in terms of associatively storing the question statement, the positive Q&A example, and the negative Q&A example in the corpus, the obtaining module 200 is specifically configured to:
[0156] Determine the cosine value between the semantic vector of the positive Q&A example and a random vector, and determine a first storage location of the positive Q&A example according to the cosine value;
[0157] Determine a second storage location of the negative Q&A example according to the first storage location;
[0158] Determine the association relationship between the positive Q&A example and the negative Q&A example according to the first storage location and the second storage location;
[0159] Associatively store the positive Q&A example, the negative Q&A example, and the question statement in the corpus according to the association relationship.
[0160] In one embodiment, before determining the positive Q&A example and negative Q&A example corresponding to the question statement in the corpus, the obtaining module 100 is specifically configured to:
[0161] Input the question statement into a preset Q&A model to obtain a query result, and obtain a score corresponding to the query result;
[0162] If the score is less than a preset threshold, then execute the step of determining the positive Q&A example and negative Q&A example corresponding to the question statement in the corpus.
[0163] In one embodiment, in terms of using the answer corpus of the positive Q&A example as the query result corresponding to the question statement, the obtaining module 100 is specifically configured to:
[0164] Use the answer corpus of the positive Q&A example and the question statement as training samples;
[0165] Retrain the preset Q&A model according to the training samples;
[0166] Save the retrained preset Q&A model.
[0167] The present invention also provides a query device for Q&A data. The query device for Q&A data includes a memory, a processor, and a Q&A data query program stored in the memory and executable on the processor. When the Q&A data query program is executed by the processor, each step of the Q&A data query method described in the above embodiments is implemented.
[0168] The present invention also provides a computer-readable storage medium storing a Q&A data query program. When the Q&A data query program is executed by a processor, each step of the Q&A data query method described in the above embodiments is implemented.
[0169] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0170] It should be noted that, in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, system, article or device including a series of elements includes not only those elements but also other elements not expressly listed, or further includes elements inherent to such process, system, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, system, article or device including the element.
[0171] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example system can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present invention, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a computer-readable storage medium as described above (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which may be a mobile phone, computer, parking management device, air conditioner, or network device, etc.) to execute the system described in each embodiment of the present invention.
[0172] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. A query method for question-and-answer data, characterized in that, The query method for the question-and-answer data includes: Obtain the question statements corresponding to the incorrect query results output by a preset question-and-answer model within a preset time interval; wherein, the preset question-and-answer model is used to output corresponding query results according to the user's question statements; Input the question statements into a preset first corpus model to obtain positive question-and-answer examples; wherein, the first corpus model is used to output positive question-and-answer examples according to the question statements, and the first corpus model is trained by a standard question-and-answer pair training set and an existing extended corpus; Input the question statements into a preset second corpus model to obtain negative question-and-answer examples; wherein, the second corpus model is used to output negative question-and-answer examples according to the question statements, and the second corpus model is trained by a standard question-and-answer pair training set and an existing extended corpus, and different questions in the existing corpus are shuffled and recombined as the training set for training; After obtaining the positive question-and-answer examples and negative question-and-answer examples corresponding to the question statements, store the question statements, the positive question-and-answer examples, and the negative question-and-answer examples in a corpus in an associated manner; Obtain the question statement input by the user, and determine the positive question-and-answer examples and negative question-and-answer examples corresponding to the question statement in the corpus according to the association relationship; If the first similarity between the question statement and the query corpus in the positive question-and-answer example is greater than a preset first threshold, and the second similarity between the question statement and the query corpus in the negative question-and-answer example is less than a preset second threshold, then use the answer corpus of the positive question-and-answer example as the query result corresponding to the question statement; wherein, the first similarity includes the cosine distance between the semantic vector of the question statement and the semantic vector of the query corpus in the positive question-and-answer example, the second similarity includes the cosine distance between the semantic vector of the question statement and the semantic vector of the query corpus in the negative question-and-answer example, and the first threshold and the second threshold are preset similarity thresholds.
2. The query method for Q&A data according to claim 1, wherein After the step of obtaining the question statements corresponding to the incorrect query results output by the preset question-and-answer model within the preset time interval, it further includes: Determine the relevance between the question statement and a preset standard question-and-answer pair; If the relevance is greater than a preset third threshold, then execute the step of inputting the question statement into the preset first corpus model to obtain the positive question-and-answer examples.
3. The query method for Q&A data according to claim 2, characterized in that, The step of determining the relevance between the question statement and the preset standard question-and-answer pair includes: Determine the first similarity value between the semantic vector corresponding to the question statement and the semantic vector of the query corpus in the standard question-and-answer pair; Generate the answer corpus corresponding to the question statement according to a preset semantic model; Determine the second similarity value between the semantic vector corresponding to the answer corpus of the question statement and the semantic vector of the answer corpus in the standard question-and-answer pair; Determine the relevance according to a preset weight value, the first similarity value, and the second similarity value.
4. The query method for Q&A data according to claim 1, wherein The step of storing the question statements, the positive question-and-answer examples, and the negative question-and-answer examples in the corpus in an associated manner includes: Determine the cosine value between the semantic vector of the positive question-and-answer example and a random vector, and determine the first storage location of the positive question-and-answer example according to the cosine value; Determine the second storage location of the negative Q&A example according to the first storage location; Determine the association relationship between the positive Q&A example and the negative Q&A example according to the first storage location and the second storage location; Associate and store the positive Q&A example, the negative Q&A example and the question statement in the corpus according to the association relationship.
5. The query method for Q&A data according to claim 1, characterized in that, Before the step of determining the positive Q&A example and the negative Q&A example corresponding to the question statement in the corpus, it further includes: Input the question statement into a preset Q&A model to obtain a query result, and obtain the score corresponding to the query result; If the score is less than a preset threshold, execute the step of determining the positive Q&A example and the negative Q&A example corresponding to the question statement in the corpus.
6. The query method for Q&A data according to claim 5, wherein After the step of using the answer corpus of the positive Q&A example as the query result corresponding to the question statement, it further includes: Use the answer corpus of the positive Q&A example and the question statement as training samples; Retrain the preset Q&A model according to the training samples; Save the retrained preset Q&A model.
7. A query device for question-and-answer data, characterized in that, The query device for the Q&A data includes: An acquisition module, configured to acquire the question statements corresponding to the incorrect query results output by the preset Q&A model within a preset time interval; wherein, the preset Q&A model is used to output corresponding query results according to the user's question statements; input the question statements into a preset first corpus model to obtain positive Q&A examples; wherein, the first corpus model is used to output positive Q&A examples according to question statements, and the first corpus model is trained by a standard Q&A pair training set and an existing extended corpus; input the question statements into a preset second corpus model to obtain negative Q&A examples; wherein, the second corpus model is used to output negative Q&A examples according to question statements, and the second corpus model is trained by a standard Q&A pair training set and an existing extended corpus, and different questions in the existing corpus are shuffled and recombined as the training set for training; after obtaining the positive Q&A examples and the negative Q&A examples corresponding to the question statements, associate and store the question statements, the positive Q&A examples and the negative Q&A examples in the corpus; acquire the question statements input by the user, and determine the positive Q&A examples and the negative Q&A examples corresponding to the question statements in the corpus according to the association relationship; A query module, configured to, if the first similarity between the question statement and the query corpus in the positive Q&A example is greater than a preset first threshold, and the second similarity between the question statement and the query corpus in the negative Q&A example is less than a preset second threshold, use the answer corpus of the positive Q&A example as the query result corresponding to the question statement; wherein, the first similarity includes the cosine distance between the semantic vector of the question statement and the semantic vector of the query corpus in the positive Q&A example, the second similarity includes the cosine distance between the semantic vector of the question statement and the semantic vector of the query corpus in the negative Q&A example, and the first threshold and the second threshold are preset similarity thresholds.
8. A query device for question-and-answer data, characterized in that, The query device for the Q&A data includes a memory, a processor, and a query program for the Q&A data stored in the memory and executable on the processor. When the query program for the Q&A data is executed by the processor, it implements each step of the query method for the Q&A data according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a query program for the Q&A data. When the query program for the Q&A data is executed by a processor, it implements each step of the query method for the Q&A data according to any one of claims 1-6.
Citation Information
Patent Citations
Smart searching method and device and computer readable memory medium
CN108763529A
Natural language processing using an ontology-based concept embedding model
US20210056168A1