Question-answer pair mining method and device

By searching participle in the target question corpus and obtaining candidate question-and-answer pairs in combination with search engines, and calculating semantic similarity for filtering, automated question-and-answer pair mining is achieved, solving the problem of low construction efficiency of question-and-answer pairs in the existing technology, and improving the mining efficiency of question-and-answer pairs.

CN113569018BActive Publication Date: 2025-05-27TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202110170206.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-02-05
Publication Date
2025-05-27
Estimated Expiration
2041-02-05

AI Technical Summary

Technical Problem

The efficiency of building Q&A pairs in the prior art is inefficient, and it is necessary to manually collect or build Q&A pairs to train Q&A models.

Method used

By searching based on the word segmentation in the target question corpus, the first sample question corpus associated with each word segmentation in the target question corpus is obtained, and searching in the search engine to obtain the first candidate question and answer pair. Calculate semantic similarity and filter candidate Q&A to get target Q&A.

Benefits of technology

Automatic Q&A pair mining is realized, improving Q&A pair mining efficiency is improved, and reducing dependence on manual construction and collection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113569018B_ABST
    Figure CN113569018B_ABST
Patent Text Reader

Abstract

This application relates to the field of artificial intelligence technology, and specifically provides a method and device for mining question-and-answer pairs. The method includes: retrieving according to the word segmentation in the target question corpus to obtain a first sample question corpus associated with each word segmentation in the target question corpus; retrieving in a search engine according to the first sample question corpus to obtain a first candidate question-and-answer pair corresponding to the first sample question corpus, where the first candidate question-and-answer pair includes a second question corpus and a second answer corpus; calculating a first semantic similarity between any two of the first sample question corpus and the second question corpus; filtering the second candidate question-and-answer pair according to the first semantic similarity to obtain a target question-and-answer pair; through this solution, the mining efficiency of question-and-answer pairs can be greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and more particularly, to a method and apparatus for mining question-and-answer pairs. Background Art

[0002] With the maturity of artificial intelligence technology, the application of automatic question-and-answer systems is becoming more and more widespread. Based on the question-and-answer models set therein, automatic question-and-answer systems automatically understand the input question corpus according to the user's input question corpus and input the corresponding answer corpus. Before the question-and-answer model is put on the line, it is necessary to train the question-and-answer model with question-and-answer pairs to ensure the accuracy of the answer corpus output by the question-and-answer model.

[0003] In the prior art, the question-and-answer pairs used to train the question-and-answer model need to be collected or constructed manually, and the efficiency of collecting question-and-answer pairs is low. Summary of the Invention

[0004] Embodiments of the present application provide a method and apparatus for mining question-and-answer pairs to solve the problem of low efficiency in constructing question-and-answer pairs in the related art.

[0005] Other features and advantages of the present application will become apparent from the following detailed description, or will be learned in part from the practice of the present application.

[0006] According to one aspect of the embodiments of the present application, a method for mining question-and-answer pairs is provided. The method includes: retrieving according to the word segmentation in the target question corpus to obtain a first sample question corpus associated with each word segmentation in the target question corpus; retrieving in a search engine according to the first sample question corpus to obtain a first candidate question-and-answer pair corresponding to the first sample question corpus, the first candidate question-and-answer pair including a second question corpus and a second answer corpus; calculating a first semantic similarity between any two of the first sample question corpus and the second question corpus; filtering the second candidate question-and-answer pairs according to the first semantic similarity to obtain target question-and-answer pairs, the second candidate question-and-answer pairs including the question-and-answer pairs formed by the first sample question corpus and the second answer corpus corresponding to the first candidate question-and-answer pairs and the first candidate question-and-answer pairs; the target question-and-answer pairs are used to train a question-and-answer model, and the question-and-answer model is used to output an answer corpus according to the input question corpus.

[0007] According to one aspect of the embodiments of the present application, a question-answer pair mining device is provided. The device includes: a first sample question corpus acquisition module, configured to retrieve according to the word segmentation in the target question corpus to obtain a first sample question corpus associated with each word segmentation in the target question corpus; a first candidate question-answer pair acquisition module, configured to retrieve in a search engine according to the first sample question corpus to obtain a first candidate question-answer pair corresponding to the first sample question corpus, where the first candidate question-answer pair includes a second question corpus and a second answer corpus; a first semantic similarity calculation module, configured to calculate a first semantic similarity between any two of the first sample question corpus and the second question corpus; a filtering module, configured to filter a second candidate question-answer pair according to the first semantic similarity to obtain a target question-answer pair, where the second candidate question-answer pair includes a question-answer pair formed by the first sample question corpus and the second answer corpus in the corresponding first candidate question-answer pair and the first candidate question-answer pair; the target question-answer pair is used to train a question-answer model, and the question-answer model is configured to output an answer corpus according to the input question corpus.

[0008] In some embodiments of the present application, based on the foregoing solution, the first sample question corpus acquisition module includes: a word segmentation unit, configured to perform word segmentation on the target question corpus to obtain a plurality of word segmentations in the target question corpus; an inverted index acquisition unit, configured to acquire an inverted index associated with each word segmentation in the target question corpus, where the inverted index indicates an index relationship between the word segmentation and the sample question corpus; a first sample question corpus acquisition unit, configured to acquire a first sample question corpus associated with each of the word segmentations according to the inverted index.

[0009] In some embodiments of the present application, based on the foregoing solution, the question-answer pair mining device further includes: a second semantic similarity calculation module, configured to calculate a second semantic similarity between any two first sample question corpora; a second filtering module, configured to filter the first sample question corpus according to the second semantic similarity.

[0010] In some embodiments of the present application, based on the foregoing solution, the second semantic similarity calculation module includes: a determination unit configured to use one of the two first sample query corpora for which the second semantic similarity is to be calculated as a standard query corpus, and the other corpus as a comparison query corpus; a second word segmentation unit configured to perform word segmentation on the comparison query corpus to obtain a plurality of word segments in the comparison query corpus; a relevance score calculation unit configured to calculate a relevance score between each word segment in the comparison query corpus and the standard query corpus; a relevance weight calculation unit configured to calculate a relevance weight corresponding to each word segment in the comparison query corpus; and a first weighting unit configured to weight the relevance scores between all the word segments in the comparison query corpus and the standard query corpus according to the relevance weights corresponding to the word segments in the comparison query corpus, to obtain the second semantic similarity between the comparison query corpus and the standard query corpus.

[0011] In some embodiments of the present application, based on the foregoing solution, the question-answer pair mining device further includes: a third semantic similarity calculation module configured to calculate a third semantic similarity between each first sample query corpus and the target query corpus; and a third filtering module configured to filter the first sample query corpus according to the third semantic similarity.

[0012] In some embodiments of the present application, based on the foregoing solution, the third semantic similarity calculation module includes: a first input unit configured to input, for each first sample query corpus, the first sample query corpus and the target query corpus into a semantic matching model; and a second output unit configured to output, by the semantic matching model, the third semantic similarity between the first sample query corpus and the target query corpus.

[0013] In some embodiments of the present application, based on the foregoing solution, the question-answer pair mining device further includes: a training data acquisition module configured to acquire training data, where the training data includes a plurality of sample corpus pairs and labels of the sample corpus pairs, and the labels are used to indicate whether the semantics of the two sample corpora in the sample corpus pair are similar; and a training module configured to train the semantic matching model according to the sample corpus pairs and the labels of the sample corpus pairs until the semantic matching model converges.

[0014] In some embodiments of the present application, based on the foregoing solution, the first semantic similarity calculation module includes: a question corpus pair set acquisition unit, configured to acquire a question corpus pair set, where the question corpus pair set includes a plurality of question corpus pairs, and the question corpus pairs are obtained by pairwise combining the corpora in the first sample question corpus and the second question corpus; a second input unit, configured to input the question corpus pairs into the semantic matching model; and a second output unit, configured to output, by the semantic matching model, a first semantic similarity between the two corpora in the question corpus pairs.

[0015] In some embodiments of the present application, based on the foregoing solution, the filtering module 940 includes: a second acquisition unit, configured to acquire a first semantic similarity and a corresponding third semantic similarity corresponding to the question corpus in each of the second candidate question-answer pairs; a second weighting unit, configured to weight the first semantic similarity and the corresponding third semantic similarity corresponding to the question corpus in the second candidate question-answer pairs to obtain a target score of the second candidate question-answer pair; and a fourth filtering unit, configured to filter the second candidate question-answer pairs whose corresponding target scores do not meet a preset score range to obtain the target question-answer pairs.

[0016] According to one aspect of the embodiments of the present application, there is provided an electronic device, including: a processor; and a memory storing computer-readable instructions thereon, where when the computer-readable instructions are executed by the processor, the question-answer pair mining method as described above is implemented.

[0017] According to one aspect of the embodiments of the present application, there is provided a computer-readable storage medium storing computer-readable instructions thereon, where when the computer-readable instructions are executed by a processor, the question-answer pair mining method as described above is implemented.

[0018] In the solution of the present application, first, retrieval is performed according to the word segmentation in the target question corpus to obtain a first sample question corpus associated with the word segmentation in the target question corpus, realizing the mining of the question corpus; then, the first sample question corpus is retrieved in a search engine, and a first candidate question-answer pair associated with the first sample question corpus is retrieved from the data facing numerous Internet users, realizing the mining of the question-answer pair based on the question corpus; and the first sample question corpus and the first answer corpus in the corresponding first candidate question-answer pair are combined to form a new question-answer pair, and the question-answer pair is further expanded according to the first candidate question-answer pair. On this basis, according to the calculated first semantic similarity, the first candidate question-answer pair and the expanded question-answer pair are filtered to obtain the target question-answer pair. Through the above process, the automatic mining of the question-answer pair is realized based on a limited sample question corpus, and compared with the situation where all question-answer corpora are manually constructed and collected, the mining efficiency of the question-answer pair is greatly improved.

[0019] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and do not limit this application. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The accompanying drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with this application, and are used together with the specification to explain the principles of this application. Obviously, the accompanying drawings in the following description are only some embodiments of this application, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts. In the drawings:

[0021] Figure 1 FIG. shows a schematic diagram of an exemplary system architecture to which the technical solution of the embodiment of this application can be applied.

[0022] Figure 2 is a flowchart of a question-and-answer pair mining method shown according to an embodiment of this application.

[0023] Figure 3 is a flowchart of the steps before step 210 shown according to an embodiment of this application.

[0024] Figure 4 is a flowchart of the steps before step 220 shown according to an embodiment of this application.

[0025] Figure 5 is a flowchart of the steps before step 220 shown according to another embodiment of this application.

[0026] Figure 6 is a schematic structural diagram of a semantic matching model shown according to a specific embodiment.

[0027] Figure 7 is a flowchart of step 230 shown according to an embodiment of this application.

[0028] Figure 8 is a flowchart of step 240 shown according to an embodiment of this application.

[0029] Figure 9 is a block diagram of a question-and-answer pair mining device shown according to an embodiment of this application.

[0030] Figure 10 shows a schematic structural diagram of a computer system of an electronic device suitable for implementing the embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0031] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art.

[0032] In addition, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of this application. However, those skilled in the art will realize that the technical solutions of this application can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be used. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of this application.

[0033] The block diagrams shown in the drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0034] The flowcharts shown in the drawings are only illustrative and do not necessarily include all the content and operations / steps, nor do they necessarily have to be executed in the order described. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.

[0035] It should be noted that: "a plurality of" as mentioned herein means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.

[0036] Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a manner similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.

[0037] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0038] In the prior art, in order to provide services to users, the application of automatic question-answering systems is becoming more and more widespread. Among them, based on the question-answering model set therein, the automatic question-answering system automatically understands the input question corpus according to the user's input question corpus and inputs the corresponding answer corpus.

[0039] Before the question-answering model is put on the line, it is necessary to train the question-answering model with training data to ensure the accuracy of the answer corpus output by the question-answering model. The training data includes a number of question-answer pairs, and the question-answer pair includes a question corpus and a reply corpus.

[0040] In the prior art, the question-answer pairs used to train the question-answering model need to be collected or constructed manually, and there is a problem of low efficiency in collecting question-answer pairs. To solve this problem, the solution of this application is proposed.

[0041] Figure 1 The schematic diagram of an exemplary system architecture to which the technical solution of the embodiment of the present application can be applied is shown.

[0042] As Figure 1 shown, the system architecture may include terminal devices (such as Figure 1 one or more of the smart phone 101, tablet computer 102, and portable computer 103 shown in

[0043] Of course, it can also be a desktop computer, etc.), network 104, and server 105. The network 104 is used to provide a medium for the communication link between the terminal device and the server 105. The network 104 may include various connection types, such as wired communication links, wireless communication links, and so on. Figure 1 It should be understood that

[0044] the numbers of the terminal devices, networks, and servers in

[0045] In some embodiments of the present application, a question-and-answer model can also be set in the server 105, and then the question-and-answer model is trained with the mined question-and-answer pairs. After the training is completed, the server 105 can automatically output answer corpus according to the trained question-and-answer model based on the question corpus input by the user.

[0046] In some embodiments of the present application, the user can initiate a question request to the server 105 based on the terminal device. The question request includes the question corpus input by the user. Then, the server 105 analyzes the semantics of the question corpus through the question-and-answer model, automatically obtains the corresponding answer corpus for the question corpus, and returns the obtained answer corpus to the terminal where the user is located.

[0047] The implementation details of the technical solutions in the embodiments of the present application are elaborated in detail below:

[0048] Figure 2 The flowchart of the question-and-answer pair mining method shown in an embodiment of the present application is shown. This method can be executed by a computer device with processing capabilities, such as a server, etc., which is not specifically limited here. Refer to Figure 2 As shown, this method at least includes steps 210 to 240, which are introduced in detail as follows:

[0049] Step 210, retrieve according to the word segmentation in the target question corpus to obtain the first sample question corpus associated with each word segmentation in the target question corpus.

[0050] The target question corpus refers to the question corpus used as the basis for mining question-and-answer pairs, which can be a question corpus selected from the known corpus.

[0051] The first sample question corpus refers to the sample question corpus associated with the word segmentation in the target question corpus.

[0052] In some embodiments of the present application, all the word segmentations in the target question corpus can be retrieved separately to obtain the first sample question corpus associated with each word segmentation; it is also possible to only retrieve some of the word segmentations in the target question corpus separately, which can be specifically set according to actual needs.

[0053] Before step 210, a sample question corpus set is constructed, and then in step 210, retrieval is performed according to the word segmentation in the sample question corpus set to obtain the first question corpus.

[0054] Before step 210, the target question corpus is first segmented to obtain multiple word segmentations included in the target question corpus. On this basis, retrieval of the question corpus is performed in the sample question corpus according to each corpus in the target question corpus to obtain the first sample question corpus associated with each word segmentation in the target question corpus.

[0055] The target question corpus can be segmented by means of a segmentation tool. The segmentation tool uses the dictionary set therein for segmentation. In a specific embodiment, it is necessary to select a segmentation tool according to the language applicable to the Q&A model. If the Q&A model is used for Chinese Q&A, a segmentation tool that constructs a Chinese dictionary is selected to segment the target question corpus; if the Q&A model is used for English Q&A, a segmentation tool that constructs an English dictionary is selected for segmentation.

[0056] In some embodiments of the present application, since the segmentation tool segments based on the dictionary set therein, if the Q&A model is used for automatic Q&A in a certain technical field, it is necessary to select a segmentation tool that includes a dictionary of professional terms in the technical field to segment the target question corpus. For example, if the Q&A model is used in the medical Q&A field, a segmentation tool whose dictionary includes terms in the medical field is selected for segmentation. By selecting a segmentation tool corresponding to the technical field to which the Q&A model applies to segment the target question corpus, the accuracy of segmentation can be guaranteed, and thus the effectiveness of Q&A pair mining can be guaranteed.

[0057] In some embodiments of the present application, before step 210, as Figure 3 shown, the method further includes:

[0058] Step 310, segment the target question corpus to obtain a plurality of segments in the target question corpus.

[0059] Step 320, obtain an inverted index associated with each segment in the target question corpus, where the inverted index indicates the index relationship between the segment and the sample question corpus.

[0060] Step 330, obtain the first sample question corpus associated with each segment according to the inverted index.

[0061] In this embodiment, before step 320, an inverted index is first constructed according to the segments in each sample question corpus. For example, if a sample question corpus A is: How to report the loss of an ID card? Its segmentation result is: How / report / loss of ID card, that is, the segments in the sample question corpus A include: How, report, ID card, and loss. Then, a mapping from the segment to the corpus address is established according to the segments included in each sample question corpus and the addresses of each sample question corpus, that is, the inverted index. The result of the established inverted index can be as follows:

[0062] "How": the address of "sample question corpus A", the address of "sample question corpus B", the address of "sample question corpus C"...

[0063] "Processing": The addresses of "Sample Question Corpus A", "Sample Question Corpus C", "Sample Question Corpus D",...

[0064] "Identity Card": The addresses of "Sample Question Corpus A", "Sample Question Corpus E", "Sample Question Corpus F",...

[0065] "Reporting Loss": The addresses of "Sample Question Corpus A", "Sample Question Corpus B", "Sample Question Corpus G",...

[0066] In step 320, based on the established inverted index, the word segments in the target question corpus are retrieved in the inverted index library, and the inverted index including the word segment is retrieved. Since the inverted index includes the address of the sample question corpus associated with the word segment, thus, according to the retrieved inverted index, the first sample question corpus associated with the obtained word segment can be correspondingly obtained.

[0067] Please continue to refer to Figure 2 , step 220, retrieve in the search engine according to the first sample question corpus, and obtain the first candidate question-answer pair corresponding to the first sample question corpus, where the first candidate question-answer pair includes a second question corpus and a second answer corpus.

[0068] Since the retrieval basis of the search engine is the questions and answers of numerous users on the Internet, therefore, by retrieving in the search engine according to the first sample corpus, the first sample question corpus can be further enriched, and diverse corpora can be mined. Moreover, since the search engine includes not only question corpora but also answer corpora for the question corpora, thus, question-answer pairs can be mined from the search engine based on the first sample question corpus.

[0069] It is worth mentioning that in the solution of this application, during the process of retrieving in the search engine through each first sample question corpus, question corpora semantically similar to the first sample question corpus can be retrieved. That is to say, in each retrieval process, the question-answer pairs in the retrieval results correspond to the first sample question corpus used as the retrieval basis.

[0070] Step 230, calculate the first semantic similarity between any two of the first sample question corpus and the second question corpus.

[0071] In some embodiments of this application, semantic vectors of the first sample question corpus and the second question corpus can be respectively constructed, and then the semantic similarity between any two corpora can be calculated according to the semantic vectors of each question corpus. For the sake of distinction, the semantic similarity calculated here is called the first semantic similarity.

[0072] In some embodiments of the present application, the cosine distance between two semantic vectors can be calculated based on the semantic vectors respectively corresponding to two corpora, and this cosine distance is used as the first semantic similarity between the two corpora. In other embodiments, the first semantic similarity between the two corpora can also be determined by calculating the Euclidean distance between the semantic vectors corresponding to the two corpora.

[0073] Step 240: Filter the second candidate Q&A pairs according to the first semantic similarity to obtain target Q&A pairs. The second candidate Q&A pairs include the Q&A pairs composed of the first sample query corpus and the second answer corpus in the corresponding first candidate Q&A pairs, and the first candidate Q&A pairs. The target Q&A pairs are used to train a Q&A model, and the Q&A model is used to output an answer corpus according to the input query corpus.

[0074] As described above, the semantics of the query corpus in the multiple first candidate Q&A pairs retrieved based on a first sample query corpus is relevant to the semantics of the first sample query corpus used as the retrieval basis. Therefore, the second answer corpus in the first candidate Q&A pairs can be combined, and the second answer corpus is combined with the first sample query corpus corresponding to the first candidate Q&A pair where it is located to form a new Q&A pair, thereby enriching the Q&A pairs.

[0075] In some embodiments of the present application, in step 240, based on the calculated first semantic similarity, the second candidate Q&A pairs where the query corpus with a first semantic similarity less than a first similarity threshold is located can also be selected as the target Q&A pairs. Through this process, the second candidate Q&A pairs where the query corpus with a relatively high semantic similarity is located are filtered out, thereby filtering out the target Q&A pairs with highly similar semantics of the query corpus.

[0076] In some embodiments of the present application, a second similarity threshold can be set, and then according to multiple first semantic similarities related to the query corpus in the second candidate Q&A pairs (i.e., the first sample query corpus or the second query corpus in the first candidate Q&A pairs), the number of similarities exceeding the second similarity threshold among the multiple similarities obtained by statistics is counted, and then the second candidate Q&A pairs where the query corpus with the counted number exceeding the set number is located are determined.

[0077] In some embodiments of the present application, the second candidate Q&A pairs where the query corpus with the counted number exceeding the set number is located can be determined as the target Q&A pairs.

[0078] In this embodiment, the second candidate Q&A pairs are screened according to the calculated first semantic similarity by means of the set second similarity threshold and the set quantity. Since the target Q&A pairs selected are those second candidate Q&A pairs where the quantity of the first semantic similarities exceeding the second similarity threshold among the relevant multiple first semantic similarities is more than the set quantity, and the quantity of the first semantic similarities exceeding the second similarity threshold is more than the set quantity, it indicates that the question corpus is a common question corpus on the Internet. Therefore, if the second candidate Q&A pairs where the question corpus is located are determined as the target Q&A pairs, it can be ensured that after training the Q&A model with the target Q&A pairs, the Q&A model can perform automatic Q&A for common questions.

[0079] In some other embodiments of the present application, in order to avoid a high similarity between the question corpora in two target Q&A pairs, it is also possible to, on the basis of determining the second candidate Q&A pairs where the quantity of the question corpora statistically exceeds the set quantity, select the second candidate Q&A pairs where the question corpora with the first semantic similarity threshold lower than the third similarity threshold are located as the target Q&A pairs, where the third similarity threshold is greater than the first similarity threshold.

[0080] In the solution of the present application, first, retrieval is performed according to the word segmentation in the target question corpus to obtain the first sample question corpus associated with the word segmentation in the target question corpus, realizing the mining of the question corpus; then the first sample question corpus is retrieved in the search engine, and the first candidate Q&A pairs corresponding to the first sample question corpus are retrieved from the data facing numerous Internet users, realizing the mining of the Q&A pairs based on the question corpus; and the first sample question corpus and the first answer corpus in the corresponding first candidate Q&A pairs are combined to form new Q&A pairs, further expanding the Q&A pairs according to the first candidate Q&A pairs. On this basis, according to the calculated first semantic similarity, the first candidate Q&A pairs and the expanded Q&A pairs are filtered to obtain the target Q&A pairs. Through the above process, the automatic mining of the Q&A pairs is realized based on the limited sample question corpus, which greatly improves the mining efficiency of the Q&A pairs compared with the situation where all the Q&A corpora are manually constructed and collected.

[0081] In an application scenario, an entity object may have multiple attributes. For example, for the entity object of an ID card, its attributes include the validity period of the ID card, the handling process, the handling location, the time from applying for the ID card to obtaining it, loss reporting, and so on.

[0082] In an automatic question-answering scenario, it is required that the question-answering model can provide answer corpora for questions related to multiple attributes of the entity object. Therefore, it is necessary to use question-answer pairs covering multiple attributes of the entity object to train the question-answering model. In such an application scenario, the solution of the present application can be used to mine question-answer pairs related to multiple attributes of an entity object.

[0083] In such an application scenario, through the process of step 210, multiple first sample question corpora associated with the word segmentation in the target question corpus can be retrieved. Since the word segmentation in the target question corpus includes the word segmentation used to describe the entity object targeted by the target question corpus, and on the basis that the corpora in the sample question corpus can cover multiple attributes of an entity object, through the process of step 210, multiple questions related to the entity object can be retrieved. Then, through the retrieval in step 220, the mining and enrichment of question-answer pairs are further carried out through the first sample question corpora, and then the target question-answer pairs are determined. Since the retrieved first sample question corpora include questions related to multiple attributes of the entity object targeted by the target question corpus, it can be ensured that the obtained target question-answer pairs include question-answer pairs for multiple attributes of the entity object targeted by the target question corpus.

[0084] In some embodiments of the present application, as Figure 4 shown, before step 220, the method further includes:

[0085] Step 410, calculating the second semantic similarity between any two first sample question corpora.

[0086] Step 420, filtering the first sample question corpora according to the second semantic similarity.

[0087] As described above, the first sample corpus is retrieved according to the word segmentation in the target question corpus, that is, the acquisition of the first sample corpus only focuses on whether a sample corpus includes a word segmentation in the target question corpus. There may be sample question corpora with highly similar semantics in the obtained first sample question corpora associated with the word segmentation in the target corpus. If the semantics of two first sample question corpora are highly similar, the second candidate pairs obtained by retrieving through these two first sample question corpora respectively may be the same. Therefore, in order to avoid such a situation, the first sample question corpora can be filtered first.

[0088] Among them, the semantic similarity between the two first-sample query corpora (for the sake of distinction, the semantic similarity calculated this time is called the second semantic similarity) can be calculated based on the semantic vectors corresponding to the two first-sample corpora in the same process as the calculation of the first semantic similarity, such as cosine distance, Euclidean distance, etc., so as to determine the second semantic similarity of the corresponding pair.

[0089] In step 420, a fourth similarity threshold can be set, and based on multiple second semantic similarities related to a first-sample query corpus (for the sake of description, it is called the target sample query corpus), the first-sample query corpora with a second semantic similarity higher than the fourth similarity threshold are filtered out, and the first-sample query corpora with a second semantic similarity not higher than the fourth similarity threshold to the target query corpus are retained.

[0090] Through the process of steps 410-420 above, among multiple first-sample query corpora with high semantic similarity, only one first-sample query corpus is retained, and other first-sample query corpora are filtered out, so as to avoid the situation where the second candidate question pairs in the multiple retrieval results obtained by retrieving with first-sample query corpora with high semantic similarity are highly identical.

[0091] Please continue to refer to Figure 4 As shown, in some embodiments of the present application, step 410 further includes:

[0092] Step 411, using one of the two first-sample query corpora to be calculated for the second semantic similarity as the standard query corpus, and the other corpus as the control query corpus.

[0093] Step 412, perform word segmentation on the control query corpus to obtain multiple word segments in the control query corpus.

[0094] The method of word segmentation can refer to the process of word segmentation of the target query corpus in the above text, and will not be elaborated here.

[0095] Step 413, calculate the correlation score between each word segment in the control query corpus and the standard query corpus.

[0096] In some embodiments of the present application, the correlation score between each word segment in the control query corpus and the standard query corpus can be calculated according to the following formula:

[0097]

[0098]

[0099] where k 1 、k 2, b is a regulation factor, which can usually be set according to experience. For example, set k 1 = 2, k 2 = 1, b = 0.75. f i is the occurrence frequency of the i-th word segment q i in the standard question corpus d; qf i is the occurrence frequency of the i-th word segment q i in the control question corpus; dl is the text length of the standard question corpus d; avgdl is the average text length of all the first sample question corpora.

[0100] Under normal circumstances, the i-th word segment q i in the control question corpus only appears once. Therefore, qf i = 1. On this basis, the above formula 1 can be further simplified as:

[0101]

[0102] Step 414, calculate the relevance weight corresponding to each word segment in the control question corpus.

[0103] In some embodiments of the present application, the relevance weight can be calculated according to the following formula:

[0104]

[0105] where N is the total number of the first sample question corpora; n(q i ) is the number of the first sample question corpora containing the i-th word segment q i in the control question corpus.

[0106] Step 415, according to the relevance weights corresponding to the respective word segments in the control question corpus, weight the relevance scores between all the word segments in the control question corpus and the standard question corpus, to obtain the second semantic similarity between the control question corpus and the standard question corpus.

[0107] After determining the relevance scores between each word segment in the control question corpus and the standard question corpus and the relevance weights corresponding to each word segment in the control question corpus through the above steps 413 and 414, use the relevance weights corresponding to each word segment in the control question corpus as the weighting coefficients to weight the relevance scores between all the word segments and the standard question corpus, that is, perform weighting according to the following formula:

[0108]

[0109] Wherein, n is the number of word segments included in the control query corpus Q, and Score(Q, d) is the second semantic similarity between the standard query corpus d and the control query corpus Q.

[0110] In the solution of this embodiment, the semantic relevance between the control query corpus and the standard query corpus is reflected by means of the relevance score between each word segment in the control query corpus and the standard query corpus, and then the semantic similarity between the control query corpus and the standard query corpus is calculated, realizing the calculation of semantic similarity using an unsupervised algorithm. Compared with the supervised algorithm, its calculation amount is less.

[0111] In some embodiments of the present application, as Figure 5 shown, before step 220, the method further includes:

[0112] Step 510, calculating the third semantic similarity between each first sample query corpus and the target query corpus.

[0113] Step 520, filtering the first sample query corpus according to the third semantic similarity.

[0114] The third semantic similarity between the first sample query corpus and the target query corpus can be calculated by the above method of first constructing the semantic vector of the corpus and then calculating the distance between the semantic vectors. It can also be calculated according to the above method based on the relevance weight and the relevance score, which is not specifically limited herein.

[0115] In step 520, the first sample query corpus can be filtered according to the set filtering range. If the third semantic similarity corresponding to a first sample query corpus is within the similarity range defined by the filtering range, then the first sample query corpus is filtered out. On the contrary, if the third semantic similarity corresponding to a first sample query corpus exceeds the similarity range defined by the filtering range, then the first sample query corpus is retained.

[0116] In a specific embodiment, the filtering range can be defined based on a set filtering threshold, and the filtering range is a range less than the filtering threshold, that is, if the third semantic similarity corresponding to a first sample query corpus is less than the filtering threshold, then the first sample query corpus is filtered out.

[0117] In some embodiments of the present application, step 510 may further include: for each first sample query corpus, inputting the first sample query corpus and the target query corpus into a semantic matching model; and outputting the third semantic similarity between the first sample query corpus and the target query corpus by the semantic matching model.

[0118] In the solution of this embodiment, a semantic matching model is used to calculate the third similarity between the first sample query corpus and the target query corpus. The semantic matching model can respectively construct vector representations of the first sample query corpus and the target query corpus, and then output the third similarity between the two query corpora according to the vectors corresponding thereto. The semantic matching model can be a model constructed based on a neural network, such as a recurrent neural network, a convolutional neural network, etc.

[0119] Figure 6 is a schematic structural diagram of a semantic matching model shown according to a specific embodiment, as Figure 6 shown, the semantic matching model includes a representation layer 610, an interaction layer 620, an aggregation layer 630, and an output layer 640. Among them, the representation layer 610 is constructed based on a Bi-LSTM (Bi-directional Long-Short Term Memory), and is used to output a hidden state sequence according to the word vectors of the text corpus.

[0120] For convenience of description, the word vectors of each word segment in the first sample query corpus are sequentially represented as: a 1 、a 2 、a 3 ……a l ; the word vectors of each word segment in the target query corpus are sequentially represented as: b 1 、b 2 、b 3 ……b m .

[0121] After the action of the Bi-LSTM in the representation layer, the hidden state vectors corresponding to each word segment in the first sample query corpus are respectively output and the hidden state vectors of each word segment in the target query corpus

[0122] The interaction layer 620 is used to perform information interaction between the hidden state vectors of each word segment in the first sample query corpus and the hidden state vectors of each word segment in the target query corpus to generate vectors after interaction.

[0123] Specifically, the interaction layer 620 first multiplies the hidden state vector corresponding to the word segment in the first sample query corpus by the hidden state vector of the word segment in the target query corpus to obtain a product vector e ij :

[0124]

[0125] Then, the interaction layer 620 calculates the interaction vectors corresponding to each word segment in the two corpora based on the product vector and the softmax function:

[0126]

[0127]

[0128] On the basis of obtaining the hidden state sequences and interaction vectors corresponding to each word segment in the first sample question corpus and the target question corpus, by performing difference and product operations on the hidden state sequences and interaction vectors, and integrating various vectors, an integrated sequence corresponding to each word segment is obtained:

[0129]

[0130]

[0131] Then, the activation layer 621 in the interaction layer 620 is used to activate the integrated sequences of each word segment obtained. The activation function of the activation layer can be the Relu function. Specifically, the activated integrated sequence can be expressed as:

[0132]

[0133]

[0134] The aggregation layer 630 includes a bidirectional long short-term memory network (Bi-LSTM) layer and a pooling layer 631. Among them, the bidirectional long short-term memory network layer is used to comprehensively analyze the information of all word segments for global analysis. After being processed by the Bi-LSTM in the aggregation layer, the sequence corresponding to the word segment in the first sample question corpus is expressed as The sequence corresponding to the word segment in the target question corpus is expressed as

[0135] The pooling layer 631 in the aggregation layer 630 includes an average pooling layer and a max pooling layer. The average pooling layer is used to perform average pooling operations on the sequence corresponding to the word segment in the first sample question corpus and the sequence corresponding to the word segment in the target question corpus. The vector obtained by the average pooling operation is denoted as v a,avg and v b,avg ; The max pooling layer is used to perform max pooling operations on the sequence corresponding to the word segment in the first sample question corpus and the sequence corresponding to the word segment in the target question corpus. The vectors obtained by the max pooling operations are respectively denoted as v a,max and v b,max . Then, the vectors obtained by the average pooling operation and the max pooling operation are integrated to generate the target vector v: v = [v a,avg , v a,max , v b,avg , vb,max .

[0136] Finally, the output layer 640 outputs the third semantic similarity y = G(v) between the first sample query corpus and the target query corpus according to the target vector v.

[0137] Of course, Figure 6 This is only an exemplary example of the structure of the semantic matching model. In other embodiments, the third semantic similarity between the first sample query corpus and the target query corpus can also be output through other semantic matching models. For example, the semantic matching model can be a BiMPM (Bilateral Multi-Perspective Matching) model, etc.

[0138] To ensure the accuracy of the results output by the semantic matching model, the semantic matching model also needs to be trained. Specifically, the semantic matching model can be trained through the following process: obtaining training data, where the training data includes a number of sample corpus pairs and the labels of the sample corpus pairs, and the labels are used to indicate whether the semantics of the two sample corpora in the sample corpus pair are similar; training the semantic matching model according to the sample corpus pair and the label of the sample corpus pair until the semantic matching model converges.

[0139] In some embodiments of the present application, if the semantics of the two corpora in the sample corpus pair are similar, the label of the sample corpus pair can be marked as "1", and if the semantics of the two corpora in the sample corpus pair are not similar, the label of the sample corpus pair can be marked as "0".

[0140] In some embodiments of the present application, in order to ensure that the trained semantic matching model matches its application, the sample corpus in the sample corpus pair can be the collected query corpus.

[0141] In some embodiments of the present application, if the sample corpus in the sample corpus pair is a query corpus, the sample corpus in the training data can be used as the sample query corpus in the solution of the present application. Furthermore, in step 210, based on the word segmentation in the target query corpus, a first sample query corpus associated with the word segmentation is retrieved in the sample query corpus.

[0142] During the process of training the semantic matching model, the sample corpus pair is input into the semantic matching model, and the semantic matching model outputs the semantic similarity between the two corpora in the sample corpus pair. If the output semantic similarity does not match the label of the sample corpus pair, the parameters of the semantic matching model are adjusted, and then the semantic similarity between the two corpora in the sample corpus pair is output again through the semantic matching model with adjusted parameters until the output semantic similarity matches the label of the sample corpus pair; then continue to train the semantic matching model with the next sample corpus pair. During the training process, the loss function value of the semantic matching model is calculated. If the calculated loss function value indicates that the semantic matching model converges, the training of the semantic matching model is ended.

[0143] In an embodiment of the present application, it is possible to determine whether the semantic similarity output by the semantic matching model for the sample corpus pair matches the label of the sample corpus pair based on a set threshold. Specifically, if the semantic similarity is greater than the set threshold, it is considered that the two corpora in the sample corpus pair are similar; otherwise, it is considered that the two corpora in the sample corpus pair are not similar.

[0144] For example, if the label of a sample corpus pair indicates that the two corpora in the sample corpus pair are similar, and the semantic similarity output by the semantic matching model for the sample corpus pair is greater than the set threshold, it indicates that the semantic similarity output by the semantic matching model for the sample corpus pair matches the label of the sample corpus pair; if the semantic similarity output by the semantic matching model for the sample corpus pair is not greater than the set threshold, it indicates that the semantic similarity output by the semantic matching model for the sample corpus pair does not match the label of the sample corpus pair.

[0145] In some embodiments of the present application, it is possible to simultaneously combine Figure 4 the filtering process shown and Figure 5 the filtering process shown to filter the first sample query corpus obtained in step 210. Specifically, it is possible to first perform a primary filter on the first sample query corpus obtained in step 210 according to Figure 4 the process shown, and then perform a secondary filter on the primarily filtered first sample query corpus according to Figure 5 the process shown.

[0146] In Figure 4 the embodiment shown, the second semantic similarity between two first sample query corpora is calculated based on the relevance score and the relevance weight, and an unsupervised algorithm is used to filter the first sample query corpus, with higher filtering efficiency. First, a rough filter is performed according to Figure 4 the process, and then a fine filter is performed in combination with Figure 5 the process shown, so that it is not necessary to perform the Figure 5 matching process shown for each first sample query corpus, reducing the matching calculation amount.

[0147] In some embodiments of the present application, as Figure 7 shown, step 230 includes:

[0148] Step 710, obtaining a set of question-and-answer corpus pairs, where the set of question-and-answer corpus pairs includes a number of question-and-answer corpus pairs, and the question-and-answer corpus pairs are obtained by pairwise combining the corpora in the first sample question corpus and the second question corpus.

[0149] Step 720, inputting the question-and-answer corpus pairs into the semantic matching model.

[0150] Step 730, outputting, by the semantic matching model, a first semantic similarity between the two corpora in the question-and-answer corpus pair.

[0151] In the solution of this embodiment, the first semantic similarity of the question-and-answer corpus pair is calculated through Figure 5 the semantic matching in the shown embodiment. The structural diagram of the semantic matching model can be as Figure 6 shown. The specific calculation process of the first semantic similarity refers to Figure 5 and Figure 6 the descriptions of the corresponding embodiments, which will not be elaborated here.

[0152] In some embodiments of the present application, as Figure 8 shown, step 240 includes:

[0153] Step 810, obtaining the first semantic similarity and the corresponding third semantic similarity corresponding to the question corpus in each of the second candidate question-and-answer pairs.

[0154] Step 820, weighting the first semantic similarity and the corresponding third semantic similarity corresponding to the question corpus in the second candidate question-and-answer pair to obtain a target score of the second candidate question-and-answer pair.

[0155] Step 830, filtering out the second candidate question-and-answer pairs whose corresponding target scores do not meet the preset score range to obtain the target question-and-answer pairs.

[0156] Among them, the weighting weights corresponding to the first semantic similarity and the second semantic similarity can be set according to actual needs, and no specific limitation is made here.

[0157] In the solution of this embodiment, by combining the first semantic similarity and the third semantic similarity corresponding to the question corpus in the second candidate question-and-answer pair to comprehensively calculate the target score, and then filtering the second candidate question-and-answer pairs according to the target score, it realizes filtering the second candidate question-and-answer pairs by combining the first semantic similarity and the third semantic similarity.

[0158] The following introduces the device embodiments of the present application, which can be used to execute the methods in the above embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the above method embodiments of the present application.

[0159] Figure 9 is a block diagram of a question-and-answer pair mining device shown according to an embodiment, as Figure 9 shown, the question-and-answer pair mining device includes:

[0160] A first sample question corpus acquisition module 910, configured to retrieve according to the word segmentation in the target question corpus, and obtain a first sample question corpus associated with each word segmentation in the target question corpus;

[0161] A first candidate question-and-answer pair acquisition module 920, configured to retrieve in a search engine according to the first sample question corpus, and obtain a first candidate question-and-answer pair corresponding to the first sample question corpus, where the first candidate question-and-answer pair includes a second question corpus and a second answer corpus;

[0162] A first semantic similarity calculation module 930, configured to calculate a first semantic similarity between any two corpora in the first sample question corpus and the second question corpus;

[0163] A filtering module 940, configured to filter a second candidate question-and-answer pair according to the first semantic similarity to obtain a target question-and-answer pair, where the second candidate question-and-answer pair includes a question-and-answer pair composed of the first sample question corpus and the second answer corpus in the corresponding first candidate question-and-answer pair and the first candidate question-and-answer pair; the target question-and-answer pair is used to train a question-and-answer model, and the question-and-answer model is configured to output an answer corpus according to an input question corpus.

[0164] In some embodiments of the present application, the first sample question corpus acquisition module 910 includes: a word segmentation unit, configured to perform word segmentation on the target question corpus to obtain a plurality of word segmentations in the target question corpus; an inverted index acquisition unit, configured to obtain an inverted index associated with each word segmentation in the target question corpus, where the inverted index indicates an index relationship between the word segmentation and the sample question corpus; a first sample question corpus acquisition unit, configured to obtain a first sample question corpus associated with each of the word segmentations according to the inverted index.

[0165] In some embodiments of the present application, the question-and-answer pair mining device further includes: a second semantic similarity calculation module, configured to calculate a second semantic similarity between any two first sample question corpora; a second filtering module, configured to filter the first sample question corpus according to the second semantic similarity.

[0166] In some embodiments of the present application, the second semantic similarity calculation module includes: a determination unit configured to use one of the two first sample query corpora for which the second semantic similarity is to be calculated as a standard query corpus, and the other corpus as a comparison query corpus; a second word segmentation unit configured to perform word segmentation on the comparison query corpus to obtain a plurality of word segments in the comparison query corpus; a relevance score calculation unit configured to calculate the relevance score between each word segment in the comparison query corpus and the standard query corpus; a relevance weight calculation unit configured to calculate the relevance weight corresponding to each word segment in the comparison query corpus; and a first weighting unit configured to weight the relevance scores between all the word segments in the comparison query corpus and the standard query corpus according to the relevance weights corresponding to the word segments in the comparison query corpus, to obtain the second semantic similarity between the comparison query corpus and the standard query corpus.

[0167] In some embodiments of the present application, the Q&A pair mining device further includes: a third semantic similarity calculation module configured to calculate the third semantic similarity between each first sample query corpus and the target query corpus; and a third filtering module configured to filter the first sample query corpus according to the third semantic similarity.

[0168] In some embodiments of the present application, the third semantic similarity calculation module includes: a first input unit configured to input the first sample query corpus and the target query corpus into a semantic matching model for each first sample query corpus; and a second output unit configured to output, by the semantic matching model, the third semantic similarity between the first sample query corpus and the target query corpus.

[0169] In some embodiments of the present application, the Q&A pair mining device further includes: a training data acquisition module configured to acquire training data, where the training data includes a plurality of sample corpus pairs and labels of the sample corpus pairs, and the labels are used to indicate whether the semantics of the two sample corpora in the sample corpus pair are similar; and a training module configured to train the semantic matching model according to the sample corpus pairs and the labels of the sample corpus pairs until the semantic matching model converges.

[0170] In some embodiments of the present application, the first semantic similarity calculation module 930 includes: a query corpus pair set acquisition unit configured to acquire a query corpus pair set, where the query corpus pair set includes a plurality of query corpus pairs, and the query corpus pairs are obtained by pairwise combining the corpora in the first sample query corpus and the second query corpus; a second input unit configured to input the query corpus pairs into the semantic matching model; and a second output unit configured to output, by the semantic matching model, the first semantic similarity between the two corpora in the query corpus pair.

[0171] In some embodiments of the present application, the filtering module 940 includes: a second acquisition unit configured to acquire a first semantic similarity and a corresponding third semantic similarity corresponding to the query corpus in each of the second candidate question-answer pairs; a second weighting unit configured to weight the first semantic similarity and the corresponding third semantic similarity corresponding to the query corpus in the second candidate question-answer pairs to obtain a target score of the second candidate question-answer pairs; and a fourth filtering unit configured to filter the second candidate question-answer pairs whose corresponding target scores do not meet a preset score range to obtain the target question-answer pairs.

[0172] Figure 10 FIG. shows a schematic structural diagram of a computer system of an electronic device suitable for implementing embodiments of the present application.

[0173] It should be noted that Figure 10 The illustrated computer system 1000 of the electronic device is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present application.

[0174] As Figure 10 shown, the computer system 1000 includes a central processing unit (CPU) 1001, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage section 1008 into a random access memory (RAM) 1003, such as executing the method in the above embodiments. In the RAM 1003, various programs and data required for system operations are also stored. The CPU 1001, the ROM 1002, and the RAM 1003 are connected to each other through a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.

[0175] The following components are connected to the I / O interface 1005: an input section 1006 including a keyboard, a mouse, etc.; an output section 1007 including such as a Cathode Ray Tube (CRT), a Liquid Crystal Display (LCD), etc. and a speaker, etc.; a storage section 1008 including a hard disk, etc.; and a communication section 1009 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 1009 performs communication processing via a network such as the Internet. The drive 1010 is also connected to the I / O interface 1005 as required. A removable medium 1011, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1010 as required so that a computer program read from it is installed into the storage section 1008 as required.

[0176] Specifically, according to an embodiment of the present application, the processes described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through the communication section 1009, and / or installed from the removable medium 1011. When the computer program is executed by a central processing unit (CPU) 1001, various functions defined in the system of the present application are executed.

[0177] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0178] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. Among them, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0179] The units involved in the embodiments of the present application can be implemented in software or in hardware, and the described units can also be provided in a processor. In some cases, the names of these units do not constitute a limitation on the units themselves.

[0180] On the other hand, the present application also provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiments; or may exist separately without being assembled into the electronic device. The above computer-readable storage medium carries computer-readable instructions, and when the computer-readable storage instructions are executed by a processor, the methods in any of the above embodiments are implemented.

[0181] According to one aspect of the present application, an electronic device is further provided, which includes: a processor; a memory, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by the processor, the methods in any of the above embodiments are implemented.

[0182] According to one aspect of the embodiments of the present application, a computer program product or a computer program is provided, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods in any of the above embodiments.

[0183] It should be noted that although several modules or units of a device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more of the above-described modules or units can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0184] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (such as a personal computer, a server, a touch terminal, or a network device, etc.) to execute the methods according to the embodiments of the present application.

[0185] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the embodiments disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include the common general knowledge or conventional technical means in the technical field not disclosed in the present application.

[0186] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.

Claims

1. A method for mining question-and-answer pairs, characterized in that, the method includes: retrieving according to the word segmentation in the target question corpus to obtain a first sample question corpus associated with each word segmentation in the target question corpus; calculating a third semantic similarity between each first sample question corpus and the target question corpus; filtering the first sample question corpus according to the third semantic similarity; retrieving in a search engine according to the first sample question corpus to obtain a first candidate question-and-answer pair corresponding to the first sample question corpus, the first candidate question-and-answer pair including a second question corpus and a second answer corpus; calculating a first semantic similarity between any two of the first sample question corpus and the second question corpus; obtaining the first semantic similarity and the corresponding third semantic similarity corresponding to the question corpus in each second candidate question-and-answer pair; weighting the first semantic similarity and the corresponding third semantic similarity corresponding to the question corpus in the second candidate question-and-answer pair to obtain the target score of the second candidate question-and-answer pair; filtering the second candidate question-and-answer pairs whose corresponding target scores do not meet the preset score range to obtain target question-and-answer pairs, the second candidate question-and-answer pairs including the question-and-answer pairs composed of the first sample question corpus and the second answer corpus corresponding to the first candidate question-and-answer pair and the first candidate question-and-answer pair; the target question-and-answer pairs are used to train a question-and-answer model, and the question-and-answer model is used to output an answer corpus according to the input question corpus.

2. The method according to claim 1, characterized in that, the retrieving according to the word segmentation in the target question corpus to obtain a first sample question corpus associated with each word segmentation in the target question corpus includes: performing word segmentation on the target question corpus to obtain a plurality of word segmentations in the target question corpus; obtaining an inverted index associated with each word segmentation in the target question corpus, the inverted index indicating the index relationship between the word segmentation and the sample question corpus; obtaining a first sample question corpus associated with each of the word segmentations according to the inverted index.

3. The method according to claim 2, characterized in that, before the retrieving in a search engine according to the first sample question corpus to obtain a first candidate question-and-answer pair corresponding to the first sample question corpus, the method further includes: calculating a second semantic similarity between any two first sample question corpora; filtering the first sample question corpus according to the second semantic similarity.

4. The method according to claim 3, characterized in that, the calculating a second semantic similarity between any two first sample question corpora includes: taking one of the two first sample question corpora to be calculated for the second semantic similarity as a standard question corpus, and the other corpus as a control question corpus; performing word segmentation on the control question corpus to obtain a plurality of word segmentations in the control question corpus; calculating a correlation score between each word segmentation in the control question corpus and the standard question corpus; calculating a correlation weight corresponding to each word segmentation in the control question corpus; According to the relevance weights corresponding to each participle in the control question corpus, weight the relevance scores between all the participles in the control question corpus and the standard question corpus to obtain the second semantic similarity between the control question corpus and the standard question corpus.

5. The method according to claim 1, wherein, the calculation of the third semantic similarity between each first sample question corpus and the target question corpus includes: For each first sample question corpus, input the first sample question corpus and the target question corpus into a semantic matching model; The semantic matching model outputs the third semantic similarity between the first sample question corpus and the target question corpus.

6. The method according to claim 5, wherein, before inputting the first sample question corpus and the target question corpus into the semantic matching model, the method further includes: Obtain training data, the training data includes a number of sample corpus pairs and labels of the sample corpus pairs, and the labels are used to indicate whether the semantics of the two sample corpora in the sample corpus pair are similar; Train the semantic matching model according to the sample corpus pairs and the labels of the sample corpus pairs until the semantic matching model converges.

7. The method according to claim 5, wherein, the calculation of the first semantic similarity between any two corpora in the first sample question corpus and the second question corpus includes: Obtain a set of question corpus pairs, the set of question corpus pairs includes a number of question corpus pairs, and the question corpus pairs are obtained by combining the corpora in the first sample question corpus and the second question corpus pairwise; Input the question corpus pairs into the semantic matching model; The semantic matching model outputs the first semantic similarity between the two corpora in the question corpus pair.

8. A question-answer pair mining device, wherein, the device includes: A first sample question corpus acquisition module, configured to retrieve according to the participles in the target question corpus to obtain a first sample question corpus associated with each participle in the target question corpus; calculate the third semantic similarity between each first sample question corpus and the target question corpus; filter the first sample question corpus according to the third semantic similarity; A first candidate question-answer pair acquisition module, configured to retrieve in a search engine according to the first sample question corpus to obtain a first candidate question-answer pair corresponding to the first sample question corpus, and the first candidate question-answer pair includes a second question corpus and a second answer corpus; A first semantic similarity calculation module, configured to calculate the first semantic similarity between any two corpora in the first sample question corpus and the second question corpus; A filtering module, which obtains the first semantic similarity and the corresponding third semantic similarity of the question corpus in each second candidate question-answer pair; weights the first semantic similarity and the corresponding third semantic similarity of the question corpus in the second candidate question-answer pair to obtain the target score of the second candidate question-answer pair; filters the second candidate question-answer pairs whose corresponding target scores do not meet the preset score range to obtain target question-answer pairs, where the second candidate question-answer pairs include the question-answer pairs composed of the first sample question corpus and the second answer corpus in the corresponding first candidate question-answer pair and the first candidate question-answer pairs; the target question-answer pairs are used to train a question-answer model, and the question-answer model is used to output an answer corpus according to the input question corpus.

9. A computer device, characterized in that it includes: a memory in which a computer program is stored; a processor for loading the computer program to implement the question-answer pair mining method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that the computer-readable storage medium stores a computer program, and the computer program is adapted to be loaded and executed by a processor to implement the question-answer pair mining method according to any one of claims 1-7.

11. A computer program product, characterized in that the computer program product includes a computer program, and the computer program is adapted to be loaded and executed by a processor to implement the question-answer pair mining method according to any one of claims 1-7.

Citation Information

Patent Citations

  • A method for implementing a question answering system based on a question-answer pair

    CN109271505A

  • Method for obtaining question and answer pairs from unstructured text based on deep learning

    CN110110054A

  • Question and answer library expansion method and device, server and storage medium

    CN110413755A

  • FAQ question similarity calculation method and system

    CN111581354A