A method for constructing a Q&A library, a terminal device, and a storage medium

By using locally sensitive hashing algorithms and semantic similarity calculation algorithms, the Q&A library is built, and problems such as inefficient construction of Q&A databases in the existing technology are solved, and a fast, efficient and widely covered Q&A library construction is achieved.

CN115982346BActive Publication Date: 2025-05-30XIAMEN KUAISHANGTONG TECH CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111196484.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-14
Publication Date
2025-05-30
Estimated Expiration
2041-10-14

AI Technical Summary

Technical Problem

The existing Q&A database construction methods have problems such as inefficiency, low cost performance, artificial deviation, difficulty in ensuring corpus coverage and richness, and difficulty in achieving personalized needs by unsupervised algorithms.

Method used

The local sensitive hash algorithm and semantic similarity calculation algorithm are used to obtain sentences with similar forms and semantic similarities in each sentence in the corpus, form similar sentence pairs, and recursively find sentences that mark the same category, filter standard sentences and representative sentences, and store them into the Q&A library.

Benefits of technology

It has realized the rapid construction of a question-and-answer library with wide coverage and rich semantic information, which has improved the construction speed and efficiency, reduced artificial deviations, and enhanced the coverage and richness of the corpus.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115982346B_ABST
    Figure CN115982346B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for constructing a Q&A library, a terminal device, and a storage medium. The method includes: S1: Collecting the corpus containing Q&A data and storing it in the corpus; S2: Based on the locality-sensitive hashing algorithm and the semantic similarity calculation algorithm, obtaining the sentences that are formally and semantically similar to each sentence in the corpus and forming similar sentence pairs; S3: For all similar sentence pairs, recursively searching for all sentences that are mutually similar sentence pairs and marking them as the same category. For all sentences in each category, screening a standard sentence and multiple representative sentences corresponding to the standard sentence, and classifying and storing the standard sentences and representative sentences corresponding to each category in the Q&A library. The operation of the present invention is simple, with a low time complexity, high calculation efficiency, and finer granularity, and can quickly construct a Q&A library with a wide coverage and rich semantic information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of question and answer, and particularly to a method for constructing a question and answer library, a terminal device, and a storage medium. Background Art

[0002] Mining question and answer data in big data has become a very important topic in the field of dialogue. The common methods for constructing the current question and answer database are as follows: (1) Artificial intelligence trainers traverse and search in a large amount of databases to select the question and answer pairs they think are common; (2) Using an unsupervised clustering method, using a pre-trained model to embed a large amount of data into a semantic space, and identifying the sentence clusters adjacent in the semantic space as the same class. These common construction methods all have deficiencies: (1) Manual operation cannot effectively traverse all data, so it takes a lot of time, manpower, and money. In the era of rapidly changing big data, it has the disadvantages of low efficiency and low cost performance, which is not conducive to the efficient cold start and continuous change of the business; (2) When manually collecting common question and answer pairs, there are often human biases, and it is difficult to guarantee the coverage rate and richness of the corpus; (3) Using an unsupervised clustering algorithm can save a lot of manpower to a certain extent, but it relies too much on the accuracy of the semantic space. When there are deviations in the semantic space, the sentences that the algorithm considers to be in the same class are often not relevant. In fact, there will always be some sentences that are not relevant to the category in the sentence classes obtained by this method; (4) Clustering algorithms often need to determine hyperparameters such as the number of categories, and these parameters vary in different data sets and are difficult to adjust; (5) The actual business scenarios are often complex and diverse. For example, Company A may think that "How much is the mobile phone of Company B" and "How much is the mobile phone of Company C" belong to the same common questions, while obviously Company C may not think so. And for such personalized needs, unsupervised algorithms are often difficult to achieve. Summary of the Invention

[0003] In order to solve the above problems, the present invention proposes a method for constructing a question and answer library, a terminal device, and a storage medium.

[0004] The specific solutions are as follows:

[0005] A method for constructing a question and answer library includes the following steps:

[0006] S1: Collect the corpus containing question and answer data and store it in the corpus;

[0007] S2: Based on the locality-sensitive hashing algorithm and the semantic similarity calculation algorithm, obtain the sentences with similar forms and similar semantics for each sentence in the corpus and form similar sentence pairs;

[0008] S3: For all similar sentence pairs, recursively search for all sentences that are mutually similar sentence pairs and mark them as the same category. For all sentences in each category, select a standard sentence and multiple representative sentences corresponding to the standard sentence, and classify and store the standard sentences and representative sentences corresponding to each category into the Q&A database.

[0009] Further, step S1 also includes cleaning the data of the corpus and then storing it in the corpus.

[0010] Further, the specific process in step S2 includes:

[0011] S21: Calculate the hash value of each sentence in the corpus through the locality-sensitive hashing algorithm and store it in the database;

[0012] S22: Traverse all sentences in the corpus. For each sentence, query from the database the sentences whose similarity based on the hash value with this sentence is greater than the first similarity threshold as the formally similar sentences of this sentence, and based on the semantic similarity calculation algorithm, query from all the formally similar sentences of this sentence the sentences whose semantic similarity with this sentence is greater than the second similarity threshold as the semantic similar sentences of this sentence, and form a similar sentence pair by corresponding this sentence with each of its semantic similar sentences.

[0013] Further, the method for screening the standard sentence in step S3 is: Set one sentence in the similar sentence pair as the recall sentence of the other sentence. According to the similar sentence pairs obtained in step S2, extract the number of recall sentences corresponding to each sentence in all sentences of a category, and use the sentence with the largest number of recall sentences as the standard sentence of this category.

[0014] Further, the method for screening multiple representative sentences corresponding to the standard sentence in step S3 is: Among all the recall sentences of the standard sentence, calculate the number of similar sentence pairs corresponding to each recall sentence, and according to the preset quantity threshold, use the recall sentences with the number of similar sentence pairs greater than the quantity threshold as the representative sentences of the standard sentence.

[0015] Further, it also includes S4: Configure the standard answer for each sentence in the Q&A database.

[0016] A Q&A database construction terminal device includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the method in the above embodiments of the present invention.

[0017] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps of the method in the above embodiments of the present invention.

[0018] The present invention adopts the above technical solution, which is simple to operate, has a low time complexity, high computing efficiency, and finer granularity, and can quickly construct a question-answering database with a wide coverage and rich semantic information. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 The flowchart of the first embodiment of the present invention is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] To further illustrate the embodiments, the present invention provides drawings. These drawings are part of the disclosure of the present invention, mainly used to illustrate the embodiments, and can be used in conjunction with the relevant descriptions in the specification to explain the operation principle of the embodiments. With reference to these contents, those of ordinary skill in the art should be able to understand other possible implementation manners and the advantages of the present invention.

[0021] The present invention will be further described below in conjunction with the drawings and specific embodiments.

[0022] Embodiment 1:

[0023] The embodiment of the present invention provides a method for constructing a question-answering database, as Figure 1 shown, the method includes the following steps:

[0024] S1: Collect the corpus containing question-and-answer data and store it in the corpus.

[0025] Since the corpus often includes some invalid characters, such as stop words, irrelevant spaces, emoticons, etc., and these invalid characters may affect the subsequent similarity calculation, therefore, in this embodiment, the corpus is also subjected to data cleaning before being stored in the corpus. For example, in this embodiment, the data “I'm three months pregnant. What's the suitable abortion method??? ” is cleaned to “I'm three months pregnant. What's the suitable abortion method? ”.

[0026] S2: Based on the Locality-Sensitive Hashing (LSH) algorithm and the semantic similarity calculation algorithm, obtain the sentences with similar forms and similar semantics in the corpus and form similar sentence pairs.

[0027] Locality-Sensitive Hashing is an approximate nearest neighbor fast search technology for massive high-dimensional data, which can quickly find one or more sentences in a massive high-dimensional data set that are most similar (closest in distance) to a certain sentence. The sublinear time complexity adopted by the Locality-Sensitive Hashing algorithm can avoid the similarity comparison between every two sentences, greatly accelerating the calculation speed. Since the Locality-Sensitive Hashing algorithm only calculates the similarity between sentences based on literal similarity and does not combine semantic similarity for query, the semantic similarity calculation still needs to be performed through the semantic similarity calculation algorithm.

[0028] The specific process in step S2 includes:

[0029] S21: Calculate the hash value of each sentence in the corpus through the locality-sensitive hashing algorithm and store it in the database.

[0030] S22: Traverse all sentences in the corpus. For each sentence, query from the database the sentences whose similarity based on the hash value with this sentence is greater than the first similarity threshold as the formally similar sentences of this sentence, and based on the semantic similarity calculation algorithm, query from all the formally similar sentences of this sentence the sentences whose semantic similarity with this sentence is greater than the second similarity threshold as the semantic similar sentences of this sentence, and form a similar sentence pair by corresponding this sentence with each of its semantic similar sentences.

[0031] Those skilled in the art can set the magnitudes of the first similarity threshold and the second similarity threshold according to requirements by themselves, and no limitation is made here. The semantic similarity calculation algorithm can adopt a common semantic similarity calculation algorithm.

[0032] For example, if the sentences in the corpus are (A, B, C, D, E, F, G, H, I), the sentences whose similarity is greater than the first similarity threshold queried from the database are: ① AC, AD, AF; ② CD, CF; ③ DC, DF; ④ EG; ⑤ FA, FD, FC; ⑥ GE. The result after secondary screening based on semantic similarity is: ① AC, AD, AF; ② CD, DF; ③ DC; ④ EG; ⑤ FD; ⑥ GE.

[0033] S3: For all similar sentence pairs, recursively find all the sentences that are mutually similar sentence pairs and mark them as the same category. For all the sentences in each category, screen a standard sentence and multiple representative sentences corresponding to the standard sentence, and classify and store the standard sentences and representative sentences corresponding to each category into the Q&A database.

[0034] In this embodiment, through recursive search, it is obtained that A, D, F, and C are in one category, and E and G are in one category.

[0035] Since under a large amount of data, one category may include thousands of sentences, in this embodiment, a method of screening out a standard sentence and multiple representative sentences from them is adopted to represent a category of data. Through this processing method, the storage space of the database can be greatly saved and the query speed can be accelerated.

[0036] The screening method for standard sentences in this embodiment is as follows: Set one sentence in a pair of similar sentences as the recalled sentence of the other sentence. According to the pair of similar sentences obtained in step S2, extract the number of recalled sentences corresponding to each sentence among all sentences of a category, and use the sentence with the largest number of recalled sentences as the standard sentence of this category. The screening method for multiple representative sentences corresponding to the standard sentence is as follows: Among all the recalled sentences of the standard sentence, calculate the number of pairs of similar sentences corresponding to each recalled sentence. According to a preset quantity threshold, use the recalled sentences with the number of pairs of similar sentences greater than the quantity threshold as the representative sentences of the standard sentence. Those skilled in the art can set the quantity threshold according to actual needs and are not limited here.

[0037] In the actual use process, according to business requirements, the standard sentences and representative sentences of multiple classifications belonging to one business requirement can also be stored as the same large category.

[0038] Further, there may be only questions but no answers in the standard sentences and representative sentences stored in the Q&A library. Therefore, step S4 needs to be further performed: Configure standard answers for each sentence in the Q&A library. The specific configuration method can adopt an existing QA matching model or can be configured manually, and is not limited here.

[0039] The embodiment of the present invention uses a supervised similarity algorithm to extract common questions, ensuring that the sentences within each category are similar to each other, and the number of sentences within the category represents the common degree of the problem, truly realizing the extraction of common questions. In this embodiment, not only standard questions can be extracted, but also similar questions to the standard questions can be extracted. In real life, there are usually many ways to express the same meaning. Obviously, it is far from enough for a category of questions to be represented by only one standard question. The supplement of representative questions that are similar but different from the standard question in the database can greatly improve the coverage rate of the answers of the question-and-answer robot. The two-stage similarity calculation method adopted in this embodiment avoids the complexity of O(n 2 ) and greatly improves the automatic construction speed of the conventional Q&A library.

[0040] Embodiment 2:

[0041] The present invention also provides a Q&A library construction terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps in the above method embodiment of Embodiment 1 of the present invention are implemented.

[0042] Further, as an executable solution, the Q&A library building terminal device may be a computing device such as a desktop computer, a notebook, a palm computer, or a cloud server. The Q&A library building terminal device may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that the composition structure of the above Q&A library building terminal device is only an example of the Q&A library building terminal device, and does not constitute a limitation on the Q&A library building terminal device. It may include more or fewer components than the above, or combine some components, or different components. For example, the Q&A library building terminal device may further include input / output devices, network access devices, buses, etc. The embodiments of the present invention do not make limitations in this regard.

[0043] Further, as an executable solution, the so-called processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the Q&A library building terminal device, and connects various parts of the entire Q&A library building terminal device through various interfaces and lines.

[0044] The memory may be used to store the computer programs and / or modules. The processor realizes various functions of the Q&A library building terminal device by running or executing the computer programs and / or modules stored in the memory, and by calling the data stored in the memory. The memory may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the mobile phone, etc. In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0045] The present invention also provides a computer-readable storage medium storing a computer program, which when executed by a processor, implements the steps of the above-mentioned method of the embodiments of the present invention.

[0046] If the module / unit integrated in the Q&A library construction terminal device is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned method embodiments of the present invention, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), and software distribution medium, etc.

[0047] Although the present invention has been specifically shown and described in conjunction with the preferred embodiments, those skilled in the art should understand that various changes can be made to the present invention in terms of form and details without departing from the spirit and scope of the present invention defined by the appended claims, and all of them are within the protection scope of the present invention.

Claims

1. A method for constructing a Q&A library, characterized in that, it includes the following steps: S1: Collect the corpus containing Q&A data and store it in the corpus; S2: Based on the locality-sensitive hashing algorithm and the semantic similarity calculation algorithm, obtain the sentences with similar forms and semantics in each sentence of the corpus and form similar sentence pairs; S3: For all similar sentence pairs, recursively search for all sentences that are mutually similar sentence pairs and mark them as the same category. For all sentences in each category, select a standard sentence and multiple representative sentences corresponding to the standard sentence, and classify and store the standard sentences and representative sentences corresponding to each category in the Q&A library.

2. The method for constructing a Q&A library according to claim 1, characterized in that: In step S1, it also includes cleaning the data of the corpus and then storing it in the corpus.

3. The method for constructing a Q&A library according to claim 1, characterized in that: The specific process in step S2 includes: S21: Calculate the hash value of each sentence in the corpus through the locality-sensitive hashing algorithm and store it in the database; S22: Traverse all sentences in the corpus. For each sentence, query the sentences in the database whose similarity based on the hash value with this sentence is greater than the first similarity threshold as the sentences with similar forms to this sentence, and based on the semantic similarity calculation algorithm, query the sentences whose semantic similarity with this sentence is greater than the second similarity threshold from all the sentences with similar forms to this sentence as the sentences with similar semantics to this sentence, and form a similar sentence pair by corresponding this sentence with each of its sentences with similar semantics.

4. The method for constructing a Q&A library according to claim 1, characterized in that: The screening method for the standard sentence in step S3 is: Set one sentence in the similar sentence pair as the recall sentence of the other sentence. According to the similar sentence pairs obtained in step S2, extract the number of recall sentences corresponding to each sentence in all sentences of a category, and take the sentence with the largest number of recall sentences as the standard sentence of this category.

5. The method for constructing a Q&A library according to claim 4, characterized in that: The screening method for multiple representative sentences corresponding to the standard sentence in step S3 is: In all the recall sentences of the standard sentence, calculate the number of similar sentence pairs corresponding to each recall sentence, and according to the preset number threshold, take the recall sentences with the number of similar sentence pairs greater than the number threshold as the representative sentences of the standard sentence.

6. The method for constructing a Q&A library according to claim 1, characterized in that: It also includes S4: Configure the standard answer for each sentence in the Q&A library.

7. A Q&A library construction terminal device, characterized in that: It includes a processor, a memory, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.

8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Text content rapid deduplication method and device, computer equipment and storage medium

    CN110309446A

  • FAQ question similarity calculation method and system

    CN111581354A