Answer tracing method and device for document question-answering system, terminal equipment and storage medium
By calculating the keyword coverage and semantic similarity between candidate documents and answers in the document question and answer system, determining the correlation degree to trace the source of the answer, the transparency and accuracy of the existing system in the answer traceability are solved, and the reliability of the system is improved.
Patent Information
- Application Number
- CN202411849275.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-16
- Publication Date
- 2025-05-16
AI Technical Summary
The existing document Q&A system lacks transparency and accuracy in answer tracing, making it difficult to trace the source of answer errors, and the small number of embedded model parameters leads to insufficient understanding of complex semantic relationships.
By obtaining keywords in candidate documents and answers, calculating keyword coverage and semantic similarity, and combining weight formulas to determine the degree of correlation between candidate documents and answers, thereby traceing the source of the answer.
It improves the accuracy of answer tracing, avoids the problem of insufficient understanding of semantic relationships caused by the small amount of embedded model parameters, and enhances the transparency and reliability of the system.
Smart Images

Figure CN120011490A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of natural language processing technology, and in particular to a method, device, terminal device and storage medium for tracing answers to a document question-and-answer system. Background Art
[0002] Most existing document question answering systems work in a black-box mode, that is, the system directly outputs the final answer without providing sufficient explanation to explain how the answer was obtained. This lack of transparency not only reduces user trust, but also brings difficulties when it is necessary to verify the correctness of the answer or understand the logic behind the answer. In addition, when the answer is wrong, it is difficult to trace the source of the error, which poses a challenge to improving the reliability and maintainability of the system.
[0003] The current mainstream technology for traceability tagging is: through each context retrieved from the database and each sentence in the answer, the embedding model calculates the score of each sentence in the context and the answer, and the embedding similarity of each sentence in the context and the answer is matched. When the similarity exceeds a threshold, it is determined that the sentence in the answer is inferred from the context.
[0004] The disadvantage of the above-mentioned prior art is that, because the embedding model has a relatively small number of parameters, it sometimes fails to fully understand complex semantic relationships. Relying only on the embedding model, it may appear that a sentence in the context and the answer is very similar, but in fact there is no connection, resulting in low accuracy in answer tracing. Summary of the invention
[0005] The embodiments of the present invention provide a method, apparatus, terminal device and storage medium for tracing the answer of a document question and answer system, which can improve the accuracy of answer tracing of the document question and answer system.
[0006] An embodiment of the present invention provides a method for tracing the source of an answer in a document question-answering system, comprising:
[0007] Obtaining a number of candidate documents and answers generated by the document question answering system for questions input by the user;
[0008] Extracting keywords from the candidate documents and the answers, and calculating coverage of the keywords in each candidate document to the keywords in the answer based on the extracted keywords;
[0009] Calculate the semantic similarity between the answer and each candidate document;
[0010] The relevance between each candidate document and the answer is determined according to the coverage and the semantic similarity, and the candidate document with the highest relevance is used as the source of the answer.
[0011] Furthermore, the obtaining of several candidate documents includes:
[0012] The relevance between the question and each document stored in the database is detected through a preset large language model, and the documents whose relevance exceeds a preset threshold are used as the candidate documents.
[0013] Furthermore, the calculation of the coverage of the keywords in each candidate document to the keywords in the answer based on the extracted keywords includes:
[0014] For each candidate document, compare the keywords in the candidate document with the keywords in the answer, and take the same keywords as common keywords;
[0015] The coverage rate is obtained by calculating the ratio of the number of common keywords to the number of all keywords in the candidate documents.
[0016] Furthermore, the calculation of the semantic similarity between the answer and each candidate document includes:
[0017] For each candidate document, use the preset embedding model to vectorize the sentences in the answer and the candidate document respectively, and obtain the first vector corresponding to the answer and the second vector corresponding to the candidate document;
[0018] The cosine similarity between the first vector and the second vector is calculated to obtain the semantic similarity.
[0019] Furthermore, determining the relevance of each candidate document to the answer according to the coverage and the semantic similarity includes:
[0020] For each candidate document, the relevance between the candidate document and the answer is calculated using the following formula:
[0021] Y = A*a+B*b;
[0022] Among them, Y is the relevance between the candidate document and the answer, A is the coverage of the keywords in the candidate document to the keywords in the answer, a is the preset first weight, B is the semantic similarity between the answer and the candidate document, and b is the preset second weight.
[0023] Based on the above method embodiment, the present invention provides a corresponding device embodiment;
[0024] An embodiment of the present invention provides an answer tracing device for a document question-answering system, comprising: a data acquisition module, a coverage calculation module, a similarity comparison module, and an answer source determination module;
[0025] The data acquisition module is used to acquire a number of candidate documents and answers generated by the document question-answering system for questions input by the user;
[0026] The coverage calculation module is used to extract keywords from the candidate documents and the answers, and calculate the coverage of the keywords in each candidate document to the keywords in the answer based on the extracted keywords;
[0027] The similarity comparison module is used to calculate the semantic similarity between the answer and each candidate document;
[0028] The answer source determination module is used to determine the relevance between each candidate document and the answer based on the coverage and the semantic similarity, and use the candidate document with the highest relevance as the source of the answer.
[0029] Furthermore, the data acquisition module acquires several candidate documents in the following manner:
[0030] The relevance between the question and each document stored in the database is detected through a preset large language model, and the documents whose relevance exceeds a preset threshold are used as the candidate documents.
[0031] Furthermore, the coverage calculation module calculates the coverage of the keywords in the candidate document to the keywords in the answer in the following manner:
[0032] Compare the keywords in the candidate documents with the keywords in the answers, and take the same keywords as common keywords;
[0033] The coverage rate is obtained by calculating the ratio of the number of common keywords to the number of all keywords in the candidate documents.
[0034] Based on the above method embodiment, the present invention provides a corresponding terminal device embodiment;
[0035] Another embodiment of the present invention provides a terminal device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, and when the processor executes the computer program, it implements a method for tracing the answer of a document question and answer system as described in any one of the embodiments.
[0036] Based on the above method embodiment, the present invention provides a storage medium embodiment;
[0037] Another embodiment of the present invention provides a storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the storage medium is located is controlled to execute a method for tracing the answer of a document question and answer system as described in any one of the above embodiments.
[0038] The following beneficial effects are achieved by implementing the present invention:
[0039] The present invention discloses a method, device, terminal device and storage medium for tracing the answer of a document question-and-answer system. The method first obtains a number of candidate documents and answers generated by the document question-and-answer system for questions input by a user; extracts keywords from the candidate documents and the answers, and calculates the coverage of keywords in each candidate document to keywords in the answer based on the extracted keywords; calculates the semantic similarity between the answer and each candidate document; determines the relevance of each candidate document to the answer based on the coverage and the semantic similarity, and uses the candidate document with the highest relevance as the source of the answer. Compared with the prior art, the present application traces the answer based on the coverage of keywords in the candidate documents to keywords in the answer, and the semantic similarity between the answer and the candidate documents, thereby improving the accuracy of answer tracing and avoiding the problem that the answer tracing is wrong because the embedding model has a relatively small number of parameters and sometimes fails to fully understand the complex semantic relationship. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 It is a flowchart of an answer tracing method of a document question and answer system provided by an embodiment of the present invention.
[0041] Figure 2 It is a structural diagram of an answer tracing device of a document question and answer system provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0042] In order to make the purpose, technical solutions and advantages of this application clearer, the technical solutions in this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0043] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by technicians in the technical field to which this application belongs; the terms used herein are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" in the specification and claims of this application and the above-mentioned figure descriptions and any variations thereof are intended to cover non-exclusive inclusions.
[0044] In the description of the embodiments of the present application, the technical terms "first", "second", etc. are only used to distinguish different objects, and cannot be understood as indicating or implying relative importance or implicitly indicating the number, specific order or primary and secondary relationship of the indicated technical features. In the description of the embodiments of the present application, the meaning of "multiple" is more than two, unless otherwise clearly and specifically defined.
[0045] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0046] In the description of the embodiments of the present application, the term "and / or" is only a description of the association relationship of the associated objects, indicating that there may be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.
[0047] In the description of the embodiments of the present application, the term "multiple" refers to more than two (including two). Similarly, "multiple groups" refers to more than two groups (including two groups), and "multiple pieces" refers to more than two pieces (including two pieces).
[0048] In the description of the embodiments of the present application, unless otherwise clearly specified and limited, technical terms such as "installed", "connected", "connected", "fixed" and the like should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, and it can be the internal connection of two elements or the interaction relationship between two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the embodiments of the present application can be understood according to the specific circumstances.
[0049] like Figure 1 As shown, an embodiment of the present invention provides an answer tracing method for a document question-answering system, comprising the following steps:
[0050] S1. Obtain several candidate documents and answers generated by the document question answering system for questions input by the user.
[0051] In a preferred embodiment, the step of obtaining a plurality of candidate documents includes:
[0052] The relevance between the question and each document stored in the database is detected through a preset large language model, and the documents whose relevance exceeds a preset threshold are used as the candidate documents.
[0053] Specifically, various documents are stored in the database of the document question-and-answer system. The open source large model Qwen2.0-7B is used to detect the correlation between the questions input by the user and the various documents. When the correlation exceeds the preset threshold, it is determined that the document is related to the question, and the relevant documents are extracted to obtain the above-mentioned candidate documents.
[0054] S2. Extract keywords from the candidate documents and the answers, and calculate coverage of the keywords in each candidate document to the keywords in the answer based on the extracted keywords.
[0055] In a preferred embodiment, the calculation of the coverage of the keywords in each candidate document to the keywords in the answer based on the extracted keywords includes:
[0056] For each candidate document, compare the keywords in the candidate document with the keywords in the answer, and take the same keywords as common keywords;
[0057] The coverage rate is obtained by calculating the ratio of the number of common keywords to the number of all keywords in the candidate documents.
[0058] Specifically, TF-IDF or other keyword extraction algorithms are used to extract keywords from candidate documents and answers, and keywords of candidate documents and keywords of answers are obtained respectively. Then the two are compared to find keywords that both candidate documents and answers have, that is, the above-mentioned common keywords. Finally, the number of all keywords in the selected document is used as the denominator, and the number of common keywords is used as the numerator to obtain the above-mentioned keyword coverage. The keyword coverage can characterize the relevance of the answer to the candidate document to a certain extent. The higher the keyword coverage, the stronger the relevance of the answer to the candidate document.
[0059] S3. Calculate the semantic similarity between the answer and each candidate document.
[0060] In a preferred embodiment, the calculating of the semantic similarity between the answer and each candidate document includes:
[0061] For each candidate document, use the preset embedding model to vectorize the sentences in the answer and the candidate document respectively, and obtain the first vector corresponding to the answer and the second vector corresponding to the candidate document;
[0062] The cosine similarity between the first vector and the second vector is calculated to obtain the semantic similarity.
[0063] Specifically, the answer sentence and the candidate document are vectorized by using the embedding model, and the cosine similarity between the two vectors is calculated as the similarity between the answer sentence and the candidate document, so as to obtain the above semantic similarity.
[0064] S4. Determine the relevance between each candidate document and the answer based on the coverage and the semantic similarity, and use the candidate document with the highest relevance as the source of the answer.
[0065] In a preferred embodiment, determining the relevance of each candidate document to the answer based on the coverage and the semantic similarity includes:
[0066] For each candidate document, the relevance between the candidate document and the answer is calculated using the following formula:
[0067] Y = A*a+B*b;
[0068] Among them, Y is the relevance between the candidate document and the answer, A is the coverage of the keywords in the candidate document to the keywords in the answer, a is the preset first weight, B is the semantic similarity between the answer and the candidate document, and b is the preset second weight.
[0069] Specifically, after obtaining the keyword coverage and the semantic similarity between the answer and the candidate document through the above steps, the relevance between the candidate document and the answer can be calculated according to the above formula. Finally, the candidate document with the highest relevance is used as the source of the answer to complete the tracing of the answer.
[0070] Preferably, the preset first weight a can be set to 0.4, and the preset second weight b can be set to 0.6. It can be understood that the actual values of the first weight and the second weight can be adjusted according to actual conditions.
[0071] By measuring the relevance of candidate documents to answers using the two dimensions of keyword coverage and semantic similarity, we can avoid the problem of answer tracing errors caused by the failure to fully understand complex semantic relationships due to the relatively small number of parameters in the embedding model, thereby improving the accuracy of answer tracing.
[0072] Based on the implementation of the above training method, another embodiment of the present invention provides an answer tracing device for a document question and answer system.
[0073] like Figure 2 As shown, an embodiment of the present invention provides an answer tracing device for a document question-answering system, a data acquisition module, a coverage calculation module, a similarity comparison module, and an answer source determination module;
[0074] The data acquisition module is used to acquire a number of candidate documents and answers generated by the document question-answering system for questions input by the user;
[0075] The coverage calculation module is used to extract keywords from the candidate documents and the answers, and calculate the coverage of the keywords in each candidate document to the keywords in the answer based on the extracted keywords;
[0076] The similarity comparison module is used to calculate the semantic similarity between the answer and each candidate document;
[0077] The answer source determination module is used to determine the relevance between each candidate document and the answer based on the coverage and the semantic similarity, and use the candidate document with the highest relevance as the source of the answer.
[0078] In a preferred embodiment, the data acquisition module acquires several candidate documents in the following manner:
[0079] The relevance between the question and each document stored in the database is detected through a preset large language model, and the documents whose relevance exceeds a preset threshold are used as the candidate documents.
[0080] In a preferred embodiment, the coverage calculation module calculates the coverage of the keywords in the candidate document to the keywords in the answer in the following manner:
[0081] Compare the keywords in the candidate documents with the keywords in the answers, and take the same keywords as common keywords;
[0082] The coverage rate is obtained by calculating the ratio of the number of common keywords to the number of all keywords in the candidate documents.
[0083] In a preferred embodiment, the similarity comparison module calculates the semantic similarity between the answer and each candidate document in the following manner:
[0084] For each candidate document, use the preset embedding model to vectorize the sentences in the answer and the candidate document respectively, and obtain the first vector corresponding to the answer and the second vector corresponding to the candidate document;
[0085] The cosine similarity between the first vector and the second vector is calculated to obtain the semantic similarity.
[0086] In a preferred embodiment, the answer source determination module determines the relevance of each candidate document to the answer in the following manner:
[0087] For each candidate document, the relevance between the candidate document and the answer is calculated using the following formula:
[0088] Y = A*a+B*b;
[0089] Among them, Y is the relevance between the candidate document and the answer, A is the coverage of the keywords in the candidate document to the keywords in the answer, a is the preset first weight, B is the semantic similarity between the answer and the candidate document, and b is the preset second weight.
[0090] It should be noted that the device embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, in the accompanying drawings of the device embodiments provided by the present invention, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines. A person of ordinary skill in the art may understand and implement it without paying any creative effort.
[0091] Those skilled in the art can clearly understand that for the sake of convenience and brevity, the specific working process of the device described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0092] Another preferred embodiment of the present invention provides a terminal device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, it implements a method for tracing the answer of a document question and answer system as described in any one of the above embodiments.
[0093] The terminal device may be a computing device such as a desktop computer, a notebook, a PDA, a cloud server, etc. The terminal device may include, but is not limited to, a processor and a memory.
[0094] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the terminal device, and uses various interfaces and lines to connect various parts of the entire terminal device.
[0095] The memory can be used to store the computer program, and the processor realizes various functions of the terminal device by running or executing the computer program stored in the memory and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application required for a function, etc.; the data storage area can store data created according to the use of the mobile phone, etc. In addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Med i aCard, SMC), a secure digital (Secure Digital, SD) card, a flash card (F l ash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0096] Another preferred embodiment of the present invention provides a storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the storage medium is located is controlled to execute the answer tracing method of the document question and answer system described in any one of the present inventions.
[0097] The storage medium is a computer-readable storage medium, and the computer program is stored in the computer-readable storage medium. When the computer program is executed by the processor, the steps of each of the above-mentioned method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium.
[0098] The above is a preferred embodiment of the present invention. It should be pointed out that a person skilled in the art can make several improvements and modifications without departing from the principle of the present invention. These improvements and modifications are also considered to be within the scope of protection of the present invention.
Claims
1. A method for tracing the answer source of a document question-answering system, characterized in that: include: Obtaining a number of candidate documents and answers generated by the document question answering system for questions input by the user; Extracting keywords from the candidate documents and the answers, and calculating coverage of the keywords in each candidate document to the keywords in the answer based on the extracted keywords; Calculate the semantic similarity between the answer and each candidate document; The relevance between each candidate document and the answer is determined according to the coverage and the semantic similarity, and the candidate document with the highest relevance is used as the source of the answer.
2. The answer tracing method of the document question-answering system according to claim 1, characterized in that: The obtaining of several candidate documents includes: The relevance between the question and each document stored in the database is detected through a preset large language model, and the documents whose relevance exceeds a preset threshold are used as the candidate documents.
3. The answer tracing method of the document question-answering system according to claim 2, characterized in that: The step of calculating the coverage of the keywords in each candidate document to the keywords in the answer according to the extracted keywords includes: For each candidate document, compare the keywords in the candidate document with the keywords in the answer, and take the same keywords as common keywords; The coverage rate is obtained by calculating the ratio of the number of common keywords to the number of all keywords in the candidate documents.
4. The answer tracing method of the document question-answering system according to claim 3, characterized in that: The calculation of the semantic similarity between the answer and each candidate document includes: For each candidate document, use the preset embedding model to vectorize the sentences in the answer and the candidate document respectively, and obtain the first vector corresponding to the answer and the second vector corresponding to the candidate document; The cosine similarity between the first vector and the second vector is calculated to obtain the semantic similarity.
5. The answer tracing method of the document question-answering system according to claim 4, characterized in that: Determining the relevance between each candidate document and the answer according to the coverage and the semantic similarity includes: For each candidate document, the relevance between the candidate document and the answer is calculated using the following formula: Y = A*a+B*b; Among them, Y is the relevance between the candidate document and the answer, A is the coverage of the keywords in the candidate document to the keywords in the answer, a is the preset first weight, B is the semantic similarity between the answer and the candidate document, and b is the preset second weight.
6. A device for tracing the source of answers in a document question-answering system, characterized in that: include: Data acquisition module, coverage calculation module, similarity comparison module and answer source determination module; The data acquisition module is used to acquire a number of candidate documents and answers generated by the document question-answering system for questions input by the user; The coverage calculation module is used to extract keywords from the candidate documents and the answers, and calculate the coverage of the keywords in each candidate document to the keywords in the answer based on the extracted keywords; The similarity comparison module is used to calculate the semantic similarity between the answer and each candidate document; The answer source determination module is used to determine the relevance between each candidate document and the answer based on the coverage and the semantic similarity, and use the candidate document with the highest relevance as the source of the answer.
7. The answer tracing device of the document question-answering system according to claim 6, characterized in that: The data acquisition module acquires several candidate documents in the following manner: The relevance between the question and each document stored in the database is detected through a preset large language model, and the documents whose relevance exceeds a preset threshold are used as the candidate documents.
8. The answer tracing device of the document question-answering system according to claim 7, characterized in that: The coverage calculation module calculates the coverage of the keywords in the candidate document to the keywords in the answer in the following way: Compare the keywords in the candidate documents with the keywords in the answers, and take the same keywords as common keywords; The coverage rate is obtained by calculating the ratio of the number of common keywords to the number of all keywords in the candidate documents.
9. A terminal device, characterized in that: It includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, and when the processor executes the computer program, it implements the answer tracing method of the document question and answer system as described in any one of claims 1 to 5.
10. A storage medium, characterized in that: The storage medium includes a stored computer program, wherein when the computer program is running, the device where the storage medium is located is controlled to execute the answer tracing method of the document question and answer system as described in any one of claims 1 to 5.