Document answering method and apparatus, and electronic device and non-volatile storage medium
Patent Information
- Application Number
- PCT/CN2025/121639
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-26
- Filing Date
- 2025-09-16
- Publication Date
- 2026-10-01
Smart Images

Figure CN2025121639_01102026_PF_FP_ABST
Abstract
Description
Document response methods, devices, electronic equipment, and non-volatile storage media
[0001] Related applications
[0002] This application claims priority to Chinese patent application filed on March 26, 2025, with application number 202510370300.7, entitled "Document Response Method, Apparatus, Electronic Device and Non-Volatile Storage Medium", the entire contents of which are incorporated herein by reference. Technical Field
[0003] This application relates to the field of artificial intelligence, and more specifically, to a document response method, apparatus, electronic device, and non-volatile storage medium. Background Technology
[0004] With the rapid development of artificial intelligence technology, especially the continuous innovation in the field of natural language processing, document response systems in related technologies mainly rely on rule-based algorithms to parse document content and generate responses. While this method is efficient in processing structured data and simple documents, its accuracy and intelligence are clearly insufficient when dealing with unstructured, lengthy, or highly specialized documents. In recent years, intelligent document response systems integrating Retrieval Augmented Generation (RAG) technology have begun to emerge. RAG technology, by combining document retrieval and language model generation capabilities, can generate more accurate and natural responses based on understanding document content, significantly improving user experience and system performance. However, the response results are often still unsatisfactory. The document response methods in related technologies have limited document understanding capabilities and cannot accurately identify key information in the document, resulting in low accuracy of document responses.
[0005] There is currently no effective solution to the above problems. Summary of the Invention
[0006] This application provides a document response method, apparatus, electronic device, and non-volatile storage medium.
[0007] According to one aspect of the embodiments of this application, a document response method is provided, comprising: obtaining prompt words input by a user; the prompt words including at least a user question to be answered and a document file corresponding to the user question; determining a first matching score between the user question and each paragraph in the document file using a first method; the first method including: determining the first matching score based on keyword matching; determining a second matching score between the user question and each paragraph in the document file using a second method; the second method including: determining the second matching score based on question-answer pair matching; determining a third matching score between the user question and each paragraph in the document file using a third method; the third method including: determining the third matching score based on expanding the user question; and determining a target answer corresponding to the user question based on the first matching score, second matching score, and third matching score of each paragraph.
[0008] In some embodiments of this application, determining the first matching score based on keyword matching includes: extracting a first preset number of question keywords from the user's question to obtain a question keyword set; for each paragraph in the document file, extracting a second preset number of paragraph keywords to obtain a paragraph keyword set for each paragraph; and determining the first matching score for each paragraph based on the question keyword set and the paragraph keyword set for each paragraph.
[0009] In some embodiments of this application, determining the first matching score of each paragraph based on the question keyword set and the paragraph keyword set of each paragraph includes: determining the matching set of each paragraph by the intersection of the question keyword set and the paragraph keyword set of each paragraph; determining the keyword matching score based on the number of matching keywords in the matching set, the total number of question keywords in the question keyword set, and the total number of paragraph keywords in the paragraph keyword set; determining the total number of matching characters corresponding to all keywords in the matching set; recording repeated characters only once in the total number of matching characters; determining the character matching score by the total number of matching characters, the total number of first characters, and the total number of second characters; the total number of first characters is the total number of characters corresponding to all question keywords in the question keyword set, and the total number of second characters is the total number of characters corresponding to all paragraph keywords in the paragraph keyword set of each paragraph; recording repeated characters only once in the total number of first characters and the total number of second characters; and determining the first matching score by the keyword matching score and the character matching score of each paragraph.
[0010] In some embodiments of this application, determining the second matching score based on question-answer pair matching includes: extracting summary information of the document file; determining the target audience of the document file based on the summary information; determining a preset number of question-answer pairs corresponding to each paragraph based on the summary information and the target audience; one question-answer pair corresponds to one reference question and one reference answer; clustering the preset number of question-answer pairs for each paragraph to obtain a target question-answer pair set, and storing the target question-answer pair set corresponding to each paragraph in a vector library; determining the similarity score between the user's question and the reference questions in the target question-answer pair set corresponding to each paragraph in the vector library, and determining the maximum similarity score of each paragraph as the second matching score of each paragraph.
[0011] In some embodiments of this application, determining a preset number of question-answer pairs for each paragraph based on summary information and the target audience includes: identifying entities, actions of entities, and relationships between entities in the summary information, and determining the topics and contexts covered by each paragraph based on the entities, actions of entities, and relationships between entities; determining an initial list of question-answer pairs corresponding to the topics and contexts covered by each paragraph based on the attributes of the target audience; the attributes include at least professional background, interests and preferences, and common contexts; and semantically expanding the initial list of question-answer pairs for each paragraph to obtain a preset number of question-answer pairs for each paragraph.
[0012] In some embodiments of this application, the third matching score is determined based on the expansion of the user question, including: extracting paragraph summary information of each paragraph in the document file; expanding the user question to obtain a third preset number of similar questions, and determining the user question and similar questions as a question set; determining the semantic similarity between the questions in the question set and the paragraph summary information of each paragraph, and determining the maximum semantic similarity as the third matching score of each paragraph.
[0013] In some embodiments of this application, the user question is expanded to obtain a third preset number of similar questions, and the user question and the similar questions are determined to form a question set, including: performing semantic parsing on the user question to extract the core intent; determining a third preset number of similar questions based on the core intent and question keywords; the similar questions have the same core intent as the user question, and the similar questions cover different expressions and different questioning angles of the user question; performing grammatical verification on the similar questions, and forming a question set with the user question after grammatical verification of the similar questions.
[0014] In some embodiments of this application, determining the target answer to a user question based on a first matching score, a second matching score, and a third matching score for each paragraph includes: determining the target paragraph in the document file based on each first matching score, second matching score, and third matching score; and generating the target answer to the user question based on the target paragraph.
[0015] In some embodiments of this application, determining the target paragraph in the document file based on each first matching score, second matching score, and third matching score includes: determining the comprehensive score of each paragraph based on the first matching score, second matching score, and third matching score of each paragraph; and filtering the target paragraph from the document file based on the comprehensive score of each paragraph.
[0016] In some embodiments of this application, generating a target answer to a user question based on a target paragraph includes: processing the target paragraph based on a pre-trained large language model and outputting the target answer to the user question.
[0017] According to another aspect of the embodiments of this application, a document response apparatus is also provided, comprising: an acquisition module, configured to acquire prompt words input by a user; the prompt words include at least a user question to be answered and a document file corresponding to the user question; a keyword matching determination module, configured to determine a first matching score between the user question and each paragraph in the document file using a first method; the first method includes determining the first matching score based on keyword matching; a question-answer pair matching determination module, configured to determine a second matching score between the user question and each paragraph in the document file using a second method; the second method includes determining the second matching score based on question-answer pair matching; a user question expansion determination module, configured to determine a third matching score between the user question and each paragraph in the document file using a third method; the third method includes determining the third matching score based on expanding the user question; and a target answer determination module, configured to determine a target answer corresponding to the user question based on the first matching score, second matching score, and third matching score of each paragraph.
[0018] According to another aspect of the embodiments of this application, a non-volatile storage medium is also provided, wherein a program is stored in the non-volatile storage medium; the program, when running, controls the device where the non-volatile storage medium is located to execute the document response method of any one of the above.
[0019] According to another aspect of the embodiments of this application, an electronic device is also provided, including: a memory and a processor, wherein the processor is configured to run a program stored in the memory; and the program executes the steps of any of the above document response methods when it runs.
[0020] According to another aspect of the embodiments of this application, a computer program product is also provided, including computer instructions, which, when executed by a processor, implement the steps of the document response method described above. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of this application or the conventional technology, the drawings used in the description of the embodiments or the conventional technology will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the disclosed drawings without creative effort.
[0022] Figure 1 is a hardware structure block diagram of a computer terminal for implementing a document response method according to an embodiment of this application;
[0023] Figure 2 is a flowchart of a document response method according to an embodiment of this application;
[0024] Figure 3 is a flowchart of a method for determining a first matching score between a user question and each paragraph in a document file, according to an embodiment of this application.
[0025] Figure 4 is a flowchart of a method for determining a second matching score between a user question and each paragraph in a document file, according to an embodiment of this application.
[0026] Figure 5 is a flowchart illustrating a method for determining a third matching score between a user question and each paragraph in a document file, according to an embodiment of this application.
[0027] Figure 6 is a schematic diagram of a document response device provided according to an embodiment of this application. Detailed Implementation
[0028] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0029] The information collected in this application embodiment is information and data authorized by the user or fully authorized by all parties. The collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data all comply with the relevant laws, regulations and standards of the relevant regions, and necessary confidentiality measures have been taken. It does not violate public order and good morals, and provides corresponding operation entry points for users to choose to authorize or reject the automated decision results. If the user chooses to reject, the process will proceed to the expert decision-making process.
[0030] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0031] Current document response systems based on Retrieval Augmented Generation (RAG) still face a series of challenges in practical applications. First, document processing and paragraph segmentation technologies are not yet mature, making it difficult to accurately capture the semantic structure of document content. This is especially true when dealing with long documents, multilingual documents, or documents containing a large amount of technical terminology. Segmentation methods often fail to effectively preserve the complete information and contextual relationships of paragraphs, resulting in limited document comprehension and affecting the accuracy of subsequent retrieval and response. Second, the synergy between the retrieval component and the generation model in RAG technology is insufficient. The relevance calculation method between retrieved paragraphs and user questions is relatively simplistic, leading to recalled paragraphs that may not match the deeper semantic requirements of the question, failing to accurately identify key information in the document. This results in a lack of flexibility and scalability in handling user questions, making it difficult to meet the diverse needs of users in different scenarios. Based on these issues, this application proposes a document response method, apparatus, electronic device, and non-volatile storage medium.
[0032] To better understand the embodiments of this application, the technical terms involved in the embodiments of this application are explained below:
[0033] Retrieval Augmented Generation (RAG) is a natural language processing method that combines Information Retrieval (IR) and Neural Language Generation (NLG) techniques. It is primarily used to generate high-quality responses based on documents or large amounts of text data. RAG is designed to address the knowledge limitations that generative AI relying solely on pre-trained model knowledge bases may encounter during the response process.
[0034] LLM (Large Language Model): This is a significant technological breakthrough in the field of Natural Language Processing (NLP). Its core characteristic is the sheer number of model parameters, typically reaching billions or even trillions. Models of this scale, pre-trained on massive amounts of text data, can learn extremely rich language patterns and knowledge, thus exhibiting performance superior to conventional models in various NLP tasks.
[0035] In related technologies, document processing and paragraph segmentation techniques in document response methods are not yet mature, making it difficult to accurately capture the semantic structure of document content. This is especially true when dealing with long documents, multilingual documents, or documents containing a large amount of technical terminology. Segmentation methods often fail to effectively preserve the complete information and contextual relationships of paragraphs, resulting in limited document comprehension and impacting the accuracy of subsequent retrieval and response. Secondly, the synergy between the retrieval component and the generation model in RAG technology is insufficient, and the method for calculating the relevance of retrieved paragraphs to user questions is relatively simplistic. This leads to recalled paragraphs potentially not matching the deeper semantic needs of the question, failing to accurately identify key information in the document, and lacking flexibility and scalability in handling user questions, making it difficult to meet the diverse needs of users in different scenarios. Therefore, related document response methods suffer from limited document comprehension capabilities, failing to accurately identify key information in documents, resulting in low accuracy in document responses. To address this problem, this application provides relevant solutions, which are detailed below.
[0036] According to an embodiment of this application, an embodiment of a document response method is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0037] The method embodiments provided in this application can be executed in a computer terminal or similar computing device. Figure 1 shows a hardware structure block diagram of a computer terminal for implementing a document response method. As shown in Figure 1, the computer terminal 10 may include one or more (shown as 102a, 102b, ..., 102n in the figure) processors 102 (processors 102 may include, but are not limited to, microprocessors such as MCUs or programmable logic devices such as FPGAs), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, the computer terminal 10 may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of a BUS bus), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that the structure shown in Figure 1 is merely illustrative and does not limit the structure of the above-described electronic device. For example, the computer terminal 10 may also include more or fewer components than shown in Figure 1, or have a different configuration than shown in Figure 1.
[0038] It should be noted that the aforementioned one or more processors 102 and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within the computer terminal 10. As involved in the embodiments of this application, the data processing circuits serve as a form of processor control (e.g., selection of a variable resistor termination path connected to an interface).
[0039] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the document response method in this embodiment. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, thereby realizing the document response method described above. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the computer terminal 10 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0040] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the computer terminal 10. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module, used for wireless communication with the Internet.
[0041] The display may be, for example, a touchscreen liquid crystal display (LCD) that allows the user to interact with the user interface of the computer terminal 10.
[0042] In the above operating environment, this application provides an embodiment of a document response method. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0043] Figure 2 shows a flowchart of a document response method according to an embodiment of this application, including:
[0044] Step S202: Obtain the prompt words input by the user; the prompt words include at least the user question to be answered and the document file corresponding to the user question.
[0045] In step S202, the large model of the document response system obtains the prompt words input by the user. The prompt words include at least the user question to be answered and the document file corresponding to the user question. For example, the following is a prompt word example: User question: "How can unemployed people withdraw their housing provident fund?" The document file corresponding to the user question is a PDF file about housing provident fund policies, or "housing provident fund policy" is entered as a keyword.
[0046] Step S204: Determine the first match score between the user question and each paragraph in the document file using a first method; the first method includes: determining the first match score based on keyword matching.
[0047] In step S204, there are several ways to determine the first matching score based on keyword matching. For example, extract a first preset number of question keywords from the user's question to obtain a question keyword set; for each paragraph in the document file, extract a second preset number of paragraph keywords to obtain a paragraph keyword set for each paragraph; and determine the first matching score for each paragraph based on the question keyword set and the paragraph keyword set for each paragraph. There are several ways to determine the first matching score for each paragraph based on the question keyword set and the paragraph keyword set for each paragraph. For example: The intersection of the question keyword set and the paragraph keyword set for each paragraph is used to determine the matching set for each paragraph; the keyword matching score is determined based on the number of matching keywords in the matching set, the total number of question keywords in the question keyword set, and the total number of paragraph keywords in the paragraph keyword set; the total number of matching characters corresponding to all keywords in the matching set is determined; repeated characters are recorded only once in the total number of matching characters; the character matching score is determined by the total number of matching characters, the total number of first characters, and the total number of second characters; the total number of first characters is the total number of characters corresponding to all question keywords in the question keyword set, and the total number of second characters is the total number of characters corresponding to all paragraph keywords in the paragraph keyword set for each paragraph; repeated characters are recorded only once in the total number of first characters and the total number of second characters; the first matching score is determined by the keyword matching score and the character matching score for each paragraph.
[0048] The following are specific examples:
[0049] The document response system's large model receives prompts corresponding to keyword matching methods, such as "Based on the information I provide, please summarize and extract 3 to 5 keywords that can express the core semantics of the document." It segments the user's question, removes stop words, and finally extracts a first preset number (e.g., 3 to 5) of question keywords from the user's question to obtain a question keyword set. For example, for "How can unemployed people withdraw their housing provident fund?", the final extracted keywords are "unemployed people", "housing provident fund withdrawal", and "housing provident fund" to form the question keyword set.
[0050] The following method can be used to segment user questions and remove stop words: The question text corresponding to the user question is segmented into a series of words according to semantic units, and stop words are removed. Then, based on the semantics and context of the question, a keyword extraction algorithm is used to find the most important and relevant words from the word list. Keyword extraction can be based on techniques such as term frequency, term frequency-inverse document frequency (TF-IDF), and term frequency-inverse document frequency weighted word embedding. Alternatively, deep learning-based methods can be used for keyword scoring. The top N keywords (N being a preset number, e.g., 3 to 5) from the extracted keyword list are selected to form the question keyword set. The selected number usually needs to be adjusted according to the complexity of the question and the application scenario to ensure that the set contains words that best represent the core of the question. Similarly, for each paragraph in the document file, a second preset number (e.g., 3-5) of paragraph keywords are extracted to obtain a paragraph keyword set for each paragraph. Then, the intersection between the question keyword set and each paragraph keyword set is determined, i.e., keywords that appear in both the question keyword set and the paragraph keyword set are found to form a matching set for each paragraph. The number of matching keywords in the matching set is then determined. The keyword matching score for each paragraph is determined based on the number of matching keywords in the matching set, the total number of question keywords in the question keyword set, and the total number of paragraph keywords in the paragraph keyword set. For example, the keyword matching score can be determined using the following formula:
[0051] score1 (keyword matching score) = number of matched keywords / set (total number of question keywords + total number of paragraph keywords), where set refers to the sum of the total number of question keywords in the question keyword set and the total number of paragraph keywords in the paragraph keyword set, while removing duplicate elements.
[0052] Then, determine the total number of matching characters corresponding to all keywords in the matching set. Duplicate characters are recorded only once in the total number of matching characters, ensuring each character is counted only once. Simultaneously, determine the total number of characters corresponding to all question keywords in the question keyword set (first total number of characters), and the total number of characters corresponding to all paragraph keywords in each paragraph's paragraph keyword set (second total number of characters), removing duplicate characters. Determine the character matching score for each paragraph using the total number of matching characters, the first total number of characters, and the second total number of characters, for example, using the following formula:
[0053] score2 (i.e., character matching score) = total number of matched characters / length of set(question keyword set + paragraph keyword set). The length of set(question keyword set + paragraph keyword set) represents the union of the characters in the question keyword set + paragraph keyword set, while removing duplicate characters. It is the sum of the total number of the first character and the total number of the second character.
[0054] The first match score for each paragraph is determined based on the character match score, keyword match score, and preset weights for each paragraph. For example, it can be determined by the following formula: score (first match score) = score1 * first preset weight (e.g., 50%) + score2 * second preset weight (e.g., 50%).
[0055] Figure 3 shows a flowchart of determining the first matching score between a user question and each paragraph in a document file using a first method according to an embodiment of this application. First, the user question is segmented into words and stop words are removed. Then, keyword 1 and keyword 2 are extracted (i.e., the first preset number of question keywords extracted from the user question are obtained to form a question keyword set). Then, for each paragraph in the document file (paragraph 1 to paragraph N, where N is an integer greater than 1), keyword extraction is performed (i.e., for each paragraph in the document file, the second preset number of paragraph keywords are extracted to obtain a paragraph keyword set for each paragraph). The keywords and keyword co-occurrence frequency of the question and paragraph are determined (keywords and keyword co-occurrence frequency are the keyword matching score and character matching score mentioned above). Finally, the final score is generated (i.e., the first matching score is determined by the keyword matching score and character matching score of each paragraph mentioned above).
[0056] Step S206: Use a second method to determine the second matching score between the user's question and each paragraph in the document file; the second method includes: determining the second matching score based on question-and-answer pair matching.
[0057] In step S206, there are several ways to determine the second matching score based on question-answer pair matching, such as: extracting the summary information of the document file; determining the target audience of the document file based on the summary information; determining a preset number of question-answer pairs for each paragraph based on the summary information and the target audience; one question-answer pair corresponds to one reference question and one reference answer; clustering the preset number of question-answer pairs for each paragraph to obtain a target question-answer pair set, and storing the target question-answer pair set for each paragraph in a vector library; determining the similarity score between the user's question and the reference questions in the target question-answer pair set for each paragraph in the vector library, and determining the similarity score of each paragraph as the second matching score for each paragraph.
[0058] It is important to note that there are multiple ways to determine the preset number of question-answer pairs for each paragraph based on the summary information and the target audience in the above steps. For example: identify entities, actions of entities, and relationships between entities in the summary information, and determine the topics and contexts covered by each paragraph based on the entities, actions of entities, and relationships between entities; determine an initial list of question-answer pairs corresponding to the topics and contexts covered by each paragraph based on the attributes of the target audience, where the attributes include at least professional background, interests and preferences, and common contexts; and semantically expand the initial list of question-answer pairs for each paragraph to obtain the preset number of question-answer pairs for each paragraph.
[0059] The following is a specific implementation: Large models (e.g., LLM) convert document files (such as PDF, Word, etc.) into plain text format. For non-text documents, optical character recognition (OCR) technology can be used to scan and convert them into text. Irrelevant information in the text, such as headers, footers, annotations, image descriptions, etc., is removed, retaining the main text content. The document text is segmented according to paragraphs, sections, and other structures for subsequent processing. Sentences in the text are broken down into words, and the part of speech of each word is labeled, providing a foundation for subsequent semantic understanding. Named entities in the text, such as names of people, places, organizations, etc., are identified; these entities are often the information that needs to be highlighted in the summary. Each sentence is scored based on the frequency of paragraph keywords, positional information (such as headings, paragraph beginnings), and syntactic structure, identifying the sentence that contributes the most to the summary. The top-scoring sentences are selected and concatenated in the original text order to form the document summary. To ensure the coherence and completeness of the abstract, and to make necessary sentence adjustments, a comprehensive and coherent abstract text is regenerated. The abstract does not have to be exactly the same as the sentences in the original text; it is a condensed expression of the original meaning.
[0060] The process involves identifying entities (such as names of people, organizations, and places) and their actions (the actions performed by the entities within the summary information), and extracting the relationships between these entities (such as affiliation, influence, causality, etc.). Based on the entities, their actions, and the relationships between them, the process determines the topic and context covered by each paragraph. The topic is the core theme of the paragraph, while the context includes background information and detailed descriptions of the paragraph's content. Audience attribute analysis collects attribute information about the audience, including but not limited to professional background, interests, and common usage contexts. This attribute information can be obtained through user registration information, historical interaction data, or questionnaires. Based on the audience attributes, an initial question-and-answer pair list related to the topic and context of each paragraph is generated. For example, for an article about health insurance targeting "office workers," possible initial question-and-answer pairs might include: "What is the special significance of health insurance for office workers?" and "How to choose a suitable health insurance plan for office workers?" Finally, the initial question-and-answer pairs are semantically expanded to generate more related questions, ensuring that the questions cover multiple aspects of the paragraph. A preset number of question-and-answer pairs (e.g., 10) are obtained for each paragraph. Then, all questions for each paragraph are clustered, with the number of clusters being a preset number (e.g., 10) of question-answer pairs divided by the total number of the audience. The target question-answer pair set for each paragraph is then integrated using a large model (or a fusion model) and stored in a vector library.
[0061] Determine the similarity score between the user's question and the reference questions in the target question-answer pair set corresponding to each paragraph in the vector library, and determine the maximum similarity score of each paragraph as the second matching score for each paragraph, for example, through the following formula:
[0062] Second matching score = max(cosin(bert_embedding(user question), bert_embedding(reference question))).
[0063] Wherein, (cosin(bert_embedding(user question), bert_embedding(reference question)) represents the similarity score of the reference questions in the target question-answer pair set corresponding to each paragraph in the database, where bert_embedding is a vector representation, and max(cosin(bert_embedding(user question), bert_embedding(reference question))) means taking the similarity score of the reference question with the highest similarity to each paragraph as the second matching score of that paragraph.
[0064] Figure 4 shows a flowchart illustrating a method for determining the second matching score between a user's question and each paragraph in a document file, according to an embodiment of this application. The method involves extracting an article summary based on LLM (i.e., extracting the summary information of the document file as described above), extracting the article audience based on LLM (i.e., determining the audience of the document file based on the summary information as described above), extracting question-answer pairs for each paragraph (paragraph 1 to paragraph N) based on paragraph + role + summary (i.e., determining a preset number of question-answer pairs corresponding to each paragraph based on the summary information and the audience as described above), clustering all questions, and integrating the final question-answer pairs using a fusion model and storing them in a vector library (i.e., clustering the preset number of question-answer pairs for each paragraph to obtain a target question-answer pair set, and storing the target question-answer pair set corresponding to each paragraph in the vector library). The similarity score between the user's question and the reference questions in the target question-answer pair set corresponding to each paragraph in the vector library is determined, and the maximum similarity score for each paragraph is determined as the second matching score for each paragraph.
[0065] Step S208 is to determine the third matching score between the user question and each paragraph in the document file using a third method according to the embodiments of this application; the third method includes: determining the third matching score based on the method of expanding the user question.
[0066] In step S208, there are several ways to determine the third matching score based on the expansion of the user question. For example, extract the paragraph summary information of each paragraph in the document file; expand the user question to obtain a third preset number of similar questions, and determine the user question and similar questions as a question set; determine the semantic similarity between the questions in the question set and the paragraph summary information of each paragraph, and determine the maximum semantic similarity as the third matching score of each paragraph.
[0067] In the above steps, expanding the user question to obtain a third preset number of similar questions and determining the user question and similar questions as a question set can be achieved in several ways, such as: performing semantic parsing on the user question to extract the core intent; determining a third preset number of similar questions based on the core intent and question keywords, wherein the similar questions have the same core intent as the user question, and the similar questions cover different expressions and different question angles of the user question; performing grammatical validation on the similar questions, and forming a question set with the user question after grammatical validation of the similar questions.
[0068] The following are specific implementation methods and examples:
[0069] Using NLP techniques, such as summarization based on Bidirectional Encoder Representations from Transformers (BERT), a summary is extracted from each paragraph in the document file to generate paragraph summary information. A deep learning model (such as BERT) is used to semantically parse the user question, identifying and extracting its core intent. Based on the core intent and keywords, a pre-trained language model (such as BART) or rule-based methods are used to generate a third preset number (e.g., 5-10) of similar questions. These similar questions should maintain the same core intent as the original question but cover different expressions and questioning angles to improve the coverage and flexibility of the questions. The generated similar questions are grammatically validated to ensure that each question is grammatically correct and fluent, avoiding grammatical errors that could affect subsequent semantic similarity determination. The grammatically validated similar questions are combined with the original user question to form a question set. The BERT model is used to convert the summary information of each question in the question set and each paragraph in the document into a vector representation. The semantic similarity between each question vector and each paragraph summary information vector in the question set is determined. The semantic similarity scores of each question with the summary information of the same paragraph are summed, and the maximum value is used as the third matching score for that paragraph. The maximum semantic similarity is the maximum value among the semantic similarities; for example, the third matching score can be determined using the following formula:
[0070] The third matching score = max(cosin(bert_embedding(question set), bert_embedding(paragraph summary information))), where cosin(bert_embedding(question set), bert_embedding(paragraph summary information)) represents determining the semantic similarity between each question vector in the question set and each paragraph summary information vector, and max(cosin(bert_embedding(question set), bert_embedding(paragraph summary information))) represents determining the third matching score for each paragraph by setting the maximum semantic similarity (i.e., the maximum value of semantic similarity).
[0071] Figure 5 shows a flowchart of a method for determining the third matching score between a user question and each paragraph in a document file, according to an embodiment of this application. First, the user question is expanded using a fusion model (i.e., the user question is expanded to obtain a third preset number of similar questions). Then, for each paragraph (paragraph 1 to paragraph N), a fusion model is used to summarize the paragraphs (i.e., the paragraph summary information of each paragraph in the document file is extracted). The vector distance between the expanded question and the paragraph summary is determined (i.e., the user question and similar questions are determined as a question set; the semantic similarity between the questions in the question set and the paragraph summary information of each paragraph is determined). Finally, a score is generated (i.e., the maximum semantic similarity is determined as the third matching score of each paragraph).
[0072] Step S210: Determine the target answer corresponding to the user's question based on the first matching score, second matching score, and third matching score of each paragraph.
[0073] Specifically, the target paragraph in the document file can be determined based on the first matching score, the second matching score, and the third matching score, and the target answer can be generated based on the target paragraph. For example, the comprehensive score of each paragraph can be determined based on the first matching score, the second matching score, and the third matching score; based on the above comprehensive score, the target paragraph most relevant to the user's question can be selected from all paragraphs in the document file. For example, the target paragraph in the document file can be determined by setting a preset comprehensive score threshold; a pre-trained large language model can be used to process the target paragraph and output the target answer corresponding to the user's question.
[0074] In this embodiment, a method is employed to obtain prompts input by the user. These prompts include at least the user question to be answered and the document file corresponding to the user question. A first matching score is determined between the user question and each paragraph in the document file using a first method. This first method includes determining the first matching score based on keyword matching. A second matching score is determined between the user question and each paragraph in the document file using a second method. This second method includes determining the second matching score based on question-and-answer pair matching. A third matching score is determined between the user question and each paragraph in the document file using a third method. This third method includes determining the third matching score based on expanding the user question. The target answer for the user question is determined based on the first, second, and third matching scores of each paragraph. By using three different methods to determine the matching score of the document file corresponding to the user question, and finally determining the target answer based on the three matching scores, this method achieves a deeper understanding of the document file and the user question, accurately identifies key information in the document file, and thus improves the accuracy of question-and-answer responses. This solves the technical problem in related technologies where document response methods have limited document understanding capabilities, cannot accurately identify key information in the document, and result in low accuracy in document responses.
[0075] This application also provides a document response device, as shown in FIG6, including:
[0076] The acquisition module 602 is used to acquire the prompt words input by the user; the prompt words include at least the user question to be answered and the document file corresponding to the user question.
[0077] Keyword matching determination module 604 is used to determine the first match score between the user question and each paragraph in the document file using a first method; the first method includes: determining the first match score based on keyword matching.
[0078] The question-and-answer pair matching determination module 606 is used to determine a second matching score between the user's question and each paragraph in the document file using a second method; the second method includes: determining the second matching score based on question-and-answer pair matching.
[0079] User question expansion determination module 608 is used to determine the third match score between the user question and each paragraph in the document file using a third method; the third method includes: determining the third match score based on the method of expanding the user question.
[0080] The target answer determination module 610 is used to determine the target answer corresponding to the user's question based on the first matching score, the second matching score, and the third matching score of each paragraph.
[0081] The keyword matching determination module 604 is also used to: extract a first preset number of question keywords from the user's question to obtain a question keyword set; for each paragraph in the document file, extract a second preset number of paragraph keywords to obtain a paragraph keyword set for each paragraph; and determine the first matching score for each paragraph based on the question keyword set and the paragraph keyword set for each paragraph.
[0082] The keyword matching determination module 604 is further configured to: determine the matching set of each paragraph by the intersection of the question keyword set and the paragraph keyword set of each paragraph; determine the keyword matching score based on the number of matching keywords in the matching set, the total number of question keywords in the question keyword set, and the total number of paragraph keywords in the paragraph keyword set; determine the total number of matching characters corresponding to all keywords in the matching set; record repeated characters only once in the total number of matching characters; determine the character matching score by the total number of matching characters, the total number of first characters, and the total number of second characters; the total number of first characters is the total number of characters corresponding to all question keywords in the question keyword set, and the total number of second characters is the total number of characters corresponding to all paragraph keywords in the paragraph keyword set of each paragraph; record repeated characters only once in the total number of first characters and the total number of second characters; determine the first matching score by the keyword matching score and the character matching score of each paragraph.
[0083] The question-answer pair matching determination module 606 is also used for: extracting summary information of the document file; determining the target audience of the document file based on the summary information; determining a preset number of question-answer pairs for each paragraph based on the summary information and the target audience; one question-answer pair corresponds to one reference question and one reference answer; clustering the preset number of question-answer pairs for each paragraph to obtain a target question-answer pair set, and storing the target question-answer pair set for each paragraph in a vector library; determining the similarity score between the user's question and the reference question in the target question-answer pair set for each paragraph in the vector library, and determining the maximum similarity score for each paragraph as the second matching score for each paragraph.
[0084] The question-answer pair matching determination module 606 is also used to: identify entities, actions of entities, and relationships between entities in the summary information, and determine the topics and contexts covered by each paragraph based on the entities, actions of entities, and relationships between entities; determine an initial question-answer pair list corresponding to the topics and contexts covered by each paragraph based on the attributes of the audience; the attributes include at least professional background, interest preferences, and common contexts; and semantically expand the initial question-answer pair list corresponding to each paragraph to obtain a preset number of question-answer pairs corresponding to each paragraph.
[0085] The user question expansion and determination module 608 is also used to: extract paragraph summary information of each paragraph in the document file; expand the user question to obtain a third preset number of similar questions, and determine the user question and similar questions as a question set; determine the semantic similarity between the questions in the question set and the paragraph summary information of each paragraph, and determine the maximum semantic similarity as the third matching score of each paragraph.
[0086] The user question expansion determination module 608 is also used for: performing semantic parsing on the user question to extract the core intent; determining a third preset number of similar questions based on the core intent and question keywords; the similar questions have the same core intent as the user question, and the similar questions cover different expressions and different question angles of the user question; performing grammatical verification on the similar questions, and combining the grammatically verified similar questions with the user question to form a question set.
[0087] The target answer determination module 610 is also used to: determine the target paragraph in the document file based on each first matching score, second matching score and third matching score; and generate the target answer corresponding to the user question based on the target paragraph.
[0088] The target answer determination module 610 is also used to: determine the overall score of each paragraph based on the first matching score, the second matching score, and the third matching score of each paragraph; and filter the target paragraphs from the document file based on the overall score of each paragraph.
[0089] The target answer determination module 610 is also used to: process the target paragraph based on the pre-trained large language model and output the target answer corresponding to the user's question.
[0090] It should be noted that the document response device shown in Figure 6 is used to execute the document response method shown in Figure 2. Therefore, the relevant explanations in the document response method in Figure 2 also apply to this document response device, and will not be repeated here.
[0091] It should be noted that each module in the above-mentioned document response device can be a program module (e.g., a set of program instructions to implement a certain function) or a hardware module. For the latter, it can be manifested in the following forms, but is not limited to them: each of the above modules is manifested as a processor, or the functions of each of the above modules are implemented by a processor.
[0092] This application also provides a non-volatile storage medium, which includes a stored program; during program execution, the device where the non-volatile storage medium is located executes the above-described document response method. For example, it obtains prompt words input by the user; the prompt words include at least a user question to be answered and a document file corresponding to the user question; it uses a first method to determine a first matching score between the user question and each paragraph in the document file; the first method includes: determining the first matching score based on keyword matching; it uses a second method to determine a second matching score between the user question and each paragraph in the document file; the second method includes: determining the second matching score based on question-answer pair matching; it uses a third method to determine a third matching score between the user question and each paragraph in the document file; the third method includes: determining the third matching score based on expanding the user question; and it determines the target answer corresponding to the user question based on the first matching score, second matching score, and third matching score of each paragraph.
[0093] This application also provides an electronic device, which includes a processor for running a program; during program execution, the above-described document response method is performed. For example, the method involves: acquiring prompts input by a user; the prompts include at least a user question to be answered and a document file corresponding to the user question; determining a first matching score between the user question and each paragraph in the document file using a first method; the first method includes: determining the first matching score based on keyword matching; determining a second matching score between the user question and each paragraph in the document file using a second method; the second method includes: determining the second matching score based on question-answer pair matching; determining a third matching score between the user question and each paragraph in the document file using a third method; the third method includes: determining the third matching score based on expanding the user question; and determining the target answer corresponding to the user question based on the first matching score, second matching score, and third matching score of each paragraph.
[0094] According to another aspect of the embodiments of this application, a computer program product is also provided, including a computer program that, when executed by a processor, implements the above-described document response method. For example, it acquires prompt words input by a user; the prompt words include at least a user question to be answered and a document file corresponding to the user question; it determines a first matching score between the user question and each paragraph in the document file using a first method; the first method includes: determining the first matching score based on keyword matching; it determines a second matching score between the user question and each paragraph in the document file using a second method; the second method includes: determining the second matching score based on question-answer pair matching; it determines a third matching score between the user question and each paragraph in the document file using a third method; the third method includes: determining the third matching score based on expanding the user question; and it determines the target answer corresponding to the user question based on the first matching score, second matching score, and third matching score of each paragraph.
[0095] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0096] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For instance, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.
[0097] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0098] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0099] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to related technologies, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.
[0100] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0101] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. A document response method, comprising: Get the suggestion words entered by the user; The prompt words include at least the user question to be answered and the document file corresponding to the user question; The first method is used to determine the first matching score between the user question and each paragraph in the document file; The first method includes: determining the first matching score based on keyword matching; A second method is used to determine a second matching score between the user question and each paragraph in the document file; the second method includes: determining the second matching score based on question-and-answer pair matching. A third method is used to determine the third match score between the user question and each paragraph in the document file; the third method includes: determining the third match score based on an expansion of the user question; The target answer to the user question is determined based on the first matching score, the second matching score, and the third matching score for each paragraph.
2. The method of claim 1, wherein, The method of determining the first matching score based on keyword matching includes: Extract a first preset number of question keywords from the user's questions to obtain a question keyword set; For each paragraph in the document file, extract a second preset number of paragraph keywords to obtain a set of paragraph keywords for each paragraph; The first matching score for each paragraph is determined based on the set of keywords for the question and the set of keywords for each paragraph.
3. The method of claim 2, wherein, The step of determining the first matching score for each paragraph based on the question keyword set and the paragraph keyword set for each paragraph includes: The intersection of the question keyword set and the paragraph keyword set of each paragraph is determined as the matching set for each paragraph; The keyword matching score is determined based on the number of matching keywords in the matching set, the total number of question keywords in the question keyword set, and the total number of paragraph keywords in the paragraph keyword set. Determine the total number of matching characters corresponding to all keywords in the matching set; repeating characters are recorded only once in the total number of matching characters. The character matching score is determined by the total number of matched characters, the first total number of characters, and the second total number of characters; the first total number of characters is the total number of characters corresponding to all question keywords in the question keyword set, and the second total number of characters is the total number of characters corresponding to all paragraph keywords in the paragraph keyword set of each paragraph; repeated characters are recorded only once in the first total number of characters and the second total number of characters. The first matching score is determined by the keyword matching score and the character matching score of each paragraph.
4. The method of claim 1, wherein, The method of determining the second matching score based on question-and-answer pair matching includes: Extract the summary information from the document file; The target audience for the document file is determined based on the summary information; Based on the summary information and the target audience, a preset number of question-and-answer pairs are determined for each paragraph; each question-and-answer pair corresponds to one reference question and one reference answer. Cluster the predetermined number of question-answer pairs for each paragraph to obtain a target question-answer pair set, and store the target question-answer pair set corresponding to each paragraph into a vector library; Determine the similarity score between the user question and the reference questions in the target question-answer pair set corresponding to each paragraph in the vector library, and determine the maximum similarity score of each paragraph as the second matching score of each paragraph.
5. The method of claim 4, wherein, The step of determining a preset number of question-answer pairs for each paragraph based on the summary information and the target audience includes: Identify entities, actions of entities, and relationships between entities in the summary information, and determine the topics and context covered by each paragraph based on the entities, actions of entities, and relationships between entities; Based on the attributes of the target audience, an initial list of question-and-answer pairs is determined, corresponding to the topics and contexts covered by each paragraph; the attributes include at least professional background, interests and preferences, and common usage contexts. The initial question-and-answer pair list corresponding to each paragraph is semantically expanded to obtain a preset number of question-and-answer pairs corresponding to each paragraph.
6. The method of claim 1, wherein, The method of determining the third matching score based on expanding the user question includes: Extract paragraph summary information for each paragraph in the document file; The user question is expanded to obtain a third preset number of similar questions, and the user question and the similar questions are determined as a question set. Determine the semantic similarity between the questions in the question set and the paragraph summary information of each paragraph, and determine the maximum semantic similarity as the third matching score of each paragraph.
7. The method of claim 6, wherein, The step of expanding the user question to obtain a third preset number of similar questions, and determining the user question and the similar questions as a question set, includes: Semantic parsing is performed on the user's question to extract the core intent; Based on the core intent and the question keywords, a third preset number of similar questions are determined; the similar questions have the same core intent as the user question, and the similar questions cover different ways of expressing the user question and different angles of inquiry; The similar questions are subjected to syntax validation, and the similar questions after syntax validation are combined with the user questions to form the question set.
8. The method of claim 1, wherein, Determining the target answer corresponding to the user question based on the first matching score, the second matching score, and the third matching score of each paragraph includes: Based on the first matching score, the second matching score, and the third matching score, the target paragraph in the document file is determined; Based on the target paragraph, generate the target answer corresponding to the user question.
9. The method of claim 8, wherein, The step of determining the target paragraph in the document file based on each of the first matching score, the second matching score, and the third matching score includes: A comprehensive score for each paragraph is determined based on the first matching score, the second matching score, and the third matching score for each paragraph. The target paragraphs are selected from the document file based on the overall score of each paragraph.
10. The method of claim 8, wherein, The step of generating the target answer corresponding to the user question based on the target paragraph includes: The target paragraph is processed based on a pre-trained large language model, and the target answer corresponding to the user's question is output.
11. A document response device, comprising: The acquisition module is used to acquire the prompt words entered by the user; The prompt words include at least the user question to be answered and the document file corresponding to the user question; The keyword matching determination module is used to determine the first matching score between the user question and each paragraph in the document file using a first method; The first method includes: determining the first matching score based on keyword matching; The question-and-answer pair matching determination module is used to determine a second matching score between the user question and each paragraph in the document file using a second method; the second method includes: determining the second matching score based on question-and-answer pair matching. The user question expansion and determination module is used to determine a third matching score between the user question and each paragraph in the document file using a third method; the third method includes: determining the third matching score based on the method of expanding the user question; The target answer determination module is used to determine the target answer corresponding to the user question based on the first matching score, the second matching score, and the third matching score of each paragraph.
12. The apparatus of claim 11, wherein, The keyword matching and determination module is also used for: Extract a first preset number of question keywords from the user's questions to obtain a question keyword set; For each paragraph in the document file, extract a second preset number of paragraph keywords to obtain a set of paragraph keywords for each paragraph; The first matching score for each paragraph is determined based on the set of keywords for the question and the set of keywords for each paragraph.
13. A non-transitory storage medium having stored therein a program, wherein, When the program is running, it controls the device containing the non-volatile storage medium to perform the steps of the document response method according to any one of claims 1 to 10.
14. An electronic device comprising: A memory and a processor, the processor being configured to run a program stored in the memory, wherein the program, when running, performs the steps of the document response method according to any one of claims 1 to 10.
15. A computer program product comprising computer instructions that, when executed by a processor, implement the steps of the document response method of any one of claims 1 to 10.