Text Question Answering Method, System, Electronic Device and Medium Based on Semantic Retrieval
By building a hierarchical knowledge graph and selecting a search strategy based on the type of problem, the existing system's inaccurate information and high cost in global problems are solved, and efficient and accurate text Q&A is achieved.
Patent Information
- Application Number
- CN202510511942.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-23
AI Technical Summary
When dealing with global problems, the existing text question-and-answer system based on large language models is not accurate enough and has high computational cost, so it cannot effectively utilize the potential of hierarchical structures.
Build a hierarchical knowledge graph with highly cohesive communities, and use different search strategies to search information by judging that the problem type is a local problem or a global problem. For local problems, use the local search module to directly calculate the similarity; for global problems, use the global search module to search in the hierarchical knowledge graph for community, and describe the community as a context input to the large language model to generate solution text.
It improves the accuracy and efficiency of searching information of global questions, reduces calculation costs, is suitable for a wide range of knowledge-intensive Q&A tasks, and improves the accuracy and efficiency of text Q&A.
Smart Images

Figure CN120045685B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of natural language processing technology, and in particular, to a text question-answering method, system, electronic device, and medium based on semantic retrieval. Background Art
[0002] Knowledge-based Question Answering (KBQA) is an important research direction in the field of Natural Language Processing (NLP), aiming to answer users' questions by extracting information from structured knowledge bases or unstructured texts. With the rapid development of large language models (LLMs) such as GPT, LLaMA, and Gemini, these models have demonstrated powerful few-shot learning capabilities and can generate accurate answers based on a small amount of context information. However, LLMs still face some challenges in practical applications, especially their limited context windows, which restrict their application in a wider range of scenarios.
[0003] To overcome these limitations, Retrieval-Augmented Generation (RAG) technology has emerged. RAG combines an external knowledge base with LLMs, enabling the model to access and utilize knowledge beyond its context window when generating answers without the need for additional fine-tuning. Traditional RAG methods assume that the required information is concentrated in a specific area of the knowledge base and retrieve relevant information by selecting the top K text chunks with the highest semantic similarity to the question. This method is effective for local questions that only need to focus on a specific part of the knowledge base, but for global questions that require integrating information from multiple parts of the knowledge base, such as Query-Focused Summarization (QFS), the retrieved information is not accurate enough. Summary of the Invention
[0004] This application aims to propose a text question-answering method, system, electronic device, and medium based on semantic retrieval, which can improve the accuracy and retrieval efficiency of the retrieved information corresponding to global questions, thereby improving the precision and efficiency of text question-answering.
[0005] In a first aspect, an embodiment of this application provides a text question-answering method based on semantic retrieval, and the method includes:
[0006] Obtain the question to be answered;
[0007] Construct a hierarchical knowledge graph with highly cohesive communities;
[0008] Determine whether the problem to be solved belongs to a local problem or a global problem;
[0009] If the problem to be solved belongs to a local problem, perform a local search through the local retrieval module to obtain multiple answer texts similar to the problem to be solved;
[0010] If the problem to be solved belongs to a global problem, perform a global search in the hierarchical knowledge graph through the global retrieval module to obtain multiple communities related to the problem to be solved, and describe the communities as context and input them into the large language model to generate answer texts.
[0011] Compared with the prior art, the first aspect of the present application has the following beneficial effects:
[0012] This method obtains the problem to be solved; constructs a hierarchical knowledge graph with highly cohesive communities; determines whether the problem to be solved belongs to a local problem or a global problem; if the problem to be solved belongs to a local problem, perform a local search through the local retrieval module to obtain multiple answer texts similar to the problem to be solved; if the problem to be solved belongs to a global problem, perform a global search in the hierarchical knowledge graph through the global retrieval module to obtain multiple communities related to the problem to be solved, and describe the communities as context and input them into the large language model to generate answer texts. In this way, by determining whether the problem to be solved belongs to a local problem or a global problem, different methods are used to search for answer texts for different types of problems. Especially, by performing a global search in the hierarchical knowledge graph through the global retrieval module, it can be highly efficient in terms of time and computational cost, can significantly improve the processing ability for complex problems without significantly increasing the system complexity, can improve the accuracy and retrieval efficiency of the retrieved information corresponding to global problems, and thus improve the accuracy and efficiency of text Q&A.
[0013] In some embodiments, the determining whether the problem to be solved belongs to a local problem or a global problem includes:
[0014] Construct a similarity distribution between the problem to be solved and multiple text blocks;
[0015] Calculate the ratio between the peak value and the secondary peak value of the similarity distribution to obtain the peak intensity;
[0016] Calculate the second derivative of the similarity distribution at the peak point to obtain the distribution sharpness;
[0017] Based on the peak intensity and the distribution sharpness, determine whether the problem to be solved belongs to a local problem or a global problem.
[0018] In some embodiments, constructing the similarity distribution between the question to be answered and multiple text blocks includes:
[0019] ;
[0020] wherein, represents the similarity distribution, represents the number of text blocks, and is a positive integer, represents the bandwidth parameter, represents the similarity score between the question to be answered and any text block, represents the similarity score between the question to be answered and the th text block, represents the kernel function.
[0021] In some embodiments, calculating the second derivative of the similarity distribution at the peak point to obtain the sharpness of the distribution includes:
[0022] ;
[0023] ;
[0024] wherein, represents the sharpness of the distribution, represents the similarity distribution, represents the peak point, represents the first derivative of the similarity distribution at the peak point, represents the second derivative of the similarity distribution at the peak point, represents the similarity score between the question to be answered and the th text block, represents the similarity score between the question to be answered and the th text block.
[0025] In some embodiments, based on the peak intensity and the sharpness of the distribution, determining whether the question to be answered belongs to a local problem or a global problem includes:
[0026] Preset a first threshold and a second threshold;
[0027] If the peak intensity is greater than or equal to the first threshold and the sharpness of the distribution is greater than or equal to the second threshold, then determine that the question to be answered belongs to a local problem;
[0028] If the peak intensity is less than the first threshold or the sharpness of the distribution is less than the second threshold, then determine that the question to be answered belongs to a global problem.
[0029] In some embodiments, the global search is performed by the global retrieval module in the hierarchical knowledge graph to obtain a plurality of communities related to the question to be answered, including:
[0030] In the hierarchical knowledge graph, each node corresponds to a community;
[0031] Calculate the similarity between the question to be answered and each of the nodes, and use the similarity as the initial weight of each of the nodes;
[0032] Calculate the weighted community weights according to the initial weights;
[0033] Select a plurality of communities related to the question to be answered based on the weighted community weights.
[0034] In some embodiments, the calculating the weighted community weights according to the initial weights includes:
[0035] ;
[0036] Wherein, represents the weighted community weight, represents a coefficient for balancing the direct similarity and the inheritance weight, represents the weight of the parent node, represents the weight of the child node, represents a coefficient for adjusting the contribution of each child node to the weight of the parent node, represents the number of child nodes.
[0037] In a second aspect, an embodiment of the present application further provides a text question-answering system based on semantic retrieval, the system includes:
[0038] A data acquisition unit, configured to acquire a question to be answered;
[0039] A graph construction unit, configured to construct a hierarchical knowledge graph with highly cohesive communities;
[0040] A question judgment unit, configured to judge whether the question to be answered belongs to a local question or a global question;
[0041] A local retrieval unit, configured to, if the question to be answered belongs to a local question, perform a local search through a local retrieval module to obtain a plurality of answer texts similar to the question to be answered;
[0042] A global retrieval unit, which is configured to, if the question to be answered belongs to a global question, perform a global search in the hierarchical knowledge graph through a global retrieval module to obtain multiple communities related to the question to be answered, and describe the communities as context and input them into a large language model to generate an answer text.
[0043] In a third aspect, an embodiment of the present application further provides an electronic device, including at least one control processor and a memory communicatively connected to the at least one control processor; the memory stores instructions executable by the at least one control processor, and the instructions are executed by the at least one control processor to enable the at least one control processor to execute a text question-answering method based on semantic retrieval as described above.
[0044] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, which stores computer-executable instructions for causing a computer to execute a text question-answering method based on semantic retrieval as described above.
[0045] It can be understood that the beneficial effects of the above second aspect to the fourth aspect compared with the related art are the same as those of the above first aspect compared with the related art. For relevant descriptions, reference can be made to the relevant descriptions in the above first aspect, and details will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the following description of the embodiments in conjunction with the accompanying drawings, where:
[0047] Figure 1 is a schematic flowchart of an embodiment of a text question-answering method based on semantic retrieval provided by the present application;
[0048] Figure 2 is a schematic overall method flowchart of the best embodiment of a text question-answering method based on semantic retrieval provided by the present application;
[0049] Figure 3 is a schematic structural diagram of an embodiment of a text question-answering system based on semantic retrieval provided by the present application;
[0050] Figure 4 is a schematic structural diagram of an embodiment of an electronic device provided by the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0051] Embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the drawings are exemplary and are only used to explain the present application and should not be construed as a limitation of the present application.
[0052] In the description of the present application, if the first, second, etc. are described, it is only for the purpose of distinguishing technical features and should not be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence relationship of the indicated technical features.
[0053] It should also be understood that the reference to "one embodiment" or "some embodiments" etc. in the description of the embodiments of the present application means that a specific feature, structure, or characteristic described in connection with the embodiment is included in one or more embodiments of the embodiments of the present application. Thus, the statements "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprise", "include", "have" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0054] In the description of the present application, it should be noted that unless otherwise clearly defined, words such as "set", "installed", "connected", etc. should be understood in a broad sense, and those skilled in the art can reasonably determine the specific meanings of the above words in the present application in combination with the specific content of the technical solution.
[0055] First, several terms involved in the present application are analyzed:
[0056] Retrieval-Augmented Generation technology: This technology was proposed by the Facebook Research Institute. This method relies on semantic similarity retrieval. By calculating the similarity between the embedding vectors of the question and the text chunks in the knowledge base, the most relevant TopK text chunks are selected as the basis for answer generation. This method performs well in dealing with local problems, but when dealing with global problems that require integrating multiple knowledge fragments, it often cannot provide sufficient context information.
[0057] Large language model: It refers to a deep learning model trained using a large amount of text data, enabling the model to generate natural language text or understand the meaning of language text. These models can provide in-depth knowledge and language production on various topics by training on a large dataset. Its core idea is to learn the patterns and structures of natural language through large-scale unsupervised training, and to simulate the human language cognition and generation process to a certain extent.
[0058] GraphRAG: A new retrieval method that uses a knowledge graph and vector search in the basic RAG architecture. Thus, it can integrate and understand various knowledge, providing a broader and more comprehensive view of the data. Proposed by Microsoft Research, this method further expands the application of graph structures, constructing a knowledge graph with communities. GraphRAG relies on LLMs to generate summaries of community content sequentially when retrieving knowledge. Although it can handle global problems, it has a high cost, and the final retrieval result still "unfolds" the hierarchical structure into a linear structure, failing to fully utilize the potential of the hierarchical structure.
[0059] Natural Language Processing: An important research direction in the field of artificial intelligence, integrating knowledge from multiple disciplinary fields such as linguistics, computer science, machine learning, mathematics, and cognitive psychology. It is an interdisciplinary subject integrating computer science, artificial intelligence, and linguistics. It includes two main aspects: natural language understanding and natural language generation. The research content includes various levels such as characters, words, phrases, sentences, paragraphs, and texts, serving as a bridge between machine language and human language.
[0060] Knowledge-based Question Answering (KBQA) is an important research direction in the field of Natural Language Processing (NLP), aiming to answer users' questions by extracting information from structured knowledge bases or unstructured texts. With the rapid development of large language models (LLMs) such as GPT, LLaMA, and Gemini, these models have demonstrated powerful few-shot learning capabilities and can generate accurate answers based on a small amount of context information. However, LLMs still face some challenges in practical applications, especially their limited context windows, which restrict their application in a wider range of scenarios.
[0061] To overcome these limitations, Retrieval-Augmented Generation (RAG) technology has emerged. By combining an external knowledge base with LLMs, RAG enables the model to access and utilize knowledge beyond its context window when generating answers without the need for additional fine-tuning. Traditional RAG methods assume that the required information is concentrated in a specific area of the knowledge base and retrieve relevant information by selecting the TopK text chunks with the highest semantic similarity to the question. This method is effective for local questions that only need to focus on a specific part of the knowledge base, but for global questions that require integrating information from multiple parts of the knowledge base, such as query-focused summarization (QFS), the retrieved information is not accurate enough.
[0062] To solve the problem that the information retrieved by the above-mentioned existing related technologies for global questions is not accurate enough, this application proposes a text question-answering method, system, electronic device, and medium based on semantic retrieval.
[0063] Refer to Figure 1 , an embodiment of this application provides a text question-answering method based on semantic retrieval, and the method includes the following steps:
[0064] Step S100, obtain the question to be answered;
[0065] Step S200, construct a hierarchical knowledge graph with highly cohesive communities;
[0066] Step S300, determine whether the question to be answered is a local question or a global question;
[0067] Step S400, if the question to be answered is a local question, perform a local search through the local retrieval module to obtain multiple answer texts similar to the question to be answered;
[0068] Step S500, if the question to be answered is a global question, perform a global search in the hierarchical knowledge graph through the global retrieval module to obtain multiple communities related to the question to be answered, and describe the communities as context and input them into the large language model to generate answer texts.
[0069] In this embodiment, by obtaining the question to be answered; constructing a hierarchical knowledge graph with highly cohesive communities; determining whether the question to be answered is a local problem or a global problem; if the question to be answered is a local problem, then performing a local search through the local retrieval module to obtain multiple answer texts similar to the question to be answered; if the question to be answered is a global problem, then performing a global search in the hierarchical knowledge graph through the global retrieval module to obtain multiple communities related to the question to be answered, and describing the communities as context and inputting them into the large language model to generate an answer text. In this way, by determining whether the question to be answered is a local problem or a global problem, different methods are used to search for answer texts for different types of questions. Especially by performing a global search in the hierarchical knowledge graph through the global retrieval module, it can be highly efficient in terms of time and computational cost, can significantly improve the processing ability of complex questions without significantly increasing the system complexity, can improve the accuracy and retrieval efficiency of the retrieved information corresponding to global questions, and thus improve the accuracy and efficiency of text Q&A.
[0070] The above-mentioned construction of a hierarchical knowledge graph with highly cohesive communities can be to merge entities with the same name and type in the sub-graphs, and to merge relationships with the same source entity and target entity, thereby creating cross-sub-graph relationships, and finally generating a hierarchical knowledge graph with highly cohesive communities, which can more accurately reflect the global structure of the source document.
[0071] The above-mentioned local search through the local retrieval module to obtain multiple answer texts similar to the question to be answered can be to directly adopt the most classic retrieval method based on the K-Nearest Neighbors (KNN) algorithm. The local retrieval module calculates the similarity between the question to be answered and multiple texts using a similarity calculation method to obtain multiple answer texts similar to the question to be answered.
[0072] In some embodiments, determining whether the question to be answered is a local problem or a global problem includes:
[0073] Constructing the similarity distribution between the question to be answered and multiple text blocks;
[0074] Calculating the ratio between the peak value of the similarity distribution and the secondary peak value of the similarity distribution to obtain the peak intensity;
[0075] Calculating the second derivative of the similarity distribution at the peak point to obtain the distribution sharpness;
[0076] Based on the peak intensity and the distribution sharpness, determining whether the question to be answered is a local problem or a global problem.
[0077] In this embodiment, by constructing a similarity distribution between the question to be answered and multiple text blocks; calculating the ratio between the peak value of the similarity distribution and the second peak value of the similarity distribution to obtain the peak intensity; calculating the second derivative of the similarity distribution at the peak point to obtain the distribution sharpness; and based on the peak intensity and the distribution sharpness, determining whether the question to be answered belongs to a local problem or a global problem. In this way, by calculating the semantic similarity between the question to be answered and each text block, the similarity distribution between the question to be answered and multiple text blocks is constructed. Since there are significant differences in the similarity distributions of local problems and global problems, by analyzing the peak intensity and the distribution sharpness, it is possible to accurately determine whether the question to be answered belongs to a local problem or a global problem, laying a good data foundation for accurately retrieving information in the later stage.
[0078] In some embodiments, constructing a similarity distribution between the question to be answered and multiple text blocks includes:
[0079] ;
[0080] wherein, represents the similarity distribution, represents the number of text blocks, and is a positive integer, represents the bandwidth parameter, represents the similarity score between the question to be answered and any one text block, represents the similarity score between the question to be answered and the th text block, represents the kernel function.
[0081] In some embodiments, calculating the second derivative of the similarity distribution at the peak point to obtain the distribution sharpness includes:
[0082] ;
[0083] ;
[0084] wherein, represents the distribution sharpness, represents the similarity distribution, represents the peak point, represents the first derivative of the similarity distribution at the peak point, represents the second derivative of the similarity distribution at the peak point, represents the similarity score between the question to be answered and the th text block, represents the similarity score between the question to be answered and the th text block.
[0085] In some embodiments, based on the peak intensity and the sharpness of the distribution, determining whether the question to be answered belongs to a local problem or a global problem includes:
[0086] Presetting a first threshold and a second threshold;
[0087] If the peak intensity is greater than or equal to the first threshold and the sharpness of the distribution is greater than or equal to the second threshold, then determine that the question to be answered belongs to a local problem;
[0088] If the peak intensity is less than the first threshold or the sharpness of the distribution is less than the second threshold, then determine that the question to be answered belongs to a global problem.
[0089] In this embodiment, preset the first threshold and the second threshold; if the peak intensity is greater than or equal to the first threshold and the sharpness of the distribution is greater than or equal to the second threshold, then determine that the question to be answered belongs to a local problem; if the peak intensity is less than the first threshold or the sharpness of the distribution is less than the second threshold, then determine that the question to be answered belongs to a global problem. In this way, by comprehensively considering the peak intensity and the sharpness of the distribution to determine whether the question to be answered belongs to a local problem or a global problem, it is possible to accurately determine whether the question to be answered belongs to a local problem or a global problem, laying a good data foundation for searching for accurate retrieval information later.
[0090] In some embodiments, performing a global search in the hierarchical knowledge graph through a global search module to obtain multiple communities related to the question to be answered, including:
[0091] In the hierarchical knowledge graph, each node corresponds to a community;
[0092] Calculating the similarity between the question to be answered and each node, and using the similarity as the initial weight of each node;
[0093] Calculating the weighted community weight according to the initial weight;
[0094] Selecting multiple communities related to the question to be answered based on the weighted community weight.
[0095] In this embodiment, in the hierarchical knowledge graph, each node corresponds to a community; calculating the similarity between the question to be answered and each node, and using the similarity as the initial weight of each node; calculating the weighted community weight according to the initial weight; selecting multiple communities related to the question to be answered based on the weighted community weight. In this way, by performing a global search in the hierarchical knowledge graph through the global search module, it is possible to have high efficiency in terms of time and computational cost, and to significantly improve the processing ability of complex problems without significantly increasing the system complexity, thereby improving the accuracy and retrieval efficiency of the retrieval information corresponding to global problems.
[0096] In some embodiments, calculating the weighted community weight according to the initial weight value includes:
[0097] ;
[0098] wherein, represents the weighted community weight, represents the coefficient for balancing the direct similarity and the inheritance weight, represents the weight of the parent node, represents the weight of the child node, represents the coefficient for adjusting the contribution of each child node to the weight of the parent node, represents the number of child nodes.
[0099] For the convenience of those skilled in the art to understand, a set of best embodiments are provided below:
[0100] Knowledge-based Question Answering (KBQA) is an important research direction in the field of Natural Language Processing (NLP), aiming to answer users' questions by extracting information from structured knowledge bases or unstructured texts. With the rapid development of large language models (LLMs) such as GPT, LLaMA, and Gemini, these models have demonstrated powerful few-shot learning capabilities and can generate accurate answers based on a small amount of context information. However, LLMs still face some challenges in practical applications, especially their limited context window, which restricts their application in a wider range of scenarios.
[0101] To overcome these limitations, Retrieval-Augmented Generation (RAG) technology has emerged. RAG combines an external knowledge base with LLMs, enabling the model to access and utilize knowledge beyond its context window when generating answers without additional fine-tuning. Traditional RAG methods assume that the required information is concentrated in a specific area of the knowledge base and retrieve relevant information by selecting the TopK text chunks with the highest semantic similarity to the question. This method is effective for local questions that only need to focus on a specific part of the knowledge base, but performs poorly for global questions that require integrating information from multiple parts of the knowledge base, such as Query-Focused Summarization (QFS). For example, when using existing RAG methods to handle local or global questions, the following limitations still exist:
[0102] 1. Imbalance in handling local and global problems: Existing methods often can only retrieve local or global information, lacking a general and efficient information retrieval paradigm to handle these two types of problems simultaneously.
[0103] 2. High computational cost: Some methods (such as GraphRAG) rely on LLMs for multi-round summary generation when retrieving information, resulting in high Q&A costs and long response times.
[0104] 3. Overgeneralization of information: During the hierarchical retrieval process, existing methods often "unfold" the hierarchical structure into a linear structure, leading to the inability to distinguish high-level semantic information or overgeneralization and unable to fully utilize the advantages of the hierarchical structure.
[0105] To address the limitations of the above existing technologies, this embodiment proposes a simple, general, and efficient RAG retrieval paradigm that can uniformly retrieve the information required for local and global problems. Compared with the existing technologies, the main differences in the technical solutions of this embodiment are as follows:
[0106] 1. Problem classification algorithm: This embodiment proposes a problem classification algorithm based on semantic similarity distribution, which can accurately judge the problem type without relying on LLMs.
[0107] 2. Hierarchical graph structure: This embodiment introduces a hierarchical graph structure to represent the original knowledge base, which can better capture the global and local relationships of knowledge.
[0108] 3. Hierarchical retrieval strategy: This embodiment proposes a hierarchical retrieval strategy that can minimize information loss in terms of time and cost efficiency and is applicable to both local and global problems.
[0109] Through the above innovations, this embodiment can retrieve more accurate and comprehensive information while maintaining a low computational cost, and is applicable to a wide range of knowledge-intensive Q&A tasks. Referring to Figure 2 , the technical solutions of this embodiment specifically include the following contents:
[0110] I. Knowledge graph construction module.
[0111] The method of this embodiment first improves the knowledge graph construction technology of GraphRAG. Similar to GraphRAG, this embodiment extracts entities and relationships from source documents. Each entity contains a name, a description, and a type (from a predefined set of types), while each relationship includes a source entity, a target entity, a brief description, and a score representing the strength of the relationship. With these components, this embodiment constructs a weighted undirected knowledge graph to capture the structural semantics of the source document. This embodiment also detects hierarchical communities in the graph and generates a summary for each community based on the descriptions of the entities and relationships within the community. However, this embodiment introduces a key improvement to address an important limitation of GraphRAG in practical applications.
[0112] In an actual RAG scenario, the source document usually exceeds the context window limit of the LLM, so the document needs to be chunked. This results in the generation of multiple sub-graphs, each representing only a part of the document. Without integrating these sub-graphs, there will be a lack of global consistency between them, leading to fragmentation of the community representation. To address this issue, this embodiment proposes a sub-graph integration method. First, this embodiment merges entities with the same name and type in different sub-graphs and uses the LLM to generate a unified description for each merged entity to ensure consistency and integrity. Second, for relationships with the same source entity and target entity, this embodiment also merges them and uses the LLM to generate the integrated description and strength score. Through this integration process, this embodiment can create relationships across sub-graphs and finally generate a hierarchical knowledge graph with highly cohesive communities, thus more accurately reflecting the global structure of the source document.
[0113] II. Gating module.
[0114] The core function of the gating module is to determine whether the current question (i.e., the question to be answered, which is also the user's question) is a local question or a global question, so as to guide the selection of the subsequent retrieval module. Its implementation is based on the following principle:
[0115] This embodiment first divides the original text into several text chunks and calculates the semantic similarity between the question and each text chunk. For local questions, their answers are usually concentrated in a specific area of the text, so there will be a significant peak in the similarity distribution, indicating that the question is highly relevant to a certain text chunk. For global questions, the required information is often scattered in multiple parts of the text, so the similarity distribution will be relatively smooth without an obvious peak.
[0116] By analyzing the morphological features of the similarity distribution, the gating module can effectively distinguish the types of questions: if there is a significant peak in the similarity distribution, it is determined as a local problem; if the similarity distribution is relatively smooth, it is determined as a global problem. This mechanism provides a reliable basis for the selection of the subsequent retrieval module, thus optimizing the overall performance of the question-answering system. The specific theoretical proof is as follows:
[0117] In this embodiment, kernel density estimation (KDE) is first used to model the similarity distribution, so as to more accurately describe the distribution characteristics of local and global problems. Given a question and text blocks with semantic similarity scores of , the similarity distribution is modeled as:
[0118] (1);
[0119] where represents the similarity distribution, represents the kernel function, satisfying , represents the bandwidth parameter, controlling the smoothness degree, represents the similarity score between the question and any text block, represents the similarity score between the question to be answered and the th text block. Based on the similarity distribution obtained from the above kernel density estimation, two statistical quantities of peak intensity and distribution curvature are defined as follows:
[0120] 1. Peak Strength: Used to calculate the ratio of the peak to the sub-peak of the similarity distribution . The larger this ratio is, the more concentrated the distribution is.
[0121] (2);
[0122] where represents the peak of the similarity distribution , represents the sub-peak of the similarity distribution .
[0123] 2. Distribution Curvature: Calculate the second derivative of the similarity distribution at the peak point to quantify the sharpness of the distribution:
[0124] (3);
[0125] Next, the central difference method is used to approximate the second derivative:
[0126] (4);
[0127] Wherein, represents the first derivative of the similarity distribution at the peak point, represents the second derivative of the similarity distribution at the peak point, represents the similarity score between the question to be answered and the th text block, represents the similarity score between the question to be answered and the th text block. In this embodiment, it is further assumed that the distribution characteristics of the two types of questions satisfy: local question: and ; global question: or . Wherein the first threshold and the second threshold are preset thresholds determined by statistical learning.
[0128] The following proves the theorem:
[0129] If the answers to local questions are concentrated in a single text block and the information of global questions is evenly distributed, then through and the two types of questions can be effectively distinguished.
[0130] Proof:
[0131] If it is a local question, there exists a unique such that is significantly higher than other , making the KDE distribution present a sharp peak at : (the denominator approaches 0) and (the curvature is large and the distribution is sharp); if it is a global question, there exists evenly distributed, making the KDE distribution close to flat: (no obvious peak) and ( (the curvature is small and the distribution is smooth).
[0132] Conclusion: There are significant differences in and between the two types of questions. Through threshold division, it can be ensured that:
[0133] (5);
[0134] Wherein, represents the probability of correctly classifying whether the question is a global question or a local question.
[0135] Supplementary proof:
[0136] 1. Bandwidth selection.
[0137] Bandwidth selection is crucial for KDE modeling. The method of this embodiment is optimized by the Silverman criterion:
[0138] (6);
[0139] where is the sample standard deviation.
[0140] 2. Noise robustness.
[0141] In the actual scenario, the similarity distribution may contain noise. The signal-to-noise ratio is defined as follows:
[0142] (7);
[0143] where represents the highest semantic similarity score between the question and all text blocks , that is, the signal strength. It reflects the matching degree between the question and the most relevant text block. The larger the value, the stronger the relevance between the question and a certain text block. represents the standard deviation of the semantic similarity scores, which is used to measure the dispersion degree of the similarity scores, that is, the noise strength. It is obtained by calculating the square root of the average of the sum of the squared deviations of each similarity score from the mean . The larger is, the more dispersed the similarity distribution is and the more noise there is; the smaller is, the more concentrated the similarity distribution is and the less noise there is. It can be calculated that for the method of this embodiment , the influence of noise on the threshold determination can be ignored.
[0144] III. Local retrieval module.
[0145] The core task of this module is to quickly locate the local information most relevant to the question from the knowledge base to support subsequent answer generation. Here, the most classic retrieval method based on the K-Nearest Neighbors (KNN) algorithm is directly adopted. The core idea of the KNN algorithm is to select the most relevant text blocks as the retrieval results by calculating the semantic similarity between the question and the text blocks in the knowledge base. Specifically, given the question and the text block in the knowledge base, first, the question and text chunks are respectively encoded into semantic vectors through a pre-trained Embedding generation model and , then calculate the cosine similarity between the question vector and each text chunk vector:
[0146] (8);
[0147] Finally, select the text chunk with the highest similarity as the retrieval result. The KNN algorithm can match the question and text fragments at a fine-grained level, so as to more accurately locate relevant information. This method is simple and efficient, and is suitable for local questions where the answers are concentrated in a local area of the knowledge base.
[0148] IV. Global Retrieval Module.
[0149] The core task of the global retrieval module is to integrate multi-part information from the hierarchical knowledge graph to answer global questions that require comprehensive knowledge. Traditional retrieval methods based on cosine similarity have limitations in dealing with global questions because the similarity scores of high-level nodes (representing abstract concepts) are usually low, which may cause them to be ignored, and relying solely on parent-child relationships may also miss relevant information. To solve this problem, this embodiment proposes a weighted retrieval mechanism, which combines the position information of nodes in the hierarchical structure on the basis of cosine similarity, so as to more comprehensively capture the requirements of global questions. In the hierarchical tree structure constructed by Graph Indexer, each node corresponds to a community. This embodiment first calculates the initial weight of each node based on cosine similarity:
[0150] (9);
[0151] where, is the embedding vector of the question, is the embedding vector of the node description. To incorporate hierarchical context information, this embodiment updates the weight of each parent node. Among them, is the weight of the parent node, is the weight of the child node, is the number of child nodes, is used to adjust the contribution of each child node to the weight of the parent node, (usually set to 0.3) is used to balance the direct similarity and the inherited weight.
[0152] (10);
[0153] Based on the weighted community weights, this embodiment selects the communities with the highest weights (in the experiment ), and use its description as the context to input into the LLM to generate answers. Compared with directly using the original description, the sub-answers can reduce noise and optimize the utilization efficiency of the context window. These sub-answers, as retrieved information, provide high-quality knowledge support for subsequent answer generation.
[0154] Compared with the prior art, the technical solution of this embodiment has the following advantages:
[0155] The retrieval method proposed in this embodiment significantly improves the performance of the question-answering system in dealing with local and global problems by combining a hierarchical knowledge graph and a weighted retrieval mechanism. For local problems, the traditional KNN retrieval method can quickly locate the text block most relevant to the problem, ensuring the accuracy and retrieval efficiency of the answer; while for global problems, the hierarchical weighted retrieval mechanism integrates the abstract information of high-level nodes and the detailed information of low-level nodes, avoiding the limitations of a single similarity metric, and thus capturing more comprehensively the key information scattered in multiple parts of the knowledge base. In addition, by generating sub-answers instead of directly using the original description, noise is further reduced and the utilization rate of the context window is optimized.
[0156] The core advantages of the method of this embodiment lie in its generality and high efficiency. Through simple similarity distribution analysis, the gating module can automatically judge the question type and flexibly select local or global retrieval strategies without relying on a complex large language model (LLM) proxy. At the same time, the hierarchical weighted retrieval mechanism is highly efficient in terms of time and computational cost, and can significantly improve the processing ability of complex problems without significantly increasing the system complexity. Experimental results show that the method of this embodiment outperforms existing state-of-the-art methods in various knowledge-intensive question-answering tasks, providing reliable technical support for practical applications.
[0157] Referring to Figure 3 , the embodiment of the present application further provides a text question-answering system based on semantic retrieval, which includes a data acquisition unit 100, a graph construction unit 200, a question judgment unit 300, a local retrieval unit 400, and a global retrieval unit 500, where:
[0158] The data acquisition unit 100 is used to acquire the question to be answered;
[0159] The graph construction unit 200 is used to construct a hierarchical knowledge graph with highly cohesive communities;
[0160] The question judgment unit 300 is used to judge whether the question to be answered belongs to a local problem or a global problem;
[0161] The local retrieval unit 400 is used to, if the question to be answered belongs to a local problem, perform local search through the local retrieval module to obtain multiple answer texts similar to the question to be answered;
[0162] The global retrieval unit 500 is used to, if the question to be answered belongs to a global question, perform a global search in the hierarchical knowledge graph through the global retrieval module to obtain multiple communities related to the question to be answered, and describe the communities as context and input them into the large language model to generate an answer text.
[0163] In some embodiments, the question judgment unit 300 may specifically be used for:
[0164] Construct a similarity distribution between the question to be answered and multiple text blocks;
[0165] Calculate the ratio between the peak value of the similarity distribution and the sub-peak value of the similarity distribution to obtain the peak intensity;
[0166] Calculate the second derivative of the similarity distribution at the peak point to obtain the distribution sharpness;
[0167] Based on the peak intensity and the distribution sharpness, determine whether the question to be answered belongs to a local question or a global question.
[0168] In some embodiments, the question judgment unit 300 may specifically be used for:
[0169] ;
[0170] Wherein, represents the similarity distribution, represents the number of text blocks, and is a positive integer, represents the bandwidth parameter, represents the similarity score between the question to be answered and any text block, represents the similarity score between the question to be answered and the th text block, represents the kernel function.
[0171] In some embodiments, the question judgment unit 300 may specifically be used for:
[0172] ;
[0173] ;
[0174] Wherein, represents the distribution sharpness, represents the similarity distribution, represents the peak point, represents the second derivative of the similarity distribution at the peak point, represents the similarity score between the question to be answered and the th text block, Indicates the similarity score between the question to be answered and the nth text block.
[0175] In some embodiments, the question determination unit 300 may be specifically configured to:
[0176] Preset a first threshold and a second threshold;
[0177] If the peak intensity is greater than or equal to the first threshold and the distribution sharpness is greater than or equal to the second threshold, it is determined that the question to be answered belongs to a local problem;
[0178] If the peak intensity is less than the first threshold or the distribution sharpness is less than the second threshold, it is determined that the question to be answered belongs to a global problem.
[0179] In some embodiments, the global retrieval unit 500 may be specifically configured to:
[0180] In the hierarchical knowledge graph, each node corresponds to a community;
[0181] Calculate the similarity between the question to be answered and each node, and use the similarity as the initial weight of each node;
[0182] Calculate the weighted community weights according to the initial weights;
[0183] Select multiple communities related to the question to be answered based on the weighted community weights.
[0184] In some embodiments, the global retrieval unit 500 may be specifically configured to:
[0185] ;
[0186] Wherein, represents the weighted community weight, represents a coefficient for balancing the direct similarity and the inheritance weight, represents the weight of the parent node, represents the weight of the child node, represents a coefficient for adjusting the contribution of each child node to the weight of the parent node, represents the number of child nodes.
[0187] It should be noted that since the text question-answering system based on semantic retrieval in this embodiment and the above-mentioned text question-answering method based on semantic retrieval are based on the same inventive concept, the corresponding content in the method embodiment also applies to this system embodiment, and will not be elaborated here.
[0188] Referring to Figure 4 , this application embodiment also provides an electronic device, and this electronic device includes:
[0189] At least one memory;
[0190] At least one processor;
[0191] At least one program;
[0192] The program is stored in the memory, and the processor executes at least one program to implement the method for text question answering based on semantic retrieval described above in the present disclosure.
[0193] The electronic device can be any intelligent terminal including a mobile phone, a tablet computer, a personal digital assistant (PDA), an in-vehicle computer, etc.
[0194] The electronic device according to the embodiments of the present application will be described in detail below.
[0195] The processor 1600 can be implemented by using a general-purpose central processing unit (CPU), a microprocessor, an application specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present disclosure;
[0196] The memory 1700 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 1700 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1700 and are called by the processor 1600 to execute the method for text question answering based on semantic retrieval in the embodiments of the present disclosure.
[0197] The input / output interface 1800 is used to implement information input and output;
[0198] The communication interface 1900 is used to implement communication interaction between this device and other devices, and can implement communication through a wired method (such as USB, network cable, etc.) or through a wireless method (such as mobile network, WIFI, Bluetooth, etc.);
[0199] The bus 2000 transmits information between various components of the device (such as the processor 1600, the memory 1700, the input / output interface 1800, and the communication interface 1900);
[0200] Among them, the processor 1600, the memory 1700, the input / output interface 1800, and the communication interface 1900 are communicatively connected to each other inside the device through the bus 2000.
[0201] An embodiment of the present disclosure also provides a storage medium, which is a computer-readable storage medium storing computer-executable instructions for causing a computer to execute the above-described text question-answering method based on semantic retrieval.
[0202] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include memories remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0203] The embodiments described in the embodiments of the present disclosure are for more clearly illustrating the technical solutions of the embodiments of the present disclosure, and do not constitute a limitation on the technical solutions provided by the embodiments of the present disclosure. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present disclosure are equally applicable to similar technical problems.
[0204] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present disclosure, and may include more or fewer steps than those shown in the figures, or combine certain steps, or different steps.
[0205] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0206] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices can be implemented as software, firmware, hardware, and appropriate combinations thereof.
[0207] In the description of this application and the above-mentioned accompanying drawings, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0208] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or a similar expression means any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0209] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.
[0210] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0211] In addition, in each embodiment of the present application, each functional unit can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0212] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes multiple instructions for causing an electronic device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in the various embodiments of the present application. The foregoing storage medium includes: various media that can store programs such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs. The embodiments of the present application have been described in detail above with reference to the accompanying drawings, but the present application is not limited to the above embodiments. Within the knowledge scope of those of ordinary skill in the art to which the present application pertains, various changes can be made without departing from the purpose of the present application.
[0213] The embodiments of the present application have been described in detail above with reference to the accompanying drawings, but the present application is not limited to the above embodiments. Within the knowledge scope of those of ordinary skill in the art to which the present application pertains, various changes can be made without departing from the purpose of the present application.
Claims
1. A text question-answering method based on semantic retrieval, characterized in that, The method includes: Obtain the question to be answered; Construct a hierarchical knowledge graph with highly cohesive communities; Determine whether the question to be answered is a local problem or a global problem, including Construct a similarity distribution between the question to be answered and multiple text blocks; Calculate the ratio between the peak value and the sub-peak value of the similarity distribution to obtain the peak intensity; Calculate the second derivative of the similarity distribution at the peak point to obtain the distribution sharpness; Based on the peak intensity and the distribution sharpness, determine whether the question to be answered is a local problem or a global problem; If the question to be answered is a local problem, perform a local search through the local retrieval module to obtain multiple answer texts similar to the question to be answered; If the question to be answered is a global problem, perform a global search in the hierarchical knowledge graph through the global retrieval module to obtain multiple communities related to the question to be answered, and describe the communities as context and input them into the large language model to generate answer texts.
2. The text question-answering method based on semantic retrieval according to claim 1, wherein The construction of the similarity distribution between the question to be answered and multiple text blocks includes: ; Among them, represents the similarity distribution, represents the number of text blocks, and is a positive integer, represents the bandwidth parameter, represents the similarity score between the question to be answered and any text block, represents the similarity score between the question to be answered and the th text block, represents the kernel function.
3. The text question-answering method based on semantic retrieval according to claim 1, characterized in that The calculation of the second derivative of the similarity distribution at the peak point to obtain the distribution sharpness includes: ; ; Among them, represents the sharpness of the distribution, represents the similarity distribution, represents the peak point, represents the first derivative of the said similarity distribution at the peak point, represents the second derivative of the said similarity distribution at the peak point, represents the similarity score between the said question to be answered and the th text block, represents the similarity score between the said question to be answered and the th text block.
4. The text question-answering method based on semantic retrieval according to claim 1, wherein The determination of whether the question to be answered is a local problem or a global problem based on the peak intensity and the distribution sharpness includes: Preset a first threshold and a second threshold; If the peak intensity is greater than or equal to the first threshold and the distribution sharpness is greater than or equal to the second threshold, then determine that the question to be answered is a local problem; If the peak intensity is less than the first threshold or the distribution sharpness is less than the second threshold, then determine that the question to be answered is a global problem.
5. The text question-answering method based on semantic retrieval according to claim 1, wherein The global search in the hierarchical knowledge graph through the global retrieval module to obtain multiple communities related to the question to be answered includes: In the hierarchical knowledge graph, each node corresponds to a community; Calculate the similarity between the question to be answered and each node, and use the similarity as the initial weight of each node; Calculate the weighted community weights according to the initial weights; Based on the weighted community weights, select multiple communities related to the question to be answered.
6. The text question-answering method based on semantic retrieval according to claim 5, characterized in that The calculation of the weighted community weights according to the initial weights includes: ; Among them, represents the weighted community weight, represents the coefficient used to balance the direct similarity and the inheritance weight, represents the weight of the parent node, represents the weight of the child node, represents the coefficient used to adjust the contribution of each child node to the weight of the parent node, represents the number of child nodes.
7. A text question answering system based on semantic retrieval, characterized in that, The system includes: A data acquisition unit for obtaining the question to be answered; A graph construction unit for constructing a hierarchical knowledge graph with highly cohesive communities; A question judgment unit for determining whether the question to be answered is a local problem or a global problem, including Constructing a similarity distribution between the question to be answered and multiple text blocks; Calculating the ratio between the peak value and the sub-peak value of the similarity distribution to obtain the peak intensity; Calculating the second derivative of the similarity distribution at the peak point to obtain the distribution sharpness; Based on the peak intensity and the distribution sharpness, determine whether the question to be answered is a local problem or a global problem; A local retrieval unit, configured to, if the question to be answered belongs to a local question, perform local search through a local retrieval module to obtain multiple answer texts similar to the question to be answered; A global retrieval unit, configured to, if the question to be answered belongs to a global question, perform global search in the hierarchical knowledge graph through a global retrieval module to obtain multiple communities related to the question to be answered, and describe the communities as context to input into a large language model to generate an answer text.
8. An electronic device, characterized in that, Comprising at least one control processor and a memory for communicatively connecting with the at least one control processor; the memory stores instructions executable by the at least one control processor, and the instructions are executed by the at least one control processor so that the at least one control processor can execute the semantic retrieval-based text question-answering method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to cause a computer to execute the semantic retrieval-based text question-answering method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Man-machine conversation processing method and device
CN111708869A
Intelligent question answering system and method based on question decomposition and community semantic search
CN119537539A