Text question and answer method and system based on semantic retrieval, electronic equipment and medium

By constructing a hierarchical knowledge graph and judging the problem type, and using different search modules to deal with local and global problems, the problem of inaccurate global problem information in the existing technology is solved, and the search efficiency and question-and-answer accuracy are improved.

CN120045685AActive Publication Date: 2025-05-27CENT SOUTH UNIV

Patent Information

Application Number
CN202510511942.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-05-27
Estimated Expiration
2045-04-23

AI Technical Summary

Technical Problem

When handling global problems in the prior art, the information retrieved is not accurate enough, making it difficult to effectively integrate information from multiple parts of the knowledge base.

Method used

A text question-and-answer method based on semantic retrieval is proposed. By constructing a hierarchical knowledge graph with highly cohesive communities, we can determine whether the questions to be answered are local or global questions, and select different search modules for searching according to the type to generate answer text.

Benefits of technology

It improves the accuracy and retrieval efficiency of the search information corresponding to global questions, significantly improves the ability to handle complex problems, and enhances the accuracy and efficiency of text questions and answers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045685A_ABST
    Figure CN120045685A_ABST
Patent Text Reader

Abstract

The invention discloses a text question answering method and system based on semantic retrieval, electronic equipment and a medium. The method comprises the following steps: obtaining a question to be answered; constructing a hierarchical knowledge graph with a highly cohesive community; judging whether the question to be answered belongs to a local question or a global question; if the to-be-answered question belongs to a local question, performing local search through a local retrieval module to obtain a plurality of answering texts similar to the to-be-answered question; and if the to-be-answered question belongs to the global question, performing global search in the hierarchical knowledge graph through a global retrieval module to obtain a plurality of communities related to the to-be-answered question, describing the communities as contexts, inputting the contexts into the large language model, and generating an answering text. According to the method, the accuracy and the retrieval efficiency of the retrieval information corresponding to the global question can be improved, so that the accuracy and the efficiency of text question answering are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of natural language processing technology, and in particular, to a text question-answering method, system, electronic device, and medium based on semantic retrieval. Background Art

[0002] Knowledge-based Question Answering (KBQA) is an important research direction in the field of Natural Language Processing (NLP), aiming to answer users' questions by extracting information from structured knowledge bases or unstructured texts. With the rapid development of large language models (LLMs) such as GPT, LLaMA, and Gemini, these models have demonstrated powerful few-shot learning capabilities and can generate accurate answers based on a small amount of context information. However, LLMs still face some challenges in practical applications, especially their limited context window, which restricts their application in a wider range of scenarios.

[0003] To overcome these limitations, Retrieval-Augmented Generation (RAG) technology has emerged. RAG combines an external knowledge base with LLMs, enabling the model to access and utilize knowledge beyond its context window when generating answers without additional fine-tuning. Traditional RAG methods assume that the required information is concentrated in a specific area of the knowledge base and retrieve relevant information by selecting the TopK text chunks with the highest semantic similarity to the question. This method is effective for local questions that only need to focus on a specific part of the knowledge base, but for global questions that require integrating information from multiple parts of the knowledge base, such as Query-Focused Summarization (QFS), the retrieved information is not accurate enough. Summary of the Invention

[0004] This application aims to propose a text question-answering method, system, electronic device, and medium based on semantic retrieval, which can improve the accuracy and retrieval efficiency of the retrieved information corresponding to global questions, thereby improving the precision and efficiency of text question-answering.

[0005] In a first aspect, an embodiment of this application provides a text question-answering method based on semantic retrieval, and the method includes: Obtain the question to be answered; Construct a hierarchical knowledge graph with highly cohesive communities; Determine whether the question to be answered belongs to a local question or a global question; If the problem to be answered belongs to a local problem, local search is performed through the local search module to obtain multiple answer texts similar to the problem to be answered; If the problem to be answered belongs to a global problem, global search is performed through the global search module in the hierarchical knowledge graph to obtain multiple communities related to the problem to be answered, and the communities are described as context and input into the large language model to generate answer texts.

[0006] Compared with the prior art, the first aspect of the present application has the following beneficial effects: This method includes obtaining the problem to be answered; constructing a hierarchical knowledge graph with highly cohesive communities; determining whether the problem to be answered belongs to a local problem or a global problem; if the problem to be answered belongs to a local problem, local search is performed through the local search module to obtain multiple answer texts similar to the problem to be answered; if the problem to be answered belongs to a global problem, global search is performed through the global search module in the hierarchical knowledge graph to obtain multiple communities related to the problem to be answered, and the communities are described as context and input into the large language model to generate answer texts. In this way, by determining whether the problem to be answered belongs to a local problem or a global problem, different methods are used to search for answer texts for different types of problems. In particular, global search is performed through the global search module in the hierarchical knowledge graph, which can be highly efficient in terms of time and computational cost, can significantly improve the processing ability of complex problems without significantly increasing the system complexity, can improve the accuracy and retrieval efficiency of the retrieved information corresponding to global problems, and thus improve the accuracy and efficiency of text Q&A.

[0007] In some embodiments, the determining whether the problem to be answered belongs to a local problem or a global problem includes: Constructing a similarity distribution between the problem to be answered and multiple text blocks; Calculating the ratio between the peak value and the secondary peak value of the similarity distribution to obtain the peak intensity; Calculating the second derivative of the similarity distribution at the peak point to obtain the distribution sharpness; Based on the peak intensity and the distribution sharpness, determining whether the problem to be answered belongs to a local problem or a global problem.

[0008] In some embodiments, the constructing a similarity distribution between the problem to be answered and multiple text blocks includes: ; wherein, represents the similarity distribution, represents the number of text blocks, and is a positive integer, represents the bandwidth parameter, represents the similarity score between the problem to be solved and any text block represents the similarity score between the problem to be solved and the th text block represents the kernel function

[0009] In some embodiments, calculating the second derivative of the similarity distribution at the peak point to obtain the sharpness of the distribution includes: ; ; wherein represents the sharpness of the distribution represents the similarity distribution represents the peak point represents the first derivative of the similarity distribution at the peak point represents the second derivative of the similarity distribution at the peak point represents the similarity score between the problem to be solved and the th text block represents the similarity score between the problem to be solved and the th text block

[0010] In some embodiments, based on the peak intensity and the sharpness of the distribution, determining whether the problem to be solved is a local problem or a global problem includes: Presetting a first threshold and a second threshold; If the peak intensity is greater than or equal to the first threshold, and the sharpness of the distribution is greater than or equal to the second threshold, then determine that the problem to be solved is a local problem; If the peak intensity is less than the first threshold, or the sharpness of the distribution is less than the second threshold, then determine that the problem to be solved is a global problem

[0011] In some embodiments, performing a global search in the hierarchical knowledge graph through a global search module to obtain multiple communities related to the problem to be solved includes: In the hierarchical knowledge graph, each node corresponds to a community; Calculating the similarity between the problem to be solved and each node, and using the similarity as the initial weight of each node; Calculating the weighted community weights according to the initial weights; Selecting multiple communities related to the problem to be solved based on the weighted community weights

[0012] In some embodiments, calculating the weighted community weight according to the initial weight value includes: ; where represents the weighted community weight, represents a coefficient for balancing the direct similarity and the inheritance weight, represents the weight of the parent node, represents the weight of the child node, represents a coefficient for adjusting the contribution of each child node to the weight of the parent node, represents the number of child nodes.

[0013] In a second aspect, an embodiment of the present application further provides a text question-answering system based on semantic retrieval. The system includes: A data acquisition unit for acquiring the question to be answered; A graph construction unit for constructing a hierarchical knowledge graph with highly cohesive communities; A question judgment unit for judging whether the question to be answered belongs to a local question or a global question; A local retrieval unit for, if the question to be answered belongs to a local question, performing a local search through a local retrieval module to obtain multiple answer texts similar to the question to be answered; A global retrieval unit for, if the question to be answered belongs to a global question, performing a global search in the hierarchical knowledge graph through a global retrieval module to obtain multiple communities related to the question to be answered, and describing the communities as context and inputting them into a large language model to generate an answer text.

[0014] In a third aspect, an embodiment of the present application further provides an electronic device, including at least one control processor and a memory communicatively connected to the at least one control processor; the memory stores instructions executable by the at least one control processor, and the instructions are executed by the at least one control processor so that the at least one control processor can execute a text question-answering method based on semantic retrieval as described above.

[0015] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, which stores computer-executable instructions for causing a computer to execute a text question-answering method based on semantic retrieval as described above.

[0016] It can be understood that the beneficial effects of the above second aspect to the fourth aspect compared with the related art are the same as those of the above first aspect compared with the related art. For the relevant descriptions, reference can be made to the relevant descriptions in the above first aspect, and details will not be repeated here. Brief Description of the Drawings

[0017] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the following description of embodiments in conjunction with the accompanying drawings, where: Figure 1 is a schematic flowchart of an embodiment of the text question-answering method based on semantic retrieval provided by the present application; Figure 2 is a schematic overall method flowchart in the best embodiment of the text question-answering method based on semantic retrieval provided by the present application; Figure 3 is a schematic structural diagram of an embodiment of the text question-answering system based on semantic retrieval provided by the present application; Figure 4 is a schematic structural diagram of an embodiment of the electronic device provided by the present application. Detailed Description of the Embodiments

[0018] The embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described by referring to the accompanying drawings below are exemplary and are only used to explain the present application and should not be construed as a limitation of the present application.

[0019] In the description of the present application, if the first, second, etc. are described only for the purpose of distinguishing technical features, they should not be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence of the indicated technical features.

[0020] It should also be understood that referring to "one embodiment" or "some embodiments" etc. in the description of the embodiments of the present application means that a specific feature, structure, or characteristic described in connection with the embodiment is included in one or more embodiments of the embodiments of the present application. Thus, the statements "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments" etc. that appear in different places in this specification are not necessarily all referring to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0021] In the description of the present application, it should be noted that unless otherwise clearly defined, words such as "set", "installed", "connected" etc. should be understood in a broad sense, and those skilled in the art can reasonably determine the specific meanings of the above words in the present application in combination with the specific content of the technical solution.

[0022] First, several nouns involved in the present application are analyzed: Retrieval-Augmented Generation Technology: This technology was proposed by Facebook Research. This method relies on semantic similarity retrieval. By calculating the similarity between the embedding vectors of the question and the text chunks in the knowledge base, it selects the top K most relevant text chunks as the basis for answer generation. This method performs well in dealing with local problems, but when dealing with global problems that require integrating multiple knowledge fragments, it often fails to provide sufficient context information.

[0023] Large Language Model: It refers to a deep learning model trained with a large amount of text data, enabling the model to generate natural language text or understand the meaning of language text. These models can provide in-depth knowledge and language production on various topics through training on a vast dataset. The core idea is to learn the patterns and structures of natural language through large-scale unsupervised training, and to a certain extent, simulate the human language cognition and generation process.

[0024] GraphRAG: It is a new retrieval method that uses a knowledge graph and vector search in the basic RAG architecture. Therefore, it can integrate and understand various knowledge, thus providing a broader and more comprehensive data view. GraphRAG was proposed by Microsoft Research. This method further expands the application of the graph structure and constructs a knowledge graph with communities. GraphRAG relies on LLMs to generate summaries of community content in sequence when retrieving knowledge. Although it can handle global problems, its cost is high, and the final retrieval result still "unfolds" the hierarchical structure into a linear structure, failing to fully utilize the potential of the hierarchical structure.

[0025] Natural Language Processing: It is an important research direction in the field of artificial intelligence, integrating knowledge from multiple disciplinary fields such as linguistics, computer science, machine learning, mathematics, and cognitive psychology. It is an interdisciplinary subject integrating computer science, artificial intelligence, and linguistics. It includes two main aspects: natural language understanding and natural language generation. The research content includes various levels such as characters, words, phrases, sentences, paragraphs, and texts. It is a bridge for communication between machine language and human language.

[0026] Knowledge-based Question Answering (KBQA) is an important research direction in the field of Natural Language Processing (NLP), aiming to answer users' questions by extracting information from structured knowledge bases or unstructured texts. With the rapid development of large language models (LLMs) such as GPT, LLaMA, and Gemini, these models have demonstrated powerful few-shot learning capabilities and can generate accurate answers based on a small amount of context information. However, LLMs still face some challenges in practical applications, especially their limited context window, which restricts their application in a wider range of scenarios.

[0027] To overcome these limitations, Retrieval-Augmented Generation (RAG) technology has emerged. RAG combines an external knowledge base with LLMs, enabling the model to access and utilize knowledge beyond its context window when generating answers without the need for additional fine-tuning. Traditional RAG methods assume that the required information is concentrated in a specific area of the knowledge base and retrieve relevant information by selecting the TopK text chunks with the highest semantic similarity to the question. This method is effective for local questions that only need to focus on a specific part of the knowledge base, but for global questions that require integrating information from multiple parts of the knowledge base, such as Query-Focused Summarization (QFS), the retrieved information is not accurate enough.

[0028] To solve the problem that the information retrieved by the above-mentioned existing related technologies for global questions is not accurate enough, this application proposes a text question-answering method, system, electronic device, and medium based on semantic retrieval.

[0029] Refer to Figure 1 This application embodiment provides a text question-answering method based on semantic retrieval, and the method includes the following steps: Step S100, obtain the question to be answered; Step S200, construct a hierarchical knowledge graph with highly cohesive communities; Step S300, determine whether the question to be answered is a local question or a global question; Step S400, if the question to be answered is a local question, perform local search through the local retrieval module to obtain multiple answer texts similar to the question to be answered; Step S500: If the question to be answered belongs to a global question, the global search module performs a global search in the hierarchical knowledge graph to obtain multiple communities related to the question to be answered, and describes the communities as context and inputs them into the large language model to generate an answer text.

[0030] In this embodiment, by obtaining the question to be answered; constructing a hierarchical knowledge graph with highly cohesive communities; determining whether the question to be answered belongs to a local question or a global question; if the question to be answered belongs to a local question, the local search module performs a local search to obtain multiple answer texts similar to the question to be answered; if the question to be answered belongs to a global question, the global search module performs a global search in the hierarchical knowledge graph to obtain multiple communities related to the question to be answered, and describes the communities as context and inputs them into the large language model to generate an answer text. In this way, by determining whether the question to be answered belongs to a local question or a global question, different methods are used to search for answer texts for different types of questions. Especially by performing a global search in the hierarchical knowledge graph through the global search module, it can be highly efficient in terms of time and computational cost, can significantly improve the processing ability for complex questions without significantly increasing the system complexity, can improve the accuracy and retrieval efficiency of the retrieved information corresponding to global questions, and thus improve the accuracy and efficiency of text Q&A.

[0031] The above-mentioned construction of a hierarchical knowledge graph with highly cohesive communities can be to merge entities with the same name and type in the sub-graphs, and merge relationships with the same source entity and target entity, thereby creating cross-sub-graph relationships, and finally generating a hierarchical knowledge graph with highly cohesive communities, which can more accurately reflect the global structure of the source document.

[0032] The above-mentioned local search through the local search module to obtain multiple answer texts similar to the question to be answered can be to directly adopt the most classic retrieval method based on the K-Nearest Neighbors (KNN) algorithm. The local search module calculates the similarity between the question to be answered and multiple texts using a similarity calculation method to obtain multiple answer texts similar to the question to be answered.

[0033] In some embodiments, determining whether the question to be answered belongs to a local question or a global question includes: Constructing a similarity distribution between the question to be answered and multiple text blocks; Calculating the ratio between the peak value of the similarity distribution and the secondary peak value of the similarity distribution to obtain the peak intensity; Calculating the second derivative of the similarity distribution at the peak point to obtain the distribution sharpness; Based on the peak intensity and the sharpness of the distribution, determine whether the problem to be solved belongs to a local problem or a global problem.

[0034] In this embodiment, by constructing the similarity distribution between the problem to be solved and multiple text blocks; calculating the ratio between the peak of the similarity distribution and the sub-peak of the similarity distribution to obtain the peak intensity; calculating the second derivative of the similarity distribution at the peak point to obtain the sharpness of the distribution; based on the peak intensity and the sharpness of the distribution, determine whether the problem to be solved belongs to a local problem or a global problem. In this way, by calculating the semantic similarity between the problem to be solved and each text block, the similarity distribution between the problem to be solved and multiple text blocks is constructed. Since there are significant differences in the similarity distributions of local problems and global problems, by analyzing the peak intensity and the sharpness of the distribution, it is possible to accurately determine whether the problem to be solved belongs to a local problem or a global problem, laying a good data foundation for accurately retrieving information in the later stage.

[0035] In some embodiments, constructing the similarity distribution between the problem to be solved and multiple text blocks includes: ; wherein, represents the similarity distribution, represents the number of text blocks, and is a positive integer, represents the bandwidth parameter, represents the similarity score between the problem to be solved and any text block, represents the similarity score between the problem to be solved and the th text block, represents the kernel function.

[0036] In some embodiments, calculating the second derivative of the similarity distribution at the peak point to obtain the sharpness of the distribution includes: ; ; wherein, represents the sharpness of the distribution, represents the similarity distribution, represents the peak point, represents the first derivative of the similarity distribution at the peak point, represents the second derivative of the similarity distribution at the peak point, represents the similarity score between the problem to be solved and the th text block, represents the similarity score between the problem to be solved and the th text block.

[0037] In some embodiments, based on the peak intensity and the sharpness of the distribution, determining whether the question to be answered belongs to a local problem or a global problem includes: Presetting a first threshold and a second threshold; If the peak intensity is greater than or equal to the first threshold and the sharpness of the distribution is greater than or equal to the second threshold, then determine that the question to be answered belongs to a local problem; If the peak intensity is less than the first threshold or the sharpness of the distribution is less than the second threshold, then determine that the question to be answered belongs to a global problem.

[0038] In this embodiment, a first threshold and a second threshold are preset; if the peak intensity is greater than or equal to the first threshold and the sharpness of the distribution is greater than or equal to the second threshold, then determine that the question to be answered belongs to a local problem; if the peak intensity is less than the first threshold or the sharpness of the distribution is less than the second threshold, then determine that the question to be answered belongs to a global problem. In this way, by comprehensively considering the peak intensity and the sharpness of the distribution to determine whether the question to be answered belongs to a local problem or a global problem, it is possible to accurately determine whether the question to be answered belongs to a local problem or a global problem, laying a good data foundation for accurately retrieving information in the later stage.

[0039] In some embodiments, the global search module performs a global search in the hierarchical knowledge graph to obtain multiple communities related to the question to be answered, including: In the hierarchical knowledge graph, each node corresponds to a community; Calculating the similarity between the question to be answered and each node, and using the similarity as the initial weight of each node; Calculating the weighted community weights according to the initial weights; Based on the weighted community weights, selecting multiple communities related to the question to be answered.

[0040] In this embodiment, in the hierarchical knowledge graph, each node corresponds to a community; calculating the similarity between the question to be answered and each node, and using the similarity as the initial weight of each node; calculating the weighted community weights according to the initial weights; based on the weighted community weights, selecting multiple communities related to the question to be answered. In this way, by performing a global search in the hierarchical knowledge graph through the global search module, it is possible to have high efficiency in terms of time and computational cost, and to significantly improve the processing ability of complex problems without significantly increasing the system complexity, thereby improving the accuracy and retrieval efficiency of the retrieval information corresponding to global problems.

[0041] In some embodiments, calculating the weighted community weights according to the initial weights includes: ; Wherein, Represents the weighted community weight, Represents the coefficient for balancing direct similarity and inheritance weight, Represents the weight of the parent node, Represents the weight of the child node, Represents the coefficient for adjusting the contribution of each child node to the parent node weight, Represents the number of child nodes.

[0042] For the convenience of those skilled in the art to understand, the following provides a set of best embodiments: Knowledge-based Question Answering (KBQA) is an important research direction in the field of Natural Language Processing (NLP), aiming to answer users' questions by extracting information from structured knowledge bases or unstructured texts. With the rapid development of large language models (LLMs) such as GPT, LLaMA, and Gemini, these models demonstrate powerful few-shot learning capabilities and can generate accurate answers based on a small amount of context information. However, LLMs still face some challenges in practical applications, especially their limited context window, which restricts their application in a wider range of scenarios.

[0043] To overcome these limitations, Retrieval-Augmented Generation (RAG) technology has emerged. RAG combines an external knowledge base with LLMs, enabling the model to access and utilize knowledge beyond its context window when generating answers without additional fine-tuning. Traditional RAG methods assume that the required information is concentrated in a specific area of the knowledge base and retrieve relevant information by selecting the TopK text chunks with the highest semantic similarity to the question. This method is effective for local questions that only need to focus on a specific part of the knowledge base, but performs poorly for global questions that require integrating information from multiple parts of the knowledge base, such as Query-Focused Summarization (QFS). For example, when using existing RAG methods to handle local or global questions, the following limitations still exist: 1. Imbalanced handling of local and global questions: Existing methods often can only retrieve local or global information and lack a general and efficient information retrieval paradigm to handle both types of questions simultaneously.

[0044] 2. High computational cost: Some methods (such as GraphRAG) rely on the LLM for multi-round summary generation when retrieving information, resulting in a high question-answering cost and a long response time.

[0045] 3. Information over-generalization: In the hierarchical retrieval process, existing methods often "unfold" the hierarchical structure into a linear structure, resulting in the inability to distinguish high-level semantic information or over-generalization, and the advantages of the hierarchical structure cannot be fully utilized.

[0046] To address the limitations of the above-mentioned prior art, this embodiment proposes a simple, general, and efficient RAG retrieval paradigm that can uniformly retrieve the information required for local and global problems. Compared with the prior art, the main differences in the technical solution of this embodiment are as follows: 1. Problem classification algorithm: This embodiment proposes a problem classification algorithm based on semantic similarity distribution, which can accurately determine the problem type without relying on LLMs.

[0047] 2. Hierarchical graph structure: This embodiment introduces a hierarchical graph structure to represent the original knowledge base, which can better capture the global and local relationships of knowledge.

[0048] 3. Hierarchical retrieval strategy: This embodiment proposes a hierarchical retrieval strategy that can minimize information loss in terms of time and cost efficiency and is applicable to both local and global problems.

[0049] Through the above innovations, this embodiment can retrieve more accurate and comprehensive information while maintaining a low computational cost, and is applicable to a wide range of knowledge-intensive question-answering tasks. Referring to Figure 2 , the technical solution of this embodiment specifically includes the following content: I. Graph construction module.

[0050] The method of this embodiment first improves the knowledge graph construction technology of GraphRAG. Similar to GraphRAG, this embodiment extracts entities and relationships from source documents. Each entity includes a name, a description, and a type (from a predefined set of types), while each relationship includes a source entity, a target entity, a brief description, and a score representing the strength of the relationship. Through these components, this embodiment constructs a weighted undirected knowledge graph to capture the structural semantics of the source documents. This embodiment also detects hierarchical communities in the graph and generates a summary for each community based on the descriptions of the entities and relationships within the community. However, this embodiment introduces a key improvement to address an important limitation of GraphRAG in practical applications.

[0051] In an actual RAG scenario, the source documents usually exceed the context window limit of the LLM. Therefore, it is necessary to chunk the documents, which results in multiple sub-graphs, each representing only a part of the document. Without integrating these sub-graphs, there will be a lack of global consistency among them, leading to fragmentation of the community representation. To solve this problem, this embodiment proposes a method for integrating sub-graphs. First, this embodiment merges entities with the same name and type in different sub-graphs and uses the LLM to generate a unified description for each merged entity to ensure consistency and integrity. Second, for relationships with the same source entity and target entity, this embodiment also merges them and uses the LLM to generate the integrated description and strength score. Through this integration process, this embodiment can create relationships across sub-graphs and finally generate a hierarchical knowledge graph with a highly cohesive community, thus more accurately reflecting the global structure of the source document.

[0052] II. Gating Module.

[0053] The core function of the gating module is to determine whether the current question (i.e., the question to be answered, also the user question) is a local question or a global question, thus guiding the selection of the subsequent retrieval module. Its implementation is based on the following principle: This embodiment first divides the original text into several text chunks and calculates the semantic similarity between the question and each text chunk. For local questions, their answers are usually concentrated in a specific area of the text, so there will be a significant peak in the similarity distribution, indicating that the question is highly relevant to a certain text chunk. For global questions, the required information is often scattered in multiple parts of the text, so the similarity distribution will be relatively smooth without an obvious peak.

[0054] By analyzing the morphological characteristics of the similarity distribution, the gating module can effectively distinguish the type of question: if there is a significant peak in the similarity distribution, it is determined as a local question; if the similarity distribution is relatively smooth, it is determined as a global question. This mechanism provides a reliable basis for the selection of the subsequent retrieval module, thus optimizing the overall performance of the question-answering system. The specific theoretical proof is as follows: This embodiment first uses kernel density estimation (KDE) to model the similarity distribution, so as to more accurately describe the distribution characteristics of local questions and global questions. Given a question and text chunks with the semantic similarity score of , the similarity distribution is modeled as: (1); where, represents the similarity distribution, represents the kernel function, satisfying , represents a bandwidth parameter that controls the smoothness degree. represents the similarity score between a question and any text block. represents the similarity score between the question to be answered and the th text block. Based on the similarity distribution of the above kernel density estimation , statistics of two features, peak intensity and distribution curvature, are defined as follows: 1. Peak Strength: Used to calculate the ratio of the peak to the sub-peak of the similarity distribution. The larger the ratio, the more concentrated the distribution indicates. where

[0055] (2); Among them, represents the peak of the similarity distribution , represents the sub-peak of the similarity distribution .

[0056] 2. Distribution Curvature: Calculate the second derivative of the similarity distribution at the peak point to quantify the sharpness of the distribution: (3); Next, the central difference method is used to approximate the second derivative: (4); Among them, represents the first derivative of the similarity distribution at the peak point, represents the second derivative of the similarity distribution at the peak point, represents the similarity score between the question to be answered and the th text block, represents the similarity score between the question to be answered and the th text block. In this embodiment, it is further assumed that the distribution characteristics of two types of questions satisfy: Local question: and ; Global question: or . Among them, the first threshold and the second threshold are preset thresholds determined by statistical learning.

[0057] The following proves the theorem: If the answers to local questions are concentrated in a single text block and the information of global questions is evenly distributed, then through and , the two types of questions can be effectively distinguished.

[0058] Proof: If it is a local problem, there exists a unique such that is significantly higher than others , making the KDE distribution at present a sharp peak: (the denominator approaches 0) and (the curvature is large and the distribution is sharp); if it is a global problem, there exists a uniform distribution, making the KDE distribution close to flat: (no obvious peak) and ( the curvature is small and the distribution is smooth).

[0059] Conclusion: There are significant differences in the and of the two types of problems. By threshold division, it can be ensured that: (5); where represents the probability of correctly classifying whether the problem is a global problem or a local problem.

[0060] Supplementary proof: 1. Bandwidth selection.

[0061] The selection of the bandwidth is crucial for KDE modeling. The method of this embodiment is optimized by the Silverman criterion: (6); where is the sample standard deviation.

[0062] 2. Noise robustness.

[0063] In the actual scenario, the similarity distribution may contain noise. The signal-to-noise ratio is defined as follows: (7); where represents the highest semantic similarity score between the problem and all text blocks , that is, the signal strength. It reflects the matching degree between the problem and the most relevant text block. The larger the value, the stronger the relevance between the problem and a certain text block. represents the standard deviation of the semantic similarity scores, which is used to measure the dispersion degree of the similarity scores, that is, the noise strength. It is obtained by taking the square root of the average of the sum of the squared deviations of each similarity score from the mean . The larger it is, the more dispersed the similarity distribution is and the more noise there is; The smaller it is, the more concentrated the similarity distribution is and the less noise there is. It can be calculated that for the method of this embodiment , the influence of noise on threshold determination can be ignored.

[0064] III. Local Retrieval Module.

[0065] The core task of this module is to quickly locate the local information most relevant to the question from the knowledge base to support subsequent answer generation. Here, the most classic retrieval method based on the K-Nearest Neighbors (KNN) algorithm is directly adopted. The core idea of the KNN algorithm is to select the most relevant text blocks as retrieval results by calculating the semantic similarity between the question and the text blocks in the knowledge base. Specifically, given a question and a text block in the knowledge base, first, the question and the text block are respectively encoded into semantic vectors and through a pre-trained embedding generation model (Embedding), and then the cosine similarity between the question vector and each text block vector is calculated: (8); Finally, the text blocks with the highest similarity are selected as retrieval results. The KNN algorithm can match the question and the text segment at a fine-grained level, so as to more accurately locate relevant information. This method is simple and efficient and is suitable for local questions where the answers are concentrated in a local area of the knowledge base.

[0066] IV. Global Retrieval Module.

[0067] The core task of the global retrieval module is to integrate multiple parts of information from the hierarchical knowledge graph to answer global questions that require comprehensive knowledge. Traditional retrieval methods based on cosine similarity have limitations in dealing with global questions because the similarity scores of high-level nodes (representing abstract concepts) are usually low, which may cause them to be ignored, and relying only on the parent-child relationship may miss relevant information. To solve this problem, this embodiment proposes a weighted retrieval mechanism that combines the position information of nodes in the hierarchical structure on the basis of cosine similarity, so as to more comprehensively capture the requirements of global questions. In the hierarchical tree structure constructed by the Graph Indexer, each node corresponds to a community. This embodiment first calculates the initial weight of each node based on cosine similarity: (9); where, is the embedding vector of the question, The embedding vector described for the node. In order to incorporate hierarchical context information, in this embodiment, the weights of each parent node are updated. Among them, is the weight of the parent node, is the weight of the child node, is the number of child nodes, is used to adjust the contribution of each child node to the weight of the parent node, (usually set to 0.3) is used to balance the direct similarity and the inherited weight.

[0068] (10); Based on the weighted community weights, in this embodiment, the communities with the highest weights are selected (in the experiment ), and their descriptions are used as context to input into the LLM to generate answers. Compared with directly using the original description, the sub-answers can reduce noise and optimize the utilization efficiency of the context window. These sub-answers, as the retrieved information, provide high-quality knowledge support for the subsequent answer generation.

[0069] Compared with the prior art, the technical solution of this embodiment has the following advantages: The retrieval method proposed in this embodiment significantly improves the performance of the question-answering system in dealing with local and global questions by combining the hierarchical knowledge graph and the weighted retrieval mechanism. For local questions, the traditional KNN retrieval method can quickly locate the text block most relevant to the question, ensuring the accuracy and retrieval efficiency of the answer; while for global questions, the hierarchical weighted retrieval mechanism avoids the limitations of a single similarity metric by integrating the abstract information of high-level nodes and the detailed information of low-level nodes, thus capturing more comprehensively the key information scattered in multiple parts of the knowledge base. In addition, by generating sub-answers instead of directly using the original description, the noise is further reduced and the utilization rate of the context window is optimized.

[0070] The core advantages of the method in this embodiment lie in its generality and efficiency. Through simple similarity distribution analysis, the gating module can automatically judge the question type and flexibly select the local retrieval or global retrieval strategy without relying on a complex large language model (LLM) proxy. At the same time, the hierarchical weighted retrieval mechanism is highly efficient in terms of time and computational cost, and can significantly improve the processing ability of complex questions without significantly increasing the system complexity. Experimental results show that the method in this embodiment outperforms the existing state-of-the-art methods in various knowledge-intensive question-answering tasks, providing reliable technical support for practical applications.

[0071] Refer to Figure 3, an embodiment of the present application also provides a text question - answering system based on semantic retrieval. The system includes a data acquisition unit 100, a knowledge graph construction unit 200, a question judgment unit 300, a local retrieval unit 400, and a global retrieval unit 500, where: The data acquisition unit 100 is used to acquire the question to be answered; The knowledge graph construction unit 200 is used to construct a hierarchical knowledge graph with highly cohesive communities; The question judgment unit 300 is used to judge whether the question to be answered belongs to a local question or a global question; The local retrieval unit 400 is used to, if the question to be answered belongs to a local question, perform local search through a local retrieval module to obtain multiple answer texts similar to the question to be answered; The global retrieval unit 500 is used to, if the question to be answered belongs to a global question, perform global search in the hierarchical knowledge graph through a global retrieval module to obtain multiple communities related to the question to be answered, and describe the communities as context and input them into a large - language model to generate answer texts.

[0072] In some embodiments, the question judgment unit 300 can specifically be used to: Construct a similarity distribution between the question to be answered and multiple text blocks; Calculate the ratio between the peak value of the similarity distribution and the secondary peak value of the similarity distribution to obtain the peak intensity; Calculate the second - order derivative of the similarity distribution at the peak point to obtain the distribution sharpness; Based on the peak intensity and the distribution sharpness, judge whether the question to be answered belongs to a local question or a global question.

[0073] In some embodiments, the question judgment unit 300 can specifically be used to: ; Wherein, represents the similarity distribution, represents the number of text blocks, and is a positive integer, represents the bandwidth parameter, represents the similarity score between the question to be answered and any text block, represents the question to be answered and the th text block similarity score, represents the kernel function.

[0074] In some embodiments, the question judgment unit 300 can specifically be used to: ; ; Among them, represents the sharpness of the distribution, represents the similarity distribution, represents the peak point, represents the second derivative of the similarity distribution at the peak point, represents the similarity score between the question to be answered and the th text block, represents the similarity score between the question to be answered and the th text block.

[0075] In some embodiments, the question judgment unit 300 may be specifically configured to: Preset a first threshold and a second threshold; If the peak intensity is greater than or equal to the first threshold, and the sharpness of the distribution is greater than or equal to the second threshold, it is determined that the question to be answered belongs to a local problem; If the peak intensity is less than the first threshold, or the sharpness of the distribution is less than the second threshold, it is determined that the question to be answered belongs to a global problem.

[0076] In some embodiments, the global retrieval unit 500 may be specifically configured to: In the hierarchical knowledge graph, each node corresponds to a community; Calculate the similarity between the question to be answered and each node, and use the similarity as the initial weight of each node; Calculate the weighted community weights according to the initial weights; Based on the weighted community weights, select multiple communities related to the question to be answered.

[0077] In some embodiments, the global retrieval unit 500 may be specifically configured to: ; Among them, represents the weighted community weight, represents the coefficient for balancing the direct similarity and the inheritance weight, represents the weight of the parent node, represents the weight of the child node, represents the coefficient for adjusting the contribution of each child node to the weight of the parent node, represents the number of child nodes.

[0078] It should be noted that since the text question-answering system based on semantic retrieval in this embodiment and the above-mentioned text question-answering method based on semantic retrieval are based on the same inventive concept, the corresponding content in the method embodiment also applies to this system embodiment and will not be elaborated here.

[0079] Refer toFigure 4 , embodiments of the present application also provide an electronic device, which includes: At least one memory; At least one processor; At least one program; The program is stored in the memory, and the processor executes at least one program to implement the text question-answering method based on semantic retrieval described above in the present disclosure.

[0080] The electronic device can be any intelligent terminal including a mobile phone, a tablet computer, a personal digital assistant (PDA), an in-vehicle computer, etc.

[0081] The electronic device of the embodiments of the present application will be introduced in detail below.

[0082] The processor 1600 can be implemented by using a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present disclosure; The memory 1700 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 1700 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1700, and the processor 1600 is called to execute the text question-answering method based on semantic retrieval of the embodiments of the present disclosure.

[0083] The input / output interface 1800 is used to implement information input and output; The communication interface 1900 is used to implement communication interaction between this device and other devices, and can implement communication through a wired method (such as USB, network cable, etc.) or through a wireless method (such as a mobile network, WIFI, Bluetooth, etc.); The bus 2000 transmits information between various components of the device (such as the processor 1600, the memory 1700, the input / output interface 1800, and the communication interface 1900); Among them, the processor 1600, the memory 1700, the input / output interface 1800, and the communication interface 1900 are communicatively connected to each other inside the device through the bus 2000.

[0084] An embodiment of the present disclosure also provides a storage medium, which is a computer-readable storage medium storing computer-executable instructions for causing a computer to execute the above-described text question-answering method based on semantic retrieval.

[0085] As a non-transitory computer-readable storage medium, a memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0086] The embodiments described in the embodiments of the present disclosure are for more clearly illustrating the technical solutions of the embodiments of the present disclosure, and do not constitute a limitation on the technical solutions provided by the embodiments of the present disclosure. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present disclosure are equally applicable to similar technical problems.

[0087] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present disclosure, and may include more or fewer steps than shown in the figures, or combine certain steps, or different steps.

[0088] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0089] Those of ordinary skill in the art can understand that all or some of the steps in the above-disclosed methods, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and appropriate combinations thereof.

[0090] In the description of this application and the above-mentioned accompanying drawings, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0091] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expression refers to any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or plural.

[0092] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical, or other forms.

[0093] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0094] In addition, in each embodiment of the present application, the functional units may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above integrated unit may be implemented in the form of hardware or in the form of a software functional unit.

[0095] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes multiple instructions for causing an electronic device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in the various embodiments of the present application. The foregoing storage medium includes: various media that can store programs, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs. The embodiments of the present application have been described in detail above with reference to the accompanying drawings, but the present application is not limited to the above embodiments. Within the knowledge scope of those of ordinary skill in the art to which the present application pertains, various changes may be made without departing from the purpose of the present application.

[0096] The embodiments of the present application have been described in detail above with reference to the accompanying drawings, but the present application is not limited to the above embodiments. Within the knowledge scope of those of ordinary skill in the art to which the present application pertains, various changes may be made without departing from the purpose of the present application.

Claims

1. A text question answering method based on semantic retrieval, characterized in that: The method comprises: Get questions to be answered; Build a hierarchical knowledge graph with highly cohesive communities; Determine whether the problem to be solved is a local problem or a global problem; If the question to be answered is a local question, a local search is performed through a local search module to obtain multiple answer texts similar to the question to be answered; If the question to be answered is a global question, a global search is performed in the hierarchical knowledge graph through a global retrieval module to obtain multiple communities related to the question to be answered, and the communities are described as contexts and input into a large language model to generate an answer text.

2. The text question answering method based on semantic retrieval according to claim 1, characterized in that: The determining whether the problem to be solved is a local problem or a global problem includes: Constructing a similarity distribution between the question to be answered and multiple text blocks; Calculating a ratio between a peak value of the similarity distribution and a sub-peak value of the similarity distribution to obtain a peak intensity; Calculating the second-order derivative of the similarity distribution at the peak point to obtain the distribution sharpness; Based on the peak intensity and the distribution sharpness, it is determined whether the problem to be solved is a local problem or a global problem.

3. The text question answering method based on semantic retrieval according to claim 2 is characterized in that: The constructing of the similarity distribution between the question to be answered and the plurality of text blocks includes: ; in, represents the similarity distribution, represents the number of text blocks, and is a positive integer, represents the bandwidth parameter, represents the similarity score between the question to be answered and any text block, Indicates the question to be answered and The similarity score between text blocks, Represents the kernel function.

4. The text question answering method based on semantic retrieval according to claim 2 is characterized in that: The calculating the second-order derivative of the similarity distribution at the peak point to obtain the distribution sharpness includes: ; ; in, Indicates the sharpness of the distribution. represents the similarity distribution, Indicates the peak point, represents the first-order derivative of the similarity distribution at the peak point, represents the second-order derivative of the similarity distribution at the peak point, Indicates the question to be answered and The similarity score between text blocks, Indicates the question to be answered and The similarity score between text blocks.

5. The text question answering method based on semantic retrieval according to claim 2, characterized in that: The determining, based on the peak intensity and the distribution sharpness, whether the problem to be solved is a local problem or a global problem includes: Presetting a first threshold and a second threshold; If the peak intensity is greater than or equal to the first threshold, and the distribution sharpness is greater than or equal to the second threshold, it is determined that the problem to be solved is a local problem; If the peak intensity is less than the first threshold, or the distribution sharpness is less than the second threshold, it is determined that the problem to be solved is a global problem.

6. The text question answering method based on semantic retrieval according to claim 1, characterized in that: The global search module is used to perform a global search in the hierarchical knowledge graph to obtain multiple communities related to the question to be answered, including: In the hierarchical knowledge graph, each node corresponds to a community; Calculating the similarity between the question to be answered and each node, and using the similarity as the initial weight of each node; Calculating a weighted community weight based on the initial weight; Based on the weighted community weights, multiple communities related to the question to be answered are selected.

7. The text question answering method based on semantic retrieval according to claim 6, characterized in that: The step of calculating the weighted community weight according to the initial weight includes: ; in, represents the weighted community weight, represents the coefficient used to balance direct similarity and inherited weight, represents the weight of the parent node, represents the weight of the child node, represents the coefficient used to adjust the contribution of each child node to the parent node weight, Indicates the number of child nodes.

8. A text question answering system based on semantic retrieval, characterized in that: The system comprises: A data acquisition unit, used for acquiring questions to be answered; Graph construction unit, used to build a hierarchical knowledge graph with highly cohesive communities; A question determination unit, used to determine whether the question to be answered is a local question or a global question; A local search unit, configured to perform a local search through a local search module to obtain a plurality of answer texts similar to the question to be answered if the question to be answered is a local question; A global retrieval unit is used to perform a global search in the hierarchical knowledge graph through a global retrieval module if the question to be answered is a global question, obtain multiple communities related to the question to be answered, and describe the communities as contexts input into a large language model to generate an answer text.

9. An electronic device, characterized in that: It includes at least one control processor and a memory for communicating with the at least one control processor; the memory stores instructions that can be executed by the at least one control processor, and the instructions are executed by the at least one control processor to enable the at least one control processor to execute the text question answering method based on semantic retrieval as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the text question answering method based on semantic retrieval as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Friend recommendation system and method based on community detection

    CN108399189A

  • Man-machine conversation processing method and device

    CN111708869A

  • Visual question and answer method based on cross-modal pre-training feature enhancement

    CN114663677A

  • Answer acquisition method, computer program product, device and storage medium

    CN119377372A

  • Intelligent question answering system and method based on question decomposition and community semantic search

    CN119537539A

Cited By

  • Intelligent student question and answer method and system based on large language model and rapid retrieval

    CN121009178A

  • Intelligent film and television media asset retrieval method and system based on image enhancement generation

    CN121524374A

  • Intelligent video and television media retrieval method and system based on graph enhancement generation

    CN121524374B