A RAG storage and retrieval method and device based on a graph database
The RAG method using graph databases for structured storage and retrieval addresses LLM inaccuracies by ensuring accurate and reliable information generation in complex text scenarios.
Patent Information
- Application Number
- CN202411426810.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-14
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2044-10-14
AI Technical Summary
Large language models have problems of hallucination and nonsense when generating text, especially when dealing with text information at complex hierarchical relationships, which leads to inaccuracy and unreliability of information retrieval.
The RAG storage and retrieval method based on graph database is adopted, and the multi-level node construction and structured storage is constructed and combined with semantic vector model and heat coefficients to achieve efficient and accurate retrieval and generation of documents.
It improves the reliability and accuracy of information generation, avoids illusions and nonsense, and ensures the relevance of search results and user attention.
Smart Images

Figure CN118939782B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of natural language processing, and particularly relates to a RAG storage and retrieval method and device based on a graph database. Background Art
[0002] In recent years, large language models (LLMs) have made remarkable progress in the field of natural language processing. With their powerful language generation and understanding capabilities, they are widely used in fields such as machine translation, dialogue systems, text generation, and information retrieval. However, although LLMs perform well in processing unstructured data, they also have some inherent problems when generating content, such as "hallucinations" and "nonsense". Specifically, when generating text, the model may generate non-existent facts out of thin air or confuse known information, resulting in inaccurate and unreliable results.
[0003] The root cause of these problems is that LLMs rely on a large amount of unstructured text data for training, and during the generation process, they cannot verify the generated content through logical reasoning or structured data. Especially when dealing with text information with complex hierarchical relationships, LLMs often have difficulty ensuring the accuracy and consistency of the generated content. This poses a great challenge to many application scenarios that require high-precision information retrieval, such as law, medicine, and scientific research literature.
[0004] To solve the above problems, the present invention proposes a RAG (Retrieve-And-Generate) retrieval method based on a graph database. Through the graph database, hierarchical and structured storage of complex text data can be achieved, thereby enabling effective management and retrieval of multi-level data such as keywords, file names, chapter contents, and paragraph contents. On this basis, by combining retrieval and generation methods, the reliability and accuracy of information generation can be significantly improved, avoiding the hallucination and nonsense problems of LLMs during the data generation process. Summary of the Invention
[0005] The purpose of the present invention is to provide a RAG (Retrieve-And-Generate) storage and retrieval method and device based on a graph database in view of the deficiencies of the prior art, aiming to solve the "hallucination" problem existing in the existing large language models when generating text, as well as the defect of low efficiency in complex document retrieval of traditional databases. The present invention realizes efficient and accurate retrieval of documents through the multi-level node construction and structured storage of the graph database, and provides a more stable and accurate input basis for large language models.
[0006] The purpose of the present invention is achieved through the following technical solutions: In the first aspect, the present invention provides a RAG storage and retrieval method based on a graph database, including two steps: storage and retrieval
[0007] The storage step includes extracting the full-text summary and keywords of the document data, as well as the chapter and section directory, performing segmentation processing, and setting the popularity coefficient of the document, which is stored in the graph database. Among them, the keywords are associated with the file name, the file name is associated with the full-text summary, chapter and section directory, and popularity coefficient, and the chapter and section directory is associated with the segmented content for subsequent retrieval processes.
[0008] The retrieval step includes matching the rewritten user question with the keywords to obtain the associated file name, calculating the file similarity based on the file name, full-text summary, and popularity coefficient, screening the files that meet the requirements, and then obtaining the chapter and section directory and segmented content that match the user question, and generating the final answer based on the large language model and returning it to the user.
[0009] Furthermore, according to the content attributes of all documents, a keyword library K i = [k0, k1, k2, … k n , where k n represents the nth keyword; and for each keyword, use the semantic vector model to convert the keyword into a vector: V k = [v0, v1, v2, … v n , where v n represents the nth keyword vector; also convert the full-text summary of the document into a vector using the semantic vector model: V f ; for each file, calculate the cosine similarity S i between the keyword vector v f and the summary vector V i , sort the cosine similarities from high to low, retain the keywords with similarities greater than the threshold, and take the Top3 as the keywords of the document.
[0010] Furthermore, the popularity coefficient of the document refers to the popularity value of the number of questions in the document that are answers to the questions asked by all users within a period of time, accounting for the total number of questions. Num hit is the number of questions in the document that are answers to the questions asked by users within a period of time, and Num total is the total number of questions asked by users within a period of time. The popularity coefficient H(f) of this document is obtained as follows. The higher the popularity coefficient, the more concerned users are about the content of this document and the more questions they ask, and vice versa;
[0011] (1)
[0012] (2)
[0013] Among them, tanh is the hyperbolic tangent function, which is used to normalize the heat value so that its range is between 0 and 1. The heat coefficient makes the more concerned documents have higher priority in the retrieval results.
[0014] Furthermore, use the LLM to extract the chapter and section directory structure of the document to obtain the structural hierarchy of the document; according to the chapter and section directory structure of the document, for each chapter content, perform segmentation processing, use the LLM to cut each chapter into multiple segmented content chunks, and convert them into vectors using the semantic vector model.
[0015] Furthermore, the keywords, file names, full text abstracts, chapters, and segmented content in the graph database are all converted into vectors using the semantic vector model and then stored, and the heat coefficient is stored using floating-point data.
[0016] Furthermore, in the retrieval step, convert the user's question into a vector using the semantic vector model, and match it with each keyword in the keyword library, retain the keywords with a similarity greater than the threshold as the candidate keyword set, traverse the file names associated with the keywords, calculate the file similarity and retain the file names with a similarity greater than the threshold as the file set; according to each file in the file set, calculate the similarity between the user's question and the chapter, and obtain the chapter with the highest similarity; and calculate the similarity between the user's question and the segmented content, and retain the segmented content with a similarity greater than the threshold as the content finally retrieved.
[0017] Furthermore, the file similarity is calculated from the file name, full text abstract, and heat coefficient. Calculate the similarity between the vectorized user's question and the vectorized file name and full text abstract respectively. The final file similarity The calculation formula is as follows:
[0018] (3)
[0019] (4)
[0020] Among them, S name is the similarity between the rewritten user's question and the file name, S summary is the similarity between the rewritten user's question and the full text abstract, c0 and c1 are coefficients, the value range is 0~1, and c0 + c1 = 1, and H is the heat coefficient.
[0021] In the second aspect, the present invention also provides a RAG storage and retrieval device based on a graph database, including a memory and one or more processors. Executable code is stored in the memory. When the processor executes the executable code, it implements the above-mentioned RAG storage and retrieval method based on a graph database.
[0022] In a third aspect, the present invention also provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, the described method for RAG storage and retrieval based on a graph database is implemented.
[0023] In a fourth aspect, the present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the described method for RAG storage and retrieval based on a graph database is implemented.
[0024] Advantages of the present invention: Through the RAG retrieval method based on a graph database of the present invention, the "hallucination" problem existing in existing large language models when processing complex text information can be effectively overcome. By using the multi-level structured storage and retrieval technology of the graph database, combined with key technologies such as vectorization processing and heat coefficient, the retrieval results are more accurate and have higher relevance, thereby providing users with more reliable and practical intelligent information services. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following-described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0026] Figure 1 It is a data storage flow chart of the present invention.
[0027] Figure 2 It is a data retrieval flow chart of the present invention.
[0028] Figure 3 It is a schematic diagram of the storage structure of the graph database of the present invention.
[0029] Figure 4 It is a schematic diagram of the heat coefficient calculation formula function of the present invention.
[0030] Figure 5 It is a RAG storage and retrieval based on a graph database. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0031] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0032] The present invention discloses a method for storing and retrieving RAG (Retrieve-And-Generate) based on a graph database. For ease of understanding and implementation, key concepts involved in the present invention are first explained, and then the implementation manners are described in detail.
[0033] Large Language Model (LLM): A large language model (LLM) is a model formed by training a large amount of natural language data, capable of understanding, generating, and processing text information. Typical large language models such as GPT (Generative Pre-trained Transformer), including ChatGPT, etc., can generate high-quality text summaries, answer questions, and perform content creation. In the present invention, the large language model is used to extract document summaries, generate table of contents structures, and process text segmentation.
[0034] RAG (Retrieve-And-Generate): RAG is a technology that combines information retrieval and text generation, capable of generating relevant text content based on retrieved documents. By combining the information retrieved with a large language model, the RAG method can generate more accurate answers when a user asks a question. The present invention utilizes RAG technology to achieve efficient document information retrieval and generation.
[0035] Graph database: A graph database is a type of database specifically designed to process graph data structures, capable of efficiently storing and querying nodes (entities) and their relationships (edges). The characteristics of a graph database are as follows:
[0036] 1. Efficiently handle complex relationships: Suitable for processing data with multi-level relationships, such as the structures of document chapters, paragraphs, keywords, etc., and the relationships between documents.
[0037] 2. Flexible query methods: Through graph traversal, it can quickly query relevant nodes and their relationships.
[0038] In the present invention, the graph database is used to store and organize keywords, file names, summaries, heat coefficient, chapters, and segmented content.
[0039] Semantic vector model: A semantic vector model is a model that converts text or other data into vectors (usually points in a high-dimensional space). In text processing, vectorization can map semantically similar texts to similar vectors. Cosine similarity is commonly used to calculate the similarity between two vectors. In the present invention, the semantic vector model is used for the processing of keywords, document summaries, and chapter content to help achieve efficient similarity calculation and matching.
[0040] Abstract: The abstract is a brief summary of the document content, which can condense the main information of the document. The abstract generated by the large language model contains the core ideas of the document, facilitating subsequent keyword extraction and content retrieval.
[0041] Popularity Coefficient: The popularity coefficient is used to measure the degree of user attention to a certain document within a period of time. By counting the content related to the document in the questions asked by users, the popularity value of the document can be calculated. The popularity coefficient is used to optimize the sorting of retrieval results, so that the content that users care more about is presented first.
[0042] Segmented Content (Chunk): Segmented content is to split long text or chapter content into smaller, semantically independent blocks. These blocks can more precisely match the user's query content and improve the accuracy of retrieval. In the present invention, Chunk is generated by the large language model and stored and retrieved through the semantic vector model.
[0043] As Figure 1 shown, this embodiment provides a RAG storage method based on a graph database, including the following steps:
[0044] S1: Document Data Collection
[0045] Collect different types of document data, supporting multiple formats (such as Word, PDF). The integrity and diversity of document data are the basis of the retrieval method of the present invention. In practical applications, data collection can be carried out through file system scanning, batch import, or automatic crawling, etc.
[0046] S2: Document Abstract Extraction
[0047] Use the large language model to generate the full-text abstract of the document. The core of abstract generation is to extract the main content of the document to form a highly summarized text, ensuring comprehensive coverage and accurate description of the content, facilitating subsequent keyword extraction and information retrieval. Abstract generation can be achieved by calling the LLM model through the API. For example, use the API of Tongyi Qianwen to generate the abstract output of the document.
[0048] S3: Keyword Extraction
[0049] Extract keywords from the full-text abstract of the document. First, construct a preset keyword library: K i = [k0, k1, k2,... k n ], k n represents the nth keyword; and use the semantic vector model (such as BGE-Large) to convert the keyword V k and the full-text abstract V f into vectors: V k = [v0, v1, v2,... v n ], vn represents the nth keyword vector. By calculating the cosine similarity S between the abstract vector and the keyword vectors i , sort the cosine similarities from high to low, and filter out the most relevant keywords with similarities greater than the threshold. The extraction of keywords is achieved by sorting and selecting the top 3 keywords, which are stored as tags of the document in the graph database.
[0050] S4: Heat coefficient setting
[0051] By calculating the ratio of the number of query results to the number of document-related questions within a specific time period for a user, determine the degree of attention of the document. The higher the heat coefficient, the higher the user's interest in the document. A schematic diagram of the heat coefficient is as shown Figure 4 . The heat coefficient of the document reflects the degree of attention of users to a specific document, and the calculation method is as follows:
[0052] S41: Association statistics between user questions and documents: Within a specific time period, count the number of questions related to the content of a certain document among the answers to the questions asked by the user (referred to as Num hit ), and at the same time count the total number of questions asked by all users during this time period (referred to as Num total ).
[0053] S42: Heat value calculation: Calculate the heat value f of the document according to Num hit and Num total :
[0054]
[0055] S43: Heat coefficient calculation: Calculate the heat coefficient H of the document through the following formula:
[0056]
[0057] where tanh is the hyperbolic tangent function, used to normalize the heat value so that its range is between 0 and 1. The calculation of the heat coefficient makes the more concerned documents have higher priorities in the retrieval results.
[0058] S5: Extraction of chapter and section directory structure
[0059] Extract the table of contents and chapter structure from the document through a large language model. In actual implementation, use the LLM to parse the table of contents information of the document (such as chapter titles and hierarchical structures), and generate the corresponding tree structure. The output of this step is the hierarchical directory structure diagram of the document, which is convenient for subsequent analysis and retrieval of chapter content.
[0060] S6: Chapter segmentation processing
[0061] According to the chapter directory structure extracted in S5, segment the content of each chapter of the document. The specific operations are as follows:
[0062] S61: Segmenting chapter content: Use a large language model to perform semantic segmentation on the content of each chapter to generate multiple segmented content Chunks. These Chunks are semantically independent text segments, facilitating subsequent fine-grained retrieval.
[0063] S62: Vectorizing Chunks: Use a semantic vector model to convert each Chunk into a vector and store it in a graph database. The vectorized Chunks can efficiently perform similarity matching during user retrieval.
[0064] Specifically, as Figure 3 shown, the construction and use of the graph database are the key to the present invention. The graph database is a data structure centered on nodes (entities) and edges (relationships). In the present invention, a storage method with multiple layers of nodes is designed. The graph database is constructed through multiple layers of nodes. The top layer is keywords, followed by file names, full text summaries, heat coefficients, chapter structures, and finally segmented content. Clear hierarchical relationships are formed between the nodes, making the storage and retrieval of document information more efficient and accurate. The nodes and their hierarchical structures in the graph database are as follows:
[0065]
[0066] The relationships in the graph database describe the associations between entities. Each relationship represents a hierarchical link between entities. The entity relationships included in this example are:
[0067] HAS_FILES: The relationship between keywords and file names, indicating that a certain file name is associated with one or more keywords, and one keyword may be associated with multiple file names.
[0068] HAS_SUMMARY: The relationship between file names and full text summaries, indicating that a certain file has a corresponding full text summary.
[0069] HAS_SECTION: The relationship between file names and chapter directories, indicating that a certain file contains one or more chapters.
[0070] HAS_HEAT: The relationship between file names and heat coefficients, indicating that a certain file has a corresponding heat coefficient.
[0071] HAS_PARAGRAPH: The relationship between chapter directories and segmented content, indicating that a certain chapter is further divided into multiple paragraphs.
[0072] All text data in the graph database, including keywords, file names, full-text abstracts, chapters, paragraphs, etc., are stored after being vectorized, and the popularity coefficient is stored as floating-point data. The graph database used can be Neo4j, etc. Through this structured storage method, the present invention realizes the effective management and efficient retrieval of large-scale text data.
[0073] The present invention not only performs structured processing on data storage, but also optimizes the retrieval process, making the retrieval results more in line with the actual needs of users. This embodiment provides a RAG retrieval method based on a graph database, as Figure 2 shown, including the following steps:
[0074] S1: Keyword retrieval.
[0075] After receiving the user's question, first rewrite the question to eliminate the ambiguity in the original question, optimize the question into a format that the system can better understand, and improve the accuracy of relevant retrieval results. This step is to improve the retrieval accuracy, ensure that the expression of the question can fully capture the user's intention, and better match the keywords and content in the document. In the retrieval system, the questions raised by users usually lead to deviations in retrieval results due to different expressions. Natural language questions may contain problems such as grammar errors, ambiguities, redundant information, etc., which will affect the system's understanding of the user's intention, thus affecting the accuracy and relevance of the retrieval. Therefore, by rewriting the user's question and standardizing or optimizing it into an easy-to-process form, it helps to improve the quality of the retrieval results. The rewritten question is more in line with the input requirements of the system, can more accurately match the document content, and improve the precision and recall rate of the system.
[0076] Convert the rewritten user question into a vector using a semantic vector model, and use the cosine similarity method to match all keywords in the keyword library to calculate the similarity. Sort according to the similarity from high to low, set a similarity threshold, and select the keywords with similarity higher than the threshold from them to form a candidate keyword set.
[0077] Question rewriting can adopt the basic method of Query Rewrite, such as:
[0078] Synonym replacement: Replace the words in the user's question with words with the same or similar meanings to ensure the extensiveness of the query. For example, replace "obtain" with "acquire".
[0079] Context completion: Complete the incomplete query according to the background input by the user. For example, the user may only enter a keyword "meteorological data", and when rewriting, it can be supplemented as "How to obtain meteorological data API".
[0080] Disambiguation: Handling polysemous words, determining their meanings in specific contexts, and avoiding ambiguity. For example, "apple" can refer to a fruit or a technology company. When rewriting, choose the appropriate meaning according to the context.
[0081] Simplify or clarify the query: Simplify complex or lengthy queries, remove unnecessary information, or clarify vague expressions. For example, change "I want to know the process of how to do project budgeting" to "Project budgeting process".
[0082] Examples of Query Rewrite are illustrated as follows:
[0083] 1) User's original question:
[0084] "How to optimize code performance?"
[0085] Query Rewrite:
[0086] "What are the effective methods and best practices for code performance optimization?"
[0087] Explanation: The original question is too broad. After rewriting, it is more specific, guiding the search results to focus on "effective methods" and "best practices".
[0088] 2) User's original question:
[0089] "What content needs to be written in the project application form?"
[0090] Query Rewrite:
[0091] "What are the key parts of the standard structure of the project application form?"
[0092] Explanation: By rewriting, it is clear that the question is about the "standard structure" and "key parts", which helps the system to return the specific writing framework of the project application form targeted.
[0093] S2: File retrieval.
[0094] For each keyword in the candidate keyword set, traverse the file names associated with the keyword, calculate the file similarity based on the file name, full text abstract, and popularity coefficient, and sort the similarities from high to low. Set a similarity threshold and retain the file names with similarities greater than the threshold as the file set. The calculation of file similarity takes into account the similarity between the user's question and the file name, full text abstract, and the popularity coefficient of the document, ensuring that the search results are both relevant and in line with the user's attention. As described in the following formula, first calculate the similarity S based on the file name and the full text abstract to obtain S temp, c0, c1 are coefficients, ranging from 0 to 1, and c0+c1=1, H is the heat coefficient formula. Then the final file similarity S is calculated according to the heat coefficient formula f .
[0095]
[0096]
[0097] S3: Chapter retrieval.
[0098] Based on each file in the final file collection, the similarity between the user's question and the file chapter is calculated. The system will retain the chapter with the highest similarity as the basis for subsequent retrieval. Chapter matching is an important step in the layer-by-layer in-depth retrieval, ensuring that the system can accurately locate the chapter content in the document that is most relevant to the user's question.
[0099] S4: Segment retrieval.
[0100] In the matched chapters, the similarity between the user's question and each segment content is further calculated. A similarity threshold is set, and the paragraph content with a similarity higher than the threshold is retained as the final retrieval result. This step ultimately determines whether the answer provided by the system to the user is specific and accurate enough.
[0101] S5: Integrated output
[0102] The question is integrated with the matched paragraph content to form a prompt word, which is provided as input to the large language model LLM. LLM generates the final answer based on these contents, generates the final text response or other processing results, and returns them to the user.
[0103] Corresponding to the aforementioned embodiment of a RAG storage and retrieval method based on a graph database, the present invention also provides an embodiment of a RAG storage and retrieval device based on a graph database.
[0104] See also Figure 5 An embodiment of the present invention provides a RAG storage and retrieval device based on a graph database, including a memory and one or more processors, wherein the memory stores executable code, and when the processor executes the executable code, it is used to implement a RAG storage and retrieval method based on a graph database in the above embodiment.
[0105] An embodiment of a RAG storage and retrieval device based on a graph database provided by the present invention can be applied to any device with data processing capabilities, and such a device with data processing capabilities can be a device or apparatus such as a computer. The device embodiment can be implemented through software, or through hardware or a combination of software and hardware. Taking software implementation as an example, as a logically meaningful device, it is formed by the processor of any device with data processing capabilities reading the corresponding computer program instructions in the non-volatile memory into the memory for operation. At the hardware level, as Figure 5 shown, it is a hardware structure diagram of any device with data processing capabilities where a RAG storage and retrieval device based on a graph database provided by the present invention is located. In addition to Figure 5 the processor, memory, network interface, and non-volatile memory shown, generally according to the actual functions of any device with data processing capabilities where the device in the embodiment is located, other hardware may also be included, which will not be elaborated here.
[0106] For the realization process of the functions and roles of each unit in the above device, please refer to the realization process of the corresponding steps in the above method for details, which will not be elaborated here.
[0107] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to the partial description of the method embodiment. The device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present invention. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0108] The embodiment of the present invention also provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements a RAG storage and retrieval method in the above embodiment.
[0109] The computer-readable storage medium may be an internal storage unit of any device with data processing capabilities described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium may also be an external storage device of any device with data processing capabilities, such as a plug-in hard disk, a Smart Media Card (SMC), an SD card, a Flash Card, etc. equipped on the device. Further, the computer-readable storage medium may also include both an internal storage unit and an external storage device of any device with data processing capabilities. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and may also be used to temporarily store the data that has been output or is to be output.
[0110] The present invention also provides a computer program product, including a computer program, which when executed by a processor, implements the RAG storage and retrieval method based on a graph database described above.
[0111] The above embodiments are used to explain the present invention, rather than to limit the present invention. Any modifications and changes made to the present invention within the spirit and scope of the claims of the present invention fall within the protection scope of the present invention.
Claims
1. A RAG storage and retrieval method based on a graph database, characterized in that: It includes two steps: storage and retrieval. The storage steps include using an LLM to extract the chapter and section directory structure of a document to obtain the hierarchical structure of the document; according to the chapter and section directory structure of the document, for each chapter content, perform segmentation processing, use the LLM to cut each chapter into multiple segmented content chunks, and convert them into vectors using a semantic vector model; extract the full text summary and keywords of the document data, as well as the chapter and section directory and perform segmentation processing, and set the popularity coefficient of the document, and store it in a graph database. The keywords, file names, full text summary, chapters, and segmented content in the database are all converted into vectors using a semantic vector model and then stored, and the popularity coefficient is stored using floating-point data; among them, according to the content attributes of all documents, a keyword library is preset; and for each keyword, use a semantic vector model to convert the keyword and the full text summary of the document into vectors; for each file, calculate the cosine similarity between the keyword vector and the summary vector, sort the cosine similarities from high to low, retain the keywords with similarities greater than the threshold, and take the top 3 keywords with the highest similarities as the keywords of the document; associate the keywords with the file names, associate the file names with the full text summary, chapter and section directory, and popularity coefficient, and associate the chapter and section directory with the segmented content for subsequent retrieval processes; the popularity coefficient of the document refers to the heat value of the number of questions whose answers are in the document among all the questions asked by all users within a period of time, Num hit Num is the number of questions whose answers are in the document among all the questions asked by users within a period of time total Num is the total number of questions asked by users within a period of time. The popularity coefficient H(f) of the document is obtained as follows. The higher the popularity coefficient, the more concerned the users are about the content of the document and the more questions they ask, and vice versa; (1) (2) Among them, tanh is the hyperbolic tangent function, which is used to normalize the heat value so that its range is between 0 and 1. The heat coefficient makes the more concerned documents have higher priority in the retrieval results. The retrieval step includes rewriting the user's question and matching it with keywords to obtain the associated file names, calculating the file similarity based on the file names, full-text abstracts, and heat coefficients, screening the files that meet the requirements, and then obtaining the chapter directory and segmented content that match the user's question. Finally, a final answer is generated based on the large language model and returned to the user. Specifically: convert the user's question into a vector using a semantic vector model, match it with each keyword in the keyword library, retain the keywords with a similarity greater than the threshold as the candidate keyword set, traverse the file names associated with the keywords, calculate the file similarity, and retain the file names with a similarity greater than the threshold as the file set; for each file in the file set, calculate the similarity between the user's question and the chapters, and obtain the chapter with the highest similarity; calculate the similarity between the user's question and the segmented content, and retain the segmented content with a similarity greater than the threshold as the content finally retrieved. The file similarity is calculated from the file name, full-text abstract, and heat coefficient. The vectorized user question is used to calculate the similarity with the vectorized file name and full-text abstract respectively, and the final file similarity The calculation formula is as follows: (3) (4) Among them, S name is the similarity between the rewritten user question and the file name, S summary is the similarity between the rewritten user question and the full text summary, c0 and c1 are coefficients, ranging from 0 to 1, and c0+c1=1, and H is the heat coefficient.
2. A RAG storage and retrieval device based on a graph database, comprising a memory and one or more processors, wherein executable code is stored in the memory, characterized in that, When the processor executes the executable code, it implements a method for RAG storage and retrieval based on a graph database as described in claim 1.
3. A computer-readable storage medium having a program stored thereon, characterized in that, When the program is executed by the processor, it implements a method for RAG storage and retrieval based on a graph database as described in claim 1.
4. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements a method for RAG storage and retrieval based on a graph database as described in claim 1.
Citation Information
Patent Citations
Question and answer method, computing equipment and storage medium
CN112417126A
Method and system for enhancing RAG questions and answers through mixed retrieval method
CN118627625A
Universal adaptive question answering method and system, storage medium and electronic equipment
CN118643144A