Cypher-stack type alignment generation method and device based on large language model
Through the semantic analytical ability and thinking chain guidance technology of the large language model, Cypher query statements are automatically generated and corrected, which solves the semantic gap between user query requirements and knowledge graphs, and realizes efficient, precise alignment and semantic consistency of Cypher query.
Patent Information
- Application Number
- CN202510165404.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-06-06
AI Technical Summary
The semantic gap between the diversity of user query requirements and the knowledge graph makes it difficult to accurately align user query intentions and knowledge graph data when generating Cypher queries.
The Cypher one-stack alignment generation method based on the large language model is adopted, and the Cypher query statement is automatically generated and corrected through a two-stage optimization strategy. The first stage uses the semantic analytic ability of the large language model to automatically map the pattern information of the knowledge graph with user query requirements into the basic framework of Cypher query. The second stage uses the entity alignment technology guided by thinking chains to deeply understand the context through a large language model, accurately match the entities and attributes in the query statement, and realize efficient and accurate alignment between user queries and knowledge graph data.
It effectively solves the problem that it is difficult to accurately align user query intentions and knowledge graph data when generating Cypher query, improves the semantic consistency and executability of query statements, and ensures that the generated Cypher query is accurately aligned with the data structure of the knowledge graph.
Smart Images

Figure CN120104110A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of knowledge graph question answering technology, and in particular to a Cypher one-stack alignment generation method and device based on a large language model. Background Art
[0002] Large Language Models (LLMs) have become a core technology in the field of natural language processing and artificial intelligence due to their excellent generation capabilities and wide range of application scenarios. These models have acquired powerful language understanding and generation capabilities by pre-training on large-scale corpora. However, despite the excellent performance of LLMs in multiple tasks, they still have limitations in knowledge absorption and updating, especially when dealing with domain-specific structured knowledge. To make up for this shortcoming, researchers have explored various ways to combine knowledge graphs (KGs) with LLMs. Knowledge graphs store a large amount of structured knowledge in the form of triples (head entity, relationship, tail entity), providing an efficient knowledge representation method that can provide LLMs with rich structured information and make up for the lack of knowledge updating.
[0003] The Knowledge Graph Question Answering (KGQA) system aims to retrieve accurate answers based on the user's natural language questions from a knowledge graph that contains a large amount of structured knowledge. The KGQA system not only promotes efficient interaction between people and knowledge graph databases, but also significantly lowers the threshold for using knowledge graphs. With the advent of the information explosion era, how to quickly and accurately extract the required knowledge from massive data has become a key issue that needs to be solved urgently. The KGQA system reduces the time and cost of information retrieval by accelerating the response to user questions, and has broad application prospects.
[0004] At present, the research methods for KGQA that combine knowledge graphs and large language models can be mainly divided into the following two categories:
[0005] (1) Information retrieval methods: This type of method focuses on retrieving relevant triples from the knowledge graph as context, and generates answers in combination with the reasoning ability of the large language model. Traditional methods use neural networks to identify key entities in the query and extract candidate answers through connections with the knowledge graph. In order to improve the interpretability and performance of reasoning, many studies have attempted to use the reasoning ability of large language models to collaborate with knowledge graphs to gradually retrieve and generate reliable evidence subgraphs. For example, the paper "Think-on-graph: Deep and responsible reasoning of large language model on knowledge graph" (Jiashuo Sun et al., In The Twelfth International Conference on Learning Representations, 2024) designed an information interaction mechanism between a large language model and a knowledge graph, and gradually found solutions to the problem through iterative reasoning. MindMap combines multi-hop reasoning and multi-path exploration to guide the large language model to generate a mind map, thereby significantly improving the accuracy of question and answering.
[0006] (2) Semantic parsing methods: This type of method converts natural language questions into logical query statements that can be executed in the knowledge graph. Compared with information retrieval methods, semantic parsing methods not only rely on factual knowledge, but also make more comprehensive use of the structural information of the knowledge graph to improve the accuracy of reasoning. For example, the paper "Few-shot in-context learning for knowledge base question answering" (Tianle Li et al., In Annual Meeting of the Association for Computational Linguistics, 2023) generates a query framework through a large language model and fills it in with the information provided by the knowledge graph, thereby improving the efficiency and accuracy of query generation. However, the challenge faced by semantic parsing methods is to deal with complex grammatical and semantic problems. The generated query statements are sometimes difficult to execute, especially when multiple entities and relationship links are involved. The accuracy is low, and any error in an entity or relationship will cause the query language to be unable to execute.
[0007] In recent years, semantic parsing methods based on large language models have made significant progress in the field of KGQA, especially in translating natural language into the Cypher query language for Neo4j database. Cypher is a query language designed for querying graph structures, and has significant performance advantages over other query languages (such as SPARQL) when processing graph data. However, due to the diversity of user query requirements and the semantic gap between knowledge graphs, it is difficult to accurately align user query intent with knowledge graph data when generating Cypher queries. Summary of the invention
[0008] The technical problem to be solved by the present invention is how to solve the problem that it is difficult to accurately align user query intentions with knowledge graph data when generating Cypher queries due to the semantic gap between the diversity of user query requirements and the knowledge graph.
[0009] The present invention solves the above technical problems through the following technical solutions: a Cypher one-stack alignment generation method based on a large language model, the method comprising:
[0010] S1. Input the user query and the pattern information of the knowledge graph into the large language model to generate the initial Cypher query statement;
[0011] S2. Based on the thought chain prompt technology, the search results related to the user's query are retrieved from the knowledge graph, and the large language model modifies the initial Cypher query statement based on the search results to obtain the modified Cypher query statement;
[0012] S3. Execute the modified Cypher query statement.
[0013] The present invention automatically generates and corrects Cypher query statements in two stages through the semantic parsing ability of the large language model and the thinking chain guidance technology. In the first stage, the semantic parsing ability of the large language model is fully utilized to automatically map the pattern information of the knowledge graph and the user query requirements into the basic framework of the Cypher query, thereby preliminarily building a query structure that meets the user's intention. In the second stage, the entity alignment technology guided by the thinking chain is used to accurately match the entities and their attributes in the query statement through the in-depth understanding of the context by the large language model, so as to achieve efficient and accurate alignment between the user query and the knowledge graph data, and effectively solve the problem that it is difficult to accurately align the user query intention and the knowledge graph data when generating Cypher queries due to the diversity of user query requirements and the semantic gap between the knowledge graph.
[0014] Preferably, the search results include entity and attribute values and their category information.
[0015] Preferably, the search results are retrieved using the Okapi BM25 algorithm, and the calculation formula of the Okapi BM25 algorithm is:
[0016]
[0017] Among them, Score(D, Q) is the relevance score of document D to user query Q, N is the total number of documents in the document collection, and q i is the i-th word in the user's query, n(q i ) is a string containing the word q i The number of documents, f(q i , D) is the word q i The frequency in document D, IDF(q i ) is the word q i is the inverse document frequency, |D| is the length of document D, avgdl is the average length of all documents, and k 1 and b are adjustable parameters.
[0018] Preferably, the retrieval results also include few-shot examples, which use the Okapi BM25 algorithm to search the constructed question-answering dataset, screen out the most relevant N questions and their corresponding Cypher query statements, and use the few-shot learning technology to input these N questions and their corresponding Cypher query statements as examples into the large language model.
[0019] Preferably, the large language model modifies the initial Cypher query statement based on the retrieval results, which means that the large language model replaces or adjusts the key fields in the initial Cypher query statement according to the retrieval results related to the user query retrieved from the knowledge graph.
[0020] Preferably, the method also includes: based on the first k entities in the retrieval results, extracting the associated information within one hop of the k entities from the knowledge graph to obtain subgraph information; when the Cypher query statement cannot be executed, using the subgraph information as the knowledge context and combining it with the user query input large language model to generate query results.
[0021] The present invention also provides a Cypher stack alignment generation device based on a large language model, the device comprising:
[0022] A query statement generation unit, used to input the user query and the pattern information of the knowledge graph into the large language model to generate an initial Cypher query statement;
[0023] A query statement correction unit is used to retrieve the search results related to the user query from the knowledge graph based on the thought chain prompt technology, and the large language model corrects the initial Cypher query statement based on the search results to obtain a corrected Cypher query statement;
[0024] The execution unit is used to execute the modified Cypher query statement.
[0025] Preferably, the search results include entity and attribute values and their category information.
[0026] Preferably, the search results are retrieved using the Okapi BM25 algorithm, and the calculation formula of the Okapi BM25 algorithm is:
[0027]
[0028] Among them, Score(D, Q) is the relevance score of document D to user query Q, N is the total number of documents in the document collection, and q i is the i-th word in the user's query, n(q i ) is a string containing the word q i The number of documents, f(q i , D) is the word q i The frequency in document D, IDF(q i ) is the word q i is the inverse document frequency, |D| is the length of document D, avgdl is the average length of all documents, and k 1 and b are adjustable parameters.
[0029] Preferably, the retrieval results also include few-shot examples, which use the Okapi BM25 algorithm to search the constructed question-answering dataset, screen out the most relevant N questions and their corresponding Cypher query statements, and use the few-shot learning technology to input these N questions and their corresponding Cypher query statements as examples into the large language model.
[0030] Preferably, the large language model modifies the initial Cypher query statement based on the retrieval results, which means that the large language model replaces or adjusts the key fields in the initial Cypher query statement according to the retrieval results related to the user query retrieved from the knowledge graph.
[0031] Preferably, the device also includes: a subgraph information extraction unit, which is used to extract the associated information within one hop of k entities from the knowledge graph based on the first k entities in the retrieval results to obtain subgraph information. When the Cypher query statement cannot be executed, the subgraph information is used as the knowledge context and combined with the user query input large language model to generate a query result.
[0032] The advantages provided by the present invention are:
[0033] (1) The present invention automatically generates and corrects Cypher query statements in two stages through the semantic parsing ability of the large language model and the thought chain guidance technology. In the first stage, the semantic parsing ability of the large language model is fully utilized to automatically map the pattern information of the knowledge graph and the user query requirements into the basic framework of the Cypher query, thereby preliminarily building a query structure that meets the user's intention. In the second stage, the entity alignment technology guided by the thought chain is used to accurately match the entities and their attributes in the query statement through the large language model's in-depth understanding of the context, so as to achieve efficient and accurate alignment between user queries and knowledge graph data, and effectively solve the problem that it is difficult to accurately align user query intentions with knowledge graph data when generating Cypher queries due to the diversity of user query requirements and the semantic gap between knowledge graphs.
[0034] (2) The present invention combines thought chain prompts and adopts the Okapi BM25 algorithm to retrieve entities and attributes, ensuring that the Cypher query generated by the large language model is accurately aligned with the data structure of the knowledge graph, thereby optimizing the semantic consistency and executability of the query statement.
[0035] (3) The retrieval results of the present invention also include few-shot examples. The few-shot examples use the Okapi BM25 algorithm to search the constructed question-answering dataset. By retrieving similar questions and their Cypher query statements from the graph database and using few-shot learning technology to help the large language model better understand the intrinsic relationship between questions, graph patterns and query statements, the accuracy and robustness of query generation are further improved.
[0036] (4) The present invention obtains subgraph information by extracting the associated information within one hop of k entities from the knowledge graph. When the generated Cypher query statement cannot be executed due to grammatical or semantic problems, the relevant subgraph information is extracted and the query generation process is optimized in combination with the problem context, thereby improving the overall stability of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 A flowchart of a Cypher stack alignment generation method based on a large language model provided by an embodiment of the present invention;
[0038] Figure 2 A schematic diagram of a Cypher stack alignment generation method based on a large language model provided in an embodiment of the present invention;
[0039] Figure 3A diagram showing an example of the practical application of the Cypher stack alignment generation method based on a large language model provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0040] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the technical solution of the present invention is clearly and completely described below in combination with specific embodiments and with reference to the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0041] Example 1
[0042] like Figure 1 and Figure 2 As shown, this embodiment provides a Cypher one-stack alignment generation method based on a large language model, comprising the following steps:
[0043] S1. Input the user query and the pattern information of the knowledge graph into the large language model to generate the initial Cypher query statement.
[0044] S2. Based on the thought chain prompt technology, the search results related to the user query are retrieved from the knowledge graph. The large language model modifies the initial Cypher query statement based on the search results to obtain the modified Cypher query statement.
[0045] S3. Execute the modified Cypher query statement. If the execution is successful, the system will combine the search results with the user's question and input it into the large language model to generate the final query result.
[0046] First, a prompt word template based on the thinking chain prompt technology is designed. The user query and the pattern information of the knowledge graph (such as node type, relationship attribute, etc.) are integrated into the prompt word template and passed to the large language model as input. The large language model gradually generates the initial Cypher query statement according to the prompt word template. By using the powerful semantic parsing ability of the large language model, combined with the user query Q and the node attributes, relationship type and attribute information S in the graph database, the preliminary logical form F of the natural language question is generated. The logical form is a structured representation of the natural language question, in which the entity E={e 1 , e 2 ,…,e m} and the relation R = {r 1 , r 2 ,…,r n} is directly generated by a large language model, but the initially generated query statements may be inconsistent with the actual data in the knowledge graph and may not be executable.
[0047] The retrieval results include entity and attribute values and their category information. The Okapi BM25 algorithm is used for information retrieval. Assume that the retrieval result is w = {(e' 1 , p' 1 ), (e′ 2 , p′ 2 ),…,(e′ k , p′ k )}, where e represents entity and p represents attribute. These results will be used for data alignment generated by subsequent Cypher queries. The BM25 algorithm is used to evaluate the relevance of document D to user query Q. Its calculation formula is as follows:
[0048]
[0049] Among them, Score(D, Q) is the relevance score of document D to user query Q, N is the total number of documents in the document collection, and q i is the i-th word in the query, n(q i ) is a string containing the word q i The number of documents, f(q i , D) is the word q i The frequency in document D, IDF(q i ),Right now It is the word q i is the inverse document frequency, |D| is the length of document D, avgdl is the average length of all documents, and k 1 and b are adjustable parameters, k 1 Typically between 1.2 and 2, b is usually set to 0.75.
[0050] In this application, k 1 is set to 1.5, b is set to 0.75, and k is set to 5, which means that the top 5 entities and attribute information most relevant to the question are retrieved from the graph database.
[0051] The present invention combines the thinking chain prompt technology to clearly inform the large language model in the prompt that the initially generated Cypher statement may have entity or relationship names that do not match the actual data in the graph database. The large language model needs to replace or adjust the key fields in the initially generated Cypher query statement based on the retrieval results most relevant to the problem entity retrieved from the graph database, including the entity name, attribute name and its category information w. This correction process combines the entity and attribute values in the actual data to ensure that the generated Cypher query statement can accurately reflect the data relationship in the knowledge graph.
[0052] The retrieval results also include few-shot examples. The few-shot examples use the Okapi BM25 algorithm to retrieve the constructed question-answering dataset, screen out the most relevant N questions and their corresponding Cypher query statements, and use the few-sample learning technology to input these N questions and their corresponding Cypher query statements as examples into the large language model. The Okapi BM25 algorithm calculates the relevance score between the user query Q and each question in the dataset to screen out the most relevant N questions and their corresponding Cypher query statements. Subsequently, the system uses the few-sample learning technology to input these N questions and their corresponding Cypher query statements as examples into the large language model. In this way, the large language model can learn the intrinsic relationship between user questions, graph database patterns, and Cypher queries from these examples, thereby generating more precise Cypher query statements and improving the accuracy of queries.
[0053] The present invention takes into account that the generated Cypher query statement may not be executed due to grammatical or semantic problems. The method also includes: based on the first k entities in the retrieval results (k=1 is set in the experiment), the associated information within the one-hop range of the k entities is extracted from the graph database to obtain subgraph information. When the Cypher query statement cannot be executed, the subgraph information is used as the knowledge context and combined with the user query content to input the large language model to generate the query result. In order to avoid information overload, the experiment sets the extraction of the first 3 pieces of associated information as the knowledge context to help the system generate more accurate query results.
[0054] The present invention aims to solve the semantic gap problem faced by traditional vector similarity matching methods in processing user queries in current specific fields (such as colleges and universities). Traditional methods usually convert user queries and entities and attributes in knowledge graphs into vector representations, and then filter the most relevant entities by calculating cosine similarity. This approach easily leads to incorrect matching when entities or attributes are semantically different but their vector representations are similar, thereby affecting the accuracy and reliability of retrieval results.
[0055] To address this limitation, the present invention proposes a Cypher one-stack alignment generation method based on a large language model, which combines a thought chain prompt word template with a two-stage optimization strategy. Specifically, in the first stage, the semantic parsing ability of the large language model is fully utilized to automatically map the pattern information of the knowledge graph and the user's query requirements into the basic framework of the Cypher query, thereby preliminarily constructing a query structure that meets the user's intention. In the second stage, the system uses the entity alignment technology guided by the thought chain, through the large language model's in-depth understanding of the context, to accurately match the entities and their attributes in the query statement, and achieve efficient and accurate alignment between user queries and knowledge graph data.
[0056] Example 2
[0057] The present invention also provides a Cypher stack alignment generation device based on a large language model, comprising:
[0058] A query statement generation unit, used to input the user query and the pattern information of the knowledge graph into the large language model to generate an initial Cypher query statement;
[0059] The query statement correction unit is used to retrieve the search results related to the user query from the knowledge graph based on the thought chain prompt technology. The large language model corrects the initial Cypher query statement based on the search results to obtain the corrected Cypher query statement; the search results include entity and attribute values and their category information. The search results use the Okapi BM25 algorithm for information retrieval. The calculation formula of the Okapi BM25 algorithm is:
[0060]
[0061] Among them, Score(D, Q) is the relevance score of document D to user query Q, N is the total number of documents in the document collection, and q i is the i-th word in the user's query, n(q i ) is a string containing the word q i The number of documents, f(q i , D) is the word q i The frequency in document D, IDF(q i ) is the word q i is the inverse document frequency, |D| is the length of document D, avgdl is the average length of all documents, and k 1 and b are adjustable parameters.
[0062] The retrieval results also include few-shot examples. The few-shot examples use the Okapi BM25 algorithm to search the constructed question-answering dataset, screen out the most relevant N questions and their corresponding Cypher query statements, and use the few-sample learning technology to input these N questions and their corresponding Cypher query statements as examples into the large language model.
[0063] The subgraph information extraction unit is used to extract the associated information within one hop of k entities from the graph database based on the first k entities in the retrieval results to obtain subgraph information. When the Cypher query statement cannot be executed, the subgraph information is used as the knowledge context and combined with the user query content to input the large language model to generate the query result.
[0064] An execution unit for executing the corrected Cypher query statement. If the execution is successful, the system combines the retrieval results with the user's question and inputs them into the large language model to generate the final query result. If the query execution fails due to syntax or semantic errors, the system uses the extracted subgraph information as the knowledge context, combines it with the user's question, and inputs it into the large language model to generate the final query result.
[0065] Actual application example
[0066] Suppose the user asks the following query: "Who are the Ph.D. supervisors in the School of Computer Science who research artificial intelligence?"
[0067] After the processing of steps 1 to 4, the system-generated prompt for generating Cypher statements based on the chain of thought is as Figure 3 shown:
[0068] The system generates the following Cypher query statement: MATCH(p:People)-[r:belong]-(d:Department) WHERE d.name = 'School of Computer Science and Technology' AND 'artificial intelligence' IN p.research_direction AND 'Ph.D. supervisor' IN p.mentor_type RETURN p.name
[0069] The query results are as follows: [{'p.name': 'Qin XX'}, {'p.name': 'Wang XX'}]
[0070] The answer returned by the large language model is: The Ph.D. supervisors in the School of Computer Science and Technology whose research direction is artificial intelligence are Qin XX and Wang XX.
[0071] In this experiment, a university knowledge graph was constructed based on texts in a specific domain of universities. This graph contains 2,433 nodes and 2,765 relationships. Based on this knowledge graph, the present invention designed a set of question-answering datasets for the university knowledge graph, which contains a total of 778 question-answering data. To evaluate the performance of the question-answering system, the present invention selects Exact Match and Hit@1 as evaluation metrics. In the experimental setup, the present invention selects the following baseline models and few-shot learning setups: GLM-4 (3-shot), GLM-4-9B-chat (5-shot), Qwen2-7B-Instruct (5-shot), Llama-3-8B-Instruct (5-shot). In addition, to comprehensively evaluate the effects of different methods, the present invention also uses the Langchain method as a comparative experiment. As shown in Table 1, the performance differences of each model in the question-answering task are further analyzed.
[0072] Table 1 Comparative experiment
[0073]
[0074] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A Cypher stack alignment generation method based on a large language model, characterized by: Methods include: S1. Input the user query and the pattern information of the knowledge graph into the large language model to generate the initial Cypher query statement; S2. Based on the thought chain prompt technology, the search results related to the user's query are retrieved from the knowledge graph, and the large language model modifies the initial Cypher query statement based on the search results to obtain the modified Cypher query statement; S3. Execute the modified Cypher query statement.
2. The Cypher stack alignment generation method based on a large language model according to claim 1, characterized in that: The retrieval results include entity and attribute values and their category information.
3. The Cypher stack alignment generation method based on a large language model according to claim 2, characterized in that: The search results are retrieved using the Okapi BM25 algorithm. The calculation formula of the Okapi BM25 algorithm is: Among them, Score(D,Q) is the relevance score of document D to user query Q, N is the total number of documents in the document collection, and q i is the i-th word in the user's query, n(q i ) is a string containing the word q i The number of documents, f(q i ,D) is the word q i The frequency in document D, IDF(q i ) is the word q i is the inverse document frequency of D, |D| is the length of document D, avgdl is the average length of all documents, k1 and b are adjustable parameters.
4. The Cypher-stack alignment generation method based on a large language model according to claim 1, characterized in that: The retrieval results also include few-shot examples. The few-shot examples use the Okapi BM25 algorithm to search the constructed question-answering dataset, screen out the most relevant N questions and their corresponding Cypher query statements, and use the few-sample learning technology to input these N questions and their corresponding Cypher query statements as examples into the large language model.
5. The Cypher-stack alignment generation method based on a large language model according to claim 1, characterized in that: The large language model modifies the initial Cypher query statement based on the retrieval results, which means that the large language model replaces or adjusts the key fields in the initial Cypher query statement based on the retrieval results related to the user query retrieved from the knowledge graph.
6. The Cypher stack alignment generation method based on a large language model according to claim 1, characterized in that: The method also includes: based on the first k entities in the retrieval results, extracting the associated information within a one-hop range of the k entities from the knowledge graph to obtain subgraph information; when the Cypher query statement cannot be executed, using the subgraph information as the knowledge context and combining it with the user query input large language model to generate query results.
7. A Cypher stack alignment generation device based on a large language model, characterized in that: The device includes: A query statement generation unit, used to input the user query and the pattern information of the knowledge graph into the large language model to generate an initial Cypher query statement; A query statement correction unit is used to retrieve the search results related to the user query from the knowledge graph based on the thought chain prompt technology, and the large language model corrects the initial Cypher query statement based on the search results to obtain a corrected Cypher query statement; The execution unit is used to execute the modified Cypher query statement.
8. The Cypher stack alignment generation device based on a large language model according to claim 7, characterized in that: The retrieval results include entity and attribute values and their category information.
9. The Cypher stack alignment generation device based on a large language model according to claim 8, characterized in that: The search results are retrieved using the Okapi BM25 algorithm. The calculation formula of the Okapi BM25 algorithm is: Among them, Score(D,Q) is the relevance score of document D to user query Q, N is the total number of documents in the document collection, and q i is the i-th word in the user's query, n(q i ) is a string containing the word q i The number of documents, f(q i ,D) is the word q i The frequency in document D, IDF(q i ) is the word q i is the inverse document frequency of D, |D| is the length of document D, avgdl is the average length of all documents, k1 and b are adjustable parameters.
10. The Cypher stack alignment generation device based on a large language model according to claim 7, characterized in that: The retrieval results also include few-shot examples. The few-shot examples use the Okapi BM25 algorithm to search the constructed question-answering dataset, screen out the most relevant N questions and their corresponding Cypher query statements, and use the few-sample learning technology to input these N questions and their corresponding Cypher query statements as examples into the large language model.
Citation Information
Cited By
Retrieval enhancement generation method and device based on knowledge graph, equipment and medium
CN120277206A
Retrieval enhancement generation method, device, equipment and medium based on knowledge graph
CN120277206B
Data retrieval method, device and system
CN120316119A
Data retrieval method, device and system
CN120316119B
Natural language and Cypher query language conversion method, system and terminal
CN120541195A