Large language model Cypher generation method and system based on graph pattern alignment enhancement
By constructing various types of training datasets, performing supervised fine-tuning and sampling on large language models, and combining them with graph pattern selection modules to generate accurate Cypher queries, we address the complexity of graph database query languages and enable efficient data analysis for non-professional users.
Patent Information
- Application Number
- CN202510932436.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2025-10-17
AI Technical Summary
Traditional data processing methods make it difficult to efficiently analyze and manage complex, multi-level relationships and dynamically changing data. The syntax of the graph database query language Cypher is complex and difficult for non-professional users to master, which limits the application of graph databases in a wider range of fields.
A large language model Cypher generation method based on graph pattern alignment enhancement is adopted. By constructing various types of training datasets, supervised fine-tuning and sampling of the large language model are performed, the model capabilities are enhanced using a direct preference optimization algorithm, and Cypher queries are generated in combination with a graph pattern selection module.
The accuracy and efficiency of Cypher query generation by large language models have been improved. Users can efficiently query and analyze graph data without in-depth understanding of Cypher syntax, improving data management and analysis efficiency.
Smart Images

Figure CN120806154A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of graph databases and large models, and relates to a large language model Cypher generation method and system based on graph pattern alignment enhancement. BACKGROUND
[0002] In today's era of data explosion, the structural relationships between data are increasingly complex, and traditional data processing methods are insufficient to efficiently analyze and manage these data. With the rapid development of information technology, the scale and complexity of data are growing continuously, which brings great challenges to data management and analysis. Traditional data processing methods, such as relational databases and simple data analysis tools, often struggle to cope with such complex data structures, especially when dealing with data that has multiple levels of association and dynamic changes.
[0003] Graph data can intuitively represent complex entity relationships and has obvious advantages in various fields. The graph data model, with its powerful expression ability and flexibility, can naturally represent and process complex entity relationships. For example, a large part of network data is now stored in knowledge graphs, which not only store a large amount of entity information but also store complex relationships between entities. In social networks, graph-based representations reveal patterns of user and community interactions, helping researchers and businesses better understand user behavior and community dynamics. In addition, graph data also plays an important role in bioinformatics, financial risk analysis, intelligent transportation, and other fields, where data often has high correlation and dynamicity, and the graph data model can effectively capture these characteristics.
[0004] These advances collectively highlight the important role of graph-based methods in addressing the challenges posed by today's interconnected data environment. With the popularity of the Internet and the development of the Internet of Things, data is becoming increasingly interconnected, and the relationships between data are becoming increasingly complex. Graph-based methods can effectively handle such complex relationships, providing deeper data insights and more efficient data management solutions. However, despite the significant advantages of graph data models and graph databases in handling complex data, effectively utilizing them remains a formidable challenge.
[0005] Graph databases are commonly used for efficient processing of graph data, providing a powerful solution for representing and storing complex, highly interconnected information. Graph databases are capable of quickly processing and analyzing large-scale graph data through their flexible graph data model and efficient query capabilities. However, the syntax of the Cypher query language for graph databases is complex, which presents an obstacle for users seeking to convert natural language into Cypher, especially for those who are not familiar with programming paradigms. Although the Cypher language is powerful, its complex syntax and query structure make it difficult for non-specialist users to master and use. This not only limits the application of graph databases in a wider range of fields, but also hinders the efficient use of graph data by ordinary users.
[0006] Therefore, there is an urgent need for a new Cypher generation method that can automatically convert user natural language into Cypher to support efficient graph database queries. This new method needs to be able to understand user natural language requirements and accurately convert them into Cypher queries that graph databases can understand and execute. This not only requires the model to have strong natural language understanding capabilities, but also requires the model to accurately map natural language to the query language of graph databases. In this way, users do not need to deeply understand the complex syntax of Cypher, and can efficiently query and analyze graph data, thereby fully utilizing the advantages of graph databases and improving the efficiency of data management and analysis. SUMMARY
[0007] The present application provides a large language model Cypher generation method and system based on graph pattern alignment enhancement to solve the above-mentioned technical problems, which specifically adopts the following technical solutions:
[0008] A large language model Cypher generation method based on graph pattern alignment enhancement, comprising the following steps:
[0009] Constructing a dataset containing multiple types of training data;
[0010] Supervised fine-tuning of the large language model;
[0011] Sampling the large language model after supervised fine-tuning to obtain a sampling dataset;
[0012] Training the large language model after supervised fine-tuning;
[0013] Obtaining user input questions and graph database patterns;
[0014] Selecting relevant subgraph patterns from complete graph patterns according to user questions and verifying them;
[0015] The large language model generates corresponding Cypher queries using the selected patterns and given questions.
[0016] Further, the data set comprises Cypher generation data, Cypher correction data and graph pattern selection data;
[0017] The input of the Cypher generation data is a user query and a graph pattern, and the output is a corresponding Cypher query;
[0018] The input of the Cypher correction data is a user query, a graph pattern and a Cypher query that is uncertain whether correct, and the output is a correct Cypher query;
[0019] The input of the graph pattern selection data is a user query and a graph pattern, and the output is a subgraph pattern required to answer the user query.
[0020] Further, in the process of supervised fine-tuning of the large language model, higher weights are assigned to words related to edge direction to enhance the attention of the large language model to the relationship direction in the graph pattern.
[0021] Further, the supervised fine-tuning process is represented as:
[0022]
[0023] Wherein, Q represents a user question, S represents a graph pattern, Ct represents the tth word, represents the weight of the tth word Ct.
[0024] Further, the specific method for obtaining the sampling data set by sampling the large language model after supervised fine-tuning is:
[0025] Given a set of question-Cypher pairs, use the large language model trained by supervised fine-tuning to generate multiple Cypher queries, and then execute these queries in the graph database. The query that produces a result matching the standard answer is labeled as a positive sample, and the query that does not match the standard answer is labeled as a negative sample.
[0026] Further, in the step of training the large language model after supervised fine-tuning, a direct preference optimization algorithm is used to train the large language model.
[0027] Further, in the process of training the large language model using the direct preference optimization algorithm, higher weights are assigned to words related to edge direction.
[0028] Further, the direct preference optimization process is represented as:
[0029]
[0030] Wherein, Q represents a user question, S represents a graph pattern, Ct represents the tth word, The weight representing the t-th word Ct, Cw is a positive sample, and Cr is a negative sample.
[0031] Further, in the step of selecting a relevant subgraph pattern from the complete graph pattern according to the user question and verifying, if the selected subgraph pattern does not satisfy the constraint of the original graph pattern, an error is returned and the subgraph pattern is reselected.
[0032] A large language model Cypher generation system based on graph pattern alignment enhancement, comprising:
[0033] A data set construction module for constructing a data set, the data set comprising a plurality of types of training data;
[0034] A supervised fine-tuning module for supervising and fine-tuning a large language model;
[0035] A sampling module for sampling the large language model after supervised fine-tuning to obtain a sampling data set;
[0036] A training module for training the large language model after supervised fine-tuning;
[0037] An acquisition module for acquiring a user input question and a graph database pattern;
[0038] A selection and verification module for selecting a relevant subgraph pattern from a complete graph pattern according to a user question and verifying;
[0039] A Cypher generation module for generating a corresponding Cypher query based on the selected pattern and the given question through the trained large language model.
[0040] The large language model Cypher generation method and system based on graph pattern alignment enhancement provided by the application has the advantages that the large language model is better adapted to the Cypher generation task after supervised fine-tuning of multiple data, and the ability of the large language model to generate Cypher is enhanced through direct preference optimization algorithm, and the large language model and the graph pattern are effectively aligned.
[0041] The large language model Cypher generation method and system based on graph pattern alignment enhancement provided by the application also has the advantages that the graph pattern selection module is introduced, the relevant subgraph pattern is selected from the complete graph pattern according to the user question, the subgraph pattern is verified, and finally the Cypher is generated according to the subgraph pattern and the question, thereby improving the performance of the model in the Cypher generation task. BRIEF DESCRIPTION OF DRAWINGS
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the accompanying drawings needed to be used in the description of the embodiments or the prior art will be briefly introduced. Obviously, the accompanying drawings in the following description only show some embodiments of the present application, and for those skilled in the art, other drawings can be obtained based on these drawings without creative labor.
[0043] Figure 1 is a schematic diagram of the Cypher generation method of the large language model based on graph pattern alignment enhancement of the present application.
[0044] Figure 2 is a schematic diagram of direct preference optimization data acquisition of the present application.
[0045] Figure 3 is a schematic diagram of graph pattern selection of the present application. DETAILED DESCRIPTION
[0046] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference signs represent the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are intended to explain the present application, and cannot be understood as a limitation of the present application.
[0047] It can be understood that the accompanying drawings are only schematic diagrams of the present application, and are not necessarily drawn to scale. Some block diagrams shown in the accompanying drawings are functional entities, which do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0048] It can be understood that the flowcharts shown in the accompanying drawings are only exemplary descriptions, and do not necessarily include all steps. For example, some steps can be further decomposed, and some steps can be combined or partially combined, so the actual execution order can be changed according to the actual situation.
[0049] As Figure 1The application is a large language model Cypher generation method based on graph pattern alignment enhancement, including the following steps: S1: constructing a data set, the data set containing multiple types of training data. S2: supervised fine-tuning of the large language model. S3: sampling the large language model after supervised fine-tuning to obtain a sampling data set. S4: training the large language model after supervised fine-tuning. S5: obtaining user input questions and graph database patterns. S6: selecting relevant subgraph patterns from complete graph patterns according to user questions and verifying. S7: the large language model generates corresponding Cypher queries using the selected patterns and given questions.
[0050] The core idea of the application is to enhance the ability of the large language model to generate Cypher by aligning the large language model and the graph pattern. Through the above steps, the large language model Cypher generation method based on graph pattern alignment enhancement of the application makes the large language model better adapt to the Cypher generation task through supervised fine-tuning of multiple data, and effectively aligns the large language model and the graph pattern by sampling the model after supervised fine-tuning. The above method is specifically introduced as follows.
[0051] For step S1: constructing a data set, the data set containing multiple types of training data.
[0052] First, the data set of different graph patterns is obtained, including movie recommendation data, network management data, forum data, retail data, Twitter data, etc., each graph data having a different graph pattern. The Cypher related data set is collected and cleaned, and multiple data types are constructed for supervised fine-tuning of the large language model.
[0053] In the embodiment of the application, the data set contains Cypher generation data, Cypher correction data and graph pattern selection data. The input of the Cypher generation data is user query and graph pattern, and the output is the corresponding Cypher query. The input of the Cypher correction data is user query, graph pattern and Cypher query which is uncertain whether correct, and the output is the correct Cypher query. The input of the graph pattern selection data is user query and graph pattern, and the output is the subgraph pattern required to answer the user query.
[0054] The construction of the data set is the basis of the whole method, and the quality and diversity of the data directly affect the performance of the model. By collecting graph pattern data in different fields, it can be ensured that the model can perform well in various scenarios. Data cleaning and preprocessing is to remove noise and inconsistent data to improve data quality and thus improve the training effect of the model.
[0055] For step S2: supervised fine-tuning of the large language model.
[0056] Supervised fine-tuning of large language models on datasets containing multiple categories of data can effectively align the large language model with the graph schema.
[0057] Cypher generation data input is user query and graph schema, output is the correct Cypher query. This data type directly guides the model on how to generate the correct Cypher query based on the user question and graph schema. The model learns from a large number of user queries and corresponding Cypher query pairs, gradually mastering the mapping relationship from natural language to Cypher. Through a large number of training samples, the model can better understand the semantics of user queries and convert them into Cypher queries that conform to the graph database schema. This helps improve the model's understanding of user questions, resulting in more accurate queries. The model learns the Cypher query generation method under different graph schemas, better adapting to various graph database schemas, improving the generality and adaptability of generated queries.
[0058] Cypher correction data input is user query, graph schema and possibly incorrect Cypher query, output is the correct Cypher query. This data type helps the model identify and correct incorrect Cypher queries. By learning incorrect Cypher queries and their corresponding correct queries, the model can identify common error patterns and learn how to correct them. When faced with incorrect input, the model can more effectively identify and correct errors, improving the robustness and reliability of generated queries. By learning from incorrect Cypher queries, the model can avoid generating similar incorrect queries, improving the accuracy and quality of generated queries.
[0059] Graph schema selection data input is user query and graph schema, output is the subgraph schema required to answer the user query. This data type helps the model select the relevant part of the complex graph schema from the user question. The model learns from user queries and corresponding subgraph schemas, and can more accurately identify which graph schema components are helpful in answering user questions. By filtering out relevant subgraph schemas, the model can focus on graph schema components directly related to user questions, avoiding irrelevant information interference, thereby improving the accuracy of generated queries. The model learns how to select subgraph schemas, enabling it to better understand the structure and constraints of graph schemas, and thus better generate Cypher queries that conform to the graph schema.
[0060] These three data types train the model from different perspectives, allowing the model to consider the semantics of user questions, the structure and constraints of graph patterns, and error correction capabilities when generating Cypher queries. Through Cypher generation data, the model learns to generate correct queries; through Cypher error correction data, the model learns to avoid errors; and through graph pattern selection data, the model learns to filter relevant graph patterns. This multi-dimensional training method makes the model more accurate, reliable and efficient when generating Cypher queries. In practical applications, user questions and graph patterns are complex. Through the training of these three data types, the model can better handle various complex situations and generate Cypher queries that meet user needs and graph pattern constraints, thereby improving the efficiency and accuracy of graph database queries.
[0061] These three data types influence the model's generation of Cypher queries in different ways, allowing the model to more accurately understand user questions, more effectively filter relevant graph patterns, and more reliably generate correct queries while avoiding common errors. These measures work together to significantly improve the model's performance in automatically generating Cypher queries for graph databases.
[0062] In this application, to clearly emphasize the direction of the relationship, higher weights are assigned to words related to edge direction during the supervised fine-tuning process. By increasing the loss weight of these words, the model is explicitly trained to pay more attention to edge direction during the generation process, thereby reducing the likelihood of direction errors and ensuring better compliance with graph pattern constraints. The supervised fine-tuning process can be represented as:
[0063]
[0064] where Q represents the user question, S represents the graph pattern, Ct represents the t-th word, represents the weight of the t-th word Ct.
[0065] For symbols related to edges (such as "->" and "<-"), after tokenization, the corresponding word units are identified and higher weights are assigned to these word units, such as a weight of 5. The weights of the remaining word units are all set to 1.
[0066] This weight allocation mechanism is similar to setting a "focus" during model training, allowing the model to pay more attention to words related to edge direction. This is like emphasizing key parts of grammar rules when learning a language to ensure the correctness of sentence structure. In this way, the model can more accurately reflect the relationship direction in the graph pattern when generating Cypher queries, reducing query failures due to direction errors.
[0067] For step S3: sampling the large language model after supervised fine-tuning to obtain a sampled data set.
[0068] In embodiments of the present application, the specific method for sampling the large language model after supervised fine-tuning to obtain a sampled data set is as follows:
[0069] As shown in Figure 2 , given a set of question-Cypher pairs, the large language model trained with supervised fine-tuning is used to generate multiple Cypher queries, which are then executed in the graph database. The queries that produce results matching the standard answers are labeled as positive samples, and the queries that do not match the standard answers are labeled as negative samples. The sampling process is to evaluate the accuracy and diversity of the model when generating Cypher queries. By generating multiple queries and evaluating their execution results, the performance of the model can be more comprehensively understood. The distinction between positive and negative samples allows the model to learn which queries are correct and which are incorrect, so that it can better adjust its parameters in subsequent training and improve the quality of generated queries.
[0070] For step S4: training the large language model after supervised fine-tuning.
[0071] In embodiments of the present application, in the step of training the large language model after supervised fine-tuning, a direct preference optimization algorithm is used to train the large language model. Supervised fine-tuning is used to train the model based on real Cypher queries, so that it can learn patterns from the data set. However, not all training queries are of the same quality, and supervised fine-tuning does not explicitly distinguish between better and worse queries. To solve this problem, in the present application, negative samples are introduced using direct preference optimization to help the large language model distinguish between correct and incorrect Cypher queries.
[0072] In embodiments of the present application, in the process of training the large language model using the direct preference optimization algorithm, higher weights are assigned to words related to the direction of the edge. For symbols representing edges (such as “->” and “<-”), after tokenization is completed, the corresponding word pieces are identified, and the weights of these word pieces are set to 5. The weights of all other word pieces are uniformly set to 1.
[0073] In embodiments of the present application, the direct preference optimization process is represented as:
[0074]
[0075] where Q represents the user question, S represents the graph pattern, Ct represents the t-th word, Ct represents the weight of the t-th word, Cw is the positive sample, Cr is the negative sample. π represents the probability distribution, ref represents the reference model (initial SFT model), β is a hyperparameter, and σ represents the sigmoid function.
[0076] Direct preference optimization algorithm is a reinforcement learning method that adjusts the parameters of the model by comparing the probabilities of positive and negative samples. This method not only optimizes the output of the model, but also enhances the model's ability to distinguish different query qualities. By assigning higher weights to words related to the edge direction, the model will pay more attention to these key parts when generating queries, thereby improving the accuracy and consistency of the generated queries.
[0077] For step S5: Obtain the user input question and the graph database schema.
[0078] Obtaining the user input question and the graph database schema is the starting point of the entire system's interaction with the user. The user's question may contain various natural language expressions, while the graph database schema defines the structure and relationships of the data. Accurate acquisition and understanding of this information is the premise of generating correct Cypher queries.
[0079] For step S6: Select relevant sub-graph schema from the complete graph schema according to the user question and verify it.
[0080] In the embodiments of the present application, if the selected sub-graph schema does not meet the constraints of the original graph schema, an error is returned and the sub-graph schema is reselected in the step of selecting relevant sub-graph schema from the complete graph schema according to the user question and verifying it.
[0081] Specifically, before generating the Cypher query, the method of the present application will select the graph schema. This process includes providing the graph schema and the user question to the large language model. As shown in Figure 3 , the large language model needs to extract the parts of the schema that are helpful to answer the question, including relevant nodes, edges and their attributes. This process can be represented as:
[0082]
[0083] Where Q represents the user question, S represents the graph schema, and S' represents the selected graph schema.
[0084] There are two main reasons for the design of the graph pattern selection. First, directly inputting the entire pattern into the large language model can cause information overload, hindering the model's ability to distinguish relevant information from irrelevant details. Pattern selection can focus the model's attention on relevant pattern components, thus solving this problem. By filtering out irrelevant information, graph pattern selection helps to more accurately map natural language queries to the corresponding nodes and edges in the graph pattern, thus improving the accuracy of the generated Cypher. In addition, the inherent ambiguity of natural language queries can be mitigated through pattern selection, ensuring that the generated query is more in line with the user's intent. Second, pattern selection has a significant advantage in terms of computational efficiency in subsequent steps. There are now many methods to improve the accuracy of code generation, such as generating multiple outputs, reflecting on the results, and planning before generation. These methods require multiple calls to the large language model. Therefore, inputting the entire pattern each time will result in a longer context, causing significant computational overhead.
[0085] Unlike rule-based systems, large language models themselves do not understand the constraints of a specific graph pattern. Therefore, they often have difficulty adhering to the given graph schema and occasionally hallucinate, i.e., the selected graph pattern deviates from the input graph pattern. Specifically, the attributes assigned to nodes and edges may not exist in the original graph pattern, and the generated direction of edges may be incorrect. Even small differences in nodes or edges can severely affect the subsequent Cypher generation step. Therefore, ensuring the accuracy of the selected pattern is crucial. Therefore, the method will check the generated graph pattern, and if the generated graph pattern does not meet the constraints of the original graph pattern, it will return an error to the large language model to reselect the graph pattern.
[0086] Graph pattern selection and verification is a critical step that ensures the graph pattern on which the model bases when generating Cypher queries is accurate and relevant. Through this process, the model can avoid generating queries that are inconsistent with the graph pattern, thus improving the accuracy and reliability of the queries.
[0087] For step S7: The large language model generates the corresponding Cypher query using the selected pattern and the given question.
[0088] After graph pattern selection, the relevant sub-graph patterns are identified and incorporated into the prompt. Specifically, patterns unrelated to the query are removed, simplifying the prompt. With the selected pattern, the prompt is greatly simplified, and the number of words in the prompt can be significantly reduced. Finally, the prompt composed of the user question and the sub-graph pattern will be used by the large language model to generate the Cypher query. This process can be described as:
[0089]
[0090] Wherein, Q represents a user question, S represents a graph pattern, S' represents a selected graph pattern, C represents Cypher, and Ct represents the tth word.
[0091] The generation of a Cypher query is the ultimate goal of the entire method. Through the previous steps, the model has obtained sub-graph patterns related to the user question, and these patterns have been verified to ensure consistency with the original graph pattern. The Cypher query generated on this basis is more likely to be accurate and effective.
[0092] As shown in Figure 2 A large language model Cypher generation system based on graph pattern alignment enhancement according to the present application, comprising:
[0093] A data set construction module for constructing a data set, the data set containing multiple types of training data.
[0094] A supervised fine-tuning module for supervising and fine-tuning a large language model.
[0095] A sampling module for sampling the large language model after supervised fine-tuning to obtain a sampling data set.
[0096] A training module for training the large language model after supervised fine-tuning.
[0097] An acquisition module for acquiring a user input question and a graph database pattern.
[0098] A selection and verification module for selecting relevant sub-graph patterns from the complete graph pattern according to the user question and verifying them.
[0099] A Cypher generation module for generating a corresponding Cypher query based on the selected pattern and the given question through the trained large language model.
[0100] For the system embodiment, since it basically corresponds to the method embodiment, the relevant parts are described in the method embodiment. The implementation method of the remaining modules will not be repeated here. The system embodiments described above are only illustrative. The units described as separate components can or can not be physically separated, and the components displayed as units can or can not be physical units, i.e. they can be located in one place or distributed on multiple network units. Some or all of the modules can be selected to achieve the purpose of the present application according to actual needs. Those skilled in the art can understand and implement without creative labor.
[0101] Embodiments of the system of the present application can be applied on any data processing capable device, which can be a device or apparatus such as a computer. The system embodiments can be implemented by software, or by hardware, or by a combination of software and hardware. Taking the software implementation as an example, as a logical device, it is formed by the processor of the data processing capable device where it is located reading the corresponding computer program instructions in the non-volatile memory into the memory for running.
[0102] To demonstrate the effectiveness of the present application, different data on graph databases were used for post-training, and tests were conducted on enterprise relationship graph datasets and network management datasets to demonstrate the generality of the proposed method. Network management graphs focus on dependency and root cause analysis in network management, enabling users to perform impact analysis from routing level to application management and track dependencies. It helps to comprehensively supervise network management and impact assessment, including dependency relationships from routing protocols to application level. Enterprise relationship graphs are based on real-world use cases, and the dataset contains enterprise entities, including companies, individuals, group factions, and various relationships such as investment, legal person, and company subsidiaries.
[0103] Evaluation index: The evaluation index of the present application is to evaluate whether the execution result of the generated Cypher in the graph database is consistent with the execution result of the standard answer, using execution accuracy as the index, the higher the index, the better the performance.
[0104] Comparison method: The following methods are used as the baseline for testing: 1) R3-NLGQL: It integrates small and large models for sorting, rewriting and refining Cypher queries. 2) Vanilla Prompt: Vanilla Prompt refers to a simple and clear prompt provided to the LLM without additional optimization. 3) Self-Refine: LLM generates initial output, then evaluates its response and iteratively improves based on self-provided feedback. 4) Best-of-N Sampling: LLM will sample N different outputs and select the one with the highest evaluation score. To ensure the comprehensiveness of the evaluation, three widely used LLM testing methods are used for testing: Qwen2.5-Coder-32B-Instruct, DeepSeek-v3 and GPT-4o. The test results are shown in Table 1,
[0105] Table 1 Cypher generation effect
[0106]
[0107] Table 1 lists the performance of the method proposed in this patent and the baseline method on the network management dataset. The method of the present invention has achieved significant performance improvements in execution accuracy. It is worth noting that when using Qwen2.5-Coder-32B-Instruct as a model, the execution of the present invention is improved by 70.60% compared with Vanilla Prompt; compared with R3-NLGQL, the performance of the method of the patent is improved by 31.82%. In addition, the combined performance of the present invention and Qwen2.5-Coder-32B-Instruct is better than GPT-4o and DeepSeek-V3, with GPT-4o improving by 16.00% over R3-NLGQL and DeepSeek-V3 improving by 7.42% over R^3-NLGQL. These results highlight the great potential of the method of the present invention in enhancing the ability to generate cyphers.
[0108] To evaluate our approach in a real-world setting, we collected data from a real-world use case involving enterprise relationship graphs. Our approach demonstrated significant improvements in execution accuracy compared to baseline methods. Notably, our approach achieved a top-tier execution accuracy of 84.62% on this dataset, surpassing R3-NLGQL using GPT-4o by 15.79%. Furthermore, our approach demonstrated a strong performance improvement of 10.01% compared to R3-NLGQL using DeepSeek-V3.
[0109] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the above embodiments do not limit the present invention in any form, and any technical solutions obtained by equivalent replacement or equivalent transformation fall within the scope of protection of the present invention.
Claims
1. A large language model Cypher generation method based on graph pattern alignment enhancement, characterized in that: The following steps are involved: Constructing a data set, wherein the data set includes multiple types of training data; Supervised fine-tuning of large language models; Sampling the large language model after supervised fine-tuning to obtain a sampling dataset; Train the large language model after supervised fine-tuning; Get the user input question and graph database schema; Select relevant subgraph patterns from the complete graph pattern based on user questions and verify them; The large language model utilizes the selected pattern and the given question to generate the corresponding Cypher query.
2. The method for generating a large language model Cypher based on graph pattern alignment enhancement according to claim 1, characterized in that: The data set includes Cypher generated data, Cypher error correction data and graph pattern selection data; The input of the Cypher generated data is the user query and the graph pattern, and the output is the corresponding Cypher query; The input of the Cypher error correction data is a user query, a graph pattern, and a Cypher query that is uncertain whether it is correct, and the output is a correct Cypher query; The input of the graph pattern selection data is a user query and a graph pattern, and the output is a subgraph pattern required to answer the user query.
3. The method for generating a large language model Cypher based on graph pattern alignment enhancement according to claim 1, characterized in that: In the process of supervised fine-tuning of the large language model, higher weights are assigned to words related to edge directions to enhance the large language model's attention to the relationship direction in the graph pattern.
4. The method for generating a large language model Cypher based on graph pattern alignment enhancement according to claim 3 is characterized in that: The supervised fine-tuning process is expressed as: Among them, Q represents the user question, S represents the graph pattern, and Ct represents the tth word. Represents the weight of the t-th word Ct.
5. The method for generating a large language model Cypher based on graph pattern alignment enhancement according to claim 1, characterized in that: The specific method for sampling the large language model after supervised fine-tuning to obtain a sampled dataset is: Given a set of question-Cypher pairs, a large language model trained with supervised fine-tuning is used to generate multiple Cypher queries, which are then executed in a graph database. Queries that produce results that match the standard answers are labeled as positive samples, and queries that do not match the standard answers are labeled as negative samples.
6. The method for generating a large language model Cypher based on graph pattern alignment enhancement according to claim 1, characterized in that: In the step of training the large language model after supervised fine-tuning, a direct preference optimization algorithm is used to train the large language model.
7. The method for generating a large language model Cypher based on graph pattern alignment enhancement according to claim 6, characterized in that: In the process of training a large language model using the direct preference optimization algorithm, higher weights are assigned to words related to edge directions.
8. The method for generating a large language model Cypher based on graph pattern alignment enhancement according to claim 7, characterized in that: The direct preference optimization process is expressed as: Among them, Q represents the user question, S represents the graph pattern, and Ct represents the tth word. represents the weight of the t-th word Ct, Cw is a positive sample, and Cr is a negative sample. π represents the probability distribution, ref represents the reference model, β is a hyperparameter, and σ represents the sigmoid function.
9. The method for generating a large language model Cypher based on graph pattern alignment enhancement according to claim 1, characterized in that: In the step of selecting relevant subgraph patterns from the complete graph pattern according to the user question and performing verification, if the selected subgraph pattern does not meet the constraints of the original graph pattern, an error is returned and a new subgraph pattern is selected.
10. A large language model Cypher generation system based on graph pattern alignment enhancement, characterized in that: Include: A data set construction module, used to construct a data set, wherein the data set includes multiple types of training data; Supervised fine-tuning module, users supervise and fine-tune the large language model; The sampling module is used to sample the large language model after supervised fine-tuning to obtain a sampling dataset; The training module is used to train the large language model after supervised fine-tuning; The acquisition module is used to obtain the user input question and graph database schema; The selection and verification module is used to select relevant subgraph patterns from the complete graph pattern according to user questions and perform verification; The Cypher generation module is used to generate corresponding Cypher queries based on the selected pattern and given questions using a trained large language model.
Citation Information
Cited By
Large language model graph query language generation method and system based on reinforcement learning
CN121681720A