Graph database Cypher query statement generation method based on thinking chain

Through a thinking chain-based method, users' natural language query is converted into Cypher query statements, solving the problem of complexity of non-technical personnel writing Cypher query, realizing seamless interaction between natural language and graph database, and improving user experience and popularity.

CN120492683AActive Publication Date: 2025-08-15JIANGSU TONGXINGBAO INTELLIGENT TRANSPORTATION TECH CO LTD

Patent Information

Application Number
CN202510598857.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-08-15
Estimated Expiration
2045-05-09

AI Technical Summary

Technical Problem

Writing Cypher query statements is complicated for non-technical personnel or beginners, and prior art is difficult to achieve seamless interaction between natural language and graph databases.

Method used

Using a thinking chain-based method, through large language model training, meta-functions are used to transform users' natural language query requirements into Cypher query statements, combining deep learning and Transformer architecture to realize the interaction between natural language and graph database.

Benefits of technology

It lowers the threshold for using graph databases, allowing beginners to easily obtain information, expands the user group, improves the popularity of applications, and improves user experience and interactive comfort.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492683A_ABST
    Figure CN120492683A_ABST
Patent Text Reader

Abstract

The invention discloses a thinking chain-based graph database Cypher query statement generation method, and relates to the technical field of graph database query. Comprising the steps of data preprocessing, model training and query generation, model evaluation, security privacy guarantee, system integration and dynamic adaptability adjustment, and a natural language query requirement of a user is converted into a Cypher query statement executable by a graph database. By defining a meta function, a query problem is decomposed into a combination of basic operations, a human thinking chain mode is simulated, and the logic habit of human to process the problem is met. By means of the decomposition mode, the problem processing process is clearer, more organized and easy to understand and master. The construction of the query logic by developers and the subsequent maintenance and optimization of the system become more efficient and convenient due to the visual problem decomposition mode.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of graph database query technology, and in particular to a method for generating graph database Cypher query statements based on thought chains. Background Art

[0002] With the continuous development of information technology, data storage and retrieval are becoming increasingly important. Graph databases are powerful tools for storing and analyzing complex relational data. Data in graph databases is represented as nodes and edges, and this data structure is particularly well-suited for representing complex relationships between entities, such as in social networks, knowledge graphs, and recommendation systems. Cypher is a query language used to query graph databases, providing powerful capabilities for retrieving and manipulating data within them. However, writing Cypher queries requires a deep understanding of database structure and query syntax, which can be a complex task for non-technical personnel or beginners.

[0003] Today, the rapid development of natural language processing technology, especially large language models, has undoubtedly provided a good solution to many text generation problems. Natural language processing (NLP) is based on deep learning and the Transformer architecture. By training on large-scale text data, computers can understand and generate natural language text, and thus perform well in various NLP tasks such as machine translation, sentiment analysis, and question answering. Large language models are usually based on the Transformer architecture and can understand and generate natural language text by pre-training on large-scale text data. They can process the grammatical, semantic, and contextual information of the text, enabling them to perform well in various NLP tasks such as text generation, question answering, sentiment analysis, and machine translation. However, in a specific field, large language models often need to be trained with a large amount of domain-specific data to achieve efficient solutions to the problem.

[0004] To address these issues, this paper proposes a thought chain-based approach that trains a large language model to automatically generate Cypher queries for interacting with data in a graph database based on user questions. This approach allows users to easily ask questions and retrieve the data they need without having to delve into Cypher query syntax. Summary of the Invention

[0005] The purpose of the present invention is to provide a method for generating Cypher query statements for a graph database based on thought chains, which solves the technical problems raised in the background technology.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a method for generating Cypher query statements for a graph database based on thought chaining, which includes data preprocessing, model training and query generation, model evaluation, security and privacy protection, system integration, and dynamic adaptive adjustment. The method converts user natural language query requirements into Cypher query statements that can be executed by the graph database, as follows:

[0007] The data preprocessing includes: integrating a meta-function system of task prompts and operation details (comments), metadata information of the graph database, and relationship pairs; cleaning natural language questions in the training data, removing noise characters, and unifying the text format; performing consistency verification and normalization on the graph database metadata; using the AES encryption algorithm to ensure data transmission and storage security; building a rights management system to control query permissions; and adding noise to the training data using a differential privacy mechanism;

[0008] The model training and query generation uses a large pre-trained model based on deep learning and Transformer architecture to map pre-processed user natural language query requirements into Python code composed of predefined meta-functions, which is then converted into Cypher query statements through the Python interpreter to achieve data interaction with the graph database;

[0009] The model evaluation uses an independent test dataset to evaluate accuracy by calculating precision, recall, and F1 value, records code execution and query response time to evaluate response efficiency, and organizes professionals to review query statements to evaluate feasibility and rationality;

[0010] The security and privacy protection mentioned above uses AES encryption algorithm, sophisticated permission management system and differential privacy mechanism to ensure data security and privacy;

[0011] The system integration design integrates interfaces with other systems to achieve seamless docking;

[0012] The dynamic adaptive adjustment establishes a graph database change monitoring mechanism, which automatically triggers the model fine-tuning mechanism when the database changes, ensuring that the model adapts to the new environment.

[0013] Preferably, in the data preprocessing stage, in addition to basic data cleaning, verification and standardization operations, the model's own language understanding ability is used to deeply analyze natural language query requirements, specifically including extracting key entities, relationships and attribute information, and annotating them according to specific classification standards. At the same time, the type and attribute information of the graph database metadata are calibrated and standardized to ensure data consistency and accuracy. In the process of converting query requirements into Python code, the preprocessed data is fully combined, and the combination of meta-functions is used to improve the adaptability and accuracy of Cypher query statements and graph databases. Based on the evaluation results of the model in the test data set, the performance differences of different types of query requirements are deeply analyzed. For query types with poor performance, the reasons are analyzed from the aspects of the quality of training data, the rationality of model parameters and the effectiveness of training algorithms. Then, targeted optimization is carried out by adjusting the distribution of training data, optimizing model parameters or improving the training algorithm. Visual tools are developed to display the model reasoning process, and explanation generation algorithms are designed to generate natural language explanations for code and query statements to enhance the interpretability of the model.

[0014] Optimally, the predefined metafunctions used to construct Python code are the basic operational units for abstracting graph database query tasks. The preprocessing phase standardizes the metafunction's input and output data formats and checks data transfer compatibility. Metafunctions perform single operations such as node query, relationship retrieval, and data filtering, and are combined to construct complex query logic. Task prompts provide macro-level guidance, while detailed operational instructions describe input and output specifications, parameter constraints, and internal logic. The metafunction's contribution to model generation results is evaluated, along with usage frequency and correct call rate analysis. Function definitions, parameter settings, or prompts are adjusted. Metafunction design considers system integration compatibility and scalability to adapt to dynamic changes in graph databases.

[0015] Preferably, it has the ability to handle complex query requirements in all scenarios of graph databases, deeply analyze natural language query requirements in the preprocessing stage, extract key information and standardize its representation, mine metadata to build a relational graph for query optimization, perform semantic analysis and logical reconstruction of the preprocessed query requirements, convert them into meta-function call sequences to generate Python code and Cypher query statements, optimize queries in combination with metadata information, use graph database indexing, storage and query mechanisms to achieve accurate retrieval and efficient analysis, evaluate the model's complex query processing capabilities, compare with expert query statements, optimize complex query generation strategies based on the gaps, and adaptively adjust the meta-function system and query generation strategies as the database changes to ensure accurate transmission and processing of complex query results during integration.

[0016] Preferably, the process of converting the user's natural language query requirements into Python code includes the following steps:

[0017] Precise metafunction design: Based on the underlying data model, topology, and common query patterns of the graph database, design metafunctions that are functionally independent, have standardized interfaces, and are reusable. Define metafunction inputs and outputs to ensure efficient interaction with the graph database, consider system integration compatibility and scalability, and adapt to dynamic database changes. After the design is complete, simulate query scenarios to evaluate the metafunction and adjust the design to ensure reliability and effectiveness.

[0018] Careful construction of sample data: Design examples based on graph database metadata and business requirements, including natural language query requirements, meta-function call sequences, and semantic interpretations. Preprocess query requirements, verify the accuracy of call sequences, select and annotate examples to cover typical query scenarios and business logic, provide high-quality training data, regularly evaluate sample data completeness, accuracy, and coverage of new scenarios, supplement and update data, integrate into system integration cases, and consider the impact of dynamic environmental data changes.

[0019] Model training and guidance: Input example data into a large pre-trained model, using supervised learning and reinforcement learning strategies to guide the model to learn the mapping relationship between natural language query requirements and Python code. Use data preprocessing techniques to enhance and standardize input data, optimize model parameters and structure, monitor training indicators, adjust training parameters according to changes, consider model interpretability during training, and guide the model to learn explainable decision-making processes;

[0020] Code generation and verification: After the model training meets the requirements, Python code is generated according to the new natural language query requirements. The code undergoes syntax checking, semantic verification, and logical reasoning, and is optimized with preprocessing information. The code verification results are incorporated into the evaluation system, and the error types and frequencies are analyzed to evaluate the reliability of code generation. The code is also verified for compatibility and correctness in system integration and stability under dynamic database changes.

[0021] Preferably, after constructing the Cypher query statement through the Python interpreter execution code, the query statement is sent to the graph database query execution engine, the engine performs lexical analysis, syntax parsing, semantic checking and query optimization, generates an execution plan to retrieve data and formats the output results, the system monitors and analyzes the query execution performance, records time and resource consumption for query optimization and system tuning, pre-processes the query results, removes redundant information and converts it into an easy-to-understand format, evaluates the model based on the query result accuracy and performance indicators, and optimizes the model with feedback results. When the query involves system integration, ensure that the results are correctly transmitted and adapted, and adjust the processing strategy when the database changes.

[0022] Preferably, the metafunction design follows the principles of modularity, hierarchy, and composability. The metafunction is functionally cohesive and is responsible for a single operation. The interface standards are unified to facilitate combination and expansion. The hierarchical design enables operations at different abstract levels to collaborate to build complex logic. Preprocessing verifies the input and output data of the metafunction based on the modular characteristics to ensure accurate and stable transmission. The metafunction design adapts to the development of graph databases and business changes, and has scalability and adaptability. The scalability of the metafunction system is evaluated, new scenarios and data changes are simulated to test the adaptability of the model and the quality of the generated results, the architectural design is optimized, and the scalability and compatibility of the metafunction in system integration, as well as the stability and adaptability of the metafunction under dynamic changes in the database, are evaluated.

[0023] Preferably, the relationship pairs of training large language models are subjected to data mining, feature engineering and quality assessment. During the mining and engineering process, natural language problem features are extracted and converted into vector representations. The function call sequence is analyzed and optimized to extract key features and patterns. The quality assessment link cleans and screens the relationship pairs, removes noise and incorrectly labeled data, and the relationship pairs cover common and complex query cases to ensure that the model learns comprehensive query knowledge and reasoning capabilities. The training data is regularly updated and expanded to adapt to dynamic changes in the database and business needs. The impact of training data on model generalization is evaluated, and the training effects of different subsets are compared. The data selection and expansion strategies are optimized, system integration cases are increased, and the impact of dynamic environmental data changes is considered.

[0024] Preferably, before the generated Python code is input into the interpreter, static code analysis, symbolic execution and model checking techniques are used to perform quality checks, preprocess variables and function calls in the code to provide accurate verification information, analyze code syntax, data types, control flow and data flow, detect errors and provide repair suggestions, use code coverage tools to evaluate test coverage, improve code reliability and stability, use code quality inspection results as an evaluation indicator of model generation capability, analyze the relationship with training parameters and data, optimize the training process to improve code quality, and check the compatibility, security and stability of the code in system integration under dynamic changes in the database.

[0025] Preferably, the generated Cypher query statements are subjected to multi-level performance optimization. In the construction phase, the query logic is transformed based on metadata and optimization rule analysis to eliminate redundant operations and subqueries. In the execution phase, the query is optimized using graph database indexing, caching, and parallel computing technologies. Optimization strategies are formulated in combination with preprocessing statistical information and historical data. Strategies are adjusted through real-time monitoring and feedback. A performance evaluation model is established. The effectiveness of the optimization strategy is evaluated by comprehensively considering execution time, resource consumption, and result accuracy. The optimal combination is selected and continuously improved. The effectiveness, adaptability, and adjustment capabilities of the optimization strategy in system integration under dynamic changes in the database are evaluated.

[0026] Compared with related technologies, the method for generating Cypher query statements for graph databases based on thought chains provided by the present invention has the following beneficial effects:

[0027] 1. This invention provides a method for generating Cypher queries for graph databases based on thought chaining. By defining metafunctions, this method decomposes query problems into a combination of basic operations, simulating the human thought chain approach and conforming to the logical habits of human problem-solving. This decomposition method makes the problem-solving process clearer, more organized, easier to understand, and more manageable. This intuitive problem-solving method makes both the developer's construction of query logic and the subsequent maintenance and optimization of the system more efficient and convenient.

[0028] 2. This invention provides a method for generating Cypher query statements for graph databases based on thought chains. This method eliminates the need for users to learn and master complex Cypher query syntax; users can interact with graph databases simply by asking questions in natural language. This significantly lowers the barrier to entry for graph database users, enabling beginners and non-technical personnel to easily obtain the information they need from graph databases. This expands the user base of graph databases and increases their widespread application across various fields.

[0029] 3. This invention provides a method for generating Cypher query statements for graph databases based on thought chains. This method uses natural language interaction, which is as natural and smooth as human conversation, aligns with users' daily communication habits and makes interacting with graph databases more comfortable and convenient. Compared to traditional methods that require memorizing and writing specific syntax, this natural language interaction method reduces user learning costs and operational difficulty, thereby improving user experience and satisfaction. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 Flow chart of the implementation of the present invention;

[0031] Figure 2 is a flow chart of the present invention;

[0032] Figure 3 This is an extended flow chart of data preprocessing of the present invention;

[0033] Figure 4 An extended flow chart for model training and query generation of the present invention;

[0034] Figure 5 An expanded flow chart for the evaluation of the model of the present invention;

[0035] Figure 6 This is an extended flow chart of the security and privacy protection of the present invention;

[0036] Figure 7An expanded flow chart for system integration of the present invention;

[0037] Figure 8 This is an extended flow chart of the dynamic adaptive adjustment of the present invention. DETAILED DESCRIPTION

[0038] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments; based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0039] Example 1:

[0040] See also Figures 1-8 The present invention provides a technical solution: a method for generating Cypher query statements for a graph database based on thought chain, including data preprocessing, model training and query generation, model evaluation, security and privacy protection, system integration, and dynamic adaptive adjustment, which converts user natural language query requirements into Cypher query statements that can be executed by the graph database, as follows:

[0041] The data preprocessing includes: integrating a meta-function system of task prompts and operation details (comments), metadata information of the graph database, and relationship pairs; cleaning natural language questions in the training data, removing noise characters, and unifying the text format; performing consistency verification and normalization on the graph database metadata; using the AES encryption algorithm to ensure data transmission and storage security; building a rights management system to control query permissions; and adding noise to the training data using a differential privacy mechanism;

[0042] The model training and query generation uses a large pre-trained model based on deep learning and Transformer architecture to map pre-processed user natural language query requirements into Python code composed of predefined meta-functions, which is then converted into Cypher query statements through the Python interpreter to achieve data interaction with the graph database;

[0043] The model evaluation uses an independent test dataset to evaluate accuracy by calculating precision, recall, and F1 value, records code execution and query response time to evaluate response efficiency, and organizes professionals to review query statements to evaluate feasibility and rationality;

[0044] The security and privacy protection mentioned above uses AES encryption algorithm, sophisticated permission management system and differential privacy mechanism to ensure data security and privacy;

[0045] The system integration design integrates interfaces with other systems to achieve seamless docking;

[0046] The dynamic adaptive adjustment establishes a graph database change monitoring mechanism, which automatically triggers the model fine-tuning mechanism when the database changes, ensuring that the model adapts to the new environment.

[0047] During the data preprocessing phase, in addition to basic data cleaning, verification, and standardization operations, the model's own language understanding capabilities are leveraged to conduct in-depth analysis of natural language query requirements. This includes extracting key entities, relationships, and attribute information, and annotating them according to specific classification standards. At the same time, the types and attribute information of the graph database metadata are calibrated and standardized to ensure data consistency and accuracy. In the process of converting query requirements into Python code, the preprocessed data is fully integrated, and a combination of meta-functions is used to improve the compatibility and accuracy of Cypher query statements with the graph database. Based on the evaluation results of the model on the test dataset, the performance differences of different types of query requirements are deeply analyzed. For query types with poor performance, the reasons are analyzed from the aspects of training data quality, model parameter rationality, and training algorithm effectiveness. Targeted optimization is then carried out by adjusting the training data distribution, optimizing model parameters, or improving the training algorithm. Visualization tools are developed to display the model reasoning process, and explanation generation algorithms are designed to generate natural language explanations for code and query statements to enhance the interpretability of the model.

[0048] Predefined metafunctions for Python code are the basic operational units for abstracting graph database query tasks. Preprocessing standardizes metafunction input and output data formats and checks data transfer compatibility. Metafunctions perform single operations such as node query, relationship retrieval, and data filtering, and can be combined to construct complex query logic. Task prompts provide high-level guidance, while detailed operational instructions describe input and output specifications, parameter constraints, and internal logic. Metafunctions are evaluated for their contribution to model generation results, analyzing usage frequency and correct call rates, and adjusting function definitions, parameter settings, and prompts. Metafunction design considers system integration compatibility and scalability to adapt to dynamic changes in graph databases.

[0049] It has the ability to handle complex query requirements in all scenarios of graph databases. In the preprocessing stage, it deeply analyzes natural language query requirements, extracts key information and standardizes its representation, mines metadata to build a relational graph for query optimization, performs semantic analysis and logical reconstruction on the preprocessed query requirements, and converts them into meta-function call sequences to generate Python code and Cypher query statements. It optimizes queries based on metadata information, uses the graph database indexing, storage and query mechanisms to achieve accurate retrieval and efficient analysis, evaluates the model's complex query processing capabilities, compares with expert query statements, and optimizes complex query generation strategies based on the gaps. The meta-function system and query generation strategies are adaptively adjusted as the database changes to ensure accurate transmission and processing of complex query results during integration.

[0050] The process of converting user natural language query requirements into Python code includes the following steps:

[0051] Precise metafunction design: Based on the underlying data model, topology, and common query patterns of the graph database, design metafunctions that are functionally independent, have standardized interfaces, and are reusable. Define metafunction inputs and outputs to ensure efficient interaction with the graph database, consider system integration compatibility and scalability, and adapt to dynamic database changes. After the design is complete, simulate query scenarios to evaluate the metafunction and adjust the design to ensure reliability and effectiveness.

[0052] Careful construction of sample data: Design examples based on graph database metadata and business requirements, including natural language query requirements, meta-function call sequences, and semantic interpretations. Preprocess query requirements, verify the accuracy of call sequences, select and annotate examples to cover typical query scenarios and business logic, provide high-quality training data, regularly evaluate sample data completeness, accuracy, and coverage of new scenarios, supplement and update data, integrate into system integration cases, and consider the impact of dynamic environmental data changes.

[0053] Model training and guidance: Input example data into a large pre-trained model, using supervised learning and reinforcement learning strategies to guide the model to learn the mapping relationship between natural language query requirements and Python code. Use data preprocessing techniques to enhance and standardize input data, optimize model parameters and structure, monitor training indicators, adjust training parameters according to changes, consider model interpretability during training, and guide the model to learn explainable decision-making processes;

[0054] Code generation and verification: After the model training meets the requirements, Python code is generated according to the new natural language query requirements. The code undergoes syntax checking, semantic verification, and logical reasoning, and is optimized with preprocessing information. The code verification results are incorporated into the evaluation system, and the error types and frequencies are analyzed to evaluate the reliability of code generation. The code is also verified for compatibility and correctness in system integration and stability under dynamic database changes.

[0055] After constructing a Cypher query statement through the Python interpreter's execution code, the query statement is sent to the graph database query execution engine. The engine performs lexical analysis, syntax parsing, semantic checking, and query optimization, generates an execution plan to retrieve data, and formats the output results. The system monitors and analyzes query execution performance, records time and resource consumption for query optimization and system tuning, preprocesses query results, removes redundant information, and converts them into an easy-to-understand format. The model is evaluated based on the query result accuracy and performance indicators, and the feedback result is used to optimize the model. When the query involves system integration, ensure that the results are correctly transmitted and adapted, and adjust the processing strategy when the database changes.

[0056] Metafunction design follows the principles of modularity, hierarchy, and composability. Metafunctions are functionally cohesive and responsible for a single operation. Interface standards are unified to facilitate combination and expansion. Hierarchical design enables operations at different abstract levels to collaborate to build complex logic. Preprocessing verifies metafunction input and output data based on modularity to ensure accurate and stable transmission. Metafunction design adapts to graph database development and business changes, and is scalable and adaptable. It evaluates the scalability of the metafunction system, simulates new scenarios and data changes to test the model's adaptability and the quality of generated results, optimizes the architectural design, and evaluates the scalability and compatibility of metafunctions in system integration, as well as their stability and adaptability under dynamic changes in the database.

[0057] The relationship pairs for training large language models undergo data mining, feature engineering, and quality assessment. During the mining and engineering process, natural language problem features are extracted and converted into vector representations. Function call sequences are analyzed and optimized to extract key features and patterns. The quality assessment process cleans and screens relationship pairs to remove noise and incorrectly labeled data. The relationship pairs cover common and complex query cases to ensure that the model learns comprehensive query knowledge and reasoning capabilities. The training data is regularly updated and expanded to adapt to dynamic changes in the database and business needs. The impact of training data on model generalization is evaluated, and the training effects of different subsets are compared. The data selection and expansion strategies are optimized, system integration cases are added, and the impact of dynamic environment data changes is considered.

[0058] Before the generated Python code is input into the interpreter, static code analysis, symbolic execution, and model checking techniques are used to perform quality checks. Variables and function calls in the code are preprocessed to provide accurate verification information. The code syntax, data types, control flow, and data flow are analyzed to detect errors and provide repair suggestions. Code coverage tools are used to evaluate test coverage to improve code reliability and stability. The code quality check results are used as an evaluation indicator for model generation capabilities. The relationship with training parameters and data is analyzed, the training process is optimized to improve code quality, and the code compatibility, security, and stability under dynamic changes in the database during system integration are checked.

[0059] Perform multi-level performance optimization on the generated Cypher query statements. During the construction phase, query logic is transformed based on metadata and optimization rules analysis to eliminate redundant operations and subqueries. During the execution phase, graph database indexing, caching, and parallel computing technologies are used to optimize queries. Optimization strategies are formulated based on preprocessing statistical information and historical data. Strategies are adjusted through real-time monitoring and feedback. A performance evaluation model is established to evaluate the effectiveness of optimization strategies by comprehensively considering execution time, resource consumption, and result accuracy. The optimal combination is selected and continuously improved. The effectiveness and adaptability of the optimization strategy in system integration and its ability to adjust under dynamic changes in the database are evaluated.

[0060] The present invention defines a series of meta-functions to decompose complex query problems into a combination of basic operations, simulating the human thought chain for problem solving. When processing a query on a social network graph database, if you want to find "people older than 30 among Alice's friends", the system will decompose this query into multiple basic operations represented by meta-functions. First, the NODE meta-function is used to locate the user node "Alice", and then the RELATION meta-function is used to find nodes with a "friend" relationship with "Alice". Finally, the FILTER meta-function is used to filter out nodes older than 30. This decomposition method makes the query logic clear and concise. When developers construct queries, they can organize various operations more intuitively, just like building blocks, combining meta-functions with different functions. During subsequent system maintenance and optimization, It can also quickly locate specific operation links for adjustment, greatly improving the efficiency of development and maintenance. Users do not need to master complex Cypher query syntax, they only need to use natural language to put forward query requirements, and the system can complete the conversion from natural language to Cypher query statements. This is due to the system's data preprocessing, model training and query generation links. In the data preprocessing stage, the system cleans and standardizes natural language problems to extract key information; in the model training and query generation links, the large-scale pre-trained model that has been pre-trained on large-scale multi-source heterogeneous text data has learned the mapping relationship between natural language and meta-function code. The system uses natural language interaction to fit the user's daily communication habits, reducing the cost and difficulty of learning specific grammar for users, and improving the user's comfort and satisfaction in using the graph database.

[0061] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a fully hardware embodiment, a fully software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code. The storage medium may be implemented by any type of volatile or non-volatile storage device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0062] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A method for generating Cypher query statements for a graph database based on thought chaining, characterized by: This includes data preprocessing, model training and query generation, model evaluation, security and privacy protection, system integration, and dynamic adaptive adjustment, converting user natural language query requirements into Cypher query statements that can be executed by the graph database. The details are as follows: The data preprocessing includes: integrating a meta-function system of task prompts and operation details (comments), metadata information of the graph database, and relationship pairs; cleaning natural language questions in the training data, removing noise characters, and unifying the text format; performing consistency verification and normalization on the graph database metadata; using the AES encryption algorithm to ensure data transmission and storage security; building a rights management system to control query permissions; and adding noise to the training data using a differential privacy mechanism; The model training and query generation uses a large pre-trained model based on deep learning and Transformer architecture to map pre-processed user natural language query requirements into Python code composed of predefined meta-functions, which is then converted into Cypher query statements through the Python interpreter to achieve data interaction with the graph database; The model evaluation uses an independent test dataset to evaluate accuracy by calculating precision, recall, and F1 value, records code execution and query response time to evaluate response efficiency, and organizes professionals to review query statements to evaluate feasibility and rationality; The security and privacy protection mentioned above uses AES encryption algorithm, sophisticated permission management system and differential privacy mechanism to ensure data security and privacy; The system integration design integrates interfaces with other systems to achieve seamless docking; The dynamic adaptive adjustment establishes a graph database change monitoring mechanism, which automatically triggers the model fine-tuning mechanism when the database changes, ensuring that the model adapts to the new environment.

2. The method for generating Cypher query statements for a graph database based on thought chaining according to claim 1, characterized in that: During the data preprocessing phase, in addition to basic data cleaning, verification, and standardization operations, the model's own language understanding capabilities are leveraged to conduct in-depth analysis of natural language query requirements. This includes extracting key entities, relationships, and attribute information, and annotating them according to specific classification standards. At the same time, the types and attribute information of the graph database metadata are calibrated and standardized to ensure data consistency and accuracy. In the process of converting query requirements into Python code, the preprocessed data is fully integrated, and a combination of meta-functions is used to improve the compatibility and accuracy of Cypher query statements with the graph database. Based on the evaluation results of the model on the test dataset, the performance differences of different types of query requirements are deeply analyzed. For query types with poor performance, the reasons are analyzed from the aspects of training data quality, model parameter rationality, and training algorithm effectiveness. Targeted optimization is then carried out by adjusting the training data distribution, optimizing model parameters, or improving the training algorithm. Visualization tools are developed to display the model reasoning process, and explanation generation algorithms are designed to generate natural language explanations for code and query statements to enhance the interpretability of the model.

3. The method for generating Cypher query statements for a graph database based on thought chaining according to claim 2 is characterized in that: Predefined metafunctions for Python code are the basic operational units for abstracting graph database query tasks. Preprocessing standardizes metafunction input and output data formats and checks data transfer compatibility. Metafunctions perform single operations such as node query, relationship retrieval, and data filtering, and can be combined to construct complex query logic. Task prompts provide high-level guidance, while detailed operational instructions describe input and output specifications, parameter constraints, and internal logic. Metafunctions are evaluated for their contribution to model generation results, analyzing usage frequency and correct call rates, and adjusting function definitions, parameter settings, and prompts. Metafunction design considers system integration compatibility and scalability to adapt to dynamic changes in graph databases.

4. The method for generating Cypher query statements for a graph database based on thought chaining according to claim 3 is characterized by: It has the ability to handle complex query requirements in all scenarios of graph databases. In the preprocessing stage, it deeply analyzes natural language query requirements, extracts key information and standardizes its representation, mines metadata to build a relational graph for query optimization, performs semantic analysis and logical reconstruction on the preprocessed query requirements, and converts them into meta-function call sequences to generate Python code and Cypher query statements. It optimizes queries based on metadata information, uses the graph database indexing, storage and query mechanisms to achieve accurate retrieval and efficient analysis, evaluates the model's complex query processing capabilities, compares with expert query statements, and optimizes complex query generation strategies based on the gaps. The meta-function system and query generation strategies are adaptively adjusted as the database changes to ensure accurate transmission and processing of complex query results during integration.

5. The method for generating Cypher query statements for a graph database based on thought chaining according to claim 4 is characterized in that: The process of converting user natural language query requirements into Python code includes the following steps: Precise metafunction design: Based on the underlying data model, topology, and common query patterns of the graph database, design metafunctions that are functionally independent, have standardized interfaces, and are reusable. Define metafunction inputs and outputs to ensure efficient interaction with the graph database, consider system integration compatibility and scalability, and adapt to dynamic database changes. After the design is complete, simulate query scenarios to evaluate the metafunction and adjust the design to ensure reliability and effectiveness. Careful construction of sample data: Design examples based on graph database metadata and business requirements, including natural language query requirements, meta-function call sequences, and semantic interpretations. Preprocess query requirements, verify the accuracy of call sequences, select and annotate examples to cover typical query scenarios and business logic, provide high-quality training data, regularly evaluate sample data completeness, accuracy, and coverage of new scenarios, supplement and update data, integrate into system integration cases, and consider the impact of dynamic environmental data changes. Model training and guidance: Input example data into a large pre-trained model, using supervised learning and reinforcement learning strategies to guide the model to learn the mapping relationship between natural language query requirements and Python code. Use data preprocessing techniques to enhance and standardize input data, optimize model parameters and structure, monitor training indicators, adjust training parameters according to changes, consider model interpretability during training, and guide the model to learn explainable decision-making processes; Code generation and verification: After the model training meets the requirements, Python code is generated according to the new natural language query requirements. The code undergoes syntax checking, semantic verification, and logical reasoning, and is optimized with preprocessing information. The code verification results are incorporated into the evaluation system, and the error types and frequencies are analyzed to evaluate the reliability of code generation. The code is also verified for compatibility and correctness in system integration and stability under dynamic database changes.

6. The method for generating Cypher query statements for a graph database based on thought chaining according to claim 5, characterized in that: After constructing a Cypher query statement through the Python interpreter's execution code, the query statement is sent to the graph database query execution engine. The engine performs lexical analysis, syntax parsing, semantic checking, and query optimization, generates an execution plan to retrieve data, and formats the output results. The system monitors and analyzes query execution performance, records time and resource consumption for query optimization and system tuning, preprocesses query results, removes redundant information, and converts them into an easy-to-understand format. The model is evaluated based on the query result accuracy and performance indicators, and the feedback result is used to optimize the model. When the query involves system integration, ensure that the results are correctly transmitted and adapted, and adjust the processing strategy when the database changes.

7. The method for generating Cypher query statements for a graph database based on thought chaining according to claim 6 is characterized in that: Metafunction design follows the principles of modularity, hierarchy, and composability. Metafunctions are functionally cohesive and responsible for a single operation. Interface standards are unified to facilitate combination and expansion. Hierarchical design enables operations at different abstract levels to collaborate to build complex logic. Preprocessing verifies metafunction input and output data based on modularity to ensure accurate and stable transmission. Metafunction design adapts to graph database development and business changes, and is scalable and adaptable. It evaluates the scalability of the metafunction system, simulates new scenarios and data changes to test the model's adaptability and the quality of generated results, optimizes the architectural design, and evaluates the scalability and compatibility of metafunctions in system integration, as well as their stability and adaptability under dynamic changes in the database.

8. The method for generating Cypher query statements for a graph database based on thought chaining according to claim 7, characterized in that: The relationship pairs for training large language models undergo data mining, feature engineering, and quality assessment. During the mining and engineering process, natural language problem features are extracted and converted into vector representations. Function call sequences are analyzed and optimized to extract key features and patterns. The quality assessment process cleans and screens relationship pairs to remove noise and incorrectly labeled data. The relationship pairs cover common and complex query cases to ensure that the model learns comprehensive query knowledge and reasoning capabilities. The training data is regularly updated and expanded to adapt to dynamic changes in the database and business needs. The impact of training data on model generalization is evaluated, and the training effects of different subsets are compared. The data selection and expansion strategies are optimized, system integration cases are added, and the impact of dynamic environment data changes is considered.

9. The method for generating Cypher query statements for a graph database based on thought chaining according to claim 8, characterized in that: Before the generated Python code is input into the interpreter, static code analysis, symbolic execution, and model checking techniques are used to perform quality checks. Variables and function calls in the code are preprocessed to provide accurate verification information. The code syntax, data types, control flow, and data flow are analyzed to detect errors and provide repair suggestions. Code coverage tools are used to evaluate test coverage to improve code reliability and stability. The code quality check results are used as an evaluation indicator for model generation capabilities. The relationship with training parameters and data is analyzed, the training process is optimized to improve code quality, and the code compatibility, security, and stability under dynamic changes in the database during system integration are checked.

10. The method for generating Cypher query statements for a graph database based on thought chaining according to claim 9, characterized in that: Perform multi-level performance optimization on the generated Cypher query statements. During the construction phase, query logic is transformed based on metadata and optimization rules analysis to eliminate redundant operations and subqueries. During the execution phase, graph database indexing, caching, and parallel computing technologies are used to optimize queries. Optimization strategies are formulated based on preprocessing statistical information and historical data. Strategies are adjusted through real-time monitoring and feedback. A performance evaluation model is established to evaluate the effectiveness of optimization strategies by comprehensively considering execution time, resource consumption, and result accuracy. The optimal combination is selected and continuously improved. The effectiveness and adaptability of the optimization strategy in system integration and its ability to adjust under dynamic changes in the database are evaluated.

Citation Information

Patent Citations

  • Improved NL2Cypher generation method and system based on generative pre-training model

    CN118551020A

  • Methods and systems for natural language processing of graph database queries

    US20220414228A1

Cited By

  • Large language model data mining interaction method and system based on traceable thinking chain

    CN121724169A

  • Traceable thought chain-based large language model data mining interaction method and system

    CN121724169B