Database query statement generation method and device, equipment, medium and product
By acquiring user input statements and knowledge graphs, information is extracted and filtered. Query statements for the NeBulaGraph database are generated and optimized using a large language model, solving the problem that non-professional users have difficulty generating NGQL query statements and achieving efficient and accurate automatic generation of query statements.
Patent Information
- Application Number
- CN202510980766.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-16
- Publication Date
- 2025-11-25
AI Technical Summary
In existing technologies, it is difficult for non-professional users to effectively and easily generate graph database query statements that meet semantic requirements, especially NGQL query statements for the NeBulaGraph database, resulting in low query efficiency and a high risk of errors.
By acquiring user input statements and knowledge graphs, information is extracted and filtered to obtain a set of candidate entities and a pattern subgraph. Multiple candidate query statements are generated using a large language model, and the target query statement is obtained by filtering based on the execution results.
Without needing to master NGQL syntax, users can automatically generate database query statements that meet semantic requirements, improving generation efficiency and accuracy, adapting to graph structures in different fields, and handling complex and diverse natural language problems.
Smart Images

Figure CN121009212A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of computer technology, and in particular to a method, apparatus, device, medium and product for generating database query statements. Background Technology
[0002] With the rapid development of big data and artificial intelligence technologies, graph databases, as a new type of database technology, have been widely used in various fields. Graph databases are particularly suitable for processing and analyzing complex relational data, such as social networks, recommendation systems, and healthcare. NeBulaGraph is an open-source distributed graph database widely used for graph data storage and retrieval, and it supports the Nebula Graph Query Language (NGQL).
[0003] However, while NGQL, as a query language, provides rich query capabilities for professional users, it still presents challenges for non-technical users in effectively and easily querying graph database data. Users often need precise queries to obtain relevant information, but due to a lack of understanding of the NGQL language, they often struggle to efficiently retrieve the data they need.
[0004] Therefore, how to automatically generate NGQL query statements that meet semantic requirements has become an urgent problem to be solved. Summary of the Invention
[0005] This invention provides a method, apparatus, device, medium, and product for generating database query statements, which can automatically generate database query statements that meet semantic requirements.
[0006] According to one aspect of the present invention, a method for generating database query statements is provided, comprising:
[0007] Obtain user input statements and knowledge graphs;
[0008] Information is extracted from the user input statement to obtain key information;
[0009] Based on the key information, the knowledge graph is filtered to obtain a candidate entity set and a pattern subgraph;
[0010] Input the prompts, user input statements, pattern subgraphs, and candidate entity sets into the large language model to obtain multiple candidate database query statements;
[0011] Based on the pattern subgraph and the execution results corresponding to each candidate database query statement, multiple candidate database query statements are filtered to obtain the target database query statement.
[0012] According to another aspect of the present invention, a database query statement generation apparatus is provided, the apparatus comprising:
[0013] The module is used to retrieve user input statements and knowledge graphs.
[0014] The information extraction module is used to extract information from the user input statement to obtain key information;
[0015] The first filtering module is used to filter the knowledge graph based on the key information to obtain a candidate entity set and a pattern subgraph.
[0016] The candidate database query statement determination module is used to input prompt information, user input statements, pattern subgraphs and candidate entity sets into the large language model to obtain multiple candidate database query statements.
[0017] The second filtering module is used to filter multiple candidate database query statements based on the pattern subgraph and the execution results corresponding to each candidate database query statement, so as to obtain the target database query statement.
[0018] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0019] At least one processor; and
[0020] A memory communicatively connected to the at least one processor; wherein,
[0021] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to execute the database query statement generation method according to any embodiment of the present invention.
[0022] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the database query statement generation method according to any embodiment of the present invention.
[0023] According to another aspect of the present invention, a computer program product is provided, which, when executed by a processor, implements the database query statement generation method as described in any of the embodiments of the present invention.
[0024] In existing technologies, most graph database queries still rely on users manually writing query statements. Users typically need to understand the structure, entity types, relation types, and related attributes of the graph database to correctly write NGQL query statements, which is a significant obstacle for non-professional users. In complex query scenarios, manually writing database query statements is time-consuming and prone to errors. To address the above problems, this invention proposes to obtain user input statements and knowledge graphs, extract information from the user input statements to obtain key information, filter the knowledge graph based on the key information to obtain a candidate entity set and a pattern subgraph, input the prompt information, user input statements, pattern subgraphs, and candidate entity sets into a large language model to obtain multiple candidate database query statements, and filter the multiple candidate database query statements according to the pattern subgraphs and the execution results corresponding to each candidate database query statement to obtain the target database query statement. Users do not need to master NGQL syntax to automatically generate database query statements that meet semantic requirements, improving the efficiency and accuracy of database query statement generation.
[0025] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0026] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0027] Figure 1 This is a flowchart of a database query statement generation method according to an embodiment of the present invention;
[0028] Figure 2 This is a schematic diagram of the structure of a database query statement generation device according to an embodiment of the present invention;
[0029] Figure 3 This is a schematic diagram of another database query statement generation device in an embodiment of the present invention;
[0030] Figure 4 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention. Detailed Implementation
[0031] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0032] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0033] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0034] Example 1
[0035] Figure 1 This is a flowchart illustrating a database query statement generation method provided in an embodiment of the present invention. This embodiment is applicable to the generation of database query statements. The method can be executed by the database query statement generation device in this embodiment, which can be implemented in software and / or hardware, such as... Figure 1 As shown, the method specifically includes the following steps:
[0036] S110, obtain user input statements and knowledge graphs.
[0037] In this embodiment, the user input statement is a natural language question that needs to be converted into a database query statement.
[0038] In this embodiment, the knowledge graph includes: schema information, entities, relationships between entities, and entity attributes. The schema information can be schema elements. In a knowledge graph, schema elements are the basic framework for building a structured knowledge system, used to define the types and rules of entities, relationships, and their attributes.
[0039] S120, extract information from the user input statement to obtain key information.
[0040] In this embodiment, the key information includes: target entity and entity type; the key information may also include: the relationship between target entities and the type of relationship between target entities; the key information may also include: the attributes of target entities and the attribute types of target entities.
[0041] In this embodiment, the method for extracting key information from the user input statement can be as follows: Keywords in the question are extracted using named entity recognition technology, and the extracted keywords are categorized to obtain various types of key information. For example, named entity recognition technology can be used to extract keywords in the question, and the extracted keywords can be categorized to obtain target entities, relationships between target entities, and attributes of target entities.
[0042] S130, the knowledge graph is filtered based on the key information to obtain a candidate entity set and a pattern subgraph.
[0043] In this embodiment, the method of filtering the knowledge graph based on the key information to obtain a candidate entity set and a pattern subgraph can be as follows: If the key information includes: target entity and entity type, then the entities in the knowledge graph are filtered based on the target entity to obtain a candidate entity set; the pattern information of the knowledge graph is truncated based on the entity type to obtain a pattern subgraph. If the key information includes: target entity, entity type, relationship between target entities, and relationship type between target entities, then the entities in the knowledge graph are filtered based on the relationship between target entities to obtain a candidate entity set; the pattern information of the knowledge graph is truncated based on the entity type and relationship type between target entities to obtain a pattern subgraph. If the key information includes: target entity, entity type, attribute of target entity, and attribute type of target entity, then the entities in the knowledge graph are filtered based on the target entity and attribute of target entity to obtain a candidate entity set; the pattern information of the knowledge graph is truncated based on the entity type and attribute type of target entity to obtain a pattern subgraph. If the key information includes: target entity, entity type, relationship between target entities, relationship type between target entities, attribute of target entity, and attribute type of target entity, then the entities in the knowledge graph are filtered based on the target entity, relationship between target entities, and attribute of target entity to obtain a candidate entity set; the pattern information of the knowledge graph is extracted based on the entity type, relationship type between target entities, and attribute type of target entity to obtain a pattern subgraph.
[0044] In this embodiment, the entity in the knowledge graph is filtered based on the target entity and the relationship between the target entities to obtain a candidate entity set. This can be done by filtering the entities in the knowledge graph based on the semantic similarity and string matching degree between the target entity and the entity in the knowledge graph, as well as the similarity between the relationship between the target entities and the relationship between the entities in the knowledge graph.
[0045] In this embodiment, the entity in the knowledge graph is filtered based on the target entity and its attributes to obtain a candidate entity set. This can be done by filtering the entities in the knowledge graph based on the semantic similarity and string matching degree between the target entity and the entities in the knowledge graph, as well as the similarity between the attributes of the target entity and the attributes of the entities in the knowledge graph.
[0046] In this embodiment, the entity in the knowledge graph is filtered based on the target entity, the relationship between target entities, and the attributes of the target entity to obtain a candidate entity set. This can be done by filtering the entities in the knowledge graph based on the semantic similarity and string matching degree between the target entity and the entities in the knowledge graph, the similarity between the relationship between the target entity and the relationship between the entities in the knowledge graph, and the similarity between the attributes of the target entity and the attributes of the entities in the knowledge graph.
[0047] In this embodiment, the pattern information of the knowledge graph is extracted based on the relationship type between the entity type and the target entity to obtain a pattern subgraph. This can be done by retaining only the part of the pattern information of the knowledge graph that is related to the relationship type between the entity type and the target entity as the pattern subgraph.
[0048] In this embodiment, the pattern information of the knowledge graph is extracted based on the entity type, the relationship type between the target entities, and the attribute type of the target entity to obtain a pattern subgraph. This can be done by retaining only the part of the pattern information of the knowledge graph that is related to the entity type, the relationship type between the target entities, and the attribute type of the target entity as the pattern subgraph.
[0049] In this embodiment, the pattern information of the knowledge graph is extracted based on the entity type and the attribute type of the target entity to obtain a pattern subgraph. This can be done by retaining only the part of the pattern information of the knowledge graph that is related to the entity type and the attribute type of the target entity as the pattern subgraph.
[0050] Optionally, the key information includes: the target entity and the entity type;
[0051] Based on the key information, the knowledge graph is filtered to obtain a candidate entity set and a pattern subgraph, including:
[0052] Based on the target entity, the entities in the knowledge graph are filtered to obtain a candidate entity set.
[0053] In this embodiment, the method of filtering entities in the knowledge graph based on the target entity to obtain a candidate entity set can be as follows: filtering entities in the knowledge graph based on the semantic similarity and string matching degree between the target entity and the entities in the knowledge graph to obtain a candidate entity set.
[0054] Based on the entity type, the pattern information of the knowledge graph is extracted to obtain a pattern subgraph.
[0055] This embodiment uses semantic alignment and association with knowledge graphs to quickly construct subgraphs related to the problem, reducing interference from irrelevant information.
[0056] Optionally, the entities in the knowledge graph are filtered based on the target entity to obtain a candidate entity set, including:
[0057] Based on the semantic similarity and string matching degree between the target entity and the entities in the knowledge graph, the entities in the knowledge graph are filtered to obtain a candidate entity set.
[0058] In this embodiment, the entity set in the knowledge graph is filtered based on the semantic similarity and string matching degree between the target entity and the entities in the knowledge graph. This can be achieved by: generating semantic vector representations of the target entity and entities in the knowledge graph using a pre-trained language model; calculating the semantic similarity between the target entity and entities in the knowledge graph based on these representations; determining the string matching degree between the name of the target entity and the name of an entity in the knowledge graph using an inverted index technique; weighted fusion of the semantic similarity and string matching degree between the target entity and entities in the knowledge graph; and filtering entities in the knowledge graph based on the weighted fusion result to obtain the candidate entity set.
[0059] S140: Input the prompt information, user input statement, pattern subgraph, and candidate entity set into the large language model to obtain multiple candidate database query statements.
[0060] In this embodiment, the prompt information can be pre-set sample data, and the prompt information includes: at least one set of information, each set of information including: a semantic skeleton sample and a database query statement corresponding to the semantic skeleton sample.
[0061] In this embodiment, the prompt information is first input into the large language model, and then the user input statement, pattern subgraph, and candidate entity set are input into the large language model to obtain multiple candidate database query statements.
[0062] In this embodiment, the prompt information can be a Few-shot example. A schema subgraph refers to a substructure related to the user's input statement extracted based on schema information (Schema) from a knowledge graph, used to guide query statement generation; for example, the schema subgraph can be a schema subgraph. This embodiment utilizes schema subgraphs and Few-shot examples to generate multiple candidate database query statements and iteratively optimizes them using a large language model.
[0063] Optionally, before inputting the prompt information, user input statement, pattern subgraph, and candidate entity set into the large language model to obtain multiple candidate database query statements, the following steps are also included:
[0064] The target entity in the user input statement is replaced with the identifier information corresponding to the entity type to obtain the target semantic skeleton;
[0065] Example data from the training set whose similarity to the target semantic skeleton is greater than a similarity threshold is used as prompt information. The example data includes: semantic skeleton samples and database query statements corresponding to the semantic skeleton samples.
[0066] In this embodiment, the target semantic skeleton can also be obtained as follows: If the key information includes: target entity, entity type, relationship between target entities, and relationship type, then the target entity in the user input statement is replaced with the identifier information corresponding to the entity type, and the relationship between target entities in the user input statement is replaced with the identifier information corresponding to the relationship type, thus obtaining the target semantic skeleton. If the key information includes: target entity, entity type, attribute type, and attribute of the target entity, then the target entity in the user input statement is replaced with the identifier information corresponding to the entity type, and the attribute of the target entity in the user input statement is replaced with the identifier information corresponding to the attribute type, thus obtaining the target semantic skeleton. If the key information includes: target entity, entity type, attribute type, attribute of the target entity, relationship between target entities, and relationship type, then the target entity in the user input statement is replaced with the identifier information corresponding to the entity type, the attribute of the target entity in the user input statement is replaced with the identifier information corresponding to the attribute type, and the relationship between target entities in the user input statement is replaced with the identifier information corresponding to the relationship type, thus obtaining the target semantic skeleton.
[0067] In this embodiment, the example data in the training set whose similarity to the target semantic skeleton is greater than a similarity threshold can be used as prompt information by: obtaining the semantic similarity and string matching degree between each semantic skeleton sample in the training set and the target semantic skeleton; determining the similarity between each semantic skeleton sample and the target semantic skeleton based on the semantic similarity and string matching degree; and using the example data to which the semantic skeleton samples in the training set whose similarity to the target semantic skeleton is greater than a similarity threshold belong as prompt information.
[0068] This embodiment combines few-shot learning and dynamic example selection mechanisms, without relying on fixed rules, and has strong generalization and flexibility. It can adapt to the graph structure of different fields and efficiently handle complex and diverse natural language problems.
[0069] This embodiment uses a dynamic example selection strategy based on semantic skeleton similarity, which can accurately select the most relevant examples to the input question, thereby improving the accuracy and generalization ability of the generated results.
[0070] S150, based on the pattern subgraph and the execution results corresponding to each candidate database query statement, multiple candidate database query statements are filtered to obtain the target database query statement.
[0071] In this embodiment, the execution result includes execution status information. The execution status information includes execution success and execution failure; if execution fails, the execution status information may further include alarm information.
[0072] In this embodiment, the database query statement can be a query statement specific to the NebulaGraph query language. For example, it can be an NGQL (Nebula Graph Query Language) query statement, which is used to manipulate graph data models.
[0073] In this embodiment, the method of filtering multiple candidate database query statements to obtain the target database query statement based on the pattern subgraph and the execution results corresponding to each candidate database query statement can be as follows: Execute the multiple candidate database query statements to obtain the execution results corresponding to each candidate database query statement; if the execution status information in the execution results corresponding to each candidate database query statement is all successful, then the multiple candidate database query statements are filtered based on the pattern subgraph and the execution results corresponding to each candidate database query statement to obtain the target database query statement. Alternatively, the method of filtering multiple candidate database query statements to obtain the target database query statement based on the pattern subgraph and the execution results corresponding to each candidate database query statement can be as follows: If there are candidate database query statements with execution status information indicating execution failure, then the candidate database query statements with execution status information indicating successful execution are filtered based on the pattern subgraph and the execution results corresponding to the candidate database query statements with execution status information indicating successful execution to obtain the target database query statement. The method of filtering multiple candidate database query statements to obtain the target database query statement based on the pattern subgraph and the execution results corresponding to each candidate database query statement can also be as follows: If there are candidate database query statements whose execution status information is "failed", then perform an iterative update operation on the failed candidate database query statements to obtain the updated candidate database query statements corresponding to the failed candidate database query statements. Then, based on the pattern subgraph, the execution results corresponding to the successfully executed candidate database query statements, and the execution results corresponding to the updated candidate database query statements, filter the successfully executed candidate database query statements and the updated candidate database query statements to obtain the target database query statement.
[0074] Optionally, based on the pattern subgraph and the execution results corresponding to each candidate database query statement, multiple candidate database query statements are filtered to obtain the target database query statement, including:
[0075] Execute the multiple candidate database query statements to obtain the execution results corresponding to each candidate database query statement;
[0076] If the execution status information in the execution results of each candidate database query statement is successful, then the pattern subgraph, multiple candidate database query statements, and the execution results corresponding to each candidate database query statement are input into the large language model to obtain the target database query statement.
[0077] In this embodiment, the input of the large language model is the pattern subgraph, multiple candidate database query statements, and the execution results corresponding to each candidate database query statement, and the output of the large language model is the target database query statement.
[0078] In this embodiment, if the execution status information in the execution results of each candidate database query statement is successful, then the target database query statement is obtained by selecting from multiple candidate database query statements based on the large language model.
[0079] Optional, also includes:
[0080] If the execution status information in the execution result corresponding to any candidate database query statement is execution failure, then perform an iterative update operation on the candidate database query statement that failed to execute, and obtain the updated candidate database query statement corresponding to the candidate database query statement that failed to execute.
[0081] The pattern subgraph, the successfully executed candidate database query statements, the updated candidate database query statements corresponding to the failed candidate database query statements, the execution results corresponding to the successfully executed candidate database query statements, and the execution results corresponding to the updated candidate database query statements are input into the large language model to obtain the target database query statement.
[0082] In this embodiment, an iterative update operation is performed on the candidate database query statements that failed to execute. It should be noted that the iteration termination condition of the iterative update operation includes any of the following:
[0083] Reaching the maximum number of iterations;
[0084] The execution status information in the execution result corresponding to the updated candidate database query statement is "execution successful".
[0085] It should be noted that if, after performing an iterative update operation on the candidate database query statement that failed to execute, the execution status information in the execution result corresponding to the updated candidate database query statement is still "execution failed", then the pattern subgraph, the candidate database query statement that was successfully executed, and the execution result corresponding to the candidate database query statement that was successfully executed are input into the large language model to obtain the target database query statement.
[0086] In this embodiment, the iterative update operation can be implemented using a Self-Refiner.
[0087] The solution provided in this embodiment significantly improves the accuracy and reliability of the final generated statements through query optimization and evaluation mechanisms.
[0088] Optionally, an iterative update operation is performed on the candidate database query statements that failed to execute, to obtain the updated candidate database query statements corresponding to the failed candidate database query statements, including:
[0089] Based on the pattern subgraph, the failed candidate database query statements, and the execution results corresponding to the failed candidate database query statements, the updated candidate database query statements are determined.
[0090] Execute the updated candidate database query statement. If the execution status information in the execution result corresponding to the updated candidate database query statement is execution failure, then return the execution based on the updated candidate database query statement. According to the pattern subgraph, the failed candidate database query statement and the execution result corresponding to the failed candidate database query statement, determine the operation of the updated candidate database query statement until the execution status information in the execution result corresponding to the updated candidate database query statement is execution success.
[0091] The successfully executed updated candidate database query statement will be used as the updated candidate database query statement corresponding to the failed candidate database query statement.
[0092] It should be noted that for any update operation, if the execution status information in the execution result of the candidate database query statement after the update is "execution failed", then the next update operation will be executed.
[0093] In this embodiment, the updated candidate database query statement can be determined by inputting the pattern subgraph, the failed candidate database query statement, and the execution result corresponding to the failed candidate database query statement into the large language model to obtain the updated candidate database query statement.
[0094] Optionally, based on the pattern subgraph, the failed candidate database query statements, and the execution results corresponding to the failed candidate database query statements, the updated candidate database query statements are determined, including:
[0095] The pattern subgraph, the failed candidate database query statements, and the execution results corresponding to the failed candidate database query statements are input into the large language model to obtain the updated candidate database query statements.
[0096] In a specific example, such as Figure 2As shown, this invention proposes a database query statement generation device based on a large language model, comprising three main modules: a Schema Linking module, an NGQL CandidateGeneration module, and an NGQL Candidate Selection module. The Schema Linking module aims to identify entities in the user input statement (Query), or at least one of entities, relations, and attributes, based on the user input statement and the schema information of the knowledge graph, and associate them with schema elements in the knowledge graph, ultimately constructing a schema subgraph related to the user input statement. The implementation steps are as follows: 1. Question parsing and entity recognition: Parse the input natural language question (user input statement) to extract potential entity types, or at least one of entity types, relation types, and attribute types. Use named entity recognition technology to extract keywords from the natural language question and classify them as entities, or at least one of entities, relations, and attributes. 2. Entity Linking: Match the entities identified in the question with entities in the knowledge graph, calculate similarity to obtain candidate entities. Similarity calculation methods include: 1. Embedding-based semantic similarity calculation: A pre-trained language model generates semantic vector representations of the target entity and entities in the knowledge graph. The semantic similarity between the target entity and entities in the knowledge graph is calculated based on these representations. 2. Rule-based matching: Using an inverted index, the string matching degree between the name of the target entity and the name of an entity in the knowledge graph is obtained. The results of embedding and rule matching are weighted and fused (i.e., semantic similarity and string matching degree are weighted and fused) to generate the final candidate entity list. 3. Pattern subgraph construction: Based on the entity type, relation type, and attribute type in the user input statement, the original knowledge graph schema is slimmed down, retaining only the parts relevant to the user input statement, and a pattern subgraph is constructed. The construction of the subgraph reduces interference from irrelevant information and improves generation efficiency. The functional goal of the NGQLCandidate Generation module is to generate multiple candidate NGQL query statements (e.g., candidate NGQL A, candidate NGQL B, and candidate NGQL C) based on the schema subgraph and user input. Implementation steps: 1. Few-shot example selection: Employing a few-shot learning approach, the generator is provided with example data through a dynamic example selection mechanism. Dynamic example selection strategy: Skeleton similarity calculation: Entities in the question are replaced with type tags (e.g., "Google" is replaced with "...). <company>1. **Extract the target semantic skeleton of the question:** Embedding technology is used to calculate the similarity between the target semantic skeleton and semantic skeleton samples in the training set. Based on the similarity score, the example data containing the Top-K semantic skeleton samples most similar to the target semantic skeleton are selected as Few-shot examples. 2. **Generate candidate NGQL query statements:** Based on the Few-shot examples, a large language model is used to generate multiple candidate NGQL query statements, covering different grammatical styles and logical paths. 3. **NGQL Refiner Optimization:** Logical and syntactic optimization is performed on the generated candidate NGQL queries. Using a pattern subgraph, the generated candidate NGQL query statements, and their execution results as input, the large language model iteratively optimizes the candidate query statements. The functional goal of the NGQL Candidate Selection module is to select the optimal database query statement (final candidate NGQL) from multiple candidate NGQL query statements. Implementation steps: Input includes all candidate NGQL query statements, a pattern subgraph, and the execution results of all candidate NGQL query statements. The large model, combined with context learning, is used to select the candidate database query statement.
[0097] In this embodiment, the Schema Linking module proposes an entity linking method that combines embedding-based semantic similarity calculation with rule matching. By constructing a pattern subgraph related to the question, interference from irrelevant information is reduced, improving generation efficiency. The semantic and rule-based Schema Linking method and its application in graph database query generation are also discussed.
[0098] The database query statement generation method provided in this invention allows users to efficiently generate query statements using natural language without needing to master NGQL syntax. Furthermore, through query optimization and evaluation mechanisms, the accuracy and reliability of the final generated statements are significantly improved. By semantic alignment and association with knowledge graphs, problem-related subgraphs are quickly constructed, reducing interference from irrelevant information. By combining few-shot learning and dynamic example selection mechanisms, it does not rely on fixed rules, possesses strong generalization and flexibility, can adapt to different domain graph structures, and efficiently handles complex and diverse natural language problems.
[0099] To address the limitations of existing database query language generation technologies in terms of complexity and flexibility, this invention proposes a method for generating NGQL query statements based on a large language model. This method can automatically generate NGQL query statements that meet the requirements of graph database queries based on the user's input natural language question. Users do not need to master NGQL syntax to efficiently and accurately generate complex NGQL query statements. In existing technologies, most graph database queries still rely on users manually writing query statements. Users typically need to understand the structure, entity types, relation types, and related attributes of the graph database to correctly write NGQL query statements, which is a significant obstacle for non-professional users. In complex query scenarios, manually writing database query statements is time-consuming and prone to errors. To address the aforementioned issues, this invention proposes a method that involves acquiring user input statements and a knowledge graph, extracting key information from the user input statements, filtering the knowledge graph based on the key information to obtain a candidate entity set and a pattern subgraph, inputting the prompt information, user input statements, pattern subgraph, and candidate entity set into a large language model to obtain multiple candidate database query statements, and filtering these multiple candidate database query statements based on the pattern subgraph and the execution results corresponding to each candidate database query statement to obtain the target database query statement. This method automatically generates database query statements that meet semantic requirements without requiring the user to master NGQL syntax, thus improving the efficiency and accuracy of database query statement generation.
[0100] Example 2
[0101] Figure 3 This is a schematic diagram of another database query statement generation device provided in an embodiment of the present invention. This embodiment is applicable to database query statement generation. The device can be implemented using software and / or hardware, and can be integrated into any device that provides database query statement generation functionality, such as… Figure 3 As shown, the database query statement generation device specifically includes: an acquisition module 310, an information extraction module 320, a first filtering module 330, a candidate database query statement determination module 340, and a second filtering module 350.
[0102] The acquisition module is used to acquire user input statements and knowledge graphs.
[0103] The information extraction module is used to extract information from the user input statement to obtain key information;
[0104] The first filtering module is used to filter the knowledge graph based on the key information to obtain a candidate entity set and a pattern subgraph.
[0105] The candidate database query statement determination module is used to input prompt information, user input statements, pattern subgraphs and candidate entity sets into the large language model to obtain multiple candidate database query statements.
[0106] The second filtering module is used to filter multiple candidate database query statements based on the pattern subgraph and the execution results corresponding to each candidate database query statement, so as to obtain the target database query statement.
[0107] The above-described products can perform the methods provided in any embodiment of the present invention, and have the corresponding functional modules and beneficial effects for performing the methods.
[0108] Example 3
[0109] Figure 4 A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0110] like Figure 4 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0111] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0112] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as database query statement generation methods.
[0113] In some embodiments, the database query statement generation method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the database query statement generation method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to execute the database query statement generation method by any other suitable means (e.g., by means of firmware).
[0114] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0115] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0116] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0117] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0118] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0119] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0120] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0121] This invention also provides a computer program product, including a computer program that, when executed by a processor, implements the database query statement generation method according to any embodiment of the invention.
[0122] In implementing the computer program product, computer program code for performing the operations of this invention can be written in one or more programming languages or a combination thereof. Programming languages include object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0123] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.< / company>
Claims
1. A method for generating database query statements, characterized in that, include: Obtain user input statements and knowledge graphs; Information is extracted from the user input statement to obtain key information; Based on the key information, the knowledge graph is filtered to obtain a candidate entity set and a pattern subgraph; Input the prompts, user input statements, pattern subgraphs, and candidate entity sets into the large language model to obtain multiple candidate database query statements; Based on the pattern subgraph and the execution results corresponding to each candidate database query statement, multiple candidate database query statements are filtered to obtain the target database query statement.
2. The method according to claim 1, characterized in that, Based on the pattern subgraph and the execution results corresponding to each candidate database query statement, multiple candidate database query statements are filtered to obtain the target database query statement, including: Execute the multiple candidate database query statements to obtain the execution results corresponding to each candidate database query statement; If the execution status information in the execution results of each candidate database query statement is successful, then the pattern subgraph, multiple candidate database query statements, and the execution results corresponding to each candidate database query statement are input into the large language model to obtain the target database query statement.
3. The method according to claim 2, characterized in that, Also includes: If the execution status information in the execution result corresponding to any candidate database query statement is execution failure, then perform an iterative update operation on the candidate database query statement that failed to execute, and obtain the updated candidate database query statement corresponding to the candidate database query statement that failed to execute. The pattern subgraph, the successfully executed candidate database query statements, the updated candidate database query statements corresponding to the failed candidate database query statements, the execution results corresponding to the successfully executed candidate database query statements, and the execution results corresponding to the updated candidate database query statements are input into the large language model to obtain the target database query statement.
4. The method according to claim 3, characterized in that, Perform iterative update operations on the candidate database query statements that failed to execute, to obtain the updated candidate database query statements corresponding to the failed candidate database query statements, including: Based on the pattern subgraph, the failed candidate database query statements, and the execution results corresponding to the failed candidate database query statements, the updated candidate database query statements are determined. Execute the updated candidate database query statement. If the execution status information in the execution result corresponding to the updated candidate database query statement is execution failure, then return the execution based on the updated candidate database query statement. According to the pattern subgraph, the failed candidate database query statement and the execution result corresponding to the failed candidate database query statement, determine the operation of the updated candidate database query statement until the execution status information in the execution result corresponding to the updated candidate database query statement is execution success. The successfully executed updated candidate database query statement will be used as the updated candidate database query statement corresponding to the failed candidate database query statement.
5. The method according to claim 4, characterized in that, Based on the pattern subgraph, the failed candidate database query statements, and the execution results corresponding to the failed candidate database query statements, the updated candidate database query statements are determined, including: The pattern subgraph, the failed candidate database query statements, and the execution results corresponding to the failed candidate database query statements are input into the large language model to obtain the updated candidate database query statements.
6. The method according to claim 1, characterized in that, Before inputting the prompts, user input statements, pattern subgraphs, and candidate entity sets into the large language model to obtain multiple candidate database query statements, the following steps are also included: The target entity in the user input statement is replaced with the identifier information corresponding to the entity type to obtain the target semantic skeleton; Example data from the training set whose similarity to the target semantic skeleton is greater than a similarity threshold is used as prompt information. The example data includes: semantic skeleton samples and database query statements corresponding to the semantic skeleton samples.
7. The method according to claim 1, characterized in that, The key information includes: the target entity and the entity type; Based on the key information, the knowledge graph is filtered to obtain a candidate entity set and a pattern subgraph, including: Based on the target entity, the entities in the knowledge graph are filtered to obtain a candidate entity set; Based on the entity type, the pattern information of the knowledge graph is extracted to obtain a pattern subgraph.
8. The method according to claim 7, characterized in that, Based on the target entity, entities in the knowledge graph are filtered to obtain a candidate entity set, including: Based on the semantic similarity and string matching degree between the target entity and the entities in the knowledge graph, the entities in the knowledge graph are filtered to obtain a candidate entity set.
9. A database query statement generation device, characterized in that, include: The acquisition module is used to acquire user input statements and knowledge graphs; The information extraction module is used to extract information from the user input statement to obtain key information; The first filtering module is used to filter the knowledge graph based on the key information to obtain a candidate entity set and a pattern subgraph. The candidate database query statement determination module is used to input prompt information, user input statements, pattern subgraphs and candidate entity sets into the large language model to obtain multiple candidate database query statements. The second filtering module is used to filter multiple candidate database query statements based on the pattern subgraph and the execution results corresponding to each candidate database query statement, so as to obtain the target database query statement.
10. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the database query statement generation method according to any one of claims 1-8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed by a processor, implement the database query statement generation method according to any one of claims 1-8.
12. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the database query statement generation method according to any one of claims 1-8.
Citation Information
Cited By
Context-aware SQL statement generation method and device
CN121705305A
Context-aware sql statement generation method and apparatus
CN121705305B