Database query statement generation method and device, electronic equipment and storage medium
By sampling and connecting edge detection of the database model graph, the target node is built, and high-quality test data is generated, the problem of poor training effect of Text-to-SQL model is solved, and more accurate statement conversion is achieved.
Patent Information
- Application Number
- CN202510185637.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-07-01
AI Technical Summary
In the prior art, the training effect of the Text-to-SQL model is poor, and the lack of high-quality test data is lacking, resulting in insufficient accuracy of statement conversion.
By obtaining the database model diagram, performing graph sampling to obtain candidate nodes, selecting target data columns, detecting the effectiveness of connection edges, adjusting target nodes, building database model subgraphs, and using large language models to generate test data, including structured query statements and natural language query statements.
It improves the semantic integrity and rationality of generated test data, enhances the training effect of the Text-to-SQL model, and improves the accuracy of statement conversion.
Smart Images

Figure CN120234337A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a method and device for generating database query statements, an electronic device, and a storage medium. Background Art
[0002] The text-to-structured query language conversion technology (Text-to-SQL) can convert the natural expression of human language into a structured query language (SQL). This means that users do not need to deeply master complex SQL syntax and database structure knowledge. They only need to clearly describe their query requirements in natural language, and the Text-to-SQL model can automatically generate the corresponding structured query statement.
[0003] Currently, the research in related technologies focuses on using various frameworks and strategies to train the Text-to-SQL model. However, due to the lack of high-quality test data, the training effect of the model is not good, which affects the accuracy of statement conversion. Therefore, how to generate high-quality test data has become an urgent technical problem to be solved. Summary of the Invention
[0004] The main purpose of the embodiments of this application is to propose a method and device for generating database query statements, an electronic device, and a storage medium, aiming to generate high-quality test data.
[0005] To achieve the above object, a first aspect of the embodiments of this application proposes a method for generating a database query statement, and the method includes:
[0006] Obtain a database model diagram of a target database, where the target database includes data tables, the data tables include data columns, the database model diagram includes an initial node representing the data table, and the connection edges of the initial node represent the mapping relationship between the data columns;
[0007] Perform graph sampling on the initial nodes in the database model diagram to obtain a sampling result, where the sampling result includes candidate nodes;
[0008] Select target data columns from the data columns of the candidate nodes, and construct a target node based on the target data columns;
[0009] Detect the validity of the connection edges of the target node to obtain validity data;
[0010] Adjust the data columns in the target node based on the validity data;
[0011] Construct a database model sub-diagram based on the target node;
[0012] Based on the database model sub-graph, test data is generated through a large language model, and the test data includes structured query statement samples and natural language query statement samples.
[0013] In some embodiments, detecting the validity of the connection edges of the target node to obtain validity data includes:
[0014] Based on the connection edges of the target node, detecting the mapping objects of the data columns in the target node in the sampling result to obtain validity data, where the validity data indicates the existence or non-existence of the mapping objects.
[0015] In some embodiments, it further includes:
[0016] Obtaining a test object, which is used to convert a natural language query statement into a structured query statement;
[0017] Inputting the natural language query statement sample and the database model sub-graph into the test object to obtain a target structured query statement;
[0018] Based on the structured query statement sample, performing reliability verification on the target structured query statement to obtain the reliability data of the target structured query statement.
[0019] In some embodiments, the step of inputting the natural language query statement sample and the database model sub-graph into the test object to obtain a target structured query statement includes:
[0020] Performing data enhancement processing on the database model sub-graph to obtain a target database model sub-graph;
[0021] Inputting the natural language query statement sample and the target database model sub-graph into the test object to obtain a target structured query statement.
[0022] In some embodiments, the step of performing data enhancement processing on the database model sub-graph to obtain a target database model sub-graph includes:
[0023] Modifying the connection edges in the database model sub-graph to change the mapping relationship between data columns to obtain a target database model sub-graph.
[0024] In some embodiments, the step of inputting the natural language query statement sample and the database model sub-graph into the test object to obtain a target structured query statement includes:
[0025] Performing data enhancement processing on the natural language query statement to obtain a target natural language query statement;
[0026] Input the target natural language query statement and the database model sub-graph into the object under test to obtain a target structured query statement.
[0027] In some embodiments, obtaining the database model graph of the target database includes:
[0028] Obtain the database metadata of the target database;
[0029] Based on the database metadata, construct the database model graph of the target database.
[0030] To achieve the above object, a second aspect of the embodiments of the present application proposes a database query statement generation device, the device includes:
[0031] A data acquisition module, configured to acquire the database model graph of the target database, the target database includes data tables, the data tables include data columns, the database model graph includes an initial node representing the data table, and the connection edge of the initial node represents the mapping relationship between the data columns;
[0032] A graph sampling module, configured to perform graph sampling on the initial nodes in the database model graph to obtain a sampling result, the sampling result includes candidate nodes;
[0033] A target node generation module, configured to select target data columns from the data columns of the candidate nodes, and construct target nodes based on the target data columns;
[0034] A connection edge detection module, configured to detect the validity of the connection edges of the target nodes to obtain validity data;
[0035] A target node adjustment module, configured to adjust the data columns in the target nodes based on the validity data;
[0036] A database model sub-graph construction module, configured to construct a database model sub-graph based on the target nodes;
[0037] A test data generation module, configured to generate test data based on the database model sub-graph through a large language model, the test data includes structured query statement samples and natural language query statement samples.
[0038] To achieve the above object, a third aspect of the embodiments of the present application proposes an electronic device, the electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the method described in the first aspect above is implemented.
[0039] To achieve the above object, a fourth aspect of the embodiments of the present application provides a computer-readable storage medium storing a computer program, which when executed by a processor implements the method described in the first aspect above.
[0040] The database query statement generation method and apparatus, electronic device, and storage medium provided by the present application perform graph sampling on a database model diagram to obtain candidate nodes and select target data columns from the candidate nodes, and construct target nodes to reduce the database structure information to be processed subsequently. At the same time, considering that the database model diagram is different from a conventional graph structure, it is necessary to ensure the mapping relationship between data columns in the nodes, detect the validity of the connection edges of the target nodes, and adjust the target nodes based on this to improve the reliability of the database model sub-diagram, thereby improving the semantic integrity and rationality of the subsequently generated test data. Finally, based on the database model sub-diagram, a large number of test data are generated. Description of the Drawings
[0041] Figure 1 is a flowchart of the database query statement generation method provided by the embodiments of the present application;
[0042] Figure 2 is Figure 1 a flowchart of step S101 in
[0043] Figure 3 is another flowchart of the database query statement generation method provided by the embodiments of the present application;
[0044] Figure 4 is Figure 3 a flowchart of step S402 in
[0045] Figure 5 is a schematic structural diagram of the database query statement generation apparatus provided by the embodiments of the present application;
[0046] Figure 6 is a schematic hardware structure diagram of the electronic device provided by the embodiments of the present application. Detailed Embodiments
[0047] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0048] It should be noted that although the functional modules are divided in the device schematic diagram and the logical sequence is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the device or the sequence in the flowchart. Terms such as "first" and "second" in the specification, claims, and the above-mentioned drawings are used to distinguish similar objects and do not necessarily describe a specific order or sequence.
[0049] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0050] First, several nouns involved in this application are analyzed:
[0051] Database Model Diagram (Schema): A Schema is a visual diagram of the database structure, showing the data tables, data columns (fields), and the relationships between data tables. In related technologies, database management software (Navicat) is used to extract the data table structure and data table relationships to generate a Schema.
[0052] Graph Sampling (GS): GS extracts representative vertices or edges from a large-scale graph structure to construct an approximate subgraph. This technology sacrifices some accuracy in exchange for improved efficiency and is suitable for processing ultra-large-scale graph data. It is mainly applied to social network analysis (user relationship mining), traffic network optimization (path planning), recommendation systems (processing user behavior graphs), and bioinformatics (studying protein interaction networks), while maintaining the key attributes of the graph structure while ensuring computational feasibility. In related technologies, graph sampling can be classified into vertex sampling (randomly selecting nodes), edge sampling (extracting connecting edges), and hybrid sampling according to the object.
[0053] Text-to-Structured Query Language Conversion Model (Text-to-SQL): Given a natural language question and the corresponding database model, Text-to-SQL generates a structured query statement that semantically matches the natural language question. In related technologies, the Text-to-SQL model based on deep learning methods uses an encoder-decoder architecture and captures the semantic alignment between the question and the database model through attention mechanisms and graph neural networks; the Text-to-SQL model based on large language models (LLMs) utilizes the generation ability of pre-trained models to directly generate SQL queries through prompt engineering or supervised fine-tuning, significantly improving cross-domain generalization ability.
[0054] Natural Language Processing (NLP): NLP uses computers to process, understand, and apply human languages (such as Chinese, English, etc.). NLP is a branch of artificial intelligence and an interdisciplinary field of computer science and linguistics, often referred to as computational linguistics. Natural language processing includes syntactic analysis, semantic analysis, discourse understanding, etc. Natural language processing is commonly used in technical fields such as machine translation, handwritten and printed character recognition, speech recognition and text-to-speech conversion, information intent recognition, information extraction and filtering, text classification and clustering, public opinion analysis, and opinion mining. It involves data mining related to language processing, machine learning, knowledge acquisition, knowledge engineering, artificial intelligence research, and linguistic research related to language computing, etc.
[0055] Data Augmentation: Data Augmentation is a technique that generates diverse new samples by performing artificial or automated transformations (such as rotation, cropping, noise addition, etc.) on the original data. Its core goal is to expand the scale and diversity of the training dataset, thereby enhancing the generalization ability of the model and reducing the risk of overfitting.
[0056] Based on this, the embodiments of the present application provide a method and apparatus for generating database query statements, an electronic device, and a storage medium, aiming to generate high-quality test data.
[0057] A method and apparatus for generating database query statements, an electronic device, and a storage medium provided by the embodiments of the present application are specifically described through the following embodiments. First, a method for generating database query statements in the embodiments of the present application is described.
[0058] The embodiments of the present application can obtain and process relevant data based on artificial intelligence technology. Among them, Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.
[0059] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0060] A method for generating a database query statement provided by an embodiment of the present application relates to the field of artificial intelligence technology. The method for generating a database query statement provided by an embodiment of the present application can be applied to a terminal, or can be applied to a server side, or can also be software running on a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, etc.; the server side can be configured as an independent physical server, or can be configured as a server cluster or a distributed system composed of multiple physical servers, or can also be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application for implementing a method for generating a database query statement, etc., but is not limited to the above forms.
[0061] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet-type devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment, where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0062] It should be noted that in each specific embodiment of the present application, when it comes to relevant processing that needs to be based on data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when an embodiment of the present application needs to obtain the user's sensitive personal information, the user's separate permission or separate consent will be obtained through methods such as pop-up windows or jumping to a confirmation page. After clearly obtaining the user's separate permission or separate consent, the necessary user-related data for the normal operation of the embodiment of the present application will be obtained.
[0063] Figure 1 is an optional flowchart of the method for generating a database query statement provided by an embodiment of the present application. Figure 1The method in [it] may include but is not limited to steps S101 to S107.
[0064] Step S101: Obtain a database model diagram of a target database. The target database includes data tables, the data tables include data columns, the database model diagram includes initial nodes representing the data tables, and the connection edges of the initial nodes represent the mapping relationships between the data columns;
[0065] Step S102: Perform graph sampling on the initial nodes in the database model diagram to obtain a sampling result, where the sampling result includes candidate nodes;
[0066] Step S103: Select target data columns from the data columns of the candidate nodes, and construct target nodes based on the target data columns;
[0067] Step S104: Detect the validity of the connection edges of the target nodes to obtain validity data;
[0068] Step S105: Adjust the data columns in the target nodes based on the validity data;
[0069] Step S106: Construct a database model sub-diagram based on the target nodes;
[0070] Step S107: Generate test data through a large language model based on the database model sub-diagram. The test data includes structured query statement samples and natural language query statement samples.
[0071] It is easy to understand that the target database includes several data tables, and the data tables are two-dimensional structures composed of data rows (records) and data columns (fields). Each data column represents the common attributes of an object (such as "name", "age", etc.), while the data rows represent specific data instances. If a data column in a data table references a data column in another data table, it is considered that there is a mapping relationship between these two data columns, a connection edge is constructed between the corresponding nodes of the above two data tables, and the reference relationship between these two data columns is written into the edge attributes of the connection edge. Specifically, the reference relationship between two data columns can be confirmed by looking up foreign key associations.
[0072] Steps S101 to S107 illustrated in the embodiments of this application perform graph sampling on the database model diagram to obtain candidate nodes and select target data columns from the candidate nodes, and construct target nodes to reduce the database structure information to be processed subsequently; at the same time, considering that the database model diagram is different from a conventional graph structure, it is necessary to ensure the mapping relationship between the data columns in the nodes, detect the validity of the connection edges of the target nodes, and adjust the target nodes based on this to improve the reliability of the database model sub-diagram, thereby improving the semantic integrity and rationality of the subsequently generated test data; finally, based on the database model sub-diagram, a large amount of test data is generated.
[0073] Please refer to Figure 2 , in some embodiments, step S101 may include but is not limited to steps S201 to S205:
[0074] Step S201, obtain the database metadata of the target database;
[0075] Step S202, based on the database metadata, construct a database model diagram of the target database.
[0076] Understandably, database metadata refers to the data that defines the database structure, which describes the information of various object structures in the database, such as database name, data table name, data column name, mapping relationship between data columns, user name, version name, etc. For example, in an e-commerce database, there may be data tables such as user table, order table, and product table.
[0077] Specifically, based on the database metadata, construct a database model diagram G=(V, E), where the node set V represents the data tables in the database, and the edge set E represents the mapping relationship between data columns.
[0078] Steps S201 to S202 illustrated in the embodiments of the present application, by obtaining the database metadata of the target database and constructing a database model diagram, on the one hand, can dynamically synchronize metadata changes, support real-time modeling of the database for specific application scenarios, provide reliable input for subsequent test data generation, and improve automation efficiency.
[0079] In step S102 of some embodiments, at least one initial node may be randomly selected in the database model diagram as a candidate node; or at least one initial node may be selected in the database model diagram based on the node degree; not limited thereto.
[0080] In step S103 of some embodiments, at least one data column may be randomly selected from the candidate nodes as the target data column; or at least one data column may be selected from the candidate nodes based on the data column priority; not limited thereto.
[0081] In step S104 of some embodiments, it may include but is not limited to step S301:
[0082] Step S301, based on the connection edges of the target node, detect the mapping objects of the data columns in the target node in the sampling result to obtain validity data, and the validity data indicates the existence or non-existence of the mapping object.
[0083] Step S301 shown in the embodiments of this application ensures the integrity of data column references by detecting whether the mapping objects of the data columns in the target nodes exist in the sampling results, reduces the invalidation of mapping relationships, provides high-confidence data column relationships for subsequent adjustment of test data generation, and improves the quality of test data.
[0084] In step S301 of some embodiments, each target node can be traversed in the sampling results, and all reachable nodes can be traversed along the connection edges for the currently accessed target node. If the traversal path contains the mapping object, it is marked as existing; otherwise, it is marked as non-existing. Alternatively, the foreign key associations of the target nodes can be obtained through database query instructions to get the mapping objects of the data columns in the target nodes, and the mapping objects can be queried in the sampling results. If a hit is found, it is marked as existing; otherwise, it is marked as non-existing. This is not limited to this.
[0085] In step S105 of some embodiments, each data column in the target node is traversed. If the validity data indicates that the data table associated with the foreign key of the currently accessed target data column is not in the sampling results, then the currently accessed target data column is deleted from the sampling results. If the validity data indicates that the data table associated with the foreign key of the currently accessed target data column is in the sampling results, but the data column associated with the foreign key is not in the sampling results, then the missing data column is added to the sampling results.
[0086] In step S106 of some embodiments, the target nodes are organized into a structured network according to the connection edges between the target nodes to obtain a subgraph of the database model.
[0087] In step S107 of some embodiments, a test data prompt can be constructed, and the test data prompt and the subgraph of the database model are output to a large language model. The large language model generates test data through the large model Chain-of-Thought to improve the accuracy of the generated test data.
[0088] For example, construct a test data prompt: Given the subgraph of the database model of the user table and the order table, generate an SQL query involving order amount statistics and its corresponding natural language question.
[0089] The Chain-of-Thought generated by the large language model: The number of users whose order amount exceeds a certain threshold can be counted; the user table and the order table need to be JOINed; the COUNT aggregation function is used to count the number of users; the WHERE clause is used to filter the order amount.
[0090] Output of the large language model: Sample of structured query statement: "SELECT COUNT(DISTINCT u.id) FROM users u JOIN orders o ON u.id = o.user_id WHERE o.total_amount > 1000"; Sample of natural language query statement: "How many users have orders with an amount exceeding 1000 yuan?"
[0091] In some embodiments, the method for generating a database query statement may further include: executing the above sample of the structured query statement through an SQL executor to obtain a query result; filtering out the sample of the structured query statement with a syntax error or an empty query result to further improve the quality of the generated test data.
[0092] Please refer to Figure 3 , after step 107 of some embodiments, the method for generating a database query statement may further include steps S401 to S403:
[0093] Step S401, obtaining a test object, where the test object is used to convert a natural language query statement into a structured query statement;
[0094] Step S402, inputting the sample of the natural language query statement and the sub-graph of the database model into the test object to obtain a target structured query statement;
[0095] Step S403, based on the sample of the structured query statement, performing reliability verification on the target structured query statement to obtain reliability data of the target structured query statement.
[0096] Steps S401 to S403 illustrated in the embodiments of the present application use the generated test data to perform reliability verification on the target structured query statement output by the test object to obtain reliability data of the target structured query statement, which can automatically verify the output accuracy of the test object and contribute to improving the training effect of the test object.
[0097] In some embodiments, step S402 may further include but is not limited to steps S501 to S502:
[0098] Step S501, performing data augmentation processing on the sub-graph of the database model to obtain a target sub-graph of the database model;
[0099] Step S502, inputting the sample of the natural language query statement and the target sub-graph of the database model into the test object to obtain a target structured query statement.
[0100] Steps S501 to S502 illustrated in the embodiments of this application perform data enhancement processing through the database model sub-graph to obtain the target database model sub-graph, enhancing the data diversity of the sub-graph; then input the target database model sub-graph into the test object to detect the fault tolerance ability of the test object for the noise relationships (such as redundant edges, missing attributes) in the sub-graph, and optimize the compatibility of the model for heterogeneous data sources.
[0101] In step S501 of some embodiments, the database model sub-graph can be enhanced by means such as replacing table column names with synonyms (replacing table names or column names with synonyms), replacing table column names with abbreviations (using abbreviated table names or column names), changing the table column naming specification (such as changing from camel case to underscore case), removing explicit foreign key definitions, adding tables or columns, randomly sorting table columns (randomly shuffling the order of table column definitions in the database model diagram), splitting and merging tables (splitting one table into multiple tables or merging multiple tables), etc., and is not limited to this. For example, the database model sub-graph includes a users table and a user_id column. After applying table column name synonym replacement, the target database model sub-graph includes a customers table and a customer_id column; after applying table splitting transformation: the target database model sub-graph includes a users_basic table (basic information) and a users_address table (address information).
[0102] In some embodiments, step S501 includes but is not limited to step S601:
[0103] Step S601 modifies the connection edges in the database model sub-graph to change the mapping relationship between data columns, obtaining the target database model sub-graph.
[0104] Specifically, modifying the connection edges can be deleting the mapping relationship between data columns and correspondingly deleting the connection edges between data tables. It is easy to understand that by deleting the connection edges, it is verified whether the test object can semantically judge whether there is a mutually dependent foreign key association between nodes through the table name / column name. For example, the id in the Students table and the student_id in the Scores table can be judged to have a foreign key association from the semantics.
[0105] Step S601 illustrated in the embodiments of this application dynamically adjusts the mapping relationship between data columns by modifying the connection edges in the database model sub-graph, verifying the robustness and adaptability of the test object to structural changes.
[0106] In some embodiments, step S402 can include but is not limited to steps S701 to S702:
[0107] Step S701 performs data enhancement processing on the natural language query statement to obtain the target natural language query statement;
[0108] Step S702: Input the target natural language query statement and the database model sub-graph into the test object to obtain the target structured query statement.
[0109] In steps S701 to S702 illustrated in the embodiments of the present application, data enhancement processing is performed on the natural language query statement to obtain the target natural language query statement. Diversity transformation is carried out while preserving the semantics to enhance the data diversity of the natural language query statement. Then, the target natural language query statement is input into the test object to detect the compatibility ability of the test object with different language styles and optimize the anti-interference ability of the test object.
[0110] In step S701 of some embodiments, data enhancement processing of the natural language query statement can be performed in ways such as sentence pattern variation (rewriting into different sentence patterns, such as declarative sentences, interrogative sentences, imperative sentences, etc.), replacement of subject synonyms (replacing the main noun with a synonym), replacement of predicate synonyms (replacing a verb or adjective with its synonym), spelling mistakes (introducing common spelling mistakes in words), grammar mistakes (introducing common grammar mistakes), word order adjustment (adjusting the word order while maintaining the semantics), adding redundant modifiers (adding modifiers that do not affect the core semantics), omitting keywords (deleting some unnecessary keywords), pronoun substitution (using a pronoun to replace a specific noun phrase), style transformation (rewriting the question into an expert style or an ordinary user style), abbreviation substitution (replacing the main noun with the corresponding abbreviated noun), dialect transformation (converting the question into an informal dialect expression), etc., not limited to this. For example, if the natural language query statement is "How many users have an order amount exceeding 1000 yuan?", after applying the declarative sentence pattern variation, the target natural language query statement is "I want to know the number of users whose order amount exceeds 1000 yuan"; after applying the replacement of subject synonyms, the target natural language query statement is "How many customers have an order amount exceeding 1000 yuan?".
[0111] Please refer to Figure 4 , in some embodiments, step S402 may include but is not limited to steps S801 to S803:
[0112] Step S801: Perform data enhancement processing on the natural language query statement to obtain the target natural language query statement;
[0113] Step S802: Perform data enhancement processing on the database model sub-graph to obtain the target database model sub-graph;
[0114] Step S803: Input the target natural language query statement and the target database model sub-graph into the test object to obtain the target structured query statement.
[0115] For example, the natural language query statement is: "What is the average math score of the students?", and after data augmentation processing, the target natural language query statement "I want to query the average math score of the students." is obtained.
[0116] The structured query statement Y is: SELECT AVG(math_score) FROM Students JOIN Score ON Students.id = Score.student_id.
[0117] The database model sub-graph DB is: Table "Students": [id, name, age], Table "Score": [id, student_id, math_score, physics_score]; after data augmentation processing, the target database model sub-graph DB' is obtained: Table "Students": [id, name, age, sex], Table "Exam": [id, student_id, mathematics, physics_score].
[0118] Input the target natural language query statement and the target database model sub-graph into the test object to obtain the target structured query statement Y'.
[0119] Detect whether the execution results of the structured query statement Y and the target structured query statement Y' are consistent, that is, whether Execute(Y, DB) is equal to Execute(Y', DB'), to obtain the reliability data of the target structured query statement.
[0120] In some embodiments, after step S802, the database query statement generation method further includes: performing validity verification on the target natural language query statement and / or the target database model sub-graph, and filtering out the target natural language query statement and the target database model sub-graph that do not pass the validity verification. Specifically, the target natural language query statement and / or the target database model sub-graph can be input into a large language model for validity verification through the large language model.
[0121] In step S403 of some embodiments, the reliability verification of the target structured query statement includes but is not limited to: obtaining the first execution result of the target structured query statement; obtaining the second execution result of the structured query statement sample; verifying whether the first execution result and the second execution result are consistent. If the two are consistent, output that the target structured query statement passes the reliability verification; otherwise, output that the target structured query statement does not pass the reliability verification.
[0122] In some embodiments, after step S403, the database query statement generation method further includes: calculating the accuracy evaluation EX of the test object.
[0123]
[0124] wherein, V n is the execution result of the nth structured query statement sample; is the execution result of the nth target structured query statement; N is the total number of structured query statement samples.
[0125] Please refer to Figure 5 , the embodiments of the present application further provide a database query statement generation device, which can implement the above database query statement generation method. The device includes:
[0126] A data acquisition module, configured to acquire a database model diagram of a target database. The target database includes data tables, and the data tables include data columns. The database model diagram includes an initial node representing the data table, and the connection edge of the initial node represents the mapping relationship between the data columns;
[0127] A graph sampling module, configured to perform graph sampling on the initial nodes in the database model diagram to obtain a sampling result, and the sampling result includes candidate nodes;
[0128] A target node generation module, configured to select target data columns from the data columns of the candidate nodes and construct target nodes based on the target data columns;
[0129] A connection edge detection module, configured to detect the validity of the connection edges of the target nodes to obtain validity data;
[0130] A target node adjustment module, configured to adjust the data columns in the target nodes based on the validity data;
[0131] A database model sub-graph construction module, configured to construct a database model sub-graph based on the target nodes;
[0132] A test data generation module, configured to generate test data through a large language model based on the database model sub-graph. The test data includes structured query statement samples and natural language query statement samples.
[0133] The specific implementation manner of this database query statement generation device is basically the same as the specific embodiments of the above database query statement generation method, and will not be elaborated here.
[0134] The embodiments of the present application further provide an electronic device. The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the above database query statement generation method. The electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.
[0135] Please refer to Figure 6 , Figure 6 which illustrates the hardware structure of an electronic device according to another embodiment. The electronic device includes:
[0136] A processor 901, which can be implemented in ways such as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present application;
[0137] A memory 902, which can be implemented in forms such as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 902 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 902 and are called by the processor 901 to execute the method for generating database query statements according to the embodiments of the present application;
[0138] An input / output interface 903, which is used to implement information input and output;
[0139] A communication interface 904, which is used to implement communication interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or through wireless means (such as mobile network, WIFI, Bluetooth, etc.);
[0140] A bus 905, which transmits information between various components of the device (such as the processor 901, the memory 902, the input / output interface 903, and the communication interface 904);
[0141] Among them, the processor 901, the memory 902, the input / output interface 903, and the communication interface 904 achieve communication connections with each other inside the device through the bus 905.
[0142] The embodiments of the present application also provide a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above method for generating database query statements is implemented.
[0143] The memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include memories remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0144] The database query statement generation method, database query statement generation device, electronic device, and storage medium provided by the embodiments of the present application perform graph sampling on a database model diagram to obtain candidate nodes and select target data columns from the candidate nodes, and construct target nodes to reduce the database structure information to be processed subsequently; at the same time, considering that the database model diagram is different from a conventional graph structure, it is necessary to ensure the mapping relationship between data columns in the nodes, detect the validity of the connection edges of the target nodes, and adjust the target nodes based on this to improve the reliability of the database model sub-diagram, thereby improving the semantic integrity and rationality of the subsequently generated test data; finally, based on the database model sub-diagram, a large number of test data are generated.
[0145] The embodiments described in the embodiments of the present application are for more clearly explaining the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art can know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0146] Those skilled in the art can understand that the technical solutions shown in the figure do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown in the figure, or combine certain steps, or different steps.
[0147] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0148] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices can be implemented as software, firmware, hardware, and appropriate combinations thereof.
[0149] In the description of the present application and the above-mentioned drawings, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0150] It should be understood that in the present application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expression means any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0151] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the above-mentioned unit division is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.
[0152] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0153] In addition, in each embodiment of the present application, each functional unit can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0154] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. The aforementioned storage medium includes: various media that can store programs, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0155] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings. This does not limit the scope of the rights of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of the rights of the embodiments of the present application.
Claims
1. A method for generating a database query statement, characterized in that: The method comprises: Acquire a database model diagram of a target database, wherein the target database includes a data table, the data table includes data columns, the database model diagram includes an initial node representing the data table, and the connection edges of the initial node represent mapping relationships between the data columns; Performing graph sampling on the initial nodes in the database model graph to obtain sampling results, wherein the sampling results include candidate nodes; Selecting a target data column from the data column of the candidate node, and constructing a target node based on the target data column; Detecting the validity of the connection edge of the target node to obtain validity data; Based on the validity data, adjusting the data column in the target node; Based on the target node, construct a database model subgraph; Based on the database model subgraph, test data is generated through a large language model, and the test data includes structured query statement samples and natural language query statement samples.
2. The method according to claim 1, characterized in that The detecting the validity of the connection edge of the target node to obtain validity data includes: Based on the connection edge of the target node, the mapping object of the data column in the target node is detected in the sampling result to obtain validity data, where the validity data represents the existence or non-existence of the mapping object.
3. The method according to claim 1, characterized in that Also includes: Acquire a test object, where the test object is used to convert a natural language query statement into a structured query statement; Inputting the natural language query statement sample and the database model subgraph into the test object to obtain a target structured query statement; Based on the structured query statement sample, reliability verification is performed on the target structured query statement to obtain reliability data of the target structured query statement.
4. The method according to claim 3, characterized in that The step of inputting the natural language query statement sample and the database model subgraph into the test object to obtain a target structured query statement includes: Performing data enhancement processing on the database model subgraph to obtain a target database model subgraph; The natural language query statement sample and the target database model subgraph are input into the test object to obtain a target structured query statement.
5. The method according to claim 4, characterized in that The performing data enhancement processing on the database model subgraph to obtain a target database model subgraph includes: The connection edges in the database model subgraph are modified to change the mapping relationship between data columns to obtain a target database model subgraph.
6. The method according to claim 3, characterized in that The step of inputting the natural language query statement sample and the database model subgraph into the test object to obtain a target structured query statement includes: Performing data enhancement processing on the natural language query statement to obtain a target natural language query statement; The target natural language query statement and the database model subgraph are input into the test object to obtain a target structured query statement.
7. The method according to any one of claims 1 to 6, characterized in that: The step of obtaining a database model diagram of a target database includes: Acquire database metadata of the target database; Based on the database metadata, a database model diagram of the target database is constructed.
8. A database query statement generating device, characterized in that: The device comprises: A data acquisition module is used to acquire a database model diagram of a target database, wherein the target database includes a data table, the data table includes data columns, the database model diagram includes an initial node representing the data table, and the connection edge of the initial node represents a mapping relationship between the data columns; A graph sampling module, used to perform graph sampling on the initial nodes in the database model graph to obtain sampling results, wherein the sampling results include candidate nodes; A target node generation module, used for selecting a target data column from the data column of the candidate node, and constructing a target node based on the target data column; A connection edge detection module, used to detect the validity of the connection edge of the target node and obtain validity data; A target node adjustment module, used for adjusting the data column in the target node based on the validity data; A database model subgraph construction module, used to construct a database model subgraph based on the target node; A test data generation module is used to generate test data based on the database model subgraph through a large language model, wherein the test data includes structured query statement samples and natural language query statement samples.
9. An electronic device, characterized in that: The electronic device comprises a memory and a processor, the memory stores a computer program, and the processor implements the database query statement generating method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for generating a database query statement according to any one of claims 1 to 7 is implemented.