Method and device for converting natural language statement into SQL (Structured Query Language) and storage medium
By using the semantic and structural similarity calculation methods of knowledge graphs in data query scenarios, the accuracy problem of traditional natural language conversion SQL methods in complex scenarios is solved, and higher conversion accuracy is achieved.
Patent Information
- Application Number
- CN202510787118.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-13
AI Technical Summary
传统的自然语言转换SQL方法依赖人工规则模板,面对复杂表结构和查询需求时通用性差,准确性不足。
By performing semantic analysis of the data query text, the target entity and query fields are extracted, and the comprehensive similarity is calculated using the semantic similarity and structural similarity in the knowledge graph, the target node is determined, and the data query text is finally converted into a structured query statement.
Improve the accuracy of converting natural language query text into SQL query statements, avoiding the problem of similar semantics but no business relationships, and improving the accuracy of the conversion.
Smart Images

Figure CN120296139A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and in particular, to a method, apparatus, and storage medium for converting natural language statements into SQL. Background Art
[0002] In data query scenarios, users usually express their query intentions in natural language, and the query system needs to convert them into an executable Structured Query Language (SQL). However, traditional methods for converting natural language into SQL rely on manual rule templates or template-based rule engines, which have poor generality when facing complex table structures and query requirements and are insufficient in terms of accuracy. Summary of the Invention
[0003] In view of this, this application provides a method, apparatus, and storage medium for converting natural language statements into SQL, which can improve the accuracy of converting natural language query text into SQL query statements. This application is implemented in the following aspects: In a first aspect, an embodiment of this application provides a method for converting natural language statements into SQL, including: performing semantic parsing on a data query text to extract a target entity and at least one query field, where the target entity is used to indicate the query intention of the data query text; determining a comprehensive similarity between each query field in the at least one query field and each first candidate node in a pre-stored knowledge graph according to the semantic similarity and structural similarity between each query field and each first candidate node in the pre-stored knowledge graph, where the knowledge graph is constructed based on multiple data tables, and the structural similarity is used to indicate the degree of structural association between each query field and each first candidate node in the knowledge graph; determining a target node corresponding to each query field according to the comprehensive similarity between each query field and each first candidate node; and converting the data query text into a structured query statement according to at least one target node, the connection relationship between the target entity and the at least one target node, target business rules, and the data query text.
[0004] In a possible implementation manner, the method further includes: traversing the knowledge graph to determine the path length from the target entity to each first candidate node; and calculating the structural similarity between each query field and each first candidate node according to the path length from the target entity to each first candidate node and a preset inverse function.
[0005] In a possible implementation, determining the comprehensive similarity between each query field and each of the first candidate nodes according to the semantic similarity and structural similarity between each query field in the at least one query field and each of the first candidate nodes includes: for each of the first candidate nodes among the first candidate nodes, calculating a first similarity according to the product result of the semantic similarity between each query field and each of the first candidate nodes and a first preset weight; calculating a second similarity according to the structural similarity between each query field and each of the first candidate nodes and a second preset weight; determining the comprehensive similarity between each query field and each of the first candidate nodes according to the sum of the first similarity and the second similarity, and obtaining the comprehensive similarity between each query field and each of the first candidate nodes in total.
[0006] In a possible implementation, before determining the comprehensive similarity between each query field and each of the first candidate nodes according to the semantic similarity and structural similarity between each query field in the at least one query field and each of the first candidate nodes, the method further includes: calculating the semantic similarity between each query field and each node in the knowledge graph; determining each of the first candidate nodes according to the semantic similarity between each query field and each of the nodes and a first similarity threshold.
[0007] In a possible implementation, determining the target node corresponding to each query field according to the comprehensive similarity between each query field and each of the first candidate nodes includes: determining at least one second candidate node according to the comprehensive similarity between each query field and each of the first candidate nodes and a second similarity threshold; inputting the at least one second candidate node, the data query text, and the context information of the at least one second candidate node in the knowledge graph into a pre-trained large language model to obtain the target node.
[0008] In a possible implementation, before determining the comprehensive similarity between each query field and each of the first candidate nodes according to the semantic similarity and structural similarity between each query field in the at least one query field and each of the first candidate nodes, the method further includes: extracting metadata information of each of the multiple data tables, where the metadata information includes table description information and column description information; constructing multiple nodes and attribute information of each node in the multiple nodes according to the table description information and the column description information, where one node is used to indicate a table entity or a column entity; extracting table-column relationships and column-column relationships between the multiple nodes according to the table description information and the column description information; constructing relationship edges between the multiple nodes according to the table-column relationships and the column-column relationships to obtain the knowledge graph.
[0009] In a possible implementation, the converting the data query text into a structured query statement according to at least one target node, the connection relationship between the target entity and the at least one target node, the target business rule, and the data query text includes: inputting the at least one target node, the connection relationship between the target entity and the at least one target node, the target business rule, and the data query text into a pre-trained statement conversion model to obtain the structured query statement, where the pre-trained statement conversion model is trained according to a sample query statement data set.
[0010] In a second aspect, an embodiment of the present application provides a device for converting a natural language statement into SQL, including: a parsing module, configured to perform semantic parsing on a data query text to extract a target entity and at least one query field, where the target entity is used to indicate the query intention of the data query text; a determining module, configured to determine the comprehensive similarity between each query field and each of the first candidate nodes in a pre-stored knowledge graph according to the semantic similarity and structural similarity between each query field in the at least one query field and each of the first candidate nodes, the knowledge graph is constructed according to multiple data tables, and the structural similarity is used to indicate the structural association degree between each query field and each of the first candidate nodes in the knowledge graph; the determining module is further configured to determine a target node corresponding to each query field according to the comprehensive similarity between each query field and each of the first candidate nodes; a conversion module, further configured to convert the data query text into a structured query statement according to at least one target node, the connection relationship between the target entity and the at least one target node, the target business rule, and the data query text.
[0011] In a third aspect, an embodiment of the present application provides an apparatus for converting a natural language statement into SQL, including: a memory storing executable program code; a processor coupled to the memory; the processor calls the executable program code stored in the memory and executes the method as described in the first aspect and any possible implementation manner.
[0012] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method as described in the first aspect and any possible implementation manner is implemented.
[0013] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method as described in the first aspect and any possible implementation manner is implemented.
[0014] By implementing the method provided in the embodiments of the present application, after determining the target entity and at least one query field, based on the semantic similarity between each query field and each first candidate node in the knowledge graph, nodes semantically similar to each query field in the knowledge graph can be determined. Based on the structural similarity between each query field and each first candidate node in the knowledge graph, the structural association degree between each query field and each first candidate node in the knowledge graph can be determined, and this structural association program can reflect the business association degree between each query field and each first candidate node in the knowledge graph. Therefore, by combining the semantic similarity and the structural similarity to determine the comprehensive similarity, and based on the target node determined by the comprehensive similarity, the problem that the query field and the node are semantically similar but have no business association can be effectively avoided, and further the problem that the converted SQL statement is inaccurate due to the incorrect determination of the target node can be avoided. It can be seen that the method for converting a natural language statement into SQL provided by the embodiments of the present application can improve the accuracy of converting a natural language query text into an SQL query statement. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required in the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0016] Figure 1 FIG. is a schematic diagram of an application scenario of a method for converting a natural language statement into SQL provided by an embodiment of the present application; Figure 2 FIG. is a schematic flowchart of a method for converting a natural language statement into SQL provided by an embodiment of the present application; Figure 3Schematic diagram of a knowledge graph provided by an embodiment of the present application Figure 1 ; Figure 4 Schematic diagram of a knowledge graph provided by an embodiment of the present application Figure 2 ; Figure 5 Schematic structure diagram of a device for converting natural language sentences into SQL provided by an embodiment of the present application Figure 1 ; Figure 6 Schematic structure diagram of a device for converting natural language sentences into SQL provided by an embodiment of the present application Figure 2 。 Detailed implementation manners
[0017] The technical solutions in the present application will be described below with reference to the accompanying drawings.
[0018] For the convenience of clearly describing the technical solutions of the embodiments of the present application, in the embodiments of the present application, terms such as "first" and "second" are used to distinguish identical or similar items with basically the same functions and effects. For example, the first instruction and the second instruction are used to distinguish different user instructions, and their sequence is not limited. Those skilled in the art can understand that terms such as "first" and "second" do not limit the quantity and execution order, and "first", "second", etc. do not necessarily mean different.
[0019] It should be noted that in the present application, words such as "exemplarily" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplarily" or "for example" in the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, using words such as "exemplarily" or "for example" is intended to present relevant concepts in a specific manner.
[0020] In addition, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one (item)" or its similar expression refers to any combination of these items, including any combination of single item (item) or plural items (items). For example, at least one (item) of a, b, and c can represent: a, or b, or c, or a and b, or a and c, or b and c, or a, b, and c, where a, b, and c can be single or multiple.
[0021] In addition, the terms "include" and "have" and any variations thereof in the embodiments of the present application and the accompanying drawings are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices.
[0022] In a data query scenario, a user usually expresses a query intention in natural language, and the query system needs to convert it into an executable Structured Query Language (SQL). However, traditional methods for converting natural language to SQL rely on manual experience, have poor generality when facing complex table structures and query requirements, and the accuracy of the converted SQL is relatively poor.
[0023] Based on this, an embodiment of the present application provides a method for converting a natural language statement into SQL. The method includes: performing semantic parsing on a data query text to extract a target entity and at least one query field, where the target entity is used to indicate the query intention of the data query text; determining a comprehensive similarity between each query field in the at least one query field and each first candidate node in a pre-stored knowledge graph according to the semantic similarity and structural similarity between each query field and each first candidate node in the pre-stored knowledge graph. The knowledge graph is constructed based on multiple data tables, and the structural similarity is used to indicate the degree of structural association between each query field and each first candidate node in the knowledge graph; determining a target node corresponding to each query field according to the comprehensive similarity between each query field and each first candidate node; and converting the data query text into a structured query statement according to the at least one target node, the connection relationship between the target entity and the at least one target node, the target business rule, and the data query text.
[0024] By combining the semantic similarity and structural similarity between the query field and each first candidate node in the knowledge graph to determine the comprehensive similarity, since the structural similarity can reflect the degree of business association between the query field and each first candidate node, that is to say, this comprehensive similarity takes into account both the degree of semantic association and the degree of business association. Therefore, the target node determined based on the comprehensive similarity can effectively avoid the problem that the query field and the node are semantically similar but have no business association, and further avoid the problem that the converted SQL statement is inaccurate due to the incorrect determination of the target node, that is, improve the accuracy of converting the natural language query text into an SQL query statement.
[0025] The following is an example illustration of the application scenario of a method for converting a natural language statement into SQL provided by an embodiment of the present application.
[0026] Please refer to Figure 1 , which is a schematic diagram of an application scenario of a method for converting natural language sentences into SQL provided by an embodiment of the present application. As Figure 1 shown, the schematic diagram includes a data query device 110 and an electronic device 120.
[0027] Among them, the data query device 110 may refer to a device or apparatus with data query function. When the data query device 110 is a device with data query function, the data query device 110 may be, for example, a terminal device. When the data query device 110 is a device with data query function, the data query device 110 may be a functional module in other devices. For example, the data query device 110 is a functional module in the electronic device 120. Exemplarily, the terminal devices involved in the embodiments of the present application may include general handheld electronic terminals, such as mobile phones, smart phones, portable terminals, terminals, personal digital assistants (Personal Digital Assistant, PDA), portable multimedia players (Personal MediaPlayer, PMP) devices, laptop computers, notebooks (Note Pad), wireless broadband (Wireless Broadband, Wibro) terminals, tablet computers (personal computer, PC), smart PCs, point of sales (Point of Sales, POS), and in-vehicle computers. The terminal device may also include wearable devices. Wearable devices are portable electronic devices that can be directly worn on the user's body or integrated into the user's clothes or accessories. Wearable devices are not only a hardware device, but can also achieve powerful intelligent functions through software support, data interaction, and cloud interaction, such as: computing function, positioning function, alarm function, and can also connect to mobile phones and various terminals. Wearable devices may include, but are not limited to, watch-type devices supported by the wrist (such as watches, wrists, etc.), shoes-type devices supported by the feet (such as shoes, socks, or other products worn on the legs), Glass-type devices supported by the head (such as glasses, helmets, headbands, etc.), and smart clothing, schoolbags, crutches, accessories, and other various non-mainstream product forms.
[0028] The electronic device 120 may refer to a terminal device or a server with data processing capabilities. The introduction of the terminal device can be referred to the corresponding content described above, and will not be elaborated here. Multiple data tables may also be pre-stored in the electronic device 120, or the electronic device 120 may access multiple data tables by accessing an external storage device.
[0029] Exemplarily, the data query device 110 can obtain the data query text input by the user and send the data query text to the electronic device 120. The electronic device 120 receives the data query text and processes the data query text to convert the data query text into an SQL query statement. Among them, the specific manner in which the electronic device 120 processes the data query text to convert the data query text into an SQL query statement will be described in detail below.
[0030] Please refer to Figure 2 , which is a schematic flowchart of a method for converting a natural language statement into SQL provided by an embodiment of the present application. Among them, the following is described in terms of the electronic device executing Figure 2 the steps shown, and the electronic device is, for example, Figure 1 the electronic device 120 shown. Figure 2 The method shown can be applied to Figure 1 the application scenario shown.
[0031] S201, perform semantic parsing on the data query text, and extract a target entity and at least one query field, where the target entity is used to indicate the query intent of the data query text.
[0032] Among them, the query intent can indicate the target data or target data table that the data query text expects to query. The target data can be any data in multiple pre-stored data tables. The target data table can be any data table among multiple data tables.
[0033] Specifically, the electronic device can perform semantic parsing on the data query text, analyze the query intent of the data query text, and determine the target entity and at least one query field of the data query text based on the query intent. In the embodiment of the present application, the target entity is described as the target data table involved in the data query text.
[0034] Alternatively, the electronic device can also extract the target entity and at least one query field from the data query text through the named entity recognition technology (NER). Exemplarily, the electronic device can perform named recognition on the data query text, identify the associated entity in the data query text, and use the associated entity as the target entity. Among them, the associated entity refers to the data table involved in the data query text, such as the "order amount table".
[0035] In this implementation manner, the electronic device can also input the data query text into the pre-trained named entity recognition model based on the pre-trained named entity recognition model to obtain the target entity and at least one query field. Among them, the pre-trained entity recognition model is trained based on a large number of labeled sample data query texts.
[0036] For example, the data query text is "Query the total order amount in City A in 2023". By semantic parsing, it is analyzed that the data query text needs to correspond to the data in the order amount table. Then, the target entity corresponds to "order amount", and at least one query field corresponds to "total order amount".
[0037] It should be understood that at least one query field extracted in step S201 is the initial query field.
[0038] S202. Determine the comprehensive similarity between each query field and each first candidate node in the pre-stored knowledge graph according to the semantic similarity and structural similarity between each query field in at least one query field and each first candidate node in the knowledge graph; wherein, the knowledge graph is constructed based on multiple data tables, and the structural similarity is used to indicate the degree of structural association between each query field and each first candidate node in the knowledge graph.
[0039] Among them, the semantic similarity can indicate the degree of semantic association between the query field and the node. Since the knowledge graph is constructed based on multiple data tables and includes the business associations between multiple data tables, the structural similarity can not only reflect the degree of structural association between the query field and the node, but also indirectly reflect the degree of business association between the query field and the node.
[0040] Specifically, the electronic device can calculate the comprehensive similarity between each query field and each first candidate node according to the sum of the product result of the semantic similarity between each query field and each first candidate node and the first preset weight, and the product result of the structural similarity between each query field and each first candidate node and the second preset weight.
[0041] Specifically, for each first candidate node among the first candidate nodes, the electronic device can calculate the first similarity according to the product result of the semantic similarity between each query field and each first candidate node and the first preset weight, and calculate the second similarity according to the structural similarity between each query field and each first candidate node and the second preset weight. Finally, according to the sum of the first similarity and the second similarity, determine the comprehensive similarity between each query field and each first candidate node, and obtain the comprehensive similarity between each query field and each first candidate node. Exemplarily, a calculation formula for determining the comprehensive similarity between each query field and a first candidate node is as follows: Comprehensive similarity = First similarity + Second similarity = First preset weight * Semantic similarity + Second preset weight * Structural similarity Among them, both the first preset weight and the second preset weight are pre-configured in the electronic device. The specific values of the first preset weight and the second preset weight can be set according to actual needs. For example, they can be obtained based on experience or experiments. The embodiments of the present application do not limit this. The first preset weight and the second preset weight are adjustable weights, and the specific adjustment method can be set according to actual business needs. The embodiments of the present application do not limit this.
[0042] The calculation methods of semantic similarity and structural similarity will be described separately below.
[0043] 1. Semantic similarity.
[0044] The electronic device can calculate the cosine similarity between each query field and each first candidate node to obtain the semantic similarity between each query field and each first candidate node. Specifically, the electronic device can convert each query field and each first candidate node into vector form, and then calculate the cosine similarity between the vectorized each query field and each first candidate node, and finally obtain the semantic similarity between each query field and each first candidate node.
[0045] Alternatively, the electronic device can also calculate the semantic similarity between each query field and each first candidate node according to other methods. For example, the semantic similarity between each query field and each first candidate node is output by a pre-trained model. Among them, the pre-trained model can be trained based on a large amount of sample data, and the sample data includes sample query fields and sample nodes.
[0046] It should be understood that the semantic similarity between each query field and each first candidate node includes the semantic similarity between each query field and each first candidate node among each first candidate node. For example, taking each query field as query field 1 and each first candidate node including node 1, node 2, and node 3, the semantic similarity between query field 1 and each first candidate node should include the semantic similarity between query field 1 and node 1, the semantic similarity between query field 1 and node 2, and the semantic similarity between query field 1 and node 3.
[0047] 2. Structural similarity.
[0048] The electronic device can calculate the structural similarity between each query field and each first candidate node according to the path length between each target entity and each first candidate node. The following combines Figure 4 The flowchart of a method for determining structural similarity shown, and the specific method for determining the structural similarity between each target entity and each first candidate node will be specifically described.
[0049] S401. Traverse the knowledge graph to determine the path lengths from the target entity to each first candidate node in the knowledge graph.
[0050] Among them, the path length can refer to the number of hops from the entity to the node, or the number of edges (relationships) passed from the entity to the node. Exemplarily, please refer to Figure 3 , which is a schematic diagram of a knowledge graph provided by an embodiment of the present application. As Figure 3 shown, the path length from target entity 1 to node 1 is 2, the path length from target entity 1 to node 2 is 3, and the path length from target entity 1 to node 3 is 1.
[0051] Particularly, if there is no connection relationship between a target entity and a node, for example, Figure 3 as shown in node 4, there is no connection relationship with target entity 1. In this case, the node can be excluded, that is, the comprehensive similarity calculation process between the target entity and the node is ended. Or, the path length between the node and the target entity can be set to a maximum value, such as 999, 9999, etc., where the maximum value can be pre-configured in the electronic device.
[0052] In a possible implementation manner, when determining the path length from the target entity to each first candidate node, the path length can also be calculated according to the type of the edge passed from the target entity to each first candidate node and the weight corresponding to the edge type. The edge type can include semantic relationships (such as semantic equivalence relationships), foreign key association relationships, etc. For example, if the foreign key edge weight is set to 1 and the semantic edge weight is set to 2, then the path length = 1×the number of foreign key edges + 2×the number of semantic edges.
[0053] It should be noted that there may be multiple paths from the target entity to each first candidate node. In the embodiment of the present application, the path length of the shortest path from the target entity to each first candidate node is used as the path length for calculating the structural similarity between each query field and each first candidate node. In practical applications, corresponding weights can also be set for the paths according to business rules to select the path length of the target path as the path length for calculating the structural similarity between each query field and each first candidate node.
[0054] S402. Calculate the structural similarity between each query field and each first candidate node according to the path lengths from the target entity to each first candidate node in the knowledge graph and a preset inverse function.
[0055] Among them, the preset inverse function can refer to the inverse of the path length, or the reciprocal of the path length. Exemplarily, a calculation formula for calculating the structural similarity between each query field and a first candidate node is as follows: Structural similarity = 1 / path length Since the range of cosine similarity is [-1, 1], the range of the structural similarity can be unified to [0, 1] through a preset inverse function. The structural similarity is indicated by the reciprocal of the path length. The shorter the path length, the greater the structural similarity, which can reflect that the structural association between the target entity and the node is closer, and indirectly reflects that the business association is closer.
[0056] It should be understood that the path lengths from the target entity to each first candidate node include the path lengths from the target entity to each of the first candidate nodes in the first candidate nodes. For example, taking the target entity as target entity 1 and the first candidate nodes including node 1, node 2, and node 3 as an example, the path lengths from target entity 1 to each first candidate node include the path length from target entity 1 to node 1, the path length from target entity 1 to node 2, and the path length from target entity 1 to node 3.
[0057] In a possible implementation manner, when the target entity is the same as a first candidate node, the path length between the target entity and the first candidate node is 0. At this time, the structural similarity between the target entity and the first candidate node cannot be calculated based on the reciprocal of the path length. Therefore, to avoid this situation, the preset inverse function is set to the reciprocal of (path length + 1). Exemplarily, a calculation formula for calculating the structural similarity between each target entity and a first candidate node is as follows: Structural similarity = 1 / (path length + 1) Exemplarily, if the path length from target entity 1 to node 2 is 3 and the path length from target entity 1 to node 3 is 1, then the structural similarity from target entity 1 to node 2 is 0.25, and the structural similarity from the target entity to node 3 is 0.5. It can be seen that the structural similarity between target entity 1 and node 3 is higher, and the structural association and business association between target entity 1 and node 3 in the knowledge graph are closer.
[0058] Based on the above content, it can be known that the embodiments of the present application provide two calculation formulas for comprehensive similarity, which are as follows: Formula 1: Comprehensive similarity = first preset weight * semantic similarity + second preset weight * (1 / path length) Formula 2: Comprehensive similarity = first preset weight * semantic similarity + second preset weight * [1 / (path length + 1)] Among them, the content of the first preset weight and the second preset weight can be correspondingly referred to the content described above, and will not be elaborated here.
[0059] In a possible implementation, each first candidate node is a node in the knowledge graph that is semantically similar to each query field. In this case, the electronic device can first screen out the first candidate nodes with similar semantics from the semantic similarities between each query field and each node in the knowledge graph. In this way, by excluding some nodes with significantly different semantics and retaining the first candidate nodes with similar semantics, it is beneficial to reduce the computational pressure of the electronic device and improve the computational efficiency. Specifically, the electronic device can calculate the semantic similarity between each query field and each node in the knowledge graph. Then, based on the semantic similarity between each query field and each node and the first similarity threshold, each first candidate node is determined. For example, the electronic device can use the nodes with a semantic similarity greater than or equal to the first similarity threshold as the first candidate nodes.
[0060] Among them, the first similarity threshold is used to screen out the nodes in each that are semantically closer to each target entity. The first similarity threshold can be set according to actual needs, and the embodiments of the present application do not limit this. Among them, the specific method for the electronic device to calculate the semantic similarity can be correspondingly referred to the content described above, and will not be elaborated here.
[0061] Exemplarily, taking the knowledge graph including nodes 1, 2, 3, 4, 5, the query field being query field 1, and the first similarity threshold being 0.5 as an example, the electronic device calculates that the semantic similarity between query field 1 and node 1 is 0.67, the semantic similarity between query field 1 and node 2 is 0.75, the semantic similarity between query field 1 and node 3 is 0.75, the semantic similarity between query field 1 and node 4 is 0.37, and the semantic similarity threshold between query field 1 and node 5 is 0.25. It can be seen that the semantic similarities between query field 1 and nodes 4 and 5 are less than the first similarity threshold. Therefore, the electronic device can use nodes 1, 2, and 3 as each first candidate node.
[0062] In a possible implementation, before calculating the comprehensive similarity, the electronic device can construct a knowledge graph for multiple data tables. Specifically, the electronic device can extract the metadata information of each data table in the multiple data tables, where the metadata information includes table description information and column description information. The table description information includes information such as table name, table structure information, association relationship with other tables (such as foreign key constraints or business logic associations), index type, primary key definition, etc. The column description information includes information such as column name (or field name), data type, constraint conditions, field attributes, business meaning (specific description of the data stored in the column), etc. Optionally, the table description information and column description information can also include other information, and the embodiments of the present application will not list them one by one here.
[0063] Further, based on the table description information and column description information, multiple nodes and the attribute information of each node in the multiple nodes are constructed. Among them, one node is used to indicate a table entity or a column entity. Generally speaking, the electronic device can abstract each table and each column in multiple databases into a node, store the table description information corresponding to each table in the attribute information of the corresponding table entity, and store the column description information corresponding to each column in the attribute information of the corresponding column entity.
[0064] Then, based on the table description information and column description information, the table-column relationships and column-column relationships among multiple nodes are extracted. Furthermore, based on the table-column relationships and column-column relationships, relationship edges among multiple nodes can be constructed to obtain a knowledge graph. Optionally, the edge types of the relationship edges include column belonging to table, foreign key association, aggregation logic, semantic equivalence relationship (such as "customer" and "user"), etc. When the electronic device constructs the relationship edges between nodes, it can also establish its display connection method through semantic labels. Among them, the relationships between columns include cross-table or cross-column data associations realized through foreign keys, calculation logics, business rules, etc., and the relationships between tables and columns include defining table structures and data rules through primary keys, data types, constraints, indexes, etc.
[0065] Finally, a database knowledge graph (DB-KG) is obtained, which explicitly expresses the complex associations between multiple tables and provides a structured knowledge basis for subsequent semantic reasoning and SQL generation.
[0066] After the knowledge graph is constructed, the knowledge graph can be persisted using a graph database (such as Neo4j) to ensure efficient relationship traversal and path query.
[0067] In a possible implementation manner, after the knowledge graph is constructed, the database structure changes (such as adding tables, modifying foreign keys) can be monitored in real time or at preset intervals, and synchronized to the knowledge graph in real time to ensure its consistency with the actual database.
[0068] S203. Determine the target node corresponding to each query field according to the comprehensive similarity between each query field and each first candidate node.
[0069] Among them, the target node may refer to the standard query field corresponding to the query field. As mentioned above, the query fields extracted in step S201 are initial query fields. Since there may be some differences in the descriptions of the fields extracted from the data query text and the fields in the data table, in order to improve the accuracy of SQL conversion, the query fields extracted from the data query text need to be corrected to obtain their corresponding standard query fields.
[0070] Specifically, the electronic device may use the first candidate node corresponding to the highest comprehensive similarity among the comprehensive similarities between each query field and each first candidate node as the target node.
[0071] Alternatively, the electronic device may also determine at least one second candidate node according to the comprehensive similarity between each query field and each first candidate node and the second similarity threshold. For example, the electronic device may use the first candidate node with a comprehensive similarity greater than or equal to the second similarity threshold as the second candidate node. Then, input the at least one second candidate node, the data query text, and the context information of the at least one second candidate node in the knowledge graph into a pre-trained large language model to obtain the target node. Specifically, by analyzing the semantics of the data query text through the pre-trained large language model and combining the context information of the at least one second candidate node in the knowledge graph, analyze the connection relationship between the query field and each second candidate node, and then the target node that matches the semantics of the data query text and the query field can be determined from the at least one second candidate node.
[0072] Among them, the second similarity threshold is pre-configured in the electronic device, and the second similarity threshold can be set according to actual needs, which is not limited in the embodiments of the present application. The pre-trained large language model is trained according to a sample data set, and the sample data set includes data such as labeled sample query fields, candidate nodes corresponding to the sample query fields, and the context information of the candidate nodes in the knowledge graph. The context information of each second candidate node in the knowledge graph includes the adjacent nodes of each second candidate node in the knowledge graph, as well as the connection relationship with the adjacent nodes, the attributes of the adjacent nodes, etc. Among them, the connection relationship between the second candidate node and the adjacent node includes the relationship between columns and / or the relationship between tables and columns. Among them, the relationship between columns includes cross-table or cross-column data association realized through foreign keys, calculation logics, business rules, etc., and the relationship between tables and columns includes defining table structures and data rules through primary keys, data types, constraints, indexes, etc.
[0073] Exemplarily, taking each of the first candidate nodes as Node 1, Node 2, and Node 3, each query field as Query Field 1, and both the first preset weight and the second preset weight as 0.5, the electronic device calculates that the semantic similarity between Query Field 1 and Node 1 is 0.67, the semantic similarity between Query Field 1 and Node 2 is 0.75, and the semantic similarity between Query Field 1 and Node 3 is 0.75; the structural similarity between Query Field 1 and Node 1 is 0.33, the structural similarity between Query Field 1 and Node 2 is 0.5, and the structural similarity between Query Field 1 and Node 3 is 0.2. Combining the above comprehensive similarity calculation formula, the electronic device calculates that the comprehensive similarity between Query Field 1 and Node 1 is 0.5 * (0.67 + 0.33) = 0.5, the comprehensive similarity between Query Field 1 and Node 2 is 0.5 * (0.75 + 0.5) = 0.625, and the comprehensive similarity between Query Field 1 and Node 3 is 0.5 * (0.75 + 0.2) = 0.475. Based on this, the electronic device can determine Node 2 as the target node.
[0074] Alternatively, if the second similarity threshold is 0.5, the electronic device can take Node 1 and Node 2 as the second candidate nodes. Then, Node 1, Node 2, the data query text, and the context information of Node 1 and Node 2 in the knowledge graph are input into a pre-trained large language model, and the target node is output. For example, the target node is Node 2.
[0075] To better understand the comprehensive similarity provided in the embodiments of the present application, the following takes the data query text as "Query the GDP growth rate of City B in 2024", the target entity as "Year-on-year GDP growth rate", the query field as "GDP growth rate", the first candidate nodes as "Year-on-year GDP growth rate" and "Month-on-month GDP growth rate", and the first preset weight as 0.6 and the second preset weight as 0.4 as an example to illustrate the comprehensive similarity provided in the embodiments of the present application. It should be understood that usually, "GDP growth rate" is defaulted to "Year-on-year GDP growth rate". Therefore, when the initial target entity "GDP growth rate" extracted by the electronic device from the data query text does not match the nodes in the knowledge graph in description, it can be mapped to the specific fields (nodes) in the knowledge graph according to the semantics of the target entity in the data query text to correct the description of the target entity.
[0076] It can be seen that the semantics between the first candidate nodes "year-on-year GDP growth rate" and "quarter-on-quarter GDP growth rate" and "GDP growth rate" are both similar. In this example, the electronic device takes the "year-on-year GDP growth rate" as the central node (in this example, the "year-on-year GDP growth rate" is the table node when serving as the central node), and calculates the path lengths between this central node and the first candidate nodes "year-on-year GDP growth rate" (column node) and "quarter-on-quarter GDP growth rate" respectively. The electronic device can determine that the path length between the central node and the "year-on-year GDP growth rate" is 1, and the path length between the central node and the "quarter-on-quarter GDP growth rate" is 2. The electronic device calculates that the semantic similarity between the query field "GDP growth rate" and the "year-on-year GDP growth rate" is 0.75, and the semantic similarity between the query field "GDP growth rate" and the "quarter-on-quarter GDP growth rate" is 0.75. Further, combining the semantic similarity and the path length, the comprehensive similarity between the query field and the "year-on-year GDP growth rate" is calculated to be 0.6×0.75 + 0.4×0.5 = 0.45 + 0.2 = 0.65, and the comprehensive similarity between the query field and the "quarter-on-quarter GDP growth rate" is 0.6×0.75 + 0.4×0.33 = 0.39 + 0.132 = 0.582. Based on this, the electronic device can determine that the comprehensive similarity between the query field and the "year-on-year GDP growth rate" is higher. Therefore, the electronic device takes the "year-on-year GDP growth rate" as the target node.
[0077] It should be noted that the above example is a simple example provided to better understand the role of the comprehensive similarity provided by the embodiments of the present application. In actual applications, the complexity of the data query text may be higher, such as involving multiple query fields. In this case, the descriptions of the central entity and the query fields may be different, and it is determined according to the actual situation.
[0078] In a possible implementation manner, after determining the target node, the electronic device can also determine the foreign key association relationship corresponding to multiple data tables involved in the data query text based on the data query text, the target entity, and the knowledge graph. Exemplarily, taking the data query text including "query the total order amount in City A in 2023" as an example, the electronic device determines that the standard query field corresponding to the query field "total order amount" is "order amount". The electronic device determines, according to the structure of the knowledge graph and the data query text, that the order amount table includes fields such as region ID, year, and order amount (corresponding to nodes in the knowledge graph), and the region ID is associated with the region table as a foreign key. The region table includes fields such as region ID and region name. Based on this, the electronic device can determine that the data query text involves two data tables (the order amount table and the region table), and these two data tables are associated by the region ID foreign key.
[0079] Analyze the semantics of the data query text through a pre-trained large language model, and analyze the connection relationship between the query field and each second candidate node by combining the context information of at least one second candidate node in the knowledge graph. Furthermore, a target node that semantically matches the data query text can be queried from at least one second candidate node, which is beneficial to improving the accuracy of determining the target node.
[0080] S204. According to at least one target node, the connection relationship between the target entity and at least one target node, the target business rule, and the data query text, convert the data query text into a structured query statement.
[0081] Among them, each target node in at least one target node is determined based on the methods described in the previous steps S202 - S203. At least one target node indicates the standard query fields corresponding to the data query text. The connection relationship between the target entity and at least one target node includes the relationship between tables and columns, and the relationship between columns and columns. Among them, the relationship between columns and columns includes cross-table or cross-column data association realized through foreign keys, calculation logics, business rules, etc., and the relationship between tables and columns includes defining table structures and data rules through primary keys, data types, constraints, indexes, etc. The connection relationship between the target entity and at least one target node can be understood as a JOIN condition. The target business rule can be determined according to the user's data query requirements. The target business rule can be pre-configured in the electronic device and matched from the business rule library based on the semantics of the data query text. For example, when the data queried by the user has a different data unit from the data in the data table, the target business rule can be a unit conversion formula.
[0082] Specifically, the electronic device can input at least one target node, the connection relationship between the target entity and at least one target node, the target business rule, and the data query text into a pre-trained statement conversion model to obtain a structured query statement. Among them, the pre-trained statement conversion model is trained according to a sample query statement dataset. The statement conversion module can be a large language model, a deep learning model, a machine learning model, etc.
[0083] In a possible implementation manner, after obtaining the structured query statement, the electronic device can execute the structured query statement to query the target data that matches the data query text from multiple data tables. Among them, the specific manner in which the electronic device executes the structured query statement to query the target data is the same as the traditional query manner, and this application embodiment does not limit this.
[0084] It can be seen that in the embodiments of the present application, since the target entity indicates the query intention of the data query text, when calculating the structural similarity between the query field and each first candidate node, the target entity is used as the central node (or called the starting node), and the path length between the target entity and each first candidate node is calculated to obtain the structural similarity between the query field and each first candidate node. In this way, the shorter the path length between each first candidate node and the target entity, the closer the business association between each first candidate node and the target entity, and thus the closer the business association with the query field. By combining the semantic similarity and the structural similarity to calculate the comprehensive similarity, the comprehensive similarity can reflect both the semantic association relationship between the query field and the node and the structural association relationship between the query field and the node. Based on this, the target node is selected according to the comprehensive similarity, which can effectively avoid the situation that the query field and the node are semantically similar but have no business association, thereby improving the accuracy of determining the target node, and further improving the accuracy of the converted SQL statement.
[0085] The method for converting a natural language statement into SQL provided by the embodiments of the present application can be applied to various data query scenarios, especially in data query scenarios where multiple field semantics are similar, such as medical data query scenarios, financial data query scenarios, and so on.
[0086] Based on the same inventive concept, the embodiments of the present application provide a device for converting a natural language statement into SQL. This device is used to implement the method for converting a natural language statement into SQL as described above, for example, Figure 2 the method for converting a natural language statement into SQL as shown, and furthermore, this device can also implement the functions of the electronic device in the foregoing text.
[0087] Please refer to Figure 5 , which is a schematic structural diagram of a device for converting a natural language statement into SQL provided by the embodiments of the present application. As Figure 5 shown, Figure 5 the device for converting a natural language statement into SQL as shown includes a parsing module 501, a determination module 502, and a conversion module 503.
[0088] Exemplarily, a parsing module 501 is configured to perform semantic parsing on a data query text to extract a target entity and at least one query field, where the target entity is used to indicate the query intent of the data query text; a determination module 502 is configured to determine a comprehensive similarity between each query field and each first candidate node according to the semantic similarity and the structural similarity between each query field in the at least one query field and each first candidate node in a pre-stored knowledge graph, the knowledge graph is constructed based on a plurality of data tables, and the structural similarity is used to indicate the degree of structural association between each query field and each first candidate node in the knowledge graph; the determination module 502 is further configured to determine a target node corresponding to each query field according to the comprehensive similarity between each query field and each first candidate node; a conversion module 503 is further configured to convert the data query text into a structured query statement according to the at least one target node, the connection relationship between the target entity and the at least one target node, the target business rule, and the data query text.
[0089] In a possible implementation manner, the determination module 502 is further configured to: traverse the knowledge graph to determine the path length from the target entity to each first candidate node; and calculate the structural similarity between each query field and each first candidate node according to the path length from the target entity to each first candidate node and a preset inverse function.
[0090] In a possible implementation manner, the determination module 502 is specifically configured to: for each first candidate node among the first candidate nodes, calculate a first similarity according to the product result of the semantic similarity between each query field and each first candidate node and a first preset weight; calculate a second similarity according to the structural similarity between each query field and each first candidate node and a second preset weight; and determine the comprehensive similarity between each query field and each first candidate node according to the sum of the first similarity and the second similarity, so as to obtain the comprehensive similarity between each query field and each first candidate node.
[0091] In a possible implementation manner, the determination module 502 is further configured to: before determining the comprehensive similarity between each query field and each first candidate node according to the semantic similarity and the structural similarity between each query field in the at least one query field and each first candidate node, calculate the semantic similarity between each query field and each node in the knowledge graph; and determine each first candidate node according to the semantic similarity between each query field and each node and a first similarity threshold.
[0092] In a possible implementation manner, the determining module 502 is specifically configured to: determine at least one second candidate node according to the comprehensive similarity between each query field and each first candidate node and the second similarity threshold; input the at least one second candidate node, the data query text, and the context information of the at least one second candidate node in the knowledge graph into a pre-trained large language model to obtain a target node.
[0093] In a possible implementation manner, the determining module 502 is further configured to: before determining the comprehensive similarity between each query field in the at least one query field and each first candidate node according to the semantic similarity and the structural similarity between each query field and each first candidate node, extract the metadata information of each data table in multiple data tables, where the metadata information includes table description information and column description information; construct multiple nodes and the attribute information of each node in the multiple nodes according to the table description information and the column description information, where one node is used to indicate a table entity or a column entity; extract the table-column relationship and the column-column relationship between the multiple nodes according to the table description information and the column description information; construct relationship edges between the multiple nodes according to the table-column relationship and the column-column relationship to obtain a knowledge graph.
[0094] In a possible implementation manner, the conversion module 503 is specifically configured to: input the at least one target node, the connection relationship between the target entity and the at least one target node, the target business rule, and the data query text into a pre-trained statement conversion model to obtain a structured query statement, where the pre-trained statement conversion model is trained according to a sample query statement dataset.
[0095] Based on the same inventive concept, an embodiment of the present application provides another device for converting a natural language statement into SQL. The device is used to implement any of the above methods for converting a natural language statement into SQL, for example Figure 2 the method for converting a natural language statement into SQL shown, and the device can also implement the functions of the electronic device in the foregoing.
[0096] Please refer to Figure 6 for the structural schematic diagram of another device for converting a natural language statement into SQL provided by an embodiment of the present application. Figure 6 The device for converting a natural language statement into SQL shown includes at least one processor 601 and a memory 602 communicatively connected to the at least one processor 601.
[0097] Among them, the processor 601 can be a general-purpose processor or a dedicated processor, etc. The processor 601 includes, for example: a baseband processor or a central processing unit, etc. The baseband processor can be used to process communication protocols and communication data. The central processing unit can be used to Figure 6Control the device for converting natural language statements into SQL, execute software programs and / or process data. Different processors can be independent devices or can be provided in one or more processing circuits, for example, integrated on one or more application specific integrated circuits.
[0098] In one embodiment, the memory 602 stores instructions executable by at least one processor 601. The at least one processor 601 realizes the functions of the aforementioned electronic device by executing the instructions stored in the memory 602. Correspondingly, the steps executed by the aforementioned electronic device can also be realized.
[0099] In this embodiment, Figure 6 The device for converting natural language statements into SQL shown can also realize the functions of the aforementioned Figure 5 The device for converting natural language statements into SQL shown, and Figure 6 At least one processor 601 in the device for converting natural language statements into SQL shown can also realize the functions of the aforementioned parsing module 501, determination module 502, and conversion module 503.
[0100] Based on the same inventive concept, an embodiment of the present application provides a computer-readable storage medium. The computer-readable storage medium stores computer instructions. When the computer instructions run on a computer, the computer is caused to execute the method for converting natural language statements into SQL as described in any one of the above, for example, Figure 2 The method for converting natural language statements into SQL shown.
[0101] Based on the same inventive concept, an embodiment of the present application provides a computer program product, which includes computer instructions. When the computer instructions run on a computer, the method for converting natural language statements into SQL as described in any one of the above is realized, for example, Figure 2 The method for converting natural language statements into SQL shown.
[0102] It should be understood that in the embodiments of the present application, the processor can be a central processing unit (CPU). The processor can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.
[0103] In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor or the instructions in the form of software. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware processor, or executed and completed by a combination of the hardware and software modules in the processor. The software module can be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, register, etc. This storage medium is located in the memory, and the processor executes the instructions in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.
[0104] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0105] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0106] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be in an electrical, mechanical, or other form.
[0107] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0108] In addition, the functional units in each embodiment of the present application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0109] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the related technology, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.
[0110] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application can easily think of changes or substitutions, which should all be covered by the protection scope of this application. Therefore, the protection scope of this application shall be subject to the protection scope of the claims.
Claims
1. A method for converting natural language statements into SQL, characterized in that, Including: Semantically parsing the data query text to extract a target entity and at least one query field, where the target entity is used to indicate the query intent of the data query text; Determining the comprehensive similarity between each query field and each first candidate node in a pre-stored knowledge graph according to the semantic similarity and structural similarity between each query field in the at least one query field and each first candidate node in the pre-stored knowledge graph. The knowledge graph is constructed based on multiple data tables, and the structural similarity is used to indicate the degree of structural association between each query field and each first candidate node in the knowledge graph; Determining a target node corresponding to each query field according to the comprehensive similarity between each query field and each first candidate node; Converting the data query text into a structured query statement according to at least one target node, the connection relationship between the target entity and the at least one target node, target business rules, and the data query text.
2. The method according to claim 1, wherein The method further includes: Traversing the knowledge graph to determine the path length from the target entity to each first candidate node; Calculating the structural similarity between each query field and each first candidate node according to the path length from the target entity to each first candidate node and a preset inverse function.
3. The method according to claim 1, characterized in that, The determining the comprehensive similarity between each query field and each first candidate node according to the semantic similarity and structural similarity between each query field in the at least one query field and each first candidate node in the pre-stored knowledge graph includes: For each first candidate node among the first candidate nodes, calculating a first similarity according to the product result of the semantic similarity between each query field and each first candidate node and a first preset weight; Calculating a second similarity according to the structural similarity between each query field and each first candidate node and a second preset weight; Determining the comprehensive similarity between each query field and each first candidate node according to the sum of the first similarity and the second similarity, and obtaining the comprehensive similarity between each query field and each first candidate node.
4. The method according to claim 1, wherein Before determining the comprehensive similarity between each query field and each first candidate node according to the semantic similarity and structural similarity between each query field in the at least one query field and each first candidate node, the method further includes: Calculating the semantic similarity between each query field and each node in the knowledge graph; Determining each first candidate node according to the semantic similarity between each query field and each node and a first similarity threshold.
5. The method according to claim 1, wherein The determining a target node corresponding to each query field according to the comprehensive similarity between each query field and each first candidate node includes: Determining at least one second candidate node according to the comprehensive similarity between each query field and each first candidate node and a second similarity threshold; Input the at least one second candidate node, the data query text, and the context information of the at least one second candidate node in the knowledge graph into a pre-trained large language model to obtain the target node.
6. The method according to any one of claims 1-5, characterized in that, Before determining the comprehensive similarity between each query field and each of the first candidate nodes according to the semantic similarity and structural similarity between each query field in the at least one query field and each of the first candidate nodes, the method further includes: Extract the metadata information of each data table in the multiple data tables, where the metadata information includes table description information and column description information; According to the table description information and the column description information, construct multiple nodes and the attribute information of each node in the multiple nodes, where one node is used to indicate one table entity or column entity; According to the table description information and the column description information, extract the table-column relationship and column-column relationship between the multiple nodes; According to the table-column relationship and the column-column relationship, construct the relationship edges between the multiple nodes to obtain the knowledge graph.
7. The method according to any one of claims 1-5, characterized in that The converting the data query text into a structured query statement according to at least one target node, the connection relationship between the target entity and the at least one target node, the target business rule, and the data query text includes: Input the at least one target node, the connection relationship between the target entity and the at least one target node, the target business rule, and the data query text into a pre-trained statement conversion model to obtain the structured query statement, where the pre-trained statement conversion model is trained according to a sample query statement data set.
8. An apparatus for converting a natural language statement into SQL, characterized in that, Includes: A parsing module, configured to perform semantic parsing on the data query text to extract a target entity and at least one query field, where the target entity is used to indicate the query intention of the data query text; A determining module, configured to determine the comprehensive similarity between each query field and each of the first candidate nodes in a pre-stored knowledge graph according to the semantic similarity and structural similarity between each query field in the at least one query field and each of the first candidate nodes, where the knowledge graph is constructed according to multiple data tables, and the structural similarity is used to indicate the degree of structural association between each query field and each of the first candidate nodes in the knowledge graph; The determining module is further configured to determine a target node corresponding to each query field according to the comprehensive similarity between each query field and each of the first candidate nodes; A conversion module is further configured to convert the data query text into a structured query statement according to at least one target node, the connection relationship between the target entity and the at least one target node, the target business rule, and the data query text.
9. An apparatus for converting a natural language statement into SQL, characterized in that, Includes: A memory storing executable program code; A processor coupled to the memory; The processor calls the executable program code stored in the memory and executes the method according to any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
A natural language query method and system oriented to a software project knowledge map
CN109033135A
Multi-skill task type dialogue system construction method fusing chat and common sense
CN114153955A
Unmanned aerial vehicle fault diagnosis method, device and equipment and storage medium
CN118445450A
Intelligent AI question and answer method and device based on domain knowledge graph and table paraphrasing enhancement generation and program product
CN120030114A
Cited By
Database table business domain division method and device, medium and product
CN120705154A
Text-to-SQL (Structured Query Language) generation method and equipment, medium and product
CN120910088A
Text-to-sql generation method, device, medium and product
CN120910088B
Inference method and device based on large language model, electronic equipment and medium
CN121009995A
Document processing method and device, electronic equipment and storage medium
CN121256012A