SQL (Structured Query Language) generation method, system and equipment based on word segmentation and knowledge graph and medium

Through the SQL generation method based on word segmentation and knowledge graphs, the shortcomings of the data question-and-answer system in terms of semantic understanding, flexibility and graph update are solved, and more accurate semantic understanding, more flexible query processing and lower maintenance costs are achieved.

CN120216617APending Publication Date: 2025-06-27BEIJING ZHITONG YUNLIAN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510263993.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing data question and answer system has problems such as insufficient semantic understanding, lack of flexibility and difficulty in graph update in SQL generation.

Method used

The SQL generation method based on word segmentation and knowledge graph is adopted, key information is extracted through word segmentation annotation method, and SQL statements are generated using knowledge graphs, and the analysis strategy is dynamically adjusted to meet diverse query needs.

Benefits of technology

Improve the accuracy of semantic understanding, enhance the flexibility of the Q&A system, and reduce the complexity of knowledge graph update and maintenance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216617A_ABST
    Figure CN120216617A_ABST
Patent Text Reader

Abstract

The invention provides an SQL (Structured Query Language) generation method, system and equipment based on word segmentation and a knowledge graph and a medium, and the method comprises the following steps: judging the type of an input question, preprocessing the input question according to the type of the input question, and generating a standard input question; key information of the standard input question is extracted through a word segmentation labeling method, and standardization processing is conducted on the key information; an SQL statement is generated based on the key information by adopting a knowledge graph, and operation is performed according to the SQL statement to generate result information; performing customization processing and display on the generated result information; wherein the key information comprises time data, metadata, main data, an attribute value limiting condition, a data return condition, a sorting word and an aggregation function condition. According to the technical scheme provided by the invention, the natural language can be converted into the SQL language.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This document relates to the field of SQL generation technology, and particularly to a method, system, device, and medium for SQL generation based on word segmentation and knowledge graph. Background Art

[0002] In the field of data processing and analysis, data question-answering systems have become important tools. With the help of natural language processing technology, users can query and analyze information in the database in an intuitive and convenient way. However, existing data question-answering systems still face many challenges.

[0003] Currently, in terms of SQL generation, the following main deficiencies exist:

[0004] Insufficient semantic understanding: When many data question-answering systems process natural language queries, it is difficult to accurately capture the semantic information in the queries, resulting in inaccurate query results or failure to meet user needs.

[0005] Lack of flexibility: Existing systems often rely on fixed query templates or rules and are difficult to adapt to diverse query needs. Especially when facing complex or unconventional queries, the performance of the systems is often unsatisfactory.

[0006] Difficulty in updating the knowledge graph: As the core component of a data question-answering system, the update and maintenance of the knowledge graph often involve cumbersome manual operations, making it difficult to keep the graph content up-to-date and accurate. Summary of the Invention

[0007] According to an embodiment of the present invention, there is provided a method, system, device, and medium for SQL generation based on word segmentation and knowledge graph, which improves the accuracy of semantic understanding, enhances the flexibility of the question-answering system, and reduces the maintenance cost.

[0008] Method Embodiment

[0009] The present invention provides a method for SQL generation based on word segmentation and knowledge graph, including:

[0010] Determine the type of the input question, preprocess the input question according to the input question type, and generate a standard input question;

[0011] Extract key information of the standard input question using the word segmentation annotation method, and standardize the key information;

[0012] Generate an SQL statement based on the key information using the knowledge graph, and perform operations according to the SQL statement to generate result information;

[0013] Perform customization processing and display on the generated result information;

[0014] Among them, the key information includes: time data, metadata, master data, attribute value restriction conditions, data return conditions, sorting terms, and aggregation function conditions.

[0015] System embodiment

[0016] The present invention provides an SQL generation system based on word segmentation and knowledge graph, including:

[0017] A problem generation module, configured to determine the type of the input problem, preprocess the input problem according to the input problem type, and generate a standard input problem;

[0018] A key information extraction module, which extracts the key information of the standard input problem by using the word segmentation annotation method and performs standardization processing on the key information;

[0019] A result generation module, which generates an SQL statement based on the knowledge graph based on the key information and generates result information according to the SQL statement;

[0020] A post-processing module, which performs customization processing and display on the generated result information;

[0021] Among them, the key information includes: time data, metadata, master data, attribute value restriction conditions, data return conditions, sorting terms, and aggregation function conditions.

[0022] Device embodiment one

[0023] The present invention provides an electronic device, including:

[0024] A processor; and,

[0025] A memory arranged to store computer-executable instructions, and the computer-executable instructions, when executed, cause the processor to execute the steps of the above-mentioned SQL generation method based on word segmentation and knowledge graph.

[0026] Device embodiment two

[0027] The present invention provides a storage medium for storing computer-executable instructions, and the computer-executable instructions, when executed, implement the steps of the above-mentioned SQL generation method based on word segmentation and knowledge graph.

[0028] A storage medium for storing computer-executable instructions, and the computer-executable instructions, when executed, implement the steps of the SQL generation method based on word segmentation and knowledge graph as described in any one of claims 1-7

[0029] By adopting the embodiments of the present invention, the input problem is deeply analyzed using the graph structure, effectively identifying and understanding complex semantic relationships in the context, including synonyms, related words, sorting words, etc., thereby greatly improving the accuracy of semantic understanding. Through the configurability of the graph, semantic rules can be customized for specific fields or application scenarios, effectively reducing the understanding ambiguity caused by polysemy or context differences, and ensuring that the problem is accurately interpreted. The graph configuration technology allows the system to dynamically adjust the parsing strategy according to the diversity of user questions. Whether it is an open question, a specific query or a compound sentence pattern, it can flexibly respond and provide accurate answers. By constructing a cross-domain semantic graph, the system can integrate and reason information across different knowledge fields, achieve accurate answers to cross-domain questions, and broaden the application scenarios of the question-answering system. In a configurable manner, the present invention reduces the complexity of knowledge graph update and maintenance and reduces the need for manual intervention. Description of the Drawings

[0030] In order to more clearly illustrate the technical solutions in one or more embodiments of this specification or in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0031] Figure 1 Flowchart of the SQL generation method based on word segmentation and knowledge graph according to the embodiments of the present invention;

[0032] Figure 2 Framework diagram of the SQL generation method based on word segmentation and knowledge graph according to the embodiments of the present invention;

[0033] Figure 3 Flowchart of entity recognition according to the embodiments of the present invention;

[0034] Figure 4 Flowchart of SQL processing according to the embodiments of the present invention;

[0035] Figure 5 Schematic diagram of the SQL generation system based on word segmentation and knowledge graph according to the embodiments of the present invention. Detailed Embodiments

[0036] To enable those skilled in the art to better understand the technical solutions in one or more embodiments of this specification, the following will clearly and completely describe the technical solutions in one or more embodiments of this specification in conjunction with the accompanying drawings in one or more embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this document.

[0037] Method Embodiment

[0038] According to an embodiment of the present invention, a method for generating SQL based on word segmentation and knowledge graph is provided. Figure 1 The following is a flowchart of the method for generating SQL based on word segmentation and knowledge graph according to an embodiment of the present invention, as Figure 2 The following is a framework diagram of the method for generating SQL based on word segmentation and knowledge graph according to an embodiment of the present invention. According to Figure 1 and Figure 2 shown, the method for generating SQL based on word segmentation and knowledge graph according to an embodiment of the present invention specifically includes:

[0039] S1. Determine the type of the input problem, preprocess the input problem according to the input problem type, and generate a standard input problem;

[0040] The preprocessing of the input problem according to the input problem type specifically includes:

[0041] If the input problem is a replacement problem, parse the fixed question method in the replacement problem, and obtain the fixed question method based on a preset rule matching method and replace it with an identifiable entity word;

[0042] If the input problem is a comparison type problem, disassemble the problem according to a preset regular matching method.

[0043] This step ensures that all input problems conform to the consistent format for internal processing of the system, thereby improving the accuracy of the entity recognition task. The embodiments of the present invention support preprocessing optimizations for two types of problems, specifically as follows:

[0044] (1) Parse the fixed question method in the replacement problem

[0045] First, design a rule matching to obtain the fixed question method, and replace the fixed question method with an identifiable entity word.

[0046] The currently supported fixed question methods are from a certain unit to a certain unit, from a certain unit to a certain single;

[0047] Question example: What is the oil production from Factory 1 to Factory 10 today?

[0048] Replacement effect: Replace "Factory 1 to Factory 10" with "Factory 1, Factory 2, Factory 3, Factory 4, Factory 5, Factory 6, Factory 7, Factory 8, Factory 9, Factory 10".

[0049] (2) Preprocessing of comparison - type questions

[0050] First, design regular expressions to match comparison - type questions, identify comparison items through rules, and then break down the questions into two questions.

[0051] Question example: What is the difference in oil production between Factory 1 and Factory 10?

[0052] The breakdown effect is as follows:

[0053] What is the difference in oil production between Factory 1 and?

[0054] What is the difference in oil production between Factory 10 and?

[0055] S2. Extract key information from the standard input question using the word - segmentation annotation method and standardize the key information; the key information includes: time data, metadata, master data, attribute value restriction conditions, data return conditions, sorting words, and aggregation function conditions. In S2, first, through word - segmentation annotation, extract key information from the input natural - language question, including time, metadata, master data, data value restrictions, data entry restrictions, sorting words, and aggregation function conditions, etc. Standardize the format of the key information (such as time, data value restrictions, etc.) extracted after word - segmentation to ensure that they conform to the syntax requirements of SQL statements. Taking time information as an example, convert the time expressed in natural language (such as "September 18, 2024") into a date format acceptable to SQL (such as "2024 - 09 - 18"). This process aims to eliminate the diversity of natural - language expressions and unify them into a standard format recognizable by SQL queries. Figure 3 This is the entity - recognition flowchart of the embodiment of the present invention.

[0056] In the embodiment of the present invention, the combination word - segmentation method is used to extract metadata and master data from the standard input question, and based on the metadata and master data, locate specific SQL tables and fields. Metadata is the header of the actual data table, and master data is the real - entity data. First, for the pre - processed question, extract metadata and master data through the combination word - segmentation method, and locate specific SQL tables and fields accordingly. The specific extraction steps are as follows:

[0057] (1) Preliminary word - segmentation and keyword arrangement

[0058] Use the master data dictionary and the metadata dictionary to perform word segmentation on the preprocessed question with long words prioritized; obtain a keyword list containing metadata and master data. And obtain the metadata words corresponding to the master data through dictionary matching; at the same time, obtain the remaining text after removing these keywords, that is, the remaining string of the question after removing keywords.

[0059] Question example: What is the oil production of Factory 1?

[0060] Keyword list: {"Factory 1": "Z", "oil production": "Y"}, where Z represents the master data and Y represents the metadata;

[0061] Master data mapping metadata dictionary: {"Factory 1": ["unit", "production unit"]}, indicating the metadata words corresponding to the unit of Factory 1

[0062] Remaining text: What is?

[0063] (2) Perform metadata word segmentation, including: separately use the metadata dictionary to perform word segmentation on the question again with long words prioritized, and obtain an independent metadata word list.

[0064] Question example: What is the oil production of Factory 1?

[0065] Metadata word list: {"oil production": "Y"}

[0066] (3) Perform common word segmentation: Use the jieba tokenizer to perform common word segmentation on the remaining string, and obtain the list of common words in it.

[0067] Question example: What is the oil production of Factory 1?

[0068] Remaining text: What is?

[0069] Common word list: ["and", "of", "what is", "?"];

[0070] (4) Optimize metadata words with common word combinations

[0071] Combine the words in the common word list with the metadata words in the metadata word list, and check whether these combined words are in the metadata dictionary; if the combined words exist, add the combined words to the keyword list and the metadata word list. Through this step, mine implicit relationships and accurately locate fields.

[0072] Before optimization:

[0073] Question example: What are the reasons for the wells not being opened;

[0074] Keyword list: {"wells not being opened": "Z"}, where "wells not being opened" is a master data word;

[0075] Master data mapping metadata dictionary: {"Well not opened": [Production type]}, where the metadata word corresponding to the production type of the well not opened

[0076] Remaining text: What are the reasons?

[0077] Metadata word list: {}

[0078] Common word list: ["of", "reason", "what are", "?"];

[0079] After optimization:

[0080] Keyword list: {"Well not opened": "Z", "Well not opened - reason": "Y"};

[0081] Metadata word list: {"Well not opened - reason": "Y"};

[0082] (5) Optimize the metadata word list for the metadata combination corresponding to the master data

[0083] Combine the metadata corresponding to the master data in the keyword list with the metadata words in the metadata word list, and check whether these combined words exist in the metadata dictionary; if the combined words exist, replace the corresponding metadata words in the keyword list and the metadata word list; through this step, mine the implicit relationships and accurately locate the fields.

[0084] Before optimization:

[0085] Question example: What is the production of layer Q9?

[0086] Keyword list: {"Layer Q9": "Z", "Production": "Y"}, where Q9 is a master data word. Master data mapping metadata dictionary: {"Layer Q9": ["Layer position", "Layer name", "Layer title"]}

[0087] Remaining text: is what?

[0088] Metadata word list: {"Production": "Y"}

[0089] Common word list: ["of", "is what", "?"]

[0090] After optimization:

[0091] Keyword list: {"Layer Q9": "Z", "Layer position - production": "Y"};

[0092] Metadata word list: {"Layer position - production": "Y"};

[0093] Adopt a method combining oral dictionary recognition and regular expression recognition to extract time data and attribute value limit conditions;

[0094] The specific steps for extracting time data are as follows:

[0095] For time extraction, a method combining spoken language dictionary recognition and regular expression recognition is adopted. First, through a pre-constructed spoken language dictionary, time-related words in the question can be recognized, such as "today", "this month", etc. Second, using regular expressions, fixed-format dates in the question can be matched and extracted, such as "August 8, 2024" or "2024-09-08", etc. Finally, the time is standardized and processed into a conforming SQL format.

[0096] Question example: Today's daily oil production situation;

[0097] Extraction result: {Today: TO_DATE(RQ,'YYYY-MM-DD') = TRUNC(SYSDATE)};

[0098] The steps for extracting attribute value restriction conditions are as follows:

[0099] For time extraction, a method combining spoken language dictionary recognition and regular expression recognition is also adopted. Through the spoken language dictionary, non-numerical words expressing ranges or boundaries, such as "more than", "above", etc., can be captured. They are often accompanied by specific numerical values and jointly define the restriction range of the attribute. This step ensures that even in the face of complex and variable natural language expressions, key restriction information can be accurately extracted. Second, using regular expressions, symbols such as ">" and "<=" are recognized, and they and the numerical values following them are recognized and extracted to form clear numerical restriction conditions. Finally, the recognized restriction conditions are converted into the corresponding SQL condition format.

[0100] Question example: Which oil production plants have a daily oil production > 300 today?

[0101] Extraction result: {compare: > 300}

[0102] Use regular expressions to extract the number of data returned and the sorting term;

[0103] The steps for extracting the number of data returned using regular expressions are as follows:

[0104] When extracting the number of data returned, use regular expressions to capture ordinal phrases such as "the first few", "the last few", "the first several", "the last several", etc., and analyze the direction words "before" and "after" and the specific quantity. Finally, replace the direction words with the corresponding SQL sorting methods.

[0105] Question example: The top three units with the highest daily oil production today;

[0106] Extraction result: {Restriction word: the first three, Direction word: before, Quantity: 3};

[0107] The steps to extract sorting words using regular expressions are as follows: When extracting sorting words, use regular expressions to capture sorting words such as "maximum" and "minimum", and convert them into sorting methods.

[0108] Example of a question: What is the unit with the maximum daily oil production today?

[0109] Extraction result: {maximum: DESC};

[0110] The steps to extract aggregate function conditions using an oral dictionary are as follows:

[0111] Aggregate function extraction is captured through an oral dictionary, which can identify oral words such as summation, average value, and how many in the question. And obtain their corresponding aggregate functions.

[0112] Example of a question: What is the sum of the daily oil production of the oil production plant today?

[0113] Extraction result: {sum: SUM()}.

[0114] S3. Use the knowledge graph to generate an SQL statement based on the key information, and perform operations according to the SQL statement to generate result information, specifically including:

[0115] Locate the table according to the knowledge graph: First, traverse the list of metadata words, check each metadata word one by one to see if all metadata words are connected to the same table. Find the table to which all metadata words are connected (if there are multiple such tables, each table is regarded as a candidate table); if there is at least one such table, the location is successful; otherwise, the location fails because no table to which all metadata words are connected is found. Thus, all candidate tables are obtained.

[0116] Locate fields in the graph: Traverse the list of keywords and check the connection between the keywords and the candidate tables. If a keyword can be connected to a candidate table, record the path length from it to the candidate table; otherwise, the location fails;

[0117] Graph path screening specifically includes: Calculate the number of connected tables: For each connection situation between keywords and candidate tables, calculate the number of directly connected candidate tables. Retain the situation with the fewest connected tables: Select the combination with the fewest connected tables. Calculate the path length: For the retained combination, calculate the total length of the paths from all keywords to the candidate tables. Retain the situation with the shortest sum of paths: From the combination with the fewest connected tables, select the combination with the shortest total path length.

[0118] SQL element recognition: First, traverse all the above paths, make a judgment on the paths. If a path contains a field node or a master data node, process this path for element addition; otherwise, skip this path without processing. If it contains a field or master data, determine whether it only contains a field. If it only contains a field, add it to the SELECT element; otherwise, add it to the WHERE element. Then, if the path only contains a field, determine whether the field value is equal to the value of the starting node of the path. If not, add it to the synonym list. If the path also contains master data, determine whether the master data value is equal to the starting node value. If not, also add it to the synonym list. Finally, if all paths are not processed, traverse all paths again and add the table fields in the paths to the TABLE element.

[0119] Figure 4 This is the SQL processing flow chart of the embodiment of the present invention. The SQL processing flow is as follows:

[0120] First, determine whether it is a multi-table query or a single-table query. If it is a multi-table query, disassemble it into single tables through table splitting for query, and finally splice them;

[0121] For each single-table SQL splicing, first search for the corresponding relationship of the aggregation function, determine which field the restriction condition identified in the word segmentation annotation corresponds to, and find its corresponding relationship. Search logic: For numerical restriction conditions, search for the nearest field in SELECT and WHERE; for aggregation functions, search for the nearest field in SELECT;

[0122] Then replace the synonyms in the WHERE condition with the standard words in the library, and ensure that all fields in the Where condition are in the SELECT condition. If the configured primary key does not exist in SELECT, add it to SELECT as well;

[0123] Sort the conditions in Select according to the REDIS column sorting dictionary;

[0124] Perform single-table SQL splicing according to the SELECT and WHERE conditions;

[0125] After assembling all the single-table SQLs, perform multi-table SQL splicing, and by default, use the RQ field for splicing;

[0126] Generate the ORDER BY statement and the LIMIT statement through the identified restriction conditions and the corresponding fields;

[0127] If it is an aggregation function, the GROUP BY statement also needs to be added.

[0128] S4. Perform customized processing and display on the generated result information;

[0129] The SQL post - processing mainly performs special processing on the query results of complex problems. This link is also responsible for encapsulating the result information of the SQL query and formatting the data required for front - end display to ensure that the data can be accurately and clearly presented to users.

[0130] For the special processing of comparison - type problems, first obtain the SELECT elements and WHERE conditions of the two problems, display the fields with different conditions separately, display only one of the fields with the same conditions, display the fields that only exist in the SELECT elements separately, and display their differences.

[0131] Encapsulate the aliases, field names, and field codes in the SQL and return the results.

[0132] By adopting the embodiments of the present invention, the following beneficial effects are achieved:

[0133] 1. Improve the accuracy of semantic understanding

[0134] Precisely analyze complex semantics: Use the graph structure to deeply analyze the input problem, effectively identify and understand the complex semantic relationships in the context, including synonyms, related words, sorting words, etc., thereby greatly improving the accuracy of semantic understanding.

[0135] Reduce ambiguous interpretations: Through the configurability of the graph, semantic rules can be customized for specific fields or application scenarios, effectively reducing the understanding ambiguities caused by polysemy or context differences, and ensuring that the problem is accurately interpreted.

[0136] 2. Enhance the flexibility of the question - answering system

[0137] Dynamically adapt to diverse questions: The graph configuration technology allows the system to dynamically adjust the parsing strategy according to the diversity of user questions. Whether it is an open - ended question, a specific query, or a complex sentence pattern, it can flexibly respond and provide accurate answers.

[0138] Support cross - domain question - answering: By constructing a cross - domain semantic graph, the system can integrate and reason information across different knowledge fields, achieve accurate answers to cross - domain questions, and broaden the application scenarios of the question - answering system.

[0139] 3. Reduce maintenance costs:

[0140] In a configurable manner, the present invention reduces the complexity of knowledge graph update and maintenance and reduces the need for manual intervention.

[0141] System embodiments

[0142] According to the embodiments of the present invention, a SQL generation system based on word segmentation and knowledge graph is provided. Figure 5Schematic diagram of the SQL generation system based on word segmentation and knowledge graph according to an embodiment of the present invention. According to Figure 5 As shown, the SQL generation system based on word segmentation and knowledge graph according to an embodiment of the present invention specifically includes:

[0143] A problem generation module 50, configured to determine the type of the input problem, preprocess the input problem according to the input problem type, and generate a standard input problem;

[0144] A key information extraction module 52, which extracts the key information of the standard input problem by using the word segmentation annotation method and performs standardization processing on the key information;

[0145] A result generation module 54, which generates an SQL statement based on the knowledge graph based on the key information and generates result information according to the SQL statement;

[0146] A post-processing module 56, which performs customization processing and display on the generated result information;

[0147] Among them, the key information includes: time data, metadata, master data, attribute value restriction conditions, data return conditions, sorting terms, and aggregation function conditions.

[0148] Device Embodiment 1

[0149] According to an embodiment of the present invention, an electronic device is provided, including:

[0150] A processor; and,

[0151] A memory arranged to store computer-executable instructions, and when the computer-executable instructions are executed, the processor executes the steps of the SQL generation method based on word segmentation and knowledge graph as described above.

[0152] Device Embodiment 2

[0153] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than limiting them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A SQL generation method based on word segmentation and knowledge graph, characterized in that include: Determine the type of input question, pre-process the input question according to the input question type, and generate a standard input question; The key information of the standard input question is extracted by using the word segmentation and tagging method, and the key information is standardized; Generate SQL statements based on the key information using a knowledge graph, and generate result information by performing operations according to the SQL statements; Customize and display the generated result information; The key information includes: time data, metadata, master data, attribute value restriction conditions, data return conditions, sorting terms and aggregation function conditions.

2. The method according to claim 1, characterized in that The preprocessing of the input question according to the input question type specifically includes: If the input question is a replacement question, the fixed question in the replacement question is parsed, and the fixed question is obtained based on a preset rule matching method and replaced with a recognizable entity word; If the input question is a comparison question, the question is broken down according to a preset regular matching method.

3. The method according to claim 1, characterized in that The key information of the standard input problem extracted by word segmentation and tagging method specifically includes: Extract metadata and main data from standard input questions by combining word segmentation, and locate specific SQL tables and fields based on the metadata and main data; The method combining spoken dictionary recognition and regular expression recognition is used to extract time data and attribute value constraints; Use regular expressions to extract data to return the number of entries and sorting words; Use spoken language dictionary to extract aggregation function conditions.

4. The method according to claim 3, characterized in that The extracting metadata and master data specifically includes: Preliminary word segmentation and keyword sorting: Use the main data dictionary and metadata dictionary to perform word segmentation with long words first on the preprocessed questions to obtain a keyword list containing metadata and main data. Determine the metadata corresponding to the main data through dictionary matching, and save the remaining text content after removing the main data and metadata. Metadata word segmentation: Use the metadata dictionary to segment the question again with long words first, thus obtaining a list of metadata words; Common word segmentation: Use the Jieba word segmenter to segment the remaining text content obtained after preliminary segmentation and keyword sorting to obtain a common word list; Optimizing metadata words by combining common words: combining words in the common word list with metadata words in the metadata word list in sequence to obtain combined words, and determining whether the combined words exist in the metadata dictionary. If so, adding the combined words to the keyword list and the metadata word list respectively; The metadata corresponding to the master data is combined to optimize the metadata word list: the metadata corresponding to the master data in the keyword list is combined with the metadata words in the metadata word list, and the combined word is checked to see if it exists in the metadata dictionary. If so, the combined word is used to replace the corresponding metadata words in the keyword list and the metadata word list.

5. The method according to claim 3, characterized in that: The method of extracting time data by combining spoken dictionary recognition with regular expression recognition specifically includes: Identify time words in questions using a pre-built spoken dictionary; Use regular expressions to match and extract fixed-format dates in the question; Standardize the time and process it into SQL format.

6. The method according to claim 1, characterized in that The generating of SQL statements based on the key information by using the knowledge graph and generating result information by operating according to the SQL statements specifically include: Locate the table based on the knowledge graph: traverse the metadata word list and find the table to which all metadata words are connected. If it exists, the positioning is successful and the table is obtained as a candidate table. If it does not exist, the positioning fails. Graph positioning field: traverse the keyword list and check the connection between the keyword and the candidate table. If the keyword can be connected to the candidate table, the path length to the candidate table is recorded, otherwise the positioning fails; Graph path screening: Calculate the number of connection tables for each keyword and candidate table, retain the combination with the least number of connection tables, then calculate the total length of the paths from all keywords to the candidate tables in the retained combination, and select the combination with the shortest total path length as the candidate path; SQL element recognition: traverse candidate paths, and if the candidate path contains field nodes or master data nodes, add elements; determine the path content, add those containing only fields to the SELECT element, and add others to the WHERE element; check the field value or master data value and the starting node value at the same time, and add those that are not equal to the synonym list; if all paths are not processed, traverse the path again and add the table fields to the TABLE element.

7. The method according to claim 6, characterized in that After generating the SQL statement, the method further includes: identifying the query type involved in the SQL statement, and if it is a complex query, breaking it down into multiple simple query units; optimizing the query conditions for each simple query unit, including but not limited to integrating conditions with similar semantics, refining fuzzy conditions, and determining the correspondence between query conditions and data fields; processing the fields of the query results to ensure the integrity and accuracy of the query fields, and sorting the fields according to preset sorting rules.

8. A SQL generation system based on word segmentation and knowledge graph, characterized by include: A question generation module is used to determine the type of input question, pre-process the input question according to the input question type, and generate a standard input question; The key information extraction module uses a word segmentation and tagging method to extract key information of the standard input question and performs standardization processing on the key information; A result generation module, which uses a knowledge graph to generate SQL statements based on the key information, and generates result information by performing operations according to the SQL statements; The post-processing module customizes and displays the generated result information; The key information includes: time data, metadata, master data, attribute value restriction conditions, data return conditions, sorting terms and aggregation function conditions.

9. An electronic device, comprising: processor; as well as, A memory arranged to store computer executable instructions, which, when executed, cause the processor to perform the steps of the SQL generation method based on word segmentation and knowledge graph as described in any one of claims 1 to 7.

10. A storage medium for storing computer executable instructions, which, when executed, implement the steps of the SQL generation method based on word segmentation and knowledge graph as described in any one of claims 1 to 7.