A Method for Enhancing NL2SQL Questions Driven by Multivariate Knowledge Linkage
By building knowledge of power data mode and business-mode mapping relationship, combining large language models for problem analysis and knowledge retrieval, designing problems to enhance Prompt templates, reconstructing and optimizing user problems, solving the problem that the existing NL2SQL technology cannot be effectively and practical in the power field, and significantly improving the accuracy and reliability of SQL generation.
Patent Information
- Application Number
- CN202510237478.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-03-03
AI Technical Summary
The existing NL2SQL technology cannot be effective and practical in the power field. The main reason is that the problem of natural language description is not enough to allow the big model to accurately understand the query intentions of business personnel, resulting in low accuracy of generated SQL.
The NL2SQL problem enhancement method driven by multi-knowledge link is adopted. By constructing the power data mode and business-mode mapping relationship knowledge, combining large language models for problem analysis and knowledge retrieval, designing problem enhancement Prompt templates, reconstructing and optimizing user problems, to improve the accuracy of SQL generation.
It significantly improves the accuracy and reliability of SQL generation in large-scale models, reduces ambiguity and ambiguity, enhances domain specificity, and enables users' natural language data queries to better match power database tables.
Smart Images

Figure CN119739837B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of power natural language data question answering, and particularly relates to a method for enhancing NL2SQL (Natural Language to SQL) questions driven by multi-source knowledge links. Background Art
[0002] The current situation that business personnel in the power field need to rely on the structured query language SQL written by system developers to access data greatly restricts the exertion of the value of power data. Recently, the NL2SQL technology that converts natural language into the structured query language SQL understandable by the database based on large language models has made great progress, but this technology still cannot be effectively applied in the power field. An important reason is that the questions described in natural language are not sufficient for the large model to accurately understand the query intentions of business personnel:
[0003] Firstly, the professionalism of power business is extremely strong, and general large language models are difficult to effectively understand various terms in the power field and cannot fully understand the data query requirements expressed in business terms, thus generating incorrect data query results.
[0004] Secondly, for business personnel who are not familiar with the database architecture and SQL language, the information in their data query requirements expressed in natural language is often insufficient. The large language model cannot obtain enough information to effectively map the query demands to data entities, thus unable to generate the correct SQL.
[0005] Thirdly, business personnel are used to making short data queries based on the business knowledge they have mastered. However, the large language model cannot understand the complex business knowledge hidden behind these queries through literal descriptions, thus unable to correctly understand the query requirements, and ultimately resulting in the accuracy of the generated SQL not meeting the requirements.
[0006] Finally, the complexity of power business and system diversity lead to extremely complex data structures, a very large number of library tables, and a high degree of similarity between table names and field names. This makes the large language model extremely prone to confusing data entities when generating SQL, resulting in low accuracy of the generated SQL. Summary of the Invention
[0007] In order to overcome the problems in the prior art, the present invention proposes a method for enhancing NL2SQL questions driven by multi-source knowledge links.
[0008] The technical solution of the present invention to solve the above technical problems is as follows:
[0009] The present invention provides a method for enhancing NL2SQL questions driven by multi-source knowledge links, including the following steps:
[0010] Step 100: Construct a power data schema using database table information and sort out the knowledge in the power field, where the knowledge in the power field includes business-schema mapping relationship knowledge and data coding relationship knowledge;
[0011] Step 200: Construct a problem analysis Prompt template and use a large language model to analyze the structure of the original problem, and extract key entities from the original problem;
[0012] Step 300: Based on the sorted out knowledge in the power field and the key entities extracted from the original problem, retrieve the knowledge in the power field by means of hybrid similarity retrieval;
[0013] Step 400: Obtain the database tables and data schemas related to the original problem through a multi-level schema linking method;
[0014] Step 500: Based on the retrieved knowledge in the power field and the obtained data schema, perform knowledge standardization, design a problem enhancement Prompt template, and use a large language model to reconstruct and enhance the original problem.
[0015] Further, in the step 100, the construction of the power data schema includes:
[0016] Extract the data definition language (DDL) of all database tables related to power operations from the power database, where each DDL contains the creation information of the database table, and extract the required database table information by parsing the DDL; collect the data example content in each database table; reorganize the database table information and data examples extracted from the database to construct a data schema.
[0017] Further, in the step 100, the business-schema mapping relationship knowledge includes:
[0018] Sort out the business terms in the power field and form a corresponding relationship between the business terms and the data schema, that is, the business-schema mapping relationship;
[0019] Extract the characteristics of business terms, where the characteristics of business terms include the vectorized business term names and the n - gram characteristics;
[0020] Store all the business-schema mapping relationship knowledge in a JSON file; store all the vectorized business term name knowledge in a vector database; store all the n - gram characteristic knowledge of the extracted business term names in an ES database.
[0021] Further, in the step 100, the data coding relationship knowledge includes:
[0022] Extract the required information from the table storing the encoding correspondence in the database, including the encoding name, encoding value, and corresponding column name;
[0023] For each encoding relationship, extract its vector features and n - gram features, and store the extracted features in the encoding relationship vector database and the encoding relationship ES database.
[0024] Furthermore, in the step 200, construct a problem parsing Prompt template, and the problem parsing Prompt template is defined as a tuple , including:
[0025] ;
[0026] Among them, G represents the defined task role goal, that is, defining a task-related role for the large language model and clarifying the task goal; I represents important reminders, including important tips and constraints when completing the task; O represents the output format of the task; C represents the provided reference output case; Q represents the original problem input.
[0027] Furthermore, in the step 300, the method of retrieving power domain knowledge by using the hybrid similarity retrieval method includes jointly retrieving business - pattern mapping relationship knowledge by using the method of fusing vector similarity and BM25 similarity:
[0028] Calculate the cosine similarity between the entity and business terms to retrieve the business term names with similar semantics to the entity in the vector database;
[0029] Use BM25 similarity to retrieve business term names with the same keyword fragments in the ES database;
[0030] According to the vector cosine similarity score, obtain the most relevant business term name of the key entity in the vector database; according to the BM25 similarity score, also obtain the most relevant business term name in the ES database; merge the two business term names as the business term name knowledge set of the entity.
[0031] Furthermore, in the step 300, the method of retrieving power domain knowledge by using the hybrid similarity retrieval method includes jointly retrieving encoding relationship knowledge by using the method of fusing vector similarity and BM25 similarity:
[0032] The cosine similarity and BM25 similarity methods are used to calculate the similarity of the content of each key entity in the encoded relationship vector library and the encoded relationship ES library, and the encoded relationship knowledge set corresponding to all key entities is obtained.
[0033] Further, in the step 400, through the multi-level mode linking method, the database tables and data modes related to the original problem are obtained, including:
[0034] Based on the fast linking method of double hash value search, the extracted original problem entities are respectively linked with the tables and fields in the database to obtain the database tables and data modes related to the original problem.
[0035] Further, in the step 500, the knowledge standardization includes:
[0036] Based on the encoded relationship knowledge retrieval result, the corresponding encoded column names are extracted according to its organizational form, and the encoded relationships corresponding to the column names not in the relevant data mode are removed; and the final representation forms of the data mode, data encoding mapping relationship and business-mode mapping relationship are standardized.
[0037] Further, in the step 500, the problem enhancement Prompt template is defined as P G , including:
[0038] ;
[0039] Among them, G represents the defined task role goal, that is, a task-related role is assigned to the large language model and the task goal is clarified; is the content of the formatted data mode; EK is the content of the formatted encoded relationship knowledge; T K is the content of the formatted business-mode mapping relationship knowledge; W represents important reminders, including important tips and constraints when completing the task; O represents the output format of the task; Q represents the original problem input; Q A is the obtained problem analysis content.
[0040] Compared with the prior art, the present invention has the following technical effects:
[0041] (1)The present invention aims to improve the accuracy of SQL generation by large models through the following steps: First, by parsing the original question, key entities and the structure information of the original question in the original question are obtained; Subsequently, using the hybrid similarity retrieval technology, business knowledge related to the key entities is retrieved; Then, relevant table information knowledge is obtained through the pattern linking method. On this basis, through knowledge standardization and problem enhancement Prompt design, the understanding of the large language model for power domain knowledge is enhanced, and finally a question with clear expression and clear logic is generated, effectively excluding confusion and interference factors, better matching the power database structure, thereby improving the accuracy and reliability of SQL queries.
[0042] (2)The present invention proposes a construction method for data patterns and business-pattern mapping relationships only for the power domain. Among them, the data pattern involves database table structure parsing, data example extraction, and pattern organization, ensuring the clarity and usability of library table knowledge; The business-pattern mapping relationship involves the collation of common businesses and explanations, the sorting of the corresponding relationships between businesses and library tables and fields, and the organizational storage of knowledge, etc. This content not only provides a detailed description of power business, but also clarifies the relationship between business and specific tables and fields, providing necessary knowledge support for the positioning of power domain tables and problem enhancement.
[0043] (3)The present invention proposes a multi-level pattern linking method. On the one hand, it adopts a data pattern linking method based on double hash value search. Combining the key entity extraction results in the original question, through the double search strategy of tables and fields, it can quickly and accurately match the user's natural language question with the most relevant table in the database. This method improves the accuracy of pattern linking and greatly reduces redundant content; On the other hand, based on the business-pattern mapping relationship knowledge obtained by retrieving the key entities of the original question, the association logic between relevant businesses and library tables in the user's question is obtained, and more accurate positioning results are provided in the way of the association between business and library tables. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required to be used in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to these drawings without creative efforts.
[0045] Figure 1 It is a flowchart of the problem enhancement method proposed by the present invention;
[0046] Figure 2 It is a schematic diagram of the data pattern and business knowledge content organized and constructed by the present invention;
[0047] Figure 3It is the specific result of the semantic parsing step of the original problem in the experiment;
[0048] Figure 4 It is the specific result of the entity-based business knowledge linking step in the experiment;
[0049] Figure 5 It is the specific result of double hash value search and table filtering in the experiment;
[0050] Figure 6 It is the specific result of the problem enhancement step in the experiment;
[0051] Figure 7 It is the effect comparison chart of generating SQL results in the experiment;
[0052] Figure 8 It shows the specific content of the problem parsing Prompt template;
[0053] Figure 9 It shows the specific content of the library table filtering Prompt template;
[0054] Figure 10 It shows the specific content of the problem enhancement Prompt template. Specific implementation manners
[0055] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the specific implementation manners, structures, features and their effects of the technical solutions proposed according to the present invention are described in detail below with reference to the accompanying drawings and preferred embodiments. Specific features, structures or characteristics in one or more embodiments may be combined in any suitable form. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs.
[0056] The present invention proposes a method for enhancing NL2SQL problems driven by multi-source knowledge linkage, which is used to improve the accuracy of business personnel in the power field when using natural language to achieve reliable data access. Through a series of key steps, including power data schema construction and business knowledge sorting, original problem semantic parsing, entity-based business knowledge linkage, multi-level data schema linkage and other strategies, the necessary knowledge is obtained, and the large language model is combined to optimize and enhance the expression mode and structure of the natural language requirements of business personnel for querying data. The enhanced user data query requirements enhance domain specificity, significantly reduce ambiguity and vagueness, and increase the data schema mapping of the business background, enabling better matching with power database tables. In other words, the core creativity of the present invention lies in enhancing and reconstructing the natural language data queries of business personnel through multi-source knowledge linkage, transforming the short and ambiguous query expressions of business personnel into forms that better conform to business logic and data schemas, and eliminating confusing and interfering factors, so that the large model can accurately understand the data query requirements of business personnel, thereby improving the quality of generated SQL and finally obtaining data that meets the requirements.
[0057] In one embodiment of the present invention, referring to Figure 1 , a method for enhancing NL2SQL problems driven by multi-source knowledge linkage is provided, including the following steps:
[0058] Step 100: Construct a power data schema using database table information, and sort out power domain knowledge, where the power domain knowledge includes business-schema mapping relationship knowledge and data coding relationship knowledge;
[0059] Step 200: Construct a problem parsing Prompt template, and use the large language model to analyze the structure of the original problem to extract key entities from the original problem;
[0060] Step 300: Based on the sorted power domain knowledge and the key entities extracted from the original problem, retrieve power domain knowledge by means of hybrid similarity retrieval;
[0061] Step 400: Through a multi-level schema linkage method, obtain database tables and data schemas related to the original problem;
[0062] Step 500: Based on the retrieved power domain knowledge and the obtained data schemas, perform knowledge standardization; and design a problem enhancement Prompt template, and use the large language model to reconstruct and enhance the original problem.
[0063] The above steps are detailed as follows:
[0064] Step 100: Construct a power data schema and sort out the knowledge in the power field, where the knowledge in the power field includes business - schema mapping relationship knowledge and data coding relationship knowledge.
[0065] This embodiment analyzes the characteristics of power data queries in terms of business term descriptions and database design patterns, constructs a power data schema using database table information, and sorts out domain knowledge such as business - schema mapping relationships and data coding relationships. In addition, a unique knowledge organization form is designed according to different contents, and finally the knowledge is stored in a database or a JSON (JavaScript Object Notation, lightweight data interchange format) file to improve the retrievability of the knowledge.
[0066] As an example, this Step 100 may include the following sub - steps:
[0067] Step 110: Construct a power data schema.
[0068] The description method of database table information is crucial for the large - language model to understand the database structure and content. Therefore, in this embodiment, the data definition languages (DDL) of all database tables related to power services are extracted from the power database, where each DDL contains the creation information of the database table. The required database table information is extracted by parsing the DDL information, and the database table information includes table name, table comment, column name, column type, etc.
[0069] In addition, for each database table, the data example content therein is collected. The data example content provides more accurate and intuitive database context information for the large - language model by showing the actual values of the fields. The first 2 example values of each field in each table are obtained by executing the following SQL statement:
[0070] SELECT DISTINCT col mn FROM t m WHERE col mn IS NOT NULL LIMIT 2 ;
[0071] where t m represents the table name of the m th table, and col mn represents the column name of the m th column of the n th table.
[0072] Reorganize the database table information and data examples extracted from the database to construct a data schema, which is more concise and structured than DDL statements, can better express the hierarchical relationship and field meaning of the database table, and is stored in JSON format for subsequent acquisition of table information. The specific form is as follows:
[0073] {
[0074] "table1":
[0075] "table1(col1([col1_ch_name], [type1], [data 11 , data 12 ), col2([col2_ch_name], [type2], [data 21 ,data 22 ), / / Other field information),
[0076] table1_ch_name, table1_comment",
[0077] … / / Structure information of other tables
[0078] }
[0079] Among them, table1 is the name of the table in the database, usually an English abbreviation; table1_ch_name represents the Chinese name of the table; table1_comment represents the comment of the table; 'col1' represents the name of the first field of the table in the database; col1_ch_name, type1, and [data 11 , data 12 represent the Chinese name, type, and two data examples of the field respectively.
[0080] Step 120: Construct business-schema mapping relationship knowledge.
[0081] As an example, this step 120 may include the following sub-steps:
[0082] Step 121: Sort out the business terms in the power field and establish a corresponding relationship between the business terms and the data schema, that is, the business-schema mapping relationship.
[0083] There are a large number of concise and refined proprietary business terms and industry languages in the field of electricity, and some terms are deeply associated with specific database tables and fields. Therefore, in this embodiment, the specific descriptions of common electricity business terms are sorted out, and the content involved is corresponded with the database structure, including relevant tables, fields, and the multi-table association relationships existing therein. Finally, it is organized into a specific JSON form. By establishing the correspondence between "business terms → data patterns" and providing relevant explanations, it is to assist large language models in understanding key business knowledge and the required database table knowledge.
[0084] The knowledge organization form of the business - mode mapping relationship is as follows:
[0085] {
[0086] "nt1": {
[0087] "description": "xxx",
[0088] "tables": ["t1", "t2", …],
[0089] "columns": ["c1", "c2", …],
[0090] "relations": ["t 1。 c1 = t 2。 c1", ......]
[0091] },
[0092] ... / / Other business term information
[0093] }
[0094] Among them, nt represents the name of the business term; the description contains the explanation of the business term; tables, columns, and relations respectively contain the table names, column names, and inter - table association relationships involved in this business term, stored in list form. If there is no definite corresponding database table content, it is represented as an empty list.
[0095] Step 122: Extract business term features, where the business term features include the vectorized business term name and the n - gram features of the business term name.
[0096] As an example, this step 122 may include the following sub - steps:
[0097] Step 1221: Perform vector embedding processing on each business term name to form a vectorized business term name.
[0098] Semantic similarity calculation is performed through embedding vectors, and similar business terms are closer in the vector space. In this embodiment, a pre-trained embedding model based on BERT (Bidirectional Encoder Representations from Transformers) is used to embed the text, which can capture the deep semantic meaning. The process is expressed as:
[0099] ;
[0100] Among them, represents the name of the i th business term; is the vector embedding result of this business term; Embedding represents the embedding model used. In this way, each mapping relationship can be represented as a multi-dimensional vector, which can be used for text similarity retrieval.
[0101] Step 1222: Extract the n - gram features of each business term name.
[0102] n - gram By splitting the text into consecutive n character combinations to capture the phrase structure in the text, which is used for keyword extraction to match similar content. In this embodiment, 2- gram is used for feature extraction. For example, for the text "arrears is not 0", its feature extraction is "arrears, fee not, not for, for 0", etc. The process is expressed as:
[0103] ;
[0104] Among them, is the 2-gram feature of this business term name.
[0105] Step 123: Store the business - mode mapping relationship knowledge, the vectorized business term name knowledge, and the 2-gram feature knowledge of the business term name.
[0106] All business - mode mapping relationship knowledge is stored in a JSON file.
[0107] All vectorized business term name knowledge is stored in a vector database. The vector database has specialized optimizations for high-dimensional vector storage structures to support the efficient storage and retrieval of large-scale vector data.
[0108] Store the 2-gram feature knowledge of all the extracted business term names in the Elasticsearch (ES) database. The ES database supports efficient full-text search functions and can quickly find relevant content based on the complete text or partial features.
[0109] Step 130: Construct data encoding relationship knowledge.
[0110] As an example, this step 130 may include the following sub-steps:
[0111] Step 131: Organization of data encoding relationship knowledge.
[0112] The design of the power database has highly standardized characteristics. Most business classification texts are usually stored in the database in encoded form. Therefore, there is a situation where the text description cannot correspond to the data structure.
[0113] Therefore, in this embodiment, the required information is extracted from the table storing the encoding correspondence relationship in the database, including:
[0114] (1) Encoding name: The specific business meaning of this encoding item;
[0115] (2) Encoding value: The value stored in the database for this encoding name;
[0116] (3) Corresponding column name: The database column name to which this encoding belongs;
[0117] Organize this information in a specific format to facilitate the effective utilization of knowledge in the future. The knowledge organization form of each encoding relationship is as follows:
[0118] Column: name---value;
[0119] Among them, Column represents the column name corresponding to the encoding, name represents the encoding item name, and value represents the encoding value. For example, in the "measurement method code (JLFSDM)" column, the encoding value corresponding to the encoding name "high supply and high metering" is "1", and its encoding relationship is organized as "JLFSDM: high supply and high metering---1;".
[0120] Step 132: Feature extraction and storage of encoding relationship knowledge.
[0121] For each encoding relationship, use the method described in step 122 to extract its vector features and 2-gram features respectively, and store the extracted features in the encoding relationship vector database and the encoding relationship ES database.
[0122] Step 200: Construct a problem parsing Prompt template and use a large language model to analyze the structure of the original problem to extract key entities from the original problem.
[0123] The semantic parsing of the original question is a key step in question enhancement, which provides a basis for understanding the user's query intention. Therefore, construct a question parsing Prompt and use a large language model to analyze the structure of the question. Through this process, identify the key elements of the question, including key entities, query conditions, and query objectives, to provide accurate information support for subsequent business knowledge linking and question enhancement. The specific steps are as follows:
[0124] Step 210: Design a question parsing Prompt template.
[0125] By designing a reasonable question parsing Prompt template, guide the large language model to analyze the question and effectively extract key information from the question. The question parsing Prompt template respectively instructs the large language model to extract three aspects of content, namely key entities, query conditions, and query objectives. Among them, the query objective is the specific information or data that the user hopes to obtain through the query, and the query condition is the content used to limit the query scope in the question. Extracting the query condition and query objective helps to disassemble the structure of the question and better identify the query intention; extracting key entities helps to retrieve database knowledge and business knowledge related to the entity.
[0126] Define the question parsing Prompt template as a tuple , including the following parts:
[0127] ;
[0128] Among them, G represents the defined task role objective, that is, define a task-related role for the large language model and clarify the task objective; represents important reminders, including important tips and constraints when completing the task; O represents the output format of the task; C represents the provided reference output case; Q represents the original question ({question}) input, and "{}" indicates that the parameter to be input is inside. The specific question parsing Prompt template is shown in Figure 8 .
[0129] Step 220: Use the large language model to analyze the structure of the original question and extract key elements from the original question. The key information includes key entities, query conditions, and query objectives.
[0130] Fill the user's original input question into the question parsing Prompt template constructed in Step 210, input it into the large language model and obtain the output result of the model. By parsing the JSON-formatted result output by the large language model, extract the key entity list En , query objective Qg and query condition listQc for subsequent use.
[0131] Step 300: Based on the sorted-out power domain knowledge and the key entities extracted from the original question, retrieve the power domain knowledge by using a method that fuses vector similarity and BM25 similarity. The power domain knowledge includes business - mode mapping relationship knowledge and data coding relationship knowledge.
[0132] Step 300: Retrieve the power domain knowledge by using a hybrid similarity retrieval method based on the sorted-out power domain knowledge and the key entities extracted from the original question.
[0133] To enhance the model's understanding ability of user questions and provide necessary information support for question enhancement, based on the key entities extracted from the original question, use two methods, vector similarity and BM25 similarity, to jointly retrieve domain - specific knowledge, so as to improve the coverage, relevance, and accuracy of knowledge retrieval. The retrieved content includes business - mode mapping relationship knowledge and data coding relationship knowledge, and associate the entities in the question with the corresponding knowledge. The specific steps are as follows:
[0134] Step 310: Extract key entity features, including vector features and 2 - gram features.
[0135] According to the list of key entities obtained in step 220 En= { e 1, e 2,...}, for each key entity e i , use the method described in step 122 to extract its vector features and 2 - gram features respectively, so that the key entity can be matched with the knowledge, expressed as:
[0136] ;
[0137] ;
[0138] where e i represents the i th key entity, is the vector feature of the key entity e i , is the 2 - gram feature of the key entity e i .
[0139] Step 320: Jointly retrieve the business - mode mapping relationship knowledge by using a method that fuses vector similarity and BM25 similarity.
[0140] To enable the large language model to have knowledge of relevant business terms during question enhancement and be able to obtain the database table connections related to the question business, retrieve the corresponding business term names from the vector database and ES library constructed in step 122 according to the entity name. This embodiment adopts a retrieval strategy of a hybrid method, and the retrieval strategy includes:
[0141] Step 3201: Calculate the cosine similarity between the entity and the business term to retrieve the business term names with similar semantics to the entity in the vector database.
[0142] The cosine similarity calculation formula is:
[0143] ;
[0144] where, represents the vector cosine similarity score between the entity and the business term, represents the norm of the vector.
[0145] Step 3202: Use the BM25 algorithm to retrieve the business term names with the same keyword fragments in the ES database.
[0146] The BM25 (Best Matching 25) algorithm calculates the occurrence frequency of keywords and their importance in the domain documents based on term frequency and inverse document frequency to evaluate their relevance.
[0147] The BM25 calculation formula is:
[0148] ;
[0149] ;
[0150] where, represents the BM25 similarity score; is the l th feature content in; N is the total number of knowledge in the ES library of business term names; is the number of documents containing the query term ; represents the frequency of occurrence in the feature T ei ; represents the feature length; L avg represents the average length of all features; k 1, k 3, b are respectively hyperparameters for adjusting the score size.
[0151] Step 3203: According to the vector cosine similarity score, obtain the most relevant business term name in the vector database for the key entity e i ; According to the BM25 similarity score, also obtain the most relevant business term name in the ES database; After merging the two business term names, use it as the business term name knowledge set of this entity.
[0152] To ensure a strong correlation between the business term and the entity, further filter out the results with a cosine similarity score less than 0.7; According to the BM25 similarity score, also obtain the most relevant business term name in the ES database, and after merging, use it as the business term name knowledge set of this entity, expressed as:
[0153] ;
[0154] where represents the first relevant business term name of the key entity e i , and if there is no relevant content, it represents an empty list.
[0155] For each retrieved business term name nt , obtain from the JSON file saved in Step 122 that has the business - schema mapping relationship nt the corresponding knowledge content, that is, the relevant business term explanation and the mapping relationship with the data schema K t , expressed as:
[0156] ;
[0157] where represents the first relevant business - schema mapping relationship knowledge of the key entity e i .
[0158] Step 330: Calculate the similarity between each key entity and the content in the coding relationship vector library and the coding relationship ES library, and obtain the coding relationship knowledge set corresponding to all key entities.
[0159] To enable the large - language model to map the problem entity to the relevant coding values and coding column names, use the cosine similarity and BM25 similarity calculation methods described in Step 320 to calculate the similarity between each key entity and the content in the coding relationship vector library and the ES library constructed in Step 132, and obtain the coding relationship knowledge set corresponding to all key entities K m , expressed as:
[0160] ;
[0161] Among them, represents the first relevant coding relationship knowledge of entity e i , and if there is no relevant knowledge, it represents an empty list.
[0162] Step 400: Obtain database tables and data patterns related to the original problem through a multi-level pattern linking method.
[0163] In order to correspond the user's original problem with specific tables and fields in the database, this embodiment adopts a multi-level data pattern linking method, which can efficiently, accurately and comprehensively obtain data patterns related to the user's original problem. The multi-level data pattern linking method, on the one hand, adopts a data pattern linking method based on double hash value search. Combining the key entity extraction results in the original problem, through the double search strategy of database tables and fields, it can quickly and accurately match the user's natural language problem with the most relevant table in the database. This method improves the accuracy of pattern linking and greatly reduces redundant content; on the other hand, based on the business-pattern mapping relationship knowledge obtained by retrieving the key entities of the original problem, obtain the association logic between relevant businesses and database tables in the user's problem, and provide more accurate positioning results in the way of business and database table association. The specific steps of step 400 are as follows:
[0164] Step 410: Table-level pattern linking;
[0165] LSH (Locality Sensitive Hashing) search is an efficient approximate nearest neighbor search technology suitable for big data, which can quickly complete similarity retrieval tasks. Based on this method, redundant information can be efficiently screened and excluded from a large number of database tables, and relevant data tables can be quickly locked.
[0166] Step 411: Table-level LSH construction.
[0167] As an example, this step 411 may include the following sub-steps:
[0168] Step 4111: Construct a table name dictionary;
[0169] Obtain the table name and Chinese name of each table in the database from the data definition language DDL and format them into a dictionary D T , and the specific format is as follows:
[0170] {"table name 1": "table Chinese name 1", / / Mapping of the remaining table names and their Chinese names};
[0171] Step 4112: Create a table-level hash index;
[0172] The dictionary D constructed according to step 4111 T Create an LSH index, and the specific formula is:
[0173] ;
[0174] where lsh t represents the D T initialized LSH data structure or object for subsequent similarity query operations of table names; MinHashes t used to store D T a set of the minimum hash signatures and their metadata of each data item so that the original data item can be traced and identified after similarity query; function create _ lsh is responsible for constructing the LSH index and generating a mapping set between signatures and metadata; the parameters involved sig, n_gram, thred respectively represent the signature length, n-gram the size of the segmentation, and the threshold of hash similarity matching. In the case of creating a table-level LSH, the three parameters are set to 80, 2, and 0.3 respectively.
[0175] Step 412: Retrieve similar table names.
[0176] Based on the list of key entities obtained in step 220 En= { e 1, e 2,...} and the table-level lsh t created in step 4112 e i Perform the following retrieval operations on each key entity
[0177] ;
[0178] where represents the matching function of the key entity e i in the lsh t model; M t represents the set of table names similar to the input entity; M t in the format of: {" e 1": ["Table Chinese Name 1",...], " e2”: [“Table Chinese Name 1”,...], , / / Other entity and table name information}。
[0179] Step 413: Table name screening.
[0180] To improve the matching accuracy between key entities and table names, it is necessary to further screen the set of similar table names retrieved in step 412 M t This is achieved by calculating the text similarity between each key entity and its similar table Chinese name in M t using SequenceMatcher. The specific calculation process is as follows:
[0181] ;
[0182] Among them, SimT represents the result of text similarity calculation; T j represents the list of table Chinese names obtained from M t corresponding to the key entity e i while t j is a single table Chinese name element in this list; SequenceMatcher is a text similarity calculation method in the difflib library, used to calculate the similarity between two strings and output a matching ratio value between 0 and 1. The larger the value, the higher the matching degree. ratio The method represents obtaining e i and t j similarity scores.
[0183] Step 414: Candidate table screening and pattern linking.
[0184] Based on the table name screening results in step 413 SimT , for each list of elements corresponding to a key entity, sort them in descending order according to the similarity results, and select the top top-k tables for each entity as the tables most relevant to the user's needs and retain them to form a set of relevant table names T s .
[0185] Based on the power data pattern constructed in step 110, obtain the table information corresponding to each table name in T c to complete the pattern linking at the problem entity and table level.
[0186] Step 420: Field-level pattern linking.
[0187] As an example, this step 420 may include the following sub-steps:
[0188] Step 421: Field-level LSH construction.
[0189] As an example, this step 421 may include the following sub-steps:
[0190] Step 4211: Screening of valuable fields in the library table;
[0191] For some fields in the database table, if they are not suitable for building an index, they are excluded to improve the efficiency of index construction. The specific screening method is as follows:
[0192] ;
[0193] Among them, T represents the set of library tables, F(T) represents the set of fields belonging to table T and Type(f) indicates the storage type of field f . F filtered represents the set of valuable fields after preliminary screening, and this set only contains fields with storage types of character types (VARCHAR and CHAR) and integer types (INT).
[0194] Step 4212: Screening out fields with specific names;
[0195] Perform secondary elimination on unnecessary or duplicate fields in the library table to enhance the performance of the index. The specific screening method is as follows:
[0196] ;
[0197] ;
[0198] Among them, Name(f) represents the name of field f . F final represents the set of fields after secondary screening. Based on the result of step 4211, fields with names ending with specific suffixes (such as "bh", "mc", "dz", "jc", "lj", "dh", "hm", "bs", "zh", "id") are further excluded. These fields mostly involve information such as numbers, names, addresses, abbreviations, paths, phones, numbers, identifiers, accounts, IDs, etc.
[0199] Step 4213: Construct a table-field dictionary;
[0200] Reformat the result of step 4212F final , construct a table-field dictionary D F , and the dictionary format is as follows:
[0201] {
[0202] "Table Name 1":
[0203] {"Field Name 1": "Chinese Name of Field 1", "Field Name 2": "Chinese Name of Field 2", / / Other field information},
[0204] "Table Name 2":
[0205] {"Field Name 1": "Chinese Name of Field 1", "Field Name 2": "Chinese Name of Field 2", / / Other field information},
[0206] / / Other table-field information
[0207] }
[0208] Step 4214: Create a field-level hash index;
[0209] Based on the constructed table name dictionary D F Create an LSH index:
[0210] ;
[0211] Among them, lsh f represents the D F initialized LSH data structure or object for subsequent similarity query operations of field names; MinHashes f used to store D F a set of the minimum hash signatures of each data item and its metadata so that the original data item can be traced and identified after a similarity query; the function create_lsh is responsible for constructing the LSH index and generating a mapping set between the signatures and the metadata. The parameters involved sig, n_gram, thred respectively represent the signature length, n-gram the size of the partition, and the threshold for hash similarity matching. In the case of creating a field-level LSH, the three parameters are set to 100, 2, and 0.05 respectively.
[0212] Step 422: Retrieve similar field names.
[0213] Based on the key entity list obtained in Step 220 En= { e 1,e 2, ...} and the field - level lsh f model constructed in step 421. For each key entity e i perform the following retrieval operations to find field names similar to the entity. The retrieval process can be expressed as:
[0214] ;
[0215] where represents the matching function of the key entity e i in the lsh f model, M f represents the set of field names similar to the input entity, M f in the format of:
[0216] {
[0217] "e1": {"table name 1": ["Chinese name of field 1", "Chinese name of field 2",...], / / Other table information},
[0218] "e2": {"table name 1": ["Chinese name of field 1", "Chinese name of field 2",...], / / Other table information},
[0219] / / Other entity information
[0220] }
[0221] Step 423: Similar table and field name filtering.
[0222] Obtain the result of step 422, merge the table names and field lists corresponding to each entity element in the original dictionary M f to form a new dictionary, where each table name corresponds to a list containing all non - duplicate fields. At the same time, obtain the corresponding field names from F final according to the field names of the tables involved in the new dictionary and format them into the new dictionary D h , in the specific format of:
[0223] {"table name 1": {"field name 1": "Chinese name of field 1", "field name 2": "Chinese name of field 2", / / Other field information}, / / Other table information};
[0224] Step 4231: Design of the Prompt template for library table filtering;
[0225] Design library table filtering Prompt template, combined with large language models, filters the tables and fields involved in the new dictionary according to the matching degree with the user's question to obtain the relevant tables that are most relevant to the question and have the most consistent semantics. Define the library table filtering Prompt template as D h and includes the following parts: P F which contains the following parts:
[0226] ;
[0227] Among them, G, I, O, Q has the same meaning as the question parsing Prompt template in step 210; KP represents the key thinking process prompt for filtering tables; TF represents the table and field information to be filtered. "{}" in the Prompt indicates that the parameter to be input is inside. See the specific library table filtering Prompt template in Figure 9 .
[0228] Step 4232: Library table information filtering;
[0229] Fill the user input question and the dictionary D h into the library table filtering Prompt template constructed in step 4231, input it into the large language model and obtain the output result of the model. By parsing the JSON-formatted result output by the model, extract the candidate table T h to which the similar fields belong for subsequent use.
[0230] Step 424: Candidate table schema linking.
[0231] Similarly, based on the power data schema constructed in step 110, obtain the table information corresponding to each table name in T h to complete the schema linking at the problem entity and field levels.
[0232] Step 430: Knowledge-based schema linking.
[0233] For the set of business term relationships Kt obtained in step 320, since each piece of knowledge kt has been organized according to the business-schema mapping relationship knowledge described in step 121, only need to parse the corresponding content in kt and obtain the required relevant table names. The process is expressed as:
[0234] ;
[0235] ;
[0236] Among them, tn ( kt i ) represents the table contained in the 'tables' primary key of knowledge kt i and is the set of table names contained in all knowledge in T n For Kt Retrieve the corresponding table information according to the set of table names as described in steps 414 and 424.
[0237] As described in steps 414 and 424, obtain the corresponding table information based on the set of table names T n Step 500: Based on the retrieved knowledge in the power field and the obtained data patterns, perform knowledge standardization, design a question enhancement Prompt template, and use a large language model to reconstruct and enhance the original question.
[0238] Based on the question parsing content, domain knowledge, and data patterns obtained in steps 200 to 400, perform knowledge standardization and question enhancement. By designing a knowledge formatting method, filter out irrelevant or redundant knowledge content; then, by designing a question enhancement Prompt template, use a large language model to reconstruct and enhance the original question, so that it can better understand the true meaning of the user's question and generate a more accurate SQL query based on this. The specific steps are as follows:
[0239] Step 510: Knowledge standardization: Filter out irrelevant or redundant knowledge content by designing a knowledge formatting method.
[0240] As an example, this step 510 may include the following sub-steps:
[0241] Step 511: Filter data encoding mapping relationships.
[0242] Since the encoding relationships retrieved in step 330 may contain content that is completely irrelevant to the data patterns obtained in step 400, it is necessary to filter the retrieval results of step 330 and only retain the information related to the actual database structure to avoid introducing unnecessary information.
[0243] Specifically, for each retrieval result of the data encoding mapping relationship, extract the corresponding encoded column names according to its organizational form, and remove the encoding relationships corresponding to the column names that are not in the relevant data patterns, which is expressed as:
[0244] Specifically, for each retrieval result of the data encoding mapping relationship, extract the corresponding encoded column names according to its organizational form, and remove the encoding relationships corresponding to the column names that are not in the relevant data patterns, which is expressed as:
[0245] ;
[0246] Among them, is the set of encoded relationship knowledge after filtering, represents The encoded column name, SC All data patterns obtained in Step 414, Step 424, and Step 430.
[0247] Step 512: Normalize the final representation forms of knowledge such as data patterns, data encoding mapping relationships, and business-pattern mapping relationships.
[0248] To make it easier for large language models to understand the provided knowledge, a standardized and clear format is designed for the final representation forms of various knowledge.
[0249] First, format the content of the data patterns obtained in Step 400, clearly distinguish the structure and content of each table, and clearly define the boundaries of different tables. The normalized table schema knowledge representation is designed as follows:
[0250] “- TABLE_i : xxx \n …
[0251] - TABLE_k : xxx \n … ”
[0252] Among them, TABLE_i represents the name of the i th table, and “ xxx ” is the specific content of the data pattern corresponding to this table.
[0253] Secondly, format the encoding mapping relationship knowledge form to ensure that each entity is in one-to-one correspondence with its related encoding relationship to avoid confusion. The normalized encoding relationship knowledge representation is designed as follows:
[0254] “# Entity name i: xxx \n Related encoding knowledge : xxx \n …
[0255] # Entity name k: xxx \n Related encoding knowledge : xxx \n … ”
[0256] Among them, the content after ‘Entity name i ’ is the th entity in i , and the content after ‘Related encoding knowledge’ represents the corresponding encoding knowledge part of this entity in ei . If there is no related encoding knowledge, it will not be formatted and output.
[0257] Finally, format the business-pattern mapping relationship knowledge, also making the key entities and knowledge in one-to-one correspondence. The normalized business-pattern mapping relationship knowledge representation is designed as follows:
[0258] i: xxx \n “# Entity name : xxx \n … Related business knowledge : xxx \n …
[0259] # Entity Name k: xxx \n Relevant business knowledge : xxx \n … ”
[0260] Among them, the content after 'Entity Name' i is the K t th key entity in i ; the content after 'Relevant business knowledge' represents the corresponding business term knowledge part of this entity in e i ; if there is no relevant knowledge, no formatting output will be performed. K t
[0261] Step 520: Design a question-enhanced Prompt template.
[0262] By designing a question-enhanced Prompt template, it is possible to guide the large language model to reconstruct and optimize the original question. Since there may still be redundant information in the knowledge and data patterns obtained in steps 300 and 400, the knowledge content obtained needs to be fully expressed in the question-enhanced Prompt template, and it should have a clear structure and necessary reminder content to ensure that the large language model can more accurately match the question with the relevant knowledge and data patterns. Define the question-enhanced Prompt template as P G , which includes the following parts:
[0263] ;
[0264] Among them, G, O, Q has the same meaning as the question parsing Prompt template in step 210; SK is the data pattern content formatted through step 512; EK is the formatted encoding relationship knowledge content; T K is the formatted business-pattern mapping relationship knowledge content; Q A is the question parsing content obtained in step 220; W represents important reminders, including important tips and constraints when completing the task. The specific content of the question-enhanced Prompt template is shown in Figure 10 .
[0265] Step 530: Use the large language model to reconstruct and enhance the original question, optimize its expression and match the database structure to better meet the query requirements of power domain data.
[0266] Fill the required knowledge content into the question enhancement Prompt template to guide the large language model to eliminate redundant information, understand the question structure, and optimize the expression of the question, so that the enhanced question can be associated with the correct business knowledge and data schema, and more accurately express the query requirements.
[0267] Interact with the large language model based on the data schema, domain knowledge, and the enhanced question to generate SQL queries, reducing the ambiguity or errors when generating SQL using the original question, so that users can obtain more accurate results in data queries in the power domain.
[0268] In this embodiment, a data query question related to the power domain is selected, and the above method is used to enhance the question and generate SQL using the enhanced question. The experimental results are as follows:
[0269] Question: Query the information of electricity users who are meter-read monthly but have no meter-reading records in April 2022 under a certain power supply station.
[0270] For the construction example content of the power data schema and related knowledge, see Figure 2 .
[0271] The result of question parsing in step 200 is shown in Figure 3 , it can be seen that under the guidance of the question parsing Prompt, the large language model successfully extracts the key entities, query targets, and query conditions of the original question respectively.
[0272] The encoded relationship knowledge and business-schema mapping relationship knowledge obtained through entity retrieval and association in step 300 are shown in Figure 4 , through entity retrieval in the question, the encoded relationships related to entities such as 'power supply station' are obtained, as well as the explanations related to the 'monthly meter reading' business and the associated table and field information.
[0273] Figure 5 The above shows the search results of the similar table hash values and similar field hash values in steps 412 and 422 of the present invention. It can be seen that based on the key entities, the table names related to the question and the tables containing the relevant table fields are obtained respectively through dual searches. Figure 5 The following is the result of filtering in step 430. The large language model is used to filter the retrieval results and screen out the two most relevant tables ("HS_KFYHXX" and "KH_YDKH").
[0274] Figure 6After standardizing the above knowledge in step 500 and filling it into the question-enhanced Prompt, the question-enhanced result output by the large language model. It can be seen that through the injection of knowledge and Prompt design, the large language model carefully analyzes the original question step by step and matches it with the relevant content in the database, and finally outputs a question statement closely related to the database content. For example, the condition for "xx power supply station" is that the value of the "GDDWBM" field is "030000", and "power user information" corresponds to the detailed information of "power consumption customers (`KH_YDKH` table)".
[0275] Figure 7 It is a comparison of the SQL results generated by the present invention and other methods. Among them, Figure 7 The result of a is the SQL generated by using the domain knowledge sorted out by the present invention and the knowledge obtained in step 250, and using the enhanced question to input the large language model; Figure 7 b in it is the SQL generated by only using database knowledge and the original question to input the large language model; Figure 7 The result of c in it is the SQL generated by using the domain knowledge sorted out by the present invention and the knowledge obtained in step 250, but combined with the original question to input the large language model. It can be seen that Figure 7 The SQL result of b in it lacks relevant domain knowledge and a method for locating library tables, resulting in incorrect use of database tables and fields in the SQL, and unable to convert "xx power supply station" and "monthly meter reading" into correct business language. Figure 7 Since c contains the required domain knowledge, it can be converted into business language in the SQL, such as "GDDWBM = '030000'", "CBZQ = '1'", but the simple expression of the user's question results in the generated SQL not conforming to the actual logic and unable to query the required data. The method result based on the present invention can correctly understand the user's query intention, generate SQL that meets the requirements and has correct logic.
[0276] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A multi-knowledge link driven NL2SQL question enhancement method, characterized in that: The following steps are involved: Step 100: constructing a power data model using database table information, and sorting out power domain knowledge, wherein the power domain knowledge includes business-model mapping relationship knowledge and data encoding relationship knowledge; Step 200: construct a question parsing prompt template, and use a large language model to analyze the structure of the original question and extract key entities from the original question; Step 300: Based on the sorted out power field knowledge and the key entities extracted from the original question, the business-mode mapping relationship knowledge is jointly retrieved by using the method of fusion of vector similarity and BM25 similarity: the cosine similarity between the entity and the business term is calculated to retrieve the business term name with similar semantics to the entity in the vector database; the business term name with the same keyword fragment is retrieved in the ES database using BM25 similarity; according to the vector cosine similarity score, the most relevant business term name is obtained in the vector database; according to the BM25 similarity score, the most relevant business term name is also obtained in the ES database; the two business term names are merged as the business term name knowledge set of the entity; The cosine similarity and BM25 similarity methods are used to calculate the similarity of the content of each key entity in the encoding relationship vector library and the encoding relationship ES library, and obtain the encoding relationship knowledge set corresponding to all key entities; Step 400: Obtain the database table and data schema related to the original question through a multi-level schema linking method; Based on the fast linking method of double hash value search, the extracted original question entities are linked to the tables and fields in the database respectively to obtain the database tables and data models related to the original question; a library table filtering prompt template is designed to filter to obtain the relevant tables that are most relevant and semantically consistent with the original question; Establish a link with the database table, including: Table-level LSH construction, including: building a table name dictionary; obtaining the table name and Chinese name of each table in the database from the data definition language DDL, and formatting it into a dictionary D T ; Create a table-level hash index based on the constructed dictionary D T Create a table-level LSH index; Similar table name retrieval, including: Obtained key entity list En= { e 1, e 2, ...} and the table-level LSH index created; for each key entity e i Perform the following retrieval operation to find out the table names that are similar to the entity. The retrieval process is expressed as: ; in, Representing key entities e i In lsh t Matching functions in the model; M t Represents a set of table names similar to the input entity; Table name filtering, including: a collection of similar table names retrieved M t Perform further screening; calculate through SequenceMatcher M t This is achieved by measuring the text similarity between each key entity and its Chinese name in the similarity table; Candidate table screening and pattern linking include: based on the text similarity calculation results, sorting the element list corresponding to each key entity in descending order according to the similarity results, and selecting the top element in each entity. top-k The tables are reserved as the most relevant tables to user needs, forming a set of related table names; The library table filtering prompt template is defined as P F , contains the following parts: ; in, G Represents the defined mission role objectives; Indicates important reminders, including important tips and constraints when completing tasks; O Indicates the output format of the task; Q Represents the original input question; KP Represents the key thinking process prompts for filtering tables; TF Indicates the table and field information to be filtered; Step 500: Based on the retrieved power field knowledge and the acquired data model, knowledge standardization is performed, and a question enhancement prompt template is designed, and the original question is reconstructed and enhanced using a large language model.
2. The NL2SQL question enhancement method driven by multiple knowledge links according to claim 1, characterized in that: In step 100, constructing the power data model includes: The data definition language DDL of all power business-related database tables is extracted from the power database, where each DDL contains the creation information of the database table. The required database table information is extracted by parsing the DDL; the data example content is collected for each database table; the database table information and data examples extracted from the database are reorganized to build a data model.
3. The NL2SQL question enhancement method driven by multiple knowledge links according to claim 2 is characterized in that: In step 100, the business-mode mapping relationship knowledge includes: Organize business terms in the power field and form a correspondence between the business terms and data models, that is, a business-model mapping relationship; Extract business term features, including vectorized business term names and business term names n - gram feature; Store all business-mode mapping relationship knowledge in JSON files; store all vectorized business term name knowledge in vector databases; n - gram Feature knowledge is stored in the ES database.
4. The NL2SQL question enhancement method driven by multiple knowledge links according to claim 3 is characterized in that: In step 100, the data encoding relationship knowledge includes: Extract the required information from the table storing the coding correspondence in the database, including the coding name, coding value and corresponding column name; For each encoding relationship, extract its vector features and n - gram Features, and store the extracted features in the encoding relationship vector database and the encoding relationship ES database.
5. The NL2SQL question enhancement method driven by multiple knowledge links according to claim 1, characterized in that: In step 200, a question resolution prompt template is constructed, and the question resolution prompt template is defined as a tuple. ,include: ; in, G Indicates the defined task role goal, that is, to give the large language model a task-related role and clarify the task goal; I Indicates important reminders, including important tips and constraints when completing tasks; O Indicates the output format of the task; C Indicates the reference output case provided; Q Represents the original question of the input.
6. The NL2SQL question enhancement method driven by multiple knowledge links according to claim 1, characterized in that: In step 500, performing knowledge standardization includes: Based on the coding relationship knowledge retrieval results, the corresponding coding column names are extracted according to their organizational form, and the coding relationships corresponding to the column names that are not in the relevant data model are removed; and the final representation of the data model, data coding mapping relationship and business-model mapping relationship is standardized and designed.
7. The NL2SQL question enhancement method driven by multiple knowledge links according to claim 1, characterized in that: In step 500, the question enhancement prompt template is defined as P G ,include: ; in, G Indicates the defined task role goal, that is, to give the large language model a task-related role and clarify the task goal; SK The formatted data mode content; EK For formatted coded relational knowledge content; T K It is the formatted business-pattern mapping relationship knowledge content; W represents important reminders, including important tips and constraints when completing tasks; O Indicates the output format of the task; Q Represents the original input question; Q A Parse the content for the obtained question.
Citation Information
Patent Citations
Text2SQL semantic parsing method for domain large language model
CN118377796A
Power field SQL intelligent agent construction method based on KMDI chain
CN119166662A
Cited By
NL2SQL method and system based on large language model and retrieval enhancement
CN121326959A
NL2SQL method and system based on large language model and retrieval enhancement
CN121326959B
Electric power NL2SQL exploration optimization method based on Monte Carlo tree search
CN121560933A