Model training method and device, electronic equipment and storage medium

By training the table selection model and the SQL statement generation model, combining SQL syntax rules and natural statement generation models, the accuracy problem of user query related tables in enterprise-level data warehouses is solved, and friendly interaction and efficient query of non-technical personnel are achieved.

CN120256947APending Publication Date: 2025-07-04MINSHENG BANKING CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510171034.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing NL2SQL technology is difficult to accurately find the tables related to user query in enterprise-level data warehouses. Users need professional skills to write SQL statements. The threshold is high, and the existing technology has high labeling costs, making it difficult to achieve universality.

Method used

By training the table selection model and SQL statement generation model, combining SQL syntax rules, statement alignment templates and natural statement generation models, constructing aligned natural statements and SQL statements, filtering related tables and generating correct SQL query statements.

Benefits of technology

Improve the generalization ability of the model, accurately understand user query intentions, reduce ambiguity and diversity, and non-technical personnel can also easily interact with the database, simplifying the query process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256947A_ABST
    Figure CN120256947A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a model training method and device, electronic equipment and a storage medium. The method comprises the following steps: processing business data in a preset database table based on an SQL grammar rule to obtain an SQL statement; processing the SQL statements according to a preset data generation mode, and constructing aligned natural statements and SQL statements; according to the preset database table, the natural statements and the SQL statements, determining the aligned database table, natural statements and SQL statements; based on the aligned natural statement and database table, training to obtain a table selection model; and based on the aligned natural statement, database table and SQL statement, training to obtain an SQL statement generation model. According to the embodiment of the invention, the generalization ability of the model can be improved, the accuracy of user intention understanding is improved, and the influence of ambiguity and diversity is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular, to a model training method, apparatus, electronic device, and storage medium. Background Art

[0002] In daily work, Excel tables are everywhere. In an APP (Application) or a web page, tables are a clear and friendly way to convey information. In enterprises, relational databases are ubiquitous. Due to the clear data structure of tables, easy maintenance, and friendliness to both human understanding and machine understanding, tables / relational databases are the most common structured knowledge storage forms in all walks of life.

[0003] However, in the query interaction of table knowledge, the threshold is not low: dialogue systems or search engines cannot well query table knowledge as answers, and the query of relational databases requires professional technical personnel to write query statements (such as SQL (Structured Query Language) statements) to complete, which is a higher threshold for most users. In related academic research, a series of works and breakthroughs have been made in the NL2SQL (Natural Language to SQL) technology. However, existing NL2SQL technologies are often based on a specific table or a small range of databases. In the actual use of enterprises, a large number of table data are usually laid out in a data warehouse, and users do not know in advance which tables need to be queried, bringing new challenges to the popularization and application of related technologies. Summary of the Invention

[0004] The technical problem to be solved by the embodiments of this application is to provide a model training method, apparatus, electronic device, and storage medium, so that by training a table selection model, relevant tables related to user queries can be found more accurately, thereby improving the generalization ability of the model. Through the collaborative work of combining large and small models, this solution can more accurately understand the query intent of users. The small model is responsible for quickly screening relevant tables, and the large model is responsible for deeply understanding user questions and generating correct SQL query statements, improving the accuracy of user intent understanding and reducing the impact of ambiguity and diversity.

[0005] In a first aspect, an embodiment of this application provides a model training method, and the method includes:

[0006] Process the business data in a preset database table based on SQL syntax rules to obtain SQL statements;

[0007] Process the SQL statements according to a preset data generation method to construct aligned natural statements and SQL statements;

[0008] Determine the aligned database table, natural statement, and SQL statement according to the preset database table, the natural statement, and the SQL statement;

[0009] Train a table selection model based on the aligned natural statement and the database table;

[0010] Train an SQL statement generation model based on the aligned natural statement, the database table, and the SQL statement.

[0011] Optionally, the processing of the SQL statement according to the preset data generation method to construct the aligned natural statement and SQL statement includes:

[0012] Fill the SQL statement into a pre-configured statement alignment template;

[0013] Determine the aligned natural statement and the SQL statement based on the statement alignment template.

[0014] Optionally, the processing of the SQL statement according to the preset data generation method to construct the aligned natural statement and SQL statement includes:

[0015] Input the SQL statement into a pre-trained natural statement generation model;

[0016] Obtain the natural statement corresponding to the SQL statement output by the natural statement generation model;

[0017] Determine the aligned natural statement and the SQL statement.

[0018] Optionally, the processing of the SQL statement according to the preset data generation method to construct the aligned natural statement and SQL statement includes:

[0019] Obtain the table description information of the data table, and select specified row data from the data table as example data;

[0020] Input the table description information and the example data into a pre-trained statement alignment model;

[0021] Obtain the aligned natural statement and the SQL statement output by the statement alignment model.

[0022] Optionally, the training of the table selection model based on the aligned natural statement and the database table includes:

[0023] Obtain the data table information of the database table;

[0024] Based on the natural statement and the data table information, construct a model training data set, where the model training data set includes: positive sample data and negative sample data. Among them, the positive sample data is the aligned natural statement and the data table information of the database table, and the negative sample data is the unaligned natural statement and the data table information of the database table;

[0025] Train the table selection model based on the model training data set.

[0026] Optionally, after training the SQL statement generation model based on the aligned natural statement, the database table, and the SQL statement, it further includes:

[0027] Obtain the question statement input by the user;

[0028] Calculate the similarity between the question statement and the column names of each database table in multiple database tables;

[0029] Exclude the data columns of each database table in the multiple database tables whose similarity is less than the similarity threshold;

[0030] Call the table selection model to process the question statement and the data table information of the multiple database tables after excluding the data columns, and obtain the target database table in the multiple database tables. Among them, the data table information after excluding the data columns is the information obtained by splicing the database table names of the multiple database tables and the column names of the remaining data columns after excluding the data columns;

[0031] Call the SQL statement generation model to process the question statement and the data table information of the target database table, and generate the target SQL statement corresponding to the question statement.

[0032] Optionally, the calculating the similarity between the question statement and the column names of each database table in multiple database tables includes:

[0033] Obtain the first vector corresponding to the question statement, and the second vector corresponding to the column name of each database table in the multiple database tables;

[0034] Calculate the similarity between the first vector and the second vector.

[0035] In a second aspect, an embodiment of the present application provides a model training device, and the device includes:

[0036] An SQL statement acquisition module, configured to process the business data in a preset database table based on SQL syntax rules to obtain an SQL statement;

[0037] An alignment statement construction module for processing the SQL statement according to a preset data generation method to construct an aligned natural statement and SQL statement;

[0038] An alignment statement determination module for determining the aligned database table, natural statement, and SQL statement according to the preset database table, the natural statement, and the SQL statement;

[0039] A table selection model training module for training a table selection model based on the aligned natural statement and the database table;

[0040] An SQL statement generation model training module for training an SQL statement generation model based on the aligned natural statement, the database table, and the SQL statement.

[0041] Optionally, the alignment statement construction module includes:

[0042] An SQL statement filling unit for filling the SQL statement into a pre-configured statement alignment template;

[0043] A first alignment statement determination unit for determining the aligned natural statement and the SQL statement based on the statement alignment template.

[0044] Optionally, the alignment statement construction module includes:

[0045] An SQL statement input unit for inputting the SQL statement into a pre-trained natural statement generation model;

[0046] A natural statement acquisition unit for acquiring the natural statement corresponding to the SQL statement output by the natural statement generation model;

[0047] A first alignment statement determination unit for determining the aligned natural statement and the SQL statement.

[0048] Optionally, the alignment statement construction module includes:

[0049] An example data selection unit for acquiring the table description information of the data table and selecting the specified row data from the data table as example data;

[0050] An example data input unit for inputting the table description information and the example data into a pre-trained statement alignment model;

[0051] An alignment statement acquisition unit for acquiring the aligned natural statement and the SQL statement output by the statement alignment model.

[0052] Optionally, the table selection model training module includes:

[0053] A data table information acquisition unit for acquiring the data table information of the database table;

[0054] A data set construction unit for constructing a model training data set based on the natural statement and the data table information, where the model training data set includes: positive sample data and negative sample data, and the positive sample data is the aligned natural statement and the data table information of the database table, and the negative sample data is the non-aligned natural statement and the data table information of the database table;

[0055] A table selection model training unit for training the table selection model based on the model training data set.

[0056] Optionally, the included device further includes:

[0057] A problem statement acquisition module for acquiring a problem statement input by a user;

[0058] A similarity calculation module for calculating the similarity between the problem statement and the column names of each database table in multiple database tables;

[0059] A data column elimination module for eliminating the data columns in each database table in the multiple database tables whose similarity is less than the similarity threshold;

[0060] A target database table acquisition module for calling the table selection model to process the problem statement and the data table information of the multiple database tables after eliminating the data columns, and obtaining the target database table in the multiple database tables, where the data table information after eliminating the data columns is the information obtained by splicing the database table names of the multiple database tables and the column names of the remaining data columns after eliminating the data columns;

[0061] A target SQL statement generation module for calling the SQL statement generation model to process the problem statement and the data table information of the target database table, and generating a target SQL statement corresponding to the problem statement.

[0062] Optionally, the similarity calculation module includes:

[0063] A vector acquisition unit for acquiring a first vector corresponding to the problem statement and a second vector corresponding to the column name of each database table in the multiple database tables;

[0064] A similarity calculation unit for calculating the similarity between the first vector and the second vector.

[0065] In a third aspect, an embodiment of the present application provides an electronic device, including:

[0066] A processor, a memory, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, the model training method described in any one of the above is implemented.

[0067] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium. When instructions in the storage medium are executed by a processor of an electronic device, the electronic device can execute the model training method described in any one of the above.

[0068] Compared with the prior art, the embodiments of the present application include the following advantages:

[0069] In the embodiments of the present application, business data in a preset database table is processed based on SQL syntax rules to obtain SQL statements. The SQL statements are processed according to a preset data generation method to construct aligned natural statements and SQL statements. Based on the preset database table, natural statements, and SQL statements, the aligned database table, natural statements, and SQL statements are determined. A table selection model is trained based on the aligned natural statements and database table. An SQL statement generation model is trained based on the aligned natural statements, database table, and SQL statements. By training the table selection model in the embodiments of the present application, the table related to the user query can be found more accurately, thereby improving the generalization ability of the model. By combining the collaborative work of large and small models, this solution can understand the user's query intention more accurately. The small model is responsible for quickly screening relevant tables, and the large model is responsible for deeply understanding the user's question and generating the correct SQL query statement, improving the accuracy of user intention understanding and reducing the influence of ambiguity and diversity.

[0070] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. Description of the Drawings

[0071] Figure 1 It is a flowchart of the steps of a model training method provided by an embodiment of the present application;

[0072] Figure 2 It is a schematic diagram of sampling to generate natural statements and SQL statements provided by an embodiment of the present application;

[0073] Figure 3 It is a schematic diagram of a model training process provided by an embodiment of the present application;

[0074] Figure 4 It is a schematic diagram of an SQL statement generation process provided by an embodiment of the present application;

[0075] Figure 5 It is a schematic diagram of a table selection model processing process provided by an embodiment of the present application;

[0076] Figure 6 A structural schematic diagram of a model training device provided by an embodiment of the present application;

[0077] Figure 7 A structural schematic diagram of an electronic device provided by an embodiment of the present application. Specific implementation manners

[0078] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below with reference to the accompanying drawings and specific implementation manners.

[0079] The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. The singular forms "a", "said", and "the" used in the embodiments of the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.

[0080] Referring to Figure 1 , a step flowchart of a model training method provided by an embodiment of the present application is shown. As Figure 1 shown, the model training method may include: step 101, step 102, step 103, step 104, and step 105.

[0081] Step 101: Process the business data in a preset database table based on SQL syntax rules to obtain an SQL statement.

[0082] In this embodiment, a database table is the basic structural unit for storing and organizing data in a database. It organizes data according to a specific data model (such as a relational model, a hierarchical model, or a network model), and the most common one is the relational model, that is, the table in a relational database.

[0083] In a relational database, a table consists of rows and columns. Rows: Each row represents a record, usually corresponding to an entity or event in the real world. Columns: Each column represents a field, usually corresponding to an attribute of an entity or event.

[0084] SQL (Structured Query Language) is the standard programming language for managing and operating relational databases. SQL statements are used to perform various operations on the database, such as querying, updating, inserting, and deleting data.

[0085] A preset database table refers to a database table that is pre-selected for SQL statement sampling.

[0086] SQL syntax rules refer to a series of guiding principles that need to be followed when writing valid SQL statements. These rules ensure the correctness and understandability of SQL statements, enabling the database management system (DBMS) to accurately parse and execute these statements.

[0087] When training the table selection model and the SQL statement generation model, preset database tables can be filtered out first, and then the business data in the preset database tables can be processed based on SQL syntax rules to obtain SQL statements. Specifically, the business requirements can be clarified first, that is, which data to retrieve, update, insert, or delete from the database, and the business requirements will guide the writing of corresponding SQL statements. Secondly, according to the business requirements, a suitable SQL operation type can be selected. Common operation types include: SELECT (used to retrieve data from the database), INSERT (used to insert new data into the database), UPDATE (used to update existing data in the database), DELETE (used to delete data from the database). Finally, according to the selected SQL operation type, SQL statements can be written according to SQL syntax rules. This includes specifying table names, column names, conditions (if necessary), and the data values to be operated on.

[0088] After processing the business data in the preset database tables based on SQL syntax rules to obtain SQL statements, step 102 is executed.

[0089] Step 102: Process the SQL statement according to the preset data generation method to construct an aligned natural statement and SQL statement.

[0090] The preset data generation method refers to the method used to generate an aligned natural statement and SQL statement. In this embodiment, "alignment" refers to result alignment. For example, for a query operation, the expected result in the natural language description should correspond to the data set returned after the SQL statement is executed, which means that the SQL statement should be able to accurately return the data required in the natural language description.

[0091] After processing the business data in the preset database tables based on SQL syntax rules to obtain SQL statements, the SQL statement can be processed according to the preset data generation method to construct an aligned natural statement and SQL statement.

[0092] In this embodiment, an aligned natural statement and SQL statement can be constructed through a statement alignment template, and the implementation process can be described in detail in combination with the following specific implementation methods.

[0093] In a specific implementation manner of this application, the above step 102 may include:

[0094] Sub-step A1: Fill the SQL statement into a pre-configured statement alignment template.

[0095] In this embodiment, after obtaining the SQL statement, the SQL statement can be filled into a pre-configured statement alignment template.

[0096] After filling the SQL statement into the pre-configured statement alignment template, execute sub-step A2.

[0097] Sub-step A2: Based on the statement alignment template, determine the aligned natural statement and the SQL statement.

[0098] After filling the SQL statement into the pre-configured statement alignment template, the aligned natural statement and the SQL statement can be determined based on the statement alignment template. Specifically, the statement alignment template is a structured document or data structure for storing and presenting the correspondence between natural statements and SQL statements. The template may include: a natural statement column (for inputting or displaying natural language sentences describing business requirements or query intents), an SQL statement column (for inputting or displaying SQL query statements corresponding to the natural statements), and alignment markers (the template may contain some markers or fields for indicating which elements in the SQL statement correspond to specific parts in the natural statement (e.g., through indexes, color coding, or other identifiers, etc.)). After obtaining the SQL statement, the SQL statement can be copied into the SQL statement column of the template. Finally, the aligned natural statement can be determined, that is, based on the business requirements, write a clear and concise natural language sentence to describe the business requirements, ensuring that the description in the natural statement matches the operations, filtering conditions, and result sets of the SQL statement. This may require repeatedly modifying the natural statement until alignment.

[0099] In this embodiment, the aligned natural statement and the SQL statement can be constructed through a natural statement generation model, and the implementation process for this can be described in detail in combination with the following specific implementation methods.

[0100] In another specific implementation manner of this application, the above step 102 may include:

[0101] Sub-step B1: Input the SQL statement into a pre-trained natural statement generation model.

[0102] In this embodiment, the natural statement generation model is capable of understanding the structure and meaning of the SQL statement and converting it into a natural language sentence accurately describing the same query intent. The training of the model usually involves a large amount of paired data of SQL statements and natural language descriptions, and these data are used to optimize the model's conversion ability from SQL to natural language. For the training process of the natural statement generation model, this embodiment does not describe it in detail here.

[0103] After obtaining the SQL statement, the SQL statement can be input into a pre-trained natural language generation model.

[0104] After inputting the SQL statement into the pre-trained natural language generation model, sub-step B2 is executed.

[0105] Sub-step B2: Obtain the natural language statement corresponding to the SQL statement output by the natural language generation model.

[0106] After inputting the SQL statement into the pre-trained natural language generation model, the natural language statement corresponding to the SQL statement output by the natural language generation model can be obtained.

[0107] Sub-step B3: Determine the aligned natural language statement and the SQL statement.

[0108] According to the natural language statement corresponding to the SQL statement output by the natural language generation model, the aligned natural language statement and the SQL statement can be determined. For example, the SQL statements input by the user to the natural language generation model include: SQL statement A, SQL statement B, and SQL statement C, and the natural language statements corresponding to SQL statement A, SQL statement B, and SQL statement C output by the natural language generation model are respectively: natural language statement 1, natural language statement 2, and natural language statement 3. Then, the aligned statements can be determined: that is, the aligned SQL statement A and natural language statement 1, the aligned SQL statement B and natural language statement 2, the aligned SQL statement and natural language statement 3, etc.

[0109] It can be understood that the above example is only an example listed for better understanding of the technical solution of the embodiments of the present application, and does not serve as the only limitation of this embodiment.

[0110] The embodiments of the present application adopt a sampling generation method (i.e., the method of template generation and natural language generation model generation) to construct the alignment data of natural language statements and SQL statements, which can be as Figure 2 shown. The sampling generation method first samples SQL statements from the tabular data based on the SQL syntax rules, and then obtains the corresponding natural language questions based on the template and the large model generation method. The problem in the prior art is that: the NL2SQL data obtained through scenarios has a large difference from the in-line scenarios, and it is difficult to achieve generality, while the NL2SQL data annotation cost is relatively high, time-consuming and laborious. Therefore, the sampling generation method is adopted to construct the alignment data of natural language statements and SQL statements. As Figure 2As shown, first, SQL statements can be sampled from Tables (data tables) based on SQL syntax rules, namely, SELECT product name, MIN (starting amount) WHERE (yield>3.1%) AND income type = principal guaranteed. Template generation can generate sales talk (i.e., natural statements): What are the products with a yield greater than 3.1 and principal guaranteed, and what is the minimum starting amount. Model generation can generate sales talk: What is the minimum starting amount for an ideal product with an interest rate greater than 3.1% and principal guaranteed.

[0111] It can be understood that the above examples are merely examples listed for a better understanding of the technical solutions of the embodiments of the present application, and are not intended to be the sole limitation to the embodiments.

[0112] In this embodiment, a few-sample prompting method can also be used to construct the alignment data of natural language and SQL. The few-sample prompting method refers to inputting the description of the table information and sample data into the large model, and the large model directly generates<NL,SQL> The implementation process can be described in detail in conjunction with the following specific implementation method.

[0113] In another specific implementation of the present application, the above step 102 may include:

[0114] Sub-step C1: Acquire table description information of a data table, and select a specified row of data from the data table as sample data.

[0115] In this embodiment, the table description information of the data table can be obtained first, and the specified row data can be selected from the data table as sample data. Specifically, first, the metadata or description information of the data table needs to be obtained. This usually includes the column name, data type, constraints (such as primary key, foreign key, non-empty constraint, etc.) and a brief description of the column of the table. This information helps to understand the structure and content of the table and is the basis for generating natural statements and SQL statements. Then, one or more rows of data can be selected from the data table as sample data. These sample data should represent the typical characteristics of the data in the table. The sample data is used to show how to reference the data in the table in a natural statement and help the statement alignment model understand the actual content of the data table.

[0116] After obtaining the table description information of the data table and selecting a specified row of data from the data table as sample data, sub-step C2 is executed.

[0117] Sub-step C2: Input the table description information and the example data into a pre-trained sentence alignment model.

[0118] After obtaining the table description information of the data table and selecting the specified row data from the data table as example data, the table description information and the example data can be input into a pre-trained statement alignment model. Specifically, the table description information and the example data can be organized into a format acceptable to the model, such as converting the table description information into a structured JSON or XML format, and converting the example data into a format easy for the model to process, etc. The organized input data is input into the pre-trained statement alignment model. This model has learned how to align natural statements with SQL statements, that is, to understand the intention in the natural statement and convert it into an equivalent SQL query.

[0119] Sub-step C3: Obtain the aligned natural statement and the SQL statement output by the statement alignment model.

[0120] Based on the input table description information and example data, the statement alignment model generates an aligned natural statement and SQL statement. For example, the natural statement is a descriptive query such as "Find all employees whose age is greater than 30 years old", and the corresponding SQL statement may be SELECT * FROM employees WHERE age > 30, etc.

[0121] In the embodiments of the present application, the alignment data of natural language and SQL is constructed by means of sampling generation and few-shot generation, which can solve the problem that the cost will be extremely high if manual annotation is adopted, and reduce the cost of constructing alignment data.

[0122] After processing the SQL statement according to the preset data generation method to construct the aligned natural statement and SQL statement, step 103 is executed.

[0123] Step 103: Determine the aligned database table, natural statement and SQL statement according to the preset database table, the natural statement and the SQL statement.

[0124] After processing the SQL statement according to the preset data generation method to construct the aligned natural statement and SQL statement, the aligned database table, natural statement and SQL statement can be determined according to the preset database table, the aligned natural statement and the SQL statement. That is, each SQL statement corresponds to a database table. After obtaining the aligned natural statement and SQL statement, the aligned database table, natural statement and SQL statement can be obtained.

[0125] After obtaining the aligned database table, natural statement and SQL statement, the aligned database table, natural statement and SQL statement can be manually reviewed. After the manual review is successful, it can be put into use.

[0126] After determining the aligned database table, natural statement, and SQL statement according to the preset database table, natural statement, and SQL statement, steps 104 and 105 are executed.

[0127] Step 104: Based on the aligned natural statement and the database table, a table selection model is trained.

[0128] After determining the aligned database table, natural statement, and SQL statement according to the preset database table, natural statement, and SQL statement, a table selection model can be trained based on the aligned natural statement and database table. The training process of the table selection model can be described in detail in combination with the following specific implementation methods.

[0129] In a specific implementation of the present application, the above step 104 may include:

[0130] Sub-step D1: Obtain the data table information of the database table.

[0131] In this embodiment, after obtaining the aligned database table, natural statement, and SQL statement, the data table information of the database table can be obtained. Among them, the data table information may include the name of the table, the name and data type of the columns, the primary key, and the foreign key, etc.

[0132] After obtaining the data table information of the database table, sub-step D2 is executed.

[0133] Sub-step D2: Based on the natural statement and the data table information, a model training data set is constructed, and the model training data set includes: positive sample data and negative sample data, where the positive sample data is the aligned natural statement and the data table information of the database table, and the negative sample data is the non-aligned natural statement and the data table information of the database table.

[0134] After obtaining the data table information of the database table, a model training data set can be constructed based on the aligned natural statement and the data table information of the database table. The model training data set may include: positive sample data and negative sample data, where the positive sample data is the aligned natural statement and the data table information of the database table, and the negative sample data is the non-aligned natural statement and the data table information of the database table.

[0135] In specific implementation, preparation of positive sample data: 1. Aligned natural statements: Collect or write natural statements aligned with the data table information. These natural statements should be able to accurately describe or query specific information in the data table. 2. Matching of data table information: Match each aligned natural statement with the corresponding data table information to form positive sample data.

[0136] Preparation of negative sample data: 1. Unaligned natural sentences: Collect or write natural sentences that are not aligned with the data table information. These sentences may contain information unrelated to the data table or may not accurately describe or query the information in the data table. 2. Random pairing of data table information: Pair each unaligned natural sentence with randomly selected data table information to form negative sample data.

[0137] When constructing negative sample data, custom phrases (such as chat phrases, etc.) can be added to replace natural sentences. Using this negative sample data for the model can improve the generalization and rejection ability of the model.

[0138] After constructing the model training dataset based on natural sentences and data table information, perform sub-step D3.

[0139] Sub-step D3: Train the table selection model based on the model training dataset.

[0140] After constructing the model training dataset based on natural sentences and data table information, the table selection model can be trained based on the model training dataset. The model training process can include the following steps:

[0141] Step 1. Selection of the model.

[0142] According to the specific application scenario and requirements, select a suitable machine learning or deep learning model as the table selection model. This model should be able to process natural sentences and data table information and output a prediction result on whether they are aligned.

[0143] Step 2. Training of the model.

[0144] Use the model training dataset to train the table selection model. During the training process, the model will learn how to distinguish positive sample data and negative sample data and gradually improve the prediction accuracy.

[0145] Step 3. Verification and tuning of the model.

[0146] After the training is completed, use the validation dataset to verify the model to evaluate its performance. If the performance of the model is not good, it may be necessary to adjust the model parameters, increase the training data, or try other models.

[0147] Step 4. Deployment of the model.

[0148] Once the model reaches a satisfactory performance, it can be deployed to the production environment for actual application scenarios.

[0149] In the embodiments of this application, 22,000 <NL, SQL, TABLE> aligned data were constructed through experiments. Positive and negative example data matching <NL, TABLE> can be sampled from them. Based on BERT, a matching model for NL and Table was trained and tested, and the accuracy, recall rate, and F1 (the harmonic mean of Precision and Recall, which is a comprehensive manifestation of these two metrics and is used to more comprehensively evaluate the performance of the model) were all greatly improved, as shown in Table 1 below:

[0150] Table 1

[0151] Model Dataset Accuracy Recall F1 BERT 4 tables / 4 tables 92.7 91.4 92.01 BERT 4 tables / hold1 table 91.6 90.8 91.19 BERT 30 tables / hold18 tables 90.7 90.4 90.5

[0152] In the embodiments of this application, a table selection model was constructed and trained. This model can determine whether a given natural statement is aligned with specific database table information. This helps to improve the accuracy and efficiency of data query in practical applications.

[0153] Step 105: Based on the aligned natural statement, the database table, and the SQL statement, train to obtain an SQL statement generation model.

[0154] After obtaining the aligned natural statement, database table, and SQL statement, an SQL statement generation model can be trained based on the aligned natural statement, database table, and SQL statement. The training process of the SQL statement generation model can include the following steps:

[0155] Step 1: Prepare training data.

[0156] It is necessary to ensure that each natural statement is completely aligned with a specific database table information and the corresponding SQL statement, and organize these data into a format easy for model training. For example, each sample can be a JSON object containing instructions, natural statements, database table metadata (such as column names and data types), and SQL statements, etc.

[0157] Step 2: Select a model architecture.

[0158] LLM-Qwen model: The LLM-Qwen model is based on a deep learning framework, such as Transformer, and integrates the capabilities of language understanding and generation. LLM-Qwen is a super-large-scale pre-trained language model designed to provide strong support for various natural language processing tasks. It is pre-trained on a large amount of text data, which enables it to have extensive knowledge and language understanding capabilities. And it can support the NL2SQL function, enabling users to interact with the database in a more natural way by converting natural language queries into SQL query statements.

[0159] Step 3: Model training.

[0160] 1. Define the loss function: For the SQL statement generation task, the commonly used loss function is the Cross-Entropy Loss, which measures the difference between the SQL statements generated by the model and the true SQL statements.

[0161] 2. Optimizer selection: Select a suitable optimizer to update the weights of the model, such as Adam, SGD, or RMSprop, etc.

[0162] 3. Training process: Batch the organized training data and input it into the model, and perform iterative training using the selected loss function and optimizer. At the same time, through SFT (Supervised Fine-Tuning), use the labeled dataset for further training to optimize the model, so that the model can learn more professional domain knowledge or follow a specific instruction format, thereby improving the accuracy and relevance of task completion. During the training process, the performance of the model on the validation set can be monitored to adjust hyperparameters such as the learning rate and batch size, and prevent overfitting.

[0163] Step 4. Model evaluation and tuning.

[0164] 1. Evaluation metrics: Use metrics such as accuracy to evaluate the performance of the model on the test set. For the SQL statement generation task, evaluation metrics commonly used in natural language processing such as BLEU and ROUGE can also be considered to measure the quality of the generated SQL statements.

[0165] 2. Model tuning: Adjust the model's architecture, hyperparameters, or training strategy according to the evaluation results to improve its performance.

[0166] Step 5. Model deployment and application.

[0167] 1. Model export and deployment: Once the model training is completed and reaches satisfactory performance, it can be exported in a deployable format (such as TensorFlow SavedModel, ONNX format, etc.) and deployed to the production environment.

[0168] 2. User interaction and feedback: In actual applications, users can query the database by inputting natural statements. The model will generate corresponding SQL statements and execute the query to return the results. Users can provide feedback to improve the model. For example, by marking whether the generated SQL statements are correct or providing additional training data.

[0169] In the embodiments of the present application, by training and deploying a model capable of generating SQL statements based on natural sentences and database table information, the process of writing database queries can be greatly simplified, and non-technical personnel can also easily interact with the database.

[0170] In this embodiment, after obtaining the aligned data of natural language, table, and SQL, on the one hand, a matching model for natural language and table (i.e., the table selection model in this embodiment) can be trained, and on the other hand, an NL2SQL model (i.e., the SQL statement generation model in this embodiment) can be trained. As Figure 3 shown, the information of natural sentences and the data table Table can be used to train an NL-Table matching model (this model can be a BERT model), and the NL-Table matching model can output a result of 0 or 1. 0 indicates non-matching, and 1 indicates matching. The information of natural sentences and the data table Table can be used to train an NL2SQL / SQL2NL model (this model can be an LLM-Qwen model), and this model can automatically generate SQL statements based on the input natural sentences and the information of the data table Table.

[0171] After training the table selection model and the SQL statement generation model, they can be used in the scenario of generating SQL statements. The process of model inference can be described in detail in combination with the following specific implementation methods.

[0172] In a specific implementation manner of the present application, after the above step 105, it further includes:

[0173] Step E1: Obtain the problem statement input by the user.

[0174] In this embodiment, the problem statement, that is, the natural sentence, refers to the problem or requirement for the user to query the database through natural language.

[0175] When the user performs operations such as querying the database, the problem statement input by the user can be obtained. In a specific implementation, it can be implemented through a user interface (such as a search box, chat window, etc.), where the user can input the problem statement.

[0176] After obtaining the problem statement input by the user, step E2 is executed.

[0177] Step E2: Calculate the similarity between the problem statement and the column names of each database table in multiple database tables.

[0178] After obtaining the question statement input by the user, the similarity between the question statement and the column names of each database table in multiple database tables can be calculated. Specifically, the first vector corresponding to the question statement and the second vectors corresponding to the column names of multiple database tables can be obtained, and the similarity between the first vector and the second vectors can be calculated.

[0179] The purpose of this step is to determine that the columns of the database table have nothing to do with the user's question. In this example, text similarity algorithms (such as cosine similarity, Jaccard similarity, etc.) can be used to calculate the similarity between the question statement and the column names of the database table. Of course, a pre-trained embedding model (such as BERT, Word2Vec, etc.) can also be used to convert the question statement and the database table column names into vectors and calculate the similarity between these vectors.

[0180] After calculating the similarity between the question statement and the column names of each database table in multiple database tables, step E3 is executed.

[0181] Step E3: Remove the data columns in each database table of the multiple database tables whose similarity is less than the similarity threshold.

[0182] The similarity threshold refers to the threshold of the similarity preset for column screening of the database table. The specific value of the similarity threshold can be determined according to business requirements, and this embodiment does not limit it.

[0183] After calculating the similarity between the question statement and the column names of each database table in multiple database tables, the data columns in each database table of the multiple database tables whose similarity is less than the similarity threshold can be removed.

[0184] After removing the data columns in each table of the multiple database tables whose similarity is less than the similarity threshold, step E4 is executed.

[0185] Step E4: Call the table selection model to process the question statement and the table information of the multiple database tables after removing the database tables and / or data columns to obtain the target database table in the multiple database tables.

[0186] After removing the data columns in each database table of the multiple database tables whose similarity is less than the similarity threshold, the table selection model can be called to process the question statement and the table information of the multiple database tables after removing the data columns to obtain the target database table in the multiple database tables, where the table information after removing the data columns is the information obtained by concatenating the database table names of the multiple database tables and the column names of the remaining data columns after removing the data columns.

[0187] For the above implementation process, it can be as Figure 5As shown, the user inputs an NL statement, and filters the columns of the database table through the column name filtering method (i.e., similarity matching filtering), so that irrelevant columns in multiple database tables can be eliminated. That is, after the user poses a question, data preprocessing (similarity matching of column names and natural language statements, filtering out data less than the threshold) is performed based on the selected multiple tables. Then, the data table information of multiple database tables after column elimination (such as table names, column names (the column names of the remaining columns after eliminating columns with lower similarity in the database table)) can be input into the NL-Table matching model (i.e., the table selection model). Through this model, it can be output whether the question statement input by the user is associated with the table name (i.e., the name of the database table) and the column name (i.e., the column name of the filtered database table), and the output result is 0 / 1. 0 indicates not associated, and 1 indicates associated. That is, based on the table selection model, natural language and tables are matched, and whether they are associated is matched according to the similarity threshold.

[0188] After calling the table selection model to process the question statement and the data table information of multiple database tables after eliminating data columns, and obtaining the target database tables in multiple database tables, step E5 is executed.

[0189] Step E5: Call the SQL statement generation model to process the question statement and the data table information of the target database table, and generate the target SQL statement corresponding to the question statement.

[0190] After calling the table selection model to process the question statement and the data table information of multiple database tables after eliminating data columns, and obtaining the target database tables in multiple database tables, the SQL statement generation model can be called to process the question statement and the data table information of the target database table, and generate the target SQL statement corresponding to the question statement.

[0191] For the generation process of the SQL statement, it can be as Figure 4 shown. Through the table selection model, the input natural statement and the data table information of the database table (i.e., <NL, Tables>) can be processed to screen the database table. Through the table selection model, the associated database tables can be identified, and the unassociated database tables (i.e., rejection recognition) can be eliminated. Then, the database tables screened by the table selection model and the natural statement input by the user can be input into the NL2SQL model (i.e., the SQL statement generation model) for processing to generate the final SQL statement.

[0192] The text-to-SQL method based on the collaboration of large models (i.e., the table selection model and the SQL statement generation model) in the embodiments of the present application first finds the most relevant tables in the data warehouse tables based on the user's query (i.e., the table selection model screens the database tables), and then inputs the relevant content into the large model (i.e., the SQL statement generation model) to generate the final SQL statement.

[0193] The model training method provided by the embodiment of the present application processes the business data in the preset database table based on the SQL syntax rules to obtain SQL statements. Process the SQL statements according to the preset data generation method to construct aligned natural statements and SQL statements. Determine the aligned database table, natural statement, and SQL statement according to the preset database table, natural statement, and SQL statement. Train a table selection model based on the aligned natural statement and database table. Train an SQL statement generation model based on the aligned natural statement, database table, and SQL statement. By training the table selection model in the embodiment of the present application, the table related to the user query can be found more accurately, thereby improving the generalization ability of the model. By combining the collaborative work of the large and small models, this solution can understand the user's query intention more accurately. The small model is responsible for quickly screening relevant tables, and the large model is responsible for deeply understanding the user's question and generating the correct SQL query statement, improving the accuracy of user intention understanding and reducing the influence of ambiguity and diversity.

[0194] Refer to Figure 6 , which shows a schematic structural diagram of a model training device provided by the embodiment of the present application. As Figure 6 shown, the model training device 600 may include the following modules:

[0195] The SQL statement acquisition module 610 is used to process the business data in the preset database table based on the SQL syntax rules to obtain SQL statements;

[0196] The aligned statement construction module 620 is used to process the SQL statements according to the preset data generation method to construct aligned natural statements and SQL statements;

[0197] The aligned statement determination module 630 is used to determine the aligned database table, natural statement, and SQL statement according to the preset database table, natural statement, and SQL statement;

[0198] The table selection model training module 640 is used to train a table selection model based on the aligned natural statement and database table;

[0199] The SQL statement generation model training module 650 is used to train an SQL statement generation model based on the aligned natural statement, database table, and SQL statement.

[0200] Optionally, the aligned statement construction module includes:

[0201] The SQL statement filling unit is used to fill the SQL statements into a pre-configured statement alignment template;

[0202] The first alignment statement determination unit is configured to determine the aligned natural statement and the SQL statement based on the statement alignment template.

[0203] Optionally, the alignment statement construction module includes:

[0204] The SQL statement input unit is configured to input the SQL statement into a pre-trained natural statement generation model;

[0205] The natural statement acquisition unit is configured to acquire the natural statement corresponding to the SQL statement output by the natural statement generation model;

[0206] The first alignment statement determination unit is configured to determine the aligned natural statement and the SQL statement.

[0207] Optionally, the alignment statement construction module includes:

[0208] The example data selection unit is configured to acquire the table description information of the data table and select the specified row data from the data table as the example data;

[0209] The example data input unit is configured to input the table description information and the example data into a pre-trained statement alignment model;

[0210] The alignment statement acquisition unit is configured to acquire the aligned natural statement and the SQL statement output by the statement alignment model.

[0211] Optionally, the table selection model training module includes:

[0212] The data table information acquisition unit is configured to acquire the data table information of the database table;

[0213] The dataset construction unit is configured to construct a model training dataset based on the natural statement and the data table information, and the model training dataset includes: positive sample data and negative sample data, where the positive sample data is the aligned natural statement and the data table information of the database table, and the negative sample data is the non-aligned natural statement and the data table information of the database table;

[0214] The table selection model training unit is configured to train the table selection model based on the model training dataset.

[0215] Optionally, the included device further includes:

[0216] The problem statement acquisition module is configured to acquire the problem statement input by the user;

[0217] The similarity calculation module is configured to calculate the similarity between the problem statement and the column names of each database table in multiple database tables;

[0218] A data column elimination module, configured to eliminate data columns in each of the multiple database tables with a similarity less than a similarity threshold;

[0219] A target database table acquisition module, configured to call the table selection model to process the problem statement and the data table information of the multiple database tables after eliminating data columns, and obtain a target database table in the multiple database tables, where the data table information after eliminating data columns is information obtained by concatenating the database table names of the multiple database tables and the column names of the remaining data columns after eliminating data columns;

[0220] A target SQL statement generation module, configured to call the SQL statement generation model to process the problem statement and the data table information of the target database table, and generate a target SQL statement corresponding to the problem statement.

[0221] Optionally, the similarity calculation module includes:

[0222] A vector acquisition unit, configured to acquire a first vector corresponding to the problem statement and a second vector corresponding to the column name of each database table in the multiple database tables;

[0223] A similarity calculation unit, configured to calculate the similarity between the first vector and the second vector.

[0224] The model training device provided by the embodiments of the present application processes business data in a preset database table based on SQL syntax rules to obtain an SQL statement. Processes the SQL statement according to a preset data generation method to construct an aligned natural statement and SQL statement. Determines the aligned database table, natural statement, and SQL statement according to the preset database table, natural statement, and SQL statement. Trains a table selection model based on the aligned natural statement and database table. Trains an SQL statement generation model based on the aligned natural statement, database table, and SQL statement. By training the table selection model, the embodiments of the present application can more accurately find a table related to a user query, thereby improving the generalization ability of the model. By combining the collaborative work of large and small models, this solution can more accurately understand the user's query intention. The small model is responsible for quickly screening relevant tables, and the large model is responsible for deeply understanding the user's problem and generating a correct SQL query statement, improving the accuracy of user intention understanding and reducing the impact of ambiguity and diversity.

[0225] The embodiments of the present application further provide an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the computer program is executed by the processor, the above-mentioned model training method is implemented.

[0226] Figure 7 The structural schematic diagram of an electronic device 700 according to an embodiment of the present invention is shown. As Figure 7 shown, the electronic device 700 includes a central processing unit (CPU) 701, which can execute various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) 702 or computer program instructions loaded from a storage unit 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the electronic device 700 can also be stored. The CPU 701, ROM 702, and RAM 703 are connected to each other through a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0227] Multiple components in the electronic device 700 are connected to the I / O interface 705, including: an input unit 706, such as a keyboard, a mouse, a microphone, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, an optical disc, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the electronic device 700 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0228] Each of the processes and processes described above can be executed by the processing unit 701. For example, the method of any of the above embodiments can be implemented as a computer software program, which is tangibly included in a computer-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the CPU 701, one or more actions in the method described above can be executed.

[0229] Additionally, an embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the above model training method is implemented.

[0230] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is the difference from other embodiments. The same or similar parts between the embodiments can be referred to each other.

[0231] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, devices, or computer program products. Therefore, the embodiments of the present application can take the form of all-hardware embodiments, all-software embodiments, or embodiments combining software and hardware aspects. Moreover, the embodiments of the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0232] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of methods, terminals (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminals to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing terminals generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0233] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminals to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0234] These computer program instructions can also be loaded onto a computer or other programmable data processing terminals, such that a series of operation steps are executed on the computer or other programmable terminals to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable terminals provide steps for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0235] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications to these embodiments once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present application.

[0236] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or terminal. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or terminal comprising said element.

[0237] The above has introduced in detail a model training method, a model training device, an electronic device and a computer-readable storage medium provided by the present application. Specific examples are used in this text to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A model training method, characterized in that, The method includes: Processing business data in a preset database table based on SQL syntax rules to obtain an SQL statement; Processing the SQL statement according to a preset data generation method to construct an aligned natural statement and SQL statement; Determining the aligned database table, natural statement, and SQL statement based on the preset database table, the natural statement, and the SQL statement; Training a table selection model based on the aligned natural statement and the database table; Training an SQL statement generation model based on the aligned natural statement, the database table, and the SQL statement.

2. The method according to claim 1, characterized in that The processing the SQL statement according to a preset data generation method to construct an aligned natural statement and SQL statement includes: Filling the SQL statement into a pre-configured statement alignment template; Determining the aligned natural statement and the SQL statement based on the statement alignment template.

3. The method according to claim 1, wherein The processing the SQL statement according to a preset data generation method to construct an aligned natural statement and SQL statement includes: Inputting the SQL statement into a pre-trained natural statement generation model; Obtaining the natural statement corresponding to the SQL statement output by the natural statement generation model; Determining the aligned natural statement and the SQL statement.

4. The method according to claim 1, wherein The processing the SQL statement according to a preset data generation method to construct an aligned natural statement and SQL statement includes: Obtaining the table description information of a data table and selecting specified row data from the data table as example data; Inputting the table description information and the example data into a pre-trained statement alignment model; Obtaining the aligned natural statement and the SQL statement output by the statement alignment model.

5. The method according to claim 1, wherein The training a table selection model based on the aligned natural statement and the database table includes: Obtaining the data table information of the database table; Constructing a model training data set based on the natural statement and the data table information, where the model training data set includes: positive sample data and negative sample data, where the positive sample data is the aligned natural statement and the data table information of the database table, and the negative sample data is the unaligned natural statement and the data table information of the database table; Training the table selection model based on the model training data set.

6. The method according to claim 1, characterized in that After training the SQL statement generation model based on the aligned natural statement, the database table, and the SQL statement, it further includes: Obtaining a problem statement input by a user; Calculating the similarity between the problem statement and the column names of each database table in a plurality of database tables; Eliminating the data columns of each database table in the plurality of database tables with a similarity less than a similarity threshold; Invoking the table selection model to process the problem statement and the data table information of the plurality of database tables after eliminating the data columns to obtain a target database table in the plurality of database tables, where the data table information after eliminating the data columns is the information obtained by concatenating the database table names of the plurality of database tables and the column names of the remaining data columns after eliminating the data columns; Call the generated model of the SQL statement to process the problem statement and the data table information of the target database table, and generate the target SQL statement corresponding to the problem statement.

7. The method according to claim 6, wherein The calculation of the similarity between the problem statement and the column names of each table in multiple database tables includes: Obtain the first vector corresponding to the problem statement and the second vector corresponding to the column names of each database table in the multiple database tables; Calculate the similarity between the first vector and the second vector.

8. A model training device, characterized in that, The device includes: An SQL statement acquisition module, configured to process the business data in a preset database table based on SQL syntax rules to obtain an SQL statement; An alignment statement construction module, configured to process the SQL statement according to a preset data generation method to construct an aligned natural statement and SQL statement; An alignment statement determination module, configured to determine the aligned database table, natural statement, and SQL statement according to the preset database table, the natural statement, and the SQL statement; A table selection model training module, configured to train a table selection model based on the aligned natural statement and the database table; An SQL statement generation model training module, configured to train an SQL statement generation model based on the aligned natural statement, the database table, and the SQL statement.

9. An electronic device, characterized in that, Includes: A processor, a memory, and a computer program stored on the memory and executable on the processor, where when the processor executes the program, it implements the model training method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, When the instructions in the storage medium are executed by the processor of the electronic device, the electronic device is enabled to execute the model training method according to any one of claims 1 to 7.