Method for training NL2SQL model in three stages

By employing a three-stage training method and LoRa fine-tuning, the accuracy and stability issues of the NL2SQL model in complex scenarios were resolved, achieving efficient SQL generation and execution.

CN121009948APending Publication Date: 2025-11-2510TH RES INST OF CETC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511129065.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-13
Publication Date
2025-11-25

AI Technical Summary

Technical Problem

Traditional NL2SQL models have low schema linking accuracy in multi-table join and nested query scenarios, suffer from format errors and non-executable statements, and are unstable in training, making convergence difficult.

Method used

A three-stage training method is adopted, including cold start, GRPO and feedback stages. Combined with LoRa fine-tuning and multi-dimensional dynamic reward function, the model is optimized through supervised learning and feedback data, thereby gradually improving the accuracy and stability of the generated SQL.

Benefits of technology

It improves the accuracy and execution efficiency of natural language to SQL statement conversion, can quickly output correct SQL answers, reduces the number of training parameters, and improves the robustness of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121009948A_ABST
    Figure CN121009948A_ABST
Patent Text Reader

Abstract

The invention discloses a three-stage NL2SQL model training method, which comprises the following steps: preprocessing original data to obtain a training data set, and initializing a model; the original data comprises a database table name, an sql statement for creating a database, a user question, a standard SQL statement and retrieval content; in the cold start stage, loading the Base model, the Lorainit model and the training data set, and obtaining a LoraIce model after training is completed; the LoraIce model is used as an initial Lora model of the GRPO stage; in the GRPO stage, the Base model, the LoraIce model and the training data set are loaded, training is completed, a LoraGRPO model is obtained, and feedback data are collected in the training process; in the feedback stage, the Base model, the Lorainit model and the feedback data set are loaded, training is completed, and a LoraFeedback model is obtained. According to the method, the conversion accuracy and execution efficiency from the natural language to the SQL statement are improved, and the correct answer can be quickly output according to the question input by the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of natural language processing and database interaction technology, and in particular to a method for training an NL2SQL model in three stages. Background Technology

[0002] Traditional NL2SQL (Natural Language to Structured Query Language) models suffer from low schema linking accuracy in complex scenarios such as multi-table joins and nested queries. Traditional SQL generation and evaluation methods either reference standard SQL information or rely solely on the final execution result of the SQL, lacking process supervision and leading to formatting errors (such as missing closing tags) or unexecutable statements. They also exhibit high error rates in understanding long-context table structures. Large models directly generating SQL may contain fictitious table names and fields. Traditional PPO algorithms require establishing independent value functions, increasing memory consumption, and model training is often unstable and difficult to converge. Summary of the Invention

[0003] In view of this, this application provides a three-stage method for training an NL2SQL model.

[0004] This application discloses a three-stage method for training an NL2SQL model, which includes: Step 1: Preprocess the raw data to obtain the training dataset and initialize the model; the raw data includes database table names, SQL statements for creating the database, user questions, standard SQL statements, and search content; Step 2: Cold start phase, load the Base model, Lora_init model and training dataset, training is completed, and the Lora_Ice model is obtained; the Lora_Ice model serves as the initial Lora model for the GRPO phase; Step 3: GRPO phase, load the Base model, Lora_Ice model and training dataset, complete training to obtain the Lora_GRPO model, and collect feedback data during the training process; Step 4: Feedback phase. Load the Base model, Lora_init model and feedback dataset, complete training, and obtain the Lora_Feedback model.

[0005] Further, step 1 includes: Step 1.1: Perform data preprocessing on the original data to obtain the training dataset; the preprocessing includes injecting noise and data format conversion; Step 1.2: Randomly sample the training data and determine the training datasets for the cold start phase and the GRPO phase according to a preset ratio; the number of training data in the training dataset of the cold start phase is less than the number of training data in the training dataset of the GRPO phase. Step 1.3: Initialize matrices A and B in LoRa. Matrix A follows a normal distribution, and matrix B is a matrix of all zeros. Set the hyperparameter r of LoRa to a positive integer and the scaling parameter alpha to a positive integer. Step 1.4: Select the Base model as the initial model for training the Lora_Ice model; the Lora_Ice model refers to the model trained using the Lora method during the cold start phase.

[0006] Furthermore, it also includes: Both the cold start and feedback phases employ supervised learning methods. The formula for the output of LoRa fine-tuning is:

[0007]

[0008]

[0009] In the formula, h is the output of the model's Attention layer after the input x passes through it. W in the Pre-trained module refers to the Wq, Wk, and Wv matrices of the original model's Attention layer, which do not participate in parameter gradient updates during training. Matrix A is randomly Gaussian initialized, and matrix B is zero initialized. During training, the parameters are updated iteratively according to the gradient.

[0010] Further, step 3 includes: Step 31: Load the Policy Model and the Reference Model; the Reference Model consists of the base model and Lora_ice; set the hyperparameters, including training period, learning rate, batch size, number of generators per group, temperature sampling, and kernel sampling coefficient top_p; Step 32: Multiple Response Generation: For each question Q, generate multiple responses by using temperature sampling and top_p sampling to increase the diversity of the generated responses; Step 33: Calculate the reward value for a single response based on the reward function, and calculate the advantage based on the reward value; Step 34: Calculate the gradient value based on the loss function; Step 35: During the training period, update the policy model parameters based on the gradient values ​​until the model training is complete.

[0011] Further, step 33 includes: The dominance is calculated based on the average reward and standard deviation of the within-group response, using the following formula:

[0012] Where: G refers to the number of items generated within the group.

[0013] Further, step 34 includes: The gradient value is obtained using the following formula: in, Used to measure the generation of new strategies The probability changes compared to the old strategy; clip() is the clipping function, which limits the gradient change range to... To prevent drastic changes from causing the model to crash; For KL divergence penalty term, β For the coefficient term of KL, Refers to the updated strategy. The reference strategy aims to prevent the strategy from deviating from the initial model.

[0014] Furthermore, the design method of the reward function includes: The scores for the first to fifth reward functions are all specified values, and the scoring criteria are as follows: The first reward function is format accuracy, which determines whether the model's output has a proprietary format. <think>,<\think>, <sql>The first sub-scoring rule: checks if each proprietary format appears and if its frequency is 1; if so, the score increases by 0.2. The second sub-scoring rule: when the score of the first sub-scoring rule reaches 0.8, it checks the order of the proprietary formats to ensure that... L( <think> )< / think> Less than L(<think>) Less than L( <sql> )< / sql> Less than L(<\sql>) ,in, L(x) Refers to each proprietary string x The starting position in the generated string; The second reward function is based on SQL syntax correctness. The third sub-scoring rule is to verify the correctness of the SQL statement using Python's built-in sqlparse library; if the syntax is correct, the score is 1. The fourth sub-scoring rule is to determine if the third sub-scoring rule is not met. "or" "Whether the format appears or not, the score is 0.2 if it appears, otherwise it is 0; The third reward function is SQL executability. For the established SQL database table, the system connects to the database and executes the SQL statement. Successful execution scores 1, otherwise it scores 0. The fourth reward function is the accuracy of the query results. Based on the data returned by the executed SQL statement, a consistency check is performed between the data and the original data. The score is 1 if the data content is consistent, and 0 otherwise. Fifth reward function: Execution efficiency, where the execution times of the generated SQL and the standard SQL are g(t) and e(t), respectively. The scoring formula is:

[0015] The first to fifth reward functions are normalized and summed according to their weights to obtain the final reward function.

[0016] Furthermore, the formula for the weight is:

[0017] The final reward function formula is:

[0018] in, Let be the i-th reward function.

[0019] Furthermore, in the cold start phase and the GRPO phase, the model's input consists of a user-inputted question, and the SQL database table includes the table name: {table_name}, database table creation information: {table_ddl}, and the user question: {user_question}; the model's output is... <think>{think module}<\think>{explanation} <sql>{SQL module}<\sql>; where the think module represents the model's reasoning process; and the SQL module contains the SQL statements generated by the model, used to execute the SQL.

[0020] Furthermore, in the feedback phase, the feedback data all come from data in the GRPO phase where the SQL statement is correct but execution fails; The SQL database includes the table name: {table_name}, table creation information: {table_ddl}, executed SQL statement: {user_question}, and error information: {error_info}; the model output is {new_SQL}.

[0021] Due to the adoption of the above technical solution, this application has the following advantages: This application improves the accuracy and execution efficiency of the conversion from natural language to SQL statements, and can quickly output the correct answer based on the user's input question. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments recorded in the embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings.

[0023] Figure 1 This is a schematic diagram of the Lora fine-tuning principle in an embodiment of this application; Figure 2 This is a schematic diagram of the GRPO training principle in an embodiment of this application; Figure 3 This is a schematic diagram of the three-stage NL2SQL training principle in an embodiment of this application. Detailed Implementation

[0024] The present application will be further described in conjunction with the accompanying drawings and embodiments. The described embodiments are only some, not all, of the embodiments of the present application. All other embodiments obtained by those skilled in the art should fall within the protection scope of the embodiments of the present application.

[0025] See Figures 1 to 3 This application provides an embodiment of a three-stage training method for an NL2SQL model, comprising: Step 1: Preprocess the raw data to obtain the training dataset and initialize the model; the raw data includes database table names, SQL statements for creating the database, user questions, standard SQL statements, and search content; Step 2: Cold start phase, load the Base model, Lora_init model and training dataset, training is completed, and the Lora_Ice model is obtained; the Lora_Ice model serves as the initial Lora model for the GRPO phase; Step 3: GRPO phase, load the Base model, Lora_Ice model and training dataset, complete training to obtain the Lora_GRPO model, and collect feedback data during the training process; Step 4: Feedback phase. Load the Base model, Lora_init model and feedback dataset, complete training, and obtain the Lora_Feedback model.

[0026] The three-stage training process of this application (cold start - GRPO - feedback): Cold start phase: Supervised fine-tuning method is adopted, mainly to lay the foundation for the GRPO (Group Relative Policy Optimization) phase and solve the instability problem in the early stage of training; GRPO phase: Eliminate the requirement of traditional value function and enhance the model's capabilities based on multi-dimensional rewards; Feedback phase: Supervised fine-tuning method is adopted to address SQL execution errors and, based on error feedback information, focus on solving the problem of fictitious table names and table fields in SQL statements.

[0027] LoRa parameter training for this application: The entire training process uses the LoRa method, without changing the parameters and performance of the original model; compared with full training, it significantly reduces the number of training parameters.

[0028] Data structure for this application: Noise is injected into the raw data (randomly deleting table annotation fields) to improve the model's robustness. Simultaneously, training data from the feedback phase is collected during the GRPO training phase to improve data utilization.

[0029] Optionally, step 1 includes: Step 1.1: Perform data preprocessing on the original data to obtain the training dataset; the preprocessing includes injecting noise and data format conversion; Step 1.2: Randomly sample the training data and determine the training datasets for the cold start phase and the GRPO phase according to a preset ratio; the number of training data in the training dataset of the cold start phase is less than the number of training data in the training dataset of the GRPO phase; the amount of training data in the cold start phase is less than the amount of training data in the GRPO phase, and can be obtained by random sampling from the original data at a ratio of 10% and 90% respectively. Step 1.3: Initialize matrices A and B in LoRa. Matrix A follows a normal distribution, and matrix B is a matrix of all zeros. Set the hyperparameter r of LoRa to a positive integer and the scaling parameter alpha to a positive integer. Step 1.4: Select the Base model as the initial model for training the Lora_Ice model; the Lora_Ice model refers to the model trained using the Lora method during the cold start phase.

[0030] Optionally, it also includes: Both the cold start and feedback phases employ supervised learning methods. The formula for the output of LoRa fine-tuning is:

[0031]

[0032]

[0033] In the formula, h is the output of the model's Attention layer after the input x passes through it. W in the Pre-trained module refers to the Wq, Wk, and Wv matrices of the original model's Attention layer, which do not participate in parameter gradient updates during training. Matrix A is randomly Gaussian initialized, and matrix B is zero initialized. During training, the parameters are updated iteratively according to the gradient.

[0034] Optionally, see Figure 2 Step 3 includes: Step 31: Load the Policy Model ( Figure 2 The Policy Model (trainable) and the Reference Model (in the system) Figure 2 The Reference Model in the code has fixed parameters and can be either the base model or Lora_ice. The Reference Model consists of the base model and Lora_ice. Hyperparameters are set, including training period, learning rate, batch size, number of generators per group, temperature sampling, and kernel sampling coefficients top_p. Step 32: Multiple Response Generation: For each question Q, generate multiple responses by using temperature sampling and top_p sampling to increase the diversity of the generated responses; Step 33: Calculate the reward value for a single response based on the reward function, and calculate the advantage based on the reward value; Step 34: Calculate the gradient value based on the loss function; Step 35: During the training period, update the policy model parameters based on the gradient values ​​until the model training is complete.

[0035] Optionally, step 33 includes: The dominance is calculated based on the average reward and standard deviation of the within-group response, using the following formula:

[0036] Where: G refers to the number of items generated within the group.

[0037] Optionally, step 34 includes: The gradient value is obtained using the following formula: in, Used to measure the generation of new strategies The probability of the new policy changes relative to the old policy; if the ratio is far apart, it indicates that the policy update is too large, which may lead to instability. clip() It is a clipping function that restricts the range of gradient change to... To prevent drastic changes from causing the model to crash; For KL divergence penalty term, β For the coefficient term of KL, Refers to the updated strategy. The reference strategy aims to prevent the strategy from deviating from the initial model.

[0038] The multi-dimensional dynamic reward function of this application: The design incorporates a tiered reward mechanism (total reward = format score + syntax score + executable score + result score + execution efficiency score). In the early stages of training, the model primarily rewards format and syntax scores, while in the later stages, the focus shifts to the quality of the generated SQL. The score weights are dynamically adjusted according to the training cycle, remaining constant after training has reached halfway point.

[0039] Optionally, the design method of the reward function includes: The scores for the first to fifth reward functions are all specified values ​​(e.g., a specified value of 1), and the scoring criteria are as follows: The first reward function is format accuracy, which determines whether the model's output has a proprietary format. <think>,<\think>, <sql>The first sub-scoring rule: checks if each proprietary format appears and if its frequency is 1; if so, the score increases by 0.2. The second sub-scoring rule: when the score of the first sub-scoring rule reaches 0.8, it checks the order of the proprietary formats to ensure that... L( <think> )< / think> Less than L(<think>) Less than L( <sql> )< / sql> Less than L(<\sql>) ,in, L(x) Refers to each proprietary string x The starting position in the generated string; The second reward function is based on SQL syntax correctness. The third sub-scoring rule is to verify the correctness of the SQL statement using Python's built-in sqlparse library; if the syntax is correct, the score is 1. The fourth sub-scoring rule is to determine if the third sub-scoring rule is not met. "or" "Whether the format appears or not, the score is 0.2 if it appears, otherwise it is 0; The third reward function is SQL executability. For the established SQL database table, the system connects to the database and executes the SQL statement. Successful execution scores 1, otherwise it scores 0. The fourth reward function is the accuracy of the query results. Based on the data returned by the executed SQL statement, a consistency check is performed between the data and the original data. The score is 1 if the data content is consistent, and 0 otherwise. Fifth reward function: Execution efficiency, where the execution times of the generated SQL and the standard SQL are g(t) and e(t), respectively. The scoring formula is:

[0040] The first to fifth reward functions are normalized and summed according to their weights to obtain the final reward function. The multi-dimensional dynamic reward mechanism is shown in Table 1.

[0041] Table 1 Multi-dimensional Dynamic Reward Mechanism

[0042] Optionally, the formula for the weight is:

[0043] The final reward function formula is:

[0044] in, Let be the i-th reward function.

[0045] Optionally, in the cold start phase and the GRPO phase, the input to the model is a question input by the user (e.g., you are an NL2SQL expert, please generate the SQL language required by the user based on the input SQL database table information and the user's question), where the SQL database table includes table name: {table_name}, database table creation information: {table_ddl}, and user question: {user_question}; the output of the model is... <think>{think module}<\think>{explanation} <sql>{SQL module}<\sql>; where the think module represents the model's reasoning process; and the SQL module contains the SQL statements generated by the model, used to execute the SQL.

[0046] Optionally, in the feedback phase, the feedback data all come from the data in the GRPO phase where the SQL statement is correct but execution fails; for example, the model's role field: You are an NL2SQL expert, please correct and regenerate the SQL based on the input SQL database table information, the executed SQL statement and the error message; The SQL database includes the table name: {table_name}, table creation information: {table_ddl}, executed SQL statement: {user_question}, and error information: {error_info}; the model output is {new_SQL}.

[0047] Specific examples of training data: 1. Example of original data: { "Database name":"students", "DDL":"""CREATE TABLE students ( student_id INT PRIMARY KEY AUTO_INCREMENT, name VARCHAR(50) NOT NULL, gender ENUM('male', 'female') DEFAULT 'male', age TINYINT CHECK (age BETWEEN 17 AND 20), department VARCHAR(50) NOT NULL COMMIT 'Institution', enrollment_date DATE NOT NULL COMMIT 'Enrollment Date' ) ENGINE=InnoDB DEFAULT CHARSET=utf8mb4;""", User Question: "Please count the number of male students aged 18 to 19." SQL search statement: "SELECT COUNT( ) FROM studens WHERE gender = 'male' ANDage BETWEEN 18 AND 19", "Search content":"((9,))" } 2. Cold start and GRPO data examples: Model Input [{ "content": "You are an NL2SQL expert. Please generate the SQL language required by the user based on the input SQL database table information and the user's question." "role":"system" }, { "content":"""There are the following SQL database tables, and the information for each table is as follows:" ## Table name: students ##Table creation information: CREATE TABLE students ( student_id INT PRIMARY KEY AUTO_INCREMENT, name VARCHAR(50) NOT NULL, gender ENUM('male', 'female') DEFAULT 'male', age TINYINT CHECK (age BETWEEN 17 AND 20), department VARCHAR(50) NOT NULL COMMIT 'Institution', enrollment_date DATE NOT NULL COMMIT 'Enrollment Date' ) ENGINE=InnoDB DEFAULT CHARSET=utf8mb4; ## User Question: Please count the number of male students aged 18 to 19. "role":"user" }] Model output { "content":" <think>Okay, I now need to help a user solve an SQL query problem. The user's problem is to count the number of male students aged 18 to 19, and they have provided a student table structure. First, I need to carefully look at the table structure to make sure I understand the meaning and constraints of each field.

[0048] First, the table name is students, and the fields include student_id (primary key, auto-incrementing), name (not null), gender (enumeration, default is male), age (tinyint, range 17 to 20), department (school / institution), and enrollment_date (enrollment date). The user's question is to count the number of male students aged 18 to 19.

[0049] Next, I need to generate the correct SQL statement. Based on the question, the conditions should be gender = 'male' and age between 18 and 19. Note that the user might want to include both 18 and 19 years old, so using BETWEEN 18AND 19 is appropriate. Then use COUNT( To count the number.

[0050] Next, I need to check for any possible errors. For example, are the enumerated values ​​for the `gender` field correct, and is it case-sensitive? According to the DDL, `gender` is an ENUM('male', 'female'), so the value should be 'male' and not any other form. Additionally, the `age` field has a CHECK constraint, ensuring it's between 17 and 20, so 18 and 19 are valid values. Based on the user's needs, the following is the answer to this question. <answer> SELECT COUNT( AS male_students_count FROM students WHERE gender = 'male' AND age BETWEEN 18 AND 19; <\answer>" } 3. Examples of Feedback Data Model Input [{ "content":"You are an NL2SQL expert. Please correct and regenerate the SQL based on the input SQL database table information, the executed SQL statement, and the error message." "role":"system" }, { "content":"""There are the following SQL database tables, and the information for each table is as follows:" ## Table name: students ##Table creation information: CREATE TABLE students ( student_id INT PRIMARY KEY AUTO_INCREMENT, name VARCHAR(50) NOT NULL, gender ENUM('male', 'female') DEFAULT 'male', age TINYINT CHECK (age BETWEEN 17 AND 20), department VARCHAR(50) NOT NULL COMMIT 'Institution', enrollment_date DATE NOT NULL COMMIT 'Enrollment Date' ) ENGINE=InnoDB DEFAULT CHARSET=utf8mb4; ##SQL:SELECT name FROM students WHERE enrollment_time<'2025-03-01'; ## Error message: 1054 - Unknown column 'enrollment_time' in 'where clause' "role":"user" }] Model output: { "content":"SELECT name FROM students WHERE enrollment_date<'2025-03-01';" } Detailed training process S1: Data preparation and model initialization S1.1: Perform data preprocessing on the original data, including injecting noise (randomly deleting table comment fields) and data format conversion, to obtain training data. S1.2: Randomly sample the training data according to a 1:9 ratio to determine the training data for the cold start and GRPO phases.

[0051] S1.3: Initialize the A and B matrices in Lora. Specifically, A is normally distributed, B is a zero matrix, the Lora hyperparameter r is set to 4, and the scaling parameter alpha is set to 4.

[0052] S1.4: Select the Base model. The application recommends using Qwen2-7b.

[0053] S2: Cold start phase, load Base model, Lora_init model and training data set, training is completed, get Lora_Ice model.

[0054] S3: GRPO phase, load Base model, Lora_Ice model and training data set, training is completed, get Lora_GRPO model, collect feedback data in the training process.

[0055] S4: Feedback phase, load Base model, Lora_init model and feedback data set, training is completed, get Lora_Feedback model.

[0056] It should be pointed out finally that the above embodiments are only used to illustrate the technical solutions of the present application but not to limit it. Although the present application has been described in detail with reference to the above embodiments, it should be understood by those skilled in the art that the specific embodiments of the present application can be modified or equivalently replaced without departing from the spirit and scope of the present application, and any modification or equivalent replacement without departing from the spirit and scope of the present application should be covered in the protection scope of the claims of the present application.< / answer> < / think> < / sql> < / think> < / sql> < / think> < / sql> < / think> < / sql> < / think>

Claims

1. A method for training an NL2SQL model in three stages, characterized in that, include: Step 1: Preprocess the raw data to obtain the training dataset and initialize the model; the raw data includes database table names, SQL statements for creating the database, user questions, standard SQL statements, and search content; Step 2: Cold start phase, load the Base model, Lora_init model and training dataset, training is completed, and the Lora_Ice model is obtained; the Lora_Ice model serves as the initial Lora model for the GRPO phase; Step 3: GRPO phase, load the Base model, Lora_Ice model and training dataset, complete training to obtain the Lora_GRPO model, and collect feedback data during the training process; Step 4: Feedback phase. Load the Base model, Lora_init model and feedback dataset, complete training, and obtain the Lora_Feedback model.

2. The method according to claim 1, characterized in that, Step 1 includes: Step 1.1: Perform data preprocessing on the original data to obtain the training dataset; the preprocessing includes injecting noise and data format conversion; Step 1.2: Randomly sample the training data and determine the training datasets for the cold start phase and the GRPO phase according to a preset ratio; the number of training data in the training dataset of the cold start phase is less than the number of training data in the training dataset of the GRPO phase. Step 1.3: Initialize matrices A and B in LoRa. Matrix A follows a normal distribution, and matrix B is a matrix of all zeros. Set the hyperparameter r of LoRa to a positive integer and the scaling parameter alpha to a positive integer. Step 1.4: Select the Base model as the initial model for training the Lora_Ice model; the Lora_Ice model refers to the model trained using the Lora method during the cold start phase.

3. The method according to claim 1, characterized in that, Also includes: Both the cold start and feedback phases employ supervised learning methods. The formula for the output of LoRa fine-tuning is: In the formula, h is the output of the model's Attention layer after the input x passes through it. W in the Pre-trained module refers to the Wq, Wk, and Wv matrices of the original model's Attention layer, which do not participate in parameter gradient updates during training. Matrix A is randomly Gaussian initialized, and matrix B is zero initialized. During training, the parameters are updated iteratively according to the gradient.

4. The method according to claim 1, characterized in that, Step 3 includes: Step 31: Load the Policy Model and the Reference Model; the Reference Model consists of the base model and Lora_ice; set the hyperparameters, including training period, learning rate, batch size, number of generators per group, temperature sampling, and kernel sampling coefficient top_p; Step 32: Multiple Response Generation: For each question Q, generate multiple responses by using temperature sampling and top_p sampling to increase the diversity of the generated responses; Step 33: Calculate the reward value for a single response based on the reward function, and calculate the advantage based on the reward value; Step 34: Calculate the gradient value based on the loss function; Step 35: During the training period, update the policy model parameters based on the gradient values ​​until the model training is complete.

5. The method according to claim 4, characterized in that, Step 33 includes: The dominance is calculated based on the average reward and standard deviation of the within-group response, using the following formula: Where: G refers to the number of items generated within the group.

6. The method according to claim 4, characterized in that, Step 34 includes: The gradient value is obtained using the following formula: in, Used to measure the generation of new strategies The probability changes compared to the old strategy; clip() is the clipping function, which limits the gradient change range to... To prevent drastic changes from causing the model to crash; For KL divergence penalty term, β For the coefficient term of KL, Refers to the updated strategy. The reference strategy aims to prevent the strategy from deviating from the initial model.

7. The method according to claim 4, characterized in that, The design method of the reward function includes: The scores for the first to fifth reward functions are all specified values, and the scoring criteria are as follows: The first reward function is format accuracy, which determines whether the model's output has a proprietary format. <think>,<\think>, <sql>The first sub-scoring rule: checks if each proprietary format appears and if its frequency is 1; if so, the score increases by 0.

2. The second sub-scoring rule: when the score of the first sub-scoring rule reaches 0.8, it checks the order of the proprietary formats to ensure that... L( <think> )< / think> Less than L(<think>) Less than L( <sql> )< / sql> Less than L(<\sql>) ,in, L(x) Refers to each proprietary string x The starting position in the generated string;< / sql> < / think> The second reward function is based on SQL syntax correctness. The third sub-scoring rule is to verify the correctness of the SQL statement using Python's built-in sqlparse library; if the syntax is correct, the score is 1. The fourth sub-scoring rule is to determine if the third sub-scoring rule is not met. "or" "Whether the format appears or not, the score is 0.2 if it appears, otherwise it is 0; The third reward function is SQL executability. For the established SQL database table, the system connects to the database and executes the SQL statement. Successful execution scores 1, otherwise it scores 0. The fourth reward function is the accuracy of the query results. Based on the data returned by the executed SQL statement, a consistency check is performed between the data and the original data. The score is 1 if the data content is consistent, and 0 otherwise. Fifth reward function: Execution efficiency, where the execution times of the generated SQL and the standard SQL are g(t) and e(t), respectively. The scoring formula is: The first to fifth reward functions are normalized and summed according to their weights to obtain the final reward function.

8. The method according to claim 7, characterized in that, The formula for the weight is: The final reward function formula is: in, Let be the i-th reward function.

9. The method according to claim 1, characterized in that, During the cold start and GRPO phases, the model's input consists of user-inputted questions, and the SQL database includes table name: {table_name}, table creation information: {table_ddl}, and user questions: {user_question}; the model's output is... <think>{think module}<\think>{explanation} <sql> {SQL module}<\sql>; where the think module represents the model's reasoning process; and the SQL module contains the SQL statements generated by the model, used to execute the SQL.< / sql> < / think> 10. The method according to claim 1, characterized in that, In the feedback phase, the feedback data all come from data in the GRPO phase where the SQL statement is correct but execution fails; The SQL database includes the table name: {table_name}, table creation information: {table_ddl}, executed SQL statement: {user_question}, and error information: {error_info}; the model output is {new_SQL}.