Large model prompt project text2sql method based on homogeneous graph recognition

By combining homogeneous graph recognition and the Performance model, similar triples are screened out and accurate SQL statements are generated, which solves the problems of few model parameters and low generation accuracy in the existing technology and realizes efficient SQL generation and verification.

CN120849448APending Publication Date: 2025-10-28HAIKOU PORT COMM TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510909337.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

In existing text2sql tasks based on large models, the limited number of model parameters and training corpora leads to deviations in localized deployment performance, while the accuracy of SQL statements generated by prompting engineering techniques based on large models is not high.

Method used

Using a homogeneous graph recognition algorithm and a performance model, the Schema-linking method is used to filter out table and column names irrelevant to the question, build a relationship graph, screen out similar triples, generate question skeletons and SQL statement skeletons, and combine few-shot learning and large model hint engineering technology to generate accurate SQL statements.

Benefits of technology

The accuracy of SQL statements generated by large models has been improved. The accuracy of the final answer has been verified through similarity matching. Comprehensive screening has been achieved under similar questions, table structures and SQL factors, which has improved the accuracy of generated SQL statements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120849448A_ABST
    Figure CN120849448A_ABST
Patent Text Reader

Abstract

The invention provides a large model prompt project text2sql method based on homogeneous graph recognition, which comprises the following steps: S101, screening table names and column names related to an original problem by adopting a large model Schema-linking method; s102, a large model Schema-linking method is adopted, and table names related to the original problem are screened out; s103, constructing a triple corpus; s104, constructing a relation graph according to the original question related table names, and screening similar schema triples by adopting homogeneous graph identification; s105, generating a problem skeleton, and screening a similar problem skeleton triple set; s106, generating a model by adopting an sql (Structured Query Language), generating a preminary sql query statement, and screening a similar sql skeleton triple set; s107, weighting and scoring are carried out on the screened similar triads by adopting a Performance model, and top-3 similar triads are screened out; s108, according to the top-3 similar triple, forming a feed-shot learning, and generating an accurate sql (structured query language) statement; and S109, translating the sql statements into text descriptions, performing similarity matching on the text descriptions and the original questions, and taking the sql statements with the text descriptions most similar to the original questions as final answers, thereby improving the sql generation accuracy of the large model cues.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of large model technology, and in particular to a text2sql method for large model hinting based on homogeneous graph recognition. Background Technology

[0002] Currently, text2sql tasks based on large model technology can be divided into two main categories. One category is efficient supervised fine-tuning methods for large models. These techniques require less computing resources and can be deployed locally. However, because the number of model parameters and training corpora are relatively small, the overall performance is still inferior compared to using a full-fledged online large model. The other category is suggestion engineering techniques based on large models. These techniques search for similar triples by using similar question skeletons and similar SQL statements to guide the large model to generate accurate SQL statements. However, the overall performance is also inferior, and the accuracy of SQL statements generated by the large model is not high, requiring further optimization. Summary of the Invention

[0003] Therefore, the purpose of this invention is to provide a text2sql method for large model hinting engineering based on homogeneous graph recognition, so as to solve or at least partially solve the above-mentioned problems existing in the prior art.

[0004] To achieve the above objectives, this invention provides a text2sql method for large model hinting projects based on homogeneous graph recognition, the method comprising the following steps:

[0005] S101. Input the original question, and use the schema-linking method of the large model to filter out some table names and column names that are not related to the original question from the database schema information based on the original question. Then, use the filtered table names and column names that are related to the question as input prompt words for in-context learning.

[0006] S102. Input the original problem, use the schema-linking method of the large model to filter out some table names that are not related to the original problem from the database schema information based on the original problem, and filter out the table names that are related to the original problem through the large model.

[0007] S103. Construct a triplet corpus, which is composed of the training set data of text2sql;

[0008] S104. Construct a relational graph from the selected table names related to the original question, and use a homogeneous graph recognition algorithm to select triples ScheT with similar schemas from the triple corpus.

[0009] S105. Generate question skeletons using name matching, value matching, and large model hinting engineering methods respectively, and then filter out the triple set QueryT with similar question skeletons from the triple corpus.

[0010] S106. Using an SQL generation model, generate a preliminary SQL query statement, and then use the preliminary SQL query statement to filter out the top-n triple set PreS with similar SQL skeletons from the triple corpus;

[0011] S107. The similar triplet selected in steps S104, S105 and S106 is weighted and scored using the Performance model. The top-3 similar triplets are then selected using the Performance model. The top-3 similar triplets meet the three factors of similarity in problem, similarity in table structure and similarity in SQL.

[0012] S108. Based on the selected top-3 similar triplets, form the few-shot learning in the prompt words, and based on the table names and column names generated in step S101, guide the large model to generate accurate SQL statements.

[0013] S109. Generate SQL statements using an odd number of large models, and use large model hinting engineering techniques to translate the SQL statements into text descriptions. Perform similarity matching between the text descriptions and the original question, and finally take the SQL statement with the most similar text descriptions to the original question as the final answer.

[0014] Furthermore, step S101 specifically includes the following steps:

[0015] S11. Use the LoRa algorithm to fine-tune the large model, allowing it to generate table and column names related to the original problem based on the input problem and database schema information, as shown below:

[0016]

[0017] in, The optimal model found through fine-tuning training using the LoRa algorithm. Make the probability under this model The maximum, T and C are the sets of relevant table names and the sets of relevant column names generated based on the original problem q and database schema information D of the current input large model, respectively;

[0018] S12. The Beam-search method is used to output different candidate table names and column names for the large model. When generating a partial sequence at each step, Beam-search retains K optimal candidate sequences. By gradually expanding the optimal candidate sequences, a complete output sequence is generated. Finally, the K optimal candidate sequences are deduplicated and merged as a stage output of the large model's schema-linking, as shown below:

[0019]

[0020] Among them, B t Let Top-k be the set of candidate sequences at time step k, and Top-k be the operator. This is the output of the i-th sequence at time step 1 to t. Given a large model input X, candidate sequences The generation probability, where X is the original question and database schema information in the text2sql task;

[0021] S13. Use the n-gram algorithm to match the original problem with the table and column names in the database, and output the table and column names mentioned in the original problem.

[0022] S14. Merge the outputs of steps S12 and S13 to form a schema-linking result set.

[0023] Further, step S102 includes:

[0024] The LoRa algorithm is used to fine-tune the training of a large model, enabling it to generate table names related to the input question and database schema information, forming a set of tables, as shown below:

[0025]

[0026] in, The optimal model found through fine-tuning training using the LoRa algorithm. Make the probability under this model Maximum, T is the set of related table names generated based on the original problem q of the current input large model and the database schema information C.

[0027] Furthermore, the triplet corpus is a {original question, ground truth schema, ground truth sql} triplet corpus. The ground truth schema in the triplet corpus is constructed into a relation graph. The table names selected in step S102 are constructed into entity points, and the foreign key relationships between tables are constructed into edges.

[0028] Furthermore, step S104 specifically includes the following steps:

[0029] S41. Construct a relational graph from the table names related to the original problem selected in step S102, with table names as entity points and foreign key relationships between tables as edges.

[0030] S42. Using a homogeneous graph identification algorithm, compare the relation graph constructed in step S41 with the ground truth schema relation graph of each record in the triplet corpus, and select the set of triples ScheT that has the same structure as the relation graph constructed in step S41, as shown below:

[0031] ScheT={g∣g∈G iso}

[0032] G iso ={g i ∣f(g,g i )=1}

[0033] Where g is the relational graph constructed in step S31, G iso To determine whether two spectra are homogeneous, g i Let f(g,g) be the relational graph composed of schemas in the i-th sample of the ternary corpus. i To identify graph g and graph g' using a homogeneous graph identification algorithm. i Is it an isomorphic graph?

[0034] Furthermore, step S105 specifically includes the following steps:

[0035] S51. Generate the problem skeleton using name matching, value matching, and large model hinting engineering methods respectively;

[0036] S52. Based on the generated question skeleton, select a set of triples QueryT with similar question skeletons from the triple corpus, represented as follows:

[0037] QueryT = {(query i ∣f(query,query i ))}

[0038] Here, query is the skeleton of the original problem input to the large model after de-semanticization. iThe question skeleton of the i-th sample in the triplet corpus after de-semanticization is defined as follows: the de-semanticization process involves replacing the fields in the sentence related to table names and column names in the database with the [mask] character, leaving only the question skeleton. f(query, query) i The similarity score is used to measure the similarity between the skeletons of the two questions.

[0039] Furthermore, the generation of the problem skeleton using name matching, value matching, and large model hinting engineering methods respectively includes the following steps:

[0040] S5.1. Use name matching to match each token in the question with the table name and column name in the database. If the names are the same, the token is replaced with [mask].

[0041] S5.2. Use value matching to match each token in the problem with the value of all records in the database. If a value with the same value as the token exists in the database, then the token is replaced with [mask].

[0042] S5.3. Employ large model prompting engineering technology, by inputting the problem, database schema information, and output sample data, the large model performs de-semanticization processing on the problem;

[0043] S5.4. Merge the results of steps S5.1-S5.3 to form the final problem skeleton after de-semanticization.

[0044] Furthermore, step S106 specifically includes the following steps:

[0045] S61. Use the SQL generation model to generate preliminary SQL query statements;

[0046] S62. Using the pattern linking method, filter the table names, column names and query values ​​in the SQL statement generated in step S51 to generate the preliminary SQL statement skeleton;

[0047] S63. Based on the preliminary SQL statement skeleton generated in step S62, select the top-n triples set PreS with similar SQL skeletons from the triple corpus, as shown below:

[0048] PreS = {(preSQL) i |f(preSQL,preSQL) i ))}

[0049] Wherein, preSQL is the preliminary SQL statement skeleton generated in step S62. i Let f(preSQL, preSQL) be the SQL statement skeleton of the i-th sample in the triplet corpus. i The similarity score is used to measure the similarity between the skeletons of two SQL statements.

[0050] Furthermore, step S107 specifically includes the following steps:

[0051] S71. Based on the ScheT set, find the intersection of ScheT and QueryT, and the intersection of ScheT and PreS. Then, combine the two intersections to form a new set FinC. If ScheT and QueryT have no intersection, take QueryT; if ScheT and PreS have no intersection, take PreS. Finally, QueryT and PreS form the FinC set, as shown below:

[0052]

[0053] S72. The similar triplet set FinC selected in step S71 is weighted and scored using the Performance model. Then, the top-3 similar triplets are finally selected using the Performance model, and the comprehensive similarity is calculated based on the selected top-3 similar triplets, as shown below:

[0054] S(t)=ω sch *S sch (t)+ω q *S q (t)+ω sql *S sql (t)

[0055] Where S(t) is the comprehensive similarity, S sch (t) represents the similarity score of ScheT, S q (t) represents the similarity score of QueryT, S sql (t) represents the similarity score of PreS, ω sch ω q ω sql S sch (t), S q (t), S sql The corresponding weight value of (t).

[0056] Furthermore, step 108 includes:

[0057] By optimizing ωsch ω q ω sql The weights, based on the selected top-3 similar triples and the overall similarity, form few-shot learning prompt words, and based on the table names and column names generated in step S101, guide the large model to generate accurate SQL statements.

[0058] Compared with the prior art, the beneficial effects of the present invention are:

[0059] This invention proposes a text2sql method for large-scale model hinting engineering based on homogeneous graph recognition. Based on the homogeneous graph recognition algorithm and a newly proposed performance model, it implements a novel few-shot learning generation method. It achieves a new comprehensive screening method for sample data in the triplet corpus under three factors: similar questions, similar table structures, and similar SQL statements, thereby improving the accuracy of SQL statements generated by the large model. Finally, the accuracy is verified by translating the SQL statements into text descriptions and then performing similarity matching between the text descriptions and the original questions. Attached Figure Description

[0060] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only preferred embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0061] Figure 1 This is a schematic diagram of the text2sql method for large model hinting based on homogeneous graph recognition, provided in an embodiment of the present invention. Detailed Implementation

[0062] The principles and features of the present invention are described below with reference to the accompanying drawings. The listed embodiments are only used to explain the present invention and are not intended to limit the scope of the present invention.

[0063] Reference Figure 1 This embodiment provides a text2sql method for large model hinting projects based on homogeneous graph recognition. The method includes the following steps:

[0064] S101. Input the original problem. Using the schema-linking method for efficient fine-tuning of large models, filter out some table names and column names that are irrelevant to the original problem from a large amount of database schema information. Use the filtered table names and column names that are more relevant to the problem as input prompts for in-context learning. The specific steps include:

[0065] S11. Use the LoRa efficient fine-tuning algorithm to fine-tune the large model, allowing the large model to generate table names and column names that are highly relevant to the original problem based on the input original problem and the database schema information, as shown below:

[0066]

[0067] in, To find an optimal model through fine-tuning training using the LoRa algorithm. Make the probability under this model The maximum, T and C are the sets of relevant table names and the sets of relevant column names generated based on the original problem q and database schema information D of the current input large model, respectively;

[0068] S12. The Beam-search method is used to output different candidate table names and column names for the large model. When generating a partial sequence at each step, Beam-search retains K optimal candidate sequences. By gradually expanding the optimal candidate sequences, a complete output sequence is generated. Finally, the K optimal candidate sequences are deduplicated and merged as a stage output of the large model's schema-linking, as shown below:

[0069]

[0070] Among them, B t Let be the set of candidate sequences at time step t. Top-k is an operator used to select the k most likely sequences from all possible candidate sequences. Each candidate sequence consists of the sequence itself and the probability of generating the sequence. The Top-k operation only retains those candidate sequences with the highest generation probability, so that these most promising sequences can be expanded in subsequent time steps. This is the output of the i-th sequence at time step 1 to t. Given a large model input X, candidate sequences The generation probability, where X is the original question and database schema information in the text2sql task;

[0071] S13. Use the n-gram algorithm to match the original problem with the table and column names in the database, and output the table and column names mentioned in the original problem.

[0072] S14. Merge the outputs of steps S12 and S13 to form a schema-linking result set.

[0073] S102. Input the original problem. Using the schema-linking method of the large model, filter out some table names that are irrelevant to the original problem from a large amount of database schema information, and then use the large model to select the table names most relevant to the original problem. Specifically, this includes:

[0074] The LoRa algorithm is used to fine-tune the training of a large model, enabling it to generate table names related to the input original question and database schema information, forming a set of tables. Experiments have verified that the tables related to the original question selected by the large model after fine-tuning have relatively high accuracy. Compared to step S101, which requires generating both table names and column names related to the question, this step achieves higher accuracy, as shown below:

[0075]

[0076] in, The optimal model found through fine-tuning training using the LoRa algorithm. Make the probability under this model Maximum, T is the set of related table names generated based on the original problem q of the current input large model and the database schema information C.

[0077] S103. Construct a triplet corpus, which consists of the training set data of text2sql, specifically including:

[0078] The triplet corpus is a {original question, ground truth schema, ground truth sql} triplet corpus. The ground truth schema in the triplet corpus is used to construct a relation graph. The table names selected in step S102 are used to construct entity points, and the foreign key relationships between tables are used to construct edges.

[0079] S104. Construct a relational graph from the selected table names that are highly relevant to the original question, and use a homogeneous graph recognition algorithm to select triples ScheT with similar schemas from the triple corpus. This includes the following steps:

[0080] S41. Construct a relational graph from the table names that are highly relevant to the original problem selected in step S102, with table names as entity points and foreign key relationships between tables as edges.

[0081] S42. Using a homogeneous graph identification algorithm, the relation graph constructed in step S31 is compared with the ground truth schema relation graph of each record in the triplet corpus. The set of triples ScheT, which has the same structure as the relation graph constructed in step S31, is selected and represented as follows:

[0082] ScheT={g∣g∈G iso}

[0083] G iso ={g i ∣f(g,g i )=1}

[0084] Where g is the relational graph constructed in step S31, G iso To determine whether two spectra are homogeneous, g i Let f(g,g) be the relational graph composed of schemas in the i-th sample of the ternary corpus. i To identify graph g and graph g' using a homogeneous graph identification algorithm. i Is it an isomorphic graph, ScheT={g∣g∈G}? iso} represents all graphs in the samples of the ternary corpus that have the same structure as the relational graph constructed in step S31.

[0085] S105. Generate question skeletons using name matching, value matching, and large model hinting engineering methods respectively. Then, select a set of triples QueryT with similar question skeletons from the triple corpus. This includes the following steps:

[0086] S51. Use name matching, value matching and large model hinting engineering methods to generate the problem skeleton. For example, the original sentence is "Find the names of all employees with a salary higher than 5000". After de-semanticization, the sentence is "Find all [mask] higher than [mask] of [mask]".

[0087] S52. Based on the generated question skeleton, select the top-n triplet set QueryT with similar question skeletons from the triplet corpus, represented as follows:

[0088] QueryT = {(query i ∣f(query,query i ))}

[0089] Here, query is the skeleton of the original problem input to the large model after de-semanticization. iThe question skeleton of the i-th sample in the triplet corpus after de-semanticization is defined as follows: the de-semanticization process involves replacing the fields in the sentence related to table names and column names in the database with the [mask] character, leaving only the question skeleton. f(query, query) i The similarity score is used to measure the similarity between the skeletons of the two questions.

[0090] The process of generating the problem skeleton using name matching, value matching, and large model hinting engineering methods includes the following steps:

[0091] S5.1. Use name matching to match each token in the question with the table name and column name in the database. If the names are the same, the token is replaced with [mask].

[0092] S5.2. Use value matching to match each token in the problem with the value of all records in the database. If a value with the same value as the token exists in the database, then the token is replaced with [mask].

[0093] S5.3. Employ large model prompting engineering technology, by inputting the problem, database schema information, and output sample data, the large model performs de-semanticization processing on the problem;

[0094] S5.4. Merge the results of steps S5.1-S5.3 to form the final problem skeleton after de-semanticization.

[0095] S106. Using an SQL generation model, generate a preliminary SQL query statement. Then, use the preliminary SQL query statement to filter out the top-n triplet set PreS with similar SQL skeletons from the triplet corpus. This includes the following steps:

[0096] S61. Use SQL generation models other than the large model hinting engineering method (such as the RSDSQL model) to generate preliminary SQL query statements;

[0097] S62. Using the pattern linking method, filter the table names, column names, and query values ​​in the SQL statement generated in step S51 to generate a preliminary SQL statement skeleton. For example, [select name from singer where age>18], the SQL statement skeleton after pattern linking is: [select from where].

[0098] S63. Based on the preliminary SQL statement skeleton generated in step S62, select the top-n triples set PreS with similar SQL skeletons from the triple corpus, as shown below:

[0099] PreS = {(preSQL) i |f(preSQL,preSQL) i ))}

[0100] Wherein, preSQL is the preliminary SQL statement skeleton generated in step S62. i Let f(preSQL, preSQL) be the SQL statement skeleton of the i-th sample in the triplet corpus. i The similarity score is used to measure the similarity between the skeletons of two SQL statements.

[0101] S107. The similar triples selected in steps S104, S105, and S106 are weighted and scored using a performance model. Then, the top-3 similar triples are finally selected using the performance model. These top-3 similar triples simultaneously meet three criteria: similarity in problem type, similarity in table structure, and similarity in SQL. Specifically, this includes the following steps:

[0102] S71. Based on the ScheT set, find the intersection of ScheT and QueryT, and the intersection of ScheT and PreS. Then, combine the two intersections to form a new set FinC. If ScheT and QueryT have no intersection, take QueryT; if ScheT and PreS have no intersection, take PreS. Finally, QueryT and PreS form the FinC set, as shown below:

[0103]

[0104] S72. The similar triplet set FinC selected in step S71 is weighted and scored using the Performance model. Then, the top-3 similar triplets are finally selected using the Performance model, and the comprehensive similarity is calculated based on the selected top-3 similar triplets, as shown below:

[0105] Let T be the sample dataset in the triplet corpus, and for the sample data t∈T in the triplet set FinC generated in step S61;

[0106] S(t)=ω sch *S sch(t)+ω q *S q (t)+ω sql *S sql (t)

[0107] Where S(t) is the comprehensive similarity, S sch (t) represents the similarity score between ScheT and other languages, with a default value of 1 for all values; S q (t) represents the similarity score of QueryT, S sql (t) represents the similarity score of PreS, ω sch ω q ω sql S sch (t), S q (t), S sql The corresponding weight value of (t) ranges from [0, 1].

[0108] The objective function of the performance model is expressed as follows:

[0109]

[0110] Among them, T θ ∈T is a subset of sample data selected based on the comprehensive similarity S(t).

[0111] S108. Based on the selected top-3 similar triplets, construct the few-shot learning in the prompt words to guide the large model to generate accurate SQL statements, specifically including:

[0112] By optimizing ω sch ω q ω sql Weights are assigned based on the top-3 similar triplets selected and the overall similarity, forming few-shot learning prompts. These few-shot learning prompts maximize the accuracy of the large model's generated SQL statements (sql_exec_match(T)). θ Based on the table names and column names generated in step S101, the large model is guided to generate accurate SQL statements. The few-shot learning can be understood as an implementation of in-context learning.

[0113] Where the weight ω sch v q ω sql The learning process employs Bayesian optimization.

[0114] S109. Generate SQL statements using an odd number of large models, and use large model hinting engineering techniques to translate the SQL statements into text descriptions. Perform similarity matching between the text descriptions and the original question, and finally take the SQL statement with the most similar text descriptions to the original question as the final answer.

[0115] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A text2sql method for large-scale model hinting based on homogeneous graph recognition, characterized in that, The method includes the following steps: S101. Input the original question, and use the schema-linking method of the large model to filter out some table names and column names that are not related to the original question from the database schema information based on the original question. Then, use the filtered table names and column names that are related to the question as input prompt words for in-context learning. S102. Input the original problem, use the schema-linking method of the large model to filter out some table names that are not related to the original problem from the database schema information based on the original problem, and filter out the table names that are related to the original problem through the large model. S103. Construct a triplet corpus, which is composed of the training set data of text2sql; S104. Construct a relational graph from the selected table names related to the original question, and use a homogeneous graph recognition algorithm to select triples ScheT with similar schemas from the triple corpus. S105. Generate question skeletons using name matching, value matching, and large model hinting engineering methods respectively, and then filter out the triple set QueryT with similar question skeletons from the triple corpus. S106. Using an SQL generation model, generate preliminary SQL query statements, and then use the preliminary SQL query statements to filter out the top-n triplet set PreS with similar SQL skeletons from the triplet corpus; S107. The similar triplet selected in steps S104, S105 and S106 is weighted and scored using the Performance model. The top-3 similar triplets are then selected using the Performance model. The top-3 similar triplets meet the three factors of similarity in problem, similarity in table structure and similarity in SQL. S108. Based on the selected top-3 similar triplets, form the few-shot learning in the prompt words, and based on the table names and column names generated in step S101, guide the large model to generate accurate SQL statements. S109. Generate SQL statements using an odd number of large models, and use large model hinting engineering techniques to translate the SQL statements into text descriptions. Perform similarity matching between the text descriptions and the original question, and finally take the SQL statement with the most similar text descriptions to the original question as the final answer.

2. The text2sql method for large model hinting projects based on homogeneous graph recognition according to claim 1, wherein step S101 specifically includes the following steps: S11. Use the LoRa algorithm to fine-tune the large model, allowing it to generate table and column names related to the original problem based on the input problem and database schema information, as shown below: in, The optimal model found through fine-tuning training using the LoRa algorithm. Make the probability under this model The maximum, T and C are the sets of relevant table names and the sets of relevant column names generated based on the original problem q and database schema information D of the current input large model, respectively; S12. The Beam-search method is used to output different candidate table names and column names for the large model. When generating a partial sequence at each step, Beam-search retains K optimal candidate sequences. By gradually expanding the optimal candidate sequences, a complete output sequence is generated. Finally, the K optimal candidate sequences are deduplicated and merged as a stage output of the large model's schema-linking, as shown below: Among them, B t Let be the set of candidate sequences at time step t, and Top-k be the operators. This is the output of the i-th sequence at time step 1 to t. Given a large model input X, candidate sequences The generation probability, where X is the original question and database schema information in the text2sql task; S13. Use the n-gram algorithm to match the original problem with the table and column names in the database, and output the table and column names mentioned in the original problem. S14. Merge the outputs of steps S12 and S13 to form a schema-linking result set.

3. The large model hinting engineering text2sql method based on homogeneous graph recognition according to claim 1, characterized in that, Step S102 includes: The LoRa algorithm is used to fine-tune the training of a large model, enabling it to generate table names related to the input question and database schema information, forming a set of tables, as shown below: in, The optimal model found through fine-tuning training using the LoRa algorithm. Make the probability under this model Maximum, T is the set of related table names generated based on the original problem q of the current input large model and the database schema information C.

4. The text2sql method for large model hinting projects based on homogeneous graph recognition according to claim 3, characterized in that, The triplet corpus is a {original question, ground truth schema, ground truth sql} triplet corpus. The ground truth schema in the triplet corpus is used to construct a relation graph. The table names selected in step S102 are used to construct entity points, and the foreign key relationships between tables are used to construct edges.

5. The large model hinting engineering text2sql method based on homogeneous graph recognition according to claim 4, characterized in that, Step S104 specifically includes the following steps: S41. Construct a relational graph from the table names related to the original problem selected in step S102, with table names as entity points and foreign key relationships between tables as edges. S42. Using a homogeneous graph identification algorithm, compare the relation graph constructed in step S41 with the ground truth schema relation graph of each record in the triplet corpus, and select the set of triples ScheT that has the same structure as the relation graph constructed in step S41, as shown below: ScheT={g∣g∈G iso } G iso ={g i ∣f(g,g i )=1} Where g is the relational graph constructed in step S31, G iso To determine whether two spectra are homogeneous, g i Let f(g,g) be the relational graph composed of schemas in the i-th sample of the ternary corpus. i To identify graph g and graph g' using a homogeneous graph identification algorithm. i Is it an isomorphic graph? 6. The large model hinting engineering text2sql method based on homogeneous graph recognition according to claim 1, characterized in that, Step S105 specifically includes the following steps: S51. Generate the problem skeleton using name matching, value matching, and large model hinting engineering methods respectively; S52. Based on the generated question skeleton, select a set of triples QueryT with similar question skeletons from the triple corpus, represented as follows: QueryT={(query i ∣f(query,query i ))} Here, query is the skeleton of the original problem input to the large model after de-semanticization. i The question skeleton of the i-th sample in the triplet corpus after de-semanticization is defined as follows: the de-semanticization process involves replacing the fields in the sentence related to table names and column names in the database with the [mask] character, leaving only the question skeleton. f(query, query) i The similarity score is used to measure the similarity between the skeletons of the two questions.

7. The large model hinting engineering text2sql method based on homogeneous graph recognition according to claim 6, characterized in that, The process of generating the problem skeleton using name matching, value matching, and large model hinting engineering methods includes the following steps: S5.

1. Use name matching to match each token in the question with the table name and column name in the database. If the names are the same, the token is replaced with [mask]. S5.

2. Use value matching to match each token in the problem with the value of all records in the database. If a value with the same value as the token exists in the database, then the token is replaced with [mask]. S5.

3. Employ large model prompting engineering technology, by inputting the problem, database schema information, and output sample data, the large model performs de-semanticization processing on the problem; S5.

4. Merge the results of steps S5.1-S5.3 to form the final problem skeleton after de-semanticization.

8. The large model hinting engineering text2sql method based on homogeneous graph recognition according to claim 1, characterized in that, Step S106 specifically includes the following steps: S61. Use the SQL generation model to generate preliminary SQL query statements; S62. Using the pattern linking method, filter the table names, column names and query values ​​in the SQL statement generated in step S51 to generate the preliminary SQL statement skeleton; S63. Based on the preliminary SQL statement skeleton generated in step S62, select the top-n triples set PreS with similar SQL skeletons from the triple corpus, as shown below: PreS={(preSQL i ∣f(preSQL,preSQL i ))} Wherein, preSQL is the preliminary SQL statement skeleton generated in step S62. i Let f(preSQL, preSQL) be the SQL statement skeleton of the i-th sample in the triplet corpus. i The similarity score is used to measure the similarity between the skeletons of two SQL statements.

9. The text2sql method for large model hinting projects based on homogeneous graph recognition according to claim 1, characterized in that, Step S107 specifically includes the following steps: S71. Based on the ScheT set, find the intersection of ScheT and QueryT, and the intersection of ScheT and PreS. Then, combine the two intersections to form a new set FinC. If ScheT and QueryT have no intersection, take QueryT; if ScheT and PreS have no intersection, take PreS. Finally, QueryT and PreS form the FinC set, as shown below: S72. The similar triplet set FinC selected in step S71 is weighted and scored using the Performance model. Then, the top-3 similar triplets are finally selected using the Performance model, and the comprehensive similarity is calculated based on the selected top-3 similar triplets, as shown below: S(t)=ω sch *S sch (t)+ω q *S q (t)+ω sql *S sql (t) Where S(t) is the comprehensive similarity, S sch (t) represents the similarity score of ScheT, S q (t) represents the similarity score of QueryT, S sql (t) represents the similarity score of PreS, ω sch ω q ω sql S sch (t), S q (t), S sql The corresponding weight value of (t).

10. The text2sql method for large model hinting projects based on homogeneous graph recognition according to claim 9, characterized in that, Step S108 includes: By optimizing ω sch ω q ω sql The weights, based on the selected top-3 similar triples and the overall similarity, form few-shot learning prompt words, and based on the table names and column names generated in step S101, guide the large model to generate accurate SQL statements.