Text2SQL (Structured Query Language) conversion method in specific scene based on large model
Through the large model-based Text2SQL conversion method, the complexity and professionalism of case clue queries in legal institutions were solved, efficient and accurate query results were achieved, and the efficiency and quality of legal work were improved.
Patent Information
- Application Number
- CN202510664778.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-09-19
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The case clue query of legal institutions is highly professional and has complex data structure. The existing general Text2SQL technology is difficult to meet the accuracy and efficiency requirements of the query, resulting in time-consuming and error-prone queries.
A Text2SQL conversion method based on a large model is used in specific scenarios, including data preprocessing and knowledge base construction in the legal field, fine-tuning the large model, introducing the concept of generative adversarial networks, executing SQL queries and optimizing results, evaluating query results and continuous learning.
It has greatly improved the efficiency of clue queries, lowered the technical threshold, and enabled non-professionals to quickly obtain accurate results. It has significantly improved query accuracy and legal work efficiency, and ensured the timeliness and quality of case handling.
Smart Images

Figure CN120670458A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of querying and screening case clues, and in particular to a Text2SQL conversion method in a specific scenario based on a large model. Background Art
[0002] In the daily work of legal institutions, searching and screening case leads is a crucial and tedious task. With the development of information technology, legal institutions have accumulated a large amount of data related to various cases. This data is stored in complex database systems, including multiple data tables and fields such as basic case information, party information, evidence materials, legal documents, etc. Traditional query methods require staff to have professional database query knowledge and be able to write accurate SQL query statements to obtain the required clue information. However, staff are often more focused on legal expertise and case investigation practices, and have limited mastery of complex database query syntax. This results in a considerable amount of time and energy spent on searching for leads, and it is easy to fail to obtain accurate and comprehensive clues due to SQL writing errors, seriously affecting the efficiency and quality of case handling.
[0003] While existing general-purpose Text2SQL technologies have addressed the natural language to SQL conversion problem to some extent, they still face numerous shortcomings when applied to the specific legal field. Firstly, legal agency data is highly specialized, has unique formatting specifications, complex semantics, and involves a large number of legal terminology and procedurally specific expressions. General-purpose Text2SQL methods struggle to accurately understand and process natural language queries in these specific fields. Secondly, legal agency databases are complex, with numerous and deeply layered relationships between tables. Generating effective SQL queries requires a deep understanding and precise parsing of the database structure. Existing technologies are limited in this regard and cannot meet the accuracy and efficiency requirements of legal agencies for lead queries. Therefore, we propose a Text2SQL conversion method based on large models for specific scenarios. Summary of the Invention
[0004] This application provides a Text2SQL conversion method for specific scenarios based on large models to solve the above-mentioned problems.
[0005] This application provides a Text2SQL conversion method for specific scenarios based on large models, including the following methods:
[0006] S1. Data preprocessing and knowledge base construction in the field of legal institutions;
[0007] S2, fine-tune the large model;
[0008] S3, introduce the concept of generative adversarial network;
[0009] S4. Execute SQL query statements and optimize query results;
[0010] S5. Evaluate query results and continue learning.
[0011] Preferably, the legal institution field data preprocessing and knowledge base construction includes the following steps:
[0012] S11. Data Collection: Collect a wide range of natural language queries related to case clues and their corresponding SQL query statements from multiple key data sources within the legal agency. n}, n = 1, 2, ... k;
[0013] S12. Obtain database table information: Completely obtain the metadata set of the legal agency database, E = {e1, e2, ...e n}, n = 1, 2, ... k; including detailed structure definitions, field information, data types, primary key and foreign key relationships, and index information of all tables;
[0014] S13. Develop data specifications: Perform comprehensive and detailed cleaning operations on all collected data C.
[0015]
[0016] Among them: Clean represents the data cleaning function;
[0017] S14. Construct vector database: Use the extracted semantic feature vectors to construct a vector database, and convert each legal term t∈C into a corresponding vector representation through the word vector model. The mapping function corresponding to the word vector model is:
[0018]
[0019] in: Where C represents all data.
[0020] Preferably, the fine-tuning of the large model comprises the following steps:
[0021] S21 Corpus Design: Design a legal terminology corpus by collecting data and professionals’ experience;
[0022] S22 Select pre-trained model base: Select pre-trained models that are suitable for processing complex natural language tasks and have excellent scalability, and design a pre-trained model set M = {m1, m2, ..., m k}, where each m l Represents a pre-training model, for each m l , define the evaluation index set I = {i1, i2, ..., ij}, the comprehensive score of the model is S l :
[0023] S l =ScoreModel(m l , i)
[0024] Where: l = 1, 2, ..., k;
[0025] S23 fine-tunes the details of the large model: By fine-tuning the large model, the model can simultaneously focus on different semantic units in natural language and the database structure elements represented by the related semantic vectors in the vector database, and accurately calculate the association weights between them, accurately constructing a complex and accurate semantic mapping relationship between natural language queries and database structures.
[0026] Preferably, the introduction of the generative adversarial network concept includes further enhancing the robustness of the model and its adaptability to complex and changing environments through adversarial training.
[0027] Preferably, further enhancing the robustness of the model and its adaptability to complex and changing environments through adversarial training comprises the following steps:
[0028] S31. Build the generator G and set the generator's loss function.
[0029] S32. Build the discriminator D and design the loss function of the discriminator;
[0030] S33. Adversarial training allows the generator and discriminator to compete with each other during the training process.
[0031] Preferably, the loss function of the generator is:
[0032]
[0033] Where: z is the input noise of the generator, P Z (Z) is the distribution of noise;
[0034] The loss function of the discriminator is:
[0035]
[0036] Where: x is a real SQL statement, P data (x) is the distribution of real SQL statements.
[0037] Preferably, executing the SQL query statement and optimizing the query results includes the following steps:
[0038] S41. The staff of the Procuratorate inputs the natural language query into the model to generate candidate queries;
[0039] S42. Analyze database indexes and optimize query statements;
[0040] S43. Execute the optimized SQL query statement on the database.
[0041] Preferably, evaluating the query results and continuously learning comprises the following steps:
[0042] S51. Obtain user satisfaction and collect feedback;
[0043] S52, save the valid query;
[0044] S53. Match similar questions.
[0045] The above technical solution provided by the embodiment of the present application has the following advantages compared with the prior art:
[0046] Compared with the prior art, the overall structure provided by the embodiment of the present application achieves:
[0047] Significantly improves lead retrieval efficiency: By allowing staff to directly use natural language to query leads, the tedious process of writing complex SQL statements is completely eliminated, significantly reducing the time required to search for leads. Compared with traditional query methods based on manually written SQL, query efficiency can be increased by several times or even dozens of times, allowing staff to quickly obtain the required lead information, significantly accelerating the handling of cases, and greatly improving the overall efficiency of legal work, providing a strong guarantee for the timely and accurate fight against crime and the maintenance of social fairness and justice.
[0048] Effectively lowering the technical threshold for staff: Even staff without specialized database knowledge can easily perform complex lead searches. Simply by clearly describing their query requirements in natural language, they can quickly obtain accurate results. This significantly reduces the learning burden for staff, allowing them to focus more time and energy on core tasks such as legal analysis, investigation and evidence collection, and evidence review. This allows them to fully leverage their professional strengths, improve the professionalism and quality of legal work, promote the deep integration of legal institutions' information development and business operations, and advance the professional development of legal work.
[0049] Significantly improve the accuracy of clue queries: Leveraging the large model's powerful ability to understand natural language, the introduction of advanced concepts such as generative adversarial networks (GANs), and optimized training for legal institutions, the system can accurately analyze staff members' query intent and generate highly accurate SQL query statements, effectively avoiding inaccurate query results caused by SQL writing errors, semantic understanding deviations, or insufficient understanding of database structure. At the same time, through rigorous result verification and intelligent feedback and correction mechanisms, the performance of the model is continuously optimized, further improving the accuracy and reliability of clue queries, ensuring that staff can obtain comprehensive, accurate, and reliable case clues, providing a solid basis for the correct handling of cases and legal decision-making, and effectively improving the quality and credibility of legal work. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0051] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0052] Figure 1 It is the overall principle diagram of the present invention;
[0053] Figure 2 Schematic diagram of data preprocessing and knowledge base construction in the legal affairs field of the present invention;
[0054] Figure 3 This is a schematic diagram of the fine-tuning large model of the present invention;
[0055] Figure 4 This is a schematic diagram of the principle of introducing the concept of generative adversarial network in the present invention;
[0056] Figure 5 A schematic diagram of executing SQL query statements and optimizing query results of the present invention;
[0057] Figure 6 Schematic diagram of the principles of evaluating query results and continuous learning of the present invention. DETAILED DESCRIPTION
[0058] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0059] The various embodiments of the present application may be presented in the form of a range. It should be understood that the description in the form of a range is merely for convenience and brevity and should not be construed as a rigid limitation on the scope of the present application. Therefore, it should be considered that the range description has specifically disclosed all possible sub-ranges and single numerical values within the range. For example, it should be considered that the range description from 1 to 6 has specifically disclosed sub-ranges, such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6, etc., as well as single numbers within the range, such as 1, 2, 3, 4, 5 and 6, regardless of the range. In addition, whenever a numerical range is indicated in this application, it is intended to include any quoted number (fraction or integer) within the indicated range. Unless otherwise specified, the various raw materials, reagents, instruments and equipment used in this application are all commercially available or can be prepared using existing equipment.
[0060] In this application, unless otherwise specified, the directional words used, such as "upper" and "lower", specifically refer to the directions of the drawings in the accompanying drawings. In addition, in this application, the terms "including", "comprising", etc. mean "including but not limited to". In this application, relational terms such as "first" and "second" are merely used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. In this application, "and / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. Wherein A and B can be singular or plural. In this application, "at least one" means one or more, and "plurality" means two or more. "At least one", "at least one of the following" or similar expressions refer to any combination of these items, including any combination of singular or plural items. For example, "at least one of a, b, or c" or "at least one of a, b and c" can both mean: a, b, c, ab, i.e. a and b, ac, bc or abc, where a, b, c can be single or multiple.
[0061] like Figure 1 and Figure 6As shown: The embodiment of the present application provides a Text2SQL conversion method for a specific scenario based on a large model, including the following methods:
[0062] S1. Data preprocessing and knowledge base construction in the field of legal institutions;
[0063] S2, fine-tune the large model;
[0064] S3, introduce the concept of generative adversarial network;
[0065] S4. Execute SQL query statements and optimize query results;
[0066] S5. Evaluate query results and continue learning.
[0067] The data preprocessing and knowledge base construction in the legal institution field include the following steps:
[0068] S11. Data Collection: Collect a wide range of natural language queries related to case clues and their corresponding SQL query statements from multiple key data sources within the legal agency. n}, n = 1, 2, ... k;
[0069] S12. Obtain database table information: Completely obtain the metadata set of the legal agency database, E = {e1, e2, ...e n}, n = 1, 2, ... k; including detailed structure definitions, field information, data types, primary key and foreign key relationships, and index information of all tables;
[0070] S13. Develop data specifications: Perform comprehensive and detailed cleaning operations on all collected data C.
[0071]
[0072] Among them: Clean represents the data cleaning function;
[0073] Specifically: The data cleaning function is used to correct common spelling and grammatical errors, and convert legal terms and common abbreviations into standard form to ensure high consistency and accuracy of the data.
[0074] S14. Construct vector database: Use the extracted semantic feature vectors to construct a vector database, and convert each legal term t∈C into a corresponding vector representation through the word vector model. The mapping function corresponding to the word vector model is:
[0075]
[0076] in: Where C represents all data.
[0077] The fine-tuning of the large model includes the following steps:
[0078] S21 Corpus Design: Design a legal terminology corpus by collecting data and professionals’ experience;
[0079] S22 Select pre-trained model base: Select pre-trained models that are suitable for processing complex natural language tasks and have excellent scalability, and design a pre-trained model set M = {m1, m2, ..., m k}, where each m l Represents a pre-training model, for each m l , define the evaluation index set I = {i1, i2, ..., i j}, the comprehensive score of the model is S l :
[0080] S l =ScoreModel(m l , i)
[0081] Where: l = 1, 2, ..., k;
[0082] S23 fine-tunes the details of the large model: By fine-tuning the large model, the model can simultaneously focus on different semantic units in natural language and the database structure elements represented by the related semantic vectors in the vector database, and accurately calculate the association weights between them, accurately constructing a complex and accurate semantic mapping relationship between natural language queries and database structures.
[0083] The introduction of the concept of generative adversarial networks includes further enhancing the robustness of the model and its adaptability to complex and changing environments through adversarial training.
[0084] The adversarial training to further enhance the model's robustness and adaptability to complex and changing environments includes the following steps:
[0085] S31. Build the generator G and set the generator's loss function.
[0086] S32. Build the discriminator D and design the loss function of the discriminator;
[0087] S33. Adversarial training allows the generator and discriminator to compete with each other during the training process.
[0088] The loss function of the generator is:
[0089]
[0090] Where: z is the input noise of the generator, P Z (Z) is the distribution of noise;
[0091] Specifically: Each time the generator parameters are updated, the gradient descent algorithm is used to minimize its loss function. G , the update formula is based on the gradient after the chain rule is derived, as follows:
[0092]
[0093] in is the learning rate, which determines the step size of each parameter update. is the loss function L G Gradients with respect to the generator parameters.
[0094] The loss function of the discriminator is:
[0095]
[0096] Where: x is a real SQL statement, P data (x) is the distribution of real SQL statements.
[0097] Specifically: For the discriminator parameter θ D , the update formula is based on the gradient after the chain rule is derived, as follows:
[0098]
[0099] in is the learning rate, which determines the step size of each parameter update. is the loss function L D Gradients with respect to the generator parameters.
[0100] Executing the SQL query statement and optimizing the query results includes the following steps:
[0101] S41. The staff of the Procuratorate inputs the natural language query into the model to generate candidate queries;
[0102] S42. Analyze database indexes and optimize query statements;
[0103] S43. Execute the optimized SQL query statement on the database.
[0104] Specifically: executing SQL query statements and optimizing query results include:
[0105] S41 generates SQL query candidate statements. The model generates corresponding SQL query statement candidates based on the precise semantic mapping relationship, complex query logic, and enhanced semantic information obtained from the vector database learned during training.
[0106] S42 makes fine-tuned optimization adjustments to the connection order, condition screening order, index usage strategy, etc. in SQL statements based on the index distribution of the legal agency's database, the size of the table, and the actual distribution characteristics of the data.
[0107] S43 executes the SQL query statement: executes the optimized SQL query statement on the legal agency database to obtain a query result set.
[0108] S44 conducts a comprehensive and in-depth evaluation of the execution results based on the expected query results and specific business needs. The specific evaluation method is as follows:
[0109]
[0110] Here, Rideal represents the ideal result set, and R represents the query result set.
[0111] The evaluation of query results and continuous learning include the following steps:
[0112] S51. Obtain user satisfaction and collect feedback;
[0113] S52, save the valid query;
[0114] S53. Match similar questions.
[0115] Specifically, evaluating query results and continuously learning specifically include:
[0116] S51 obtains user satisfaction visualization query results and collects user satisfaction.
[0117] S52 sets the SQL query statements of the query results with high user satisfaction as SQL={sql1, sql2, ..., sql n} and the corresponding natural language question set Q = {q1, q2, ..., q n}Save it to the vector library VDB, as follows:
[0118] VDB=VDB∪{(sql n ,q n )|k∈I high}
[0119] Among them, I high Represents a set of queries with high user satisfaction.
[0120] S53 priority matching: the input natural language query is q new , convert it into a vector representation Calculate the similarity sim with the inventory problem in the vector library k .
[0121]
[0122] Principle: First, key data sources within the legal agency are collected, and the data is screened and cleaned, and a knowledge vector library is constructed based on the processed data. Secondly, a legal terminology corpus is designed by collecting data and the experience of professionals. At the same time, a basic pre-trained large model is selected for fine-tuning. The concept of generative adversarial networks (GANs) is introduced to further enhance the robustness of the model and its adaptability to complex and changing environments through adversarial training. Finally, the staff inputs natural language, and the model generates candidate SQL query statements. At the same time, based on the index distribution of the legal agency's database, the size of the table, and the actual distribution characteristics of the data, the connection order, condition screening order, index usage strategy, etc. in the SQL statement are fine-tuned and adjusted. Based on the query results, the user's satisfaction with the results is collected, and the satisfactory SQL results are stored in the knowledge vector library.
[0123] The foregoing is merely a list of specific embodiments of the present application, intended to enable those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but rather is intended to conform to the broadest scope consistent with the principles and novel features claimed herein.
Claims
1. The Text2SQL conversion method in a specific scenario based on a large model is characterized by: The following methods are included: S1. Data preprocessing and knowledge base construction in the field of legal institutions; S2, fine-tune the large model; S3, introduce the concept of generative adversarial network; S4. Execute SQL query statements and optimize query results; S5. Evaluate query results and continue learning.
2. The Text2SQL conversion method for a specific scenario based on a large model according to claim 1, characterized in that: The data preprocessing and knowledge base construction in the legal institution field include the following steps: S11. Data Collection: Collect a wide range of natural language queries related to case clues and their corresponding SQL query statements from multiple key data sources within the legal agency. n }, n = 1, 2, ... k; S12. Obtain database table information: Completely obtain the metadata set of the legal agency database, E = {e1, e2, ...e n }, n = 1, 2, ... k; including detailed structure definitions, field information, data types, primary key and foreign key relationships, and index information of all tables; S13. Develop data specifications: Perform comprehensive and detailed cleaning operations on all collected data C. Among them: Clean represents the data cleaning function; S14. Construct vector database: Use the extracted semantic feature vectors to construct a vector database, and convert each legal term t∈C into a corresponding vector representation through the word vector model. The mapping function corresponding to the word vector model is: in: Where C represents all data.
3. The Text2SQL conversion method for a specific scenario based on a large model according to claim 1, characterized in that: The fine-tuning of the large model includes the following steps: S21 Corpus Design: Design a legal terminology corpus by collecting data and professionals’ experience; S22 Select pre-trained model base: Select pre-trained models that are suitable for processing complex natural language tasks and have excellent scalability, and design a pre-trained model set M = {m1, m2, ..., m k }, where each m l Represents a pre-training model, for each m l , define the evaluation index set I = {i1, i2, ..., i j }, the comprehensive score of the model is S l : S l =ScoreModel(m l ,i) Where: l = 1, 2, ..., k; S23 fine-tunes the details of the large model: By fine-tuning the large model, the model can simultaneously focus on different semantic units in natural language and the database structure elements represented by the related semantic vectors in the vector database, and accurately calculate the association weights between them, accurately constructing a complex and accurate semantic mapping relationship between natural language queries and database structures.
4. The Text2SQL conversion method for a specific scenario based on a large model according to claim 1, characterized in that: The introduction of the concept of generative adversarial networks includes further enhancing the robustness of the model and its adaptability to complex and changing environments through adversarial training.
5. The Text2SQL conversion method for a specific scenario based on a large model according to claim 4, characterized in that: The adversarial training to further enhance the model's robustness and adaptability to complex and changing environments includes the following steps: S31. Build the generator G and set the generator's loss function. S32. Build the discriminator D and design the loss function of the discriminator; S33. Adversarial training allows the generator and discriminator to compete with each other during the training process.
6. The Text2SQL conversion method for a specific scenario based on a large model according to claim 5, characterized in that: The loss function of the generator is: Where: z is the input noise of the generator, P Z (Z) is the distribution of noise; The loss function of the discriminator is: Where: x is a real SQL statement, P data (x) is the distribution of real SQL statements.
7. The Text2SQL conversion method for a specific scenario based on a large model according to claim 1, characterized in that: Executing the SQL query statement and optimizing the query results includes the following steps: S41. The staff of the Procuratorate inputs the natural language query into the model to generate candidate queries; S42. Analyze database indexes and optimize query statements; S43. Execute the optimized SQL query statement on the database.
8. The Text2SQL conversion method for a specific scenario based on a large model according to claim 1, characterized in that: The evaluation of query results and continuous learning include the following steps: S51. Obtain user satisfaction and collect feedback; S52, save the valid query; S53. Match similar questions.