Training data generation method and related equipment
By generating accurate and diverse training data, the problems of insufficient training data and insufficient diversity in NL2SQL systems are solved, and the recognition accuracy and response efficiency are improved.
Patent Information
- Application Number
- CN202411992979.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-23
AI Technical Summary
Due to insufficient training data or insufficient diversity in existing NL2SQL systems, the recognition accuracy rate and increased response time are problems.
Provide a training data generation method, by receiving query questions and SQL statement relationship pairs as input corpus, and combining prompt information, generate query templates and SQL template relationship pairs, standard entities and generalized entities, determine the location of slots to be filled, combine the training data to be generated, and perform verification to ensure the accuracy and diversity of the data.
The recognition accuracy and response efficiency of the NL2SQL system are improved, and the performance of the language model is enhanced by generating more accurate and diverse training data.
Smart Images

Figure CN120030345A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence, and in particular to a training data generation method and related equipment. Background Art
[0002] Natural Language to SQL (NL2SQL) is a technology that converts users' natural language queries into SQL queries. The main goal of this technology is to enable users to interact with the database using natural language without having to master the complex SQL syntax. In NL2SQL, the process of converting natural language into SQL statements requires sufficient diverse and accurate data, and the recognition accuracy of NL2SQL can be improved by training the NL2SQL model.
[0003] Therefore, how to generate more accurate and diverse training data to improve the recognition accuracy of NL2SQL has become an urgent problem to be solved. Summary of the invention
[0004] The present application provides a training data generation method and related equipment, which are used to solve the problem that in the process of the NL2SQL system using a language model to perform data query, the recognition accuracy of the NL2SQL system using the language model is low and the response time of the NL2SQL system is increased due to insufficient amount or insufficient diversity of the language model training data.
[0005] In a first aspect, a method for generating training data is provided, wherein the training data is used to train a second language model, the second language model is used to convert natural language into SQL statements, and a number-checking problem and an SQL statement relationship pair are received, wherein the number-checking problem and the SQL statement relationship pair are input corpus of the first language model; one or more number-checking templates and SQL template relationship pairs, standard entities and generalized entities are generated according to the input corpus and prompt information; the positions of slots to be filled in the one or more number-checking templates and SQL template relationship pairs are determined; a number-checking template and SQL template relationship pair includes a number-checking template and an SQL template; the positions of slots to be filled in the one or more number-checking templates and SQL template relationship pairs are used to indicate the positions of standard entities and generalized entities in the number-checking problem and SQL statement relationship pairs; the generalized entity and the standard entity are respectively combined with the number-checking template and the SQL template in a preset manner to generate a combined number-checking problem and SQL statement relationship pair; the combined number-checking problem and SQL statement relationship pair are verified, and training data is generated according to the verification result, wherein the training data is the number-checking problem and SQL statement relationship pair after the combination and verification.
[0006] The method described in the first aspect helps the first language model to better understand the task requirements and improve the performance of the first language model in executing the task of generating training data by adding prompt information while inputting corpus into the first language model, thereby ensuring the accuracy of the generated training data; in addition, a number-checking template, an SQL template, a standard entity, and a generalized entity are generated by the first language model, and the number-checking template, the SQL template, the standard entity, and the generalized entity are combined in a preset manner to finally obtain a combined number-checking problem and an SQL statement. Through the combination of the preset methods proposed in the present application, the diversity of the generated training data is ensured; in addition, the present application scheme also verifies the combined number-checking problem and SQL statement, further screens out unqualified training data, further improves the accuracy of the generated training data, and thus improves the recognition accuracy of NL2SQL.
[0007] In a possible implementation, the generalized entity and the standard entity are combined with a number-looking template and an SQL template respectively in a preset manner to generate a combined number-looking problem and SQL statement relationship pair, including: filling the generalized entity into a slot to be filled in the number-looking template to generate a number-looking problem; filling the standard entity into the SQL template to generate an SQL statement; and determining a combined number-looking problem and SQL statement relationship pair based on the generated number-looking problem and the generated SQL statement; wherein the combined number-looking problem and SQL statement relationship pair includes a number-looking problem and an SQL statement corresponding to the number-looking problem.
[0008] The above implementation further illustrates the implementation process of the preset combination method, and the diversity of generated data is guaranteed through the preset combination method.
[0009] In a possible implementation, the step of filling the generalized entities into the slots to be filled in the counting template to generate the counting problems further includes: when there are multiple slots to be filled in the counting template, filling any one of the generalized entities of the same type into the slot to be filled corresponding to the generalized entity of that type until each slot to be filled is filled with the generalized entity, thereby generating a counting problem; traversing each generalized entity of the same type, filling the slots to be filled corresponding to the generalized entity of that type until each slot to be filled is filled with the generalized entity, thereby generating multiple counting problems; and the number of counting problems generated satisfies the Cartesian product.
[0010] The above implementation further illustrates the generation process of the number-checking problem. By traversing each slot position to be filled, each slot position to be filled traverses each generalized entity corresponding to the position, thereby ensuring the diversity of the generated number-checking problems, thereby ensuring the diversity of the number-checking problems and SQL statement relationship pairs as training data.
[0011] In a possible implementation, the prompt information includes: role definition (Role), background (Background), process (Process), examples (Examples), and specified format (Output form) of the large model.
[0012] The above implementation method further illustrates the content of the prompt information, and further improves the performance of the first language model through the content and type of the prompt information, thereby improving the first language model's ability to understand the input corpus and generating more accurate number-checking templates, SQL templates, generalized entities, and standard entities.
[0013] In a possible implementation, the checking of the combined number-checking question and SQL statement relationship pair includes: language model checking and / or SQL statement checking.
[0014] The above implementation further screens out unqualified training data by verifying the combined number-checking problem and SQL statement, further improving the accuracy of the generated training data, thereby improving the recognition accuracy of NL2SQL.
[0015] In a second aspect, a method is provided for training a second language model using the training data obtained by the training data generating method provided by the first aspect, wherein the parameters of the second language model are optimized according to the training data provided by the first aspect to obtain a second language model with optimized parameters; wherein the second language model is used to assist users in completing data queries.
[0016] The method of the second aspect is implemented to train the second language model through training data, thereby optimizing the parameters of the second language model, thereby improving the accuracy and response efficiency of executing NL2SQL query tasks using the second language model.
[0017] According to a third aspect, a method for executing data query according to the second language model provided by the first aspect or the second aspect is provided, characterized in that a natural language input of a user is received, and the natural language input is converted into an SQL query statement according to the NL2SQL system and the second language model; the second language model is trained with training data generated by the training data generation method provided by the first aspect; and a data query result is determined in a database management system according to the SQL query statement.
[0018] By implementing the method of the third aspect, since the second language model is trained with the training data provided by the first language model and the parameters of the second language model are optimized, the accuracy and efficiency of data query using the second language model are improved.
[0019] In a fourth aspect, a chip system is provided, the chip system comprising a processor and a power supply circuit, the power supply circuit being used to supply power to the processor, and the processor being used to execute the operating steps of the method described in the first aspect.
[0020] In a fifth aspect, a computing device is provided, comprising a processor and a memory; the processor is used to execute instructions stored in the memory so that the computing device performs the operating steps of the method described in the first aspect.
[0021] In a sixth aspect, a computing device cluster is provided, comprising at least one computing device, each computing device comprising a processor and a memory;
[0022] The processor of the at least one computing device is used to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster performs the operating steps of the method described in the first aspect.
[0023] In a seventh aspect, a computer program product comprising instructions is provided. When the instructions are executed by a computing device cluster, the computing device cluster executes the operation steps of the method described in the first aspect.
[0024] In an eighth aspect, a computer-readable storage medium is provided, comprising computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster performs the operating steps of the method described in the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 This is a schematic diagram of the architecture of an NL2SQL system and its corresponding flow chart provided by this application;
[0026] Figure 2 This is an example diagram of an NL2SQL system deployed in a cloud environment provided by this application;
[0027] Figure 3 This application provides an architecture diagram of a training data generation system;
[0028] Figure 4 It is a schematic diagram of a training data generation method provided by this application;
[0029] Figure 5 is a structural diagram of a computing device provided by the present application;
[0030] Figure 6 is a structural diagram of another computing device provided by the present application;
[0031] Figure 7 It is a structural diagram of a computer cluster provided by this application. DETAILED DESCRIPTION
[0032] The terms "first", "second", and "third" etc. in the specification embodiments and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequential order. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, including a series of steps or units. Methods, systems, products or equipment are not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or equipment. "And / or" is used to represent the selection of one or all between the two objects connected thereto. For example, "A and / or B" means A, B or A+B.
[0033] The application scenarios involved in this application are explained below.
[0034] In the field of machine learning, training data is used as the input corpus of the language model. The performance of the language model is improved by training the language model with training data. However, due to the insufficient amount and diversity of training data, the model's generalization ability is poor, or because the training data is not representative, the language model may learn the wrong pattern, resulting in poor performance of the language model in practical applications. For example, if there is less data in certain categories in the training data set, the model may not be able to accurately identify these categories. In addition, when the amount of training data is insufficient, the impact of noise and outliers on the model will be greater, which may cause the model to learn wrong information. Some complex machine learning models require a large amount of data for effective training. If the training data is insufficient, the model may not be able to fully learn, and more computing resources are needed to make up for the problem of insufficient data.
[0035] Furthermore, in machine learning in a specific field, a language model can be used to perform recognition in a specific field. A language model can be used to assist in the execution of NL2SQL query tasks. NL2SQL is a technology that converts a user's natural language query into an SQL query. The main goal of this technology is to enable users to interact with the database using natural language without having to master complex SQL syntax. In NL2SQL, in the process of converting natural language into SQL statements, there needs to be sufficient diverse and accurate data. By training the language model used in the NL2SQL (Natural Language to SQL) system, the recognition accuracy of NL2SQL (Natural Language to SQL) can be improved. Therefore, how to generate training data and thus improve the recognition accuracy of NL2SQL (Natural Language to SQL) has become an urgent problem to be solved.
[0036] In order to solve the problem of low recognition accuracy caused by insufficient training data, the present application provides a method for generating training data, wherein the training data is generated by a first large model, and the generated training data is used to train a second large model. The second large model is parameter optimized after training, and the second large model with optimized parameters is used to execute the NL2SQL task, thereby improving the accuracy of NL2SQL recognition.
[0037] The following is an explanation of the terms related to this application.
[0038] (1) Counting questions and SQL statement pairs: Counting questions refer to query questions input by users, such as "Please query the income in 2024". SQL (Structured Query Language) statements are a standard language used to manage and operate relational databases and are used to query data in databases. For example, when performing a database query, the SQL statement accessed may be select*. Counting questions and SQL statement pairs refer to two counting questions and SQL statements that express the same semantics and have a corresponding relationship. Generally speaking, one counting question corresponds to one SQL statement. For example, the SQL statement corresponding to the counting question "Please query the income of this month" is "SELECT income From Table WHEREDate='202409'", so [{"Question":"This month's income","SQL":"SELECT income FromTable WHERE Date='202409'"}] is a set of counting questions and SQL statement pairs.
[0039] (2) Counting template and SQL template relationship pair: Counting template refers to a general template used to express counting problems. For example, the counting problem
[0040] The query template corresponding to "Please query the income in 2024" is "Please help me query <date>of <item>SQL template refers to a general template used to express SQL statements. For example, the SQL template corresponding to the SQL statement "SELECT income From TableWHERE Date='202409'" is: <item>From Table WHERE Date= <date>". The relationship pair of number-checking template and SQL template is two general templates used to express the relationship pair of number-checking problem and SQL statement, that is, two number-checking templates and SQL templates with corresponding relationship and expressing the same semantics.
[0041] (3) Standard entity: A standard entity is a value that is queried in the database. For example, in the SQL statement "SELECT <item>From Table WHERE Date= <date>", a standard entity is <item>and <date>Generally speaking, standard entities include income, amount, user ID, date, etc. This application will not list them here, nor will it limit them. For SQL databases, the expression form of standard entities usually cannot be Chinese characters, but must conform to the expression form of the database language (such as SQL language). For example, standard entities can be: sev_income, 202402. For example, in the actual SQL statement "SELECTincome From Table WHERE Date='202409'" the standard entities are income and 202409. In this application, standard entities are represented by st.
[0042] (4) Generalized entity: A generalized entity refers to a key entity that has been generalized in the number-checking problem. In databases and knowledge graphs, key entities refer to objects that have independent existence and can be clearly defined. They are usually abstract representations of people, things, objects, and relationships in the real world. For example, in a student management system, students, courses, and teachers are all key entities; for another example, in "Please help me check September's income", September and income are both key entities. The generalization process of key entities can be understood as expanding and / or listing any expression that can express key entities. For example, September is generalized to include: September, September, September 2024, this month, etc. For example, for the standard entity sev_income, its corresponding generalized entities can be: "income", "income progress". In this application, generalized entities are represented by au.
[0043] The following is an example in the form of a table.
[0044]
[0045] (5) NL2SQL is a technology that converts natural language queries into SQL queries. It combines natural language processing (NLP) and database queries, and aims to enable users to interact directly with the database through natural language.
[0046] (6) Prompt information: The prompt information of the language model is the input text used to guide and control the language model to generate output.
[0047] The prompt can be a question, a command, or a context paragraph that clarifies the model's task and the expected output format.
[0048] In order to make the purpose, technical solution and principle of the present application clearer, the implementation mode of the present application will be further described in detail below with reference to the accompanying drawings.
[0049] First, the NL2SQL technology implemented by the second language model provided by the present application will be introduced.
[0050] like Figure 1 shown. Figure 1 : is the architecture of the NL2SQL system provided in the embodiment of the present application and its corresponding flow chart, such as Figure 1 As shown, the architecture includes a client 100 , an NL2SQL system 200 , a second language model 300 , and a database system 400 .
[0051] First of all, Figure 1 The architecture of the client 100, NL2SQL system 200, second language model 300, and database system 400 are explained and illustrated.
[0052] A communication connection is established between the client 100, the NL2SQL system 200, the second language model 300, and the database system 400 through a network. The communication connection can be a wired connection or a wireless connection. The network can be the public Internet, an internal local area network (LAN), a virtual private network (VPN), a dedicated line such as a fiber optic line, a copper wire, a satellite connection, etc., or a wireless network such as a wireless local area network (Wi-Fi), a cellular network, etc., which is not specifically limited in this application.
[0053] The client 100 is deployed on a terminal device to realize human-computer interaction. The terminal device includes a personal computer, a smart phone, a wearable device, a handheld processing device, a tablet computer, a mobile notebook, an augmented reality (AR) device, a virtual reality (VR) device, an intelligent conference device, etc., which are not specifically limited here. The client 100 can also be deployed on a physical server, such as an ARM server or an X86 server, which is not specifically limited in this application.
[0054] The NL2SQL system 200 can be deployed on a server, which can be a bare metal server (BMS), a virtual machine, a container, or an edge computing device. BMS refers to a general physical server, such as an ARM server or an X86 server; a virtual machine refers to a complete computer system with complete hardware system functions and running in a completely isolated environment simulated by software. All work that can be done in a physical computer can be done in a virtual machine. When creating a virtual machine in a computing device, part of the hard disk and memory capacity of the physical machine needs to be used as the hard disk and memory capacity of the virtual machine. Each virtual machine has an independent basic input / output system (BIOS), hard disk, and operating system, and can operate the virtual machine like a physical machine; a container is a portable software unit that can merge an application and all its dependencies into a software package that is not limited by the underlying host operating system, so there is no need to build a complex environment, simplifying the process from application development to deployment; an edge computing device refers to a device that is closer to the data source and end user, with low latency and high bandwidth characteristics, such as intelligent routing, edge servers, etc. The server can be an edge server, or a local server in an enterprise's local data center. This application is not specifically limited.
[0055] The second language model 300 includes at least one of a rule-based model, a machine learning-based model, and a pre-trained language model. The rule-based model converts natural language into SQL statements through predefined rules and templates, and is suitable for structured and simple query scenarios. The machine learning-based model uses supervised learning methods to train the model through a large amount of labeled data, and can handle more complex queries, but requires a large amount of training data. Models based on deep learning such as Seq2Seq, Transformer and other models can capture the complex relationship between natural language and SQL, and have high accuracy and generalization capabilities. Pre-trained language models, such as BERT, GPT, etc., are pre-trained on large-scale corpora and then fine-tuned, and can perform well in NL2SQL tasks. The number of parameters is usually between hundreds of millions and billions.
[0056] The database system 400 includes a database management system (DBMS) and at least one database (DB) ( Figure 1 (not shown). In the database system, the application program can transparently operate the database through the database management system, and the data in the database is managed by the database management system.
[0057] Optionally, the aforementioned database may be a relational database, which refers to a database that uses a relational model to organize data. It stores data in the form of rows and columns. Each relational model may be referred to as a relational table. According to different storage principles, relational databases may be divided into distributed relational databases and non-distributed relational databases.
[0058] In a possible implementation, the database system 400 receives an SQL statement, determines data according to the SQL statement, and returns the data to the client 100 .
[0059] Optionally, the client 100, the NL2SQL system 200, the second language model 300, and the database system 400 are respectively deployed on different computing devices or computing device clusters, for example Figure 1 Alternatively, the NL2SQL system 200 and the second language model 300 are deployed on the same computing device or computing device cluster.
[0060] In a specific implementation, the client 100 can be software or an application running on a terminal device or computing device controlled by a user, such as a personal computer (PC) client, a world wide web (web) client accessed through a browser, an application (APP) client running on a mobile terminal, or a console of a cloud platform, which is not specifically limited in this application. The user holding the client 100 is a person who manages transaction business of an enterprise, such as a financial employee or an information technology (IT) employee of the enterprise, which is not specifically limited in this application.
[0061] Optionally, the client 100 may be a client specifically used for performing natural language to SQL data query, and the NL2SQL system 200 and / or the second language model 300 are integrated with the client 100 on the same computing device.
[0062] Optionally, the client 100 may also be a comprehensive client including a natural language to SQL data query function, such as a tool in which the client has a data query function. These comprehensive clients include not only data query functions, but also other functions, such as dialogue, navigation, etc., in addition to the query function, and are used in different scenarios. The above examples are for illustration only and are not specifically limited in this application.
[0063] Optionally, client 100 can also be a client of the cloud platform, used for users to purchase and rent various cloud services. The natural language to SQL query provided in this application can be one of the cloud services, and users can purchase this cloud service separately to realize data query; or, the cloud platform provides users with a comprehensive service, and the above-mentioned data query function through natural language can be used as a sub-service in the comprehensive cloud service. For example, when a user purchases a database cloud service, the above-mentioned user input to the SQL query service can be a sub-service in the cloud service. This application does not make any specific limitations.
[0064] The above describes in detail the possible deployment methods of the client 100, the NL2SQL system 200, the second language model 300, and the database system 400. In actual deployment, they can be flexibly deployed in combination with specific application scenarios and business requirements. The following examples are given to illustrate the actual deployment methods of the client 100, the NL2SQL system 200, the second language model 300, and the database system 400 in combination with specific application scenarios.
[0065] In one application scenario, the client 100, the NL2SQL system 200, the second language model 300, and the database system 400 can be deployed on office equipment within the enterprise. For example, the NL2SQL system 200 is deployed on a server or server cluster purchased by the enterprise, the client 100 is deployed on the enterprise's office computer, and the database system 400 is deployed on the enterprise's storage device. The database systems can be different databases on the same storage device or different databases on different storage devices, and this application does not make any specific limitations.
[0066] In another application scenario, the client 100, the NL2SQL system 200, the second language model 300, and the database system 400 can be deployed in a cloud environment. The data query system composed of the NL2SQL system 200 and the second language model 300 is the console of the cloud platform. For example, Figure 2 This is an example diagram of an NL2SQL system 200 deployed in a cloud environment provided by the present application. Figure 2 As shown, the user can initiate a data query request through the client 100. After the client 100 sends the data query request to the cloud platform, the cloud platform can provide the client 100 with the right to use the NL2SQL system 200, so that the user can use the NL2SQL system through the client 100 to implement the data query function based on the user's natural language input.
[0067] In the specific implementation, the cloud platform also maintains various basic resources, including computing resources, storage resources, network resources, security resources, etc., to meet the computing needs of the NL2SQL system under different scales and loads, and these computing resources can be dynamically scaled according to the usage requirements of the NL2SQL system to ensure the stable operation of the NL2SQL system. Among them, the database system 400 can also be a cloud service of the data center, such as elastic cloud services, cloud storage services, etc. After purchasing the cloud service, the user configures the database system 400 and stores the data in the database. When the user has data query needs, the data query cloud service provided by the purchased NL2SQL system can be used. After converting the natural language input by the user into an SQL statement, the data stored in the database, such as the income table, is retrieved according to the SQL statement, and the results are finally returned to the client 100. The above examples are for illustration, and this application does not make specific limitations.
[0068] It should be understood that the above application scenarios are used for illustration purposes only. The client 100, server 200, second language model 300, and database system 400 can be flexibly deployed according to actual business needs, and examples are not given one by one here.
[0069] The above describes Figure 1 The deployment method and actual deployment example of the client 100, NL2SQL system 200, second language model 300, and database system 400 shown in FIG. Figure 1 The operation flow corresponding to the client 100, NL2SQL system 200, second language model 300, and database system 400 is shown.
[0070] Step S100 generates a query request according to the natural language input by the user;
[0071] Specifically, step S100 is performed by the client 100. The client 100 is used to receive user input, which may be a query question, for example, a number-checking question. After receiving the number-checking question, the client 100 generates a query request and sends it to the NL2SQL system 200.
[0072] Optionally, the user input also includes various query questions expressed in natural language, and the present application is not limited to number-checking questions. At the same time, the user input natural language includes Chinese, English, and German, and the present application does not specifically limit the language type of the natural language.
[0073] Step S101 calls the second language model to perform conversion from natural language to SQL statements.
[0074] Specifically, step S101 is performed by the NL2SQL system 200. The NL2SQL system 200 calls the second language model 300 through the API interface to achieve the conversion of natural language into SQL statements. When the NL2SQL system 200 receives the query request, it calls the second language model. The NL2SQL system 200 can be deployed on a server, such as AWS, Google Cloud, Microsoft Azure, etc., which is not limited in this application.
[0075] Step S102 converts the natural language input by the user into SQL statements.
[0076] Specifically, step S102 is performed by the second language model 300. For example, when the content represented by the natural language is a number query question, the number query question is: please query the income in September 2024. The second language model 300 converts the number query question into a corresponding SQL statement "SELECT income From Table WHERE Date = '202409'".
[0077] In a possible implementation, the NL2SQL system 200 and the second language model 300 can be a module (the module can be implemented by software or hardware) respectively, and deployed on the same computing device, which can be a server or a terminal device. For example, the second language model 300 can be used as one of the functions or tools of the NL2SQL system 200. When the NL2SQL system 200 needs to perform the conversion from natural language to SQL statements, the second language model is called to perform related operations and return the results to the NL2SQL system 200.
[0078] In another possible implementation, the NL2SQL system 200 and the second language model 300 are deployed on different computing devices.
[0079] The SQL statement generated in step S103 will be sent to the database system 400 (such as MySQL, PostgreSQL, SQL Server, etc.) to determine the data content and return the result.
[0080] Specifically, step S103 is performed by the database system 400. The database system 400 may be MySQL, PostgreSQL, SQL Server, etc. The generated SQL statement will be sent to the database management system 400, where the data content corresponding to the SQL statement is determined, and finally the data content is returned to the client 100 as a query result.
[0081] Step S104 obtains query questions and SQL statements that have a paired relationship as training corpus for the first language model.
[0082] Specifically, step S103 is performed by client 100. A query question and an SQL statement having a pairing relationship means that when the query question and the SQL statement express the same meaning, the query question and the SQL statement have a pairing relationship. For example, the query question is: please query the income in September 2024; the SQL statement is "SELECT income From Table WHERE Date = '202409'", then it is considered that the above query question and the SQL statement have a pairing relationship. The two with a pairing relationship are divided as alternative training corpus, which can be used for training the first language model.
[0083] It should be understood that the client 100, the NL2SQL system 200, the second language model 300, and the database system 400 can be implemented by software or by hardware. As an example, the implementation of the client 100 is described below. Similarly, the implementation of the NL2SQL system 200, the second language model 300, and the database system 400 can refer to the implementation of the client 100.
[0084] As an example of a software functional unit, the client 100 may include code running on a computing instance. The computing instance may be at least one of a physical host (computing device), a virtual machine, a container, and other computing devices. Furthermore, the computing device may be one or more. For example, the YY device may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the application may be distributed in the same region or in different regions. The multiple hosts / virtual machines / containers used to run the code may be distributed in the same AZ or in different AZs, and each AZ includes a data center or multiple data centers with close geographical locations. Generally, a region may include multiple AZs.
[0085] Similarly, multiple hosts / virtual machines / containers used to run the code can be distributed in the same VPC or in multiple VPCs. Usually, a VPC is set up in a region. For cross-region communication between two VPCs in the same region and between VPCs in different regions, a communication gateway must be set up in each VPC to achieve interconnection between VPCs through the communication gateway.
[0086] As an example of a hardware functional unit, the client 100 may include at least one computing device, such as a server, etc. Alternatively, the YY device may also be a device implemented using a CPU, ASIC, PLD, CPLD, FPGA, GAL, DPU, NPU, SoC, offload card, accelerator card, etc. Among them, the above-mentioned PLD may be implemented by CPLD, FPGA, GAL or any combination thereof.
[0087] The multiple computing devices included in the client 100 can be distributed in the same region or in different regions. The multiple computing devices included in the client 100 can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the client 100 can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, GALs, DPUs, NPUs, SoCs, offload cards, acceleration cards, etc.
[0088] The above describes the method and architecture for identifying NL2SQL using the second language model 300. The following describes the process of parameter tuning of the second language model 300. The parameter tuning process is actually the process of training the language model, that is, the process of training the second language model using training data. The process of identifying NL2SQL using the second language model 300 is an inference process.
[0089] The second language model 300 involves a training process and an inference process.
[0090] In the training phase, the training data is input into the second language model, and the parameters of the second language model are optimized to obtain an optimized second language model. It should be noted that the training data is generated by the first language model, and the first language model is a part of the training data generating device 500. Therefore, it can also be said that the generated training data is generated by the training data generating device 500. The process of parameter optimization belongs to the prior art and will not be described in detail in this application.
[0091] In the inference phase, when the client receives the natural language input by the user using the second language model after parameter optimization, it sends a query request to the server. The NL2SQL system deployed on the server calls the second language model to convert the natural language into the corresponding SQL statement. The database management system (DBMS) executes the SQL statement and completes the query operation and returns the data. The returned data is received by the client.
[0092] Combined with the above Figure 1-2 The operation flow and system framework of the NL2SQL number query technology using the second language model are introduced; then, the process of training the second language model 300 using training data is introduced through text description. Next, the training data generation process and its corresponding system architecture will be introduced. Figure 3 , attached Figure 3 is a diagram of the architecture of the training data generation system 500. Figure 3 The architecture of the training data generation system 500 will be introduced in detail, and the training data generation process will be explained.
[0093] First, the content of the training data and the application scenarios are explained.
[0094] Specifically, the task of the training data generation system 500 is to generate more rich training data after the received limited input corpus is processed by the first language model. That is, the task of the training data generation system 500 is to generate training data that can ensure diversity and preparation. The generated training data is also called output corpus, which is used to train the second language model, thereby improving the recognition accuracy and efficiency of using the second language model to assist in executing NL2SQL query tasks.
[0095] The content of the training data includes:
[0096] Optionally, the training data is number-checking problems and their corresponding SQL statements;
[0097] Optionally, the generated training data is a pair of number-checking problems and SQL statements;
[0098] Optionally, the training data is used to assist users in querying the SQL database using natural language.
[0099] Optionally, the generated training data can be used in actual scenarios such as revenue query and transaction risk control.
[0100] It should be understood that training data helps to improve the recognition efficiency and accuracy of NL2SQL technology in executing tasks, thereby improving the recognition efficiency and accuracy in multiple scenarios such as income query, transaction data query, risk data query, and financial data query.
[0101] The generated training data can be applied to any one or more of the following real-world scenarios:
[0102] (1) Revenue query: The user enters "Query the total monthly revenue in 2023", and the system generates the SQL statement: SELECT SUM (revenue) FROM sales WHERE year = 2023 GROUP BY month.
[0103] (2) Transaction data query: The user enters "Calculate the total amount of all transactions last month", and the system generates the SQL statement: SELECT SUM(amount)FROM transactions WHERE date>='2023-10-01'AND date<'2023-11-01'.
[0104] (3) Risk data query: When the user inputs "Find all high-risk customers", the system generates the SQL statement: SELECT * FROM customers WHERE risk_level = 'high'.
[0105] (4) Financial data query: The user enters "Get stock trading records for the past year", and the system generates the SQL statement: SELECT * FROM stock_trades WHERE trade_date>='2023-01-01'AND trade_date<'2024-01-01'.
[0106] The above application scenarios are examples of applicable scenarios for training data and are not limited in this application.
[0107] It should be understood that the generated training data can effectively improve the effect of NL2SQL technology on data query in the above-mentioned actual scenarios, reduce the dependence on SQL knowledge, and enable business personnel to obtain the required data more conveniently directly through natural language. At the same time, the generated training data can further optimize the parameters of the second language model, thereby improving the accuracy of using the second language model when performing tasks.
[0108] Combine the following Figure 3 , the framework of the training data generation system is introduced. Figure 3 As shown, Figure 3 An architectural diagram of a training data generation system is provided. Figure 3 It can be determined from the system framework shown that the training data generating system includes a database 600 and a training data generating device 500 .
[0109] The training data generating device 500 includes a first language model 510, an information extraction module 520, a rule combination module 530, and a verification module 540. The training data generating system 500 includes the first language model 510, the information extraction module 520, the rule combination module 530, and the verification module 540, which is an exemplary division method. In a specific implementation, the training data generating system 500 may also not be divided according to Figure 3 The division method shown divides the training data generation system 500 into units. For example, the first language model 510 and the rule combination module 520 can be merged, or the rule combination module 520 and the verification module 530 can be merged. This application does not make specific limitations.
[0110] The first language model 510 is used to receive a number-checking question and an SQL statement relationship pair, which is the input corpus of the first language model; and generate one or more number-checking templates and SQL template relationship pairs, standard entities and generalized entities according to the input corpus and prompt information.
[0111] The information extraction module 520 is used to determine the positions of the slots to be filled in the relationship pairs of the one or more number-checking templates and SQL templates.
[0112] Specifically, the information extraction module 520 determines the position of the slot to be filled in the one or more number-checking templates and SQL templates relationship pair according to the regular expression. A number-checking template and SQL template relationship pair includes a number-checking template and an SQL template; the position of the slot to be filled in the one or more number-checking template and SQL template relationship pair is used to indicate the position of the standard entity and the generalized entity in the number-checking problem and SQL statement relationship pair.
[0113] The rule combination module 530 is used to combine the generalized entity and the standard entity with the number-checking template and the SQL template respectively according to a preset manner, and generate a combined number-checking problem and SQL statement relationship pair.
[0114] Specifically, the rule combination module 530 is used to determine the positions of the to-be-filled slots in the relationship pairs of one or more number-looking-up templates and SQL templates according to regular expressions; according to a preset rule combination method, the generalized entity and the standard entity are respectively combined with the to-be-filled slots in the number-looking-up template and the SQL template to generate one or more new number-looking-up question and SQL statement relationship pairs as enhanced corpus.
[0115] As an example, the preset rule combination method includes: filling the generalized entity into the to-be-filled slot position of the number-looking template to generate a number-looking problem; filling the standard entity into the SQL template to generate an SQL statement; determining a combined number-looking problem and SQL statement relationship pair based on the generated number-looking problem and the generated SQL statement; wherein the combined number-looking problem and SQL statement relationship pair includes a number-looking problem and an SQL statement corresponding to the number-looking problem.
[0116] As an example, when there are multiple slots to be filled in the count template, any one of the generalized entities of the same type is filled into the slot to be filled corresponding to the generalized entity of that type until each slot to be filled is filled with the generalized entity, thereby generating a count problem; each generalized entity of the same type is traversed, and the slot to be filled corresponding to the generalized entity of that type is filled until each slot to be filled is filled with the generalized entity, thereby generating multiple count problems; and the number of count problems generated satisfies the Cartesian product.
[0117] The verification module 540 is used to verify the number-checking problem and SQL statement relationship pair after the combination, and generate training data according to the verification result. The training data is the number-checking problem and SQL statement relationship pair after the combination and verification.
[0118] Specifically, the verification module 540 is used to verify the enhanced corpus to obtain multiple verified number-checking problem and SQL statement relationship pairs, and the verified number-checking problem and SQL statement relationship pairs are used as generated training data. The verification method includes: language model verification and / or SQL statement verification.
[0119] It should be noted that the training data generated according to the first language model 510 is used to train the second language model 300 .
[0120] Optionally, the number of parameters of the first language model used for training data is greater than the number of parameters of the second language model 510 .
[0121] Specifically, the number of parameters in a language model refers to the total number of adjustable weights and biases in the model. These parameters are adjusted during the training process so that the model can better complete specific tasks (such as language understanding, image recognition, NL2SQL tasks, etc.).
[0122] Optionally, the parameter amount may be, for example, 7B or 13B.
[0123] It should be understood that the more parameters there are, the more expressive and complex the model is. For example, GPT-3 has 175 billion parameters, while some smaller models may only have a few hundred million parameters. Models with large parameter counts usually require more computing resources to train and run, but they also tend to have better generalization capabilities and can handle more diverse data and tasks.
[0124] It should be understood that a model with a large number of parameters can utilize a large amount of data and computing resources during the training phase to learn rich language features and complex patterns. These models can capture more subtle semantic relationships and contextual information, thereby achieving higher accuracy and generalization capabilities during the training process and improving the accuracy of the generated training data. In practical applications, using a model with a small number of parameters for reasoning (or recognition) can significantly reduce the consumption of computing resources and improve the response time. This is particularly important for NL2SQL scenarios that require real-time processing. Therefore, the first language model parameter amount is greater than the second language model 510 parameter amount, which has the above-mentioned beneficial effects.
[0125] In a specific implementation, the first language model 510, the information extraction module 520, the rule combination module 530, and the verification module 540 can be implemented by software or by hardware. For example, the implementation of the first language model 510 is described below by taking the first language model 510 as an example. Similarly, the implementation of the information extraction module 520, the rule combination module 530, and the verification module 540 can also refer to the implementation of the first language model 510.
[0126] The first language model 510 is taken as an example of a software functional unit, and the first language model 510 may include code running on a computing instance. Among them, the computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Further, the above-mentioned computing instance may be one or more. For example, the first language model 510 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region (region) or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one data center or multiple data centers with close geographical locations. Among them, usually a region may include multiple AZs.
[0127] Similarly, multiple hosts / virtual machines / containers used to run the code can be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Usually, a VPC is set up in a region. For cross-region communication between two VPCs in the same region and between VPCs in different regions, a communication gateway needs to be set up in each VPC to achieve interconnection between VPCs through the communication gateway.
[0128] The first language model 510 is taken as an example of a hardware functional unit, and the first language model 510 may include at least one computing device, such as a server, etc. Alternatively, the first language model 510 may also be implemented using a central processing unit (CPU), an application-specific integrated circuit (ASIC), or a device implemented by a programmable logic device (PLD). The above-mentioned PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), a data processing unit (DPU), a neural network processing unit (NPU), a system on chip (SoC), an offload card, an acceleration card, or any combination thereof.
[0129] The multiple computing devices included in the first language model 510 can be distributed in the same region or in different regions. The multiple computing devices included in the first language model 510 can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the first language model 510 can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, GALs, DPUs, NPUs, SoCs, offload cards, acceleration cards, etc.
[0130] Combine the following Figure 4 , the training data generation method provided in this application is explained.
[0131] Attached Figure 4 The training data generating apparatus 500 provided in the present application corresponds to the method steps. The training data obtained by the training data generating method can be used for training such as Figure 1 The NL2SQL technology implements the second language model in the architecture, thereby optimizing the parameters of the second language model and improving the efficiency and accuracy of data query operations performed based on the second language model.
[0132] First, we will combine Figure 3 , the training data generation method provided in this application is explained.
[0133] Step S200 receives a relation pair of a number-checking question and an SQL statement, and the relation pair of the number-checking question and the SQL statement is used as input corpus of the first language model.
[0134] Specifically, the input corpus of the first language model can be a pair of number-checking questions and SQL statements. For example, the number-checking question is: "Please tell me the income in May 2024"; the SQL statement is: "SELECT income FROM table WHERE Date = 202407". The number-checking question "Please tell me the income in May 2024" and the SQL statement "SELECT income FROMtable WHERE Date = 202407" have a corresponding relationship, and the two are a pair of number-checking questions and SQL statements.
[0135] Optionally, the input corpus of the first language model is root such as Figure 1 The second language model 300 shown is generated, and the second language model performs the NL2SQL task to convert the number-checking question input by the user into an SQL statement. For example, the NL2SQL system calls the second language model, converts the number-checking question input by the user into an SQL statement through the second language model, and uses the number-checking question and SQL statement relationship pair with a corresponding relationship as the input corpus of the first language model, which is used for the first language model to generate new training data.
[0136] Step S210 generates one or more number-checking templates and SQL template relationship pairs, standard entities and generalized entities according to the input corpus and prompt information, wherein the standard entity refers to the item to be queried in the SQL database; the generalized entity refers to the item to be queried in the number-checking problem;
[0137] Optionally, the standard entity corresponds to one or more generalized entities with the same semantics.
[0138] Optionally, the prompt information includes: role definition (Role), background (Background), process (Process), examples (Examples), and specified format (Output form) for the large model.
[0139] As an example, the role definition (Role) of the big model in the prompt information of inputting the big model includes: #Role: As a data expert, please understand the NL2SQL task and assist in completing the NL2SQL template generation and entity generalization tasks;
[0140] As an example, the background (Background) in the input large model prompt information includes: #Background: NL2SQL refers to converting user number-checking problems into executable SQL codes. This task requires you to understand the user-entered problems and the corresponding SQL statements (also known as the relationship between number-checking problems and SQL statements), disassemble the key entities, and provide diverse templates and generalized entities.
[0141] As an example, the process of inputting large model prompt information can be as follows: #Process: (1) Understand the user's question and SQL pair, and find the corresponding relationship between the key entities. The key entities can be extracted from the date and query items in the user's question, or inferred from the SELECT and WHERE in the SQL code. (2) Extract the key entities and convert the user's question into multiple generalized slot query templates to be filled. (3) Generalize the extracted key entities accordingly and ensure accuracy. (4) Output the user query template, SQL template, and generalized entity in a structured form.
[0142] As an example, the examples (Examples) in the prompt information of the input large model include: examples of number-looking problems of the input large model and their corresponding SQL statements (also known as pairs of number-looking problems and SQL statements), number-looking templates of the output large model, SQL templates, standard entities, and examples of generalized entities.
[0143] For specific data, please see the table below.
[0144]
[0145]
[0146]
[0147] As a possible implementation, the standard entity is matched with the SQL template to generate an SQL statement. The method of matching the standard entity with the SQL template includes filling the standard entity into the slot to be filled in the number-checking template to generate one or more SQL statements. For example, the slot to be filled in the number-checking template is a date, so the standard entity 202409 can be filled into the number-checking template to obtain the SQL statement "SELECT sev_income From Table WHERE Date = '202409'.
[0148] As a possible implementation, the generalized entity is matched with the number-checking template to generate a number-checking question. The generalized entity is matched with the number-checking template in a manner that includes filling the generalized entity into a slot to be filled in the number-checking template to generate one or more number-checking questions.
[0149] Specifically, the generalized entity is filled into the slot to be filled in the counting template to generate a counting problem; the counting problem here can be multiple or one. When there are multiple slots to be filled in the counting template, any one of the generalized entities of the same type is filled into the slot to be filled corresponding to the generalized entity of the type, until each slot to be filled is filled with the generalized entity, generating a counting problem; traversing each generalized entity of the same type, filling the slot to be filled corresponding to the generalized entity of the type, until each slot to be filled is filled with the generalized entity, generating multiple counting problems; the number of generated counting problems satisfies the Cartesian product.
[0150] As an example, the generalized entity includes two different types of generalized entities, one is a generalized entity representing a date, such as today, December 4, 2024, 20241204, and the other is a generalized entity representing income, such as income, revenue, and profit. The number query template is to answer the following question: <date>of <item>After matching the generalized counting entity with the counting template, there are 9 counting questions generated, as shown in the following table:
[0151]
[0152]
[0153] As a possible implementation, the standard entity and the generalized entity have a corresponding relationship. For example, the generalized entity corresponding to the standard entity income is income and revenue.
[0154] Optionally, a standard entity corresponds to one or more generalized entities that express the same semantics. For example, "Entity":{" <date>":{, this example indicates that when the entity is a date, the standard entity has one expression form, i.e. "202409"; the generalized entity has one or more expressions, and preferably all the generalized entity expressions are output through the large model, for example, ["this month", "September this year", "September 2024", "September", "September"].
[0155] Optionally, one SQL statement corresponds to one or more number-checking questions. For example, the nine number-checking questions generated in the above table can correspond to the same SQL statement, which is "SELECT sev_income From Table WHERE Date='20241204'". Therefore, a total of nine pairs of number-checking questions and SQL statements are obtained.
[0156] As a possible implementation, the prompt information includes a specified format (Output form), and the specified format and / or the specified format example are used to indicate the format of the training data, so that the output content of the first language model is output in a fixed format. The following table illustrates the specified format by taking the output in Json format as an example.
[0157]
[0158]
[0159] From the examples in Table 2, we can see that by specifying the output format of the large model in the form of prompt information, the format of the entities output by the large model is limited. The output formats of standard entities and generalized entities are given as examples, for example, "Entity":{" <date>":{(the following content is omitted), this example indicates that when the entity is a date, the standard entity and generalized entity expressions corresponding to "date" are output according to the format in the example.
[0160] It should be understood that the generalized entity is used to fill in the number-checking template and further generate diversified number-checking questions. By filling multiple generalized entities with the same expression content into the same number-checking template, it is possible to ensure that diverse and comprehensive number-checking questions are generated, and then the first language model generates accurate and comprehensive corpus.
[0161] Step S220 determines the slot positions to be filled in the one or more query templates and SQL template relationship pairs.
[0162] Specifically, the position of the slot to be filled in the pair of the count-check template and the SQL template relationship refers to the position of the slot to be filled in the two general templates with a corresponding relationship in the pair of the count-check template and the SQL template relationship. Generally speaking, the slot to be filled in the pair of the count-check template and the SQL template relationship refers to the position where the generalized entity and / or the standard entity is located. For the count-check template in the pair of the count-check template and the SQL template relationship, the position of the slot to be filled refers to the position where the generalized entity is located. For the SQL template in the pair of the count-check template and the SQL template relationship, the position of the slot to be filled refers to the position where the standard entity is located.
[0163] As a possible implementation manner, the positions of the to-be-filled slots of the one or more number-lookup template and SQL template relationship pairs are determined according to a regular expression.
[0164] As an example, the regular expression may be (? <= ( <item><)).+? (?=(>)), this regular expression can recognize <item><dev_income>The position of the entity in the text box and extract the entity involved in the position; similarly, the regular expression (? <= ( <date><)).+? (?=(>)), this regular expression can recognize <date><202410> The location of the entity in the data and extract the entities involved in the location.
[0165] As a possible implementation, the process of searching for a marker position using a regular expression includes (the marker position here refers to the slot position to be filled in the query template and / or SQL template): first, define a regular expression pattern to match a specific position in the target string. For example, the regular expression (? <= ( <date><)).+? (?=(>)), the target string is: "SELECT <item>From Table WHERE Date= <date>". Next, the specific position in the target string is marked. The specific position here refers to the position of the slot to be filled. Finally, the standard entity and / or generalized entity content corresponding to the position of the slot to be filled is extracted.
[0166] It should be understood that the entity position and the entity content in the number-checking problem can be determined according to the regular expression.
[0167] Step S230 combines the generalized entity and the standard entity with a number-checking template and an SQL template respectively in a preset manner to generate a combined number-checking problem and SQL statement relationship pair.
[0168] Specifically, the generated combined number-checking question and SQL statement relationship pair is also called enhanced corpus.
[0169] Fill the generalized entity into the slot to be filled in the number-checking template to generate a number-checking problem; fill the standard entity into the SQL template to generate an SQL statement; determine the combined number-checking problem and SQL statement relationship pair according to the generated number-checking problem and the generated SQL statement; wherein the combined number-checking problem and SQL statement relationship pair includes a number-checking problem and an SQL statement corresponding to the number-checking problem. The above-mentioned slot-filling process can be in the form of a code language such as Python, C language, etc., which is not limited in this application.
[0170] The following table lists the number-checking questions and their corresponding SQL statements that can be used as enhanced corpus after filling the number-checking template and SQL template through rule combination. Examples are shown in the following table:
[0171]
[0172]
[0173] It can be seen from the above table that the number of generated enumeration problems satisfies the Cartesian product. For example, the number of generalized entities representing income shown in the above table is 2, and the number of generalized entities representing dates is 5. The number of generated enumeration problems is equal to the product of the number of generalized entities representing the same content, that is, the number of generalized entities representing income × the number of generalized entities representing dates = 2 × 5 = 10.
[0174] It should be understood that through the combination of the above entities and templates, comprehensive and diverse number-checking questions and SQL languages are generated. Since the generalized entities comprehensively express various expressions that users may input, it helps the data query system to better understand the questions input by users and improve the efficiency of data query.
[0175] It should be understood that in the relevant fields of financial data query, users usually only need to input natural language, for example, please query the income in September. After the data query system is optimized by the above training data, the data query system's ability to understand natural language is improved, so that the data corresponding to the natural language query problem input by the user can be quickly and accurately obtained from the SQL database and returned to the user. The method for generating training data described in this application helps to improve the efficiency of user queries.
[0176] In a possible implementation, the rule combination method between the standard entity, the generalized entity, the SQL template and the number-checking template includes at least one or more of the following rules: (1) using regular expressions to identify tags in the number-checking template and the SQL template, and confirming fillable slots in the number-checking template and the SQL template; (2) replacing the fillable slots in the number-checking template with generalized entities; (3) replacing the fillable slots in the SQL template with standard entities; (4) outputting the number-checking problem and SQL statement relationship pairs after rule combination according to structured requirements.
[0177] The following will use a practical example to explain the process of rule combination between standard entities, generalized entities, SQL templates and number-checking templates.
[0178] First, as an example, identify the positions that can be filled in the number-checking template and its corresponding SQL template according to the regular expression. When the program runs the regular expression, the system determines the marked or identified positions in the SQL template according to the regular expression. Generally speaking, the marked or identified positions in the SQL template are usually the positions where the entities are located, or the positions that can be filled in by the entities. The entities here can be generalized entities or standard entities, which are not limited in this application. For example, the regular expression: -(? <=( <item><)).+? (?=(>)) can convert the SQL template <item><dev_income>Two positions are marked and identified, or the regular expression - (? <= ( <item><)).+? (?=(>)) can convert the SQL template <date><202410> The two positions are marked or identified. Then, after executing the regular expression, the running program completes marking the entity positions of the number-checking template and the SQL template, and then continues to output the marked position by running the code, for example, python: outputs the position of dev_income. Finally, the running program executes the slot-filling process, which specifically includes: filling the standard entity into the position that can be filled in the SQL template, filling the generalized entity into the position that can be filled in the SQL template, and obtaining multiple filled-in number-checking problems and their corresponding SQL statements. The number-checking problems and SQL statements with corresponding relationships are used as number-checking problems and SQL statements. The relationship pairs are output as generated training data.
[0179] It should be understood that by filling the standard entities into the positions that can be filled in the SQL template and filling the generalized entities into the positions that can be filled in the SQL template, multiple filled-in number-checking problems and their corresponding SQL statements are obtained, thereby improving the accuracy and efficiency of users' queries / searches to the SQL database through natural language.
[0180] In another possible implementation, replacing the slot to be filled in the count template with a generalized entity satisfies a rule combination method, and the rule combination method includes: when there are multiple slots to be filled in the count template, filling any one of the generalized entities of the same type into the slot to be filled corresponding to the generalized entity of that type, until each slot to be filled is filled with the generalized entity, generating a count problem; traversing each generalized entity of the same type, filling the slot to be filled corresponding to the generalized entity of that type, until each slot to be filled is filled with the generalized entity, generating multiple count problems.
[0181] In order to better illustrate the process of combining the above rules, the following example will be given. For example, when the query template is: "Please help me query <date>of <item>Situation", here is <date>and <item>There are two positions that need to be filled. The generalized entities corresponding to the two positions need to be filled in the corresponding positions one by one. <date>The generalized entities corresponding to the position include "this month", "September of this year", "September of 2024", "September", "September", <item>The generalized entities corresponding to the positions include: "income" and "total income". After filling in the generalized entities of the two positions that need to be filled in one by one, a total of 6×2=12 combinations are obtained, so 12 counting problems are obtained. For example, one of the counting problems is "Please help me check the income of this month."
[0182] Optionally, the generalized entities corresponding to the same slot position to be filled express the same meaning.
[0183] Optionally, the same slot position to be filled corresponds to the same type of generalized entity. The generalized entity includes a first generalized entity, a second generalized entity, ..., an Nth generalized entity (where N is a positive integer), the first generalized entity is filled into the first slot position in the number-checking template, the second generalized entity is filled into the second slot position in the number-checking template, and so on.
[0184] It should be understood that by setting a combination of rules, a limited number of number-checking problems and SQL statements can be used to generate a larger number of number-checking problems and their corresponding SQL statements through the combination of the above rules. Since there are many different ways of expressing number-checking problems with the same meaning, a variety of different ways of expressing are generated as comprehensive and accurate as possible according to the above rules, so that the system can understand the user's natural language input more comprehensively and accurately. At the same time, since the number-checking problems and SQL statements are paired as training data, the system can accurately pair the number-checking problems expressed in natural language with SQL statements when it understands the user's natural language input more comprehensively and accurately, and complete the data retrieval in the SQL database through the SQL language, thereby improving the accuracy and efficiency of data query.
[0185] Step S240 verifies the number-checking problem and the SQL statement relationship pair after rule combination, and determines the training data according to the verification result.
[0186] Specifically, the relationship between the number-checking question and the SQL statement is [{"Question":"This month's income","SQL":"SELECT income From Table WHERE Date='202409'"}] before the rule combination; then, the relationship between the number-checking question and the SQL statement after the rule combination includes [{"Question":"Income in September 2024","SQL":"SELECT income From Table WHERE Date='202409'"}], [{"Question":"Income in September this year","SQL":"SELECT income From Table WHERE Date='202409'"}]. That is to say, after the rule combination, a relationship between the number-checking question and the SQL statement will generate one or more relationship between the number-checking question and the SQL statement. Verifying the relationship between the number-checking question and the SQL statement includes language model verification and / or SQL statement verification. The process of language model verification includes a large model verification process, and the language model verification and SQL statement verification processes are respectively explained in the form of examples below.
[0187] As one example, the language model verification process includes: according to the prompt information, the number query problem and its corresponding SQL statement are verified through the large model. The verification process is as follows:
[0188] The information for entering the large model is shown in the following table:
[0189]
[0190] It should be understood that the purpose of the language model verification process is to use the big model to see whether the query problem and the SQL can match. At the same time, it can also make a preliminary judgment on whether the SQL statement can be executed.
[0191] In one possible implementation, language model verification mainly utilizes prompt engineering technology to guide the language model to complete query problems and SQL verification tasks. The above example only shows one of the verifications. In actual scenarios, multiple rounds of verification are performed on a pair of number-checking problems and SQL codes. For example, the number of multiple rounds of verification is 3 or 4 times, which is not limited in this application.
[0192] It should be understood that multiple rounds of verification help improve the accuracy of the verification results, thereby improving the accuracy of the generated training data and further improving the recognition efficiency of NL2SQL.
[0193] As an example, SQL validation is a syntax validation. The SQL validation process includes the following steps:
[0194] (1) Parsing SQL statements:
[0195] Use a SQL parser to parse SQL statements into an abstract syntax tree (AST). Commonly used parsers include JSqlParser, Apache Calcite, etc. The parser will decompose SQL statements into different components, such as query items, table names, conditions, etc.
[0196] (2) Syntax tree analysis:
[0197] By traversing the abstract syntax tree, check whether each part of the SQL statement conforms to the syntax rules.
[0198] Different rule checkers are applied for different types of SQL statements (such as SELECT, UPDATE, DELETE, etc.).
[0199] (3) Rule checking:
[0200] Define a series of grammatical rules, such as "query statements must contain WHERE conditions", "SELECT* cannot be used", etc. According to these rules, check the parsed syntax tree and generate a check report.
[0201] (4) Error handling:
[0202] If a syntax error is found, the error information is recorded and a corresponding error report is generated.
[0203] Provides detailed error descriptions and suggested correction methods to help developers quickly locate and fix problems.
[0204] (5) Integration and testing:
[0205] Integrate the syntax check function into the development environment or CI / CD pipeline to automatically detect syntax errors in SQL statements. Ensure the accuracy and stability of the syntax check function through unit testing and integration testing.
[0206] It should be understood that syntax checking only targets SQL codes to see if there are syntax errors or unexecutable errors. The training data that has been checked by the language model and SQL is output as the final corpus and stored in the database for training the second language model to improve the accuracy and response efficiency of the second language model in the NL2SQL recognition process.
[0207] The above describes in detail the method for generating training data provided by the present application. The following explains the purpose of the training data provided by the present application.
[0208] The present application also provides a method for applying the training data generated by the training data generation method as described above to train a second language model. The method comprises: Figure 4 The training data generated by the corresponding method flow is used to optimize the parameters of the second language model to obtain a second language model with optimized parameters; wherein the second language model is used to assist users in completing data queries.
[0209] The present application also provides a method according to the attached Figure 3 The method for executing data query using the second language model in the corresponding training data generating device is characterized in that the natural language input of the user is received, and the natural language input is converted into a SQL query statement according to the NL2SQL system and the second language model; the second language model is Figure 4 The training data generated by the method flow shown is obtained through training; the data query result is determined according to the SQL query statement.
[0210] Combine the following Figure 5-7 The computing device provided by the present application is explained.
[0211] The present application also provides a chip system, which includes a processor and a power supply circuit, wherein the power supply circuit is used to supply power to the processor, and the processor is used to execute the following steps: Figure 1 or Figure 4 The corresponding operation steps of the corresponding method are not described here for brevity. The processor can be implemented by a GPU, or by a computing device such as a DPU, NPU, XPU, SoC, offload card, accelerator card, etc.
[0212] The present application also provides a computing device 10. Figure 5 As shown, the computing device 10 includes: a bus 102, a processor 104, a memory 106, and a communication interface 108. The processor 104, the memory 106, and the communication interface 108 communicate with each other through the bus 102. The computing device 10 can be a server or a terminal device. It should be understood that the present application does not limit the number of processors and memories in the computing device 10.
[0213] Based on the received user input, one or more application programming interface (API) calls may be made between the input / output device of the computing device and the computing device. In some embodiments, the API calls may be configured for a particular API and may be interpreted and / or converted to API calls configured for a different API. As used herein, an API may refer to a defined (e.g., according to an API specification) interface or connection between computers or between computer programs.
[0214] The bus 102 may be a peripheral component interconnect Express (PCIe) bus or an extended industry standard architecture (EISA) bus, a unified bus (Ubus or UB), a compute express link (CXL), a cache coherent interconnect for accelerators (CCIX), etc. Among them, the unified bus is also called a Lingqu bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 5 The bus 104 is represented by only one line, but it does not mean that there is only one bus or one type of bus. The bus 104 may include a path for transmitting information between various components of the computing device 10 (e.g., the memory 106, the processor 104, and the communication interface 108). Among them, the unified bus may also be called a Lingqu bus.
[0215] The processor 104 may include any one or more computing devices such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP) or a digital signal processor (DSP), an ASIC, an FPGA, a CPLD, an NPU, a SoC, an offload card, an acceleration card, etc.
[0216] The memory 106 may include a volatile memory, such as a random access memory (RAM). The processor 104 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD) or a solid state drive (SSD). In addition, the memory 106 may also be implemented by a storage class memory (SCM), a phase change memory (PCM), or other types of storage media.
[0217] It is worth noting that the same type of storage medium can be configured in the same computing device to implement the function of memory 106, or two or more types of storage media can be configured to implement the function of memory 106, and this application does not limit this.
[0218] The memory 106 stores executable program codes, and the processor 104 executes the executable program codes to respectively implement the functions of the aforementioned module A, module B, and module C, thereby implementing the following: Figure 1 That is, the memory 106 stores a method for executing Figure 1 The corresponding method instruction.
[0219] Alternatively, the memory 106 stores executable codes, and the processor 104 executes the executable codes to respectively implement the functions of the aforementioned YY device, AA device, and BB device, thereby implementing the following: Figure 4 That is, the memory 106 stores a method for executing Figure 4 The corresponding method instruction.
[0220] The communication interface 103 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 10 and other devices or a communication network.
[0221] As a possible implementation, the computing device 10 may also include a chip system, the chip system includes a processor and a power supply circuit, the power supply circuit is used to supply power to the processor, and the processor is used to perform the following steps: Figure 1 or Figure 4 The operation steps corresponding to the method shown are not described here for the sake of brevity. The processor can be implemented by a GPU, or by a computing device or AI chip such as a DPU, NPU, XPU, SoC, offload card, accelerator card, etc.
[0222] As a possible implementation, the computing device 10 may include multiple types of processors 104, that is, the computing device 10 is a heterogeneous device. For example, the computing device 10 includes a CPU and a GPU, and at least one of the processors 104 may execute the following operations: Figure 1 or Figure 4 The operation steps corresponding to the method shown are not repeated here for the sake of brevity.
[0223] The embodiment of the present application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.
[0224] like Figure 6 As shown, the computing device cluster includes at least one computing device 10. The memory 106 in one or more computing devices 10 in the computing device cluster may store the same Figure 1 or Figure 4 Instructions for the method shown.
[0225] In some possible implementations, the memory 106 of one or more computing devices 10 in the computing device cluster may also store a program for executing the following steps: Figure 1 or Figure 4 In other words, a combination of one or more computing devices 10 can jointly execute instructions for performing the following method: Figure 1 or Figure 4 Instructions for the method shown.
[0226] It should be noted that the memory 106 in different computing devices 10 in the computing device cluster may store different instructions, respectively used to execute the following instructions: Figure 4 That is, the instructions stored in the memory 106 in different computing devices 10 can implement the following Figure 4 The functionality of one or more modules is shown.
[0227] In some possible implementations, one or more computing devices in the computing device cluster may be connected via a network, which may be a wide area network or a local area network. Figure 7 A possible implementation is shown. Figure 7 As shown, two computing devices 10A and 100B are connected via a network. Specifically, the network is connected via a communication interface in each computing device. In this type of possible implementation, the memory 106 in the computing device 10A stores instructions for executing the functions of the first language model 510. At the same time, the memory 106 in the computing device 10B stores instructions for executing the functions of the information extraction module 520, the rule combination module 530, and the verification module 540.
[0228] Figure 7 The connection mode between the computing device clusters shown can be considered as follows: Figure 4 The training data generation method shown in the figure requires a lot of computing resources, so consider Figure 4 Some modules in the illustrated architecture are executed by computing device 10B.
[0229] It should be understood that Figure 7 The functions of the computing device 10A shown in FIG. 10A may also be completed by multiple computing devices 10. Similarly, the functions of the computing device 10B may also be completed by multiple computing devices 10.
[0230] The present application also provides a computer program product including instructions. The computer program product may be software or a program product including instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device executes the following steps: Figure 1 The method shown or Figure 4 The method shown.
[0231] The present application also provides a computer-readable storage medium. The computer-readable storage medium may be any available medium that can be stored by a computing device or a data storage device such as a data center that contains one or more available media. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk). The computer-readable storage medium includes instructions that instruct the computing device to execute the following: Figure 1 The method shown, or instructing a computing device to execute Figure 4 The method shown.
[0232] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present invention.< / item> < / date> < / item> < / date> < / item> < / date> < / date> < / item> < / item> < / item> < / date> < / item> < / date> < / date> < / date> < / item> < / item> < / date> < / date> < / item> < / date> < / date> < / item> < / date> < / item> < / date> < / item> < / item> < / date>
Claims
1. A method for generating training data, wherein the training data is used to train a second language model, and the second language model is used to convert natural language into SQL statements, characterized in that : receiving a number-checking question and an SQL statement relationship pair, wherein the number-checking question and the SQL statement relationship pair are input corpus of a first language model; Generate one or more number-checking templates and SQL template relationship pairs, standard entities and generalized entities according to the input corpus and prompt information; Determine the position of the slot to be filled in the relationship between the one or more number-checking templates and the SQL template; A number-checking template and SQL template relationship pair includes a number-checking template and an SQL template; the positions of the slots to be filled in the one or more number-checking template and SQL template relationship pairs are used to indicate the positions of the standard entities and the generalized entities in the number-checking problem and SQL statement relationship pairs; Combining the generalized entity and the standard entity with a number-checking template and an SQL template respectively in a preset manner to generate a combined number-checking problem and SQL statement relationship pair; The number-checking problem and the SQL statement relationship pair after the combination are verified, and training data is generated according to the verification result. The training data is the number-checking problem and the SQL statement relationship pair after the combination and verification.
2. The method according to claim 1, characterized in that The generalized entity and the standard entity are combined with a number-checking template and an SQL template respectively in a preset manner to generate a combined number-checking problem and SQL statement relationship pair, including: Fill the generalized entity into the slot to be filled in the number-checking template to generate a number-checking problem; Filling the standard entity into the SQL template to generate an SQL statement; According to the generated number-checking problem and the generated SQL statement, a combined number-checking problem and SQL statement relationship pair is determined; wherein the combined number-checking problem and SQL statement relationship pair includes a number-checking problem and an SQL statement corresponding to the number-checking problem.
3. The method according to claim 2, characterized in that The step of filling the generalized entity into the slot to be filled in the number-checking template to generate the number-checking problem further includes: When there are multiple slots to be filled in the number-checking template, any one of the generalized entities of the same type is filled into the slot to be filled corresponding to the generalized entity of the same type, until each slot to be filled is filled with the generalized entity, thereby generating a number-checking problem; Traversing each generalized entity in the generalized entities of the same type, filling in the slot positions to be filled corresponding to the generalized entities of the type, until each slot position to be filled is filled with the generalized entity, generating a plurality of number-checking problems; The number of generated number-finding problems satisfies the Cartesian product.
4. The method according to claim 1, characterized in that: The prompt information includes: role definition (Role), background (Background), process (Process), examples (Examples), and specified format (Output form) for the large model.
5. The method according to claim 1, characterized in that: The checking of the combined number-checking problem and SQL statement relationship pair includes: language model checking and / or SQL statement checking.
6. A method for training a second language model using the training data according to claims 1-5, characterized in that: Optimizing the parameters of the second language model according to the training data according to claims 1-5 to obtain a second language model with optimized parameters; wherein the second language model is used to assist users in completing data queries.
7. A method for executing data query according to the second language model as claimed in claims 1-5, characterized in that: Receive natural language input from the user, converting the natural language input into a SQL query statement according to the NL2SQL system and the second language model; the second language model is trained with the training data generated by the generation method according to claims 1-5; Determine data query results in the database management system based on SQL query statements.
8. A chip system, characterized in that: The chip system includes a processor and a power supply circuit, wherein the power supply circuit is used to supply power to the processor, and the processor is used to execute the operation steps of any method described in claims 1-5.
9. A computing device, characterized in that The computing device includes a processor and a memory; the processor is used to execute instructions stored in the memory, so that the computing device performs the operating steps of any one of the methods described in claims 1-5.
10. A computing device cluster, characterized in that: comprising at least one computing device, each computing device comprising a processor and a memory; The processor of the at least one computing device is used to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster performs the operation steps of any one of the methods according to claims 1-5.
11. A computer program product comprising instructions, characterized in that When the instruction is executed by the computing device cluster, the computing device cluster executes the operation steps of any one of the methods described in claims 1-5.
12. A computer-readable storage medium, characterized in that: The method comprises computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster executes the operation steps of the method according to any one of claims 1 to 5.
Citation Information
Cited By
Text-to-SQL (Structured Query Language) full-link acquisition method
CN120705250A
Training data set synthesis method, device and equipment based on error driving
CN121434795A