Database query method, system and equipment based on large model data distillation and medium
By applying large-model data distillation technology in database queries, the knowledge of large-models is migrated to lightweight models, and the problems of database query efficiency and resource consumption are solved, and efficient and accurate query performance and low resource consumption are achieved.
Patent Information
- Application Number
- CN202510152720.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art is inefficient and poorly adaptable in database queries, especially in complex queries and massive data environments, and large-scale model calculations are complex and resource consumption is high, making it difficult to apply in real-time queries and resource-constrained environments.
Using a method based on large model data distillation, the knowledge of the large model is transferred to a lightweight model, and the student model is transferred through the teacher model and trained in the loss function, efficient SQL statements are generated and database query execution plan is optimized.
It realizes efficient and accurate database queries, adapts to different databases and query scenarios, especially in resource-constrained environments, and provides an efficient and resource-friendly solution, which significantly reduces the hardware resource requirements.
Smart Images

Figure CN120104633A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of database technology, and in particular to a database query method, system, device and medium based on large model data distillation. Background Art
[0002] With the advent of the big data era, database queries have become crucial in applications such as data analysis and real-time retrieval. However, traditional SQL query methods often show problems of inefficiency and poor adaptability when faced with complex queries and massive data. Large-scale pre-trained language models (such as GPT, etc.) can convert natural language queries into SQL statements and optimize query plans due to their powerful semantic understanding capabilities, providing a smarter query method. However, existing large models are computationally complex and resource-intensive, making them difficult to be widely used in real-time queries and resource-constrained environments.
[0003] To solve this problem, model distillation technology came into being. By migrating the knowledge of large models to lightweight models, distillation technology reduces computing resource consumption while maintaining high performance. Although this technology has achieved certain results in image recognition, natural language processing and other fields, its application in database queries is still relatively limited. Therefore, how to use distillation technology to extract high-quality training data, optimize complex queries, and improve execution efficiency is a technical problem that needs to be solved urgently. Summary of the invention
[0004] The technical task of the present invention is to provide a database query method, system, device and medium based on large model data distillation to solve the problem of how to use distillation technology to extract high-quality training data, optimize complex queries, and thus improve execution efficiency.
[0005] The technical task of the present invention is achieved in the following way: a database query method based on large model data distillation, the method is as follows:
[0006] Based on the pre-trained large language model, knowledge related to database queries is extracted from multiple data sources to generate high-quality training datasets;
[0007] Distillation technology is used to transfer knowledge from the teacher model (large model) to the student model (lightweight model), and the lightweight model is trained by optimizing the loss function, so that the student model has stronger reasoning ability in database query tasks;
[0008] The student model parses the natural language query input by the user and generates the corresponding structured query language statement, i.e., SQL statement;
[0009] Based on the prediction results of the student model, the database execution plan is optimized to improve query efficiency;
[0010] The content returned by the query is verified and adjusted through the model prediction results to ensure the accuracy of the query results and the matching degree with user needs.
[0011] As a preferred method, based on the pre-trained large language model, knowledge related to database queries is extracted from multiple data sources to generate a high-quality training data set as follows:
[0012] Data extraction: Collect raw data from multiple data sources including open datasets, business logs, and domain-specific corpora;
[0013] Data annotation: annotate the original data using annotation tools, that is, add semantic labels and structured information to the data to generate training samples;
[0014] Data cleaning and filtering: remove duplicate data, correct data noise and errors from the original data, and filter out samples that are irrelevant to database queries or are of low quality to obtain high-quality training samples;
[0015] Data enhancement: Generate diverse samples through data expansion techniques such as synonym replacement, data reorganization, and fuzzy matching to improve the generalization ability of the model.
[0016] Preferably, the student model is constructed as follows:
[0017] A lightweight neural network architecture is used as the student model to reduce model complexity and inference time; the lightweight neural network architecture uses a Transformer variant or TinyBERT;
[0018] The student model is further compressed through pruning and quantization techniques to improve operating efficiency.
[0019] A database query system based on large model data distillation, the system comprising:
[0020] A data set generation module is used to extract knowledge related to database queries from multiple data sources and generate training data sets;
[0021] The distillation module is used to transfer the knowledge of the large model to the lightweight model, and improve the query capability of the model by optimizing the training process;
[0022] The query parsing module is used to convert the natural language query input by the user into a structured query language, i.e., SQL statement;
[0023] Query optimization module, used to optimize database query execution plan based on generated SQL statements;
[0024] The result verification module is used to verify the accuracy of the query results and the matching degree with user requirements.
[0025] Preferably, the data set generation module includes:
[0026] The data extraction submodule is used to collect raw data from multiple data sources such as open datasets, business logs, and domain-specific corpora;
[0027] The data annotation submodule is used to annotate the original data through annotation tools, that is, to add semantic labels and structured information to the data and generate training samples;
[0028] The data cleaning and filtering submodule is used to remove duplicate data, correct data noise and error correction from the original data, and filter out samples that are irrelevant to database queries or are of low quality to obtain high-quality training samples;
[0029] The data enhancement submodule is used to generate diverse samples through data expansion techniques such as synonymous replacement, data reorganization and fuzzy matching, thereby improving the generalization ability of the model.
[0030] Preferably, the lightweight model is constructed as follows:
[0031] A lightweight neural network architecture is used as the student model to reduce model complexity and inference time; the lightweight neural network architecture uses a Transformer variant or TinyBERT;
[0032] The student model is further compressed through pruning and quantization techniques to improve operating efficiency.
[0033] An electronic device comprising: a memory and at least one processor;
[0034] Wherein, the memory stores a computer program;
[0035] The at least one processor executes the computer program stored in the memory, so that the at least one processor performs the database query method based on large model data distillation as described above.
[0036] A computer-readable storage medium having a computer program stored therein, wherein the computer program can be executed by a processor to implement the database query method based on large model data distillation as described above.
[0037] The database query method, system, device and medium based on large model data distillation of the present invention have the following advantages:
[0038] (i) The present invention transforms the knowledge of large models into lightweight models through distillation technology, realizes efficient and accurate query execution, adapts to different databases and query scenarios, and provides an efficient and resource-friendly solution, especially in resource-constrained environments;
[0039] (ii) The present invention overcomes the shortcomings of the prior art in terms of database query efficiency, accuracy and model complexity. By utilizing the powerful representation ability and knowledge transfer mechanism of the large language model, the knowledge of the large model is refined and transferred to the lightweight model through data distillation technology, thereby achieving efficient, accurate and low computing resource consumption support for database query;
[0040] (III) While optimizing database query performance, the present invention significantly reduces hardware resource requirements, making it suitable for data query tasks in multiple scenarios and fields;
[0041] (IV) The present invention also has the following advantages:
[0042] ① High efficiency distillation technology significantly reduces model complexity, enabling lightweight models to quickly parse and execute database queries, shortening response time, and query optimization effectively reducing database resource consumption and improving overall operating efficiency;
[0043] ② Accuracy: The knowledge representation capability of the large model is inherited during the distillation process. The lightweight model performs well in understanding user query intent and generating high-quality SQL statements, and further improves the reliability of query results through the result verification module;
[0044] ③ Low resource consumption: The reasoning requirements of lightweight models are significantly lower than those of large models. They are suitable for resource-constrained devices (such as edge computing devices and mobile devices). They also optimize model size and computing efficiency through pruning and quantization techniques, thereby reducing hardware costs.
[0045] ④Multi-scenario adaptability: It supports multiple database types and query scenarios, and can meet the diverse needs of enterprise data analysis, real-time retrieval, etc. The dynamic learning mechanism enables the model to be continuously optimized as user needs and data changes;
[0046] ⑤ Easy to promote: The method proposed in the present invention does not require retraining of large language models, but only requires a one-time distillation process. The generated lightweight model is universal and easy to use, and is suitable for large-scale promotion and application;
[0047] (V) The lightweight model of the present invention further compresses the computational complexity of the student model by means of pruning and quantization;
[0048] (vi) The present invention uses natural language processing technology based on deep learning to convert users' natural language queries into corresponding SQL statements;
[0049] (VII) The present invention performs consistency check by comparing the query results generated by the lightweight model with the actual returned results to ensure the accuracy of the query results;
[0050] (VIII) The present invention performs online learning based on user query logs and continuously optimizes the lightweight model to adapt to changes in different fields and query scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] The present invention is further described below in conjunction with the accompanying drawings.
[0052] Attached Figure 1 The flowchart of the database query method based on large model data distillation. DETAILED DESCRIPTION
[0053] The database query method, system, device and medium based on large model data distillation of the present invention are described in detail below with reference to the drawings and specific embodiments of the specification.
[0054] Embodiment 1:
[0055] As attached Figure 1 As shown, this embodiment provides a database query method based on large model data distillation, and the method is specifically as follows:
[0056] S1. Based on the pre-trained large language model, knowledge related to database queries is extracted from multiple data sources to generate high-quality training datasets;
[0057] S2. Use distillation technology to transfer knowledge from the teacher model (large model) to the student model (lightweight model), and train the lightweight model by optimizing the loss function, so that the student model has strong reasoning ability in database query tasks;
[0058] S3, parsing the natural language query input by the user through the student model and generating the corresponding structured query language statement, i.e., SQL statement;
[0059] S4. Based on the prediction results of the student model, the database execution plan is optimized to improve query efficiency;
[0060] S5. Verify and adjust the query returned content through model prediction results to ensure the accuracy of the query results and the matching degree with user needs.
[0061] In step S1 of this embodiment, based on the pre-trained large language model, knowledge related to database query is extracted from multiple data sources to generate a high-quality training data set as follows:
[0062] S101, Data Extraction: Collect raw data from multiple data sources including open datasets, business logs, and domain-specific corpora;
[0063] S102, data annotation: annotate the original data using annotation tools, that is, add semantic labels and structured information to the data to generate training samples;
[0064] S103, data cleaning and filtering: remove duplicate data, correct data noise and error correction from the original data, and filter out samples that are irrelevant to database query or low-quality samples to obtain high-quality training samples;
[0065] S104, Data enhancement: Generate diverse samples through data expansion techniques such as synonym replacement, data reorganization and fuzzy matching to improve the generalization ability of the model.
[0066] The student model construction in step S2 of this embodiment is specifically as follows:
[0067] S201. Use a lightweight neural network architecture as the student model to reduce model complexity and inference time; the lightweight neural network architecture uses a Transformer variant or TinyBERT;
[0068] S202. Further compress the student model through pruning and quantization technology to improve operation efficiency.
[0069] This embodiment is also suitable for multi-scenario adaptation and iterative optimization, as follows:
[0070] a) Support multiple database types (such as relational databases, NoSQL databases) and multiple application scenarios (such as data analysis, text retrieval).
[0071] b) Combined with the online learning mechanism, the lightweight model is dynamically updated according to the user query log to improve the adaptability of the model to specific fields.
[0072] Embodiment 2:
[0073] This embodiment provides a database query system based on large model data distillation, the system comprising:
[0074] A data set generation module is used to extract knowledge related to database queries from multiple data sources and generate training data sets;
[0075] The distillation module is used to transfer the knowledge of the large model to the lightweight model, and improve the query capability of the model by optimizing the training process;
[0076] The query parsing module is used to convert the natural language query input by the user into a structured query language, i.e., SQL statement;
[0077] Query optimization module, used to optimize database query execution plan based on generated SQL statements;
[0078] The result verification module is used to verify the accuracy of the query results and the matching degree with user requirements.
[0079] The data set generation module in this embodiment includes:
[0080] The data extraction submodule is used to collect raw data from multiple data sources such as open datasets, business logs, and domain-specific corpora;
[0081] The data annotation submodule is used to annotate the original data through annotation tools, that is, to add semantic labels and structured information to the data and generate training samples;
[0082] The data cleaning and filtering submodule is used to remove duplicate data, correct data noise and error correction from the original data, and filter out samples that are irrelevant to database queries or are of low quality to obtain high-quality training samples;
[0083] The data enhancement submodule is used to generate diverse samples through data expansion techniques such as synonymous replacement, data reorganization and fuzzy matching, thereby improving the generalization ability of the model.
[0084] The lightweight model construction in this embodiment is specifically as follows:
[0085] ① Use a lightweight neural network architecture as the student model to reduce model complexity and inference time; the lightweight neural network architecture uses a Transformer variant or TinyBERT;
[0086] ② Further compress the student model through pruning and quantization technology to improve operation efficiency.
[0087] Embodiment 3:
[0088] This embodiment also provides an electronic device, including: a memory and at least one processor;
[0089] Wherein, the memory stores computer-executable instructions;
[0090] The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the database query method based on large model data distillation in any embodiment of the present invention.
[0091] The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor may be a microprocessor or any conventional processor, etc.
[0092] The memory can be used to store computer programs and / or modules. The processor realizes various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory, and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application required for a function, etc.; the data storage area can store data created according to the use of the terminal, etc. In addition, the memory can also include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (SMC), a secure digital (SD) card, a flash memory card, at least one disk storage period, a flash memory device, or other volatile solid-state storage devices.
[0093] Embodiment 4:
[0094] This embodiment also provides a computer-readable storage medium, which stores a plurality of instructions, which are loaded by a processor to enable the processor to execute the database query method based on large model data distillation in any embodiment of the present invention. Specifically, a system or device equipped with a storage medium can be provided, on which a software program code that implements the functions of any of the above embodiments is stored, and a computer (or CPU or MPU) of the system or device reads and executes the program code stored in the storage medium.
[0095] In this case, the program code itself read from the storage medium can realize the function of any one of the above-mentioned embodiments, and thus the program code and the storage medium storing the program code constitute a part of the present invention.
[0096] The storage medium embodiments for providing the program code include a floppy disk, a hard disk, a magneto-optical disk, an optical disk (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RYM, DVD-RW, DVD+RW), a magnetic tape, a non-volatile memory card, and a ROM. Alternatively, the program code can be downloaded from a server computer via a communication network.
[0097] In addition, it should be clear that the functions of any of the above embodiments can be implemented not only by executing the program code read by the computer, but also by enabling an operating system operating on the computer to complete part or all of the actual operations based on instructions from the program code.
[0098] In addition, it can be understood that the program code read from the storage medium is written to a memory provided in an expansion board inserted into the computer or written to a memory provided in an expansion unit connected to the computer, and then based on the instructions of the program code, a CPU installed on the expansion board or the expansion unit is enabled to perform part or all of the actual operations, thereby realizing the functions of any of the above-mentioned embodiments.
[0099] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A database query method based on large model data distillation, characterized in that: The method is as follows: Based on the pre-trained large language model, knowledge related to database queries is extracted from multiple data sources to generate high-quality training datasets; Use distillation technology to transfer knowledge from the teacher model to the student model, and train the lightweight model by optimizing the loss function; The student model parses the natural language query input by the user and generates the corresponding structured query language statement, i.e., SQL statement; Based on the prediction results of the student model, the database execution plan is optimized to improve query efficiency; The content returned by the query is verified and adjusted through the model prediction results to ensure the accuracy of the query results and the matching degree with user needs.
2. The database query method based on large model data distillation according to claim 1, characterized in that: Based on the pre-trained large language model, knowledge related to database queries is extracted from multiple data sources to generate high-quality training data sets as follows: Data extraction: Collect raw data from multiple data sources including open datasets, business logs, and domain-specific corpora; Data annotation: annotate the original data using annotation tools, that is, add semantic labels and structured information to the data to generate training samples; Data cleaning and filtering: remove duplicate data, correct data noise and errors from the original data, and filter out samples that are irrelevant to database queries or are of low quality to obtain high-quality training samples; Data enhancement: Generate diverse samples through data expansion techniques such as synonym replacement, data reorganization, and fuzzy matching.
3. The database query method based on large model data distillation according to claim 1 or 2, characterized in that: The student model is constructed as follows: A lightweight neural network architecture is used as the student model; wherein the lightweight neural network architecture uses a Transformer variant or TinyBERT; The student model is further compressed through pruning and quantization techniques.
4. A database query system based on large model data distillation, characterized in that: The system includes: A data set generation module is used to extract knowledge related to database queries from multiple data sources and generate training data sets; The distillation module is used to transfer the knowledge of the large model to the lightweight model, and improve the query capability of the model by optimizing the training process; The query parsing module is used to convert the natural language query input by the user into a structured query language, i.e., SQL statement; Query optimization module, used to optimize database query execution plan based on generated SQL statements; The result verification module is used to verify the accuracy of the query results and the matching degree with user requirements.
5. The database query system based on large model data distillation according to claim 4 is characterized in that: The dataset generation module includes: The data extraction submodule is used to collect raw data from multiple data sources such as open datasets, business logs, and domain-specific corpora; The data annotation submodule is used to annotate the original data through annotation tools, that is, to add semantic labels and structured information to the data and generate training samples; The data cleaning and filtering submodule is used to remove duplicate data, correct data noise and error correction from the original data, and filter out samples that are irrelevant to database queries or are of low quality to obtain high-quality training samples; The data enhancement submodule is used to generate diverse samples through data expansion techniques such as synonymous replacement, data reorganization, and fuzzy matching.
6. The database query system based on large model data distillation according to claim 4 or 5, characterized in that: The lightweight model is constructed as follows: A lightweight neural network architecture is used as the student model; wherein the lightweight neural network architecture uses a Transformer variant or TinyBERT; The student model is further compressed through pruning and quantization techniques.
7. An electronic device, characterized in that: include: memory and at least one processor; Wherein, the memory stores a computer program; The at least one processor executes the computer program stored in the memory, so that the at least one processor performs the database query method based on large model data distillation as described in any one of claims 1 to 4.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which can be executed by a processor to implement the database query method based on large model data distillation as described in any one of claims 1 to 4.
Citation Information
Cited By
Knowledge distillation-based database test method, system, equipment and medium
CN122019397A
Database testing method, system, device and medium based on knowledge distillation
CN122019397B
Steel industry-oriented large model training data construction method and system
CN122491431A