Database execution code automatic generation method and system based on large language model

By introducing a large language model, the database execution code is automatically generated and adaptively optimized, solving the problems of generalized execution code and unutilized runtime information in existing technologies. This improves query execution efficiency and intelligence, and reduces system complexity and maintenance costs.

CN121478802APending Publication Date: 2026-02-06ZHEJIANG UNIV

Patent Information

Application Number
CN202610031027.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-12
Publication Date
2026-02-06

AI Technical Summary

Technical Problem

Existing database execution frameworks suffer from severe code generalization, an inability to dynamically generate optimal execution logic for specific queries, underutilization of runtime information, a lack of adaptive code generation mechanisms, and difficulty in scaling operator fusion technology, leading to code bloat and increased system complexity.

Method used

By introducing large language models (such as ChatGPT and Qwen) to establish intelligent mapping, the system can adaptively generate query-specific execution code and achieve automatic code generation, semantic consistency verification, and performance optimization through performance feedback and self-optimization mechanisms.

Benefits of technology

It improves database execution efficiency and intelligence, reduces development and maintenance costs, has continuous self-learning and performance evolution capabilities, and solves the performance bottlenecks and code bloat problems in existing technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121478802A_ABST
    Figure CN121478802A_ABST
Patent Text Reader

Abstract

The invention discloses a database execution code automatic generation method and system based on a large language model, and the method specifically comprises the steps: carrying out the semantic analysis and logic mapping of an SQL query statement through a large language model, and generating a corresponding executable query code through a code template, and self-adaptive optimization of codes is realized by combining semantic consistency verification and a performance feedback mechanism. According to the method, special execution logic can be dynamically generated according to different query semantics and data features, intelligent generation and performance evolution of database execution codes are achieved, a system can automatically generate specific high-performance execution logic according to specific queries, and a closed loop is formed; the query efficiency and the system expandability are remarkably improved, the manual development and maintenance cost is reduced, and the method has good universality and reliability.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of artificial intelligence and database technology, and particularly relates to a database execution code automatic generation method and system based on a large language model (LLM). BACKGROUND

[0002] In modern database systems, the execution of query statements usually relies on the operator and function call framework predefined in the database kernel. For any SQL query, the system will call the corresponding operator module to complete the execution according to the logical plan and physical plan generated by the optimizer. In order to cope with different query forms and data characteristics, the execution code of the database kernel is usually designed in a highly generalized form, supporting diversified execution scenarios through multi-layer function calls, template distribution and branch judgment. However, this universal design has a natural contradiction between flexibility and performance.

[0003] For different types of queries (such as aggregation, join, sorting, etc.), the database usually uses the same operator framework to complete the execution logic through condition judgment, function dispatching and template expansion. In previous studies, in order to reduce the overhead of function call level and branch judgment, the academia and industry have proposed JIT (Just-In-Time) code generation technology, which dynamically compiles part of the operator logic into machine code at runtime to speed up the execution process. Although JIT technology can effectively reduce the function call and branch jump of the general layer, it still stays at the level of local optimization, i.e. only the internal flow of the general operator is dynamically compiled, and it cannot customize a set of optimal execution code structure for specific queries. Therefore, when facing complex queries or multi-operator fusion scenarios, the existing system still introduces a lot of redundant logic and execution overhead.

[0004] During the query execution phase, the database can often obtain rich runtime information, however, the existing database execution framework mostly runs in a general execution path, lacking a mechanism to directly map these runtime information to the code layer. This results in the system being unable to generate truly specialized execution code for each query, and still performing calculations in a general mode, thus wasting the potential performance optimization space.

[0005] Meanwhile, in recent years, to further reduce the boundary overhead between operators and intermediate materialization, a large number of operator fusion (Operator Fusion) technologies have appeared in database research, such as GroupJoin, AggJoin, etc. These methods fuse multiple operators into a single execution unit at the physical layer, thereby improving the efficiency of pipelined execution. However, when such fusion logic is directly integrated into the database kernel, it will cause serious code bloat problems. Each fusion mode needs to define new operator implementations or template instances in the source code, which leads to rapid increase in kernel complexity, significant increase in maintenance and debugging costs, and difficulty in sharing and reusing these optimization logic between different database environments or versions. Therefore, although the operator fusion technology has significant acceleration potential in theory, due to the complexity of implementation and poor maintainability, it has not been widely used in mainstream database systems.

[0006] In summary, the existing database execution framework mainly has the following limitations: first, the execution code is severely generalized and cannot dynamically generate optimal execution logic for specific queries; second, the runtime information is not fully utilized, and there is a lack of adaptive code generation mechanism based on data distribution or statistical characteristics; third, the fused operator mechanism is difficult to extend, leading to code bloat and system complexity while pursuing performance optimization. SUMMARY

[0007] In view of the deficiencies of the prior art, the purpose of the present application is to provide a database execution code automatic generation method and system based on a large language model. The method specifically introduces a large language model (such as ChatGPT, Qwen, etc.) with strong semantic understanding and code generation capability, and the present application can establish an intelligent mapping between the query layer, the semantic layer and the code generation layer, realize the automatic generation of database execution logic, semantic consistency verification and performance self-optimization, thereby improving the execution efficiency and intelligent level of the database.

[0008] The purpose of the present application is achieved by the following technical solutions: The first aspect of the present application provides a database execution code automatic generation method based on a large language model, comprising the following steps: Step one: according to the pre-defined query template in the user workload, decompose the query logic of the query template and generate an execution plan, collect database schema information, data distribution information and index feature information corresponding to the query template; Step two: use a large language model to adaptively generate database execution code matching the query semantics in combination with the query template and the collected information; Step three: error analysis and feedback for errors occurred during the execution of the generated query-specific code, including compilation errors, execution failures and execution result errors, input the feedback optimization instructions into the large language model for adjustment or re-generation of the code, and output the corrected code; Step four: monitor the running performance of the corrected code, when detecting that the performance of the code is still lower than that of the database native execution engine, input the performance indicators and performance bottleneck characteristics as feedback into the large language model, generate a performance optimization scheme by the model and re-generate the optimized execution code to realize performance feedback and self-optimization.

[0009] Further, the step one specifically includes the following sub-steps: (1.1) Query template input: receiving a pre-defined structured query template as input basis; the template is a set of structured, parameterized and query-intended query templates with basic structure and framework, wherein one template represents a query logical structure, and contains placeholder variables inside; the placeholder variables include filter conditions, aggregation functions, projection columns, grouping keys, sorting or limiting parameters for describing query behavior and instantiable parameters; (1.2) Query logic decomposition and execution plan generation: parsing query logic and identifying semantic units and operation sequences; generating an intermediate expression structure, and formulating a preliminary query execution path according to the logical structure to guide the subsequent code generation process; (1.3) Database structure perception and schema collection: connecting the database instance through the database schema collection module, extracting meta-information including current data table structure, field type and index distribution; providing context reference for code generation to ensure compatibility between the code and the actual database structure.

[0010] Further, the step two is implemented through the following sub-steps: (2.1) Input query execution plan: receiving the query execution plan, which includes the following logical operations: node scanning, filtering, joining, sorting, aggregation, projection, deduplication, set operation, limitation, window function and output; (2.2) Dynamic workflow generation: according to the operation types in the query execution plan, automatically dispatching task agents Agent for each logical operation, each Agent is responsible for scanning, filtering, joining, sorting, aggregation, projection, deduplication, set operation, window function and output stage workflow modeling, and submitting the structured operation process to the next step for code generation; (2.3) Execution code snippet generation: generating corresponding code snippets according to the task definition of each Agent, and automatically adding necessary logical processing statements during the generation process to ensure semantic integrity; (2.4) Execute code combination: the generated multiple code snippets are spliced into a complete execution function according to the preset structure sequence by the template engine, and the combination process includes fragment injection, template driving and structure integration to ensure that the generated code meets the compilation specification; (2.5) Embedding and execution: the complete execution code combined is automatically embedded in the project source code, and the compilation is completed through the project build system. After successful compilation, the structured query request is automatically executed and the result is output.

[0011] Further, each code snippet in step (2.4) usually includes the following five parts: (a) Headers: import library or structure definition; (b) Execute_code: main execution logic; (c) Finalize_code: result tail and cleanup; (d) Support_code: auxiliary function; (e) Bind_code: variable and parameter binding logic.

[0012] Further, the combination process in step (2.4) includes fragment injection, template driving and structure integration, which specifically includes: the fragment injection, embedding each functional fragment into the code template; the template driving, ensuring that the function structure, sequence, and nesting relationship meet the compilation standard; the structure integration, generating the final executable complete query function.

[0013] Further, step three is implemented through the following sub-steps: (3.1) Execute query function code: execute the query-specific execution code generated by the automatic generation process and obtain the query result; (3.2) Error detection and exception identification: when execution fails or the result is abnormal, collect error logs, exception stacks and corresponding error types, identify error root causes, including syntax errors, field mismatches or null pointer exceptions, and extract correction features as feedback input; (3.3) Large language model optimization instruction generation: feedback the error correction features to the large language model for optimization, that is, construct optimization instructions according to the feedback feature content, and call the large language model to correct the query function code.

[0014] Further, step four is implemented through the following sub-steps: (4.1) Performance data collection and analysis: when the execution is successful, collect the performance indicators of the code running, including response time, CPU occupancy and memory consumption, and evaluate the efficiency of the execution plan, identify inefficient parts, and extract the corresponding optimization features; the inefficient parts include redundant sorting, full table scanning or non-optimal JOIN strategy; (4.2) Large language model optimization instruction generation: feedback the error correction features and performance optimization features to the large language model for optimization, that is, construct optimization instructions according to the feature content, and call the large language model to rewrite or optimize the query function code; (4.3) Output optimized code and replace old version: output the optimized query execution code as a new version, replace the original inefficient or error version, and re-submit the execution process for verification; if the verification result still does not meet the performance standard, repeat the above steps until the preset performance threshold is reached, thereby realizing the code self-learning and continuous optimization process based on performance feedback.

[0015] The second aspect of the present application: a large language model-based database execution code automatic generation system is provided, which is used to realize the large language model-based database execution code automatic generation method, and the system specifically comprises an SQL query statement input unit, a pre-defined query template set module, a large language model module, a database execution code automatic generation module, a query-specific execution code template module, a semantic consistency and safety verification module, a performance feedback and self-optimization module, a database instance module and a query result output module: The SQL query statement input unit supports users to input standard SQL statements based on pre-defined query templates through text boxes, command lines or API interfaces; the input unit has basic syntax detection and format standardization processing capability to ensure that the input content meets the analysis conditions; The pre-defined query template set module maintains a set of structured and parameterized query templates, each template representing a query logic structure and containing placeholder variables inside; The large language model module performs semantic analysis on the user input query statement, including extracting query purpose, fields and conditions; and maps the query to the pre-defined executable function in the system, while outputting the required parameters for subsequent calling; in addition to semantic recognition, the module is also used to handle abnormal situations, and can automatically generate modification suggestions or corrected code when the query statement has errors or unclear semantics; when the query has performance bottlenecks, it can propose targeted optimization suggestions based on historical execution features or query structure;

[0016] The database execution code automatic generation module is responsible for constructing a prompt for an input large language model, generating input content for calling the large language model in combination with a semantic structure of a user query statement and pre-defined template information; through a carefully designed prompt structure, guiding the large language model to generate an executable query function code, and realizing automatic logical expression of a query request; the generated function will encapsulate query logic, parameter interface and return structure according to system requirements, and has direct invocability; The query-specific execution code template module is used for storing a code function template generated by the large language model according to a query type, the template defining a function structure, a parameter form and a logic framework, and not containing specific parameters; the template serves as a basis for generating an executable query function, and is executed after filling parameters according to a query statement, ensuring code structure specification, consistency and adaptation to query semantics; The semantic consistency and security checking module is used for checking whether a code function generated by the large language model accurately expresses an original query intention; first, the generated function is actually executed, and an execution result of the same query by a database native execution engine is compared to judge semantic consistency; the module also includes a security checking mechanism, detecting whether there is a SQL injection risk or a security risk of permission boundary access in the code, to ensure that the automatically generated code can be trusted in terms of logic and security; The performance feedback and self-optimization module performs performance monitoring and evaluation on an execution process of the generated query function, including response time, resource consumption and execution plan; if the generated code function is inferior to a database native execution result in performance, the system will reconstruct a prompt according to performance data, and call the large language model again to generate an optimized code; the module also builds a closed-loop optimization mechanism based on feedback, for continuously improving running efficiency and resource utilization of the system generated code; The database instance module is a database system actually deployed, used for executing query statements and generating functions, and supporting mainstream databases; the database instance not only bears an execution benchmark of original SQL statements, but also serves as a running environment of the generated functions, providing standardized execution results for comparison and optimization; The query result output module is responsible for unified display of query result outputs of the system, including two parts: one is a query result of a database native execution engine, serving as a benchmark for accuracy and performance; the other is an output result of a code function generated and executed by the large language model, used for evaluating effectiveness and replaceability of the generated logic; the result supports multiple format display, and is accompanied by performance comparison reports and semantic explanation information.

[0017] The third aspect of the present application provides an electronic device, comprising a memory and a processor, wherein the memory is coupled with the processor; the memory is used for storing program data, and the processor is used for executing the program data to realize the database execution code automatic generation method based on a large language model.

[0018] The fourth aspect of the present application provides a computer readable storage medium, which stores a computer program, wherein the program is executed by a processor to realize the database execution code automatic generation method based on a large language model.

[0019] The present application has the following beneficial effects: By introducing a large language model with semantic understanding and code generation capability, the automatic generation and adaptive optimization of database query execution code are realized, which can dynamically generate special execution logic according to different SQL query semantics and data characteristics, thereby breaking through the performance bottleneck caused by the generalization of the traditional database execution framework. This method not only significantly improves the query execution efficiency and the intelligent level of code generation, but also ensures the correctness and reliability of the generated code through semantic consistency comparison and security verification mechanism. At the same time, the system constructs a closed-loop optimization mechanism based on performance feedback, so that the generated code has the ability of continuous self-learning and performance evolution, reducing the development and maintenance cost of the database system, and having good universality and scalability. BRIEF DESCRIPTION OF DRAWINGS

[0020] Figure 1 is the overall structure diagram of the database execution code automatic generation system of the present application; Figure 2 is the flowchart of the SQL execution code automatic generation method of the present application; Figure 3 is the flowchart of the query execution plan driven dynamic workflow and execution code generation of the present application; Figure 4 is the performance feedback and self-optimization module diagram of the present application. DETAILED DESCRIPTION

[0021] In order to make the purpose, features and advantages of the present application more clear and easy to understand, the specific embodiments of the present application are described in detail below with reference to the drawings. It should be understood that the following embodiments are only used to illustrate the present application, and do not constitute a limitation on the protection scope of the present application. Equivalent replacements or modifications made by those skilled in the art without departing from the spirit and essence of the present application shall be covered within the protection scope of the present application.

[0022] This embodiment provides a method and system for automatically generating database execution code by combining a large language model. Through large language model analysis, code template generation, and performance feedback mechanisms, it effectively improves the generation efficiency and running performance of database query code.

[0023] Example 1: Automatic Database Execution Code Generation System Based on Large Language Model

[0024] like Figure 1 As shown, this system mainly includes the following modules: (1) SQL query statement input unit This module allows users to input standard SQL statements based on predefined query templates via text boxes, command lines, or API interfaces. The input unit has basic syntax checking and format standardization capabilities to ensure that the input content meets parsing conditions and prevent semantic ambiguity or format abnormalities from affecting subsequent processing. (2) Predefined query template set module This module maintains a set of structured, parameterized query templates. Each template represents a common query logic structure and contains placeholder variables (such as filtering conditions). (3) Large Language Model Module This module performs semantic parsing on user-input queries, extracting elements such as query purpose, fields, and conditions, and mapping the query to predefined executable functions in the system. It also outputs the necessary parameters for subsequent calls. In addition to semantic recognition, this module handles exceptions. When a query contains errors or is semantically unclear, it can automatically generate modification suggestions or corrected code. When query performance bottlenecks occur, it can also provide targeted optimization suggestions based on historical execution characteristics or query structure, helping the system improve execution efficiency and stability. (4) Database execution code automatic generation module This module is responsible for constructing prompts for the input large language model. Combining the semantic structure of the user's query with predefined template information, it generates input content for calling the large language model. Through a carefully designed prompt structure, it guides the large language model to generate executable query function code, achieving automated logical expression of the query request. The generated function will encapsulate the query logic, parameter interfaces, and return structure according to system requirements, ensuring direct callability. (5) Query specific executable code template module This module stores code function templates generated by the large language model based on the query type. The templates define the function structure, parameter formats, and logical framework, but do not yet include specific parameters. These templates serve as the basis for generating executable query functions. After filling in the parameters according to the query statement, the functions are executed, ensuring that the code structure is standardized, consistent, and adapted to the query semantics. (6) Semantic consistency and security checking module This module is used to check whether the code function generated by the large language model accurately represents the original query intent. Its method is to actually execute the generated function and compare it with the execution result of the same query by the database native execution engine to judge semantic consistency. At the same time, this module also includes a security checking mechanism to detect whether there are SQL injection risks, permission boundary access and other security risks in the code, to ensure that the automatically generated code can be trusted in terms of logic and security; (7) Performance feedback and self-optimization module This module monitors and evaluates the performance of the generated query function execution process, including response time, resource consumption, execution plan and other indicators. If the generated code function is inferior to the database native execution result in performance, the system will reconstruct the prompt according to the performance data and call the large language model again to generate optimized code. This module builds a closed-loop optimization mechanism based on feedback, continuously improving the running efficiency and resource utilization of the system generated code; (8) Database instance module This module is for the actual deployed database system to execute query statements and generate functions. It supports mainstream databases such as DuckDB, MySQL, PostgreSQL, SQL Server, etc. The database instance not only serves as the benchmark for the execution of the original SQL statement, but also as the running environment for the generated function, providing standardized execution results for comparison and optimization; (9) Query result output module This module is responsible for unified display of the query result output of the system, including two parts: one is the query result of the database native execution engine, which serves as the accuracy and performance benchmark; the other is the output result of the code function generated and executed by the large language model, which is used to evaluate the effectiveness and replaceability of the generation logic. The results can be displayed in multiple formats such as tables, JSON, logs, etc., and can also be accompanied by performance comparison reports, semantic interpretation information, etc., to improve the transparency and user interpretability of the system.

[0025] II. The specific operation process of the system includes query template driven code function generation and user query statement driven execution and optimization process: The query template driven code function generation process is as follows:

[0026] The system selects a structured query template from the query template set as input to the database execution code automatic generation module, which is responsible for driving the entire function generation process.

[0027] The database execution code automatic generation module constructs a Prompt based on the received query template content and calls the large language model module. The model generates a standard code function template corresponding to the query template. The generated result includes the function name, parameter structure, and query logic body, and is stored in the query-specific execution code template module for subsequent invocation and instantiation in specific queries.

[0028] The execution and optimization process driven by the user query statement is as follows: When a user submits a specific SQL query statement through the query statement input unit, the system automatically matches the corresponding query template and its generated function template, extracts the actual parameters from the SQL statement, instantiates the function template, and generates an executable function with complete logic and parameters.

[0029] The system calls the instantiated query function, executes the function in the database instance module, obtains the query results, and outputs the results to the query result output module for subsequent display and comparative analysis.

[0030] The semantic consistency and security verification module verifies the execution results by comparing the function execution results with the results of executing the same SQL statement directly in the database's native engine. It determines the consistency between the two in terms of semantic expression and data results, and detects the security of the function code, including risks such as SQL injection and privilege escalation.

[0031] The performance feedback and self-optimization module further analyzes the performance metrics of function execution, such as response time and resource consumption. If the generated function performs worse than the database's native execution path, the system will reconstruct the Prompt based on the feedback, call the large language model to optimize and rewrite the function structure, forming a closed-loop performance self-optimization process.

[0032] Example 2: Figure 2 As shown, the implementation flow of the above system's automatic code generation method is as follows:

[0033] This embodiment provides a method for automatically generating executable code for structured query templates. Its core objective is to automatically construct highly executable, semantically consistent, and high-performance database query function templates through query structure analysis, execution plan formulation, code template generation, and multi-round verification and optimization. The specific method includes the following steps: Query template input: The system receives a predefined structured query template as the input basis for the process. This template is a set of structured, parameterized, and query-logic-structure-based templates with basic intentions and structural frameworks, where a template represents a query logical structure, containing placeholder variables that describe the query behavior, such as filter conditions, aggregation functions, projection columns, grouping keys, sorting or limiting parameters. Query logic decomposition and execution plan generation: The query logic is parsed by the query logic decomposition module to identify semantic units and operation sequences, generating intermediate expression structures. Then, the execution plan generation module formulates a preliminary query execution path based on the logical structure to guide the subsequent code generation process.

[0034] Database structure awareness and schema collection: The system connects the database instance through the database schema collection module to extract meta-information such as current data table structure, field type, index distribution, etc., providing context reference for code generation to ensure compatibility between code and actual database structure.

[0035] Code generation process starts: In the database execution code automatic generation system, first, the create dynamic workflow graph submodule constructs the flowchart of the query task, identifying the functions and dependencies of each execution node; then, the execution code generation submodule generates corresponding code segments based on the flowchart nodes and execution plan; finally, the execution code combination submodule combines each code segment into a complete query function structure.

[0036] Semantic consistency and security verification: The generated query function is submitted to the semantic consistency and security verification module, which compares it with the original query template's intention and compares the actual execution result with the database native query result to confirm semantic accuracy. At the same time, risk detection such as SQL injection, field boundary crossing, and illegal access is performed to ensure the safety and reliability of the generated function.

[0037] Performance evaluation and optimization feedback: If the semantic verification passes, the system enters the performance feedback and self-optimization module to evaluate the execution efficiency of the query function, including response time, resource usage, and execution plan complexity. If the semantic verification fails, i.e., the performance does not meet the threshold, the system will return to the code generation process, automatically optimizing the execution strategy or reconstructing the function structure until the preset performance threshold is met.

[0038] Output query-specific execution code template: The code function after semantic verification and performance optimization is finally output as a query-specific execution code template for subsequent specific query instantiation and use. This template has the characteristics of structural specification, semantic accuracy, and high performance, and can directly serve the database query automatic execution task. Figure 3 and Figure 4As shown, the specific implementation process of the code generation and combination based on the query execution plan and the specific implementation process of the code optimization based on the performance feedback of the large language model provided by the embodiment, specifically through the task agent module to decompose the execution plan, generate the corresponding code fragments, and automatically complete the splicing and integration of the code, finally embed the combined code into the project source code to realize the query function, and through the execution monitoring, error analysis and performance evaluation of the combined code embedded in the project source code, identify the problem points in the query code, and automatically repair or optimize the code logic with the help of the large language model, realize the efficient operation and continuous evolution of the query function. Specifically, it includes the following steps: Step 1: Input query execution plan Receive the query execution plan output by the upstream module (such as the execution plan generation module), which contains the following logical operation nodes: Scan, Filter, Join, Sort, Agg, Project, Distinct, SetOp, Limit / TopN, Window, and Output, as the basic input of this process.

[0039] Step 2: Dynamic workflow generation Call the dynamic workflow generation module according to the operation type in the query execution plan, automatically dispatch the task agent (Agent) module, including Scan Agent, Filter Agent, Join Agent, Sort Agent, Agg Agent, Project Agent, Distinct Agent, SetOp Agent, Window Agent, and Output Agent. Each Agent is responsible for the workflow modeling of the scan, aggregation and output stages, and submits the structured operation process to the code generation module.

[0040] Step 3: Execution code generation In the execution code generation phase, according to the task definition of each Agent, the corresponding code fragment is generated. Each code fragment usually contains multiple parts, such as: (a) Headers: Import library or structure definition; (b) Execute_code: Main execution logic; (c) Finalize_code: Result tail and cleanup; (d) Support_code: Auxiliary function; (e) Bind_code: variable and parameter binding logic.

[0041] Step 4: Code combination All generated code snippets enter the code combination module. This module uses a template engine to combine multiple snippets into a complete execution code according to the preset template. The combination process includes: Step 4-1: Snippet injection: embedding each functional snippet into the code template; Step 4-2: Template-driven: ensuring that the function structure, order, and nesting relationship meet the compilation standards; Step 4-3: Code integration: generating the final executable complete query function.

[0042] Step 5: Embedding project source code and compiling execution The complete code after combination will be automatically inserted into the target project source code, forming a new query logic function, and completing the compilation through the project build system. After successful compilation, the function can be used as an execution module to respond to query calls and implement automatic response and data processing of structured query requests.

[0043] Step 6: Execute query function code Execute the query-specific execution code generated by the automatic generation process and attempt to obtain the query results.

[0044] Step 7: Error detection and exception identification If the execution fails or the query results are abnormal, error logs, running exception information, and their corresponding error types will be collected. Then the code analysis module identifies the root cause of the error (such as syntax error, field mismatch, null pointer exception, etc.) and extracts representative correction features as feedback basis; Step 8: Large language model optimization instruction generation The error correction features are fed back to the large language model for optimization, i.e., optimization instructions are constructed according to the feedback feature content, and the large language model is called to correct the query function code.

[0045] Step 9: Performance data collection and analysis If the execution is successful, the key performance indicators of the function during runtime will be collected for performance monitoring, including response time, CPU occupancy, memory consumption, etc. Then the performance analysis module evaluates and identifies inefficient parts (such as unreasonable JOIN, redundant sorting, full table scanning, etc.) and extracts corresponding optimization features.

[0046] Step 10: Large language model optimization instruction generation The extracted performance optimization features will be fed back to the large language model optimization module respectively. The module constructs optimization instructions according to the feature content, and calls the large language model to rewrite the code, and automatically generates the modified query execution code, aiming to improve its semantic correctness and running efficiency.

[0047] Step 1: output the optimized code and replace the old version The optimized query-specific execution code is output as a new version, replacing the original inefficient or erroneous function. Steps 4-1 and 4-2 are repeated, and it is resubmitted to the execution process to verify its result correctness and performance improvement effect. If it still does not meet the standard, it will continue to enter the next round of optimization iteration until it reaches the expected standard.

[0048] In summary, the present application uses a large language model to perform semantic analysis and function encapsulation on database query statements, enabling automatic generation of database execution code, semantic consistency verification, and performance self-optimization, and is suitable for intelligent programming, database automation management, and AI-assisted software development, etc. technical scenarios.

[0049] The embodiment of the present application also provides a computer readable storage medium having a program stored thereon, which is executed by a processor to implement the database execution code automatic generation method based on a large language model in the above embodiment.

[0050] The computer readable storage medium can be an internal storage unit of any data processing capable device, such as a hard disk or memory, as described in any of the preceding embodiments. The computer readable storage medium can also be any data processing capable device, such as a plug-in hard disk, smart media card (SMC), SD card, flash card, etc. equipped on the device. Further, the computer readable storage medium can include both the internal storage unit of any data processing capable device and the external storage device. The computer readable storage medium is used to store the computer program and other programs and data required by the data processing capable device, and can also be used to temporarily store data that has been output or will be output.

[0051] The above is only a preferred embodiment of the present application and does not limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

[0052] The above examples are only used for illustrating the design idea and characteristics of the present application, and the purpose is to enable the person skilled in the art to understand the present application and to implement it, and the protection scope of the present application is not limited to the above examples. Therefore, any equivalent changes or modifications made according to the disclosed principles and design ideas of the present application are within the protection scope of the present application.

Claims

1. A method for automatically generating database execution code based on a large language model, characterized in that, Includes the following steps: Step 1: Based on the predefined query templates in the user's workload, decompose the query logic of the query templates and generate... The execution plan includes collecting database schema information, data distribution information, and index feature information corresponding to the query template. Step 2: Using a large language model combined with the query template and collected information, adaptively generate database execution code that matches the query semantics; Step 3: For errors that occur during the execution of the generated query-specific code, including compilation errors, execution failures, and execution result errors, perform error analysis and feedback, input the feedback optimization instructions into the large language model for adjustment or regeneration of the code, and output the corrected code; Step 4: Monitor the performance of the corrected code. When the code's performance is still lower than that of the database's native execution engine, input the performance metrics and performance bottleneck characteristics as feedback into the large language model. The model will then generate a performance optimization plan and regenerate the optimized execution code to achieve performance feedback and self-optimization.

2. The method for automatically generating database execution code based on a large language model according to claim 1, characterized in that, Step one specifically includes the following sub-steps: (1.1) Query template input: Receive a predefined structured query template as the input basis; the template is a set of structured, parameterized query templates with basic query intent and structural framework, one of which represents a query logic structure and contains placeholder variables; the placeholder variables include filter conditions, aggregate functions, projection columns, grouping keys, sorting or restriction parameters and instantiable parameters used to describe query behavior; (1.2) Query logic decomposition and execution plan generation: parse the query logic and identify semantic units and operation sequences; Generate intermediate representation structures and formulate preliminary query execution paths based on the logical structure to guide the subsequent code generation process; (1.3) Database structure awareness and schema collection: The database schema collection module connects to the database instance and extracts metadata, including the current data table structure, field types, and index distribution; it provides contextual reference for code generation and ensures that the code is compatible with the actual database structure.

3. The method for automatically generating database execution code based on a large language model according to claim 1, characterized in that, Step two is achieved through the following sub-steps: (2.1) Input query execution plan: Receive query execution plan, which includes the following logical operations: node scanning, filtering, joining, sorting, aggregation, projection, deduplication, set operations, constraints, window functions and output; (2.2) Dynamic workflow generation: Based on the operation type in the query execution plan, the task agent for each logical operation is automatically dispatched. Each agent is responsible for the workflow modeling of the scanning, filtering, joining, sorting, aggregation, projection, deduplication, set operation, window function and output stages, and submits the structured operation process to the next step for code generation. (2.3) Code snippet generation: Generate corresponding code snippets based on the task definition of each Agent. During the generation process, necessary logical processing statements are automatically added according to the query logic to ensure semantic integrity. (2.4) Execution code assembly: The template engine concatenates multiple generated code fragments into a complete execution function according to a preset structure order. The assembly process includes fragment injection, template-driven and structure integration to ensure that the generated code conforms to the compilation specifications. (2.5) Embedding and execution: The complete executable code is automatically embedded into the project source code and compiled through the project build system. After successful compilation, the structured query request is automatically executed and the results are output.

4. The method for automatically generating database execution code based on a large language model according to claim 3, characterized in that, Each code snippet in step (2.4) typically contains the following five parts: (a) Headers: Imported library or structure definition; (b) Execute_code: Main execution logic; (c) Finalize_code: Finalize and clean up the results; (d) Support_code: Auxiliary function; (e) Bind_code: Logic for binding variables and parameters.

5. The method for automatically generating database execution code based on a large language model according to claim 3, characterized in that, The assembly process in step (2.4) includes fragment injection, template-driven assembly, and structure integration, specifically including: The fragment injection embeds each functional fragment into a code template; The template-driven approach ensures that the function structure, order, and nesting relationships conform to compilation standards. The structure is integrated to generate a final, executable, complete query function.

6. The method for automatically generating database execution code based on a large language model according to claim 1, characterized in that, Step three is achieved through the following sub-steps: (3.1) Execute query function code: Execute the query-specific execution code generated by the automatic generation process and obtain the query results; (3.2) Error detection and anomaly identification: When execution fails or results are abnormal, collect error logs, exception stacks and corresponding error types, identify the root cause of the error, including syntax errors, field mismatches or null pointer exceptions, and extract correction features as feedback input; (3.3) Generation of large language model optimization instructions: The error correction features are fed back to the large language model for optimization, that is, optimization instructions are constructed based on the feedback feature content, and the large language model is called to correct the query function code.

7. The method for automatically generating database execution code based on a large language model according to claim 1, characterized in that, Step four is achieved through the following sub-steps: (4.1) Performance data collection and analysis: When the execution is successful, collect the performance indicators of the code running, including response time, CPU utilization and memory consumption, and evaluate the efficiency of the execution plan, identify inefficient parts, and extract corresponding optimization features; The inefficient parts include redundant sorting, full table scans, or suboptimal JOIN strategies. (4.2) Generation of large language model optimization instructions: The error correction features and performance optimization features are fed back to the large language model for optimization, that is, optimization instructions are constructed according to the feature content, and the large language model is called to rewrite or optimize the query function code; (4.3) Output optimized code and replace the old version: Output the optimized query execution code as the new version, replace the original inefficient or erroneous version, and resubmit the execution process for verification; If the verification results still do not meet the performance standards, repeat the above steps until the preset performance threshold is reached, thereby realizing the code self-learning and continuous optimization process based on performance feedback.

8. A database execution code automatic generation system based on a large language model, characterized in that, This system is used to implement the method for automatically generating database execution code based on a large language model as described in any one of claims 1-7, and the system specifically includes: The SQL query statement input unit allows users to input standard SQL statements based on predefined query templates via text boxes, command lines, or API interfaces. The predefined query template set module is a set of structured, parameterized query templates. Each template represents a query logic structure and contains placeholder variables. The large language model module parses query semantics, maps them to preset functions and outputs parameters, while handling exceptions and optimizing performance. It can automatically generate corrected code or propose optimization suggestions based on historical features. The database execution code automatic generation module is responsible for constructing prompts for inputting into the large language model. It combines the semantic structure of the user's query statement with predefined template information to generate input content for calling the large language model. Through the prompt structure, it guides the large language model to generate executable query function code. The query-specific execution code template module stores code function templates generated by the large language model based on the query type. This template serves as the basis for generating executable query functions, which are executed after being populated with parameters according to the query statement and adapted to the query semantics. The semantic consistency and security verification module is used to compare the generated code with the database engine results to ensure consistency of query intent; the built-in security check mechanism is used to detect whether the code has the risk of SQL injection and privilege escalation. The performance feedback and self-optimization module monitors query execution metrics and compares and evaluates performance differences. If it is inferior to the native engine, the prompt message is refactored to call the large model to generate optimized code, forming a feedback-based closed loop. A feedback-based closed-loop optimization mechanism is then built to continuously improve execution efficiency and resource utilization. The database instance module is used to execute query statements and generate functions, and supports mainstream databases; it also provides standardized execution results for comparison and optimization. The query results output module is responsible for uniformly displaying the system's query results, including the query results from the database's native execution engine and the output results of the generated and executed code functions.

9. An electronic device comprising a memory and a processor, wherein, The memory is coupled to the processor; characterized in that the memory is used to store program data, and the processor is used to execute the program data to implement the database execution code automatic generation method based on a large language model as described in any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the database execution code automatic generation method based on any one of claims 1-7.

Citation Information

Patent Citations

  • Code generation method and system based on sql query statement

    CN117873443A

  • Code automatic generation method and device, equipment and medium

    CN118796180A

  • SQL (Structured Query Language) statement generation method based on large-model multi-stage iteration

    CN120086244A

  • Data quality detection and improvement system for large language model

    CN120315718A

  • SQL (Structured Query Language) statement generation method and system based on large language model

    CN120892445A

Cited By

  • MDSplus data access code self-correction method based on execution feedback

    CN122594063A