Database data query method, device and equipment based on large model and storage medium
Through the database data query method based on the big model, a single table view is generated and a task sequence is executed, which solves the efficiency and accuracy problems of the NL2SQL system in complex calculations, and realizes efficient and accurate data query.
Patent Information
- Application Number
- CN202510482516.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-08
AI Technical Summary
When handling complex calculations such as year-on-year, month-on-month, and compound growth rates, the existing NL2SQL system faces the challenges of data preprocessing complexity, benchmark determination, percentage calculation complexity and data missing processing, resulting in low computing efficiency and insufficient accuracy.
A database data query method based on the big model is used to generate a single table view through preset data table governance rules, a problem parser in the big model is used for intent recognition and parameter extraction, a task executor is used to execute task sequences, and a predesign calculation engine calls functions to ensure the accuracy and efficiency of query statements.
It improves the speed and quality of data query, improves the user experience, and ensures the accuracy and efficiency of complex calculations.
Smart Images

Figure CN120277096A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly relates to a method, device, equipment and storage medium for querying database data based on a large model. Background Art
[0002] Currently, as data-driven decision-making becomes more and more common, the importance of the technology for converting natural language to SQL (Structured Query Language) statements is increasing day by day. Among them, the NL2SQL technology allows non-technical users to obtain information in the database through natural language queries. However, when it comes to complex calculations, traditional NL2SQL (Natural Language to Structured Query Language, a technology for automatically converting natural language into SQL query statements) systems face challenges in logical decomposition, semantic understanding and execution control.
[0003] That is, when dealing with complex calculations such as year-on-year, month-on-month, and compound growth rate, the NL2SQL system needs to solve the following difficulties:
[0004] First, before performing year-on-year and month-on-month calculations, it is necessary to preprocess the data, such as time series filling and null value filling of numerical fields, which not only increases the complexity of the calculation but also affects the accuracy of the final result.
[0005] Second, when performing year-on-year and month-on-month calculations, it is necessary to determine a comparison benchmark. For example, for year-on-year, it is necessary to find the data of the same month in the previous year, and for month-on-month, it is necessary to find the data in the previous period or the previous time period. During the data processing process, complex logical judgment operations and data extraction operations are required.
[0006] Third, complex percentage calculations are involved during the process of year-on-year and month-on-month calculations, and it is necessary to ensure that the calculation formula is correct.
[0007] Fourth, during the calculation process, if data is missing, appropriate processing is required, such as filling in missing values or marking as null, thus affecting the accuracy and integrity of the calculation result.
[0008] Fifth, when dealing with large-scale data sets, it is necessary to optimize SQL statements to improve the efficiency and accuracy of the calculation to ensure that the calculation can be completed within a reasonable time and the result is accurate.
[0009] As can be seen from the above, how to improve the speed and quality of data query during the process of querying database data based on a large model is an urgent problem to be solved at present. Summary of the Invention
[0010] In view of this, the purpose of the present invention is to provide a method, device, equipment and storage medium for querying database data based on a large model, which can improve the efficiency of data querying during the process of querying database data based on a large model. The specific scheme is as follows:
[0011] In the first aspect, the present application provides a method for querying database data based on a large model, including:
[0012] Processing the data corresponding to several related data tables in the database by using a preset data table governance rule to obtain corresponding single-table views;
[0013] Using the problem resolver in the large model and based on the single-table view, performing intent recognition and parameter extraction on the natural language query information to obtain corresponding query types and parameter information respectively, and performing a task generation operation based on the query types and the parameter information to obtain a sequence of tasks to be executed;
[0014] Using the preset task executor in the large model and sequentially performing execution operations on each of the tasks to be executed in the sequence of tasks to be executed based on the task execution order to obtain corresponding conversation contents respectively;
[0015] Using a pre-designed computing engine and based on each of the conversation contents, determining the function to be called and the function parameters corresponding to the function to be called, then calling the function to be called and determining a database query statement based on the function parameters, so as to query the database data by using the database query statement.
[0016] Optionally, the processing the data corresponding to several related data tables in the database by using a preset data table governance rule to obtain corresponding single-table views includes:
[0017] Performing a screening operation on each of the data tables in the database based on business requirements to obtain screened data tables, and determining related data tables related to the screened data tables based on each of the screened data tables;
[0018] Performing a field information extraction operation from each of the screened data tables and the corresponding related data tables based on a preset field value rule to obtain corresponding field information, and creating a single-table view based on each of the field information.
[0019] Optionally, the using the problem resolver in the large model and based on the single-table view, performing intent recognition and parameter extraction on the natural language query information to obtain corresponding query types and parameter information respectively, and performing a task generation operation based on the query types and the parameter information to obtain a sequence of tasks to be executed includes:
[0020] Use the problem parser in the large model to determine the query type corresponding to the natural language query information, and obtain the query type; the query type includes an index value query type and an index value comparison type;
[0021] Use the problem parser and extract parameter information from the natural language query information according to the preset parameter information extraction rules and the single-table view to obtain parameter information; the parameter information includes a query time range and a query index name;
[0022] Use the preset task sequence generation rules and perform a task sequence generation operation based on the query type and the parameter information corresponding to the natural language query information to obtain a to-be-executed task sequence; the to-be-executed task sequence is a structured task sequence.
[0023] Optionally, use the preset task executor in the large model, and sequentially perform execution operations on each to-be-executed task in the to-be-executed task sequence based on the task execution order to obtain the corresponding dialogue content, including:
[0024] Use the preset task executor in the large model, and sequentially perform data selection operations, data filtering operations, data joining operations, data grouping operations, and data sorting operations on each to-be-executed task in the to-be-executed task sequence based on the task execution order to obtain the dialogue content corresponding to each to-be-executed task.
[0025] Optionally, use the preset task executor in the large model, and sequentially perform execution operations on each to-be-executed task in the to-be-executed task sequence based on the task execution order, including:
[0026] Use the preset task executor in the large model, and perform an execution operation on the current to-be-executed task in the to-be-executed task sequence to obtain the current execution result, and update the current number of executed tasks;
[0027] Judge whether the current number of executed tasks is less than the number of to-be-executed tasks in the to-be-executed task sequence. If the current number of executed tasks is less than the number of tasks, determine a new current to-be-executed task according to the preset task execution order, then use the current to-be-executed task and determine a new current execution result based on the current execution result, and trigger the step of judging whether the current number of executed tasks is less than the number of to-be-executed tasks in the to-be-executed task sequence until the current number of executed tasks is not less than the number of tasks.
[0028] Optionally, use the preset task executor in the large model, and sequentially perform execution operations on each to-be-executed task in the to-be-executed task sequence based on the task execution order, including:
[0029] Collect all user feedback information according to a preset sampling algorithm to obtain target user feedback information; the target user feedback information includes structured data and unstructured data;
[0030] Use a preset reinforcement learning algorithm and based on the target user feedback information to determine the execution logic corresponding to the task execution rules; the execution logic is used to execute each of the to-be-executed tasks in the to-be-executed task sequence;
[0031] Use a preset test tool to verify the execution logic, and after the verification result indicates that the verification is passed, use the preset task executor in the large model, and based on the task execution order and the task execution rules corresponding to the execution logic, perform execution operations on each of the to-be-executed tasks in the to-be-executed task sequence.
[0032] Optionally, the using a preset computing engine and based on each of the conversation contents to determine a to-be-called function and function parameters corresponding to the to-be-called function, and then calling the to-be-called function and based on the function parameters to determine a database query statement, so as to query database data using the database query statement, includes:
[0033] Use a preset natural language processing technology to parse user requirements to obtain a parsing result, extract key information from the parsing result, and then match a corresponding to-be-called function from a preset built-in function library based on the key information; the user requirements include year-on-year growth rate, month-on-month growth rate, and compound growth rate;
[0034] Based on each of the to-be-called functions, determine respectively corresponding function parameters, call each of the to-be-called functions and based on the corresponding function parameters to determine respectively corresponding query statement fragments, and based on each of the query statement fragments to determine a database query statement, so as to query data in the database using the database query statement.
[0035] In a second aspect, the present application provides a database data query device based on a large model, including:
[0036] A single-table view determination module, configured to use preset data table governance rules to process data corresponding to several related data tables in the database to obtain corresponding single-table views;
[0037] A task sequence determination module, configured to use the problem resolver in the large model and based on the single-table view to perform intention recognition and parameter extraction on natural language query information, obtain respectively corresponding query types and parameter information, and perform task generation operations based on the query types and the parameter information to obtain a to-be-executed task sequence;
[0038] A dialogue content determination module, configured to utilize a preset task executor in the large model and sequentially perform execution operations on each of the to-be-executed tasks in the to-be-executed task sequence based on the task execution order, so as to obtain corresponding dialogue contents respectively;
[0039] A data query module, configured to utilize a pre-designed computing engine and determine a to-be-called function and function parameters corresponding to the to-be-called function based on each of the dialogue contents, then call the to-be-called function and determine a database query statement based on the function parameters, so as to query database data by using the database query statement.
[0040] Thirdly, the present application provides an electronic device, including:
[0041] A memory, configured to store a computer program;
[0042] A processor, configured to execute the computer program to implement the foregoing method for querying database data based on a large model.
[0043] Fourthly, the present application provides a computer-readable storage medium, configured to store a computer program, wherein the computer program, when executed by a processor, implements the foregoing method for querying database data based on a large model.
[0044] As can be seen from the above, before querying database data based on a large model in the present application, it is necessary to process the data corresponding to several related data tables in the database by using preset data table governance rules to obtain a single-table view; utilize a problem parser in the large model and perform intent recognition and parameter extraction on natural language query information based on the single-table view to obtain corresponding query types and parameter information respectively, and perform a task generation operation based on the query types and parameter information to obtain a to-be-executed task sequence; utilize a preset task executor in the large model and sequentially perform execution operations on each of the to-be-executed tasks in the to-be-executed task sequence based on the task execution order to obtain corresponding dialogue contents respectively; utilize a pre-designed computing engine and determine a to-be-called function and function parameters corresponding to the to-be-called function based on each of the dialogue contents, then call the to-be-called function and determine a database query statement based on the function parameters, so as to query database data by using the database query statement.
[0045] It can be seen that in this application, first, it is necessary to process the data corresponding to several related data tables in the database using the preset data table governance rules to obtain a single-table view. Subsequently, the problem resolver in the large model is used to identify the intent and extract parameters from the natural language query information based on the single-table view, obtaining the corresponding query type and parameter information respectively, and performing a task generation operation based on the query type and parameter information to obtain a sequence of tasks to be executed. Furthermore, the preset task executor in the large model is used to execute each task to be executed in the sequence of tasks to be executed in turn based on the task execution order, obtaining the corresponding conversation content respectively. Finally, the preset calculation engine is used to determine the function to be called and the function parameters corresponding to the function to be called based on each conversation content, and then the function to be called is called and the database query statement is determined based on the function parameters to query the database data using the database query statement. In this way, the speed of language conversion is improved, thereby improving the speed and quality of data query, and further enhancing the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other accompanying drawings can be obtained according to the provided drawings without creative efforts.
[0047] Figure 1 It is a flowchart of a method for querying database data based on a large model disclosed in this application;
[0048] Figure 2 It is a schematic diagram of a specific method for parsing natural language using a problem resolver disclosed in this application;
[0049] Figure 3 It is a flowchart of a specific method for querying database data based on a large model disclosed in this application;
[0050] Figure 4 It is a schematic diagram of the structure of a device for querying database data based on a large model disclosed in this application;
[0051] Figure 5 It is a structural diagram of an electronic device disclosed in this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0052] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0053] Currently, as data-driven decision-making becomes increasingly common, the importance of natural language to SQL technology is increasing day by day. Among them, NL2SQL technology allows non-technical users to obtain information in the database through natural language queries. However, when it comes to complex calculations, when dealing with complex calculations such as year-on-year, month-on-month, and compound growth rate, the NL2SQL system needs to solve the following difficulties: First, before performing year-on-year and month-on-month calculations, the data needs to be preprocessed. Second, a comparison benchmark needs to be determined when performing year-on-year and month-on-month calculations. Third, complex percentage calculations are involved in the process of year-on-year and month-on-month calculations, and the calculation formula needs to be ensured to be correct. Fourth, if the situation of data missing is encountered, appropriate processing needs to be carried out. Fifth, when dealing with large-scale data sets, the SQL statement needs to be optimized to improve the efficiency and accuracy of the calculation. For this reason, this application provides a database data query method based on a large model, which can improve the speed and quality of data query.
[0054] See Figure 1 As shown, the embodiments of the present invention disclose a database data query method based on a large model, including:
[0055] Step S11, processing the data corresponding to several related data tables in the database by using preset data table governance rules to obtain corresponding single-table views.
[0056] In this embodiment, since the database contains multiple tables, and there are complex association relationships among the above tables. Therefore, in order to perform data queries more intuitively and efficiently, the embodiment of the present application creates the data required for the query as a view based on business requirements, and constructs the created view into a single-table view. It is worth mentioning that the construction process of the single-table view involves operations such as screening, joining, and field selection and calculation on the relevant tables in the database, so that the data originally scattered in different tables is integrated into one view, so that there is no need to care about the complex structure of the underlying database when performing queries. In addition, the single-table view can be adjusted and extended according to changes in business requirements, thus ensuring the flexibility and adaptability of data queries. Specifically, using the preset data table governance rules to process the data corresponding to several related data tables in the database to obtain the corresponding single-table view may include: performing a screening operation on each data table in the database based on business requirements to obtain the screened data tables, and determining the relevant data tables related to the screened data tables based on each screened data table; performing a field information extraction operation from each screened data table and the corresponding relevant data tables based on the preset field value rules to obtain the corresponding field information, and creating a single-table view based on each field information.
[0057] Step S12: Use the problem parser in the large model and based on the single-table view to perform intent recognition and parameter extraction on the natural language query information, obtain the corresponding query type and parameter information respectively, and perform a task generation operation based on the query type and the parameter information to obtain a task sequence to be executed.
[0058] In this embodiment, the main function of the problem parser is to receive the natural language question of the user, and output a structured task sequence after a series of complex processing processes on the natural language question. In a specific implementation manner, the file format of the task sequence is JSON (JavaScript Object Notation, that is, a lightweight data exchange format), and the task sequence includes user query intent recognition, key parameter extraction, and tasks required to generate the key parameters.
[0059] It is worth mentioning that intent recognition is the primary task of the question parser, which needs to accurately judge the type of query the user wants to perform, such as: obtaining the value of a certain indicator, comparing two indicators. Parameter extraction is to extract key information from the query information input by the user, such as the time range and specific indicator names. Task generation is to generate tasks based on the recognized intent and the extracted parameters, and the above tasks are used to generate SQL statements and execute query statements. Among them, the question parser built based on the LLM (Large Language Model) can understand complex natural language expressions and convert them into a sequence of tasks executable by machines. And in a specific implementation, the schematic diagram of using the question parser to parse natural language is as Figure 2 shown.
[0060] Specifically, use the question parser in the large model and based on the single-table view to perform intent recognition and parameter extraction on the natural language query information, obtain the corresponding query type and parameter information respectively, and perform task generation operations based on the query type and parameter information to obtain the task sequence to be executed, which may include: use the question parser in the large model to judge the query type corresponding to the natural language query information to obtain the query type; the query type includes the indicator value query type and the indicator value comparison type; use the question parser and extract parameter information from the natural language query information according to the preset parameter information extraction rules and the single-table view to obtain parameter information; the parameter information includes the query time range and the query indicator name; use the preset task sequence generation rules and perform task sequence generation operations based on the query type and parameter information corresponding to the natural language query information to obtain the task sequence to be executed; the task sequence to be executed is a structured task sequence.
[0061] Step S13, use the preset task executor in the large model, and sequentially execute each of the tasks to be executed in the task sequence to be executed based on the task execution order to obtain the corresponding conversation content respectively.
[0062] In this embodiment, a large language model is used to generate SQL query statements based on the parsing results of the schema. Among them, the large language model needs to parse the parsing results according to the problem parser and output a structured task sequence. It is worth mentioning that when generating SQL statements, the LLM needs to consider all aspects of the query information, including but not limited to data selection, filtering, joining, grouping, sorting, and limiting the number of returned records. It is worth mentioning that not all queries require the above operations. For example, some queries only need to retrieve data from the database without complex grouping or sorting. In this case, the LLM will remove those unnecessary parts of the SQL statement to ensure that the generated SQL statement is as concise and effective as possible. Specifically, the preset task executor in the large model is used, and each to-be-executed task in the to-be-executed task sequence is executed based on the task execution order, and the corresponding conversation content can be obtained, including: using the preset task executor in the large model, and performing data selection operations, data filtering operations, data joining operations, data grouping operations, and data sorting operations on each to-be-executed task in the to-be-executed task sequence based on the task execution order, to obtain the conversation content corresponding to each to-be-executed task respectively.
[0063] In this embodiment, the embodiment of the present application adopts a task executor to sequentially execute each to-be-executed task, and constructs a multi-round conversation based on the problem and the query results to ensure that the output of each task can be correctly passed to the next task, and the query strategy can be adjusted according to the user's feedback when needed. Specifically, using the preset task executor in the large model, and performing execution operations on each to-be-executed task in the to-be-executed task sequence based on the task execution order may include: using the preset task executor in the large model, and performing an execution operation on the current to-be-executed task in the to-be-executed task sequence to obtain the current execution result and update the current number of executed tasks; judging whether the current number of executed tasks is less than the number of to-be-executed tasks in the to-be-executed task sequence, if the current number of executed tasks is less than the number of tasks, then determine a new current to-be-executed task according to the preset task execution order, and then use the current to-be-executed task and based on the current execution result to determine a new current execution result, and trigger the step of judging whether the current number of executed tasks is less than the number of to-be-executed tasks in the to-be-executed task sequence, until the current number of executed tasks is not less than the number of tasks.
[0064] It is worth mentioning that the task executor is responsible for executing each task to be executed in sequence and constructing a multi-round dialogue based on the question and query results to ensure that the output of each task can be correctly passed to the next task and that the query strategy can be adjusted according to the user's feedback when necessary. Specifically, when using the task executor to execute a query task, the executor will execute each task to be executed in the order of execution and according to the structured task sequence generated by the question parser. In a specific implementation, the tasks to be executed include SQL query tasks, data processing tasks, or complex calculation tasks. In addition, the task executor needs to ensure that the output of each task is correct and can be correctly used by the next task. This requires the executor to have certain error handling capabilities and exception management capabilities to ensure that problems can be processed and adjusted in a timely manner when encountered. In addition, the task executor can also adjust the query strategy according to the user's feedback information. During the multi-round dialogue, the user will raise questions about the query results or request further detailed information, and the information provided by the user will be fed back to the task executor. Subsequently, the task executor needs to be able to understand the user's feedback information and perform an adjustment operation on the query strategy based on the feedback information to provide more accurate information that meets the user's needs.
[0065] Specifically, by using the preset task executor in the large model and sequentially performing execution operations on each task to be executed in the task sequence to be executed, it may include: collecting all user feedback information according to the preset sampling algorithm to obtain target user feedback information; the target user feedback information includes structured data and unstructured data; using the preset reinforcement learning algorithm and determining the execution logic corresponding to the task execution rules based on the target user feedback information; the execution logic is used to execute each task to be executed in the task sequence to be executed; using the preset test tool to verify the execution logic, and after the verification result indicates that the verification is passed, using the preset task executor in the large model and performing execution operations on each task to be executed in the task sequence to be executed based on the task execution order and the task execution rules corresponding to the execution logic.
[0066] In a specific implementation, the code for using the task executor to execute tasks is as follows:
[0067] class TaskExecutor:
[0068] def run(self, task_sequence):
[0069] convs = []
[0070] context = {}
[0071] for step in task_sequence["steps"]:
[0072] convs.append({"role":"user","content": step["query"]})
[0073] if step["type"] == "sql_query":
[0074] # Execute the SQL query and save the result
[0075] result = execute_sql(generate_sql(step["schema"]))
[0076] convs.append({"role":"assistant","content": result})
[0077] return convs
[0078] Output
[0079] # Determine the function (LLM) to be called based on the multi-round conversation content after the task execution
[0080] convs = [{"role":"user","content": "Query the net profit in April"},{"role":"assistant","content": "89"},
[0081] {"role":"user","content": "Query the net profit in May"},{"role":"assistant","content": "101"},
[0082] {"role":"user","content": "Calculate how much the net profit in May has increased compared to April"}]
[0083] Step S14: Use a pre-designed calculation engine and based on each of the conversation contents, determine the function to be called and the function parameters corresponding to the function to be called, then call the function to be called and determine a database query statement based on the function parameters, so as to query the database data using the database query statement.
[0084] In this embodiment, the calculation engine is used to complete the determination operation, parameter parsing operation, and execution operation of function calls on the basis of multi-round conversations, so as to convert complex calculation requirements into specific function calls, and execute these functions to obtain the final result, and the flow chart is as Figure 3As shown. That is, during the multi-round conversation process, users will put forward various complex calculation requirements, such as year-on-year calculation, month-on-month calculation, and compound growth rate calculation. The calculation engine needs to determine the function to be called according to the conversation content and parse the parameters required to execute the above-mentioned called function. During the process of parsing the parameters required to execute the above-mentioned called function, the embodiments of the present application need to deeply understand and analyze the conversation content and accurately grasp the calculation requirements. In this way, the calculation engine needs to have strong semantic understanding ability and calculation ability to ensure that it can correctly parse the user's intention and execute the corresponding calculation task. Subsequently, when the calculation engine determines the function and parameters to be called, the calculation engine will execute the corresponding function based on the corresponding parameters and obtain the final result according to the output of the function. Among them, during the process of function execution, continuous calls of multiple functions and parameter passing are involved. The calculation engine needs to ensure that the execution of each function is correct and can accurately pass the result to the next function. Finally, the calculation engine will present the calculation result to the user to meet the user's query needs.
[0085] Specifically, using a pre-designed calculation engine and based on each conversation content to determine the function to be called and the function parameters corresponding to the function to be called, and then calling the function to be called and determining the database query statement based on the function parameters to query the database data using the database query statement may include: using a preset natural language processing technology to parse the user's requirements, obtaining the parsing result, and extracting key information from the parsing result, and then matching the corresponding function to be called from a preset built-in function library based on the key information; the user's requirements include year-on-year growth rate, month-on-month growth rate, and compound growth rate; determining the function parameters corresponding to each function to be called respectively, calling each function to be called and determining the corresponding query statement fragments respectively based on the corresponding function parameters, and determining the database query statement based on each query statement fragment to query the data in the database using the database query statement.
[0086] In a specific implementation manner, several complex calculation functions are defined using a calculation engine, such as calculate_mom (Month-over-Month Calculation, that is, the month-on-month calculation function), calculate_yoy (Year-over-Year Calculation, that is, the year-on-year calculation function), calculate_sum (Sum Calculation, that is, the summation calculation function), calculate_difference (Difference Calculation, that is, the difference calculation function), and according to the multi-round conversation content after task execution, determine the function to be called and perform the calculation, so as to obtain the final result, and the code for calling the function to perform the calculation is as shown below:
[0087] # Calculate month-on-month growth rate
[0088] def calculate_mom(current_value: float, previous_value: float) ->dict:
[0089] if previous_value == 0:
[0090] raise ValueError("Previous value cannot be zero for MoM calculation.")
[0091] growth_rate = (current_value - previous_value) / previous_value
[0092] return {
[0093] "growth_rate": growth_rate,
[0094] "result": f"{growth_rate * 100:.2f}%"
[0095] }
[0096] # Calculate year-on-year growth rate
[0097] def calculate_yoy(current_value: float, previous_value: float) ->dict:
[0098] if previous_value == 0:
[0099] raise ValueError("Previous value cannot be zero for YoY calculation.")
[0100] growth_rate = (current_value - previous_value) / previous_value
[0101] return {
[0102] "growth_rate": growth_rate,
[0103] "result": f"{growth_rate * 100:.2f}%"
[0104] }
[0105] # Calculate the sum of a list of values
[0106] def calculate_sum(values: list) -> dict:
[0107] return {
[0108] "sum": sum(values)
[0109] }
[0110] # Calculate the difference between two values
[0111] def calculate_difference(value1: float, value2: float) -> dict:
[0112] return {
[0113] "difference": value1 - value2
[0114] }
[0115] function_call_request = {"name": "calculate_difference", "arguments":{"value1": 101, "value2": 89}}
[0116] # Parse Function Call
[0117] function_name = function_call_request["name"]
[0118] function_args = json.loads(function_call_request["arguments"])
[0119] # Call the corresponding function according to the function name
[0120] if function_name == "calculate_mom":
[0121] result = calculate_mom(**function_args)
[0122] elif function_name == "calculate_yoy":
[0123] result = calculate_yoy(**function_args)
[0124] elif function_name == "calculate_sum":
[0125] result = calculate_sum(**function_args)
[0126] elif function_name == "calculate_difference":
[0127] result = calculate_difference(**function_args)
[0128] elif function_name == "calculate_growth_rate":
[0129] result = calculate_growth_rate(**function_args)
[0130] else:
[0131] raise ValueError(f"Unknown function: {function_name}")
[0132] # Return the result
[0133] print(result)
[0134] It can be seen that, in the embodiments of the present application, first, it is necessary to process the data corresponding to several related data tables in the database by using the preset data table governance rules to obtain a single-table view; subsequently, use the problem parser in the large model and based on the single-table view to perform intent recognition and parameter extraction on the natural language query information, obtain the corresponding query type and parameter information respectively, and perform a task generation operation based on the query type and parameter information to obtain a task sequence to be executed; furthermore, use the preset task executor in the large model and sequentially perform execution operations on each task to be executed in the task sequence to be executed based on the task execution order to obtain the corresponding conversation content respectively; finally, use the preset calculation engine and based on each conversation content to determine the function to be called and the function parameters corresponding to the function to be called, then call the function to be called and determine the database query statement based on the function parameters, so as to query the database data by using the database query statement. In this way, the speed of language conversion is improved, thereby improving the efficiency of data query.
[0135] Correspondingly, as shown in Figure 4 the present application also provides a database data query device based on a large model, including:
[0136] A single-table view determination module 11, configured to process the data corresponding to several related data tables in the database by using preset data table governance rules to obtain a corresponding single-table view;
[0137] A task sequence determination module 12, configured to use the problem parser in the large model and based on the single-table view to perform intent recognition and parameter extraction on the natural language query information, obtain the corresponding query type and parameter information respectively, and perform a task generation operation based on the query type and the parameter information to obtain a task sequence to be executed;
[0138] A conversation content determination module 13, configured to use the preset task executor in the large model and sequentially perform execution operations on each task to be executed in the task sequence to be executed based on the task execution order to obtain the corresponding conversation content respectively;
[0139] A data query module 14, configured to use a preset calculation engine and based on each conversation content to determine the function to be called and the function parameters corresponding to the function to be called, then call the function to be called and determine the database query statement based on the function parameters, so as to query the database data by using the database query statement.
[0140] As can be seen from the above, before querying the database data based on the large model in the embodiments of the present application, it is first necessary to process the data corresponding to several related data tables in the database using the preset data table governance rules to obtain a single-table view; subsequently, use the problem parser in the large model and based on the single-table view to perform intent recognition and parameter extraction on the natural language query information, obtain the corresponding query type and parameter information respectively, and perform a task generation operation based on the query type and parameter information to obtain a task sequence to be executed; furthermore, use the preset task executor in the large model and sequentially execute each task to be executed in the task sequence to be executed based on the task execution order to obtain the corresponding conversation content respectively; finally, use the preset calculation engine and based on each conversation content to determine the function to be called and the function parameters corresponding to the function to be called, then call the function to be called and determine the database query statement based on the function parameters to query the database data using the database query statement. In this way, the speed of language conversion is improved, thereby improving the speed and quality of data query.
[0141] In some specific embodiments, the single-table view determination module 11 may specifically include:
[0142] A data table determination unit, configured to perform a screening operation on each of the data tables in the database based on business requirements to obtain the screened data tables, and determine the data-related data tables associated with the screened data tables based on each of the screened data tables;
[0143] A field information determination unit, configured to perform a field information extraction operation from each of the screened data tables and the corresponding data-related data tables according to the preset field value-taking rules to obtain the corresponding field information, and create a single-table view based on each of the field information.
[0144] In some specific embodiments, the task sequence determination module 12 may specifically include:
[0145] A query type determination unit, configured to use the problem parser in the large model to determine the query type corresponding to the natural language query information to obtain the query type; the query type includes an index value query type and an index value comparison type;
[0146] A parameter information determination unit, configured to use the problem parser and perform a parameter information extraction operation on the natural language query information according to the preset parameter information extraction rules and the single-table view to obtain the parameter information; the parameter information includes a query time range and a query index name;
[0147] A task sequence determination subunit, configured to generate a task sequence by using a preset task sequence generation rule and based on the query type and the parameter information corresponding to the natural language query information, so as to obtain a to-be-executed task sequence; the to-be-executed task sequence is a structured task sequence.
[0148] In some specific embodiments, the dialogue content determination module 13 may specifically include:
[0149] A dialogue content determination subunit, configured to use a preset task executor in the large model and sequentially perform data selection operations, data filtering operations, data joining operations, data grouping operations, and data sorting operations on each of the to-be-executed tasks in the to-be-executed task sequence based on the task execution order, so as to obtain dialogue content corresponding to each of the to-be-executed tasks.
[0150] In some specific embodiments, the dialogue content determination module 13 may specifically include:
[0151] A task quantity update unit, configured to use a preset task executor in the large model and perform an execution operation on the current to-be-executed task in the to-be-executed task sequence, so as to obtain a current execution result and update the current number of executed tasks;
[0152] A task quantity judgment unit, configured to judge whether the current number of executed tasks is less than the number of to-be-executed tasks in the to-be-executed task sequence. If the current number of executed tasks is less than the number of tasks, a new current to-be-executed task is determined according to the preset task execution order, and then a new current execution result is determined by using the current to-be-executed task and based on the current execution result, and the step of judging whether the current number of executed tasks is less than the number of to-be-executed tasks in the to-be-executed task sequence is triggered until the current number of executed tasks is not less than the number of tasks.
[0153] In some specific embodiments, the dialogue content determination module 13 may specifically include:
[0154] A feedback information determination unit, configured to collect all user feedback information according to a preset sampling algorithm to obtain target user feedback information; the target user feedback information includes structured data and unstructured data;
[0155] A to-be-executed task determination unit, configured to determine an execution logic corresponding to a task execution rule by using a preset reinforcement learning algorithm and based on the target user feedback information; the execution logic is used to execute each of the to-be-executed tasks in the to-be-executed task sequence;
[0156] A task execution unit is configured to verify the execution logic by using a preset test tool. After the verification result indicates that the verification is passed, the task execution unit uses a preset task executor in the large model and executes each of the to-be-executed tasks in the to-be-executed task sequence based on the task execution rules corresponding to the execution logic according to the task execution order.
[0157] In some specific embodiments, the data query module 14 may specifically include:
[0158] A to-be-called function determination unit is configured to parse user requirements by using a preset natural language processing technique to obtain a parsing result, extract key information from the parsing result, and then match a corresponding to-be-called function from a preset built-in function library based on the key information; the user requirements include year-on-year growth rate, month-on-month growth rate, and compound growth rate.
[0159] A query statement determination unit is configured to determine respectively corresponding function parameters based on each of the to-be-called functions, call each of the to-be-called functions and determine respectively corresponding query statement fragments based on the corresponding function parameters, and determine a database query statement based on each of the query statement fragments to query data in the database by using the database query statement.
[0160] Furthermore, an embodiment of the present application also discloses an electronic device. Figure 5 It is a structural diagram of an electronic device 20 shown according to an exemplary embodiment, and the content in the figure should not be considered as any limitation to the scope of use of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the database data query method based on a large model disclosed in any of the foregoing embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0161] In this embodiment, the power supply 23 is used to provide a working voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows is any communication protocol applicable to the technical solution of the present application, and no specific limitation is imposed thereon here; the input / output interface 25 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application needs, and no specific limitation is made here.
[0162] In addition, the memory 22, as a carrier for storing resources, can be a read-only memory, a random access memory, a magnetic disk, an optical disk, etc. The resources stored thereon can include an operating system 221, a computer program 222, etc. The storage method can be temporary storage or permanent storage.
[0163] Among them, the operating system 221 is used to manage and control each hardware device and the computer program 222 on the electronic device 20, and it can be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program that can be used to complete the database data query method based on the large model executed by the electronic device 20 disclosed in any of the foregoing embodiments, the computer program 222 can further include computer programs that can be used to complete other specific tasks.
[0164] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the database data query method based on the large model disclosed above is implemented. For the specific steps of this method, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be repeated here.
[0165] In this specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the device disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and reference can be made to the method part for related parts.
[0166] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this document can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0167] The steps of the method or algorithm described in combination with the embodiments disclosed in this document can be directly implemented by hardware, a software module executed by a processor, or a combination of the two. The software module can be placed in a random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the technical field.
[0168] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.
[0169] The technical solutions provided in this application have been introduced in detail above. Specific examples are used in this text to elaborate on the principles and implementation manners of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to this application.
Claims
1. A method for querying database data based on a large model, characterized in that, Including: Processing the data corresponding to several related data tables in the database using a preset data table governance rule to obtain corresponding single-table views; Using the problem resolver in the large model and based on the single-table view to perform intent recognition and parameter extraction on natural language query information, obtaining corresponding query types and parameter information respectively, and performing a task generation operation based on the query type and the parameter information to obtain a to-be-executed task sequence; Using the preset task executor in the large model and sequentially performing execution operations on each of the to-be-executed tasks in the to-be-executed task sequence based on the task execution order to obtain corresponding conversation contents respectively; Using a preset calculation engine and based on each of the conversation contents to determine a to-be-called function and function parameters corresponding to the to-be-called function, then calling the to-be-called function and determining a database query statement based on the function parameters to query the database data using the database query statement.
2. The method for querying database data based on a large model according to claim 1, wherein The step of processing the data corresponding to several related data tables in the database using a preset data table governance rule to obtain corresponding single-table views includes: Performing a screening operation on each of the data tables in the database based on business requirements to obtain screened data tables, and determining data-related data tables associated with the screened data tables based on each of the screened data tables; Performing a field information extraction operation from each of the screened data tables and the corresponding data-related data tables based on a preset field value rule to obtain corresponding field information, and creating a single-table view based on each of the field information.
3. The method for querying database data based on a large model according to claim 1, wherein, The step of using the problem resolver in the large model and based on the single-table view to perform intent recognition and parameter extraction on natural language query information, obtaining corresponding query types and parameter information respectively, and performing a task generation operation based on the query type and the parameter information to obtain a to-be-executed task sequence includes: Using the problem resolver in the large model to judge the query type corresponding to the natural language query information to obtain the query type; the query type includes an index value query type and an index value comparison type; Using the problem resolver and extracting parameter information from the natural language query information according to a preset parameter information extraction rule and the single-table view to obtain the parameter information; the parameter information includes a query time range and a query index name; Performing a task sequence generation operation using a preset task sequence generation rule and based on the query type and the parameter information corresponding to the natural language query information to obtain a to-be-executed task sequence; the to-be-executed task sequence is a structured task sequence.
4. The method for querying database data based on a large model according to claim 1, wherein The step of using the preset task executor in the large model and sequentially performing execution operations on each of the to-be-executed tasks in the to-be-executed task sequence based on the task execution order to obtain corresponding conversation contents respectively includes: Utilize the preset task executor in the large model, and sequentially perform data selection operations, data filtering operations, data joining operations, data grouping operations, and data sorting operations on each of the to-be-executed tasks in the to-be-executed task sequence based on the task execution order, to obtain the dialogue content corresponding to each of the to-be-executed tasks.
5. The method for querying database data based on a large model according to claim 1, wherein, The utilization of the preset task executor in the large model and the sequential execution operations on each of the to-be-executed tasks in the to-be-executed task sequence based on the task execution order include: Utilize the preset task executor in the large model, and perform an execution operation on the current to-be-executed task in the to-be-executed task sequence to obtain the current execution result, and update the current number of executed tasks; Judge whether the current number of executed tasks is less than the number of tasks in the to-be-executed task sequence. If the current number of executed tasks is less than the number of tasks, determine a new current to-be-executed task according to the preset task execution order, then use the current to-be-executed task and based on the current execution result determine a new current execution result, and trigger the step of judging whether the current number of executed tasks is less than the number of tasks in the to-be-executed task sequence until the current number of executed tasks is not less than the number of tasks.
6. The method for querying database data based on a large model according to claim 1, wherein The utilization of the preset task executor in the large model and the sequential execution operations on each of the to-be-executed tasks in the to-be-executed task sequence based on the task execution order include: Collect all user feedback information according to the preset sampling algorithm to obtain the target user feedback information; the target user feedback information includes structured data and unstructured data; Utilize the preset reinforcement learning algorithm and based on the target user feedback information determine the execution logic corresponding to the task execution rule; the execution logic is used to execute each of the to-be-executed tasks in the to-be-executed task sequence; Utilize the preset test tool to verify the execution logic, and after the verification result indicates that the verification is passed, utilize the preset task executor in the large model, and based on the task execution order and the task execution rule corresponding to the execution logic perform execution operations on each of the to-be-executed tasks in the to-be-executed task sequence.
7. The method for querying database data based on a large model according to claim 1, wherein The utilization of the preset computing engine and based on each of the dialogue contents determine the to-be-called function and the function parameters corresponding to the to-be-called function, then call the to-be-called function and based on the function parameters determine the database query statement, to query the database data using the database query statement, includes: Utilize the preset natural language processing technology to parse the user requirements to obtain the parsing result, and extract the key information from the parsing result, then match the corresponding to-be-called function from the preset built-in function library based on the key information; the user requirements include year-on-year growth rate, month-on-month growth rate, and compound growth rate; Determine the corresponding function parameters for each of the to-be-called functions respectively, call each of the to-be-called functions, determine the corresponding query statement fragments based on the corresponding function parameters, and determine a database query statement based on each of the query statement fragments, so as to query the data in the database using the database query statement.
8. A database data query device based on a large model, characterized in that, Including: A single-table view determination module, configured to process the data corresponding to several related data tables in the database using a preset data table governance rule to obtain a corresponding single-table view; A task sequence determination module, configured to use the problem resolver in the large model and perform intent recognition and parameter extraction on the natural language query information based on the single-table view to obtain the corresponding query type and parameter information respectively, and perform a task generation operation based on the query type and the parameter information to obtain a to-be-executed task sequence; A conversation content determination module, configured to use the preset task executor in the large model and perform execution operations on each of the to-be-executed tasks in the to-be-executed task sequence in sequence according to the task execution order to obtain the corresponding conversation content respectively; A data query module, configured to use a preset calculation engine and determine the to-be-called functions and the function parameters corresponding to the to-be-called functions based on each of the conversation contents, then call the to-be-called functions and determine a database query statement based on the function parameters, so as to query the database data using the database query statement.
9. An electronic device, characterized in that, Including: A memory, configured to store a computer program; A processor, configured to execute the computer program to implement the large model-based database data query method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, For storing a computer program, wherein the computer program, when executed by the processor, implements the large model-based database data query method according to any one of claims 1 to 7.