Big data analysis method, system and equipment based on large model and storage medium

By parsing user problems into domain-specific languages and converting them into database query commands, the problem of low accuracy in large language models when querying databases is solved, and efficient and secure database queries are achieved.

CN120448411APending Publication Date: 2025-08-08GUANGZHOU YUNCONG INFORMATION TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510595039.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

Large language models have low accuracy when querying specific database scenarios, and the cost of adjusting model parameters is high, which poses data security risks.

Method used

By obtaining user questions, we can determine whether the database needs to be queried. If necessary, use the first prompt information to parse it into a domain-specific language (DSL), and then convert it into a database query command such as SQL to generate analysis results.

Benefits of technology

It improves the accuracy of database queries, reduces customization costs, ensures the privacy and security of user data, and has strong user problem analysis capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120448411A_ABST
    Figure CN120448411A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent questions and answers, particularly provides a big data analysis method, system and device based on a large model and a storage medium, and aims to solve the problem that a large language model is low in accuracy when querying a database scene. The method comprises the steps of obtaining a user question, judging whether the user question needs to query a database or not, responding to the user question needing to query the database, analyzing the user question into a domain-specific language according to first prompt information, determining a database query command according to the domain-specific language, and obtaining a query result from a preset database according to the database query command, and generating an analysis result according to the user problem and the query result. On one hand, the user problem is analyzed into the domain-specific language by using the excellent understanding ability of the large language model to the natural language, and on the other hand, the domain-specific language is converted into the database query command, so that the database query accuracy is ensured, and meanwhile, the method has relatively high data analysis ability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of intelligent question-answering technology, and specifically to a big data analysis method, system, device and storage medium based on a large model. Background Art

[0002] Large language models (LLMs) offer excellent generalization capabilities for semantic understanding, but they also have certain drawbacks. For example, their structure is relatively closed, and feedback is based on the knowledge base from the training phase. Therefore, in certain scenarios, such as querying user-specific database content, accurate question and answer performance may be inadequate. Using a LLM to output Structured Query Language (SQL) for database queries requires high training costs.

[0003] Accordingly, this field needs a new big data analysis solution based on large models to solve the above problems. Summary of the Invention

[0004] In order to overcome the above-mentioned defects, the present application is proposed to solve or at least partially solve the technical problem of low accuracy of large language models when querying database scenarios.

[0005] In a first aspect, a big data analysis method based on a big model is provided, the method comprising:

[0006] In a technical solution of the above-mentioned big data analysis method based on a large model, the method includes: obtaining a user question, determining whether the user question requires querying a database; in response to the user question requiring querying a database, parsing the user question into a domain-specific language according to a first prompt information; determining a database query command according to the domain-specific language; obtaining a query result from a preset database according to the database query command; and generating an analysis result based on the user question and the query result.

[0007] In a technical solution of the above-mentioned big data analysis method based on a large model, the method further includes: generating an analysis result according to the user question and the second prompt information in response to the user question without querying the database.

[0008] In a technical solution of the above-mentioned big data analysis method based on a big model, the first prompt information includes the task requirements, function definitions, parameter descriptions and reference examples of the big model, wherein the parameter descriptions are used to specify the structure of the domain-specific language, and the task requirements include parsing the user questions into domain-specific language based on the function definitions, parameter descriptions and reference examples.

[0009] In a technical solution of the above-mentioned big data analysis method based on a big model, the parameter description specifies that the structure of the domain-specific language includes parameters and parameter values in the form of key-value pairs, wherein the parameters include entities, query dimensions, filtering rules, query indicators, operation operators and grouping rules.

[0010] In a technical solution of the above-mentioned big data analysis method based on a big model, the user question is parsed into a domain-specific language according to the first prompt information, including: extracting the parameter values corresponding to the entities, query dimensions, filtering rules, query indicators, operators and grouping rules from the user question according to the function definition, parameter description and reference examples; and converting the user question into a domain-specific language according to the parameter values corresponding to the entities, query dimensions, filtering rules, query indicators, operators and grouping rules.

[0011] In a technical solution of the above-mentioned big data analysis method based on a large model, the analysis results include reports or text results, and generating the analysis results based on the user questions and the query results includes: splicing the user questions and the query results to obtain a third prompt information; generating the report or text results based on the third prompt information.

[0012] In a technical solution of the above-mentioned big data analysis method based on a large model, the parameters also include the table name and table header name of the database.

[0013] In a second aspect, a big data analysis system based on a big model is provided, the system comprising: an acquisition module for acquiring user questions and determining whether the user questions require a database query; a parsing module, in response to the user question requiring a database query, for parsing the user questions into a domain-specific language according to first prompt information; a determination module for determining a database query command according to the domain-specific language; a query module for obtaining query results from a preset database according to the database query command; and a generation module for generating analysis results based on the user questions and the query results.

[0014] In a third aspect, an intelligent device is provided, comprising at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program, and when the computer program is executed by the at least one processor, the method described in any one of the technical solutions of the above-mentioned big data analysis method based on a large model is implemented.

[0015] In a fourth aspect, a computer-readable storage medium is provided, which stores a plurality of program codes, wherein the program codes are suitable for being loaded and run by a processor to execute the method described in any one of the technical solutions of the above-mentioned big data analysis method based on a large model.

[0016] The above one or more technical solutions of this application have at least one or more of the following beneficial effects:

[0017] In the technical solution for implementing the present application, the user question that needs to query the database is parsed into a domain-specific language according to the first prompt information, and then the database query command is determined according to the domain-specific language. Then, the database query command is used to obtain the query result from the preset database, and finally, the analysis result is generated according to the user question and the query result. On the one hand, the excellent understanding ability of the large language model for natural language is utilized to parse the user question into a domain-specific language. On the other hand, the domain-specific language is converted into a database query command. While ensuring the accuracy of the database query, it also has a strong user question analysis capability, can efficiently realize the analysis and mining of big data, and solves the problem of low accuracy of the large language model in the database query scenario. In addition, this method does not need to put user data into model training, avoids high customization costs, and ensures the privacy and security of user data. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The disclosure of this application will become more easily understood with reference to the accompanying drawings. Those skilled in the art will readily appreciate that these drawings are for illustrative purposes only and are not intended to limit the scope of protection of this application. Among them:

[0019] Figure 1 This is a flowchart of the main steps of a big data analysis method based on a big model according to an embodiment of the present application;

[0020] Figure 2 1 is a schematic diagram of the overall framework of a big data analysis method based on a big model according to an embodiment of the present application;

[0021] Figure 3 This is a schematic diagram of the main structural block diagram of a big data analysis system based on a big model according to an embodiment of the present application;

[0022] Figure 4 1 is a schematic diagram of the overall framework of a big data analysis system based on a large model according to an embodiment of the present application;

[0023] Figure 5 It is a schematic diagram of the main structure of a smart device according to an embodiment of the present application.

[0024] Reference numerals:

[0025] 11: Memory; 12: Processor. DETAILED DESCRIPTION

[0026] Some embodiments of the present application are described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principles of the present application and are not intended to limit the scope of protection of the present application.

[0027] In the description of this application, the terms "first", "second", etc. are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way are interchangeable where appropriate so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices. The terms "installed", "connected", and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, an indirect connection through an intermediate medium, or a communication between two elements, a wireless connection, or a wired connection.

[0028] In addition, "module" and "processor" may include hardware, software, or a combination of the two. A module may include hardware circuits, various suitable sensors, communication ports, and memory, and may also include software components, such as program code, or a combination of software and hardware. The processor may be a central processing unit, a microprocessor, an image processor, a digital signal processor, or any other suitable processor. The processor has data and / or signal processing functions. The processor may be implemented in software, hardware, or a combination of the two. Computer-readable storage media include any suitable media that can store program code, such as a magnetic disk, a hard disk, an optical disk, a flash memory, a read-only memory, a random access memory, and the like.

[0029] In addition, if the meaning of "and / or" appears in this application, it includes three parallel solutions. Taking "A and / or B" as an example, it includes solution A, solution B, or a solution in which A and B are satisfied at the same time. In addition, the technical solutions between the various embodiments can be combined with each other, but it must be based on the fact that ordinary technicians in this field can implement it. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such a combination of technical solutions does not exist and is not within the scope of protection required by this application. The term "at least one A or B" or "at least one of A and B" has a similar meaning to "A and / or B" and can include only A, only B, or A and B. The singular terms "one" and "this" can also include plural forms.

[0030] This application attaches great importance to the security of user personal information and has taken reasonable and feasible security protection measures that comply with industry standards to protect user information and prevent personal information from being accessed, disclosed, used, modified, damaged or lost without authorization.

[0031] Here we first explain some terms involved in this application.

[0032] Prompt: In an AI dialogue system or question-and-answer system, a prompt is a statement or question that initiates or advances a conversation. It is an initial command or trigger that causes the AI (Artificial Intelligence) software to respond.

[0033] A database query language is a language used to retrieve data from a database. It is designed to perform specific operations such as querying, updating, inserting, deleting, and creating or modifying database structures. Common database query languages include SQL, NoSQL, and XQuery.

[0034] SQL (Structured Query Language): is a standard programming language for managing relational databases.

[0035] Json: A lightweight data exchange format that stores and transmits data in the form of key-value pairs, where key-value refers to two interrelated data fragments, the key is unique, and the value is the content associated with the key.

[0036] A domain-specific language (DSL) is a computer language specifically used in a specific problem domain to simplify the understanding and processing of problems in that domain.

[0037] Generalization ability: An important indicator of an AI dialogue system, referring to its ability to understand and respond to unprecedented or unfamiliar dialogue environments, topics, and questions.

[0038] Semantic understanding capability: refers to the ability of an artificial intelligence dialogue system to understand human language and identify key information such as intentions, emotions, entities, etc.

[0039] Currently, large language models demonstrate excellent generalization capabilities in semantic understanding, achieving excellent generalization even with little or no annotation. However, these models also have significant drawbacks: their structure is relatively closed, and feedback is based on the knowledge system developed during the training phase. For example, in specific application scenarios such as performance query management systems, much knowledge is stored in user-specific databases. Directly querying large language models cannot accurately answer questions about specific databases. If the model is tuned to learn specific data, such as industry data or user data, the large number of parameters in large models makes the adjustment costly and poses certain data security risks. Currently, one method for querying data uses a knowledge base to match and locate user input information and construct search information. The search information and input information are combined to construct prompt words, which are then fed into the large language model to obtain returned information, thereby completing knowledge dialogue tasks in specific professional fields. This method can produce simple query results, but when user questions involve data operations, such as "What was the month-over-month growth rate of the highest-selling car last year?" the large model's analysis accuracy is significantly reduced. Moreover, questions and answers about industry data tables often involve multi-table queries. A question sometimes requires information from more than two data tables. The relevant information between tables may be partially lost during extraction, resulting in inaccurate subsequent queries.

[0040] Using a database query language like SQL as a bridge between the large language model and the database—for example, translating user questions into SQL and then querying the data results through a MySQL system—can, to a certain extent, address multi-table query and computation requirements. However, this approach places high demands on prompts, requiring them to include a large amount of information from the database's tables. Complex businesses with numerous tables can result in excessively long context, potentially exceeding the model's capabilities. Furthermore, if the database undergoes additions, deletions, or modifications, users are required to modify the prompt content, creating certain difficulties in interaction and usage.

[0041] To address the above issues, this embodiment provides a large data analysis method based on a large model. Figure 1 , Figure 1 This is a flow chart of the main steps of a big data analysis method based on a large model according to an embodiment of the present application. Figure 1 As shown, the method mainly includes the following steps S2 to S10:

[0042] Step S2: Obtain the user's question and determine whether the user's question requires a database query.

[0043] In this embodiment, user questions refer to specific questions raised by users, which are usually expressed in natural language. It is understood that user questions can be questions that require a database query, such as asking what is the best-selling product in a certain month, or they can be general conversation questions that do not require a database query.

[0044] In one embodiment, determining whether a user question requires a database query can be performed by analyzing the user question using natural language processing techniques to extract keywords and topic information. If the question involves specific domain knowledge or information that needs to be found in a database, a database query is determined to be necessary. Alternatively, question preprocessing, knowledge base matching, statistical analysis, and other methods can be used to determine whether a user question requires a database query.

[0045] Step S4: In response to the user question requiring a database query, the user question is parsed into a domain-specific language according to the first prompt information.

[0046] In this embodiment, when a user question requires a database query, the user question is parsed into a domain-specific language (DSL) based on the first prompt information. The first prompt information is a pre-set text segment used to guide the large language model to parse the user question into a DSL. This text segment can be a requirement for the large model, a task description, or any other information that can provide the context and subject matter required for generation.

[0047] In one embodiment, the first prompt information includes the task requirements, function definitions, parameter descriptions, and reference examples of the large model, wherein the parameter descriptions are used to specify the structure of the domain-specific language, and the task requirements include parsing the user questions into the domain-specific language based on the function definitions, parameter descriptions, and reference examples.

[0048] For example, suppose a user asks, "Please list the top ten car models sold at the Zhongshan Road store in 2020?" The pre-set first prompt information includes, but is not limited to, task requirements, function definitions, parameter descriptions, and reference examples. Task requirements, for example, involve parsing text (user questions) into DSL information; function definitions, for example, require the large model, such as specifying its identity, conversation focus, and tone; parameter descriptions define the structure of the domain-specific language, such as what information to extract from the user question and what form of DSL to construct; and reference examples, such as a specific example of converting a user question into DSL information, guiding the large model to learn the desired output style, thereby reducing misunderstandings and ambiguity.

[0049] In one embodiment, the parameter description specifies that the structure of the domain-specific language includes parameters and parameter values in the form of key-value pairs, where the parameters include entities, query dimensions, filtering rules, query indicators, operation operators, and grouping rules.

[0050] In one embodiment, the parameters further include the table name and table header name of the database.

[0051] In this embodiment, by including the database table names and table headers in the first prompt information, the DSL information can be more clearly parsed to identify the database information to be queried, further improving the accuracy of database queries and, consequently, the accuracy of large-scale model analysis. During the conversion of the DSL information into SQL commands, the database table names and table headers can be incorporated into the SQL command processing, thus avoiding issues such as incompatibility after database updates and data security requirements that may arise from prepending the database information.

[0052] In one implementation, for DSL information in the form of key-value pairs, one key may correspond to multiple values. For example, when the key is an entity, the corresponding value may be year, product, and sales volume.

[0053] Step S6: Determine a database query command according to the domain specific language.

[0054] In this embodiment, the DSL information obtained in step S4 includes entities, query objects, operators, and grouping rules identified from the user question, which correspond to common SQL statements such as SELECT FROM, operators, and GROUP BY. The query statement used to access the database can be obtained through the DSL to SQL conversion program.

[0055] In one embodiment, after parsing the user question into DSL information in step S4, the DSL information can be converted into database query commands, such as SQL commands, through text regularization and operator mapping. As an example, text regularization is used to match or search for keywords in the DSL information. Based on the number and usage of operators, a corresponding SQL statement template is created. Operator mapping is used to replace the obtained keywords with the corresponding SQL operators. The SQL statement template is then combined with the semantics of the original DSL to obtain the SQL query statement.

[0056] Step S8: Obtain query results from a preset database according to the database query command.

[0057] In this embodiment, the query result is obtained from a preset database such as a user database according to the SQL query statement obtained in step S6.

[0058] Step S10: Generate analysis results based on user questions and query results.

[0059] In this embodiment, the user question and query results are fed into the large language model to obtain analysis results corresponding to the user question. The analysis results can be in the form of charts, reports, text answers, etc., which are not specifically limited in this embodiment.

[0060] Based on the method described in steps S2 to S10 above, the user question that needs to be queried in the database is parsed into a domain-specific language according to the first prompt information, and then the database query command is determined based on the domain-specific language. The database query command is then used to obtain the query results from the preset database, and finally the analysis results are generated based on the user question and the query results. On the one hand, the excellent understanding of natural language by the large language model is utilized to parse the user question into a domain-specific language. On the other hand, the domain-specific language is converted into a database query command. While ensuring the accuracy of the database query, it also has a strong user question analysis capability, solving the problem of low accuracy of the large language model in the database query scenario. In addition, this method does not require user data to be put into model training, avoiding high customization costs and ensuring the privacy and security of user data.

[0061] In an optional embodiment, after the above step S2, the following step S3 may be further included:

[0062] Step S3: In response to the user question, an analysis result is generated according to the user question and the second prompt information without querying the database.

[0063] In this embodiment, when the user question does not require a database query, analysis results can be generated directly based on the user question and the second prompt information without conversion to a domain-specific language. The second prompt information includes, but is not limited to, a sentence, a question, a description, or any information that can provide the context and theme required for generation. It is understood that in actual applications, by observing the effects of different prompt information, the prompt information can be continuously adjusted and improved to better meet user expectations.

[0064] In an optional implementation, step S4 may further include the following steps S42 and S44:

[0065] Step S42: extracting parameter values corresponding to entities, query dimensions, filtering rules, query indicators, operation operators, and grouping rules from the user question according to the function definition, parameter description, and reference examples.

[0066] In this example, assume the user's question is: What is the top-selling car model at the Zhongshan Road store in 2020? The function definition is based on the pre-training convention of the large model and uses the following English description: You are an artificial intelligence assistant. You are designed to be helpful, honest, harmless, and good at acting in different roles. Please answer the user questions based on the user asks, descriptions, and your role. Think logically and give the correct answer. The parameter description is: write the entity object in the text into "Entity", write the relevant information of the query object into "Dimension", write the restriction condition into "Filter", write the query target into "Metrics", fill in the required operation method into "Operator", and write the grouping unit into "Groupby".

[0067] Specifically, the Entity in the user's question includes: 2020, Zhongshan Road store, sales volume, and the number one car model; the Dimension includes: time, sales point, and car model; the Filter includes: time: 2020, sales point: Zhongshan Road store; the Metrics includes: sales volume; the Operator includes: number one ranking; and the Groupby includes: car model.

[0068] Step S44 : converting the user question into a domain-specific language according to the parameter values corresponding to the entity, query dimension, filtering rule, query index, operation operator, and grouping rule.

[0069] In this embodiment, the DSL for parsing the above user question into JSON format is: {'Entity':['2020', 'Zhongshan Road Store', 'Sales Volume', 'No. 1 Car Model'], 'Dimension':['Time', 'Sales Point', 'Car Model'], 'Filters':{'Time': '2020', 'Sales Point': 'Zhongshan Road Store'}, 'Metrics':['Sales Volume'], 'Operator': 'No. 1', 'Groupby':['Car Model']}.

[0070] In an optional implementation, the above step S10 may further include the following steps S102 and S104:

[0071] Step S102: The user question and the query result are combined to obtain third prompt information.

[0072] In this example, suppose the user question "What was the top-selling car model at the Zhongshan Road store in 2020?" returns "Sports model." The user question and the query result are then concatenated to produce the following third prompt: "The question is 'What was the top-selling car model at the Zhongshan Road store in 2020?' The query result is 'Sports model.' Please output the complete analysis results for me."

[0073] Step S104: Generate a report or text result according to the third prompt information.

[0074] In this embodiment, the large model generates an analysis result corresponding to the user's question based on the third prompt information obtained by splicing in step S102. As an example, the large model ultimately outputs content to the interactive interface such as "The top-selling car model at the Zhongshan Road store in 2020 is a sports car," or outputs a 2020 Zhongshan Road store sales ranking report.

[0075] In an application scenario according to an embodiment of the present application, the current method of directly outputting SQL statements through natural language questions input by users, such as the DB-GPT project, depends on the accuracy of the SQL output by the large language model. In the face of rich SQL grammar, how to improve the accuracy of SQL statements is a huge challenge. The accuracy of SQL statements also has a certain impact on the robustness of the question-answering system. The basis for the use of existing methods is to embed all the user's database headers into the prompt part. The context is too lengthy, resulting in a high reasoning cost, and is not conducive to the actual use scenario where users can flexibly modify the table content format. In response to the above problems, Figure 2 This is a schematic diagram of the overall framework of a big data analysis method based on a large model according to an embodiment of the present application. It should be noted that the method can be implemented by a question-answering system, and the question-answering process can include one round of dialogue or multiple rounds of dialogue, such as Figure 2 As shown, the implementation process of this method is as follows:

[0076] After the question-answering begins, the Q&A system assistant can display "Hello, I am your office assistant, how can I help you?" and then obtain the user input text, that is,<user_input> In other words, the interface content of a single round of dialogue is as follows:

[0077] Assistant: Hello, I'm your office assistant. How can I help you?

[0078] User:<user_input> .

[0079] Note:<user_input> Is the user input text.

[0080] After obtaining the user input text, that is, the user question, it is necessary to determine whether the user question queries the database. If the user question does not need to query the database (the judgment result is no), that is, the user question is a general conversation question, then the common method used in the existing technology is used to analyze general conversation questions. If the user question needs to query the database (the judgment result is yes), then the user question is parsed into DSL. For details, please refer to Figure 1 Step S4 of the embodiment shown is not described here in detail. Optionally, the prompt message for calling the large language model for parsing can be the following prompt:

[0081] prompt=SYSTEM+“User:”+<user_input> + “Assistant:”

[0082] SYSTEM="You are an artificial intelligence assistant. You are designed to be helpful, honest, harmless, and good at acting in different roles. Please answer the user questions based on the user asks, descriptions, and your role. Think logically and give the correct answer.\nNext, please talk to me. If my question is about querying the database, please help me parse a text into JSON-formatted DSL information. Remember to write the entity objects in the text into "Entity" and the relevant information of the query object into "Dimension". Remember to write the restriction conditions into "Filter", the query target into "Metrics", the required operation method into "Operator", and the grouping unit into "Groupby". For example, the text: "Please list the top 10 car models in terms of sales at the Zhongshan Road store in 2020?" ", parsed into json: {'Entity':['2020', 'Zhongshan Road Store', 'Sales Volume', 'Top 10 Car Models'], 'Dimension':['Time', 'Sales Point', 'Car Model'], 'Filters':{'Time': '2020', 'Sales Point': 'Zhongshan Road Store'}, 'Metrics':['Sales Volume'], 'Operator': 'Top 10', 'Groupby':['Car Model']}.\n\n"

[0083] SYSTEM is a concatenated object of pre-set information, covering four parts: the model's functional definition, task requirements, parameter descriptions, and examples. The functional definition uses English instructions based on the model's pre-training conventions. The task requirements explicitly convert user input text into a DSL in JSON format. The parameter descriptions demonstrate the DSL structure and explain each key value, with examples providing a single demonstration. The DSL is in JSON format and contains the entity, query object, filter rules, operators, and grouping rules. This prompt design enables the model to concisely extract key information from user questions and fully document the logical requirements for querying the database. As described above, the DSL structure is designed as follows (string represents string format): {'Entity':['string'],'Dimension':['string'],'Filters':{'string':'string'},'Metrics':['string'],'Operator':'string','Groupby':['string']}}. This prompt indicates that the input content of the large language model includes two scenarios: general interaction and database query. On the other hand, it explains how to parse user questions and the expected output DSL format.

[0084] Next, the DSL is converted into SQL statements. Specifically, a computer program with preset functions, such as a program that converts the DSL into SQL statements through text regularization and operator mapping, can be used to convert the DSL information parsed by the large model into SQL commands. For example:

[0085] Suppose the user's question is: What is the best-selling car model in 2022?

[0086] The DSL obtained by parsing the large model is: {'Entity':['2022', 'Sales', 'Highest Model'], 'Dimension':['Time', 'Model'], 'Filters':{'Time': '2022'}, 'Metrics':['Sales'], 'Operator': 'Highest', 'Groupby':['Model']}.

[0087] The following SQL commands are obtained by using a preset computer program:

[0088] WITH AllSales AS(

[0089] SELECT

[0090] car_type_name,

[0091] SUM(amount)AS total_sales

[0092] FROM sales

[0093] WHERE EXTRACT(YEAR FROM sale_date)='2022'

[0094] GROUP BY car_type_name )

[0096] SELECT

[0097] car_type_name,

[0098] total_sales

[0099] FROM AllSales

[0100] ORDER BY total_sales desc

[0101] LIMIT 1;

[0102] Then use the above SQL command to query the database information and return the query results. Then call the analysis template to get the analysis results. Exemplarily, the user questions and the query results returned by the database can be passed into the large language model in the form of prompt questions, and organized into complete content and output to the interactive interface. For example, the prompt passed into the large language model is "The known question is'What is the best-selling car model in 2022', and the query result is'sports type'. Please help me output the complete analysis results." The content finally output to the interactive interface is "The best-selling car model in 2022 is the sports type" and / or the sales report of different car models in 2022. This round of question and answer ends, and multiple rounds of question and answer can be conducted according to the above process.

[0103] The method provided in this embodiment does not require the use of domain-specific knowledge to adjust a general language model. Instead, it designs a domain-specific language (DSL) for querying databases by organizing and designing a set of commonly used SQL queries and operators. Because it is much easier for a large language model to derive a DSL based on user questions than for the model to learn and generate accurate SQL commands, this method enables reliable, efficient, and robust database querying. This embodiment leverages the large language model's superior ability to understand natural language question requirements. For example, when asked, "Which product models are most popular with customers?" the large language model can compensate for the shortcomings of general programs and models, successfully interpreting the user's purpose as "Which product models have the highest sales volume?" Furthermore, by converting DSL into SQL to issue accurate database query instructions, the accuracy of database queries is improved, a feature that most current models lack. The method described in this application can handle a wide variety of questions while ensuring accuracy. It also eliminates the need to train user data on the model, avoiding high customization costs and ensuring user data security.

[0104] It should be pointed out that although the various steps in the above embodiments are described in a specific order, those skilled in the art will understand that in order to achieve the effect of the present application, different steps do not have to be performed in such an order. They can be performed simultaneously (in parallel) or in other orders. These adjusted solutions are equivalent to the technical solutions described in this application, and therefore will also fall within the scope of protection of this application.

[0105] Another aspect of the present application is to provide a big data analysis system based on a large model, such as Figure 3 As shown, the system includes: an acquisition module 2, which is used to obtain user questions and determine whether the user questions need to query the database; a parsing module 4, which is used to parse the user questions into a domain-specific language according to the first prompt information in response to the user questions needing to query the database; a determination module 6, which is used to determine the database query command according to the domain-specific language; a query module 8, which is used to obtain query results from a preset database according to the database query command; and a generation module 10, which is used to generate analysis results based on the user questions and the query results.

[0106] The above-mentioned big data analysis system based on big model is used to perform Figure 1 The big data analysis method embodiment based on the big model shown has similar technical principles, technical problems solved and technical effects produced. Technical personnel in this technical field can clearly understand that for the convenience and conciseness of description, the specific working process and related instructions of the system can refer to the contents described in the above method embodiment, and will not be repeated here.

[0107] In an optional embodiment, Figure 4 This is a schematic diagram of the overall framework of a big data analysis system based on a large model according to an embodiment of the present application. Figure 4 As shown, the system mainly includes the following modules:

[0108] Large model interaction module: To serve database query tasks, this module uses preset prompt content as a question template. On the one hand, it prompts the large language model input content for two scenarios: general interaction and database query. On the other hand, it explains how to parse user questions and require the output DSL format.

[0109] DSL information to SQL command module: Through the preset DSL to SQL program, the value corresponding to each key is extracted from the DSL information of the JSON structure, and accurate SQL query statements are obtained based on rule combination.

[0110] Database access module: accesses the database provided by the user by calling the preset database access program and obtains query results according to SQL statements.

[0111] Answer integration module: Combines user questions with returned query results according to preset templates to achieve complete information output.

[0112] The above system and Figure 2 The big data analysis method embodiment based on the big model shown corresponds to the embodiment of the big data analysis method. The technical principles, technical problems solved and technical effects produced by the two are similar. Technical personnel in this technical field can clearly understand that for the convenience and conciseness of description, the specific working process of the system and related instructions can refer to the contents described in the above method embodiment, and will not be repeated here.

[0113] This embodiment cleverly combines a large language model with high generalization performance with a high-accuracy computer program. It effectively leverages the large prediction model's superior semantic understanding capabilities, directly addresses user needs and effectively resolves problems. It also collaborates with the computer program to process database queries in the background, resulting in high practicality. This embodiment provides a solution for interactively querying database information with users. By utilizing a universal large language model and a series of pre-configured computer programs, it creates an intermediate path in the conversion process between natural language and SQL commands, avoiding the high training costs required to directly use a large language model to output SQL. Effective query key information can be obtained solely through prompt instructions, and the computer program's fast execution and accurate matching ensure the accuracy of SQL command output.

[0114] It will be understood by those skilled in the art that all or part of the processes in the method for implementing the above embodiment of the present application can also be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, it can implement the steps of the above-mentioned various method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file or some intermediate form. The computer-readable storage medium may include: any entity or device, medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory, random access memory, electric carrier signal, telecommunication signal and software distribution medium, etc. that can carry the computer program code.

[0115] Another aspect of the present application provides a computer-readable storage medium.

[0116] In an embodiment of a computer-readable storage medium according to the present application, the computer-readable storage medium can be configured to store a program for executing the large model-based big data analysis method of the above-mentioned method embodiment, and the program can be loaded and run by the processor to implement the above-mentioned large model-based big data analysis method. For ease of explanation, only the parts related to the embodiment of the present application are shown. For specific technical details not disclosed, please refer to the method part of the embodiment of the present application. The computer-readable storage medium can be a storage device formed by various electronic devices. Optionally, the computer-readable storage medium in the embodiment of the present application is a non-transitory computer-readable storage medium.

[0117] Another aspect of the present application provides a smart device.

[0118] In an embodiment of a smart device according to the present application, the smart device may include at least one processor; and a memory in communication with the at least one processor; wherein the memory stores a computer program, and when the computer program is executed by the at least one processor, the method described in any of the above embodiments is implemented. Figure 5 , Figure 5 exemplarily shows that the memory 11 and the processor 12 are communicatively connected via a bus.

[0119] In some embodiments of the present application, the smart device may further include at least one sensor for sensing information. The sensor is communicatively connected to any type of processor mentioned in the present application. Optionally, the smart device described in the present application may be, but is not limited to, a mobile phone, a tablet computer, a desktop, a laptop, a handheld computer, a notebook computer, a vehicle-mounted device, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), an augmented reality (AR) or virtual reality (VR) device, etc., and the embodiments of the present application are not limited thereto.

[0120] Thus far, the technical solution of the present application has been described in conjunction with an embodiment shown in the accompanying drawings. However, it is readily understood by those skilled in the art that the scope of protection of the present application is obviously not limited to these specific embodiments. Without departing from the principles of the present application, those skilled in the art may make equivalent changes or substitutions to the relevant technical features, and the technical solutions after such changes or substitutions will fall within the scope of protection of the present application.

Claims

1. A big data analysis method based on a big model, characterized in that: The method comprises: Obtaining a user question and determining whether the user question requires querying a database; In response to the user question requiring a database query, parsing the user question into a domain-specific language according to the first prompt information; determining a database query command according to the domain specific language; Obtaining query results from a preset database according to the database query command; An analysis result is generated according to the user question and the query result.

2. The method according to claim 1, characterized in that The method further comprises: In response to the user question, an analysis result is generated based on the user question and the second prompt information without querying a database.

3. The method according to claim 1, characterized in that The first prompt information includes the task requirements, function definitions, parameter descriptions and reference examples of the large model, wherein the parameter descriptions are used to specify the structure of the domain-specific language, and the task requirements include parsing the user questions into domain-specific language based on the function definitions, parameter descriptions and reference examples.

4. The method according to claim 3, characterized in that The parameter description specifies the structure of the domain-specific language, including parameters and parameter values in the form of key-value pairs, wherein the parameters include entities, query dimensions, filtering rules, query indicators, operation operators, and grouping rules.

5. The method according to claim 4, characterized in that Parsing the user question into a domain-specific language according to the first prompt information includes: Extracting parameter values corresponding to entities, query dimensions, filtering rules, query indicators, operation operators, and grouping rules from the user question according to the function definition, parameter description, and reference examples; The user question is converted into a domain-specific language according to the parameter values corresponding to the entities, query dimensions, filtering rules, query indicators, operation operators and grouping rules.

6. The method according to claim 1, characterized in that The analysis result includes a report or text result. The analysis result is generated based on the user question and the query result, including: Concatenate the user question and the query result to obtain third prompt information; The report or text result is generated according to the third prompt information.

7. The method according to claim 4, characterized in that The parameters also include the table name and table header name of the database.

8. A big data analysis system based on a big model, characterized in that: The system comprises: An acquisition module is used to obtain user questions and determine whether the user questions require a database query; a parsing module, configured to query a database in response to the user question and parse the user question into a domain-specific language according to the first prompt information; a determination module, configured to determine a database query command according to the domain specific language; A query module, configured to obtain query results from a preset database according to the database query command; A generation module is used to generate analysis results based on the user question and the query result.

9. A smart device, characterized in that: include: at least one processor; and, a memory communicatively coupled to the at least one processor; The memory stores a computer program, and when the computer program is executed by the at least one processor, the method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium storing a plurality of program codes, characterized in that: The program code is suitable for being loaded and executed by a processor to perform the method according to any one of claims 1 to 7.