Data analysis method, device and equipment and computer storage medium
By obtaining input problem information, matching target data tables and indicator information, and using big models to generate query statements, the problem of unified management of indicators in the business intelligence platform is solved, and efficient and accurate data analysis and indicator management are achieved.
Patent Information
- Application Number
- CN202510215791.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-07-18
AI Technical Summary
The existing business intelligence platform has problems such as large workload and high cost in unified management of indicators and maintenance of caliber changes, which affect data accuracy and consistency.
By obtaining input problem information, matching the target data table and pre-bound indicator information, using the big model to generate query statements, combining the indicator knowledge base and graphical page for binding and update, and generating data analysis results.
It greatly improves the generation accuracy of indicators to SQL, improves the efficiency and accuracy of data analysis, lowers the threshold for use, and adapts to the needs of specialized analysis in specific fields.
Smart Images

Figure CN120336346A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of indicator analysis, and particularly to a data analysis method, a data analysis device, a data analysis equipment, and a computer storage medium. Background Art
[0002] With the rapid development of big data technology, many fields such as enterprises and government affairs have introduced big data-based indicator statistics and analysis to evaluate and optimize the operation status. Big data analysis not only helps various organizations better understand market dynamics, user behavior, financial health, and operational efficiency, but also provides accurate and reliable data support for real-time monitoring and strategy formulation. Therefore, building and applying an efficient big data indicator system has become an indispensable part of modern organizations.
[0003] With the rise of large model technology, empowering traditional BI (Business Intelligence) analysis has become a new trend. By integrating the functions of conversational question and chart automatic generation, the new generation of big data analysis tools enables enterprise executives to directly interact with data, thereby quickly obtaining the required key information. This not only improves the efficiency of data analysis, but also significantly enhances the accuracy and quality of decision-making. With the support of this large model technology, organizations can make more accurate business decisions in a shorter time, thus gaining an advantage in the highly competitive market environment.
[0004] BI tools are designed for data analysis and display scenarios and do not solve the problem of unified management of indicators themselves. Engineers conduct modeling in the data warehouse and develop a large number of wide tables and summary tables, and then import the data sets or wide tables into BI for analysis. In this mode, indicators are usually scattered in various reports and data sets, and the same indicator may have inconsistent calibers between different data sets. Whenever the business requirements change, it is often necessary to change the indicator calibers in different BI tools or data sets. The workload and cost of changing and maintaining the indicator calibers are large, and it brings potential differences in data results, affecting the accuracy and consistency of data. Summary of the Invention
[0005] To solve the above technical problems, this application proposes a data analysis method, a data analysis device, a data analysis equipment, and a computer storage medium.
[0006] To solve the above technical problems, this application proposes a data analysis method, and the data analysis method includes:
[0007] Obtain input problem information;
[0008] Based on the input problem information, obtain a target data table;
[0009] Based on the target data table, obtain the pre-bound metric information;
[0010] Input the table information, metric information of the target data table, and the input question information into the large model to obtain the query statement generated by the large model;
[0011] Use the query statement to obtain the data analysis result.
[0012] Among them, the obtaining of the target data table based on the input question information includes:
[0013] Extract the word vectors in the input question information;
[0014] Use the word vectors to match the table information of the data tables in the database;
[0015] Determine the data tables with a matching similarity higher than the preset threshold as the target data tables.
[0016] Among them, the obtaining of the target data table based on the input question information includes:
[0017] Extract the question intent in the input question information;
[0018] Extract the intent classification results of each data table in the database;
[0019] Determine the data tables corresponding to the intent classification results that conform to the question intent as the target data tables.
[0020] Among them, after obtaining the pre-bound metric information based on the target data table, the data analysis method further includes:
[0021] Input the metric information and the input question information into the large model to obtain the relevant metric information output by the large model;
[0022] The inputting the table information, metric information of the target data table, and the input question information into the large model to obtain the query statement generated by the large model includes:
[0023] Input the table information, relevant metric information of the target data table, and the input question information into the large model to obtain the query statement generated by the large model.
[0024] Among them, the data analysis method further includes:
[0025] Define a metric knowledge base, where the metric knowledge base includes at least metric names and metric dimensions;
[0026] Obtain the summary information of the data tables to be associated;
[0027] Traverse each field of the summary information, and obtain relevant indicator names and relevant indicator dimensions whose similarity meets the preset conditions from the indicator knowledge base;
[0028] Generate a prompt message by using the summary information, the relevant indicator names, and the relevant indicator dimensions;
[0029] Input the prompt message into the large model for recognition, and obtain the associated indicator names and associated indicator dimensions output by the large model;
[0030] Bind the associated indicator names and the associated indicator dimensions to the data table to be associated.
[0031] Wherein, after binding the associated indicator names and the associated indicator dimensions to the data table to be associated, the data analysis method further includes:
[0032] Display the binding relationship on a graphical page;
[0033] Update the binding relationship based on an edit instruction;
[0034] Save the updated binding relationship based on a save instruction.
[0035] Wherein, the obtaining of the data analysis result by using the query statement includes:
[0036] Call the data warehouse interface to submit the query statement to other data development platforms to obtain the data analysis results of other data development platforms;
[0037] Display the data analysis result on a graphical page.
[0038] To solve the above technical problems, the present application also proposes a data analysis device, which includes: an input module, an extraction module, a query module, and an analysis module; wherein,
[0039] The input module is used to obtain input problem information;
[0040] The extraction module is used to obtain a target data table based on the input problem information;
[0041] The extraction module is used to obtain pre-bound indicator information based on the target data table;
[0042] The query module is used to input the table information, indicator information, and input problem information of the target data table into the large model to obtain a query statement generated by the large model;
[0043] The analysis module is used to obtain a data analysis result by using the query statement.
[0044] To solve the above technical problems, the present application also proposes a data analysis device, which includes a memory and a processor coupled to the memory; wherein, the memory is used to store program data, and the processor is used to execute the program data to implement the data analysis method as described above.
[0045] To solve the above technical problems, the present application also proposes a computer storage medium, which is used to store program data, and when the program data is executed by a computer, it is used to implement the above data analysis method.
[0046] Compared with the prior art, the beneficial effects of the present application are as follows: The data analysis device obtains input problem information; based on the input problem information, it obtains a target data table; based on the target data table, it obtains pre-bound index information; it inputs the table information, index information of the target data table, and the input problem information into a large model to obtain a query statement generated by the large model; and it uses the query statement to obtain a data analysis result. Through the above data analysis method, relevant tables and indexes are recalled from the knowledge base, and then the large model is used to generate SQL (Structured Query Language) according to the relevant tables and indexes, which greatly improves the generation accuracy from indexes to SQL and improves the data analysis effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0048] Among them:
[0049] Figure 1 is a schematic diagram of the architecture of the data analysis system based on a large model and index management provided by the present application;
[0050] Figure 2 is a schematic diagram of the composition of an embodiment of a derived index provided by the present application;
[0051] Figure 3 is a schematic diagram of the process of an embodiment of the data analysis method provided by the present application;
[0052] Figure 4 is a schematic diagram of the process of the index query and analysis solution provided by the present application;
[0053] Figure 5 is a schematic diagram of the process of another embodiment of the data analysis method provided by the present application;
[0054] Figure 6 It is a schematic flowchart of the index definition and index construction scheme provided by this application;
[0055] Figure 7 It is a schematic structural diagram of an embodiment of the data analysis device provided by this application;
[0056] Figure 8 It is a schematic structural diagram of an embodiment of the data analysis device provided by this application;
[0057] Figure 9 It is a schematic structural diagram of an embodiment of the computer storage medium provided by this application. Detailed implementation manners
[0058] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.
[0059] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application described here, for example, can be implemented in an order other than those illustrated or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products, or devices.
[0060] However, traditional Business Intelligence (BI) platforms face certain challenges during application. Although these platforms provide graphical interfaces and simplify the data query and display processes, there are still relatively high usage thresholds. Especially for senior leaders, they often lack the time and technical background to conduct data analysis themselves. This situation usually forces senior leaders to rely on the analysis and reports of the data team. However, this reliance has many problems in terms of timeliness, business professionalism, and data credibility, and there may even be a risk of data tampering.
[0061] Therefore, the present application provides a data analysis method and system based on large models and metric management. By combining large model technology with metric management capabilities, the accuracy and efficiency of analyzing specialized and customized metrics for specific domains are improved, the usage threshold is reduced, and the timeliness and flexibility of data analysis are enhanced.
[0062] For the overall architecture involved in the data analysis system based on large models and metric management provided by the present application, please refer to Figure 1 , Figure 1 which is a schematic diagram of the architecture of the data analysis system based on large models and metric management provided by the present application.
[0063] As Figure 1 shown, the overall architecture in the data analysis system of the present application can be divided into the following three layers:
[0064] 1. Metric Query and Analysis Layer: Based on large model capabilities and metric definitions, understand and parse user questions and generate query statements.
[0065] 2. Metric Storage and Management Layer: Define and manage system metrics, establish indexes, retrieve and recall metrics, and establish mapping relationships between metrics and fields, etc.
[0066] 3. Query and Calculation Layer: Corresponding to the backend databases and data warehouses of various BI analysis systems.
[0067] Regarding the definitions of the metrics involved in the data analysis system, they include but are not limited to the following three categories of metrics:
[0068] 1) Atomic Metrics
[0069] Atomic metrics refer to the smallest measurable units in data analysis, usually a numerical value or a count. Atomic metrics are the basis of data analysis. They can be used to describe a specific event, behavior, or state, such as sales amount, number of visits, conversion rate, etc. Atomic metrics are usually indivisible because they are already the smallest measurable units. Atomic metrics do not exist independently. They must be combined with the business scope and dimensions to make sense. Atomic metrics are usually the basis for other metrics, and higher-level metrics can be obtained through the analysis of atomic metrics.
[0070] 2) Derived Metrics
[0071] Derived metrics are composed of three major elements: atomic metrics, time period, and dimension within the scope defined by the business. Specifically, such as Figure 2As shown, it is used to count the numerical performance of target indicators under specific time, dimensions, and business conditions, reflecting the business status of a certain business activity of an enterprise. The indicators used in the business are all derived indicators. Different derived indicators may have the same atomic indicators, so the derived indicators define an equivalence relationship, and belonging to the same atomic indicators constitutes a partition of the indicator system. Among them, Figure 2 is a schematic diagram of the composition of an embodiment of the derived indicator provided by this application.
[0072] 3) Composite indicator
[0073] Another type of derived indicator is a composite indicator, which is also called a calculation indicator in some materials. For example:
[0074] Average selling price: The derived indicator is calculated by sales amount and sales volume, which reflects the average selling price of each product. The atomic indicators are sales amount and sales volume.
[0075] Conversion rate: The derived indicator is calculated by the number of visits and the number of conversions, which reflects the conversion effect of each channel. The atomic indicators are the number of visits and the number of conversions.
[0076] Customer lifetime value: The derived indicator is calculated by the average purchase amount of customers, purchase frequency, and customer retention rate, which reflects the contribution value of each customer to the enterprise. The atomic indicators are customer purchase amount, purchase frequency, and customer retention rate.
[0077] Based on the above system architecture and technical foundation, for details, please continue to refer to Figure 3 and Figure 4 , Figure 3 is a schematic flowchart of an embodiment of the data analysis method provided by this application, Figure 4 is a schematic flowchart of the indicator query and analysis solution provided by this application.
[0078] The data analysis method of this application is applied to a data analysis device. Among them, the data analysis device of this application can be a server, or a terminal device, or a system composed of a server and a terminal device cooperating with each other. Correspondingly, each part included in the data analysis device, such as each unit, subunit, module, and submodule, can be all set in the server, or all set in the terminal device, or respectively set in the server and the terminal device.
[0079] Furthermore, the above server can be hardware or software. When the server is hardware, it can be implemented as a distributed server cluster composed of multiple servers, or as a single server. When the server is software, it can be implemented as multiple software or software modules, such as software or software modules used to provide a distributed server, or as a single software or software module, which is not specifically limited here.
[0080] As Figure 3 shown, the specific steps are as follows:
[0081] Step S11: Obtain input problem information.
[0082] In the embodiment of the present application, the user inputs a problem through the graphical page of the data analysis device, such as "Analyze the monthly average unit price of a certain brand of mobile phone orders in the past three months."
[0083] Step S12: Based on the input problem information, obtain the target data table.
[0084] In the embodiment of the present application, the data analysis device recalls relevant table information in combination with the problem input by the user, that is, searches for relevant target data tables from the database.
[0085] Specifically, the data analysis device performs word segmentation on the problem, calculates the vector of the word segmentation, and recalls tables and fields with high similarity from the metadata of the table. The table recall methods provided in the present application include but are not limited to:
[0086] Coarse recall: First, through simple keyword matching or rule-based methods, quickly recall a batch of potentially relevant tables. For example, data tables containing keywords such as "order", "mobile phone", and "a certain brand".
[0087] Fine recall: Based on vector similarity or large models, recall the data tables that best meet the problem intent. For example, according to the intent classification results, preferentially recall data tables related to "order analysis" and "brand analysis".
[0088] Step S13: Based on the target data table, obtain the pre-bound metric information.
[0089] In the embodiment of the present application, the data analysis device continues to recall relevant metrics according to one or more target data tables determined in step S12.
[0090] Specifically, the metric recall methods provided in the present application include but are not limited to:
[0091] Metric coarse recall: Since Figure 1 as shown in the metric storage in the "metric definition and index construction" link of the management layer, the main fields of each data table in steps S12 and S13 have been associated with standard metrics (including atomic metrics, derived metrics, etc.), dimensions, and calculation formulas. Therefore, the data analysis device can directly obtain all relevant metrics, dimensions, and calculation methods of composite metrics for each data table.
[0092] Accurate recall of indicators: To further improve the accuracy of indicator recall, the recall results of the above-mentioned rough recall of indicators can be combined with the user's question and input into the large model together. The large model then selects the most relevant indicators, dimensions, and calculation formulas for the user's question again. For example, for the above problem case, the output finally given by the large model may contain the following key information:
[0093]
[0094]
[0095]
[0096] It should be noted that the above examples are for reference only to explain the necessary information that the output should contain, and any other formats containing this information are acceptable.
[0097] Step S14: Input the table information, indicator information, and input question information of the target data table into the large model to obtain the query statement generated by the large model.
[0098] In the embodiment of the present application, the data analysis device calls the large model to generate a query SQL statement. Specifically, the data analysis device can obtain the table Schema information according to the recalled table name. The user's original question, the above accurate recall results, and the table Schema information are input into the large model together, and the large model generates an execution SQL.
[0099] Step S15: Use the query statement to obtain the data analysis result.
[0100] In the embodiment of the present application, the data analysis device calls the underlying database or data warehouse interface, or docks with other data development platforms and submits task parameters to complete the execution of the SQL statement and obtain the result, including but not limited to: the results and the returned data, or failure and the corresponding error reason, and presents the final result to the user in a graphical page manner.
[0101] The analysis results of the data analysis method of the present application have been actually tested. By pre-defining and recalling indicators, as well as the dimensions and fields associated with the indicators, it is possible to accurately generate indicator query SQLs within a specific industry domain, with an accuracy rate of over 90%. And there are no special requirements for model training fine-tuning, and using general models with strong SQL generation capabilities (such as Qwen2.5-14B, Qwen2.5 Coder-7B, etc.) can basically meet the requirements.
[0102] In this application, a data analysis device obtains input question information; based on the input question information, obtains a target data table; based on the target data table, obtains pre-bound metric information; inputs the table information, metric information, and the input question information of the target data table into a large model to obtain a query statement generated by the large model; and uses the query statement to obtain a data analysis result. Through the above data analysis method, relevant tables and metrics are recalled from the knowledge base, and then the large model is allowed to generate SQL (Structured Query Language) based on the relevant tables and metrics, greatly improving the generation accuracy of metrics to SQL and enhancing the data analysis effect.
[0103] Further, for the specific data analysis method in the "metric definition and index construction" link in the above embodiment, please refer to Figure 5 and Figure 6 , Figure 5 which is a schematic flowchart of another embodiment of the data analysis method provided by this application, Figure 6 and is a schematic flowchart of the metric definition and index construction solution provided by this application.
[0104] As Figure 5 shown, the specific steps are as follows:
[0105] Step S21: Define a metric knowledge base, where the metric knowledge base includes at least a metric name and a metric dimension.
[0106] In the embodiments of this application, different fields, different industry institutions, and different enterprises have their own metrics, and the definitions of some seemingly similar metrics also vary among different organizations. Therefore, in this link, system administrators and business experts are still required to provide defined metric items, and various metric definitions can be imported through forms such as graphical pages, imported files, and docking other system interfaces.
[0107] Among them, metric definitions include but are not limited to the following dimensions: metric name; metric meaning; metric calculation formula (for computational metrics, such as average price, year-on-year growth rate, the metrics involved in the calculation are usually other existing atomic or derived metrics, or some constants); metric association dimension (such as sales time, product type, product name, etc.).
[0108] Step S22: Obtain the summary information of the data tables to be associated.
[0109] In the embodiments of this application, the data analysis device imports the metadata of the data tables to be associated, that is, the summary information.
[0110] Specifically, the data analysis device imports metadata from a database, a data warehouse, or other existing data management platforms, or manually enters metadata. Among them, the metadata includes, but is not limited to, table names, table descriptions, field names, field descriptions, field types, sample data, etc.
[0111] Step S23: Traverse each field of the feed information, and obtain relevant indicator names and relevant indicator dimensions that meet the preset conditions from the indicator knowledge base.
[0112] In the embodiment of the present application, the data analysis device establishes an association between indicators and the metadata of data tables. Among them, the data analysis device combines table metadata and atomic indicator definitions, and automatically invokes a large model to infer relevant association relationships.
[0113] Specifically, the data analysis device traverses each data table to be associated, and for each field of the data table to be associated, retrieves and recalls indicator items with higher similarity from the indicator library, including but not limited to: string matching retrieval and recall, semantic similarity retrieval and recall, etc.
[0114] Step S24: Generate a prompt message using the feed information, relevant indicator names, and relevant indicator dimensions.
[0115] In the embodiment of the present application, the data analysis device generates a prompt message using the feed information, relevant indicator names, and relevant indicator dimensions, inputs it into the large model for recognition, and recommends the association relationships between fields and indicators and indicator association dimensions.
[0116] In a specific implementation manner, the input of the large model Prompt (prompt message):
[0117] You are a data expert proficient in data governance, indicator definition and management, and BI analysis.
[0118] Currently, three types of indicators are defined in the system: atomic indicators, derived indicators, and composite indicators:
[0119] ……
[0120] Your task is to find out which fields in the table are associated with these indicators and indicator dimensions based on the given database table description and Schema definition, and relevant atomic indicators, derived indicators, and associated dimensions, and composite indicators.
[0121] Input:
[0122] 1) Table Schema:
[0123]
[0124]
[0125] 2) Related metrics and dimensions:
[0126] Atomic metrics: Sales amount, number of visits, conversion rate
[0127] Derived metrics: Monthly sales amount, quarterly sales amount, single - product sales amount
[0128] Dimensions: Monthly, quarterly, single - product ID
[0129] >>> Please start answering!
[0130] Step S25: Input the prompt information into the large - model for recognition, and obtain the associated metric names and associated metric dimensions output by the large - model.
[0131] In the embodiment of the present application, specific examples of the associated metric names and associated metric dimensions output by the large - model are as follows:
[0132] Atomic metric association relationship: total_price - Sales amount
[0133] Derived metric association relationship: total_price - Monthly sales amount, quarterly sales amount, single - product sales amount
[0134] Derived metric associated dimensions: created_at - Monthly, quarterly; product_id - Single - product ID
[0135] Step S26: Bind the associated metric names and associated metric dimensions to the data table to be associated.
[0136] In the embodiment of the present application, based on the output of the large - model in step S25 above, the data analysis device displays the above - mentioned association relationship on the graphical page for the administrator to finally confirm, or the administrator can manually modify and edit to form the final mapping relationship, and then the platform saves and records this relationship.
[0137] The data analysis method of the present application utilizes the capabilities of the large - model to automatically identify the association relationships between metrics and database tables and fields, and supports manual modification, review, and confirmation, greatly reducing the workload and improving the recognition efficiency and accuracy.
[0138] In the query stage of the data analysis method of the present application, first recall relevant tables, metrics, and the association relationships and metric calculation methods between metrics and data tables and fields from the knowledge base, and then let the large - model generate SQL, greatly improving the accuracy of generating SQL from metrics. Especially for industry - specific metrics and user - defined metrics, it can quickly adapt and accurately identify.
[0139] In the metric and table recall stage of the data analysis method of the present application, a two - stage processing of "coarse recall, precise recall" is adopted to improve the recall rate and accuracy, and ensure the accuracy of subsequent SQL generation.
[0140] The data analysis method of this application can be well compatible and interoperable with traditional indicator platforms and BI platforms. It can import indicator definitions in the original platform through a predefined format, and then establish the association relationships between indicators, tables, and fields based on the method of this proposal.
[0141] The data analysis method of this application has a low dependence on large model training and fine-tuning, and a general model can meet the vast majority of requirements.
[0142] The data analysis method of this application has wide applicability in the industry. By customizing the indicator knowledge base and using the large model to automatically identify association relationships, in theory, it can be quickly applied to any industry and institution, meeting the analysis requirements of specific fields or even enterprise-private indicators.
[0143] Those skilled in the art can understand that in the above method of the specific implementation manner, the writing order of each step does not mean a strict execution order and does not impose any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.
[0144] To implement the above data analysis method, this application also proposes a data analysis device. For details, please refer to Figure 7 , Figure 7 which is a schematic structural diagram of an embodiment of the data analysis device provided by this application.
[0145] The data analysis device 500 of this embodiment includes: an input module 51, an extraction module 52, a query module 53, and an analysis module 54.
[0146] Among them, the input module 51 is used to obtain input problem information.
[0147] The extraction module 52 is used to obtain a target data table based on the input problem information.
[0148] The extraction module 52 is used to obtain pre-bound indicator information based on the target data table.
[0149] The query module 53 is used to input the table information, indicator information of the target data table, and the input problem information into the large model to obtain a query statement generated by the large model.
[0150] The analysis module 54 is used to obtain a data analysis result by using the query statement.
[0151] To implement the above data analysis method, this application also proposes a data analysis device. For details, please refer to Figure 8 , Figure 8 which is a schematic structural diagram of an embodiment of the data analysis device provided by this application.
[0152] The data analysis device 400 in this embodiment includes a processor 41, a memory 42, an input / output device 43, and a bus 44.
[0153] The processor 41, the memory 42, and the input / output device 43 are respectively connected to the bus 44. Program data is stored in the memory 42, and the processor 41 is configured to execute the program data to implement the data analysis method described in the above embodiment.
[0154] In an embodiment of the present application, the processor 41 may also be referred to as a CPU (Central Processing Unit). The processor 41 may be an integrated circuit chip with signal processing capabilities. The processor 41 may also be a general-purpose processor, a digital signal processor (DSP, Digital Signal Process), an application specific integrated circuit (ASIC, Application Specific Integrated Circuit), a field programmable gate array (FPGA, Field Programmable Gate Array), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor may be a microprocessor or the processor 41 may also be any conventional processor, etc.
[0155] The present application also provides a computer storage medium. Please continue to refer to Figure 9 , Figure 9 which is a schematic structural diagram of an embodiment of the computer storage medium provided by the present application. A computer program 61 is stored in the computer storage medium 600. When the computer program 61 is executed by a processor, it is used to implement the data analysis method described in the above embodiment.
[0156] When the embodiment of the present application is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical disks, etc., which can store program codes.
[0157] The above are only the embodiments of the present application, and do not thus limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall similarly be included within the patent protection scope of the present application.
Claims
1. A data analysis method, characterized in that, The data analysis method includes: Obtain input problem information; Based on the input problem information, obtain a target data table; Based on the target data table, obtain pre-bound metric information; Input the table information, metric information, and input problem information of the target data table into a large model to obtain a query statement generated by the large model; Use the query statement to obtain a data analysis result.
2. The data analysis method according to claim 1, wherein the step of obtaining a target data table based on the input problem information includes: Extract the word vectors in the input problem information; Use the word vectors to match the table information of the data tables in the database; Determine the data tables with a matching similarity higher than a preset threshold as the target data table.
3. The data analysis method according to claim 1 or 2, wherein the step of obtaining a target data table based on the input problem information includes: Extract the problem intention of the input problem information; Extract the intention classification results of each data table in the database; Determine the data tables corresponding to the intention classification results that match the problem intention as the target data table.
4. The data analysis method according to claim 1, wherein after obtaining the pre-bound metric information based on the target data table, the data analysis method further includes: Input the metric information and the input problem information into a large model to obtain relevant metric information output by the large model; The step of inputting the table information, metric information, and input problem information of the target data table into a large model to obtain a query statement generated by the large model includes: Input the table information, relevant metric information, and input problem information of the target data table into a large model to obtain a query statement generated by the large model.
5. The data analysis method according to claim 1 or 4, wherein the data analysis method further includes: Define a metric knowledge base, where the metric knowledge base includes at least metric names and metric dimensions; Obtain the summary information of the data table to be associated; Traverse each field of the summary information, and obtain relevant metric names and relevant metric dimensions that meet the preset conditions from the metric knowledge base; Generate prompt information using the summary information, the relevant metric names, and the relevant metric dimensions; Input the prompt information into a large model for recognition to obtain the associated metric names and associated metric dimensions output by the large model; Bind the associated metric names and the associated metric dimensions to the data table to be associated.
6. The data analysis method according to claim 5, wherein after binding the associated metric names and the associated metric dimensions to the data table to be associated, the data analysis method further includes: Display the binding relationship on a graphical page; Update the binding relationship based on an edit instruction; Save the updated binding relationship based on a save instruction.
7. The data analysis method according to claim 1, wherein the step of using the query statement to obtain a data analysis result includes: Call the data warehouse interface to submit the query statement to other data development platforms to obtain the data analysis results of other data development platforms; Display the data analysis results on a graphical page.
8. A data analysis device, characterized in that, The data analysis device includes: an input module, an extraction module, a query module, and an analysis module; wherein, The input module is used to obtain input problem information; The extraction module is used to obtain a target data table based on the input problem information; The extraction module is used to obtain pre-bound metric information based on the target data table; The query module is used to input the table information, metric information, and input problem information of the target data table into a large model to obtain a query statement generated by the large model; The analysis module is used to obtain data analysis results using the query statement.
9. A data analysis device, characterized in that, The data analysis device includes a memory and a processor coupled to the memory; Wherein, the memory is used to store program data, and the processor is used to execute the program data to implement the data analysis method according to any one of claims 1 to 7.
10. A computer storage medium, characterized in that, The computer storage medium is used to store program data, and the program data, when executed by a computer, is used to implement the data analysis method according to any one of claims 1 to 7.
Citation Information
Cited By
BI intelligent question-answering system and method based on business rule retrieval and AI workflow
CN121117166A