Data analysis method and architecture based on AIGC and application and training of data analysis method and architecture
By using an AIGC-based architecture and methodology, large-scale language models are employed to improve the efficiency and accuracy of data analysis. This addresses the issue of data analysts requiring a high level of business understanding, enabling more efficient and accurate data analysis while reducing the need for specialized knowledge.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 张晏铭
- Filing Date
- 2023-05-24
- Publication Date
- 2026-04-24
AI Technical Summary
In current data analysis processes, data exploration is inefficient, requires data analysts to have a high level of understanding of the business, wastes a lot of time, and must be done manually.
By adopting an AIGC-based architecture and methodology, we improve the natural language processing and semantic parsing capabilities of large language models through training, guide large language models to gradually complete data analysis, and combine deep learning technology to understand user intent and needs, thereby reducing reliance on professional knowledge.
It improves the efficiency and accuracy of data analysis, enables more people to participate in data analysis, reduces reliance on professional knowledge, and allows users to participate more proactively and meet their personalized needs.
Smart Images

Figure CN121920371A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of data analysis technology, and specifically relates to data analysis methods, architectures, applications and training based on AIGC. Background Technology
[0002] In existing data workflows, when a data analyst performs data analysis on a cleaned data table, the first step is to clarify the analysis objective (e.g., exploring patterns in sales data). This involves value extraction from the data. During this process, the data analyst uses methods such as visualization and descriptive statistics to conduct preliminary exploration and analysis, understanding the data's distribution and correlations. Only after uncovering patterns do they proceed to the next stage of analysis. After the data mining analysis results are obtained, the data analyst writes corresponding data insights based on their business understanding and personal experience.
[0003] In this process, data exploration is an inefficient activity that requires data analysts to have a high level of understanding of the business, wastes a lot of time on data exploration, and must be done manually.
[0004] Terminology Explanation
[0005] - Data Analyst: A professional who performs data analysis.
[0006] - Data Exploration: Data exploration refers to the preliminary exploratory analysis of data during the data analysis process to understand its basic characteristics and patterns. Its purpose is to help data analysts better understand and utilize the data, providing foundational support for subsequent data modeling and analysis. During the data exploration phase, a series of visualization techniques are typically used to display the data's distribution, outliers, and relationships between variables. Furthermore, basic statistical measures such as the mean, median, and standard deviation can be calculated to understand the data's fundamental statistical characteristics. Through data exploration, data analysts can better understand the inherent patterns in the data, effectively identify problems and anomalies, and provide support for subsequent data preprocessing and modeling.
[0007] Data mining is the process of deeply mining and analyzing data through various data analysis techniques and algorithms. Data mining aims to discover hidden patterns, correlations, and regularities within large amounts of data to provide useful insights and information for decision-making. Data mining can be applied to various fields, such as marketing, finance, healthcare, and scientific research. In this data-driven decision-making process, data mining techniques can help businesses and organizations better understand their customers, products, and market trends, thereby better targeting and meeting needs. Summary of the Invention
[0008] The purpose of this invention is to provide a data analysis method, architecture, and its application and training based on AIGC. Through reinforcement learning of a large language model using a training set, the large language model acquires the ability to process and semantically parse user-input natural language, thereby improving the efficiency and accuracy of data analysis. Training the large language model enhances its ability to process natural language and semantically parse, strengthening its understanding of data analysis and thus improving both efficiency and accuracy in the data analysis process. A guided approach allows the large language model to gradually complete data analysis; this guidance can have a global impact or target only a specific stage, allowing users to better participate in the data analysis process and further improving accuracy. Deep learning technology is used to understand user-input natural language text and attempt to infer user intent and needs, effectively improving the efficiency and accuracy of data analysis. This makes it easier for users to express their needs without needing to master technical jargon, reducing reliance on specialized data analysis knowledge and allowing more people to participate in the data analysis process.
[0009] Firstly, this invention provides an AIGC-based architecture, including a main module, a large-scale language model, a big data computing engine, a data warehouse, a training set, and a user. The main module interacts with the large-scale language model, the big data computing engine, and the data warehouse, respectively. The big data computing engine interacts with the data warehouse, and the large-scale language model interacts with both the big data computing engine and the training set. The main module interacts with the user to schedule commands from the large-scale language model, the big data computing engine, the data warehouse, and the user. The large-scale language model uses neural networks to analyze and understand natural language and schedules the big data computing engine to achieve big data processing capabilities. The training set enables the large-scale language model to recognize patterns and relationships in language and generate coherent and grammatically correct text. The user can customize the training set to improve analysis and capabilities in specific vertical domains. The data warehouse serves as the data foundation for the big data computing engine, providing data sources for the user.
[0010] Secondly, this invention provides a data analysis method based on AIGC, comprising the following steps:
[0011] (1) User initiates data analysis request: Input the corresponding data source and the purpose of data analysis through text description into the main module;
[0012] (2) The main module returns preliminary data exploration: After receiving input from the user, the main module hands it over to the large language model to understand the data source and the purpose of data analysis; during this process, the large language model splits the data source into dimension tables and measure tables; the main module returns the dimension tables and measure tables for the user to confirm;
[0013] (3) User confirms the preliminary data exploration results: The user corrects the preliminary data exploration and analysis results of the large language model. This process is repeated until the user confirms. In this process, the user will confirm the dimension table and measure table after the preliminary data exploration of the main module and perform further data cleaning. The user can directly and manually modify the calculation scope and name of the measure and dimension or add calculation fields. The user can also guide the large language model to regenerate the measure and dimension through text description.
[0014] (4) Return data matrix: After user confirmation, the main module returns one or more data matrices consisting of dimensions and measures;
[0015] (5) User adjusts data matrix: The user obtains multiple data matrices generated by the main module; the user modifies / adds fields / adds filter conditions to these data matrices, and at the same time corrects or restarts the large language model for data analysis;
[0016] (6) After user confirmation, the main module returns: The main module receives the data matrix adjusted and confirmed by the user, schedules the big data computing engine to process the data, obtains the data analysis instance results, and the main module returns one or more data analysis instance results. If the data source is a data warehouse, the big data computing engine interacts with the data warehouse.
[0017] Preferably, the dimension table and metric table in step (2) include the following: (2.1) the identified dimension column and degree column; (2.2) the dimension column that is treated as a number after calculation; and (2.3) the new calculated metric and calculated index generated after identification.
[0018] Preferably, step (4) further includes:
[0019] (4.1) Large-scale language models combine existing data dimension columns and data measure columns to form a data matrix. At this time, the generated data matrix is one or more.
[0020] (4.2) In the generation of data matrix generation rules, when there are fewer than 10 data matrices, the large language model directly combines dimensions and measures by traversing the dimensions; when there are more than 10 data matrices, the large language model selects one or more data matrices after analyzing the user's data analysis purpose, wherein the overlap of dimensions and measures between each data matrix is no higher than 90%.
[0021] (4.3) When the main module outputs the data matrix, it needs to explain the purpose of combining these data matrices and the suggested data chart type, and output data slices to preview the data matrix. Specifically, when generating data slices of the data matrix, the dataset and the corresponding big data engine running instructions for generating slices are first output to the big data engine. The big data engine calculates the difficulty value based on the dataset size / header complexity / data matrix combination difficulty / indicators, and returns the obtained difficulty value to the main module. The main module compares the returned difficulty value with the threshold set by the user. If the difficulty value is greater than the threshold, virtual data slicing is performed. At this time, the large language model randomly generates data slices that conform to the rules according to the current data matrix format. If the difficulty value is less than the threshold, the big data engine is directly called for calculation.
[0022] Preferably, after the data processing is completed in step (6), the main module returns the following information for each data matrix confirmed by the user:
[0023] (6.1) Data slicing;
[0024] (6.2) A complete data table of the data matrix, including: a preview of the data slices, descriptive text and analytical insights for the data charts or data matrix;
[0025] (6.3) The code for generating the data matrix and the adjustment functions on this analysis module;
[0026] (6.4) Visual charts: The GUI page that connects to the BI system is for further modification of the charts;
[0027] (6.5) Code for generating visualization charts of data matrices or visualization adjustments in the GUI;
[0028] (6.6) Data description and insights of data matrices.
[0029] Preferably, in step (6), if the user is not satisfied with the generated result, textual guidance is provided for specific content, including the following steps:
[0030] (1) The user inputs the data source and data analysis purpose, which guides the large language model to generate dimension tables and measure tables;
[0031] (2) After the user generates the dimension table and measure table in the large language model, guide the large language model to explain the further requirements for the subsequent data matrix. If the user thinks that the generated dimension table and measure table do not meet the requirements, the user can also guide the large language model to regenerate the dimension table and measure table.
[0032] (3) When users generate data matrices in large language models, they can guide the large language models to meet the requirements of the generated data analysis report. If users believe that the generated data matrix does not meet their requirements, they can also guide the large language models to regenerate the data matrix.
[0033] (4) When a user generates a data analysis report using a large language model, if the user believes that the generated data analysis report does not meet their needs, the user can also guide the large language model to regenerate the data analysis report.
[0034] The preferred and specific guidance process is as follows:
[0035] (1) The user sends a text containing the data source and data analysis objectives after data preparation and integration; the main module returns the dimension table and measure table after preliminary data exploration.
[0036] (2.1) The user guides the large language model to regenerate the dimension table and measure table by inputting text on the returned data table. This process is repeated until the user confirms. The main module returns the data table generated based on the new descriptive text.
[0037] (2.2) Users can directly and manually adjust the dimensions / measures / custom metrics in the data tables generated by the large language model; the main module returns the adjusted dimension table and measure table; manual adjustment does not conflict with the input text guiding the large language model, and users can manually modify the data table after any guidance modification;
[0038] (2.3) The user confirms the dimension table and the measure table; the main module parses and understands the data analysis purpose through a large language model and returns one or more data matrices. When each data matrix is returned, it will return the following information: the data slice of the data matrix, the filtering conditions of the data matrix, and the perspective of the visualization analysis of the data matrix;
[0039] (3.1) The user guides the large language model to regenerate one or more data matrices by inputting text. This process can be repeated until the user confirms. The main module returns one or more data matrices generated based on the new descriptive text.
[0040] (3.2) Users can directly adjust the dimensions / measures and filtering conditions in the data table generated by the large language model by manually adjusting them; the main module returns one or more data matrices after adjustment; manual adjustment does not conflict with the input text guiding the large language model, and users can manually modify the data matrix after any guidance modification.
[0041] (3.3) The user selects one or more data matrices for further data analysis in the generated data matrix; the main module parses and understands the data analysis purpose through a large language model and returns one or more data analysis instance results. Each data analysis instance returns the following information: visualization chart, complete data file of the data matrix, data slice preview, descriptive text and analysis insights for the data chart or data matrix;
[0042] (4) Users can reintroduce descriptive text and analytical insights into a large language model by inputting text, and this process can be repeated.
[0043] Thirdly, the present invention provides a training method based on an AIGC architecture, comprising the following steps:
[0044] (1) Pre-training is performed by inputting a large language model into a regular training set;
[0045] (2) If the training set includes a dataset in tabular format, the data neural network converts the header data of the tabular data into text format data before performing inference.
[0046] (3) Large language models will infer a probability distribution based on the current input content and the preceding context information to obtain the prediction result;
[0047] (4) Large language models will update their parameters based on the gap between the predicted results and the actual results, continuously optimize the prediction ability of large language models, and complete pre-training.
[0048] (5) Users create custom training sets for vertical domains according to the training set template and input the custom training sets into large language models for training.
[0049] (6) Large language models will infer a probability distribution based on the current input content and the preceding context information to obtain the prediction result;
[0050] (7) Large language models will adjust their parameters based on the gap between the predicted results and the actual results, continuously optimize the prediction ability of large language models, and complete fine-tuning training.
[0051] Preferably, the logic of the training inference result includes the following steps:
[0052] (1) Processing premise: The basic premise for standardizing the response of large language models, and macro-control to ensure the harmlessness of the results generated by large language models;
[0053] (2) Semantic parsing: guiding large language models to reason for data analysis purposes;
[0054] (3) Data analysis objectives and data analysis objective dataset: The large language model combines the column fields of the data analysis objectives and the data analysis objective dataset to infer the user's data analysis purpose and the content that needs to be presented to the user;
[0055] (4) Meta-instructions and Big Data Engine Calling Commands: Meta-instructions are a set of internal instructions used to implement the scheduling and control of various modules of the system by the large language model. Meta-instructions are called by the large language model. Instructions for the big data engine are also written in the meta-instructions. The large language model uses meta-instructions to generate big data engine calling commands to control the big data engine.
[0056] (5) Analysis Results and Analysis Results Dataset: The analysis results returned by the large language model to the user involve calculations and the analysis results dataset; the content returned by the big data engine is called, and there are no restrictions on the form of the generated results, including data streams / data slices / data warehouse indexes / data files / text. The purpose is to provide the basis and calculation results for the analysis operation results and analysis results dataset of the large language model; the analysis results dataset is the analysis results dataset generated by the large language model by parsing the results called by the big data engine.
[0057] Fourthly, the present invention provides an application based on the AIGC architecture, which provides services to users as a standalone data analysis system, or is integrated into a big data platform or a BI system; when the architecture is integrated into the BI system, the architecture uses the PREP data preparation tool of the BI system to clean the data source.
[0058] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0059] (1) This invention uses deep learning technology to understand the natural language text input by users and attempts to infer the user's intentions and needs, thereby improving the efficiency of data analysis. This invention allows users to express their needs more conveniently without needing to master professional terminology, and also reduces the reliance on professional data analysis knowledge, enabling more people to participate in the data analysis process.
[0060] (2) A guided approach is adopted to allow large language models to gradually complete data analysis. This guidance can have a global impact or target only a specific stage, allowing users to better participate in the data analysis process and further improving the accuracy of the data analysis. This invention enables users to participate more actively in the data analysis process, not only improving the accuracy of the data analysis but also better meeting user needs.
[0061] (3) By pre-training a large language model on a training set, the large language model is equipped with the ability to process natural language input from users and to perform semantic parsing, thereby improving the efficiency and accuracy of data analysis. By fine-tuning the large language model, its ability to process customized data in vertical domains and to perform semantic parsing can be improved, enhancing its understanding of data analysis, thus improving its efficiency and accuracy in customized data analysis in vertical domains. Attached Figure Description
[0062] Figure 1 This is an architecture block diagram of the present invention based on AIGC;
[0063] Figure 2 This is a flowchart of the data analysis method based on AIGC of the present invention;
[0064] Figure 3 This is a schematic diagram of the guided process in the AIGC-based data analysis method of this invention;
[0065] Figure 4 This is a logical block diagram of the training and inference results based on AIGC in this invention. Detailed Implementation
[0066] The present invention will now be described in further detail with reference to the accompanying drawings:
[0067] Please see Figure 1 As shown, this invention provides an AIGC-based architecture, including a main module, a large language model, a big data computing engine, a data warehouse, a training set, and a user. The main module interacts with the large language model, the big data computing engine, and the data warehouse; the big data computing engine interacts with the data warehouse; and the large language model interacts with both the big data computing engine and the training set. The main module interacts with the user to schedule commands from the large language model, the big data computing engine, the data warehouse, and the user. The large language model uses neural networks to analyze and understand natural language and schedules the big data computing engine to achieve big data processing capabilities. The training set enables the large language model to recognize patterns and relationships in language, generate coherent and grammatically correct text, and allows users to customize the training set to improve analysis and capabilities in specific vertical domains. The data warehouse serves as the data foundation for the big data computing engine, providing data sources for users.
[0068] It's important to note that AIGC (AI-generated content) is an artificial intelligence technique that can generate new data from given input data. This new data can be in any form, such as images, audio, or text. AIGC typically uses deep learning models, such as Generative Adversarial Networks (GANs) and Variational Autoencoders (VAEs), to generate new data.
[0069] A Large Language Model (LLM) is an artificial intelligence model that uses neural networks to analyze and understand natural language. Specifically, a LLM aims to generate human-like text. Here, the role of the LLM is to analyze the user's input data source and the corresponding data analysis objective, and provide appropriate feedback.
[0070] Training set: Users use the training set for LLM training. The principle is to train the large language model by inputting the training set. Through training, the large language model can recognize patterns and relationships in the language. Then, these contents can be used to generate coherent and grammatically correct new text, thereby improving the customized analysis perspective and capabilities in vertical domains.
[0071] This invention provides an application based on an AIGC architecture, which can serve as a standalone data analysis system or be integrated into a big data platform or BI system. For example, when this architecture is embedded in a BI system, it utilizes the BI system's PREP data preparation tool to clean the data source.
[0072] A data warehouse is a centralized storage system that aggregates various data sources within an organization. It is designed to support decision-making and business analysis within an enterprise or organization. The design of a data warehouse aims to provide a consistent, integrated, and easily accessible data storage system, enabling businesses or organizations to quickly and accurately access and analyze data to make better decisions. A data warehouse typically includes multiple data sources, such as transaction systems, human resource management systems, and marketing systems. These data sources are cleaned, transformed, and integrated, and then stored in a unified location for analysis and querying.
[0073] A training set is the dataset used to train an artificial intelligence model. In the training set, data is labeled and classified so that the model can learn and recognize various types of data. The training set is a crucial component of an artificial intelligence model and has a significant impact on its performance and accuracy.
[0074] like Figure 2 As shown, this invention provides a data analysis method based on AIGC, comprising the following steps:
[0075] (1) User initiates data analysis request: Input the corresponding data source and the purpose of data analysis through text description into the main module;
[0076] (2) The main module returns preliminary data exploration: After receiving input from the user, the main module hands it over to the large language model to understand the data source and the purpose of data analysis; during this process, the large language model splits the data source into dimension tables and measure tables; the main module returns the dimension tables and measure tables for the user to confirm;
[0077] (3) User confirms the preliminary data exploration results: The user corrects the preliminary data exploration and analysis results of the large language model. This process is repeated until the user confirms. In this process, the user will confirm the dimension table and measure table after the preliminary data exploration of the main module and perform further data cleaning. The user can directly and manually modify the calculation scope and name of the measure and dimension or add calculation fields. The user can also guide the large language model to regenerate the measure and dimension through text description.
[0078] (4) Return the data matrix. After user confirmation, the main module returns one or more data matrices consisting of dimensions and measures.
[0079] (5) User adjusts data matrix: The user obtains multiple data matrices generated by the main module; the user modifies / adds fields / adds filter conditions to these data matrices, and at the same time corrects or restarts the large language model for data analysis; the main module receives the data matrices adjusted and confirmed by the user, and schedules the big data computing engine to process the data. If the source of the data is a data warehouse, the big data computing engine interacts with the data warehouse.
[0080] (6) After user confirmation, the main module returns.
[0081] It should be noted that the relationship between metrics, dimensions, and data matrices is as follows:
[0082] A measure column is a column that contains numerical data that can be used for calculations and statistical analysis. Common measure columns include quantity, amount, time, and ratio.
[0083] Dimension columns are columns that contain non-numerical data used to describe and categorize data. Common dimension columns include region, time, product, and customer. Dimension columns are typically used for grouping, filtering, and summarizing data.
[0084] A data matrix is a two-dimensional table consisting of rows and columns. In this description, a data matrix is a two-dimensional table composed of measures and dimensions.
[0085] For step (1), the specific details are as follows:
[0086] Users need to input the corresponding data source and a textual description of the data analysis objective into the main module. There are no restrictions on the source of the data source (it can be a data warehouse, an external data source, or a local file), or it can be a cleaned data table prepared from the data of an embedded BI system. This data table must be a two-dimensional table.
[0087] Data Source Description: In a typical workflow, users need to send data source information and a text description of the data analysis objective to the main module at least once. There are no restrictions on the data source choice; it can be a database query or a local file. If the system is integrated into a BI system, the BI system's data PREP (prepare) tool can be used directly to prepare the data source. Data PREP refers to the process of preparing data, which typically includes data cleaning, transformation, normalization, scaling, feature selection, and feature extraction to facilitate subsequent data analysis and machine learning tasks. After preparing the data source, the main module will hand it over to the big data computing engine, which will then store the prepared data source in a data warehouse for later use.
[0088] For example: The table below is a sales data table for an order. The user will input this data table into the main module.
[0089]
[0090]
[0091] For step (2), the dimension table and the metric table include the following: (2.1) the identified dimension and degree columns; (2.2) the dimension columns that are treated as degrees after calculation; and (2.3) the new calculated metrics and calculated indicators generated after identification.
[0092] Preliminary Data Exploration – Identifying Dimensions and Measures: Large language models identify and determine the data source of the input to record whether the current data belongs to a dimension or a measure. For example, if we take "Country / Region / Payment Method / Order Number / Ordering User_id / Order Creation Date / User Gender" from a sales data table as dimensions and "GMV / NMV" as measures, we get the following dimension and measure tables.
[0093] Dimension table:
[0094] Column names
[0095] nation
[0096] area
[0097] User gender
[0098] Payment methods
[0099] Order number
[0100] User ID of the order
[0101] Order creation date (day)
[0102] Measurement table:
[0103] Column names
[0104] GMV
[0105] NMV
[0106] Preliminary Data Exploration – Generating Columns Based on Data Analysis Objectives: After segmenting measures and dimensions, the large language model will process the current dimensions and measures in conjunction with the data analysis objectives. This processing includes, but is not limited to:
[0107] Remove overly detailed dimension columns based on the data analysis objectives, and convert detailed dimensions into measures. Example: To convert order numbers to order counts, delete the dimension column "Order Number" and add a new measure column "Order Count".
[0108] Generate custom metrics based on the data analysis objectives. Example: Combine the metrics [GMV] and [NMV] into a single metric [Order Collection Rate] (NMV / GMV).
[0109] The following dimension and measure tables were obtained:
[0110] Dimension table:
[0111] Column names
[0112] nation
[0113] area
[0114] User gender
[0115] Payment methods
[0116] Order creation date (day)
[0117] Measurement table:
[0118] Column names
[0119] GMV
[0120] NMV
[0121] Number of orders
[0122] Number of users placing orders
[0123] Order collection rate
[0124] Generate custom dimension metrics based on the data analysis objective. Example: Suppose the current data analysis objective is to analyze high-value and low-value users, then the large language model will generate the dimension column "User Level" based on the data analysis objective.
[0125] It should be noted that when large language models perform automated data processing, they are affected by the guiding text input by the user and the training set.
[0126] The main module will provide feedback to the user on the identified and generated metrics and dimensions.
[0127] For step (3), in this process, the user will confirm the dimension table and measure table after the initial data exploration of the main module. During this process, the user can further clean the data. It should be noted that during this process, the user can directly and manually modify the calculation scope and name of the measure and dimension or add new calculated fields, or guide the LLM to regenerate the measure and dimension through text description. For example, for dates, it can be requested to generate dimension columns with the order creation date (year) as the period.
[0128] For step (4), the specific details are as follows:
[0129] (4.1) In this process, the large language model combines the existing data dimension columns and data measure columns to form a data matrix. The generated data matrix may be one or more.
[0130] (4.2) In the generation of data matrix generation rules, when there are fewer than 10 data matrices, the large language model directly combines dimensions and measures by traversing the dimensions; when there are more than 10 data matrices, the large language model selects one or more data matrices after analyzing the user's data analysis purpose, wherein the overlap of dimensions and measures between each data matrix is no higher than 90%.
[0131] (4.3) When the main module outputs the data matrix, it needs to explain the purpose of combining these data matrices and the suggested data chart type, and output data slices to preview the data matrix. Specifically, when generating data slices of the data matrix, the dataset and the corresponding big data engine running instructions for generating slices are first output to the big data engine. The big data engine calculates the difficulty value based on the dataset size / header complexity / data matrix combination difficulty / indicators, and returns the obtained difficulty value to the main module. The main module compares the returned difficulty value with the threshold set by the user. If the difficulty value is greater than the threshold, virtual data slicing is performed. At this time, the large language model randomly generates data slices that conform to the rules according to the current data matrix format. If the difficulty value is less than the threshold, the big data engine is directly called for calculation.
[0132] It's important to note that data slicing is a frequently used technique in data analysis. It involves dividing large amounts of data according to specific criteria to facilitate better analysis and processing. Data slicing targets a specific time frame, region, product, or service, allowing for a better understanding of the data's inherent patterns and trends. Common data slicing techniques include time-based slicing, location-based slicing, and user behavior-based slicing. Different slicing methods help us analyze data from different perspectives and uncover its value. When performing data slicing, it's necessary to first determine the slicing objectives and conditions, and then filter and slice the data accordingly. The selection of slicing conditions should be based on actual needs while ensuring the accuracy and completeness of the data.
[0133] Example illustration:
[0134] The data matrix is arranged according to "country-payment method" because this combination helps us understand the geographic and payment method preferences of customer purchasing behavior. We can use this combination to answer the following questions:
[0135] -Which countries' users are more inclined to pay on delivery or online?
[0136] - In which countries are GMV and NMV the highest?
[0137] Which payment methods are most frequently used by users in different countries?
[0138] These questions help us better understand customer buying behavior and market demand, thereby enabling us to develop more effective marketing strategies and product strategies.
[0139] The following table shows the corresponding data slices:
[0140]
[0141]
[0142] For step (5), the user obtains multiple data matrices generated by the main module; the user modifies / adds fields / adds filter conditions to these data matrices, and at the same time corrects or re-guides the large language model for data analysis.
[0143] For step (6), during this process, the main module will receive the data matrix adjusted and confirmed by the user. At this time, the main module will schedule the big data engine to process the data and obtain the data analysis instance results. The main module will return one or more data analysis instance results. If the source of the data source is a data warehouse, the big data engine will also interact with the data warehouse during this process.
[0144] After the data is processed, the main module will return the following information to the user for each data matrix that has been confirmed by the user:
[0145] (6.1) Visualization charts of the data matrix;
[0146] (6.2) The code for generating visualization charts of the data matrix or the visualization adjustment of the GUI page of the BI system are for the purpose of further modifying the charts;
[0147] (6.3) The complete data table of the data matrix;
[0148] (6.4) The code for generating the data matrix and the adjustment functions on this analysis module;
[0149] (6.5) Data slice preview;
[0150] (6.6) Descriptive text and analytical insights for data charts or data matrices (this process can be generated iteratively to refine the content);
[0151] It should be noted that the code for generating the data charts and matrices returned here does not impose any restrictions on language or adjustment methods. This is simply to emphasize that users can further adjust the generated content.
[0152] Furthermore, there are no restrictions on the user's output results; the output content is generated through corresponding interfaces / apis to customize the content.
[0153] like Figure 3 As shown, if the user is not satisfied with the results returned by the main module, textual guidance will be provided for specific content, including the following steps:
[0154] (1) The user inputs the data source and data analysis purpose, which guides the large language model to generate dimension tables and measure tables;
[0155] (2) After the user generates the dimension table and measure table in the large language model, guide the large language model to explain the further requirements for the subsequent data matrix. If the user thinks that the generated dimension table and measure table do not meet the requirements, the user can also guide the large language model to regenerate the dimension table and measure table.
[0156] (3) When users generate data matrices in large language models, they can guide the large language models to meet the requirements of the generated data analysis report. If users believe that the generated data matrix does not meet their requirements, they can also guide the large language models to regenerate the data matrix.
[0157] (4) When a user generates a data analysis report using a large language model, if the user believes that the generated data analysis report does not meet their needs, the user can also guide the large language model to regenerate the data analysis report.
[0158] To further improve the inference capabilities of LLM and better align them with users' actual data analysis objectives, the LLM data analysis process is divided into three parts: (1) splitting the dataset into measure tables and dimension tables; (2) assembling the dimension tables and measure tables to generate a data matrix; and (3) generating a data analysis report.
[0159] Users can guide the LLM process at the beginning and end of each step.
[0160] Guidance refers to the ability of users to modify and refine the LLM reasoning for their data analysis purposes at each stage of the LLM job. Guidance can be targeted at a specific stage or have a global impact.
[0161] It's important to note that each user input guides the LLM process, causing it to continuously adjust its output based on the user's text. Only the initial input, intended for data analysis, is mandatory. Without further guidance, the LLM will automatically parse the semantics of the first input to generate the remaining data matrix and report. Whether the user guides the LLM to regenerate data or to address new requirements for the next stage is determined by the LLM's reasoning based on the user's description; there are no specific formatting requirements.
[0162] The table below shows the specific guidance process.
[0163]
[0164]
[0165] This invention provides a training architecture based on AIGC, characterized by the following steps:
[0166] (1) Pre-training is performed by inputting a large language model into a regular training set;
[0167] (2) If the training set includes a dataset in tabular format, the data neural network converts the header data of the tabular data into text format data before performing inference.
[0168] (3) Large language models will infer a probability distribution based on the current input content and the preceding context information to obtain the prediction result;
[0169] (4) Large language models will update their parameters based on the gap between the predicted results and the actual results, continuously optimize the prediction ability of large language models, and complete pre-training.
[0170] (5) Users create custom training sets for vertical domains according to the training set template and input the custom training sets into large language models for training.
[0171] (6) Large language models will infer a probability distribution based on the current input content and the preceding context information to obtain the prediction result;
[0172] (7) Large language models will adjust their parameters based on the gap between the predicted results and the actual results, continuously optimize the prediction ability of large language models, and complete fine-tuning training.
[0173] Large-scale language models use deep learning techniques in language processing to understand natural language text input by users and attempt to infer the user's intent and needs. Specifically, through pre-training and fine-tuning, large-scale language models extract features such as language habits, sentence structure, and contextual information from data, and then use these features to generate responses relevant to user input. It's important to note that large-scale language models cannot directly understand user intent; instead, they infer responses through prediction and simulation. Therefore, the quality and accuracy of a large-scale language model's response are related to the quantity of training data in the training set and the content of the user input.
[0174] The principle behind training a large-scale language model is based on a neural network. A neural network can learn the rules and patterns of language by inputting large amounts of text data, thereby constructing a large-scale language model capable of predicting the next sentence. In this way, the large-scale language model can better infer the semantics and context of text data, thus predicting the next sentence more accurately. The reason why the large-scale language model in this invention can understand the dataset, i.e., tabular data, is that the neural network converts the header data of the tabular data into text format data before inference during training. During training, the neural network continuously adjusts the parameters of the large-scale language model, enabling it to predict the next text more accurately. Specifically, the large-scale language model infers and outputs a probability distribution based on the current input content and preceding context information, indicating what the next possible text might be. Furthermore, the large-scale language model updates its parameters based on the difference between the predicted and actual results, continuously optimizing its predictive ability. Therefore, if the input training set is of sufficiently high quality and has a more complex structure, the inference results of the large-scale language model will be more in line with expectations.
[0175] Large language models continuously learn language rules and patterns during training, enabling them to better capture language patterns and useful text fragments describing data analysis objectives within text data. This allows for more accurate prediction of necessary operations such as indicator decomposition, data matrix generation, and data analysis report generation. Consequently, large language models can better meet user needs and generate results that align with language habits. For example, when users input specific data analysis terminology or common data analysis objectives, large language models can better predict the required output, generating results that conform to data analysis and language patterns. In this invention, the LLM data analysis process is broken down into three steps: splitting the dataset into measure and dimension tables, assembling dimension and measure tables to generate a data matrix, and generating a data analysis report. The purpose is to strengthen the reasoning patterns and language patterns within each step.
[0176] During pre-training, the model is trained on two tasks: Masked Language Modeling (MLM) and Next Sentence Prediction (NSP). In the MLM task, the model randomly masks some words in the input text and then attempts to predict the original values of these words. In the NSP task, the model receives two sentences as input and then attempts to predict whether the two sentences are consecutive. During pre-training, the model gradually improves its language understanding and language generation capabilities through multiple rounds of iterative learning. The pre-trained model is then fine-tuned on various tasks to adapt to different application scenarios. It is important to note that the quality of the pre-training task and data will affect the model's predictive ability; therefore, high-quality data and tasks must be selected for pre-training. Simultaneously, the pre-trained model also requires post-training optimization to achieve better results. Fine-tuning refers to applying the pre-trained model to specific tasks, adjusting the model parameters through supervised learning on labeled data to better adapt it to the task.
[0177] The large-scale language model described above has already demonstrated that the reasoning and processing logic of LLM for data analysis purposes is related to the content of the training set used by LLM. In addition to natural language recognition, the training content also needs to enable LLM to call big data engines and reason and calculate indicators and data analysis directions based on data tables and purposes. When LLM is trained, the training set used exists in the form of QA dialogues, and the training set is a collection of a large number of training examples.
[0178] It's important to note that Natural Language Processing (NLP) is a branch of artificial intelligence that focuses on enabling interaction between computers and humans through natural language. It involves teaching machines the ability to understand, interpret, and generate human language. NLP has many practical applications, including chatbots, language translation, and sentiment analysis.
[0179] In this invention, the data structure for training examples consists of the following parameters: 1. Prerequisites for processing: This instruction serves as a fundamental premise for standardizing responses from large models, and its purpose is to render the generated results harmless.
[0180] 2. Meta-instructions: The meta-instructions here are a set of internal instructions primarily used to implement the scheduling and control of various modules in the system by the LLM model. These meta-instructions can be called by the LLM model. It should be noted that instructions specific to the big data engine are also included in the meta-instructions.
[0185] 3. Rounds:
[0186] The response rounds in a single training instance should ideally be three; this corresponds to the three-step data analysis process for this architecture: splitting dimension tables and measure tables → combining data matrices → writing data analysis methods.
[0187] 4. Purpose of data analysis:
[0188] Ideally, each round should have a corresponding data analysis objective. However, except for the first round where the data analysis objective is mandatory, the data analysis objective for each subsequent round is optional.
[0189] 5. Target dataset for data analysis:
[0190] Each round requires a corresponding dataset to be used as the original dataset.
[0191] 6. Semantic parsing: Guiding LLM's reasoning regarding data analysis objectives. For example:
[0192] Question: Can you explain to me what profit margin is and how it is calculated?
[0193] Semantic analysis: The user wants to know what profit margin is and how it is calculated. I need to provide a brief explanation of the concept of profit margin and an example of how to calculate it.
[0194] 7. Big Data Engine Commands:
[0195] LLM obtains the commands to call the big data engine through semantic parsing. The reason for calling the big data engine is that performing calculations on the target dataset by LLM would be too costly and inefficient. However, not all rounds require calling the big data engine.
[0196] 8. Big Data Engine Call Results:
[0197] The content returned by the big data engine is not limited in terms of the format of the generated results, including but not limited to: data streams / data slices / data warehouse indexes / data files / text. Its main function is to provide the basis and calculation results for the LLM's [analysis job results] and [analysis result datasets].
[0198] 9. Analysis Results:
[0199] The ideal answer given for each round of data inference; this analysis result is the result that must be returned in each round.
[0200] 10. Analysis Results Dataset:
[0201] This is the dataset of analysis results generated by LLM through parsing the results of the big data engine call. It should be noted that in one round, an analysis result dataset may return multiple datasets, or it may not return any data analysis results. The number of datasets returned is related to the results of semantic parsing.
[0202] like Figure 4 As shown, the following is a logical explanation of the reasoning results of the large-scale language model training method.
[0203] The reason this training set format is needed is that after LLM obtains the user's input data analysis objective and target dataset, it requires internal interaction to implement the analysis process. First, the **processing prerequisites** provide macro-level control to ensure the harmlessness of the LLM-generated results. LLM combines the column fields of the **data analysis objective** and the **target dataset** to infer the user's data analysis purpose and the content to be presented to the user, using **meta-instructions** to generate **big data engine call commands** to control the big data engine. This process returns the computationally involved content and the **analysis result dataset** presented to the user in the **analysis results**.
[0204] The above description is only a preferred embodiment of the present invention. All equivalent changes and modifications made within the scope of the claims of the present invention should be covered by the claims of the present invention.
Claims
1. An AIGC-based architecture, characterized in that, The system includes a main module, a large-scale language model, a big data computing engine, a data warehouse, a training set, and a user. The main module interacts with the large-scale language model, the big data computing engine, and the data warehouse. The big data computing engine interacts with the data warehouse. The large-scale language model interacts with both the big data computing engine and the training set. The main module interacts with the user to schedule commands from the large-scale language model, the big data computing engine, the data warehouse, and the user. The large-scale language model uses neural networks to analyze and understand natural language and schedules the big data computing engine to perform big data processing. The training set enables the large-scale language model to recognize patterns and relationships in language and generate coherent and grammatically correct text. The user can customize the training set to improve analysis and capabilities in specific vertical domains. The data warehouse serves as the data foundation for the big data computing engine and provides data sources for the user.
2. A data analysis method based on AIGC, characterized in that, Includes the following steps: (1) User initiates data analysis request: Input the corresponding data source and the purpose of data analysis through text description into the main module; (2) The main module returns preliminary data exploration: After receiving input from the user, the main module hands it over to the large language model to understand the data source and the purpose of data analysis; in this process, the large language model splits the data source into dimension tables and measure tables; The main module returns dimension and measure tables for user confirmation. (3) User confirms preliminary data exploration results: The user corrects the preliminary data exploration and analysis results of the large language model, and repeats this process until the user confirms; In this process, the user will confirm the dimension table and measure table after the preliminary data exploration of the main module, and further clean the data. The user can directly and manually modify the calculation scope and name of the measure and dimension or add calculation fields, or guide the large language model to regenerate the measure and dimension through text description; (4) Return data matrix: After user confirmation, the main module returns one or more data matrices consisting of dimensions and measures; (5) User adjusts data matrix: The user obtains multiple data matrices generated by the main module; the user modifies / adds fields / adds filter conditions to these data matrices, and at the same time corrects or restarts the large language model for data analysis; (6) After user confirmation, the main module returns: The main module receives the data matrix adjusted and confirmed by the user, schedules the big data computing engine to process the data, obtains the data analysis instance results, and the main module returns one or more data analysis instance results. If the data source is a data warehouse, the big data computing engine interacts with the data warehouse.
3. The data analysis method based on AIGC according to claim 2, characterized in that, The dimension table and measure table in step (2) include the following: (2.1) Identification of the dimension and degree sequence; (2.2) The calculated degree is regarded as a number of dimensions; (2.3) The new calculation metric and calculation index are generated after identification.
4. The data analysis method based on AIGC according to claim 2, characterized in that, Step (4) further includes: (4.1) Large-scale language models combine existing data dimension columns and data measure columns to form a data matrix. At this time, the generated data matrix is one or more. (4.2) In the generation of data matrix generation rules, when there are fewer than 10 data matrices, the large language model directly combines dimensions and measures by traversing the dimensions; when there are more than 10 data matrices, the large language model selects one or more data matrices after analyzing the user's data analysis purpose, wherein the overlap of dimensions and measures between each data matrix is no higher than 90%. (4.3) When the main module outputs the data matrix, it needs to explain the purpose of combining these data matrices and the suggested data chart type, and output data slices to preview the data matrix. Specifically, when generating data slices of the data matrix, the dataset and the corresponding big data engine running instructions for generating slices are first output to the big data engine. The big data engine calculates the difficulty value based on the dataset size / header complexity / data matrix combination difficulty / indicators, and returns the obtained difficulty value to the main module. The main module compares the returned difficulty value with the threshold set by the user. If the difficulty value is greater than the threshold, virtual data slicing is performed. At this time, the large language model randomly generates data slices that conform to the rules according to the current data matrix format. If the difficulty value is less than the threshold, the big data engine is directly called for calculation.
5. The data analysis method based on AIGC according to claim 2, characterized in that, After the data processing is completed in step (6), the main module returns the following information to the user for each data matrix that has been confirmed by the user: (6.1) Data slicing; (6.2) A complete data table of the data matrix, including: a preview of the data slices, descriptive text and analytical insights for the data charts or data matrix; (6.3) The code for generating the data matrix and the adjustment functions on this analysis module; (6.4) Visual charts: The GUI page that connects to the BI system is for further modification of the charts; (6.5) Code for generating visualization charts of data matrices or visualization adjustments in the GUI; (6.6) Data description and insights of data matrices.
6. The data analysis method based on AIGC according to claim 2, characterized in that, In step (6), if the user is not satisfied with the generated result, textual guidance is provided for specific content, including the following steps: (1) The user inputs the data source and data analysis purpose, which guides the large language model to generate dimension tables and measure tables; (2) After the user generates the dimension table and measure table in the large language model, guide the large language model to explain the further requirements for the subsequent data matrix. If the user thinks that the generated dimension table and measure table do not meet the requirements, the user can also guide the large language model to regenerate the dimension table and measure table. (3) When users generate data matrices in large language models, they can guide the large language models to meet the requirements of the generated data analysis report. If users believe that the generated data matrix does not meet their requirements, they can also guide the large language models to regenerate the data matrix. (4) When a user generates a data analysis report using a large language model, if the user believes that the generated data analysis report does not meet their needs, the user can also guide the large language model to regenerate the data analysis report.
7. The data analysis method based on AIGC according to claim 6, characterized in that, The specific guidance process is as follows: (1) The user sends a text containing the data source and data analysis objectives after data preparation and integration; the main module returns the dimension table and measure table after preliminary data exploration. (2.1) The user guides the large language model to regenerate the dimension table and measure table by inputting text on the returned data table. This process is repeated until the user confirms. The main module returns the data table generated based on the new descriptive text. (2.2) Users can directly and manually adjust the dimensions / measures / custom metrics in the data tables generated by the large language model; the main module returns the adjusted dimension table and measure table; manual adjustment does not conflict with the input text guiding the large language model, and users can manually modify the data table after any guidance modification; (2.3) The user confirms the dimension table and the measure table; the main module parses and understands the data analysis purpose through a large language model and returns one or more data matrices. When each data matrix is returned, it will return the following information: the data slice of the data matrix, the filtering conditions of the data matrix, and the perspective of the visualization analysis of the data matrix; (3.1) The user guides the large language model to regenerate one or more data matrices by inputting text. This process can be repeated until the user confirms. The main module returns one or more data matrices generated based on the new descriptive text. (3.2) Users can directly adjust the dimensions / measures and filtering conditions in the data table generated by the large language model by manually adjusting them; the main module returns one or more data matrices after adjustment; manual adjustment does not conflict with the input text guiding the large language model, and users can manually modify the data matrix after any guidance modification. (3.3) The user selects one or more data matrices for further data analysis in the generated data matrix; the main module parses and understands the data analysis purpose through a large language model and returns one or more data analysis instance results. Each data analysis instance returns the following information: visualization chart, complete data file of the data matrix, data slice preview, descriptive text and analysis insights for the data chart or data matrix; (4) Users can reintroduce descriptive text and analytical insights into a large language model by inputting text, and this process can be repeated.
8. A training method based on an AIGC architecture, characterized in that, Includes the following steps: (1) Pre-training is performed by inputting a large language model into a regular training set; (2) If the training set includes a dataset in tabular format, the data neural network converts the header data of the tabular data into text format data before performing inference. (3) Large language models will infer a probability distribution based on the current input content and the preceding context information to obtain the prediction result; (4) Large language models will update their parameters based on the gap between the predicted results and the actual results, continuously optimize the prediction ability of large language models, and complete pre-training. (5) Users create custom training sets for vertical domains according to the training set template and input the custom training sets into large language models for training. (6) Large language models will infer a probability distribution based on the current input content and the preceding context information to obtain the prediction result; (7) Large language models will adjust their parameters based on the gap between the predicted results and the actual results, continuously optimize the prediction ability of large language models, and complete fine-tuning training.
9. The training method based on the AIGC architecture according to claim 8, characterized in that, The logic of the training inference results includes the following steps: (1) Processing premise: The basic premise for standardizing the response of large language models, and macro-control to ensure the harmlessness of the results generated by large language models; (2) Semantic parsing: guiding large language models to reason for data analysis purposes; (3) Data analysis objectives and data analysis objective dataset: The large language model combines the column fields of the data analysis objectives and the data analysis objective dataset to infer the user's data analysis purpose and the content that needs to be presented to the user; (4) Meta-instructions and Big Data Engine Calling Commands: Meta-instructions are a set of internal instructions used to implement the scheduling and control of various modules of the system by the large language model. Meta-instructions are called by the large language model. Instructions for the big data engine are also written in the meta-instructions. The large language model uses meta-instructions to generate big data engine calling commands to control the big data engine. (5) Analysis Results and Analysis Results Dataset: The analysis results returned by the large language model to the user involve calculations and the analysis results dataset; the content returned by the big data engine is called, and there are no restrictions on the form of the generated results, including data streams / data slices / data warehouse indexes / data files / text. The purpose is to provide the basis and calculation results for the analysis operation results and analysis results dataset of the large language model; the analysis results dataset is the analysis results dataset generated by the large language model by parsing the results called by the big data engine.
10. An application based on an AIGC architecture, characterized in that, The architecture described in claim 1 can be provided to users as a standalone data analysis system, or integrated into a big data platform, or integrated into a BI system; when the architecture is integrated into the BI system, the architecture uses the PREP data preparation tool of the BI system to clean the data source.