Intelligent data analysis method and system based on large language model
Patent Information
- Application Number
- CN202610737145.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-27
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2046-05-27
AI Technical Summary
[0008]本发明实施例提供了一种基于大语言模型的智能数据分析方法及系统,以解决现有技术中的上述技术的问题
1)本发明通过基于LLM Agent的智能任务分解与自动化数据分析机制,根据LLMAgent能够根据用户自然语言需求自动推断分析意图,完成数据检测、智能数据清洗、特征构建、统计建模与结果验证等一体化任务分解,并自主生成和执行分析代码。系统通过多轮推理与计划重构机制,自主发现问题并调整分析路径,使非专业用户无需具备编程技能也能完成完整的数据分析流程,从而显著提升任务完成效率与分析准确性。
Smart Images

Figure CN122285733B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of artificial intelligence and data analysis technology, and in particular to an intelligent data analysis method and system based on a large language model. Background Technology
[0002] Currently, data analysis and visualization tools such as Tableau (a visualization analytics platform), Power BI (Microsoft Business Intelligence), SPSS (Statistical Package for Social Sciences), and Stata (Statistical Data Analysis Software) are widely used in scientific research and business decision-making. They can import data, perform statistical modeling, visualize analysis, and generate basic charts and reports. They support multi-table and multi-sheet data processing and can achieve semi-automated analysis through macros or scripts, thus meeting basic data processing and display needs.
[0003] However, existing tools have low levels of intelligence and automation, heavily relying on users' experience in statistics, programming, and data analysis. Users must manually complete data cleaning, model building, and visualization, resulting in high barriers to entry and low efficiency. Ordinary users often struggle to choose appropriate data processing methods, statistical models, and chart formats, easily leading to non-standard analysis processes and insufficient reliability of results. Furthermore, traditional platforms cannot automatically plan analysis paths based on research objectives, lack intelligent interpretation and theoretical insight capabilities, and are ill-suited to the current demands for intelligent, low-barrier, and highly interpretable data in research involving larger data volumes, more complex analysis scenarios, and interdisciplinary studies. Achieving an end-to-end intelligent data analysis platform presents several technical challenges: Firstly, intelligent data cleaning and code generation: Social survey data often suffers from various problems such as missing values, logical errors, duplicate responses, and incomplete reverse coding, and frequently spans multiple data tables. The system needs to automatically generate cleaning strategies based on variable semantics and measurement levels, and then generate executable code through a large language model. After review, the code is automatically executed to ensure scientific data quality. Secondly, how to maintain strategy compatibility across different data formats (such as Excel and Stata) while ensuring the reproducibility and consistency of the cleaning results?
[0004] Secondly, natural language understanding and task planning: The system must be able to understand the user's natural language analysis needs and map them into specific data processing and analysis tasks. This involves not only text semantic parsing but also combining data structure characteristics and the context of social science research to achieve multi-level task decomposition and dynamic execution planning. How to enable the system to automatically recommend suitable data cleaning strategies and analysis methods without any instructions is the core challenge in building a fully intelligent platform.
[0005] Third, automated analysis and theoretical insight generation: The system needs to automatically perform statistical analysis based on data characteristics and research objectives, including descriptive statistics, regression analysis, group difference testing, text and sentiment analysis, etc. Simultaneously, the system must integrate professional social science theories to interpret results and generate insights, transforming statistical results into a high-level professional understanding of social science.
[0006] Fourth, interactive visualization and end-to-end closed-loop integration: To meet users' needs for rapid understanding and decision-making, the platform must dynamically present the analysis results as distribution maps, relationship diagrams, and interactive visualizations, and generate structured reports to integrate visualization with theoretical insights. Simultaneously, it must ensure an end-to-end closed-loop process from data import, cleaning, analysis to report generation, completing tasks with minimal human intervention. This requires a high degree of coordination between visualization, code generation, task scheduling, and result verification.
[0007] Therefore, how to provide an intelligent data analysis method and system based on a large language model is an urgent problem to be solved. Summary of the Invention
[0008] This invention provides an intelligent data analysis method and system based on a large language model to solve the problems mentioned above in the prior art.
[0009] According to a first aspect of the present invention, an intelligent data analysis method based on a large language model is provided.
[0010] In one embodiment, the intelligent data analysis method based on a large language model includes: Obtain the data structure features and association rules, and perform variable semantic parsing on the data structure features and association rules to obtain the data catalog and variable description document; Based on the statistical rules in the large language model, and combined with the data catalog, variable description documents, and semantic analysis to detect data quality problems, a problem diagnosis is performed on the complex anomaly reconstruction analysis process to obtain a data quality risk report. Based on the data quality risk report, a personalized data cleaning strategy is constructed. The personalized data cleaning strategy is used to reconstruct and repair the target abnormal data. The personalized data cleaning strategy is converted into executable code using a large language model. After logical verification and security testing of the executable code, the reconstructed and repaired target abnormal data is cleaned to obtain the cleaned data. Extract multi-attribute feature vectors from the cleaned data and generate a feature vector distribution map from the multi-attribute feature vectors. Based on the feature vector distribution map, the planning and analysis task process is carried out. The statistical analysis scheme is matched and dynamically adjusted based on the analysis task process. The adjusted statistical analysis scheme is converted into analysis code and executed to obtain automated statistical analysis results. Based on the results of automated statistical analysis, interactive visualization charts are generated. Logical deduction and causal explanation are then performed based on the interactive visualization charts and social science theories. The deduction results and causal explanations are integrated to obtain an analytical report with theoretical insights, thereby achieving intelligent control of the entire process of intelligent data analysis by a large language model.
[0011] In one embodiment, data structure features and association rules are obtained, and variable semantic parsing is performed on the data structure features and association rules to obtain a data catalog and variable description document, including: Extracting data structure features and association rules from various structured data formats in spreadsheets and statistical software; By performing variable semantic parsing on the structured features and association rules of the data, the variable types, text encoding, scale attributes and multi-table associations are automatically identified to obtain preliminary semantic parsing results. Based on the preliminary semantic parsing results, the data of multiple worksheets is parsed to eliminate differences in fields and structural ambiguities between tables, and standardized semantic parsing results are obtained. Based on the standardized parsing results, a large language model is used to extract variable semantics, determine the logical relationships between variables, and measure the hierarchy, resulting in a data catalog and variable description documents.
[0012] In one embodiment, based on statistical rules in the large language model, and combined with data catalogs, variable description documents, and semantic analysis to detect data quality issues, a problem diagnosis is performed on the complex anomaly reconstruction analysis process, resulting in a data quality risk report including: The large language model is used as the control terminal of the analysis process, and based on the target analysis node, the current data status, execution operation and decision basis are recorded to form an analysis path status chain. Based on the analysis path state chain, the current data state is used to determine the statistical rule constraints, and semantic analysis is performed by combining the semantic description of variables, the meaning of rules, and the analysis objectives. Based on statistical rule constraints and semantic analysis, data quality detection is performed on multiple worksheets to identify missing value patterns, logical contradictions, abnormal data, unreversed reverse options, duplicate answers, and extreme responses, and to obtain preliminary data quality diagnosis results. Based on complex anomalies identified by a single statistical rule, the analysis process is reconstructed by reconstructing variable states, historical processing paths, and related variable information to obtain the diagnostic results after analysis. The preliminary data quality diagnosis results are merged with the post-analysis diagnosis results. Based on the merged results, a quality risk report is generated and the problems and their severity are marked, thus obtaining the data quality risk report.
[0013] In one embodiment, a personalized data cleaning strategy is constructed based on a data quality risk report. This personalized strategy is then used to reconstruct and repair the target anomalous data. A large language model is used to convert the personalized data cleaning strategy into executable code. After logical verification and security checks, the executable code is executed to clean the reconstructed and repaired target anomalous data, resulting in cleaned data including: For data quality risk reports, identify target abnormal data that cannot be repaired through the established cleaning process; Using a large language model as the process control terminal, reverse reasoning is performed on the target abnormal data and the variable constraints are met to obtain the reconstructed data repair path. Based on the data repair path, variable semantics, measurement level, and data characteristics, a personalized data cleaning strategy is constructed. The personalized data cleaning strategy is converted into executable code using a large language model, and the built-in code review mechanism in the large language model is used again to perform logical verification and security checks on the executable code to obtain the approved executable code. The executable code that passes the verification is cleaned, and the cleaned results are verified a second time to obtain the cleaned data.
[0014] In one embodiment, the built-in code review mechanism includes: The built-in code review mechanism is used to perform logic, syntax and security checks on the generated executable code to ensure that the code conforms to the preset analysis path and analysis objectives. The code review mechanism uses the large language model as a process control node. After generating executable code, it determines whether the code conforms to the preset analysis path based on the current analysis target and process status. If code logic deviations, security risks, or syntax errors are detected, code execution will be halted and process adjustments will be triggered. Potential errors will be automatically repaired and retried. If code verification passes, code execution will be allowed.
[0015] In one embodiment, extracting multi-attribute feature vectors from the cleaned data and generating a feature vector distribution map from the multi-attribute feature vectors includes: Multi-attribute feature vectors are extracted from the cleaned data, and the multi-attribute feature vectors include numerical, categorical, textual, and date variables in the cleaned data. Variables are extracted from multi-attribute feature vectors, and a distribution map is generated and integrated based on the extracted variable results to obtain the feature vector distribution map.
[0016] In one embodiment, the feature vector distribution map includes: The feature vector distribution map can adopt two interaction modes, namely intelligent analysis mode and dialogue interaction mode. The intelligent analysis mode is used to automatically recommend analysis methods based on variable type, data characteristics and research objectives, and generate preliminary insights; The dialogue interaction mode is used by users to describe their analytical needs using natural language, understand instructions through a large language model, plan the analysis process, and generate executable code.
[0017] In one embodiment, the task flow for planning and analysis is based on the feature vector distribution map. The statistical analysis scheme is then matched and dynamically adjusted based on the task flow. The adjusted statistical analysis scheme is converted into analysis code and executed to obtain automated statistical analysis results, including: Based on the feature vector distribution map, and combined with data features, user research intentions and historical interaction context, the task flow is planned and analyzed. Based on the planned analysis task process, a corresponding statistical analysis scheme is matched, which includes regression analysis, group difference test, factor analysis, cluster analysis and text sentiment analysis. Based on the analysis process, the analysis path is adjusted and the statistical analysis scheme is optimized. The optimized statistical analysis scheme is then converted into analysis code and executed using a large language model to obtain automated statistical analysis results.
[0018] In one embodiment, based on the results of automated statistical analysis, interactive visualization charts are generated. Logical deduction and causal explanation are then performed based on these charts and social science theories. The deduction results and causal explanations are integrated to obtain an analytical report with theoretical insights. This achieves intelligent control over the entire process of intelligent data analysis using a large language model, including: The automated statistical analysis results will be automatically generated into interactive visual charts that allow for chart filtering, variable linkage, drill-down analysis, and backtracking. Based on social science theories, this study performs logical deduction, causal explanation, and semantic analysis on interactive visualization charts and automated statistical analysis results. It then integrates the deduction results with the causal explanations to generate an analytical report with theoretical insights.
[0019] According to a second aspect of the present invention, an intelligent data analysis system based on a large language model is provided.
[0020] In one embodiment, the intelligent data analysis system based on a large language model includes: The data semantic parsing module is used to obtain the data structure features and association rules, and to perform variable semantic parsing on the data structure features and association rules to obtain the data catalog and variable description document; The quality risk analysis module is used to detect data quality problems based on statistical rules in the large language model, combined with data catalog, variable description documents and semantic analysis, to diagnose problems in complex anomaly reconstruction analysis processes and generate data quality risk reports. The data cleaning execution module is used to build personalized data cleaning strategies based on data quality risk reports. Using these personalized data cleaning strategies, the module reconstructs and repairs the paths of the target abnormal data. It also uses a large language model to convert the personalized data cleaning strategies into executable code. After performing logical verification and security checks on the executable code, the module executes the reconstructed and repaired target abnormal data cleaning to obtain the cleaned data. The distribution feature extraction module is used to extract multi-attribute feature vectors from the cleaned data and generate a feature vector distribution map from the multi-attribute feature vectors. The statistical analysis planning module is used to plan the analysis task process based on the feature vector distribution map, match and dynamically adjust the statistical analysis scheme based on the analysis task process, convert the adjusted statistical analysis scheme into analysis code and execute it to obtain automated statistical analysis results. The insight report generation module is used to generate interactive visualization charts based on the results of automated statistical analysis. It then performs logical deduction and causal explanation based on the interactive visualization charts and social science theories, integrates the deduction results and causal explanations to obtain an analysis report with theoretical insights, so as to realize the intelligent control of the entire process of intelligent data analysis by the big language model.
[0021] According to a third aspect of the present invention, a computer device is provided.
[0022] In some embodiments, the computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps of the method described above.
[0023] According to a fourth aspect of the present invention, a computer-readable storage medium is provided.
[0024] In one embodiment, a computer program is stored on the computer-readable storage medium, which, when executed by a processor, implements the steps of the above method.
[0025] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects: 1) This invention utilizes an intelligent task decomposition and automated data analysis mechanism based on an LLM Agent. LLM Agent automatically infers the analysis intent based on the user's natural language requirements, completing an integrated task decomposition process including data detection, intelligent data cleaning, feature construction, statistical modeling, and result verification. It also autonomously generates and executes analysis code. Through multi-round reasoning and plan refactoring mechanisms, the system autonomously identifies problems and adjusts the analysis path, enabling non-professional users to complete the entire data analysis process without programming skills, thereby significantly improving task completion efficiency and analysis accuracy.
[0026] 2) This invention employs a multi-layered self-verification and data insight generation mechanism. Through multi-level self-verification strategies such as data quality review, model robustness testing, and result consistency comparison, it achieves automatic identification and repair of abnormal results during the analysis process. Simultaneously, based on the knowledge enhancement capabilities of a large model, the system automatically discovers significant features, correlation patterns, and potential logic in the data, outputting key insights and decision-making points, thus overcoming the technical shortcomings of traditional tools that only provide static results and lack proactive discovery capabilities.
[0027] 3) This invention utilizes a knowledge-based expression and interactive visualization mechanism driven by theoretical explanations to support the automatic association of analysis results with semantic systems such as social theories and industry knowledge bases. It also uses natural language to generate interpretable conclusions, including causal explanations, behavioral logic deductions, and policy implications. Simultaneously, combined with interactive visualization, it dynamically presents the structural relationships of complex data and theoretical insights, enhancing the understandability, reproducibility, and dissemination effectiveness of the results.
[0028] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit the invention. Attached Figure Description
[0029] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0030] Figure 1 This is a flowchart illustrating an intelligent data analysis method based on a large language model, according to an exemplary embodiment. Figure 2 This is a block diagram illustrating the principle of an intelligent data analysis system based on a large language model, according to an exemplary embodiment. Figure 3 This is a schematic diagram of the structure of a computer device according to an exemplary embodiment. Detailed Implementation
[0031] The following description and accompanying drawings fully illustrate specific embodiments described herein to enable those skilled in the art to practice them. Some portions and features of certain embodiments may be included in or replace portions and features of other embodiments. The scope of the embodiments herein includes the entire scope of the claims and all available equivalents thereof. The various embodiments described herein are presented in a progressive manner, with each embodiment focusing on its differences from other embodiments; similar or identical parts between embodiments can be referred to interchangeably.
[0032] The modules in the apparatus or system of this application can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0033] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.
[0034] Figure 1 An embodiment of an intelligent data analysis method based on a large language model according to the present invention is shown.
[0035] In this optional embodiment, the intelligent data analysis method based on a large language model includes: Step S101: Obtain data structure features and association rules, and perform variable semantic parsing on the data structure features and association rules to obtain data catalog and variable description document; Step S102: Based on the statistical rules and conditions in the large language model, and combined with the data catalog, variable description document, and semantic analysis to detect data quality problems, perform problem diagnosis on the complex anomaly reconstruction analysis process and obtain a data quality risk report. Step S103: Based on the data quality risk report, construct a personalized data cleaning strategy, use the personalized data cleaning strategy to reconstruct and repair the target abnormal data, and use a large language model to convert the personalized data cleaning strategy into executable code. After performing logical verification and security checks on the executable code, execute the reconstructed and repaired target abnormal data cleaning to obtain the cleaned data. Step S104: Extract multi-attribute feature vectors from the cleaned data and generate a feature vector distribution map from the multi-attribute feature vectors; Step S105: Based on the feature vector distribution map, plan the analysis task flow, match and dynamically adjust the statistical analysis scheme based on the analysis task flow, convert the adjusted statistical analysis scheme into analysis code and execute it to obtain automated statistical analysis results. Step S106: Based on the results of automated statistical analysis, generate interactive visualization charts, and perform logical deduction and causal explanation based on the interactive visualization charts and social science theories. Integrate the deduction results and causal explanations to obtain an analysis report with theoretical insights, so as to realize the full-process intelligent control of intelligent data analysis by the large language model.
[0036] In this optional embodiment, data structure features and association rules are obtained, and variable semantic parsing is performed on the data structure features and association rules to obtain a data catalog and variable description document, including: Extracting data structure features and association rules from various structured data formats in spreadsheets and statistical software; By performing variable semantic parsing on the structured features and association rules of the data, the variable types, text encoding, scale attributes and multi-table associations are automatically identified to obtain preliminary semantic parsing results. Based on the preliminary semantic parsing results, the data of multiple worksheets is parsed to eliminate differences in fields and structural ambiguities between tables, and standardized semantic parsing results are obtained. Based on the standardized parsing results, a large language model is used to extract variable semantics, determine the logical relationships between variables, and measure the hierarchy, resulting in a data catalog and variable description documents.
[0037] Specifically, the data import and structure parsing process supports various structured data formats such as Excel (spreadsheets) and Stata (statistical software), automatically identifying variable types, text encoding, scale attributes, and multi-table relationships (i.e., data structure features and association rules). It parses multi-sheet data one by one, eliminating differences in fields and structural ambiguities between tables, and automatically generating a data catalog and variable description document. During parsing, the LLM (Large Language Model) understands variable semantics, determines logical relationships and measurement levels between variables, providing a semantic foundation for subsequent cleaning and analysis. In this invention, the LLM continuously receives analysis status and makes decisions to control the process through an analysis flow control mechanism. These decisions do not directly generate analysis conclusions but rather select or adjust subsequent analysis flow actions, including but not limited to whether to proceed to the next analysis stage, whether to trigger cleaning or verification steps, whether to adjust the analysis path, or terminate the current process. Decisions are triggered by the user. All statuses are only initiated after user confirmation or feedback, thus achieving a dynamic understanding of variable semantics and logical relationships.
[0038] In this optional embodiment, based on the statistical rules in the large language model, and combined with data catalogs, variable description documents, and semantic analysis to detect data quality issues, a problem diagnosis is performed on the complex anomaly reconstruction analysis process, resulting in a data quality risk report including: The large language model is used as the control terminal of the analysis process, and based on the target analysis node, the current data status, execution operation and decision basis are recorded to form an analysis path status chain. Based on the analysis path state chain, the current data state is used to determine the statistical rule constraints, and semantic analysis is performed by combining the semantic description of variables, the meaning of rules, and the analysis objectives. Based on statistical rule constraints and semantic analysis, data quality detection is performed on multiple worksheets to identify missing value patterns, logical contradictions, abnormal data, unreversed reverse options, duplicate answers, and extreme responses, and to obtain preliminary data quality diagnosis results. Based on complex anomalies identified by a single statistical rule, the analysis process is reconstructed by reconstructing variable states, historical processing paths, and related variable information to obtain the diagnostic results after analysis. The preliminary data quality diagnosis results are merged with the post-analysis diagnosis results. Based on the merged results, a quality risk report is generated and the problems and their severity are marked, thus obtaining the data quality risk report.
[0039] Specifically, intelligent quality diagnosis: The large language model in this invention acts as an analysis process controller, recording the current data status, executed operations, and decision-making basis at each key analysis node, forming an analysis path state chain. For each sheet, the system automatically detects various data quality issues through semantic analysis and statistical rules. Statistical rules are not a fixed set of rules but rather analysis constraint units that can be enabled or disabled, such as IQR outlier detection, whose applicability is determined by the large language model based on the current data status. Semantic analysis refers to the comprehensive judgment process by which the large language model judges the relationship between the semantic description of variables, the meaning of rules, and the analysis objectives. The model selects appropriate statistical rules for quality diagnosis based on the current analysis status and dynamically adjusts rule priorities or analysis paths when there are rule conflicts or inconsistent diagnostic results. The rules are transmitted to the LLM (Low Language Management System) for dynamic decision-making. The various data quality issues mentioned above include missing value patterns, logical contradictions, anomaly checks, unreversed reverse options, duplicate answers, and extreme responses (i.e., data quality issues detected by semantic analysis). A table-by-table quality risk report is automatically generated, marking potential problems and their severity, providing a basis for subsequent processing decisions. For example, inconsistencies or potential coding errors between related variables can be identified, thus achieving higher-precision quality diagnosis. LLM can reason and judge complex anomalies. For complex anomalies that cannot be identified by a single statistical rule, the large language model acts as a process controller, treating anomaly identification as a trigger condition for process branches. When an anomaly occurs, the model reconstructs the subsequent analysis process based on the current variable state, historical processing paths, and related variable information. The auxiliary variables for analysis process reconstruction come from the original data or derived results of processed data. For example, variable combinations and aggregation calculations. Based on these results, the model re-evaluates how to continue processing, such as introducing new auxiliary variables or changing the verification method, thereby achieving reasoning and judgment on complex anomalies through process reconstruction.
[0040] In this optional embodiment, a personalized data cleaning strategy is constructed based on the data quality risk report. This personalized strategy is then used to reconstruct and repair the target anomalous data. A large language model is used to convert the personalized data cleaning strategy into executable code. After logical verification and security checks, the executable code is executed to clean the reconstructed and repaired target anomalous data, resulting in cleaned data including: For data quality risk reports, identify target abnormal data that cannot be repaired through the established cleaning process; Using a large language model as the process control terminal, reverse reasoning is performed on the target abnormal data and the variable constraints are met to obtain the reconstructed data repair path. Based on the data repair path, variable semantics, measurement level, and data characteristics, a personalized data cleaning strategy is constructed. The personalized data cleaning strategy is converted into executable code using a large language model, and the built-in code review mechanism in the large language model is used again to perform logical verification and security checks on the executable code to obtain the approved executable code. The executable code that passes the verification is cleaned, and the cleaned results are verified a second time to obtain the cleaned data.
[0041] In this optional embodiment, the built-in code review mechanism includes: The built-in code review mechanism is used to perform logic, syntax and security checks on the generated executable code to ensure that the code conforms to the preset analysis path and analysis objectives. The code review mechanism uses the large language model as a process control node. After generating executable code, it determines whether the code conforms to the preset analysis path based on the current analysis target and process status. If code logic deviations, security risks, or syntax errors are detected, code execution will be halted and process adjustments will be triggered. Potential errors will be automatically repaired and retried. If code verification passes, code execution will be allowed.
[0042] Specifically, the automatic cleaning strategy is generated and executed as follows: When abnormal data cannot be repaired through the established cleaning process (i.e., the target abnormal data), the large language model, acting as the process controller, initiates a reverse reasoning mode. Starting from the target analysis state, it infers the variable constraints that should be satisfied. The analysis constraint unit is a predefined set of rules or conditions used to limit the legality and consistency of the analysis process, such as the aggregate value of all columns equaling the total column, or certain user-preset operations. Quantitative constraints are the instantiation results of the analysis constraint unit in the specific analysis stage, and the data repair path is reconstructed accordingly. Examples include imputation and inferring missing values based on the total column. The system automatically generates personalized cleaning strategies based on variable semantics, measurement level, and data characteristics, including missing value imputation, outlier correction, classification recoding, and consistency verification. The LLM converts the strategy into executable code (Python, Stata, etc.) and performs logic verification and security checks through a built-in code review mechanism, which is implemented by the large language model as the process control node. After generating the analysis code, the model judges whether the code conforms to the established analysis path based on the current analysis target and process state. If the code logic deviates or may introduce risks, execution is stopped and process adjustment is triggered. Therefore, code review falls under process control. After review, the code automatically performs cleaning operations and undergoes secondary verification of the cleaning results. This secondary verification is a process review step controlled by the large language model and is a crucial link in a multi-layered self-verification mechanism. Whenever the analysis process enters a new stage or a critical decision node, the large language model triggers a verification process based on the current state to determine whether the analysis can continue. The initial decision is generated by large model A, and large model B comprehensively judges the rationality of A's decision based on multiple verification results, including but not limited to data quality assessment results, anomaly repair verification results, and consistency detection results, thus forming a layered and controlled verification system. This ensures that the data meets research-grade quality standards while supporting users to fine-tune strategies and code, achieving a flexible and controllable cleaning process.
[0043] In this optional embodiment, extracting multi-attribute feature vectors from the cleaned data and generating a feature vector distribution map from the multi-attribute feature vectors includes: Multi-attribute feature vectors are extracted from the cleaned data, and the multi-attribute feature vectors include numerical, categorical, textual, and date variables in the cleaned data. Variables are extracted from multi-attribute feature vectors, and a distribution map is generated and integrated based on the extracted variable results to obtain the feature vector distribution map.
[0044] In this optional embodiment, the feature vector distribution map includes: The feature vector distribution map can adopt two interaction modes, namely intelligent analysis mode and dialogue interaction mode. The intelligent analysis mode is used to automatically recommend analysis methods based on variable type, data characteristics and research objectives, and generate preliminary insights; The dialogue interaction mode is used by users to describe their analytical needs using natural language, understand instructions through a large language model, plan the analysis process, and generate executable code.
[0045] Specifically, variable extraction and preliminary visualization: After data upload, the system automatically identifies all variables, including numerical, categorical, textual, and date types (i.e., multi-attribute feature vectors). A distribution plot is generated for each variable, which can be used to quickly determine variable distribution characteristics and outliers (i.e., feature vector distribution plot). Two interactive modes are also provided: Intelligent Analysis Mode: The system automatically recommends suitable analysis methods based on variable type, data characteristics, and research objectives, and generates preliminary insights. Chat Mode (i.e., dialogue interaction mode): Users describe their analysis needs in natural language; the LLM automatically understands the instructions, plans the analysis process, and generates executable code.
[0046] In this optional embodiment, the task flow for planning and analysis is performed based on the feature vector distribution map. The statistical analysis scheme is then matched and dynamically adjusted based on the task flow. The adjusted statistical analysis scheme is converted into analysis code and executed to obtain automated statistical analysis results, including: Based on the feature vector distribution map, and combined with data features, user research intentions and historical interaction context, the task flow is planned and analyzed. Based on the planned analysis task process, a corresponding statistical analysis scheme is matched, which includes regression analysis, group difference test, factor analysis, cluster analysis and text sentiment analysis. Based on the analysis process, the analysis path is adjusted and the statistical analysis scheme is optimized. The optimized statistical analysis scheme is then converted into analysis code and executed using a large language model to obtain automated statistical analysis results.
[0047] Specifically, intelligent analysis task planning and method recommendation: LLM (Large Language Model) automatically plans the analysis task process based on data characteristics, user research intentions, and historical interaction context, including variable selection, model type, and parameter settings. In intelligent mode, the system can recommend methods and related variables (i.e., statistical analysis schemes) for analysis, including regression analysis, group difference testing, factor analysis, cluster analysis, and text sentiment analysis. The analysis path can be dynamically adjusted, such as adding interaction effect analysis, variable group comparison, or multidimensional cluster analysis, to maximize the insight value. Multiple analysis methods refer to specific statistical or modeling techniques, while semantic analysis is the control logic used by the Large Language Model to determine when and in what order to activate the aforementioned analysis methods.
[0048] Automated code generation and secure execution: Based on the analysis task, LLM automatically generates executable Python or Stata code and supports exporting Stata command scripts. A built-in code review mechanism verifies logic, syntax, and security, and automatically repairs and retries when potential errors are detected. The system supports batch processing of multi-variable, multi-model, and multi-sheet data, ensuring the efficiency and stability of the analysis process.
[0049] In this optional embodiment, interactive visualization charts are generated based on the results of automated statistical analysis. Logical deduction and causal explanation are then performed based on these charts and social science theories. The deduction results and causal explanations are integrated to obtain an analytical report with theoretical insights. This enables intelligent control of the entire intelligent data analysis process through a large language model, including: The automated statistical analysis results will be automatically generated into interactive visual charts that allow for chart filtering, variable linkage, drill-down analysis, and backtracking. Based on social science theories, this study performs logical deduction, causal explanation, and semantic analysis on interactive visualization charts and automated statistical analysis results. It then integrates the deduction results with the causal explanations to generate an analytical report with theoretical insights.
[0050] Specifically, the system offers interactive visualization and results presentation: it can automatically generate analytical charts, including variable distribution plots, scatter plots, regression result plots, group difference comparison plots, and model prediction visualizations. The front-end interface supports chart filtering, variable linkage, drill-down analysis, and backtracking operations, facilitating in-depth data exploration. Analysis results can be exported as structured reports or Markdown documents (i.e., data import and structure analysis, and dual-model decision-making documents), enabling research sharing and policy reporting.
[0051] Theoretical Insights and Knowledge-Based Interpretation: Based on LLM and combined with social science theories, the system performs logical deduction, causal explanation, and policy implications analysis on the analysis results. LLM can automatically identify potential patterns, structural regularities, and group differences in the data, and generate understandable high-level insights. The final output report includes not only statistical results and charts, but also theoretical explanations, making the data analysis results directly applicable to decision-making and academic research.
[0052] Specifically, after a user uploads a data file, the system automatically extracts all variables and generates a distribution map for each variable, enabling preliminary visualization and exploration of the data. The system supports two analysis modes: Intelligent Analysis Mode and Chat Mode. In Intelligent Analysis Mode, the platform automatically recommends the most suitable analysis method based on data characteristics and research intent, generates executable statistical code to complete the analysis, and outputs key insights. In Chat Mode, users can express their analysis needs using natural language. The LLM (Local Management Module) plans the analysis task, generates analysis code, and executes it based on user instructions, ultimately presenting interactive visualization results and theoretical insight reports, achieving end-to-end intelligent data analysis and social science interpretation.
[0053] The core innovation of this invention lies in enabling LLM (Large Language Model) to not only understand natural language task descriptions but also autonomously identify data structures and research contexts. It automatically recommends suitable data cleaning strategies and analysis methods without any user input, completing task planning and execution. This allows users without statistical or programming backgrounds to obtain professional-grade data processing results. The TaoSight end-to-end data processing and analysis platform, based on the Large Language Model (LLM) intelligent agent architecture, achieves end-to-end automation from raw data to knowledge insights through multi-layered task decomposition, intelligent code generation, automatic cleaning and verification, analysis insight generation, and interactive visualization. It provides a low-threshold, highly reliable, and interpretable intelligent data analysis solution for social science research, policy analysis, and industrial research.
[0054] Specifically, an intelligent system is built that covers the entire process from natural language task understanding to data cleaning, statistical analysis, interactive visualization, and insight report generation. Users only need to describe their research questions to automatically complete data reading, variable identification, format conversion, and anomaly detection. It also supports parsing various data file formats such as Excel, CSV, and Stata, enabling low-threshold use and full utilization of social survey data.
[0055] The intelligent modeling and dynamic analysis recommendation mechanism can automatically recommend suitable statistical methods based on data structure and research context, and autonomously generate executable analysis commands and corresponding visualization charts, including automatically drawing distribution maps of all variables, relationship diagrams, and model result diagrams. It also supports the automatic generation of Stata commands and execution scripts, allowing users to flexibly export and reproduce the analysis process. Its technical advantages are significantly reduced in terms of lowering the professional threshold for statistical analysis, reducing human error, and improving analysis efficiency.
[0056] Based on the theoretical explanation and knowledge enhancement capabilities of a large language model, the system not only provides analysis results but also integrates professional theories from the humanities and social sciences to complete conclusion interpretation, logical deduction, and policy implications analysis. The entire processing flow has self-verification and consistency detection mechanisms to ensure the credibility and reproducibility of the analysis results. In addition, the interactive visual interface supports dynamic filtering, variable linkage viewing, and conclusion backtracking, significantly improving the readability and dissemination effect of research results.
[0057] Figure 2 An embodiment of an intelligent data analysis system based on a large language model according to the present invention is shown.
[0058] In this optional embodiment, the intelligent data analysis system based on a large language model includes: The data semantic parsing module 201 is used to obtain data structured features and association rules, and to perform variable semantic parsing on the data structured features and association rules to obtain a data catalog and variable description document; The quality risk analysis module 202 is used to detect data quality problems based on the statistical rules in the large language model, combined with the data catalog, variable description document and semantic analysis, to diagnose problems in the complex anomaly reconstruction analysis process and obtain a data quality risk report. The data cleaning execution module 203 is used to build a personalized data cleaning strategy based on the data quality risk report. The personalized data cleaning strategy is used to reconstruct and repair the target abnormal data. The personalized data cleaning strategy is converted into executable code using a large language model. After logical verification and security detection of the executable code, the reconstructed and repaired target abnormal data is cleaned to obtain the cleaned data. The distribution feature extraction module 204 is used to extract multi-attribute feature vectors from the cleaned data and generate a feature vector distribution map from the multi-attribute feature vectors. The statistical analysis planning module 205 is used to plan the analysis task flow based on the feature vector distribution map, match and dynamically adjust the statistical analysis scheme based on the analysis task flow, convert the adjusted statistical analysis scheme into analysis code and execute it to obtain automated statistical analysis results. The insight report generation module 206 is used to generate interactive visualization charts based on the results of automated statistical analysis, and to perform logical deduction and causal explanation based on the interactive visualization charts and social science theories. The deduction results and causal explanations are integrated to obtain an analysis report of theoretical insights, so as to realize the intelligent control of the entire process of intelligent data analysis by the big language model.
[0059] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 3As shown, the computer device includes a processor, memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database stores static and dynamic information data. The network interface communicates with external terminals via a network connection. When the computer program is executed by the processor, it implements the steps in the above method embodiments.
[0060] Those skilled in the art will understand that Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present invention and does not constitute a limitation on the computer device to which the present invention is applied. A specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0061] In addition, the present invention also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.
[0062] In addition, the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.
[0063] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. Any references to memory, storage, databases, or other media used in the embodiments provided by this invention can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical storage, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0064] This invention is not limited to the structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this invention is limited only by the appended claims.
Claims
1. An intelligent data analysis method based on a large language model, characterized in that, include: Obtain the data structure features and association rules, and perform variable semantic parsing on the data structure features and association rules to obtain the data catalog and variable description document; Based on the statistical rules in the large language model, and combined with the data catalog, variable description documents, and semantic analysis to detect data quality problems, a problem diagnosis is performed on the complex anomaly reconstruction analysis process to obtain a data quality risk report. Based on the data quality risk report, a personalized data cleaning strategy is constructed. The personalized data cleaning strategy is used to reconstruct and repair the target abnormal data. The personalized data cleaning strategy is converted into executable code using a large language model. After logical verification and security testing of the executable code, the reconstructed and repaired target abnormal data is cleaned to obtain the cleaned data, including: the target abnormal data that cannot be repaired by the determined cleaning process based on the data quality risk report. Using a large language model as the process control terminal, reverse reasoning is performed on the target abnormal data and the variable constraints are met to obtain the reconstructed data repair path. Based on the data repair path, variable semantics, measurement level, and data characteristics, a personalized data cleaning strategy is constructed. The personalized data cleaning strategy is converted into executable code using a large language model, and the built-in code review mechanism in the large language model is used again to perform logical verification and security checks on the executable code to obtain the approved executable code. The executable code that passes the verification is cleaned, and the cleaning results are verified a second time to obtain the cleaned data. The built-in code review mechanism includes: The built-in code review mechanism is used to perform logic, syntax and security checks on the generated executable code to ensure that the code conforms to the preset analysis path and analysis objectives. The code review mechanism uses the large language model as a process control node. After generating executable code, it determines whether the code conforms to the preset analysis path based on the current analysis target and process status. If code logic deviations, security risks, or syntax errors are detected, code execution will be halted and process adjustments will be triggered, with automatic repair and retries for potential errors; if code validation passes, code execution will be allowed. Extract multi-attribute feature vectors from the cleaned data and generate a feature vector distribution map from the multi-attribute feature vectors. Based on the feature vector distribution map, the task flow is planned and analyzed. The statistical analysis scheme is matched and dynamically adjusted based on the analysis task flow. The adjusted statistical analysis scheme is converted into analysis code and executed to obtain automated statistical analysis results. Based on the results of automated statistical analysis, interactive visualization charts are generated. Logical deduction and causal explanation are then performed based on the interactive visualization charts and social science theories. The deduction results and causal explanations are integrated to obtain an analytical report with theoretical insights, thereby achieving intelligent control of the entire process of intelligent data analysis by a large language model.
2. The intelligent data analysis method based on a large language model according to claim 1, characterized in that, The process of obtaining data structured features and association rules, and performing variable semantic parsing on these features and rules to obtain a data catalog and variable description document includes: Extracting data structure features and association rules from various structured data formats in spreadsheets and statistical software; By performing variable semantic parsing on the structured features and association rules of the data, the variable types, text encoding, scale attributes and multi-table associations are automatically identified to obtain preliminary semantic parsing results. Based on the preliminary semantic parsing results, the data of multiple worksheets is parsed to eliminate differences in fields and structural ambiguities between tables, and standardized semantic parsing results are obtained. Based on the standardized parsing results, a large language model is used to extract variable semantics, determine the logical relationships between variables, and measure the hierarchy, resulting in a data catalog and variable description documents.
3. The intelligent data analysis method based on a large language model according to claim 1, characterized in that, The process involves diagnosing data quality issues based on statistical rules in the large language model, combined with data catalogs, variable description documents, and semantic analysis. This process diagnoses problems in the complex anomaly reconstruction analysis workflow and generates a data quality risk report, including: The large language model is used as the control terminal of the analysis process, and based on the target analysis node, the current data status, execution operation and decision basis are recorded to form an analysis path status chain. Based on the analysis path state chain, the current data state is used to determine the statistical rule constraints, and semantic analysis is performed by combining the semantic description of variables, the meaning of rules, and the analysis objectives. Based on statistical rule constraints and semantic analysis, data quality detection is performed on multiple worksheets to identify missing value patterns, logical contradictions, abnormal data, reversed options not reversed, duplicate answers and extreme responses, and to obtain preliminary data quality diagnosis results. Based on complex anomalies identified by a single statistical rule, the analysis process is reconstructed by reconstructing variable states, historical processing paths, and related variable information to obtain the diagnostic results after analysis. The preliminary data quality diagnosis results are merged with the post-analysis diagnosis results. Based on the merged results, a quality risk report is generated and the problems and their severity are marked, thus obtaining the data quality risk report.
4. The intelligent data analysis method based on a large language model according to claim 1, characterized in that, The step of extracting multi-attribute feature vectors from the cleaned data and generating a feature vector distribution map from the multi-attribute feature vectors includes: Multi-attribute feature vectors are extracted from the cleaned data, and the multi-attribute feature vectors include numerical, categorical, textual, and date variables in the cleaned data. Variables are extracted from multi-attribute feature vectors, and a distribution map is generated and integrated based on the extracted variable results to obtain the feature vector distribution map.
5. The intelligent data analysis method based on a large language model according to claim 4, characterized in that, The feature vector distribution map includes: The feature vector distribution map can adopt two interaction modes, namely intelligent analysis mode and dialogue interaction mode. The intelligent analysis mode is used to automatically recommend analysis methods based on variable type, data characteristics and research objectives, and generate preliminary insights; The dialogue interaction mode is used by users to describe their analytical needs using natural language, understand instructions through a large language model, plan the analysis process, and generate executable code.
6. The intelligent data analysis method based on a large language model according to claim 1, characterized in that, The process of planning and analyzing tasks based on feature vector distribution maps, matching and dynamically adjusting statistical analysis schemes based on the analysis task process, converting the adjusted statistical analysis schemes into analysis code and executing them to obtain automated statistical analysis results includes: Based on the feature vector distribution map, and combined with data features, user research intentions, and historical interaction context, the task flow is planned and analyzed. Based on the planned analysis task process, a corresponding statistical analysis scheme is matched, which includes regression analysis, group difference test, factor analysis, cluster analysis and text sentiment analysis. Based on the analysis process, the analysis path is adjusted and the statistical analysis scheme is optimized. The optimized statistical analysis scheme is then converted into analysis code and executed using a large language model to obtain automated statistical analysis results.
7. The intelligent data analysis method based on a large language model according to claim 1, characterized in that, The process involves generating interactive visualization charts based on automated statistical analysis results, and then performing logical deductions and causal explanations based on these charts and social science theories. The deduction results and causal explanations are integrated to obtain a theoretically insightful analysis report. This achieves intelligent control over the entire process of intelligent data analysis using a large language model, including: The automated statistical analysis results will be automatically generated into interactive visual charts that allow for chart filtering, variable linkage, drill-down analysis, and backtracking. Based on social science theories, this study performs logical deduction, causal explanation, and semantic analysis on interactive visualization charts and automated statistical analysis results. It then integrates the deduction results with the causal explanations to generate an analytical report with theoretical insights.
8. An intelligent data analysis system based on a large language model, used to implement the steps of the intelligent data analysis method based on a large language model according to any one of claims 1-7, characterized in that, include: The data semantic parsing module is used to obtain the data structure features and association rules, and to perform variable semantic parsing on the data structure features and association rules to obtain the data catalog and variable description document; The quality risk analysis module is used to detect data quality problems based on statistical rules in the large language model, combined with data catalog, variable description documents and semantic analysis, to diagnose problems in complex anomaly reconstruction analysis processes and generate data quality risk reports. The data cleaning execution module is used to build personalized data cleaning strategies based on data quality risk reports. Using these personalized data cleaning strategies, the module reconstructs and repairs the paths of the target abnormal data. It also uses a large language model to convert the personalized data cleaning strategies into executable code. After performing logical verification and security checks on the executable code, the module executes the reconstructed and repaired target abnormal data cleaning to obtain the cleaned data. The distribution feature extraction module is used to extract multi-attribute feature vectors from the cleaned data and generate a feature vector distribution map from the multi-attribute feature vectors. The statistical analysis planning module is used to plan the analysis task process based on the feature vector distribution map, match and dynamically adjust the statistical analysis scheme based on the analysis task process, convert the adjusted statistical analysis scheme into analysis code and execute it to obtain automated statistical analysis results. The insight report generation module is used to generate interactive visualization charts based on the results of automated statistical analysis. It then performs logical deduction and causal explanation based on the interactive visualization charts and social science theories, integrates the deduction results and causal explanations to obtain an analysis report with theoretical insights, so as to realize the intelligent control of the entire process of intelligent data analysis by the big language model.
Citation Information
Patent Citations
Biological information drawing system based on multiple agents
CN120339451A
Enterprise process intelligent analysis system based on large language model
CN121119659A