Data analysis and chart generation method, apparatus, device, and medium
By receiving structured data and natural language instructions, and using a large AI model to analyze data features and analysis objectives, compatible chart templates are generated, solving the problem of existing tools relying on manual operation and achieving flexible and efficient data analysis and visualization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA PING AN LIFE INSURANCE CO LTD
- Filing Date
- 2026-05-09
- Publication Date
- 2026-07-31
AI Technical Summary
Existing data analysis tools rely too heavily on manual operation, have poor interactivity and flexibility, and cannot dynamically adapt to data characteristics, resulting in poor display results.
By receiving structured data files, natural language instructions, and chart type preferences, the system uses a large AI model to analyze data features and analytical objectives, verifies chart type compatibility, recommends or selects the final chart template, and generates visual charts and reports.
It reduces reliance on manual intervention, improves operational flexibility and accuracy, ensures consistency between analytical conclusions and visual representations, and is suitable for non-professional users.
Smart Images

Figure CN122491235A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of data processing and artificial intelligence technology, and in particular to a data analysis and chart generation method, apparatus, device and medium, which can be applied to fields such as financial data visualization or medical data visualization. Background Technology
[0002] With the rapid development of information technology, we have fully entered the era of big data. In this context, the application value of data analysis and data visualization technologies is increasingly prominent, becoming key means for various industries to unlock data value and improve decision-making efficiency. Data visualization, as a modern information display method integrating data presentation, analysis, and interaction, has been widely used in many fields such as financial data processing and medical data processing.
[0003] However, the inventors discovered that current mainstream data analysis tools (such as traditional BI systems) have significant flaws. They rely heavily on manual implementation, requiring users to manually select chart types and design analysis logic, which demands high levels of professional knowledge. Data analysis, chart generation, and report writing are separated, requiring multiple interactive operations. More importantly, most existing analysis tools use fixed rules to generate charts, which cannot dynamically adapt to data characteristics, resulting in poor final display results and insufficient flexibility. Summary of the Invention
[0004] This invention provides a data analysis and chart generation method, apparatus, computer equipment, and medium to solve the technical problems of existing analysis tools being overly reliant on manual labor, having poor interactivity, and lacking flexibility.
[0005] Firstly, it provides a data analysis and chart generation method, including: Receive structured data files, natural language instructions, and chart type preferences uploaded by users; The system uses a pre-deployed large model to parse structured data files, extract the data and its structural features, and parses natural language instructions to obtain the analysis target. Verify the compatibility between chart type preferences, structural characteristics, and analysis objectives; If incompatible, a compatible chart type will be recommended to the user based on the preset template library, and the final chart template will be confirmed based on the user's feedback; If compatible, the final chart template will be selected from the preset template library based on the user's input chart type preference; The data is analyzed based on the analytical objectives to obtain the analytical results, and then visualization charts and analysis reports are generated based on the analytical results and the final chart template.
[0006] Secondly, a data analysis and chart generation device is provided, including: The receiving module is used to receive structured data files, natural language instructions, and chart type preferences uploaded by users; The parsing module is used to parse structured data files using a pre-deployed large model, extract the data and its structural features, and parse natural language instructions to obtain the analysis target. The validation module is used to verify the compatibility of chart type preferences with structural features and analysis objectives. The template matching module is used to recommend compatible chart types to users based on a preset template library when incompatible chart types are found, and to confirm the final chart template based on user feedback; or, when compatible chart types are found, to select the final chart template from the preset template library based on the user's input chart type preferences. The generation module is used to analyze data based on the analysis objectives, obtain analysis results, and generate visualization charts and analysis reports based on the analysis results and the final chart template.
[0007] Thirdly, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described data analysis and chart generation methods.
[0008] Fourthly, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the steps of the above-described data analysis and chart generation methods.
[0009] The aforementioned data analysis and chart generation methods, devices, computer equipment, and storage media enable the simultaneous reception of structured data files, natural language instructions, and chart type preferences. A large AI model then synchronously parses the structured data files and natural language instructions to obtain data features and analysis objectives, replacing the manual data analysis step and reducing the need for human intervention. Natural language instructions allow even non-professionals to explain their needs using text descriptions, minimizing the need for user expertise. Furthermore, the system proactively verifies the compatibility of the user's selected chart type. If compatible, it directly uses the user's preferred template to satisfy their chart type preference. In case of incompatibility, it automatically generates recommended options from a preset template library, requiring only user confirmation for modification. This eliminates the need for non-professional users to understand chart application rules, increases operational error tolerance, and provides high flexibility. Moreover, it allows for data analysis, chart generation, and report writing to be completed within the same system, eliminating information fragmentation caused by cross-platform operations and ensuring strict consistency between analytical conclusions and visual representations. Attached Figure Description
[0010] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 This is a schematic diagram of an application environment for a data analysis and chart generation method according to an embodiment of the present invention.
[0012] Figure 2 This is a schematic flowchart of a data analysis and chart generation method in one embodiment of the present invention.
[0013] Figure 3 This is a schematic diagram of a data analysis and chart generation device in one embodiment of the present invention.
[0014] Figure 4 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention. Detailed Implementation
[0015] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0016] The data analysis and chart generation method provided in this invention can be applied to, for example... Figure 1In this application environment, the client is used for users to upload structured data files, natural language instructions, and chart type preferences. The client communicates with the server via a network, specifically employing end-to-end encryption technology to ensure the security of client data during transmission. The client transmits structured data files, natural language instructions, and chart type preferences to the server via the transmission network. The server receives the user-uploaded structured data files, natural language instructions, and chart type preferences; then, it uses a pre-deployed large model to parse the structured data files, extracting the data and its structural features, and simultaneously parses the natural language instructions to obtain the analysis target. It then verifies the compatibility of the chart type preferences with the structural features and analysis target. If incompatible, it recommends compatible chart types to the user based on a preset template library and confirms the final chart template based on user feedback; if compatible, it selects the final chart template from the preset template library based on the user's input chart type preferences. Finally, it analyzes the data based on the analysis target, obtains the analysis results, and generates visual charts and analysis reports based on the analysis results and the final chart template. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server is implemented using a cloud computing platform. The present invention will now be described in detail through specific embodiments.
[0017] Please see Figure 2 As shown, Figure 2 A flowchart illustrating the data analysis and chart generation method provided in this embodiment of the invention includes the following steps: Step S1: Receive the structured data file, natural language instructions, and chart type preferences uploaded by the user.
[0018] It should be noted that the structured data file typically includes CSV, Excel, and JSON files. In this embodiment, an Excel file is preferred, as it is the most commonly used structured data carrier in business analysis. It adopts a standard two-dimensional table structure, where rows typically represent records (such as sales revenue, sales date, and financial product category for financial products), and columns represent variables (such as the specific monthly sales revenue figures for each financial product), allowing for the recording of large amounts of structured data. The natural language command refers to phrases or sentences directly entered by the user. Users can directly explain their data analysis needs through text descriptions, making operation more convenient. Chart type preference refers to the type of chart the user prefers for display. It should be understood that in this embodiment, chart types include, but are not limited to, line charts, bar charts, pie charts, scatter plots, candlestick charts, tree diagrams, and dashboards.
[0019] In some embodiments, the structured data file may be financial data, such as sales records of various financial products, sales customer information records, etc. In other embodiments, the structured data file may be medical data, such as patient medical records, drug usage records, etc.
[0020] Step S2: Use a pre-deployed large model to parse the structured data file, extract the data and its structural features, and parse natural language instructions to obtain the analysis target.
[0021] Specifically, this pre-deployed large model can be implemented using models such as Deepseek, ChatGPT, and Claude. The pre-deployed large model extracts data from the structured data file, then parses the data to obtain its structural features. These structural features are used to define the data structure; for example, the data might be based on a proportional distribution of different categories (e.g., the number of people corresponding to different health levels in the healthcare field) or a trend distribution based on time series (e.g., sales data of financial products over several consecutive months). Simultaneously with parsing the data's structural features, the user's input natural language commands are also parsed to determine the user's analytical objectives.
[0022] Furthermore, in step S2 above, the step of using a pre-deployed large model to parse the structured data file and extract the data and its structural features specifically includes: 1. Use a pre-deployed large model to parse structured data files and extract structured data.
[0023] Specifically, this pre-deployed large model needs to be able to automatically identify file types and then call a dedicated parsing engine based on the file type. For example, for Excel files, a spreadsheet library is needed to read all worksheets; for CSV files, the delimiter (comma / tab, etc.) and encoding format (UTF8 / GBK, etc.) need to be automatically detected; and for JSON files, nested data structures are directly parsed. By parsing the tables, row and column headers, data types, etc., the structured data in the structured data file is extracted.
[0024] 2. Preprocess the structured data and convert the processed structured data into standard JSON format data.
[0025] It should be noted that the preprocessing of structured data mainly includes handling missing values, duplicate values, and outliers in the structured data to complete data cleaning. Specifically, in handling missing values, for missing numerical columns, the median of the column can be used to fill the missing values; for missing categorical columns, the most frequent value (mode) of the column can be used to fill the missing values; and for missing time columns, interpolation between the preceding and following time points can be used. In handling duplicate values, full-field matching or key-field matching can be used. For full-field matching, only the first identical record can be kept; for key-field matching, duplicates can be removed based on the business primary key field. In handling outliers, outliers can first be identified based on statistical distribution (such as the IQR rule) or cluster analysis, and then deleted, or interpolation can be used to replace the outlier by combining the preceding and following data.
[0026] 3. Identify discrete variables used for classification or grouping, continuous numerical variables to be analyzed, and time series identifiers in standard JSON format data.
[0027] Discrete variables are variables with a finite number of values that are typically used for classification or grouping. They do not represent numerical values but rather categories. Specific methods for identifying discrete variables include: (1) Text fields are automatically classified as discrete variables. Any field whose data type is string (text) is classified as a discrete variable by default. For example: insurance name, user gender, claim record, etc. Text fields usually represent category information. Even if there are a large number of unique values (such as user comments), they are still regarded as discrete variables before further processing (such as natural language processing).
[0028] (2) Low-cardinality numeric fields are treated as discrete variables. Some fields, although stored as numeric types (integers or floating-point numbers), actually represent categories, such as product type codes (1 represents medical insurance, 2 represents property insurance), user levels (level 15), etc. The judgment criteria for these low-cardinality numeric fields need to be set in advance, such as: when the number of unique values in a numeric field is less than 20, it is considered to be low-cardinality and therefore treated as a discrete variable.
[0029] (3) Enhanced identification using field names containing specific semantics. Keywords in field names can be used to help identify discrete variables. For example, if a field name contains words such as "type," "category," "status," "code," "identifier," or "name," it will be preferentially identified as a discrete variable, even if its data type is numeric. Examples include: product_type (product type), user_status (user status), and country_code (country code). This method can avoid incorrectly identifying certain numeric fields representing categories as continuous variables.
[0030] A continuous numerical variable refers to a variable that can represent the magnitude of a numerical value and can theoretically take any value within a certain range. Methods for identifying continuous numerical variables include: (1) Automatic classification of floating-point / integer fields: All fields of numerical types (including integers and floating-point numbers) are first classified as continuous numerical variables. For example: sales amount, temperature, height, etc. It should be understood that numerical fields may also be discrete variables (such as in the case of low cardinality mentioned above), so further judgment is required.
[0031] (2) Enhanced identification of high-cardinality fields (unique values > 50% of the total number of records). If a numerical field has a very high cardinality (i.e., the number of unique values exceeds 50% of the total number of records), then strengthen its identification as a continuous variable. High cardinality indicates that the field has a rich range of values, a low possibility of representing categories, and is more likely to represent continuously changing numerical values. For example, in a dataset with 1000 records, if a certain numerical field has more than 500 unique values, it may be a continuous variable.
[0032] (3) Exclude ID-type fields (names containing "ID", "number", etc.). Although some numerical fields (such as user ID, order ID) may have a high cardinality (even unique), they are actually identifiers rather than continuous variables for numerical analysis. Therefore, they need to be excluded. The specific judgment method can be through the field name. When the field name contains keywords such as "ID", "number", "sequence number", "No.", etc., even if it is a numerical type and has a high cardinality, it is not regarded as a continuous variable. For example: user_id (user ID), order_number, etc.
[0033] A time series variable refers to a variable that represents a time point and is usually used for time series analysis. Methods for identifying time series variables include: (1) Format matching: Identify common time formats. The system will attempt to match the values of string fields with known time formats, including: ISO date format: "YYYYMMDD" (such as "20231001"), Chinese date format: "YYYY year MM month DD day" (such as "2023 year 10 month 01 day"), timestamp format: such as "YYYY / MM / DD" (such as "2023 / 10 / 01"), "MM / DD / YYYY" (such as "10 / 01 / 2023"), etc. Formats with time: "YYYYMMDDHH:MM:SS" (such as "2023100110:10:10"). If the value of a field mostly conforms to these formats, then it is identified as a time series variable.
[0034] (2) Semantic analysis: Field names contain time-related keywords. Keywords in field names are used to help identify time series variables. For example, if the field name contains words such as "date", "time", "month", "year", "quarter", "week", or "day", it is given priority as a time series variable. For example: "order_date" (order date), "create_time" (creation time), "month" (month).
[0035] (3) Content Validation: Values conform to time-increasing / periodic characteristics. In addition to format and name, it is also necessary to verify whether the field values have time-series characteristics: a) Increasing: In an ordered dataset, the values of time fields usually show an increasing trend (i.e., the time of the later record is no earlier than the time of the previous record); b) Periodicity: The values of time fields may show periodicity, such as daily, monthly, etc. For example, the value of a date field may cover multiple consecutive months, or there may be records every day; c) Reasonable range: The time values should be within a reasonable range (e.g., the year is between 1900 and 2100). If the field values meet these characteristics, it can be further confirmed that it is a time-series variable.
[0036] This embodiment can intelligently identify discrete variables, continuous numerical variables, and time series variables through comprehensive judgment of the above three dimensions, providing an important foundation for subsequent data analysis (such as grouping and aggregation, statistical description, time series analysis, etc.).
[0037] Step S3: Verify the compatibility of chart type preference with structural features and analysis objectives. If incompatible, proceed to step S4; if compatible, proceed to step S5.
[0038] It's important to understand that chart types include line charts, bar charts, pie charts, scatter plots, candlestick charts, tree diagrams, dashboards, etc. Different chart types can be used to display data with different structural characteristics. For example, data with a time-series distribution is suitable for line charts but not for pie charts. Therefore, when using charts to display data, it is also necessary to match the chart type selected by the user with the structural characteristics of the data and the user's analysis objectives, and confirm whether the chart type preference is compatible with the structural characteristics and analysis objectives.
[0039] Furthermore, step S3 specifically includes: 1. Call the set of chart applicable conditions and the conflict rule library in the preset rule library. The set of chart applicable conditions includes the mapping relationship between chart type and structural features and analysis objectives. The conflict rule library stores incompatible chart types and structural features and analysis objectives combinations and explanations.
[0040] Specifically, this pre-built rule base includes a set of chart applicability conditions and a conflict rule base. The set of chart applicability conditions is a user-defined mapping relationship library that defines the applicable conditions for each chart type (such as bar charts, line charts, pie charts, etc.). These conditions include structural characteristics and analytical objectives. Structural characteristics include the number of variables (e.g., a pie chart requires one discrete variable and one continuous variable), type (e.g., a scatter plot requires two continuous variables), and data scale (e.g., a pie chart is not suitable for data with more than 10 categories). Analytical objectives represent the analytical purpose the user wants to achieve through the chart, such as comparison (comparing values of different categories), distribution (showing the distribution of data), relationship (showing the relationship between two variables), composition (showing the relationship between parts and the whole), and trend (showing the trend of data changes over time). Table 1 below shows a set of chart applicability conditions and a reference example. Table 1 The conflict rule library stores incompatible chart types and combinations of structural features and analytical objectives, along with explanations. These rules are used to exclude clearly inappropriate charts. Example rules: Conflict combination 1: Pie chart + discrete variable with more than 10 categories. Explanation: Pie charts are difficult to read when there are too many categories; bar charts or column charts are recommended. Conflict combination 2: Line chart + non-time series discrete variable. Explanation: Line charts require the X-axis to be a continuous ordinal variable (usually time); otherwise, it will mislead the trend. Conflict combination 3: Heatmap + single continuous variable. Explanation: Heatmaps require two categorical variables and one numerical variable; a single variable does not meet this requirement.
[0041] 2. Verify the compatibility of chart type preferences and structural characteristics with the analysis objectives based on the chart applicability condition set and conflict rule base.
[0042] Specifically, after obtaining the user's chart type preference, data structure characteristics, and analysis objectives, firstly, rules are extracted from the applicable conditions set according to the chart type. For example, if the user prefers a pie chart, the applicable conditions for a pie chart are extracted: it requires one discrete variable and one continuous variable, and the number of discrete variable categories should ideally be between 5 and 10. Secondly, the data structure characteristics are checked to see if they meet the applicable conditions, specifically including: (1) checking if there is one discrete variable and one continuous variable in the data, and (2) checking if the number of unique values of the discrete variable is within the recommended range (e.g., 5-10). If not, it is recorded as a conflict (but it may not be an absolute conflict, but rather a warning). Then, the conflict rule base is checked, and all rules for this chart type are traversed in the conflict rule base to check if the current data characteristics trigger conflict conditions. For example, one of the conflict rules for a pie chart is that the number of discrete variable categories exceeds 10. If the current data has 15 discrete variable categories, a conflict is triggered. Finally, a verification report is generated. If the applicable conditions are met and there is no conflict, it is marked as "compatible". If the applicable conditions are met but a conflicting rule is triggered, it is marked as "incompatible" and the reason for the conflict is returned. If the applicable conditions are not met, it is marked as "applicable conditions not met" and the specific conditions that are not met are indicated.
[0043] Furthermore, in some embodiments, the chart type preference entered by the user may include multiple chart types. In the case of multiple chart types, the above verification needs to be performed on each chart type separately, and the verification results should be fed back.
[0044] Step S4: Recommend compatible chart types to users based on the preset template library, and confirm the final chart template based on user feedback.
[0045] Specifically, the preset template library is pre-built and stores templates corresponding to each chart type. When the chart type preference entered by the user is incompatible with the structural characteristics of the data and the analysis objectives, a compatible chart type is recommended to the user. Then, the final chart template is confirmed based on the user's feedback. For example, when the user confirms the use of the recommended chart type, the chart template corresponding to that chart type is retrieved from the preset template library. When the user confirms that the chart type preference is still used, the chart template corresponding to the chart type preference is retrieved from the preset template library. The final selected chart template is still selected based on the user's feedback.
[0046] Furthermore, in step S4, the step of recommending compatible chart types to the user based on the preset template library specifically includes: 1. When the analysis target involves time series, recommend line charts or area charts to users.
[0047] Specifically, line charts can clearly show the trend of data changes over time (such as the change in sales over several consecutive days), while area charts emphasize the magnitude of change on the basis of line charts, and can also show the cumulative relationship of multiple time series (such as the superposition of sales trends of different product categories). Therefore, when the analysis target involves time series, line charts or area charts are recommended to users.
[0048] 2. When analyzing the relationship between parts and the whole, recommend pie charts or donut charts to users.
[0049] Specifically, pie charts visually represent the proportion of each part to the whole (such as market share distribution), while donut charts leave a central space on top of pie charts to display additional information (such as total sales figures). Therefore, when the analysis involves the relationship between parts and the whole, pie charts or donut charts are recommended to users.
[0050] 3. When the analysis target involves the correlation of variables, recommend scatter plots or heatmaps to users.
[0051] Specifically, scatter plots demonstrate the correlation between two continuous variables (such as the relationship between advertising spending and sales), while heatmaps are suitable for demonstrating the strength of the association between two discrete variables (such as the sales popularity of different regions and product categories). Therefore, when the analysis objective involves the correlation of variables, scatter plots or heatmaps are recommended to users.
[0052] Step S5: Select the final chart template from the preset template library based on the chart type preference entered by the user.
[0053] Specifically, when the chart type preference entered by the user is compatible with the structural characteristics of the data and the analysis objectives, the corresponding template can be directly selected from the preset template library based on the chart type preference.
[0054] Step S6: Analyze the data based on the analysis objectives, obtain the analysis results, and generate visualization charts and analysis reports based on the analysis results and the final chart template.
[0055] Specifically, the analysis objectives include, but are not limited to, data statistics, modeling, mining, and data prediction. After analyzing the data according to the analysis template and obtaining the analysis results, the analysis results are embedded into the final chart template to generate a visual chart of the data set. Then, an analysis report is generated based on the analysis results, and finally, the visual chart and analysis report are presented to the user.
[0056] Furthermore, in some application scenarios, users need to make predictions by analyzing data. For example, in the sales industry, it is necessary to predict future sales based on historical sales data in order to prepare in advance. Therefore, in some embodiments, step S6 specifically includes: 1.1 When the analysis target includes a prediction request, the data is input into a pre-trained prediction model to make a prediction and obtain the prediction result.
[0057] 1.2. Combine and render the data and prediction results into the final chart template to generate a visual chart that includes both the analysis results and the prediction results.
[0058] Specifically, for some time-series data, future data can be predicted based on changes in historical data over time. For example, in the sales industry, sales data for the next month, the next quarter, and the next year can be predicted based on monthly, quarterly, and yearly sales data, thus making it easier for sales personnel to set future work plans.
[0059] Furthermore, in step S6, the steps of generating visualization charts and analysis reports based on the analysis results and the final chart template specifically include: 2.1 Generate visualization charts based on the analysis results and the final chart template.
[0060] 2.2. Generate textual conclusions based on the analysis results using a pre-deployed large model.
[0061] 2.3 Embed the visual charts into the text conclusions and output them for display.
[0062] Specifically, in order to facilitate the display of visualization charts and analysis reports, this embodiment embeds the visualization charts into the text conclusions after generating the visualization charts and text conclusions, generating a complete analysis report to be output and displayed to the user. The user only needs to read the analysis report to complete the reading of the data analysis results.
[0063] Furthermore, in some embodiments, when users upload structured data files, there may be instances where they do not input a preference for chart type. Therefore, in some embodiments, the data analysis and chart generation method further includes: 1. When the user does not input a preference for chart type, the final chart template is selected from the preset template library based on structural features and analysis objectives; 2. Embed the data into the final chart template, generate a visual chart, and output it for display.
[0064] Specifically, chart type preferences can also be obtained through natural language commands input by the user. In some cases, the user may not have inputted chart type preferences. In such cases, a compatible chart template can be selected directly from the preset template library based on the structural characteristics and analysis objectives as the final chart template. The analysis results can then be rendered and visualized based on this final chart template.
[0065] The data analysis and chart generation method in this embodiment simultaneously receives structured data files, natural language instructions, and chart type preferences. It then utilizes a large AI model to synchronously parse the structured data files and natural language instructions to obtain data features and analysis objectives. This replaces the manual data analysis step, reducing the need for human intervention. The natural language instructions allow even non-experts to explain their needs using text descriptions, minimizing the need for user expertise. Furthermore, it proactively verifies the compatibility of the user's selected chart type. If compatible, it directly uses the user's preferred template to satisfy their chart type preference. In case of incompatibility, it automatically generates recommended options from a preset template library, requiring only user confirmation for modification. This eliminates the need for non-experts to understand chart application rules, increases operational error tolerance, and provides high flexibility. Moreover, it completes data analysis, chart generation, and report writing within the same system, eliminating information fragmentation caused by cross-platform operations and ensuring strict consistency between analytical conclusions and visual representations.
[0066] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0067] In one embodiment, a data analysis and chart generation apparatus is provided, which corresponds one-to-one with the data analysis and chart generation methods described in the above embodiments. For example... Figure 3 As shown, the data analysis and chart generation device includes: a receiving module 10, a parsing module 11, a verification module 12, a template matching module 13, and a generation module 14.
[0068] The receiving module 10 is used to receive structured data files, natural language instructions, and chart type preferences uploaded by the user; Parsing module 11 is used to parse structured data files using a pre-deployed large model, extract the data and its structural features, and parse natural language instructions to obtain the analysis target; Validation module 12 is used to verify the compatibility of chart type preferences with structural features and analysis objectives; The template matching module 13 is used to recommend compatible chart types to the user based on the preset template library when incompatible, and to confirm the final chart template based on the user's feedback; or, when compatible, to select the final chart template from the preset template library based on the chart type preference entered by the user. The generation module 14 is used to analyze data based on the analysis objectives, obtain analysis results, and generate visualization charts and analysis reports based on the analysis results and the final chart template.
[0069] Optionally, the parsing module 11 performs operations to parse the structured data file using a pre-deployed large model and extract the data and its structural features. Specifically, this includes: parsing the structured data file using a pre-deployed large model and extracting structured data; preprocessing the structured data and converting the processed structured data into standard JSON format data; and identifying discrete variables used for classification or grouping, continuous numerical variables to be analyzed, and time series identifiers in the standard JSON format data.
[0070] Optionally, the verification module 12 performs an operation to verify the compatibility of chart type preferences with structural features and analysis objectives. Specifically, this includes: calling the chart applicability condition set and conflict rule set in the preset rule base. The chart applicability condition set includes the mapping relationship between chart type, structural features, and analysis objectives. The conflict rule set stores incompatible combinations of chart types, structural features, and analysis objectives, along with explanations. Based on the chart applicability condition set and the conflict rule set, the module verifies whether the chart type preferences are compatible with structural features and analysis objectives.
[0071] Optionally, the template matching module 13 performs the operation of recommending compatible chart types to the user based on a preset template library, specifically including: recommending line charts or area charts to the user when the analysis target involves time series; recommending pie charts or donut charts to the user when the analysis target involves the relationship between parts and the whole; and recommending scatter plots or heatmaps to the user when the analysis target involves the correlation of variables.
[0072] Optionally, the generation module 14 performs the following operations: analyzes the data based on the analysis objective, obtains the analysis results, and generates a visualization chart based on the analysis results and the final chart template. Specifically, when the analysis objective includes a prediction request, the data is input into a pre-trained prediction model for prediction to obtain the prediction results; the data and prediction results are merged and rendered into the final chart template to generate a visualization chart that includes the analysis results and the prediction results.
[0073] Optionally, the generation module 14 performs the operation of generating visualization charts and analysis reports based on the analysis results and the final chart template, specifically including: generating visualization charts based on the analysis results and the final chart template; generating text conclusions based on the analysis results using a pre-deployed large model; embedding the visualization charts into the text conclusions and outputting them for display.
[0074] Optionally, the generation module 14 is also used to: select a final chart template from a preset template library based on structural features and analysis objectives when the user does not input a chart type preference; embed the data into the final chart template, generate a visual chart, and output it for display.
[0075] Specific limitations regarding the data analysis and chart generation device can be found in the limitations of the data analysis and chart generation methods described above, and will not be repeated here. Each module in the aforementioned data analysis and chart generation device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.
[0076] In one embodiment, a computer device is provided, the internal structure of which can be shown in the following diagram. Figure 4 As shown. The computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When the computer program is executed by the processor, it performs the following steps: Receive structured data files, natural language instructions, and chart type preferences uploaded by users; The system uses a pre-deployed large model to parse structured data files, extract the data and its structural features, and parses natural language instructions to obtain the analysis target. Verify the compatibility between chart type preferences, structural characteristics, and analysis objectives; If incompatible, a compatible chart type will be recommended to the user based on the preset template library, and the final chart template will be confirmed based on the user's feedback; If compatible, the final chart template will be selected from the preset template library based on the user's input chart type preference; The data is analyzed based on the analytical objectives to obtain the analytical results, and then visualization charts and analysis reports are generated based on the analytical results and the final chart template.
[0077] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor: Receive structured data files, natural language instructions, and chart type preferences uploaded by users; The system uses a pre-deployed large model to parse structured data files, extract the data and its structural features, and parses natural language instructions to obtain the analysis target. Verify the compatibility between chart type preferences, structural characteristics, and analysis objectives; If incompatible, a compatible chart type will be recommended to the user based on the preset template library, and the final chart template will be confirmed based on the user's feedback; If compatible, the final chart template will be selected from the preset template library based on the user's input chart type preference; The data is analyzed based on the analytical objectives to obtain the analytical results, and then visualization charts and analysis reports are generated based on the analytical results and the final chart template.
[0078] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0079] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0080] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0081] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A method of data analysis and chart generation, characterized by, include: Receive structured data files, natural language instructions, and chart type preferences uploaded by users; The structured data file is parsed using a pre-deployed large model to extract the data and its structural features, while the natural language instructions are parsed to obtain the analysis target. Verify the compatibility between the chart type preference and the structural features and the analysis objectives; If incompatible, a compatible chart type will be recommended to the user based on the preset template library, and the final chart template will be confirmed based on the user's feedback; If compatible, the final chart template is selected from the preset template library based on the chart type preference entered by the user. The data is analyzed based on the stated analytical objectives to obtain analytical results, and visualization charts and analytical reports are generated based on the analytical results and the final chart template.
2. The data analysis and chart generation method of claim 1, wherein, The process of parsing the structured data file using a pre-deployed large model to extract the data and its structural features includes: The structured data file is parsed using a pre-deployed large model to extract the structured data; The structured data is preprocessed, and the processed structured data is converted into standard JSON format data; Identify the discrete variables used for classification or grouping, the continuous numerical variables to be analyzed, and the time series identifiers in the standard JSON format data.
3. The data analysis and chart generation method of claim 1, wherein, The verification of the compatibility between the chart type preference and the structural features and the analysis objective includes: The system calls upon the set of chart applicability conditions and the conflict rule base in the preset rule base. The set of chart applicability conditions includes the mapping relationship between chart type and structural features and analysis objectives. The conflict rule base stores incompatible chart types and structural features and analysis objectives combinations and explanations. Based on the set of applicable conditions for the chart and the conflict rule base, verify whether the chart type preference is compatible with the structural features and the analysis objectives.
4. The data analysis and chart generation method of claim 1, wherein, The method of recommending compatible chart types to the user based on a preset template library includes: When the analysis target involves time series, a line chart or area chart is recommended to the user; When the analysis objective involves the relationship between parts and the whole, a pie chart or a donut chart is recommended to the user; When the analysis objective involves variable correlation, a scatter plot or heatmap is recommended to the user.
5. The data analysis and chart generation method of claim 1, wherein, The process of analyzing the data based on the analytical objective, obtaining analytical results, and generating a visualization chart based on the analytical results and the final chart template includes: When the analysis objective includes a prediction request, the data is input into a pre-trained prediction model to make a prediction and obtain the prediction result. The data and the prediction results are merged and rendered into the final chart template to generate a visual chart that includes the analysis results and the prediction results.
6. The data analysis and chart generation method of claim 1, wherein, The process of generating visualization charts and analysis reports based on the analysis results and the final chart template includes: Generate a visual chart based on the analysis results and the final chart template; Textual conclusions are generated based on the analysis results using a pre-deployed large model. The visualization chart is embedded into the text conclusion and then displayed.
7. The data analysis and chart generation method of claim 1, wherein, The method further includes: When the user does not input a chart type preference, the final chart template is selected from the preset template library based on the structural features and the analysis objectives. The data is embedded into the final chart template to generate a visual chart, which is then output and displayed.
8. A data analysis and chart generation apparatus, characterized by, include: The receiving module is used to receive structured data files, natural language instructions, and chart type preferences uploaded by users; The parsing module is used to parse the structured data file using a pre-deployed large model, extract the data and the structural features of the data, and parse the natural language instructions to obtain the analysis target; The verification module is used to verify the compatibility between the chart type preference and the structural features and the analysis objectives; The template matching module is used to recommend compatible chart types to the user based on a preset template library when incompatibility occurs, and to confirm the final chart template based on the user's feedback. Alternatively, when compatible, the final chart template can be selected from the preset template library based on the user's input chart type preference; The generation module is used to analyze the data based on the analysis objective, obtain analysis results, and generate visualization charts and analysis reports based on the analysis results and the final chart template.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the data analysis and chart generation method as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the data analysis and chart generation method as described in any one of claims 1 to 7.