Visual display method based on multi-source economic data mining

Through the visual display method of multi-source economic data mining, the integration and analysis problems of multi-source heterogeneous data on e-commerce platforms are solved, unified data processing and multi-dimensional relationship presentation are realized, intelligent trend prediction and abnormal identification are provided, and the accuracy and user experience of data analysis are improved.

CN120277146AInactive Publication Date: 2025-07-08ZHENGZHOU QINGMAO INFORMATION TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510351108.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional economic data visualization methods are difficult to effectively integrate and analyze multi-source heterogeneous data, especially on e-commerce platforms, where data sources are diverse and types are different, and it is difficult to present complex multi-dimensional relationships and associations.

Method used

Through visual presentation methods of multi-source economic data mining, including data cleaning, standardization, multi-dimensional data dimensionality reduction, feature importance sorting, time series analysis, deep learning model prediction, abnormal fluctuation recognition, economic knowledge graph interpretability analysis and interactive visualization, combined with a responsive layout engine, multi-terminal adaptation is achieved.

Benefits of technology

It realizes unified processing and analysis of multi-source heterogeneous data, improves the accuracy and credibility of data analysis, provides intelligent trend prediction and abnormal identification, enhances decision support, and ensures a good user experience in large data volumes and multi-device scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277146A_ABST
    Figure CN120277146A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of digital economy, in particular to a visual display method based on multi-source economic data mining. The method comprises the following steps: acquiring multi-source e-commerce economic data; performing data cleaning and standardization processing on the multi-source e-commerce economic data, and performing data quality evaluation and screening to obtain standardized e-commerce economic data; performing multi-dimensional data dimension reduction analysis according to the standardized e-commerce economic data, and performing feature importance sorting to obtain key economic index data; and performing time sequence analysis on the standardized e-commerce economic data according to the key economic index data, and performing sales trend and user behavior prediction based on key economic indexes through a plurality of preset deep learning models to obtain trend prediction data. According to the method, the interpretive data generation of the economic index relation network is realized in combination with the economic knowledge graph, and the interpretability of data analysis is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of digital economy, and particularly to a visualization display method based on multi-source economic data mining. Background Art

[0002] In today's digital economy era, e-commerce platforms have developed rapidly, and a large number of data sources have emerged continuously, including platform sales data, merchant financial data, user transaction behavior data, marketing data, and external market intelligence data, etc. These multi-source heterogeneous data contain a large amount of economic information and potential value. How to effectively mine, analyze and visually present it has become a key challenge in e-commerce enterprise decision optimization and market competition. However, when facing such a large amount of diverse data, traditional economic data visualization methods have the following technical difficulties: The sources of e-commerce data are diverse and of different types, including structured data (such as financial statements), semi-structured data (such as user reviews), unstructured data (such as pictures, videos), etc. The data formats and structures vary greatly, and it is difficult for traditional methods to achieve unified processing and integration of data; E-commerce economic data involves multiple dimensions (such as time, geographical location, product category, user behavior, etc.) and various complex relationships. Traditional linear or single-dimensional analysis methods are difficult to effectively present the complex associations and multi-dimensional relationships between data. Summary of the Invention

[0003] Based on this, it is necessary for the present invention to provide a visualization display method based on multi-source economic data mining to solve at least one of the above technical problems.

[0004] To achieve the above object, a visualization display method based on multi-source economic data mining includes the following steps:

[0005] Step S1: Obtain multi-source e-commerce economic data, including platform sales data, merchant financial data, user transaction behavior data, marketing data, and external market intelligence data; perform data cleaning and standardization processing on the multi-source e-commerce economic data, and conduct data quality assessment and screening to obtain standardized e-commerce economic data; perform multi-dimensional data dimensionality reduction analysis on the standardized e-commerce economic data, and conduct feature importance ranking to obtain key economic indicator data;

[0006] Step S2: Perform time series analysis on the standardized e-commerce economic data according to the key economic indicator data, and perform sales trend and user behavior prediction based on the key economic indicators through a variety of preset deep learning models to obtain trend prediction data; identify abnormal fluctuations in the trend prediction data, and conduct correlation analysis according to the key economic indicator data to construct an economic indicator relationship network;

[0007] Step S3: Generate explanatory data based on economic phenomena for the economic indicator relationship network according to the preset economic knowledge graph to obtain economic phenomenon explanatory data; select a visualization type based on the economic phenomenon explanatory data and the trend prediction data, and implement it through an interactive exploration tool to obtain interactive visualization data;

[0008] Step S4: Perform data pre-aggregation and progressive data loading optimization on the interactive visualization data to obtain performance-optimized data, and design a responsive layout engine based on the performance-optimized data and the interactive visualization data, so as to generate multi-terminal adaptation data for display devices of different sizes.

[0009] The present invention ensures the consistency of heterogeneous data through data cleaning and standardization, so that different types of data can be uniformly processed and analyzed, avoiding analysis misleading caused by data bias and incompleteness; data quality assessment and screening help to eliminate low-quality or noise data, improve the accuracy and credibility of data analysis; feature importance ranking can screen out the most valuable key economic indicators for the business in a huge data set, avoid information redundancy, and provide a basis for subsequent analysis and prediction. Time series analysis can identify trend characteristics such as seasonal sales fluctuations and the impact of holidays on sales in e-commerce platforms, providing a reference for strategic planning; the introduction of deep learning models can process a large amount of complex data and enhance the accuracy of predictions, especially in the face of multi-dimensional and multi-variable scenarios; abnormal fluctuation identification can help timely discover sales anomalies, inventory problems or market changes, reduce business risks and enhance the ability to respond to emergencies; the correlation analysis of economic indicators helps to reveal the inherent connection between key indicators by constructing a relationship network of economic indicators, such as the linkage effect between user behavior and marketing activities. The introduction of knowledge graphs combines isolated data with economic theories and actual business scenarios, allowing decision makers to intuitively understand the reasons behind complex economic phenomena; explanatory data generation provides reasonable explanations for sales trends and changes in user behavior, helping managers find potential reasons and opportunities in business operations; visualization type selection ensures the intuitiveness and rationality of data display, and can dynamically select the most appropriate visualization method according to data characteristics, improving data readability and decision effectiveness; interactive visualization further enhances users' ability to explore data, allowing managers to dig deep into data through interactive operations and reveal hidden trends or problems. Data pre-aggregation processes some common analysis tasks in advance, reducing the pressure of real-time computing and ensuring that the system can still provide fast response under high concurrency. Progressive data loading improves the speed of front-end loading while ensuring data integrity, optimizing user experience, and is especially suitable for visualization scenarios with large amounts of data. The responsive layout engine ensures seamless adaptation of multiple devices and multiple terminals, allowing visualization to be adaptively adjusted on display devices of different sizes, improving the flexibility of visualization and cross-device availability. Through the mutual cooperation of the above steps, the present invention provides a visualization display method based on multi-source economic data mining, which solves the technical problems of heterogeneous e-commerce data sources, huge data volume, and multi-dimensional complex associations. The method has the following comprehensive effects: effectively integrate multi-source heterogeneous data to improve the consistency and reliability of data analysis; provide intelligent trend prediction and anomaly identification to help companies cope with market fluctuations and business changes; through explanatory analysis based on economic knowledge graphs, the depth of data interpretation is improved, providing managers with more explanatory decision support; provide multi-terminal adaptation High-performance visualization solutions to ensure a good user experience in large data volumes and multi-device scenarios.The invention provides strong support for e-commerce enterprises in decision-making optimization and market competition in the big data environment, and has broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Other features, objects, and advantages of the present invention will become more apparent by reading the following detailed description of non-limiting embodiments with reference to the accompanying drawings:

[0011] Figure 1 is a schematic flowchart of the steps of the visualization display method based on multi-source economic data mining of the present invention;

[0012] Figure 2 is Figure 1 a detailed flowchart of step S1 in

[0013] Figure 3 is Figure 1 a detailed flowchart of step S2 in DETAILED DESCRIPTION OF THE EMBODIMENTS

[0014] The technical method of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of the present invention.

[0015] In addition, the accompanying drawings are only schematic diagrams of the present invention and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. The functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor methods and / or microcontroller methods.

[0016] It should be understood that although terms such as "first", "second", etc. may be used herein to describe various units, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, the first unit may be referred to as the second unit, and similarly the second unit may be referred to as the first unit. The term "and / or" used herein includes any and all combinations of one or more of the listed associated items.

[0017] To achieve the above object, please refer to Figures 1 to 3 , the present invention provides a visualization display method based on multi-source economic data mining, and the method includes the following steps:

[0018] Step S1: Obtain multi-source e-commerce economic data, including platform sales data, merchant financial data, user transaction behavior data, marketing data, and external market intelligence data; perform data cleaning and standardization on the multi-source e-commerce economic data, and conduct data quality assessment and screening to obtain standardized e-commerce economic data; perform multi-dimensional data dimensionality reduction analysis on the standardized e-commerce economic data, and conduct feature importance ranking to obtain key economic indicator data;

[0019] Step S2: Perform time series analysis on the standardized e-commerce economic data based on the key economic indicator data, and use a variety of preset deep learning models to predict sales trends and user behavior based on the key economic indicators, so as to obtain trend prediction data; identify abnormal fluctuations in the trend prediction data, and conduct correlation analysis based on the key economic indicator data, so as to construct an economic indicator relationship network;

[0020] Step S3: Generate explanatory data based on economic phenomena for the economic indicator relationship network according to the preset economic knowledge graph to obtain economic phenomenon explanatory data; select a visualization type based on the economic phenomenon explanatory data and the trend prediction data, and implement it through an interactive exploration tool to obtain interactive visualization data;

[0021] Step S4: Perform data pre-aggregation and progressive data loading optimization on the interactive visualization data to obtain performance-optimized data, and design a responsive layout engine based on the performance-optimized data and the interactive visualization data, so as to generate multi-terminal adaptation data for display devices of different sizes.

[0022] In the embodiment of the present invention, refer to Figure 1 As described above, it is a schematic diagram of the step process of a visualization display method based on multi-source economic data mining of the present invention. In this example, the visualization display method based on multi-source economic data mining includes the following steps:

[0023] Step S1: Obtain multi-source e-commerce economic data, including platform sales data, merchant financial data, user transaction behavior data, marketing data, and external market intelligence data; perform data cleaning and standardization on the multi-source e-commerce economic data, and conduct data quality assessment and screening to obtain standardized e-commerce economic data; perform multi-dimensional data dimensionality reduction analysis on the standardized e-commerce economic data, and conduct feature importance ranking to obtain key economic indicator data;

[0024] In an embodiment of the present invention, multi-source data from an e-commerce platform is obtained, including structured financial data, semi-structured user review data, and unstructured image and video data. Automated data cleaning tools are used for data cleaning, and redundant, duplicate, and incorrect data are removed through rule matching. For example, in financial data, outliers and null values are removed; in user reviews, natural language processing technology (NLP) is used to denoise and clean the text. Then, standardization algorithms are used to uniformly process data from different sources, unifying numerical data into the same measurement unit and encoding text data at the same time. Data quality assessment is carried out through data consistency verification and missing value filling mechanisms, and finally high-quality data is selected. After that, dimensionality reduction algorithms such as principal component analysis (PCA) are used to perform dimensionality reduction analysis on high-dimensional economic data, removing redundant dimensions, and the feature importance of the dimensionality-reduced data is ranked through feature selection methods based on information gain or entropy, so as to extract key economic indicator data, such as product sales volume, user purchase frequency, and average order amount.

[0025] Step S2: Perform time series analysis on the standardized e-commerce economic data based on the key economic indicator data, and predict the sales trend and user behavior based on the key economic indicators through a variety of preset deep learning models, so as to obtain trend prediction data; identify abnormal fluctuations in the trend prediction data, and perform correlation analysis based on the key economic indicator data, so as to construct an economic indicator relationship network;

[0026] In an embodiment of the present invention, historical sales data and user behavior data are selected as analysis objects, and traditional time series models such as moving average and ARIMA are used to analyze the trends, periodicity, and seasonal changes therein. Then, preset deep learning models (such as long short-term memory network LSTM and recurrent neural network RNN) are applied to the key economic indicator data for sales trend and user behavior prediction. When training the model, the data is divided into a training set and a validation set, and the gradient descent optimization algorithm is used to iteratively optimize the model parameters. After the model training is completed, the trained model is used to predict the sales and user behavior trends in the future for a period of time, and trend prediction data is generated. After that, based on the set fluctuation threshold, abnormal fluctuations in the prediction results are identified, and Pearson correlation coefficient or Granger causality analysis is used to perform correlation analysis on the key economic indicator data, and a relationship network between economic indicators is constructed.

[0027] Step S3: Generate explanatory data based on economic phenomena for the economic indicator relationship network according to a preset economic knowledge graph to obtain economic phenomenon explanation data; select a visualization type according to the economic phenomenon explanation data and the trend prediction data, and implement it through an interactive exploration tool to obtain interactive visualization data;

[0028] In the embodiment of the present invention, an economic knowledge graph constructed in advance is used to perform interpretive analysis on the economic indicator relationship network in step S2. The economic knowledge graph integrates multi-source economic theories and industry rules to establish a mapping relationship between key economic indicators and economic phenomena in actual business scenarios. For example, a decline in sales is associated with negative sentiment in user reviews. Through this mapping relationship, interpretive data is generated for each key node in the relationship network. For example, it is identified that the reason for the decline in sales is an increase in negative reviews of a certain type of product. Based on the generated economic phenomenon interpretation data and combined with trend prediction data, automatic selection of the visualization type is performed. For time series data, the system selects a line chart; for economic phenomenon relationships, a force-directed graph is used for display. Finally, through an interactive visualization tool such as D3.js or Tableau, interactive visualization data is generated, allowing users to deeply explore the data through operations such as clicking and dragging.

[0029] Step S4: Perform data pre-aggregation and progressive data loading optimization on the interactive visualization data to obtain performance-optimized data, and design a responsive layout engine based on the performance-optimized data and the interactive visualization data, thereby generating multi-terminal adaptation data for display devices of different sizes.

[0030] In the embodiment of the present invention, data pre-aggregation processing is first performed on the interactive visualization data. For example, sales data is pre-grouped and summarized by time and product category to reduce the computational amount of real-time queries. Progressive data loading is implemented through a demand-based chunk loading algorithm, which preferentially loads the part of the data currently browsed by the user instead of all the data to optimize the loading speed of the front-end page and the user experience. Then, based on the resolutions and screen sizes of different display devices, a responsive layout engine is designed; this engine uses CSS media query technology and JavaScript dynamic layout adjustment to automatically adjust the display effect on different devices, thereby generating multi-terminal adaptation data, enabling the visualization content to be smoothly displayed on different devices such as mobile phones, tablets, and computers, and enhancing the multi-terminal usage experience of users.

[0031] The present invention ensures the consistency of heterogeneous data through data cleaning and standardization, so that different types of data can be uniformly processed and analyzed, avoiding analysis misleading caused by data bias and incompleteness; data quality assessment and screening help to eliminate low-quality or noise data, improve the accuracy and credibility of data analysis; feature importance ranking can screen out the most valuable key economic indicators for the business in a huge data set, avoid information redundancy, and provide a basis for subsequent analysis and prediction. Time series analysis can identify trend characteristics such as seasonal sales fluctuations and the impact of holidays on sales in e-commerce platforms, providing a reference for strategic planning; the introduction of deep learning models can process a large amount of complex data and enhance the accuracy of predictions, especially in the face of multi-dimensional and multi-variable scenarios; abnormal fluctuation identification can help timely discover sales anomalies, inventory problems or market changes, reduce business risks and enhance the ability to respond to emergencies; the correlation analysis of economic indicators helps to reveal the inherent connection between key indicators by constructing a relationship network of economic indicators, such as the linkage effect between user behavior and marketing activities. The introduction of knowledge graphs combines isolated data with economic theories and actual business scenarios, allowing decision makers to intuitively understand the reasons behind complex economic phenomena; explanatory data generation provides reasonable explanations for sales trends and changes in user behavior, helping managers find potential reasons and opportunities in business operations; visualization type selection ensures the intuitiveness and rationality of data display, and can dynamically select the most appropriate visualization method according to data characteristics, improving data readability and decision effectiveness; interactive visualization further enhances users' ability to explore data, allowing managers to dig deep into data through interactive operations and reveal hidden trends or problems. Data pre-aggregation processes some common analysis tasks in advance, reducing the pressure of real-time computing and ensuring that the system can still provide fast response under high concurrency. Progressive data loading improves the speed of front-end loading while ensuring data integrity, optimizing user experience, and is especially suitable for visualization scenarios with large amounts of data. The responsive layout engine ensures seamless adaptation of multiple devices and multiple terminals, allowing visualization to be adaptively adjusted on display devices of different sizes, improving the flexibility of visualization and cross-device availability. Through the mutual cooperation of the above steps, the present invention provides a visualization display method based on multi-source economic data mining, which solves the technical problems of heterogeneous e-commerce data sources, huge data volume, and multi-dimensional complex associations. The method has the following comprehensive effects: effectively integrate multi-source heterogeneous data to improve the consistency and reliability of data analysis; provide intelligent trend prediction and anomaly identification to help companies cope with market fluctuations and business changes; through explanatory analysis based on economic knowledge graphs, the depth of data interpretation is improved, providing managers with more explanatory decision support; provide multi-terminal adaptation High-performance visualization solutions to ensure a good user experience in large data volumes and multi-device scenarios.The invention provides strong support for e-commerce enterprises in decision-making optimization and market competition in the big data environment, and has broad application prospects.

[0032] Preferably, step S1 includes the following steps:

[0033] Step S11: Obtain multi-source e-commerce economic data, including platform sales data, merchant financial data, user transaction behavior data, marketing data, and external market intelligence data;

[0034] Step S12: Perform format unification processing on the multi-source e-commerce economic data based on the date format and numerical unit to obtain format-unified e-commerce data;

[0035] Step S13: Perform missing value processing on the format-unified e-commerce data based on moving average and median filling, and perform outlier processing through the Z-score method to obtain preprocessed e-commerce data;

[0036] Step S14: Perform data quality evaluation and screening on the preprocessed e-commerce data to obtain standardized e-commerce economic data;

[0037] Step S15: Perform data alignment and merging on the standardized e-commerce economic data from different sources based on time and region to construct a multi-dimensional data cube;

[0038] Step S16: Use the t-distributed stochastic neighbor embedding algorithm to perform data dimensionality reduction analysis on the multi-dimensional data cube, and perform principal component selection according to the preset explained variance ratio threshold to obtain dimensionality-reduced e-commerce data;

[0039] Step S17: Perform feature importance ranking on the dimensionality-reduced e-commerce data to obtain key economic indicator data.

[0040] As an embodiment of the present invention, refer to Figure 2 As shown, it is Figure 1 The detailed step flow diagram of step S1 in

[0041] Step S11: Obtain multi-source e-commerce economic data, including platform sales data, merchant financial data, user transaction behavior data, marketing data, and external market intelligence data;

[0042] In the system of the embodiments of the present invention, economic data of an e-commerce platform is obtained from multiple data sources, including internal sales data of the platform (such as daily sales volume, product sales volume, etc.), financial data provided by merchants (such as income statements, profit margins, etc.), user transaction behavior data (such as click-through rate, conversion rate, shopping cart addition behavior, etc.), marketing data (such as advertisement clicks, promotion activity effects, etc.), and external market intelligence data (such as industry reports, competitor analysis, macroeconomic indicators, etc.). To obtain this data, the system integrates multiple API interfaces or data scraping tools to regularly or real-time collect data from different data sources and import the data into a unified database system for subsequent processing and use.

[0043] Step S12: Perform format unification processing on the multi-source e-commerce economic data based on the date format and numerical unit to obtain format-unified e-commerce data;

[0044] In the embodiments of the present invention, for the e-commerce data obtained from multiple sources, the format is first unified. For example, during the processing of the date format, the system will convert different date formats (such as "YYYY / MM / DD" and "DD-MM-YYYY", etc.) into the standard ISO date format "YYYY-MM-DD". In addition, for numerical data, the system will standardize different units. For example, the price will be uniformly converted from multiple currency units to the local currency, or the product weight will be uniformly converted from kilograms to grams. Through this unification processing based on the date format and numerical unit, it is ensured that all data has a consistent format in subsequent analysis, thus avoiding processing errors caused by format differences.

[0045] Step S13: Perform missing value processing on the format-unified e-commerce data based on moving average and median filling, and perform outlier processing through the Z-score method to obtain preprocessed e-commerce data;

[0046] In the embodiments of the present invention, for the format-unified e-commerce data, the system performs missing value processing. First, for time series data, the moving average method is used to fill in the missing values, that is, the average value of the data at adjacent time points is used to replace the missing values, ensuring the continuity of the time series. For non-time series data, such as product price or sales volume, median filling is used, that is, the median of this feature is used to replace the missing values to avoid the influence of outliers on the filled data. Outlier processing adopts the Z-score method. By calculating the standard score of each data point, the data points exceeding the set threshold (such as more than 3 standard deviations) are regarded as outliers and excluded or adjusted to ensure the rationality and accuracy of the data. Finally, preprocessed e-commerce data after missing value and outlier processing is obtained.

[0047] Step S14: Evaluate and screen the preprocessed e-commerce data to obtain standardized e-commerce economic data;

[0048] In the system of the embodiment of the present invention, data quality evaluation is performed on the preprocessed e-commerce data. The quality evaluation includes multiple dimensions such as consistency check, integrity check, and accuracy evaluation. For example, consistency is evaluated by checking whether the sales volume data of the same product on different sales platforms is consistent, integrity is evaluated by detecting whether there are missing key fields (such as date, product ID, etc.), and in addition, the accuracy can be judged by comparing with historical data or external reference data. For data with poor quality, the system will automatically screen it out or relabel it as low-quality data and not participate in subsequent analysis. Finally, the screened high-quality data is defined as standardized e-commerce economic data for subsequent analysis steps.

[0049] Step S15: Align and merge the standardized e-commerce economic data from different sources based on time and region to construct a multi-dimensional data cube;

[0050] In the system of the embodiment of the present invention, for the standardized e-commerce economic data, the data is aligned and merged according to two main dimensions of time and region; for example, the system will align the sales data in different data sources according to the date to ensure that the data from different sources has the same time granularity. At the same time, based on the region dimension, the system will merge the sales situations, user behavior data, and marketing data in different regions to construct a multi-dimensional data structure. For this purpose, the system stores and organizes the data in the form of a multi-dimensional data cube, which allows querying and analysis in multiple dimensions such as time, region, and product category. Through alignment and merging, ensure that all data dimensions are consistent, laying a foundation for subsequent data dimensionality reduction and analysis.

[0051] Step S16: Use the t-distributed stochastic neighbor embedding algorithm to perform data dimensionality reduction analysis on the multi-dimensional data cube, and select the principal components according to the preset threshold of the proportion of explained variance to obtain the dimensionality-reduced e-commerce data;

[0052] In the embodiments of the present invention, in order to reduce the data dimension, improve the calculation efficiency and interpretability, the system adopts the t-distributed Stochastic Neighbor Embedding (t-SNE) algorithm to perform dimensionality reduction processing on the multi-dimensional data cube. t-SNE is a commonly used non-linear dimensionality reduction algorithm that can effectively capture the local structure and complex relationships in high-dimensional data. First, the system performs t-SNE processing on the input multi-dimensional data to generate a two-dimensional or three-dimensional embedding representation. During this process, the system will select the number of principal components according to the preset threshold of the proportion of explained variance to ensure that the data after dimensionality reduction retains sufficient interpretability. The data after dimensionality reduction will contain the most representative information and discard redundant or noisy data, thereby improving the efficiency of subsequent analysis. Finally, the system generates the e-commerce data after dimensionality reduction.

[0053] Step S17: Rank the importance of features for the e-commerce data after dimensionality reduction to obtain the key economic indicator data.

[0054] The system in the embodiments of the present invention ranks the importance of features for the e-commerce data after dimensionality reduction; First, machine learning algorithms such as Random Forest or XGBoost are used to analyze the importance of features by measuring the contribution degree of each feature during the model training process. Secondly, the system will rank each feature according to the importance score of the feature, and preferentially retain the features that have the greatest impact on the target variable (such as sales volume or user behavior prediction). During the feature ranking process, business experience and domain knowledge will also be referred to, and combined with the results of machine learning algorithms, the key economic indicator data that is most valuable to the business will be screened out, such as user purchase frequency, advertising click-through rate, and sales trends in specific regions. These key indicators will provide strong support for the decision-making of e-commerce enterprises.

[0055] The present invention obtains e-commerce economic data from multiple data sources (such as platform sales data, merchant financial data, user transaction behavior data, marketing data, and external market intelligence data), ensuring the diversity and comprehensiveness of the data; integrating multi-source data can provide a more complete business panoramic view, enabling the analysis to cover a wider range of business activities and market dynamics, thereby laying a foundation for more accurate decision-making support. The multi-source data is unified in date format and numerical unit to ensure that the data is analyzed under the same standard. This step effectively eliminates the format differences between the data, improves the compatibility and readability of the data, reduces the processing and analysis complexity caused by inconsistent formats, and makes the subsequent data processing more efficient. The missing values are filled by moving average and median to ensure the integrity of the data and avoid analysis biases caused by missing data. The Z-score method is used to process outliers, which can effectively identify and remove abnormal situations in the data, improve the accuracy and credibility of the data, and provide a more stable data basis for subsequent data analysis and modeling. Through data quality assessment and screening, it is ensured that only high-quality data enters the subsequent analysis stage. This step is comprehensively evaluated from multiple dimensions (such as integrity, consistency, timeliness, accuracy, etc.), and standardized e-commerce economic data is screened out, further enhancing the reliability of the data and the credibility of the analysis results. The data from different sources is aligned and merged according to time and region to ensure that the data is comparable in the same time and space dimensions. Constructing a multi-dimensional data cube enables the data to be analyzed and presented from multiple angles and levels, helps to identify complex cross-dimensional associations, and supports deeper business insights. By performing dimensionality reduction analysis on the multi-dimensional data cube through the t-distributed stochastic neighbor embedding (t-SNE) algorithm, the complexity of the data can be reduced while retaining the main features and structure of the data. According to the preset threshold of the proportion of explained variance, the principal components are screened to ensure that only the variables most important for the analysis are retained, reducing noise and redundancy, and improving the analysis efficiency and the interpretability of the results. Ranking the importance of features for the dimensionality-reduced data can identify the key economic indicators that have the greatest impact on the target variables (such as sales volume, user activity, etc.). This process helps to simplify the model and concentrate resources for efficient decision-making support, and also improves the interpretability and transparency of the analysis model. These steps form a complete data preprocessing and analysis process. Through a series of operations such as data acquisition, cleaning, quality assessment, dimensionality reduction analysis, and feature ranking, the accuracy, consistency, and reliability of the data are ensured to the greatest extent, laying a solid foundation for subsequent deep learning modeling and visualization display. This method not only improves the efficiency of data processing but also enhances the ability to analyze complex business problems, helping enterprises to make data-based decisions faster.

[0056] Preferably, step S14 includes the following steps:

[0057] Step S141: Calculate the field filling rates of product information, transaction records, and user behavior logs for the preprocessed e-commerce data, so as to obtain integrity detection data;

[0058] In the embodiment of the present invention, for the preprocessed e-commerce data, the system calculates the filling rates of the main fields such as product information, transaction records, and user behavior logs. First, the system analyzes the non-null rate of each field. For example, for core fields such as product name, order number, transaction amount, and user ID, it checks whether there are null values or missing data. By counting the filling rate of each field, the system generates integrity detection data. The higher the filling rate, the better the integrity of the data. For fields with a relatively high missing rate, the system will record them for subsequent quality evaluation.

[0059] Step S142: Use the preset e-commerce business rules to perform consistency checks on the preprocessed e-commerce data, so as to obtain consistency scoring data, where the consistency checks include the matching of the order amount and the product price, and the verification of the logical relationship between the inventory quantity and the sales volume;

[0060] In the embodiment of the present invention, the system performs consistency checks on the preprocessed e-commerce data based on the e-commerce business rules. First, the system checks the matching of the order amount and the product price to verify whether the total order amount correctly reflects the product of the purchased quantity and the unit price. Secondly, the system checks the logical relationship between the inventory quantity and the sales volume. For example, the inventory quantity should not be negative, and the sales volume should not exceed the inventory. Through these consistency checks, the system generates consistency scoring data. The consistency score reflects the logical rationality of the data through a scoring mechanism. The higher the score, the more the data conforms to the business rules.

[0061] Step S143: Perform real-time monitoring of the update frequency heat of the preprocessed e-commerce data, and calculate the synchronization degree of real-time transactions, so as to obtain timeliness evaluation data;

[0062] In the embodiment of the present invention, the system performs real-time monitoring of the update frequency heat of the e-commerce data. Specifically, the system analyzes the update time of the transaction data to monitor whether the update frequency of the data meets the real-time requirements of the e-commerce platform. For example, whether the transaction records within an hour can be synchronized in a timely manner. In addition, the system also calculates the synchronization degree of real-time transactions. By comparing the real-time data with the actual transaction situation, it calculates the deviation degree of the data timeliness. Based on these monitoring results, the system generates timeliness evaluation data to measure the real-time nature and the rationality of the update frequency of the data.

[0063] Step S144: Calculate the deviation rates of the conversion rate and the average order value of the preprocessed e-commerce data by comparing the historical data of the data source of the preprocessed e-commerce data pre-acquired with the industry benchmark, so as to obtain accuracy scoring data;

[0064] The system of the embodiment of the present invention calculates the deviation rate of the conversion rate and the average order value of the e-commerce data by comparing the pre-processed e-commerce data with the historical data and the industry benchmark data. The system first analyzes the historical sales data, calculates the conversion rate and the average order value in a specific time period, and then compares it with the current data. At the same time, the system refers to the industry benchmark data, calculates the deviation rate, and evaluates whether the performance of the e-commerce platform is abnormal. Through these comparative analyses, the system generates accuracy scoring data to measure the accuracy and rationality of the current data.

[0065] Step S145: using an outlier detection method to identify false transactions and fake order behaviors in the pre-processed e-commerce data, and calculating the abnormal proportion, so as to obtain the data credibility;

[0066] The system of the embodiment of the present invention uses an outlier detection method to analyze the pre-processed e-commerce data to identify false transactions and order-brushing behaviors. The system identifies possible order-brushing behaviors by analyzing abnormal transaction records, such as the same user purchasing a large number of the same products in a short period of time, or the order amount is far higher than the normal range. The outlier detection methods used include distance-based K-means clustering or density peak detection algorithms. After identifying abnormal transactions, the system calculates the abnormal proportion and generates data credibility based on the abnormal proportion, which is used to evaluate the proportion of false transactions in the data and the credibility of the overall data.

[0067] Step S146: performing a comprehensive evaluation process based on a weighted average method on the pre-processed e-commerce data according to the integrity detection data, consistency score data, timeliness evaluation data, accuracy score data, and data credibility, thereby obtaining comprehensive quality score data;

[0068] The system of the embodiment of the present invention performs a weighted average comprehensive evaluation on the above-mentioned scoring data. Specifically, the system assigns weights to the integrity detection data, consistency scoring data, timeliness evaluation data, accuracy scoring data, and data credibility, and calculates the comprehensive quality score data by weighted average method according to the importance of each score. For example, timeliness and consistency may be given higher weights, while accuracy and credibility may be given slightly lower weights. In this way, the system derives a comprehensive quality score for each piece of data to evaluate the quality of the overall data.

[0069] Step S147: The pre-processed e-commerce data is graded and screened according to the comprehensive quality score data, and low-quality data is eliminated to obtain standardized e-commerce economic data.

[0070] In the system of the embodiment of the present invention, the preprocessed e-commerce data is classified and screened according to the comprehensive quality score. The system will perform hierarchical processing on the data according to the comprehensive score, and divide it into high-quality, medium-quality, and low-quality data. The low-quality data below the set threshold will be eliminated to avoid affecting subsequent analysis and decision-making. The high-quality data will be retained and further processed to finally form standardized e-commerce economic data for subsequent in-depth analysis and visual display.

[0071] By calculating the field filling rates of commodity information, transaction records, and user behavior logs, the present invention can effectively evaluate the integrity of data; ensure that the key fields in the data are fully filled, avoid analysis biases and decision-making errors caused by missing data, thereby improving the integrity and reliability of the data set, and providing a more solid foundation for subsequent data analysis and modeling. Use preset e-commerce business rules to conduct consistency checks on the data to ensure that the data is logically consistent; for example, the order amount and the commodity price should match, and there should be a reasonable logical relationship between the inventory quantity and the sales volume. This process can promptly detect and correct logical errors or input errors in the data, enhance the internal consistency of the data, and improve the data quality and the accuracy of analysis. By monitoring the data update frequency and calculating the synchronization degree of real-time transactions, the timeliness of the data can be evaluated. This is particularly important for e-commerce data because timeliness directly affects the real-time reflection of market dynamics and the timeliness of decision-making. Efficient timeliness evaluation ensures that the data is up-to-date, helps accurately capture market trends, and quickly respond to changes. By comparing the preprocessed e-commerce data with the historical data of the data source and industry benchmarks, and calculating the deviation rates of the conversion rate and the average customer price, it helps to evaluate the accuracy of the data. This step can identify potential biases and errors, ensure the data is more accurate, reduce the risks brought by incorrect data, and improve the reliability of prediction and decision-making. Use outlier detection methods to identify false transactions and brush order behaviors, and calculate the abnormal proportion, which helps filter out untrustworthy data. This can improve the credibility of the data set, reduce the risk of incorrect analysis and decision-making caused by untrue data, and ensure that the model and analysis are based on a more real and reliable data foundation. According to the evaluation results of integrity, consistency, timeliness, accuracy, and credibility, use the weighted average method for comprehensive evaluation to obtain the comprehensive quality score data. This multi-dimensional comprehensive evaluation provides a global quality overview, helps identify high-quality data and low-quality data, ensures that the subsequent analysis process is based on high-quality data, and improves the effectiveness and accuracy of analysis and model construction. Conduct hierarchical screening according to the comprehensive quality score, eliminate low-quality data, and ensure that only high-quality standardized e-commerce economic data is retained. By removing the noise and low-quality parts in the data, the overall quality of the data set is improved, the interference factors in model training and analysis are reduced, and the effectiveness and precision of data mining and analysis results are enhanced. These steps ensure that before the e-commerce data enters the subsequent processing and analysis stage, it has undergone comprehensive data quality control and optimization, including checks and evaluations of integrity, consistency, timeliness, accuracy, and credibility. This comprehensive data quality management process helps improve the overall quality and reliability of the data, makes subsequent data mining, analysis, and visualization more accurate and meaningful, and at the same time reduces the decision-making risks brought by data problems.

[0072] Preferably, step S17 includes the following steps:

[0073] Step S171: Conduct a correlation analysis on the dimensionality-reduced e-commerce data based on the Pearson correlation coefficient, group the highly correlated features using the hierarchical clustering algorithm, and perform feature selection on each group to obtain a preliminary screening feature set;

[0074] In the system of the embodiment of the present invention, first, a correlation analysis based on the Pearson correlation coefficient is performed on the dimensionality-reduced e-commerce data to quantify the linear relationship between each feature. The system calculates the Pearson correlation coefficient of each feature pair, and the value of the correlation coefficient is between -1 and 1. The closer it is to 1 or -1, the stronger the linear correlation between the two features. Next, the system uses the hierarchical clustering algorithm to group the highly correlated features. By classifying the strongly correlated features into the same group, feature redundancy can be reduced. Finally, the system performs feature selection within each feature group, retains the features that are most important for the business objective, and forms a preliminary screening feature set. This operation ensures the efficiency and accuracy of subsequent analysis by maximizing information retention and minimizing feature redundancy.

[0075] Step S172: Calculate the average decrease in impurity of each feature pair for the target variable based on the preliminary screening feature set to obtain an initial ranking of feature importance, where the target variable specifically includes sales volume and user activity;

[0076] In the system of the embodiment of the present invention, the contribution degree of each feature to the target variable is calculated according to the preliminary screening feature set. The specific method is to use the decision tree or random forest model and measure the importance of the feature by the mean decrease in impurity (MDI). The target variables include sales volume and user activity. When training the model, the system measures the impact of each feature on the target variable according to its importance in the splitting node. The higher the MDI score, the stronger the explanatory ability of the feature for the target variable. The system performs a preliminary ranking of all features according to MDI to obtain an initial ranking of feature importance. This step helps to identify the features that have the greatest impact on sales and user behavior, laying a foundation for subsequent feature selection.

[0077] Step S173: Select the top N features whose cumulative importance reaches the threshold from the initial ranking of feature importance according to the preset cumulative importance threshold to obtain candidate key indicator data;

[0078] The system of the embodiment of the present invention selects the most explanatory feature from the initial ranking of feature importance according to a preset cumulative importance threshold. The cumulative importance threshold is usually set to 90% or 95% to ensure that the retained features can explain most of the changes in the target variable. The system sorts the features by importance and gradually accumulates the importance score of each feature until the cumulative score reaches the preset threshold. The top N features that reach the threshold are selected as candidate key indicator data. Through this method, the system ensures that the number of selected features is within a controllable range, while retaining features that have a high explanatory power for the target variable, thereby improving model performance and computational efficiency.

[0079] Step S174: Perform a multicollinearity test on the candidate key indicator data, calculate the variance inflation factor, and eliminate the features whose variance inflation factor exceeds the preset threshold, so as to obtain the key economic indicator data.

[0080] The system of the embodiment of the present invention performs a multicollinearity test on the candidate key indicator data. Collinearity will affect the stability and explanatory power of the model, so it is necessary to detect the linear dependence between features. The system calculates the variance inflation factor (VIF) of each feature. The higher the VIF value, the stronger the collinearity between the features. Generally, if the VIF of a feature exceeds a preset threshold (such as 10), it indicates that the feature has strong collinearity with other features and should be eliminated. According to this standard, the system eliminates features with high VIF, retains features with no collinearity or low collinearity, and finally obtains key economic indicator data. This process ensures the stability of the final model while improving the explanatory power and prediction accuracy of the model.

[0081] Through the correlation analysis based on the Pearson correlation coefficient, the present invention can identify the mutual relationships between features in the dimensionality-reduced e-commerce data, thereby determining which features are highly correlated; using the hierarchical clustering algorithm to group the highly correlated features helps to reduce redundant features and simplify the data structure. By feature selection to retain the most representative features, the dimensionality of the feature space can be reduced, the training speed and performance of the model can be improved, at the same time, the problem of multicollinearity can be avoided, and the stability and interpretability of the model can be enhanced. Calculating the average decrease in impurity of each feature with respect to the target variable (such as sales and user activity) can evaluate the importance of each feature in the model prediction. This method ranks based on the contribution of features to the model performance, which helps to identify and screen out those features that have a significant impact on the target variable, optimize the feature selection process, and improve the prediction accuracy and efficiency of the model. According to the preset cumulative importance threshold, selecting the top N features whose cumulative importance reaches the threshold from the initial ranking of feature importance can ensure that the final candidate key indicator dataset only contains the features that have the most significant impact on the target variable. Doing so can not only further reduce the number of features and improve the computational efficiency of the model, but also ensure that the model only uses the information that is most valuable for result prediction, avoid overfitting and improve the generalization ability of the model. Conducting a multicollinearity test on the candidate key indicator data, calculating the variance inflation factor (VIF), and removing the features that exceed the preset threshold can effectively eliminate the multicollinearity problem existing in the dataset. Multicollinearity will lead to unstable estimation in the regression model, thus affecting the interpretability and prediction performance of the model. By removing the features with VIF exceeding the threshold, it can be ensured that the features in the model are relatively independent, and the robustness and reliability of the model are improved. The effect of these steps is to ensure that only the key features that are most important for the target variable and do not have multicollinearity are retained in the e-commerce data through a systematic feature selection and optimization process; this process helps to improve the quality and efficiency of data analysis and model construction, reduce the computational burden, and at the same time enhance the interpretability and prediction ability of the model. The final result is to obtain a concise but representative feature set, providing strong support for subsequent analysis, modeling, and decision-making.

[0082] Preferably, step S2 includes the following steps:

[0083] Step S21: Extract relevant time series from the standardized e-commerce economic data according to the key economic indicator data, thereby obtaining the key indicator time series data;

[0084] Step S22: Perform seasonal decomposition on the key indicator time series data, thereby obtaining the trend component data, seasonal component data, and residual component data;

[0085] Step S23: Use a variety of preset deep learning models to perform sales trend and user behavior prediction based on key economic indicators according to the trend component data, seasonal component data, and residual component data, so as to obtain trend prediction data;

[0086] Step S24: Identify abnormal fluctuations in the trend prediction data and perform correlation analysis based on the key economic indicator data, thereby constructing an economic indicator relationship network.

[0087] As an embodiment of the present invention, referring to Figure 3 shown, for Figure 1 the detailed step flow diagram of step S2 in

[0088] Step S21: Extract relevant time series from the standardized e-commerce economic data according to the key economic indicator data, so as to obtain key indicator time series data;

[0089] The system of the embodiment of the present invention extracts relevant time series from the standardized e-commerce economic data according to the key economic indicator data. Specifically, the system will identify key economic indicators (such as sales volume, user activity, conversion rate, etc.), and according to the records of these indicators in different periods, extract their time series data. These time series data may be continuous records in units of days, weeks, months, etc. The system ensures that the extracted time series can accurately reflect the change of the indicator over time by matching timestamps, and finally forms key indicator time series data. This operation provides the basic data source for subsequent trend and seasonal analysis.

[0090] Step S22: Perform seasonal decomposition on the key indicator time series data, so as to obtain trend component data, seasonal component data, and residual component data;

[0091] The system of the embodiment of the present invention performs seasonal decomposition on the key indicator time series data. Seasonal decomposition is an important step in time series analysis, usually using the STL decomposition (Seasonal-Trend-Residual decomposition) method or other similar algorithms. The system decomposes the time series data into three main components: the trend component, which reflects the long-term change trend of the time series; the seasonal component, which reveals the periodic fluctuations of the time series (such as the sales peak during holidays every year); the residual component, which represents the random fluctuations or the unexplained part. Through this decomposition method, the system can identify and independently analyze each component in the time series, generating trend component data, seasonal component data, and residual component data.

[0092] Step S23: Use a variety of preset deep learning models to perform sales trend and user behavior prediction based on key economic indicators according to the trend component data, seasonal component data, and residual component data, so as to obtain trend prediction data;

[0093] In the system of the embodiment of the present invention, multiple preset deep learning models are used to predict the trend component, seasonal component and residual component. The system usually selects models suitable for time series prediction, such as LSTM (Long Short-Term Memory Network), GRU (Gated Recurrent Unit) or TCN (Temporal Convolutional Network). First, the system predicts the future overall sales trend or the change in user activity through the trend component data. Then, based on the seasonal component data, the system predicts the future periodic fluctuations, such as the sales peak during holidays or promotions. Finally, the system analyzes the residual component to handle unpredictable anomalies. By combining the outputs of these models, the system generates trend prediction data to provide a basis for decision-making.

[0094] Step S24: Identify abnormal fluctuations in the trend prediction data and conduct correlation analysis based on the key economic indicator data, thereby constructing an economic indicator relationship network.

[0095] The system of the embodiment of the present invention identifies abnormal fluctuations in the trend prediction data. First, the system identifies the abnormal fluctuations in the prediction through comparison with historical data using statistical methods (such as Z-score or anomaly detection models), such as sudden sales surges or changes in user behavior. Next, the system conducts correlation analysis based on the key economic indicator data, usually using Granger causality analysis or mutual information analysis, to determine the mutual influence between different economic indicators. By analyzing the correlation of different economic indicators, the system constructs an economic indicator relationship network, showing the correlation paths and influence directions between key indicators. This network can help enterprises better understand the interaction between various economic factors and provide support for strategy optimization and decision-making.

[0096] The present invention extracts time series related to key economic indicators from standardized e-commerce economic data, ensuring that the analysis focuses on the most important indicators for e-commerce business; this helps reduce data complexity and redundancy, focuses on the information most valuable for prediction and decision-making, and enhances the pertinence and efficiency of analysis. Such a method lays the foundation for subsequent time series analysis and trend prediction, enabling the model to capture the dynamic change patterns of key indicators. Seasonal decomposition of the time series data of key indicators into three components: trend, seasonality, and residuals, helps identify and understand the long-term trends, periodic patterns, and random fluctuations in the time series. Seasonal decomposition can help decision-makers more clearly see the long-term growth trends in the data (such as the continuous increase or decrease in sales), periodic changes (such as seasonal sales peaks), and unpredictable random fluctuations (such as sales changes caused by unexpected events). By clarifying these data components, more accurate predictions can be made, and strategies can be formulated and decisions optimized in a targeted manner. Using deep learning models to model and predict trend, seasonality, and residual data can fully exploit the non-linear relationships in the data and improve the accuracy and reliability of predictions. Deep learning models can automatically extract complex patterns and features in the data and adapt to the high-dimensional and dynamic change characteristics of e-commerce data. By using multiple deep learning models (such as long short-term memory networks LSTM, convolutional neural networks CNN, etc.), the advantages of different models can be combined for more comprehensive predictions, effectively enhancing the ability to grasp future sales trends and user behaviors and providing more accurate support for business decisions. Identifying abnormal fluctuations in the trend prediction data can timely detect and warn of potential market anomalies (such as sudden demand changes, market turmoil, etc.), helping enterprises to quickly respond and reduce risks. By analyzing the correlation between abnormal fluctuation data and key economic indicator data, an economic indicator relationship network can be constructed to reveal the mutual influence and dynamic relationships between various indicators. Such a network model helps to deeply understand the driving factors and interaction mechanisms of market behavior and provides a strong basis for optimizing resource allocation, adjusting marketing strategies, and improving operating efficiency. The effect of these steps is to establish a comprehensive and dynamic prediction analysis framework, combining deep learning technology to effectively decompose and predict the time series data of key economic indicators. By identifying trends, seasonality, and abnormal changes, enterprises can more accurately grasp future development trends, quickly respond to market changes, optimize decisions, and improve business performance. In addition, the construction of the economic indicator relationship network also provides a scientific basis for more in-depth strategic analysis and decision-making.

[0097] Preferably, step S23 includes the following steps:

[0098] Step S231: Perform sales trend prediction on the trend component data based on the long short-term memory network model to obtain sales trend prediction data;

[0099] In the embodiment of the present invention, sales trend prediction is performed on trend component data, and a long short-term memory network (LSTM) model is used for training. LSTM is a special recurrent neural network that can effectively capture long-term dependencies in time series. The system inputs trend component data (such as the sales trend in the past few years), and LSTM learns the long-term patterns and short-term fluctuations in the time series through its memory units. After multiple iterative trainings, the LSTM model can predict the sales trend at future time points and generate sales trend prediction data. This method can effectively handle the long-term non-linear changes in sales data and improve the accuracy of prediction.

[0100] Step S232: Perform seasonal pattern recognition on seasonal component data based on a convolutional neural network model to obtain seasonal prediction data;

[0101] In the embodiment of the present invention, a convolutional neural network (CNN) model is used to perform seasonal pattern recognition on seasonal component data. The CNN model can extract local patterns in the time series through convolutional layers, such as seasonal peaks or periodic fluctuations. The system first formats the seasonal component data and converts it into a format suitable for input into the CNN, and then extracts seasonal patterns (such as sales peaks during specific holidays every year) through multiple convolutional operations. After processing by pooling and fully connected layers, the system can identify seasonal fluctuations in future time periods and generate seasonal prediction data. This process can capture the periodic changes in e-commerce business and provide support for marketing and inventory planning.

[0102] Step S233: Perform short-term fluctuation prediction on residual component data based on an autoregressive integrated moving average model to obtain residual prediction data;

[0103] In the embodiment of the present invention, short-term fluctuation prediction is performed on residual component data based on an autoregressive integrated moving average model (ARIMA). The ARIMA model is a classic time series prediction model suitable for time series with short-term volatility. The system first performs a stationarity test on the residual data and performs differencing on non-stationary data. Then, it captures the dependencies of the residuals through the autoregressive (AR) part and eliminates the influence of random errors through the moving average (MA) part. After parameter adjustment and model training, the system can predict short-term residual fluctuations and generate residual prediction data. This model can effectively handle the influence of short-term fluctuations and noise and improve the overall prediction accuracy.

[0104] Step S234: Construct a user behavior graph based on the user transaction behavior data using a graph neural network model and perform user behavior prediction to obtain user behavior prediction data;

[0105] Embodiments of the present invention are based on user transaction behavior data and use a graph neural network (GNN) model to construct and predict a user behavior graph. First, the system models user transaction behavior data (such as purchase records and browsing habits) as a graph structure, where nodes represent users and edges represent interactions or behavioral associations between users. Then, the system uses the GNN model to train the user behavior graph. The GNN can learn the complex correlations of user behaviors through a message passing mechanism. The model can infer the future behaviors of target users from the behaviors of neighboring users, thereby predicting user behaviors. This step generates user behavior prediction data, which helps predict future user activity, purchase intention, etc.

[0106] Step S235: Combine the sales trend prediction data, seasonal prediction data, residual prediction data, and user behavior prediction data into trend prediction data.

[0107] Embodiments of the present invention combine the sales trend prediction data, seasonal prediction data, residual prediction data, and user behavior prediction data. The combination process aligns the data by time and integrates the results output by each prediction model in chronological order. The system uses a weighted average method to assign different weights to the prediction results from different sources to ensure a reasonable balance of the results of each prediction model during the combination process. Finally, the system generates complete trend prediction data, which can comprehensively reflect the sales trend, seasonal fluctuations, short-term anomalies, and user behavior changes over a period of time in the future, providing comprehensive decision-making support for business strategies.

[0108] The present invention uses a Long Short-Term Memory (LSTM) model to predict the sales trend of trend component data, which can capture long-term dependencies in time series data and is particularly suitable for trend analysis with a long time span. The LSTM model performs outstandingly in processing complex time series data with lag effects and non-linear relationships, and can effectively reduce the prediction error caused by the long-term dependence of data, improving the prediction accuracy of sales trends. This process helps enterprises accurately grasp the long-term development direction and trend of the market, providing more accurate data support for strategic planning and resource allocation. By using a Convolutional Neural Network (CNN) to perform pattern recognition on seasonal component data, it can effectively capture local spatio-temporal features in the data, such as periodic sales fluctuations and peak periods. The feature extraction ability of CNN enables it to quickly identify and learn seasonal patterns in the data, which is particularly important in the analysis of e-commerce sales data with obvious periodic characteristics. In this way, enterprises can better predict future seasonal sales trends, optimize inventory and supply chain management, and formulate more targeted marketing strategies. Using the Autoregressive Integrated Moving Average (ARIMA) model to predict short-term fluctuations of residual component data can handle random fluctuations and short-term changes in time series data. The ARIMA model is very effective for short-term prediction and is suitable for capturing random and non-stationary factors in the data. Through this step, enterprises can identify short-term market fluctuations and anomalies, adjust operation strategies in a timely manner, reduce market risks and losses. At the same time, it can also provide a basis for short-term promotional activities and marketing decisions. Using a Graph Neural Network (GNN) model to construct a user behavior graph and make predictions helps to deeply understand the complex correlations and interaction patterns of user behavior. GNN can capture the complex relationships between users and goods, and between users and users, and predict the future behavior of users, such as purchase propensity, consumption frequency, etc. by learning these relationship networks. The effect of this step is to help e-commerce enterprises better understand user needs and preferences, optimize the personalized recommendation system, improve user experience and satisfaction, and thus increase user conversion rate and stickiness. Merging the sales trend prediction data, seasonal prediction data, residual prediction data, and user behavior prediction data to form a comprehensive trend prediction dataset can comprehensively reflect the future trend of e-commerce business. This merging strategy can integrate the prediction advantages of different models, achieving higher prediction accuracy and reliability. By comprehensively considering various factors, enterprises can formulate more accurate sales forecasts and market strategies, reduce decision-making risks, and improve the flexibility and responsiveness of the business. The effect of these steps is to comprehensively predict the future trends and user behavior of e-commerce business by combining the advantages of multiple deep learning models and traditional statistical models. By deeply analyzing long-term trends, seasonal changes, short-term fluctuations, and user behavior, enterprises can more accurately understand market dynamics and user needs, optimize operation strategies, improve decision-making quality and efficiency, and thus gain a competitive advantage.

[0109] Preferably, step S24 includes the following steps:

[0110] Step S241: Calculate the local statistical features of the trend prediction data by the sliding window method to obtain local feature data;

[0111] In the embodiment of the present invention, the sliding window method is used to calculate the local statistical features of the trend prediction data. The sliding window method means sliding a window of a fixed size within a period of time (such as 7 days or 30 days) and gradually performing statistical analysis on the data within the window. The system calculates local statistical features such as mean, variance, maximum value, and minimum value for the data within each time window, thereby generating local feature data. By means of the sliding window, the system can dynamically capture the short-term fluctuations and local changes of the trend data, providing a basis for subsequent anomaly detection.

[0112] Step S242: Calculate the anomaly scores of each time point based on the Z-score method according to the local feature data to obtain anomaly score sequence data;

[0113] In the embodiment of the present invention, the Z-score method is used to calculate the anomaly scores of the local feature data. The Z-score represents the degree of standard deviation deviation of a data point relative to the mean value and is used to measure the degree of anomaly. The system calculates the Z-score of each time point to determine whether it deviates from the overall trend. For the time points with too high Z-score values, the system marks them as potential anomalies and records the corresponding anomaly score sequence data. This method can effectively identify the moments that deviate significantly from the normal trend, providing a quantitative basis for the identification of abnormal fluctuations.

[0114] Step S243: Identify the time points exceeding the threshold for the anomaly score sequence data by the dynamic threshold method to obtain anomaly fluctuation marked data;

[0115] In the embodiment of the present invention, the dynamic threshold method is used to identify the anomaly fluctuations of the anomaly score sequence data. The dynamic threshold method is to continuously adjust the threshold and set a suitable judgment criterion according to the data distribution. The system performs real-time analysis on the anomaly scores, sets an adaptive threshold, identifies the time points with scores exceeding the threshold, marks them as anomaly fluctuation moments, and generates anomaly fluctuation marked data. This method can more flexibly adapt to the fluctuation characteristics of different time periods and improve the accuracy of anomaly detection.

[0116] Step S244: Extract the index values of the corresponding time periods for the key economic indicator data according to the anomaly fluctuation marked data to obtain anomaly period index data;

[0117] In an embodiment of the present invention, according to the abnormally fluctuating marked data, a time period corresponding to the abnormal moment is extracted from the key economic indicator data. The system extracts the economic indicator values at these time points by matching the time points of the abnormally fluctuating marks, forming abnormally-period indicator data. This step ensures that the system can accurately locate which key economic indicators have abnormal changes during the abnormally fluctuating period, providing accurate input data for subsequent correlation analysis.

[0118] Step S245: Calculate the Pearson correlation coefficient matrix based on the abnormally-period indicator data to obtain indicator correlation data;

[0119] In an embodiment of the present invention, the Pearson correlation coefficient is used to perform correlation analysis on the abnormally-period indicator data. The Pearson correlation coefficient is used to measure the linear correlation between two variables, and its value range is between -1 and 1. The system generates a correlation matrix by calculating the correlation coefficients between key economic indicators. Each element in the matrix represents the correlation strength between the corresponding indicators. This step can reveal the mutual correlation between various economic indicators at the abnormal moment.

[0120] Step S246: Identify the main influencing factors based on the principal component analysis method according to the indicator correlation data, and construct a weighted undirected graph based on the main influencing factors to obtain an indicator relationship graph, where the nodes in the undirected graph represent indicators and the edge weights represent the correlation strength;

[0121] In an embodiment of the present invention, based on the principal component analysis (PCA) method, several main factors that have the greatest impact on the abnormal fluctuation are identified. The principal component analysis uses dimensionality reduction technology to screen out the principal components that explain the largest variance in the data, and these principal components represent the main influencing factors. The system constructs a weighted undirected graph according to the PCA results, where the nodes represent economic indicators and the weights of the edges represent the correlation strength between the indicators, generating an indicator relationship graph. This graph reveals the association structure between key economic indicators, providing a visual representation for understanding the influence path between indicators.

[0122] Step S247: Identify indicator groups with close associations in the indicator relationship graph and construct a hierarchical economic indicator relationship network.

[0123] In an embodiment of the present invention, group identification is performed on the indicator relationship graph to identify indicator groups with close associations. The system uses community detection algorithms, such as the Girvan-Newman algorithm, to identify highly associated indicator clusters in the relationship graph. Subsequently, the system constructs a hierarchical economic indicator relationship network based on these clusters, divides the key indicators in the group according to the influence level, and shows the hierarchical structure between various economic factors and their mutual influence paths. This network helps to reveal the internal associations and dominant factors of various economic indicators during the abnormal fluctuation.

[0124] The present invention calculates the local statistical features of trend prediction data through the sliding window method, which can capture the local change trends and features of the data in different time periods. This method can identify short-term fluctuations and local anomalies in the data, thereby providing detailed feature data support for subsequent anomaly detection. The application of the sliding window method helps to refine data analysis, enhance the accuracy and sensitivity of anomaly detection, and effectively discover potential short-term market changes and anomalies. By calculating the anomaly scores at each time point, the degree of deviation of each time point from the overall data distribution can be quantified. The Z-score method can standardize the values of different features, making anomaly detection more unified and accurate. It helps to quickly identify the abnormal time points in the trend prediction data, enabling enterprises to promptly discover potential abnormal fluctuations, respond quickly, and reduce market and business risks. The dynamic threshold method can adaptively adjust the criteria for anomaly detection, making the anomaly identification more flexible and accurate in different time periods. It can dynamically set the threshold according to the distribution characteristics of the data, effectively avoiding misjudgment problems caused by fixed thresholds. Through this method, enterprises can accurately identify the fluctuation events beyond the normal range and thus quickly take effective countermeasures when anomalies occur. By identifying the abnormal time points and extracting the indicator values of key economic indicators in the corresponding time periods, it helps to further analyze the reasons and features behind these abnormal fluctuations. This process can provide a precise data basis for subsequent correlation analysis, enabling enterprises to deeply understand the specific economic performance and market reactions during abnormal fluctuations. Using the Pearson correlation coefficient matrix calculation can quantitatively measure the correlations between key economic indicators; this step can help enterprises identify which indicators show strong correlations during abnormal fluctuations, thereby providing reliable data support for subsequent factor analysis of impacts and construction of relationship diagrams. This correlation analysis helps to reveal potential causal relationships and enables enterprises to better understand market dynamics and the complex interactions between indicators. By using the principal component analysis (PCA) method to identify the main factors affecting abnormal fluctuations, the dimensionality of the data can be reduced, focusing on the most representative economic indicators. At the same time, the constructed weighted undirected graph can visually present the mutual relationships and influence degrees among the main economic indicators, providing a clear overall view for enterprises and helping to quickly identify the key factors most influential on the market and business. Analyzing the indicator relationship diagram and identifying closely related indicator groups can construct a hierarchical economic indicator relationship network. Such a network structure can help enterprises more clearly understand the hierarchical relationships and impact paths among economic indicators, providing a more systematic and intuitive tool for the analysis of complex economic systems. By constructing this relationship network, enterprises can effectively identify core indicators and main driving factors, optimize resource allocation, and improve the accuracy and efficiency of decision-making. The effects of these steps are to comprehensively reveal the complex associations and influencing factors among key economic indicators through various methods such as detailed feature analysis, anomaly detection, correlation analysis, and graph structure construction.Through this in-depth analysis, enterprises can more precisely understand market dynamics and business performance, optimize business strategies, and enhance market responsiveness and competitive advantages.

[0125] Preferably, step S3 includes the following steps:

[0126] Step S31: Construct an economic knowledge graph including entities of goods, users, transactions, and marketing based on a preset ontology knowledge base in the e-commerce field;

[0127] In the embodiment of the present invention, an economic knowledge graph including entities of goods, users, transactions, and marketing is constructed based on a preset ontology knowledge base in the e-commerce field. This ontology knowledge base contains the definitions of various entities in the e-commerce field and their mutual relationships. For example, the relationships between goods and users, transactions, and marketing activities. The system parses the information in the e-commerce data, identifies and labels the goods entities, user behaviors, transaction records, and marketing strategies, and adds these entities and their interaction relationships to the knowledge graph. In this way, the system can construct a comprehensive economic knowledge graph that can reflect the business logic of e-commerce, providing a basis for subsequent steps.

[0128] Step S32: Map the nodes and edges in the economic indicator relationship network to the corresponding entities and relationships in the economic knowledge graph, and use BERT to generate vector representations of the entities and relationships in the knowledge graph, thereby obtaining entity-relationship vector data;

[0129] In the embodiment of the present invention, the nodes (economic indicators) and edges (relationships) in the economic indicator relationship network are mapped to the corresponding entities and relationships in the constructed economic knowledge graph. Subsequently, the system uses BERT (Bidirectional Encoder Representations from Transformers) to vectorize the entities and their relationships in the graph. The BERT model can generate high-dimensional vector representations for each entity and relationship according to the context information. These vectors can capture the semantic information between the entities and relationships, thereby generating entity-relationship vector data. This vector representation provides a basis for subsequent reasoning and causal relationship identification, enhancing the operability of the information in the graph.

[0130] Step S33: Perform reasoning on the economic indicator relationship network based on the graph attention network model according to the entity-relationship vector data, and identify potential causal chains and influence paths, thereby obtaining potential causal path data;

[0131] In the embodiments of the present invention, a graph attention network (GAT) model is used to reason about entity-relationship vector data. The graph attention network can identify important nodes and relationships by assigning different attention weights to the nodes and edges in the graph. Through reasoning about the economic knowledge graph, the system can identify potential causal chains and influence paths. For example, how the price change of a certain commodity affects user behavior, or how a certain marketing activity promotes sales growth. Finally, the system generates potential causal path data to represent these potential causal relationships and influence paths.

[0132] Step S34: Generate an explanatory text template based on the potential causal path data, and use the explanatory text template to generate text describing the relationship between economic indicators and the reasons for changes, so as to obtain preliminary explanatory text data;

[0133] In the embodiments of the present invention, an explanatory text template is generated according to the identified potential causal path data. These templates preset how to describe the relationship between different economic indicators and the reasons for their changes. For example, a certain template may describe that "the increase in commodity price leads to the decrease in user purchase volume" or "the promotion activity increases user activity". The system uses these templates to explain economic phenomena and generates preliminary explanatory text data for the causal relationship and the reasons behind it. This step generates preliminary text through fixed templates, laying a foundation for subsequent personalized explanations.

[0134] Step S35: Use natural language processing technology to fill the information in the trend prediction data into the preliminary explanatory text data, so as to obtain personalized economic phenomenon explanatory data;

[0135] In the embodiments of the present invention, through natural language processing technology, the key information in the trend prediction data is filled into the preliminary explanatory text to generate personalized economic phenomenon explanatory data. The system dynamically adjusts the explanatory text according to specific contents such as user transaction behavior, sales trends, and user behavior prediction. For example, the system can add specific numerical predictions to the explanatory text, such as "it is expected that the sales volume will increase by 15% in the next week" or "the user click-through rate is expected to decrease by 10%". This personalized filling makes the explanation more accurate and meets the needs of specific e-commerce scenarios.

[0136] Step S36: Select a visualization type according to the economic phenomenon explanatory data and the trend prediction data, and implement it through an interactive exploration tool, so as to obtain interactive visualization data.

[0137] In the embodiments of the present invention, appropriate visualization types are selected based on the generated economic phenomenon interpretation data and trend prediction data. According to different data characteristics, the system may select different visualization methods such as time series charts, association network diagrams, bar charts, heat maps, etc., to intuitively display economic phenomena and prediction results. Then, the system realizes interaction with the user through an interactive exploration tool. The user can click on different visualization elements to explore the data details behind, view specific economic indicators and association relationships, and finally generate interactive visualization data to improve the interpretability of the data and the user's participation.

[0138] The present invention can systematically and structurally organize and represent various data and knowledge in the e-commerce field by constructing an economic knowledge graph that includes entities such as commodities, users, transactions, and marketing; the establishment of the knowledge graph helps integrate different types of information together to form a comprehensive economic data semantic network, enabling better understanding of the complex relationships and interactions between entities in subsequent analyses. At the same time, such a graph can promote data sharing and reuse, improve the interpretability of data and the accuracy of model reasoning; mapping the nodes and edges in the economic indicator relationship network to the entities and relationships of the knowledge graph, and using the BERT model to generate vector representations can provide deep semantic information for the entities and relationships in the knowledge graph. In this way, entity-relationship vector data can capture the complex semantic associations between economic indicators and entities, providing rich feature representations for subsequent causal reasoning. This method helps improve the reasoning effect of the graph attention network model, making the identified causal chains and influence paths more accurate and reliable. By using the graph attention network model to reason about entity-relationship vector data, potential causal chains and influence paths in the economic system can be identified. The graph attention network model can adjust the weights of nodes and edges, enabling the model to pay more attention to important nodes and relationships and identify the key factors and paths that have important impacts on changes in economic indicators. This reasoning method can help enterprises discover hidden causal relationships, provide more insightful data support for business decisions, and enhance the scientific nature and accuracy of decisions. Generating explanatory text templates based on potential causal path data and using these templates to generate texts describing the relationships and reasons for changes between economic indicators can provide a preliminary qualitative explanation for economic phenomena. This step helps convert complex analysis results into easy-to-understand natural language descriptions, enabling non-technical personnel to understand the meaning of the analysis and the discovered causal relationships, thus better supporting cross-departmental communication and decision-making within the enterprise. Filling the information in the trend prediction data into the preliminary explanatory text through natural language processing technology to generate personalized economic phenomenon explanation data can provide more targeted analysis results and explanations for users. This method ensures the relevance and accuracy of the explanations, enabling users to obtain customized explanation content according to their own needs and concerns, thereby improving the user experience and the practicality of the analysis. Selecting the visualization type according to the economic phenomenon explanation data and trend prediction data and implementing interactive visualization through an interactive exploration tool enable users to explore the data and analysis results in an intuitive way. The interactive visualization data provides a multi-dimensional perspective, and users can perform data screening, filtering, and in-depth analysis according to their own needs. This step helps improve the readability and understandability of the data, while enhancing the user's ability to explore and gain insights from the data, providing more effective support for data-driven decision-making. These steps provide a comprehensive, interpretable, and user-friendly data analysis process by combining various technologies such as knowledge graph construction, vector representation generation, deep learning reasoning, natural language generation, and interactive visualization.This process can not only reveal the complex causal relationships in the e-commerce economic system, but also present the analysis results in a personalized and visual way, enhancing the enterprise's capabilities and efficiency in data analysis, strategic formulation, and decision-making support.

[0139] Preferably, step S36 includes the following steps:

[0140] Step S361: Extract content features based on data dimension, time span, and relationship complexity according to the economic phenomenon explanatory data, and establish a visualization type decision tree;

[0141] In the embodiment of the present invention, in-depth analysis is carried out on the economic phenomenon explanatory data, and content features such as data dimension, time span, and relationship complexity are extracted therefrom. The data dimension may include different economic indicators, such as sales volume, user activity, etc.; the time span involves the time range of the data, such as week, month, or quarter; the relationship complexity refers to the complexity of the relationship network between the data, such as the association of a single indicator with multiple factors. According to these features, the system constructs a visualization type decision tree, which includes nodes of various visualization types (such as line chart, bar chart, heat map, etc.), and each node corresponds to different feature combination conditions. By traversing the decision tree, the system can select the most suitable visualization scheme for different types of data, ensuring that the visualization effect can accurately convey the characteristics and trends of economic phenomena.

[0142] Step S362: Traverse the visualization type decision tree, select the visualization type according to the economic phenomenon explanatory data, and generate an initial visualization chart based on the selected visualization type and ECharts;

[0143] In the embodiment of the present invention, the visualization type decision tree is traversed to determine the most suitable visualization type. The system selects a suitable visualization chart type according to the specific content features in the economic phenomenon explanatory data, such as data dimension, time span, and relationship complexity. For example, if the data has strong time series characteristics, a line chart may be selected; if multi-dimensional relationships need to be displayed, a heat map or scatter plot may be selected. After the visualization type is selected, the system uses the ECharts library to generate an initial visualization chart. ECharts is an open-source visualization library based on JavaScript, which can efficiently generate various charts and support rich customization functions, such as chart styles, colors, and data labels, to ensure that the generated charts can accurately and clearly display the economic phenomenon explanatory data.

[0144] Step S363: Design interactive controls for the initial visualization chart, and implement the drilling function from macroeconomic indicators to specific commodity categories or regions step by step, so as to obtain interactive visualization data, where the interactive controls include a time slider, an index selector, and a granularity adjuster.

[0145] In the embodiments of the present invention, the design and implementation of interactive controls for the initial visualization chart are carried out. The system adds multiple interactive controls, such as a time slider, an indicator selector, and a granularity adjuster. The time slider allows users to adjust the time range for viewing data; the indicator selector provides users with the option to select different economic indicators for comparison and analysis; the granularity adjuster enables users to control the level of detail of the data, such as drilling down from macroeconomic indicators to specific commodity categories or regions step by step. The system ensures that these controls can interact dynamically with the chart data by writing corresponding JavaScript code and ECharts configurations, enabling users to explore the data in depth. Finally, the system generates interactive visualization data, allowing users to conveniently adjust the view according to their personal needs, view data from different dimensions, and thus gain comprehensive and specific insights into economic phenomena.

[0146] The present invention extracts content features based on data dimensions, time spans, and relationship complexities by interpreting data according to economic phenomena, and establishes a visualization type decision tree, which can systematically determine the most suitable visualization type for displaying data. The decision tree helps analysts quickly select the most effective visualization method when facing different data characteristics, thereby ensuring that the data is presented in the clearest and most effective way. The key to this step lies in being able to formulate corresponding visualization strategies for different data characteristics, enhancing the intuitiveness of data display and the effect of information transmission. Traverse the visualization type decision tree, and select the visualization type according to the economic phenomenon to interpret the data. Generate an initial visualization chart through ECharts, and the data can be presented in a graphical form. As a powerful visualization library, ECharts can generate high-quality charts and support a variety of interaction and customization functions. This step enables users to intuitively view data trends and patterns by generating an initial chart, helping to quickly identify important information and anomalies. The generated chart provides a basis for subsequent in-depth analysis and data exploration, and provides users with first-hand visual data feedback. Design interactive controls for the initial visualization chart and implement a drill-down function from macroeconomic indicators to specific commodity categories or regions, which can greatly enhance the interactivity between users and data. By designing controls such as a time slider, an indicator selector, and a granularity regulator, users can dynamically adjust the data view and view the data according to different time periods, indicators, or granularities. The drill-down function allows users to gradually drill down from the overall data to specific sub-datasets, thereby discovering more detailed information and insights. This interactive design not only enhances users' control over the data, but also supports more detailed analysis and decision-making, improving the flexibility and depth of data exploration. These steps provide an efficient, intuitive, and interactive display method for interpreting data according to economic phenomena by comprehensively applying technologies such as data feature extraction, decision tree establishment, chart generation, and interaction design; this method not only ensures that the data is presented in the most suitable visualization way, but also enhances the interactive experience and analysis ability of users in the data analysis process, making the data display more insightful and practical.

[0147] Preferably, step S4 includes the following steps:

[0148] Step S41: Extract the minimum data granularity units in multiple dimensions for the interactive visualization data, and perform a multi-level data aggregation strategy based on the granularity units, and pre-calculate the aggregation results at different granularity levels to obtain pre-aggregated data, where the multiple dimensions specifically include time, region, and product category;

[0149] In the embodiments of the present invention, the minimum data granularity units of interactive visualization data are extracted in multiple dimensions, and a multi-level data aggregation strategy is based on the granularity units, and the aggregation results at different granularity levels are pre-computed. First, the pandas library is used to identify and extract the minimum data granularity units of each data dimension (time, region, and product category) in the interactive visualization data. For example, the time dimension is refined to the day level, the region dimension is refined to the city level, and the product category dimension is refined to the specific product model. Then, the groupby function and the aggregate method are used to perform aggregation operations on the data at different granularity levels, generating result data such as the total sales amount aggregated by month and the user activity aggregated by province. Finally, the aggregation results at all granularity levels are stored in the database, and index optimization is performed to provide a faster access speed for subsequent query operations, thereby obtaining pre-aggregated data.

[0150] Step S42: Perform spatial data segmentation based on a quadtree index on the pre-aggregated data, and construct a multi-resolution data pyramid, thereby obtaining a multi-level data structure;

[0151] In the embodiments of the present invention, spatial data segmentation based on a quadtree index is performed on the pre-aggregated data, and a multi-resolution data pyramid is constructed. First, according to the geographical location information of the pre-aggregated data, the spatial data is segmented using the quadtree index algorithm (QuadTree) in the geopandas library, and the data is divided into different node levels of the quadtree. Then, the multi-resolution grading strategy of the data pyramid (for example, according to the amount of data and the geographical resolution) is used to downsample the data, constructing a multi-level data structure, forming a data pyramid from low resolution to high resolution; the data at the low resolution level is used for fast loading and preliminary display, and the data at the high resolution level is gradually loaded when the user requests. This hierarchical design can effectively optimize performance, reduce the time and resource consumption of data loading, thereby obtaining a multi-level data structure.

[0152] Step S43: Implement a progressive data loading strategy according to the multi-level data structure, thereby obtaining performance-optimized data, where the data loading strategy includes fast loading of the first screen, scroll loading, and asynchronous update;

[0153] Embodiments of the present invention implement a progressive data loading strategy based on a multi-level data structure to obtain performance-optimized data. First, using a front-end framework such as React or Vue.js, when the page is loaded, according to the user's network conditions and device performance, low-resolution data in the multi-level data structure is preferentially loaded to achieve fast first-screen loading. Then, when the user scrolls or zooms the map, the IntersectionObserver API is used to trigger the scroll loading mechanism, asynchronously obtain higher-resolution data from the server side and update it to the front-end display. The asynchronous update strategy uses Web Workers to execute background data requests, avoiding blocking the main thread, thereby improving the loading efficiency and achieving the goal of performance-optimized data.

[0154] Step S44: Perform data compression processing on the performance-optimized data based on integer encoding and differential encoding to obtain compressed and optimized data;

[0155] Embodiments of the present invention perform data compression processing on the performance-optimized data based on integer encoding and differential encoding. First, the NumPy library is used to convert the data into an integer type to reduce the storage space of the data. For time series data with high continuity, differential encoding (Delta Encoding) is used to calculate the difference between adjacent data, and the original data is replaced with a difference sequence to further compress the data volume. The encoded data is serialized in JSON format and stored in memory or transmitted to the client. This compression method effectively reduces the overhead of data transmission and storage, thereby obtaining compressed and optimized data.

[0156] Step S45: Design a responsive layout engine based on the compressed and optimized data, including grid system design, flexible layout implementation, and component adaptive adjustment, to obtain a responsive layout solution;

[0157] Embodiments of the present invention design a responsive layout engine based on the compressed and optimized data, including grid system design, flexible layout implementation, and component adaptive adjustment. First, CSS Grid and Flexbox technologies are used to build a grid system to ensure the responsiveness and flexibility of the layout. By dynamically calculating the width and height of the grid, an adaptive layout of the page content on different devices is achieved. Then, for display devices with different resolutions (such as desktops, tablets, and mobile phones), Media Queries are set in the front-end code to automatically adjust the display style and position of the components according to the screen size and resolution. The design of the responsive layout engine ensures that the visual interface can be optimally displayed on various devices, thereby forming a responsive layout solution.

[0158] Step S46: Set breakpoints and style adaptation for different-sized display devices for the responsive layout solution, and implement dynamic component loading and layout reorganization logic to obtain multi-terminal adaptation data.

[0159] In the embodiments of the present invention, breakpoint settings and style adaptation for different-sized display devices are performed on the responsive layout solution, and the implementation of dynamic component loading and layout reorganization logic is carried out; First, use Media Queries in CSS to set breakpoints (such as min-width and max-width) at the front end, and adjust style adaptation according to different-sized display devices (such as mobile phones, tablets, desktops). Secondly, utilize the component lazy loading function (Lazy Loading) of JavaScript and React or Vue.js frameworks. When the device first loads the page, only key components are loaded, and other secondary components are dynamically loaded when needed. Finally, combined with the user's interaction operations (such as window resizing, screen rotation), recalculate the layout logic, and dynamically adjust the arrangement order and size of components to ensure the best user experience in various environments, thereby obtaining multi-terminal adaptation data.

[0160] The present invention can significantly improve data processing efficiency by extracting the smallest data granularity units in multiple dimensions from interactive visual data and performing multi-level data aggregation based on these granularity units; pre-computing the aggregation results at different granularity levels enables the rapid loading and display of the required data views when the user interacts. This pre-aggregation strategy helps reduce the burden of real-time computing, improve the system's response speed and user experience, and also optimize the performance of data storage and processing. Performing spatial data segmentation based on quadtree indexing on the pre-aggregated data and constructing a multi-resolution data pyramid can achieve efficient data organization and access. The quadtree indexing helps divide the data space into smaller regions, thereby increasing the speed of data query and update. The multi-resolution data pyramid allows the system to load data at different resolutions according to the user's zoom level, optimizing the data loading and display efficiency, so that the user can still maintain good performance and a smooth interaction experience when viewing a large range of data. Implementing a progressive data loading strategy based on a multi-level data structure helps optimize performance and enhance the user experience. The quick loading of the first screen ensures that users can quickly see important data when they first access, while the scroll loading and asynchronous update functions allow the system to dynamically load data when the user scrolls or interacts. This can reduce the initial loading time and load additional data on demand when the user needs it, improving the system's response speed and smoothness, and also reducing the occupancy of system resources. Performing data compression processing on the performance-optimized data based on integer coding and differential coding can effectively reduce the data storage space and transmission bandwidth. The integer coding and differential coding techniques reduce redundant data by compressing the data representation method, thereby reducing the storage cost and transmission delay. This not only improves the overall performance of the system but also ensures the efficiency of data transmission over the network, which is particularly important when dealing with large amounts of data. Designing a responsive layout engine based on the compressed and optimized data, including a grid system design, flexible layout implementation, and component adaptive adjustment, can ensure good display effects of the data on different screen sizes and devices. The grid system and flexible layout allow layout components to adaptively adjust their positions and sizes according to the screen size and content, improving the flexibility and usability of the interface. The component adaptive adjustment ensures the unity and consistency of the user interface on different devices, enhancing the user experience and the operability of the interface. Setting breakpoints and style adaptation for different-sized display devices for the responsive layout scheme, and implementing the logic of dynamic component loading and layout reorganization, thus obtaining multi-terminal adapted data. This step ensures that the visual interface can be optimally displayed on various devices and screen sizes, including desktop computers, tablets, and mobile phones, etc. The dynamic component loading and layout reorganization improve the experience consistency of users on different devices, ensure the readability and interactivity of the interface on various devices, and enhance the overall user satisfaction and usage experience. These steps significantly improve the system's performance and user experience by optimizing data processing, loading, storage, and display strategies.They ensure the efficiency and consistency of data display across various devices and screen sizes, while also enhancing the system's ability to process large amounts of data, enabling users to smoothly conduct data interaction and analysis.

[0161] Therefore, in every aspect, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present invention is defined by the appended claims rather than the above description. Thus, all changes falling within the meaning and scope of the equivalent elements of the application documents are intended to be encompassed within the present invention.

[0162] The above description is only a specific implementation manner of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.

Claims

1. A visualization display method based on multi-source economic data mining, characterized in that, Including the following steps: Step S1: Obtain multi-source e-commerce economic data, including platform sales data, merchant financial data, user transaction behavior data, marketing data, and external market intelligence data; perform data cleaning and standardization processing on the multi-source e-commerce economic data, and conduct data quality assessment and screening to obtain standardized e-commerce economic data; perform multi-dimensional data dimensionality reduction analysis on the standardized e-commerce economic data, and conduct feature importance ranking to obtain key economic indicator data; Step S2: Perform time series analysis on the standardized e-commerce economic data based on the key economic indicator data, and use a variety of preset deep learning models to predict sales trends and user behaviors based on the key economic indicators, so as to obtain trend prediction data; Identify abnormal fluctuations in the trend prediction data, and conduct correlation analysis based on the key economic indicator data, so as to construct an economic indicator relationship network; Step S3: Generate explanatory data based on economic phenomena for the economic indicator relationship network according to the preset economic knowledge graph to obtain economic phenomenon explanatory data; select a visualization type based on the economic phenomenon explanatory data and the trend prediction data, and achieve interaction through an interactive exploration tool to obtain interactive visualization data; Step S4: Perform data pre-aggregation and progressive data loading optimization on the interactive visualization data to obtain performance optimization data, and design a responsive layout engine based on the performance optimization data and the interactive visualization data to generate multi-terminal adaptation data for display devices of different sizes.

2. The visualization display method based on multi-source economic data mining according to claim 1, characterized in that Step S1 includes the following steps: Step S11: Obtain multi-source e-commerce economic data, including platform sales data, merchant financial data, user transaction behavior data, marketing data, and external market intelligence data; Step S12: Perform format unification processing on the multi-source e-commerce economic data based on the date format and numerical unit to obtain format-unified e-commerce data; Step S13: Perform missing value processing on the format-unified e-commerce data based on moving average and median filling, and perform outlier processing through the Z-score method to obtain preprocessed e-commerce data; Step S14: Conduct data quality assessment and screening on the preprocessed e-commerce data to obtain standardized e-commerce economic data; Step S15: Align and merge the standardized e-commerce economic data from different sources based on time and region to construct a multi-dimensional data cube; Step S16: Use the t-distributed stochastic neighbor embedding algorithm to perform data dimensionality reduction analysis on the multi-dimensional data cube, and select the principal components according to the preset explained variance ratio threshold to obtain dimension-reduced e-commerce data; Step S17: Conduct feature importance ranking on the dimension-reduced e-commerce data to obtain key economic indicator data.

3. The visualization display method based on multi-source economic data mining according to claim 2, wherein, Step S14 includes the following steps: Step S141: Calculate the field filling rate of product information, transaction records, and user behavior logs for the preprocessed e-commerce data to obtain integrity detection data; Step S142: Use the preset e-commerce business rules to conduct a consistency check on the preprocessed e-commerce data, so as to obtain consistency scoring data, where the consistency check includes the matching of the order amount and the commodity price, and the verification of the logical relationship between the inventory quantity and the sales volume; Step S143: Conduct real-time monitoring of the update frequency fever of the preprocessed e-commerce data, and calculate the synchronization degree of real-time transactions, so as to obtain timeliness evaluation data; Step S144: Calculate the conversion rate and the deviation rate of the customer unit price of the preprocessed e-commerce data by comparing the historical data of the data source of the preprocessed e-commerce data pre-obtained with the industry benchmark, so as to obtain accuracy scoring data; Step S145: Use the outlier detection method to identify false transactions and brushing behavior in the preprocessed e-commerce data, and calculate the abnormal proportion, so as to obtain data credibility; Step S146: Conduct a comprehensive evaluation process based on the weighted average method on the preprocessed e-commerce data according to the integrity detection data, consistency scoring data, timeliness evaluation data, accuracy scoring data, and data credibility, so as to obtain comprehensive quality score data; Step S147: Conduct hierarchical screening on the preprocessed e-commerce data according to the comprehensive quality score data and eliminate low-quality data, so as to obtain standardized e-commerce economic data.

4. The visualization display method based on multi-source economic data mining according to claim 3, characterized in that Step S17 includes the following steps: Step S171: Conduct a correlation analysis on the dimensionality-reduced e-commerce data based on the Pearson correlation coefficient, use the hierarchical clustering algorithm to group highly correlated features, and perform feature selection on each group, so as to obtain a preliminary screening feature set; Step S172: Calculate the average impurity reduction amount of each feature for the target variable according to the preliminary screening feature set, so as to obtain the initial ranking of feature importance, where the target variable specifically includes sales volume and user activity; Step S173: Select the top N features whose cumulative importance reaches the threshold from the initial ranking of feature importance according to the preset cumulative importance threshold, so as to obtain candidate key index data; Step S174: Conduct a multicollinearity test on the candidate key index data, calculate the variance inflation factor, and eliminate the features whose variance inflation factor exceeds the preset threshold, so as to obtain key economic index data.

5. The visualization display method based on multi-source economic data mining according to claim 4, wherein Step S2 includes the following steps: Step S21: Extract relevant time series from the standardized e-commerce economic data according to the key economic index data, so as to obtain key index time series data; Step S22: Conduct seasonal decomposition on the key index time series data, so as to obtain trend component data, seasonal component data, and residual component data; Step S23: Use a variety of preset deep learning models to conduct sales trend and user behavior prediction based on the key economic index according to the trend component data, seasonal component data, and residual component data, so as to obtain trend prediction data; Step S24: Identify abnormal fluctuations in the trend prediction data, and conduct a correlation analysis according to the key economic index data, so as to construct an economic index relationship network.

6. The visualization display method based on multi-source economic data mining according to claim 5, wherein Step S23 includes the following steps: Step S231: Perform sales trend prediction on the trend component data based on the long short-term memory network model to obtain sales trend prediction data; Step S232: Perform seasonal pattern recognition on the seasonal component data based on the convolutional neural network model to obtain seasonal prediction data; Step S233: Perform short-term fluctuation prediction on the residual component data based on the autoregressive integrated moving average model to obtain residual prediction data; Step S234: Construct a user behavior graph based on the user transaction behavior data using the graph neural network model and perform user behavior prediction to obtain user behavior prediction data; Step S235: Merge the sales trend prediction data, seasonal prediction data, residual prediction data, and user behavior prediction data into trend prediction data.

7. The visualization display method based on multi-source economic data mining according to claim 6, characterized in that Step S24 includes the following steps: Step S241: Calculate the local statistical features of the trend prediction data by the sliding window method to obtain local feature data; Step S242: Calculate the anomaly scores of each time point based on the Z-score method according to the local feature data to obtain anomaly score sequence data; Step S243: Identify the time points exceeding the threshold for the anomaly score sequence data by the dynamic threshold method to obtain anomaly fluctuation marking data; Step S244: Extract the index values of the corresponding time periods for the key economic indicator data according to the anomaly fluctuation marking data to obtain anomaly period index data; Step S245: Calculate the Pearson correlation coefficient matrix according to the anomaly period index data to obtain index correlation data; Step S246: Identify the main influencing factors based on the principal component analysis method according to the index correlation data and construct a weighted undirected graph according to the main influencing factors to obtain an index relationship graph, where the nodes in the undirected graph represent indicators and the edge weights represent the correlation strength; Step S247: Identify the indicator groups with close associations in the index relationship graph and construct a hierarchical economic indicator relationship network.

8. The visualization display method based on multi-source economic data mining according to claim 7, characterized in that, Step S3 includes the following steps: Step S31: Construct an economic knowledge graph containing entities of goods, users, transactions, and marketing based on the preset e-commerce domain ontology knowledge base; Step S32: Map the nodes and edges in the economic indicator relationship network to the corresponding entities and relationships in the economic knowledge graph, and use BERT to generate vector representations of the entities and relationships in the knowledge graph to obtain entity-relationship vector data; Step S33: Perform reasoning on the economic indicator relationship network based on the graph attention network model according to the entity-relationship vector data, and identify potential causal chains and influence paths to obtain potential causal path data; Step S34: Generate an explanatory text template according to the potential causal path data, and use the explanatory text template to generate text describing the relationships and change reasons between economic indicators to obtain preliminary explanatory text data; Step S35: Use natural language processing technology to fill the information in the trend prediction data into the preliminary explanatory text data to obtain personalized economic phenomenon explanation data; Step S36: Select a visualization type based on the economic phenomenon interpretation data and the trend prediction data, and implement the interaction through an interactive exploration tool to obtain interactive visualization data.

9. The visualization display method based on multi-source economic data mining according to claim 8, wherein Step S36 includes the following steps: Step S361: Extract content features based on the data dimension, time span, and relationship complexity from the economic phenomenon interpretation data, and establish a visualization type decision tree. Step S362: Traverse the visualization type decision tree, select a visualization type based on the economic phenomenon interpretation data, and generate an initial visualization chart based on the selected visualization type and ECharts. Step S363: Design interactive controls for the initial visualization chart, and implement a drilling function that drills down from macroeconomic indicators to specific commodity categories or regions to obtain interactive visualization data, where the interactive controls include a time slider, an indicator selector, and a granularity regulator.

10. The visualization display method based on multi-source economic data mining according to claim 9, characterized in that, Step S4 includes the following steps: Step S41: Extract the smallest data granularity units in multiple dimensions from the interactive visualization data, and perform a multi-level data aggregation strategy based on the granularity units, and pre-compute the aggregation results at different granularity levels to obtain pre-aggregated data, where the multiple dimensions specifically include time, region, and product category. Step S42: Perform spatial data segmentation based on a quadtree index on the pre-aggregated data, and construct a multi-resolution data pyramid to obtain a multi-level data structure. Step S43: Implement a progressive data loading strategy based on the multi-level data structure to obtain performance-optimized data, where the data loading strategy includes fast first-screen loading, scroll loading, and asynchronous update. Step S44: Perform data compression processing on the performance-optimized data based on integer coding and difference coding to obtain compressed and optimized data. Step S45: Design a responsive layout engine based on the compressed and optimized data for the grid system design, elastic layout implementation, and component adaptive adjustment to obtain a responsive layout scheme. Step S46: Set breakpoints and style adaptation for different-sized display devices for the responsive layout scheme, and implement the dynamic component loading and layout reorganization logic to obtain multi-terminal adaptation data.

Citation Information

Patent Citations

  • Economic operation monitoring method based on big data

    CN111950775A

  • Personal credit report query monitoring and early warning method based on artificial intelligence

    CN117934159A

  • Digital economic data acquisition system and method based on big data, and storage medium

    CN118093687A

  • Macroeconomic investment portfolio optimization method based on deep learning

    CN118485521A

  • Knowledge graph-driven sales data multi-dimensional analysis and visualization method and system

    CN119248989A

Cited By

  • Intelligent drug purchase quantity prediction method based on medical insurance consumption data

    CN120996698A

  • Precision marketing-oriented intelligent delivery targeted penetration decision-making method

    CN121032580A

  • Unmanned retail industry intelligent data analysis method based on large language model

    CN121094860A

  • Explanatable multi-dimensional CI index dynamic scoring and risk assessment method

    CN121638942A