Method for integrating data analysis and reasonable prediction
By integrating weighted grey relational analysis, GM(1,1) model, ARIMA model and decision tree analysis, this method overcomes the limitations of traditional data analysis methods in complex data processing, achieving efficient and accurate data analysis and prediction, and supporting intelligent decision-making in various industries.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGDONG UNIV OF SCI & TECH
- Filing Date
- 2025-12-31
- Publication Date
- 2026-04-21
AI Technical Summary
Traditional data analysis and prediction methods struggle to fully capture data characteristics when faced with complex, multi-dimensional, nonlinear, and uncertain data, leading to large biases in prediction results. Furthermore, they have limited processing capabilities for small samples and data with limited information, failing to provide a reliable basis for prediction.
By employing a multi-algorithm fusion approach, including weighted grey relational analysis, GM(1,1) model, ARIMA model, and decision tree analysis, comprehensive data analysis and accurate prediction are achieved through weighted analysis, time series prediction, and decision path construction.
It improves data processing efficiency, reduces decision-making risks, provides more forward-looking and scientific decision-making basis, and helps various industries achieve intelligent and precise development.
Smart Images

Figure CN121901910A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data analysis and prediction technology, and specifically to a method for integrating data analysis and reasonable prediction. Background Technology
[0002] In today's digital age, data has become a core element driving the development of various industries. In financial markets, accurate market forecasting is crucial for investment decisions; in industrial production, quality monitoring relies on the analysis of large amounts of production data; and in the medical field, early prediction and diagnosis of diseases also depend on the mining of relevant data. However, with the rapid increase in data volume, data exhibits characteristics such as multi-dimensionality, non-linearity, and uncertainty.
[0003] Traditional data analysis and forecasting methods exhibit numerous limitations when faced with such complex data. Single models often process data only from a specific perspective, failing to comprehensively capture data characteristics. For example, when dealing with data exhibiting complex relationships, traditional methods may fail to accurately assess the weights of various factors, leading to the omission of crucial information and significant biases in predictions. Furthermore, traditional methods have limited processing capabilities for small-sample, low-information data, unable to provide a reliable basis for predictions. In time series analysis, traditional methods also struggle to effectively capture trends and periodicities in the data. Therefore, there is an urgent need for a data analysis and forecasting method that integrates the advantages of multiple algorithms and boasts superior overall performance to address increasingly complex data challenges and provide a solid foundation for decision-making across various fields. Summary of the Invention
[0004] To address the aforementioned problems, this invention provides a method integrating data analysis and reasonable prediction, aiming to solve the problem of effective analysis and accurate prediction of massive amounts of data in finance, industry, healthcare, and other fields in today's digital age. By integrating the weighted grey relational algorithm, GM(1,1) model, ARIMA model, and decision tree analysis method, it overcomes the limitations of traditional single algorithms in handling complex, multi-dimensional, nonlinear, and uncertain data, providing reliable and efficient data support and scientific basis for decision-making in various industries, and promoting the development of intelligent and precise directions in various fields.
[0005] The technical solution adopted in this invention is: A method for integrating data analysis and reasonable prediction includes the following steps: S1. Use the weighted grey relational degree algorithm to perform weight analysis and key information mining on the data. Obtain data weights and perform consistency checks through the AHP hierarchical analysis method, and calculate the grey relational coefficient and weighted grey relational degree. S2. Use the GM(1,1) model to perform preliminary predictive analysis on the data after weighted grey relational analysis, including level ratio test and establishment of time series prediction equation; S3. Use the ARIMA model to perform secondary predictive analysis on the data. Determine the stationarity and lag relationship of the time series through ADF test, autocorrelation analysis and partial autocorrelation analysis, and optimize the predictive data. S4. The data is comprehensively analyzed using decision tree analysis methods to construct decision paths, including calculating information entropy, information gain, gain ratio and Gini coefficient, performing pruning and generating a decision tree architecture diagram. S5. Finally, draw data analysis conclusions to provide a basis for decision-making.
[0006] Preferably, in the S1 weighted grey relational algorithm, the step of obtaining data weights using the AHP (Analytic Hierarchy Process) includes: Calculate the average value of each analysis item and construct the judgment matrix; Perform consistency checks and calculate consistency indices and consistency ratios; Calculate the grey relational coefficient, and consider the resolution coefficient to adjust the resolution; The grey weighted correlation degree of each evaluation index object is calculated according to the weighted grey correlation degree formula, and the average value is accumulated to avoid randomness.
[0007] Preferably, the preliminary prediction analysis steps of the GM(1,1) model in S2 include: Perform a grade ratio test and establish a time series; Solve the prediction equation to obtain the predicted value and the relative error test equation.
[0008] Preferably, the secondary prediction analysis steps of the ARIMA model in S3 include: Perform the ADF test to examine the stationarity of the time series; The optimal difference sequence and lag relationship are determined using autocorrelation analysis and partial autocorrelation analysis. Use the ARIMA model for prediction to capture trends, periodicity, and other characteristics in the data.
[0009] Preferably, the decision tree analysis method in S4 includes: Calculate the information entropy of the sample set, and define information gain and gain ratio; The Gini coefficient is used to measure data impurity. Pruning is performed to generate a decision tree architecture diagram, which is then analyzed.
[0010] Preferably, the weighted grey relational algorithm, GM(1,1) model, ARIMA model and decision tree analysis method are applied sequentially in the data processing process to achieve multi-algorithm fusion and collaborative optimization.
[0011] Preferably, in the consistency test, a consistency index and a consistency ratio are calculated. When the consistency ratio is less than a set threshold, the weight setting is considered reasonable and usable.
[0012] Preferably, in the process of solving the prediction equation of the GM(1,1) model, the predicted value is calculated by a given formula, and the prediction accuracy is checked by a relative error test equation.
[0013] Preferably, in the autocorrelation analysis of the ARIMA model, the stability and trend of the optimal difference sequence are determined by calculating the autocorrelation coefficient; partial autocorrelation analysis is used to determine the lag relationship of the difference data.
[0014] Preferably, this method is applied to data analysis and prediction in the fields of finance, industry, and healthcare, providing reliable and efficient data support and decision-making references for each field.
[0015] The integrated data analysis and reasonable prediction method of the present invention achieves significant and multifaceted beneficial effects by incorporating multiple advanced algorithms, as detailed below: This invention combines a weighted grey relational analysis algorithm with the Analytic Hierarchy Process (AHP) to accurately assess the degree of correlation between various data factors and ensures the accuracy of weight settings through consistency checks. By calculating the grey relational coefficient and weighted grey relational degree, and averaging the data over many years to avoid randomness, it effectively uncovers key information hidden in massive amounts of data, laying a solid foundation for subsequent analysis and avoiding decision-making biases caused by information omissions. This corresponds to the relevant steps of the weighted grey relational analysis algorithm in the claims.
[0016] In this invention, the GM(1,1) model excels at handling small-sample, information-poor data, providing a basic framework for prediction; the ARIMA model, on the other hand, leverages its powerful time series analysis capabilities, further optimizing the prediction data through steps such as ADF test, autocorrelation analysis, and partial autocorrelation analysis. The sequential application of these two models achieves complementary advantages, greatly improving the accuracy and reliability of the prediction results, and better adapting to different types and scales of data prediction scenarios, consistent with the prediction and analysis steps of the GM(1,1) model and the ARIMA model described in the claims.
[0017] The decision tree analysis method in this invention analyzes data from multiple dimensions, constructing an intuitive and clear decision path by calculating indicators such as information entropy, information gain, gain ratio, and Gini coefficient. After pruning and generating a decision tree architecture diagram, the comprehensiveness and accuracy of the data analysis are further ensured, consistent with the relevant content of the decision tree analysis method in the claims.
[0018] This invention's overall method, through the collaborative work of multiple algorithms, fully leverages the advantages of each algorithm, overcoming the bottlenecks of single algorithms and forming a comprehensive, efficient, and accurate data analysis and prediction system. It not only improves data processing efficiency and reduces decision-making risks, but also provides more forward-looking and scientific decision-making basis for multiple fields such as finance, industry, and healthcare, helping various industries achieve intelligent and precise development, demonstrating the innovation and practicality of this invention in the field of data processing. Attached Figure Description
[0019] Figure 1 This is a schematic diagram illustrating the principle and logic of the present invention. Detailed Implementation
[0020] The integrated data analysis and reasonable prediction method of the present invention will be described in detail below with reference to the accompanying drawings, so that those skilled in the art can implement the present invention based on the description.
[0021] like Figure 1 As shown, the integrated data analysis and reasonable prediction method of the present invention mainly includes steps such as weighted grey relational analysis, GM(1,1) prediction analysis, ARIMA prediction analysis, and decision tree analysis. These steps work together to achieve comprehensive data analysis and reasonable prediction. Specifically, the method includes the following steps: S1. Use the weighted grey relational degree algorithm to perform weight analysis and key information mining on the data. Obtain data weights and perform consistency checks through the AHP hierarchical analysis method, and calculate the grey relational coefficient and weighted grey relational degree. S2. Use the GM(1,1) model to perform preliminary predictive analysis on the data after weighted grey relational analysis, including level ratio test and establishment of time series prediction equation; S3. Use the ARIMA model to perform secondary predictive analysis on the data. Determine the stationarity and lag relationship of the time series through ADF test, autocorrelation analysis and partial autocorrelation analysis, and optimize the predictive data. S4. The data is comprehensively analyzed using decision tree analysis methods to construct decision paths, including calculating information entropy, information gain, gain ratio and Gini coefficient, performing pruning and generating a decision tree architecture diagram. S5. Finally, draw data analysis conclusions to provide a basis for decision-making.
[0022] Weighted grey relational analysis in S1 First, the original data is processed using the Analytic Hierarchy Process (AHP) to obtain the corresponding data weights. Suppose there are multiple indicators, and the weights are represented as where represents the weight of the i-th indicator. The specific steps are as follows: Calculate the average value of each analysis item, and construct a judgment matrix based on these average values, as shown in formula (1) below. The construction of the judgment matrix is the foundation of the AHP (Analytical Hierarchy Process) method. It reflects the comparison of the relative importance of each indicator. More specifically, it uses the AHP method to process the data and obtain the corresponding data weights. Let the weights be... ,in For the first The weights corresponding to each indicator are calculated first. Then, the average value of each analysis item is used to construct a judgment matrix. as follows.
[0023] (1) To ensure the accuracy of the weight settings, it is necessary to perform a consistency check on the results and determine their feasibility.
[0024] To ensure the accuracy of the weight settings, a consistency check is required. The consistency index is calculated using formula (2), and the consistency ratio is then solved using formula (3).
[0025] (2) The calculation shows that the solution can then be obtained using equation (3). size.
[0026] (3) If the calculated consistency ratio is less than the set threshold (usually 0.1), the weight setting is considered reasonable and usable; otherwise, the judgment matrix needs to be adjusted and recalculated until the consistency requirements are met.
[0027] Therefore, the final calculation The value is used to determine whether it meets the consistency test and whether the weight is available. If the weight is available, the grey relational coefficient is calculated using formula (4). Where, is the correlation coefficient between the comparison series and the reference series on the i-th index, is the resolution coefficient, the value of which affects the resolution and can usually be set according to the actual situation, and and are the minimum difference and maximum difference of the two levels, respectively.
[0028] Calculate the grey relational coefficient using equation (4): (4) in To compare sequences For the reference sequence In the Correlation coefficients on each indicator The resolution coefficient, and These represent the minimum difference and the maximum difference at two levels, respectively. Here, we assume... The larger the resolution coefficient, the greater the resolution.
[0029] Finally, the weighted grey relational degree is obtained according to formula (5). Here, the grey weighted relational degree of each evaluation index object is calculated, and to avoid randomness, the data of each selected year are summed and averaged. Through weighted grey relational degree analysis, the degree of correlation between each data factor can be accurately considered, and reasonable weights can be assigned to it to uncover the potential key information of the data, as follows: The weighted grey relational degree is obtained according to the following formula (5): (5) in For the first The grey weighted correlation degree of each evaluation index object is calculated using the AHP (Analytic Hierarchy Process) combined with weights. To avoid randomness, the data of each selected year are summed and the average is calculated.
[0030] GM(1,1) prediction analysis in S2 After completing the weighted grey relational analysis, the GM(1,1) model is used for preliminary predictive analysis of the data. The GM(1,1) model is good at handling small samples and data with limited information, providing a basic framework for prediction. The specific steps are as follows: Perform a grade ratio test and establish a time series, as shown in formula (6). Then calculate the grade ratio according to formula (7), where represents the year. The grade ratio test is to determine whether the data meets the conditions for using the GM(1,1) model, as follows: A grade ratio test was conducted, and the time series of domestic sewage discharge was established as follows. The grade ratios were then calculated. (6) (7) in, Indicates the first Year, used Satisfactory GM(1,1) algorithm prediction was performed, and the prediction algorithm and relative error test equation were obtained.
[0031] If the grade ratio falls within the acceptable coverage range, the GM(1,1) algorithm can be used for prediction. This is achieved through the formula: (8) The prediction algorithm is obtained, and this formula describes the prediction equation form of the GM(1,1) model. Simultaneously, using the formula: (9) A relative error test was performed to assess the accuracy of the prediction results. The preliminary predictions from the GM(1,1) model provide a data foundation for further analysis.
[0032] ARIMA Prediction Analysis in S3 To ensure the accuracy of the prediction results, the ARIMA model is introduced for secondary prediction. The ARIMA algorithm considers the autocorrelation and moving average characteristics of time series data, and can capture the inherent trends, periodicity, and other features of the data. The specific implementation process is as follows: First, an ADF test is performed to check whether the time series observations are stationary. Let time be , be the observation at that time, be the regression coefficient, be the time trend, be the intercept, and be the error for testing the existence of the error, as shown in formula (10). If the test result is not stationary, the data needs to be differencing. More specifically, the ARIMA algorithm considers the autocorrelation and moving average characteristics in the time series data. It captures the inherent trends, periodicity, and other characteristics in the data by generating future predictions through a linear combination of past observations.
[0033] To test whether the time series observations are stationary, an ADF test is first required. Let time be... , For this time observation value, For regression coefficients, For time trends, The intercept is... To verify the existence of errors.
[0034] (10) Autocorrelation analysis is performed using formula (11) to measure the correlation between the difference data and to determine whether the optimal difference sequence has stability and trend. Here, represents the distance between any two data in the sequence, and represents the data in the time series. More specifically, in order to measure the correlation between the difference data and to determine whether the optimal difference sequence has stability and trend, autocorrelation analysis is performed using formula (11). (11) in, This represents the distance between any two data points in the sequence. The time series is represented as The data.
[0035] In addition, to determine the lag relationship of the differenced data, the formula is used: (12) Partial autocorrelation analysis was performed, in which This represents the expectation of the data. Based on the results of autocorrelation analysis and partial autocorrelation analysis, the parameters of the ARIMA model are determined, and then the ARIMA model is used for prediction to further optimize the predicted data.
[0036] Decision tree analysis in S4 Finally, the data is comprehensively analyzed using decision tree analysis to construct a clear decision path. The specific steps are as follows: Let be a set of sample data, and be time. Assume that the proportion of samples of type in the current sample set is . Calculate the information entropy according to formula (13). The smaller the value of the information entropy, the higher the purity of the sample set, as follows: set up yes A collection of sample data, Let time be the time period, and assume the current sample set is... In the middle, the first The proportion of class samples is ,but Information entropy is defined as: (13) when The smaller the value, the better. The higher the purity.
[0037] Let the attribute be discrete, and assume it has different values. Using the partitioning of the sample set, multiple branches will be generated, where the node of the first branch is . Calculate the information gain according to formula (14). The larger the information gain, the greater the improvement in purity obtained by partitioning the sample set. Then, the gain ratio is calculated as shown in formula (15), taking into full account the diversity of attribute values, where is the inherent value of the attribute. The more possible values of the attribute, the larger the value of is usually, as shown in formula (16). The specific details are as follows: set up Assuming discrete properties, have Different values ,use For sample set Dividing the data will result in multiple branches, among which the first branch... The branch nodes are Information gain is a metric used in decision trees to select the best attribute for splitting, and it is defined as follows: (14) The larger, For sample set The greater the increase in purity obtained by the division, the gain rate was then calculated, taking into full account the diversity of attribute values.
[0038] (15) For attributes The inherent value, attribute The more possible values a value has, the better. The larger the value, the better. The definition is as follows: (16) The Gini coefficient is used to measure and determine the impurity of data, as shown in the formula: (17) Finally, the results of each branch are pruned, and a decision tree architecture diagram is generated using image rendering methods. Analyzing the decision tree architecture diagram allows for multi-dimensional data analysis, further ensuring the comprehensiveness and accuracy of the data analysis.
[0039] After a series of steps including weighted grey relational analysis, GM(1,1) predictive analysis, ARIMA predictive analysis, and decision tree analysis, the final data analysis conclusions are obtained, providing reliable and efficient data support and decision-making references for various industries. Those skilled in the art can apply the method of this invention to multiple fields such as finance, industry, and healthcare according to actual needs to achieve intelligent and precise decision-making.
[0040] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for integrating data analysis and rational prediction, characterized in that, Includes the following steps: S1. Use the weighted grey relational degree algorithm to perform weight analysis and key information mining on the data. Obtain data weights and perform consistency checks through the AHP hierarchical analysis method. Calculate the grey relational coefficient and weighted grey relational degree. S2. Use the GM(1,1) model to perform preliminary predictive analysis on the data after weighted grey relational analysis, including level ratio test and establishment of time series prediction equation; S3. Use the ARIMA model to perform secondary predictive analysis on the data. Determine the stationarity and lag relationship of the time series through ADF test, autocorrelation analysis and partial autocorrelation analysis, and optimize the predictive data. S4. The data is comprehensively analyzed using decision tree analysis methods to construct decision paths, including calculating information entropy, information gain, gain ratio and Gini coefficient, performing pruning and generating a decision tree architecture diagram. S5. Finally, draw data analysis conclusions to provide a basis for decision-making.
2. The method for integrated data analysis and reasonable prediction according to claim 1, characterized in that, In the S1 weighted grey relational algorithm, the steps for obtaining data weights using the AHP (Analytic Hierarchy Process) include: Calculate the average value of each analysis item and construct the judgment matrix; Perform consistency checks and calculate consistency indices and consistency ratios; Calculate the grey relational coefficient, and consider the resolution coefficient to adjust the resolution; The grey weighted correlation degree of each evaluation index object is calculated according to the weighted grey correlation degree formula, and the average value is accumulated to avoid randomness.
3. The method for integrated data analysis and reasonable prediction according to claim 1, characterized in that, The preliminary prediction and analysis steps of the GM(1,1) model in S2 include: Perform a grade ratio test and establish a time series; Solve the prediction equation to obtain the predicted value and the relative error test equation.
4. The method for integrated data analysis and reasonable prediction according to claim 1, characterized in that, The secondary prediction analysis steps of the ARIMA model in S3 include: Perform the ADF test to examine the stationarity of the time series; The optimal difference sequence and lag relationship are determined using autocorrelation analysis and partial autocorrelation analysis. Use the ARIMA model for prediction to capture trends, periodicity, and other characteristics in the data.
5. The method for integrated data analysis and reasonable prediction according to claim 1, characterized in that, The decision tree analysis method in S4 includes: Calculate the information entropy of the sample set, and define information gain and gain ratio; The Gini coefficient is used to measure data impurity. Pruning is performed to generate a decision tree architecture diagram, which is then analyzed.
6. The method for integrated data analysis and reasonable prediction according to claim 1, characterized in that, The weighted grey relational algorithm, GM(1,1) model, ARIMA model and decision tree analysis method are applied sequentially in the data processing process to achieve multi-algorithm fusion and collaborative optimization.
7. The method for integrated data analysis and reasonable prediction according to claim 2, characterized in that, In the consistency test, a consistency index and a consistency ratio are calculated. When the consistency ratio is less than a set threshold, the weight setting is considered reasonable and usable.
8. The method for integrated data analysis and reasonable prediction according to claim 3, characterized in that, In the process of solving the prediction equation of the GM(1,1) model, the predicted value is calculated by the given formula, and the prediction accuracy is checked by the relative error test equation.
9. The method for integrated data analysis and reasonable prediction according to claim 4, characterized in that, In the autocorrelation analysis of the ARIMA model, the stability and trend of the optimal difference sequence are determined by calculating the autocorrelation coefficient; partial autocorrelation analysis is used to determine the lag relationship of the difference data.
10. The method for integrated data analysis and reasonable prediction according to any one of claims 1-9, characterized in that, This method is applied to data analysis and forecasting in the financial, industrial, and medical fields, providing reliable and efficient data support and decision-making references for each field.