Regional Digital Economy Data Processing Method Based on ARIMA and BP Neural Network

By combining the ARIMA model and BP neural network, the Lagrangian interpolation method is used to fill in the missing data and build a rolling input data matrix to train the BP neural network, which solves the accuracy problem of regional digital economy data processing and achieves a more efficient data processing effect.

CN116244551BActive Publication Date: 2025-07-25FUZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310032908.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-10
Publication Date
2025-07-25
Estimated Expiration
2043-01-10

AI Technical Summary

Technical Problem

The existing technology has not yet formed a mature and reliable regional digital economy data processing method, and cannot effectively improve data processing accuracy.

Method used

Using a combination of ARIMA model and BP neural network, the missing data is filled by Lagrangian interpolation method, the ARIMA model is constructed and the BP neural network is trained using a mobile window to generate a rolling input data matrix, and finally the data is processed through the error correction method.

Benefits of technology

Improve data processing accuracy and improve the accuracy and reliability of data results, especially in the case of missing data processing, which significantly improves the data processing effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116244551B_ABST
    Figure CN116244551B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for processing regional digital economy data based on ARIMA and BP neural networks, comprising the following steps: Step S1: Obtain the dataset of the system database, and screen out normal data and missing data; Step S2: Fill in the missing data using the Lagrange interpolation method, and combine and process the filled data with the normal data to obtain an exclusive dataset; Step S3: Introduce three criteria, namely AIC, BIC, and HQ, to determine the model order and construct an ARIMA model; Step S4: Generate a rolling input data matrix for training through a moving window method to construct a BP neural network; Step S5: Based on the ARIMA model and the BP neural network, input the exclusive dataset to obtain the data processing result. The present invention effectively improves the data processing accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing, and particularly to a method for processing regional digital economy data based on ARIMA and BP neural networks. Background Art

[0002] The digital economy is a new economic form following the agricultural economy and the industrial economy. Therefore, it is necessary to establish a suitable model through scientific and rigorous methods to process regional digital economy data, so as to better formulate policies and take measures to promote the stable development of the regional digital economy. Regarding the problem of regional digital economy data processing, there is currently no mature and reliable technical solution, and there is an urgent need to study scientific and effective data processing technologies. Summary of the Invention

[0003] In view of this, the purpose of the present invention is to provide a method for processing regional digital economy data based on ARIMA and BP neural networks, aiming to solve the above problems.

[0004] To achieve the above purpose, the present invention adopts the following technical solutions:

[0005] A method for processing regional digital economy data based on ARIMA and BP neural networks includes the following steps:

[0006] Step S1: Obtain the dataset of the system database, and screen out the normal data and the missing data;

[0007] Step S2: Fill the missing data using the Lagrange interpolation method, and combine and process the filled data with the normal data to obtain an exclusive dataset;

[0008] Step S3: Introduce three criteria of AIC, BIC, and HQ to determine the model order, and construct an ARIMA model;

[0009] Step S4: Generate a rolling input data matrix for training by means of a moving window, and construct a BP neural network;

[0010] Step S5: Based on the ARIMA model and the BP neural network, input the exclusive dataset to obtain the data processing result.

[0011] Further, the filling of the missing data using the Lagrange interpolation method is specifically as follows:

[0012] For the known k + 1 groups of data [(x0, y0),...,(x k , y y )], its Lagrange interpolation polynomial is:

[0013]

[0014] where y i is the dependent variable;

[0015] where l j (x) represents the Lagrange interpolation basis function, and the expression is:

[0016]

[0017] Further, the specific steps of step S3 are as follows:

[0018] Step S31: Determine the autoregressive order p and the moving average order q in the ARIMA model through the ACF and PACF of the second-order difference sequence;

[0019] Step S32: Further select the most objective and appropriate p and q parameter values by using the AIC, BIC, and HQ criteria;

[0020] Step S33: Determine whether the selected ARIMA model can extract almost all the sample correlation information in the sequence and whether the selected ARIMA model is available by analyzing the Q-Q plot and Ljung-Box plot of the residual sequence.

[0021] Further, the specific steps of step S31 are as follows:

[0022] The formula for the autocorrelation function ACF is as follows:

[0023]

[0024] If the sequence is stationary, there is Then:

[0025]

[0026] where σ x represents the variance, and γ k and γ o represent the covariance;

[0027] where k represents the distance between random variables;

[0028] The formula for the partial autocorrelation function PACF of the second-order difference sequence of the regional digital economy is as follows:

[0029]

[0030] where EX t = E[X t |X t-1 ,...,X t-k+1 , EX t-k = E[X t-k |X t-1 ,...,Xt-k+1 represents the conditional expectation.

[0031] Furthermore, the AIC, BIC, and HQ criteria are specifically as follows:

[0032] ① AIC criterion:

[0033] AIC = -2ln(L) + 2K (8)

[0034] where k represents the number of model parameters and L represents the maximum likelihood function;

[0035] ② BIC criterion:

[0036] BIC = -2ln(L) + ln(n) * K (9)

[0037] where n represents the number of samples;

[0038] ③ HQ criterion:

[0039] HQ = -2ln(L) + ln(ln(n)) * K (10).

[0040] Furthermore, the specific step S4 is as follows: Generate a rolling input data matrix in a moving window manner. Select n values of the time series as a group of input data in sequence, and the subsequent m values as output data. Then, N data will slide to generate N - (n + m) + 1 groups of samples. Through training, establish the mapping relationship between the first n values and the subsequent m values;

[0041] Regarding the number of hidden layers and nodes, considering the training sample size and the problem of model overfitting, set the number of hidden layers to 1, and the number of nodes is determined to be within [3, 12] according to the principle of a ∈ [1, 10]; The activation function of the hidden layer selects the Sigmoid function, the activation function of the output layer selects the Purelin linear function, the learning rate is 0.05, the target error is 0.001, the training method is the gradient descent method, and the performance function is the residual sum of squares (RSS);

[0042] Based on the above settings, construct a BP neural network.

[0043] Furthermore, the specific step S5 is as follows: Add the processed values of the ARIMA model and the BP neural network through the error correction method to obtain the final processed result of the regional digital economy data.

[0044] A regional digital economy data processing system based on ARIMA and BP neural network, characterized in that it includes a processor, a memory, and a computer program stored on the memory. When the processor executes the computer program, it specifically executes the steps in the above-mentioned regional digital economy data processing method based on ARIMA and BP neural network.

[0045] The present invention has the following beneficial effects compared with the prior art:

[0046] The present invention combines the ARIMA model and the BP neural network model to further extract more useful information from the original data, and proposes to use the error correction method to effectively integrate the two single models to form a combined data processing method, which effectively improves the data processing accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 is the flowchart of the method of the present invention;

[0048] Figure 2 is the ACF graph of the second-order difference sequence of the regional digital economy data in the example of the present invention;

[0049] Figure 3 is the PACF graph of the second-order difference sequence of the regional digital economy data in the example of the present invention;

[0050] Figure 4 is the BP neural network structure diagram for processing regional digital economy data in the example of the present invention;

[0051] Figure 5 is the schematic diagram of the model combination based on the error correction method. DETAILED DESCRIPTION OF THE INVENTION

[0052] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0053] Please refer to Figure 1 , the present invention provides a method for processing regional digital economy data based on ARIMA and BP neural network, including the following steps:

[0054] Step S1: Obtain the data set of the system database, and screen out the normal data and the missing data;

[0055] Step S2: Fill in the missing data using the Lagrange interpolation method, and combine and process the filled data with the normal data to obtain a dedicated data set;

[0056] Step S3: Introduce three criteria of AIC, BIC, and HQ to determine the model order, and construct an ARIMA model;

[0057] Step S4: Generate a rolling input data matrix for training by moving the window, and construct a BP neural network;

[0058] Step S5: Based on the ARIMA model and the BP neural network, input the exclusive data set to obtain the data processing result.

[0059] In this embodiment, the regional digital economy data from 1993 to 2020 is used as the research object, among which the data from 1993 to 2017 is used to fit the data processing model, and the data from 2018 to 2020 is used to test the accuracy of the data processing model. An ARIMA model and a BP neural network model are established, and the linear and non-linear advantages of the ARIMA model and the BP neural network model are combined through the error correction method and applied to the processing of regional digital economy data.

[0060] In this embodiment, the Lagrange interpolation method is used to fill in the missing data in the data set. Due to reasons such as data collection failure, human statistical errors, and equipment aging and damage, the phenomenon of data missing is very common in practical applications. Ignoring or deleting missing data will reduce the amount of information of the original valuable data to a certain extent, reduce the actual data processing effect, and even affect the development of research work. For data missing, according to the size, nature, and filling effect of the missing amount, the present invention uses the Lagrange interpolation method to fill in the missing data, and the specific formula is as follows:

[0061] For the known k + 1 groups of data [(x0, y0),...,(x k , y y )], its Lagrange interpolation polynomial is:

[0062]

[0063] Among them, y i is the dependent variable;

[0064] Among them, l j (x) represents the Lagrange interpolation basis function, and the expression is:

[0065]

[0066] At the same time, in order to eliminate the differences in dimension or value range between features, accelerate the convergence speed of network training, and improve the accuracy of data results, it is necessary to standardize the cleaned data in advance. Considering the problem of regional digital economy data processing studied in the present invention, the actual data collected, and the operation specifications of relevant networks, it is decided to use Min-Max standardization in the BP neural network to process the data.

[0067] Deviation standardization (Min-Max standardization):

[0068]

[0069] In this embodiment, to verify whether the time series data is a wide-sense stationary time series, the regional digital economy data from 1993 to 2017 has no seasonality or periodicity, but shows an obvious long-term upward trend, belonging to a non-stationary time series. To eliminate the existing long-term increasing trend, it is necessary to perform stationarity processing on the original regional digital economy data using difference operation. The difference operation formula is as follows:

[0070]

[0071] where B is the lag operator, X t is the value at time t, and X t-1 is the time series value at time t - 1;

[0072] The second-order difference sequence of the regional digital economy data obtained through the second-order difference operation belongs to a stationary time series, satisfying the ARIMA model establishment principle, and data processing and fitting can be performed. Thus, it can be determined that the difference order d = 2 in the ARIMA(p, d, q) model.

[0073] In this embodiment, the formula for the autocorrelation function ACF of the second-order difference sequence of the regional digital economy is as follows:

[0074]

[0075] If the sequence is stationary, there is Then:

[0076]

[0077] where σ x represents the variance, and γ k and γ o represent the covariance;

[0078] where k represents the distance between random variables, usually called the lag. It can be seen from the above formula that the autocorrelation function reflects the correlation degree between the current period and the lag period of the time series. The calculation result range of the function is [-1, 1]. When ρ k , tending to 1 indicates that the degree of positive or negative correlation gradually strengthens; when ρ k , tending to 0 indicates that the correlation degree gradually weakens; when ρ k = 0, it indicates no correlation.

[0079] The formula for the partial autocorrelation function PACF of the second-order difference sequence of the regional digital economy is as follows:

[0080]

[0081] where EXt = E[X t |X t-1 ,...,X t-k+1 , EX t-k = E[X t-k |X t-1 ,...,X t-k+1 represents the conditional expectation.

[0082] As can be seen from the above formula, the partial autocorrelation function reflects the degree of correlation between the current period and the k-th lag period of the time series after removing the interference of the middle k - 1 random variables. The calculation result range of the function is [-1, 1]. When φ k tends to ±1, it indicates that the degree of positive or negative correlation gradually strengthens; when φ k tends to 0, it indicates that the degree of correlation gradually weakens; when φ k = 0, it indicates no correlation.

[0083] The autocorrelation function and the partial autocorrelation function are usually used for the preliminary parameter selection of time series models. The approximate range of parameters can be determined by judging the way the coefficients tend to zero. There are two ways for the coefficients to tend to zero: truncation and trailing:

[0084] Truncation: There exists a positive integer N. When k > N, it is always true that ρ k / φ k = 0.

[0085] Trailing: No matter what value k takes, ρ k / φ k is not always equal to zero and monotonically decreases or oscillates and decays at an exponential rate.

[0086] As Figure 2 and Figure 3 shown, it is found that both the ACF graph and the PACF graph of the second-order difference sequence of the regional digital economy fall completely within the standard deviation after a 1-lag delay, and neither of them completely decays to 0. It can be judged that both the ACF graph and the PACF graph show a 1-lag trailing phenomenon. Therefore, it is preliminarily judged that the p parameter in the ARIMA(p, d, q) model can take 0 or 1, and the q parameter can take 0 and 1.

[0087] After combination, three alternative models, ARIMA(1, 2, 0), ARIMA(1, 2, 1), and ARIMA(0, 2, 1), can be obtained.

[0088] In this embodiment, the AIC, BIC, and HQ criteria are specifically:

[0089] ① AIC criterion:

[0090] AIC = -2ln(L) + 2K (8)

[0091] where k represents the number of model parameters and L represents the maximum likelihood function;

[0092] ② BIC criterion:

[0093] BIC = -2ln(L) + ln(n) * K (9)

[0094] where n represents the number of samples;

[0095] ③ HQ criterion:

[0096] HQ = -2ln(L) + ln(ln(n)) * K (10)

[0097] To select the most objective and appropriate p and q parameters, this paper further adopts the AIC, BIC, and HQ criteria to further determine the parameters. When choosing the ARIMA(1,2,0) model, the AIC, BIC, and HQ values of the model are all the smallest. That is, compared with the other two alternative models, the ARIMA(1,2,0) model has the best balance between complexity and the ability to describe the dataset. Therefore, this paper selects the ARIMA(1,2,0) model to process the second-order difference sequence of regional digital economy.

[0098] In this embodiment, the ARIMA(1,2,0) is tested. By analyzing the Q-Q plot and Ljung-Box statistical table of its residual sequence, it is judged whether the selected ARIMA(1,2,0) model extracts almost all the sample correlation information in the sequence. After calculation, it can be seen that after the residual sequence of regional digital economy is fitted by the ARIMA(1,2,0) model, the P-values are all much greater than 0.05 after delays of 6 and 12 orders, that is, the null hypothesis that the residual sequence is a white noise sequence cannot be rejected. Combining the above two methods, it can be determined that the residual sequence of the ARIMA(1,2,0) model fitting the regional digital economy data has no significant correlation and belongs to a white noise sequence, that is, the model passes the test.

[0099] In this embodiment, a moving window method is used to generate a rolling input data matrix. n values of the time series are sequentially selected as a group of input data, and the subsequent m values are used as output data. Then N data will slide to generate N-(n + m)+1 groups of samples. Through training, a mapping relationship between the first n values and the subsequent m values is established. Considering the data volume of regional digital economy, this paper sets the input node n to 3 and the output node m to 1. For the number of hidden layers and nodes, considering the training sample size and the problem of model overfitting, the number of hidden layers is set to 1, and the number of nodes is based on The principle of a ∈ [1, 10] is set to [3, 12]. After multiple debuggings, when the number of hidden layer nodes is set to 6, the fitting effect is the best. In terms of other parameters, the activation function of the hidden layer selects the Sigmoid function, the activation function of the output layer selects the Purelin linear function, the learning rate is 0.05, the target error is 0.001, the training method is the gradient descent method, and the performance function is the residual sum of squares (RSS). Based on the above settings, a BP neural network for regional digital economy data processing is constructed as Figure 4 shown.

[0100] In this embodiment, the schematic diagram of the combined model based on the error correction method is as Figure 5 shown. First, use the ARIMA(1, 2, 0) model to process and fit the original regional digital economy data. Its result contains the linear law in the data sequence, and the non-linear information is retained in the residuals. Then, utilize the powerful mining ability of the BP neural network to process and fit the residual terms generated by the ARIMA(1, 2, 0) model to obtain the non-linear law of the data sequence. Finally, add the processing results of the ARIMA(1, 2, 0) model and the BP neural network to obtain the final processing result of the regional digital economy data.

[0101] Through MAPE, MAE, and R 2 to test the performance of the combined data processing model, the MAE is 1296.2, and R 2 is 0.99, close to 1, and the MAPE is only 0.75%. Both the processing and fitting performances reach the ideal situation, indicating that this method is applicable to the processing of regional digital economy data.

[0102] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0103] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, so that the instructions executed by the processors of the computer or other programmable data processing devices generate for implementation in the processFigure 1 one or more processes and / or blocks Figure 1 a device for the functions specified in one or more blocks

[0104] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction device that implements the functions in the process Figure 1 one or more processes and / or blocks Figure 1 the functions specified in one or more blocks

[0105] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the process Figure 1 one or more processes and / or blocks Figure 1 the functions specified in one or more blocks

[0106] As described above, it is only the preferred embodiment of the present invention, and it is not a limitation to the present invention in other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes. However, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention without departing from the technical solution content of the present invention still fall within the protection scope of the technical solution of the present invention.

Claims

1. A method for processing regional digital economy data based on ARIMA and BP neural network, characterized in that, It includes the following steps: Step S1: Obtain the dataset of the system database, and filter out the normal data and missing data; Step S2: Fill in the missing data using the Lagrange interpolation method, and combine and process the filled data with the normal data to obtain an exclusive dataset; Step S3: Introduce three criteria, AIC, BIC, and HQ, to determine the model order, and construct an ARIMA model; Step S4: Generate a rolling input data matrix for training in a moving window manner, and construct a BP neural network; Step S5: Based on the ARIMA model and the BP neural network, input the exclusive dataset to obtain the data processing result; The specific content of step S3 is as follows: Step S31: Determine the autoregressive order p and the moving average order q in the ARIMA model through the ACF and PACF of the second-order difference sequence; Step S32: Adopt the AIC, BIC, and HQ criteria to further select the most objective and appropriate p and q parameter values; Step S33: Judge whether the selected ARIMA model extracts almost all the sample correlation information in the sequence by analyzing the Q-Q plot and Ljung-Box plot of the residual sequence, and determine whether the selected ARIMA model is available; The specific content of step S31 is as follows: The formula of the autocorrelation function ACF is as follows: If the sequence is stationary, there is Then: where σ x represents variance, γ k and γ o represent covariance; where k represents the distance between random variables; The formula of the partial autocorrelation function PACF of the second-order difference sequence of the regional digital economy is as follows: where EX t = E[X t |X t-1 ,...,X t-k+1 , EX t-k = E[X t-k |X t-1 ,...,X t-k+1 represents the conditional expectation; The specific content of step S4 is as follows: Generate a rolling input data matrix in a moving window manner. Select n values of the time series as a group of input data in turn, and the subsequent m values as output data. Then N data will slide to generate N-(n + m)+1 groups of samples. Through training, establish the mapping relationship between the first n values and the subsequent m values; Regarding the number of hidden layers and nodes, considering the number of training samples and the problem of model overfitting, the number of hidden layers is set to 1, and the number of nodes is determined to be within [3, 12] according to the principle of a ∈ [1, 10]; the activation function of the hidden layer is selected as the Sigmoid function, the activation function of the output layer is selected as the Purelin linear function, the learning rate is 0.05, the target error is 0.001, the training method is the gradient descent method, and the performance function is the residual sum of squares (RSS); Based on the above settings, construct a BP neural network.

2. The regional digital economy data processing method based on ARIMA and BP neural network according to claim 1, wherein The specific content of filling in the missing data using the Lagrange interpolation method is as follows: For the known \(k + 1\) sets of data \([(x_0,y_0),...,(x k ,y y )]\), its Lagrange interpolation polynomial is: where y i is the dependent variable; where l j (x) represents the Lagrange interpolation basis function, and the expression is:

3. The regional digital economy data processing method based on ARIMA and BP neural network according to claim 1, wherein The AIC, BIC, and HQ criteria are specifically as follows: ① AIC criterion: AIC = -2ln(L)+2K (8) where k represents the number of model parameters, and L represents the maximum likelihood function; ② BIC criterion: BIC = -2ln(L)+ln(n)*K (9) where n represents the number of samples; ③ HQ criterion: HQ = -2ln(L)+ln(ln(n))*K (10).

4. The regional digital economy data processing method based on ARIMA and BP neural network according to claim 1, characterized in that, The specific content of step S5 is as follows: Add the processed values of the ARIMA model and the BP neural network through the error correction method to obtain the final processing result of the regional digital economy data.

5. A regional digital economy data processing system based on ARIMA and BP neural network, characterized in that, It includes a processor, a memory, and a computer program stored on the memory. When the processor executes the computer program, it specifically executes the steps in the method for processing regional digital economy data based on ARIMA and BP neural network according to any one of claims 1-4.

Citation Information

Patent Citations

  • ARIMA-BP neutral network-based bridge monitoring data prediction method

    CN106529145A

  • MIMU gyroscope random drift forecasting method based on ARMA and BPNN combination model

    CN107330149A