Public income prediction system and method based on mixed time series
By building a public income prediction system with mixed time series, combining dynamic weighting and feature fusion model, the problem of poor adaptability of existing models in multiple data sources and dynamic change scenarios is solved, and high-precision and robust public income prediction are achieved.
Patent Information
- Application Number
- CN202510294826.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-27
AI Technical Summary
When existing public revenue prediction models deal with multiple data sources and dynamic change scenarios, it is difficult to achieve high-precision prediction, and the static weight allocation strategy leads to poor model adaptability.
A public income prediction system with mixed time series is adopted to build a dynamic weight prediction model and feature fusion model through data acquisition, preprocessing, feature fusion, dynamic weight calculation, prediction and harmonization and risk assessment modules to realize the organic combination of text features and time series features.
It significantly improves prediction accuracy, enhances the adaptability and robustness of the model, can dynamically adjust the weights to adapt to different prediction scenarios, and provides credibility assessment and risk warning through the risk assessment module.
Smart Images

Figure BDA0005309690410000131 
Figure FDA0005309690400000021 
Figure FDA0005309690400000022
Abstract
Description
Technical Field
[0001] The present invention relates to the field of prediction technology, and more specifically, to a public revenue prediction system and method based on a hybrid time series. Background Art
[0002] A patent with the application publication number CN117035968A discloses a bank revenue prediction method and related device, including: decomposing a training set into a trend component, a seasonal component, and a residual component through a time series decomposition algorithm, and respectively inputting the trend component, the seasonal component, and the residual component into an LSTM neural network for training, corresponding to outputting a trend prediction, a seasonal prediction, and a residual prediction, and finally combining the trend prediction, the seasonal prediction, and the residual prediction into a prediction sequence to obtain a bank revenue prediction result. It can realize combining a group of time series in multiple aspects to predict bank revenue. In addition, the LSTM neural network has the ability of autonomous learning and can efficiently output trend prediction, seasonal prediction, and residual prediction results. Therefore, the time series decomposition algorithm combined with the LSTM neural network can significantly improve the prediction accuracy and stability of bank business revenue.
[0003] However, the combination of the time series decomposition algorithm and the LSTM neural network still has the problem of over-reliance on a single data source and is difficult to achieve high-precision prediction; currently, the public prediction field mainly relies on the autoregressive integrated moving average model and the long short-term memory network; the autoregressive integrated moving average model can only capture the linear relationship in the data and has limited processing ability for complex non-linear fluctuations; although the long short-term memory network can effectively capture the non-linear relationship in the data, it has limitations in dealing with long-term dependencies and external feature fusion. The current hybrid models generally adopt a static weight allocation strategy, and the model weights are fixed, which makes it difficult for them to adapt to the dynamic changes in the prediction scenario. The current hybrid models are still insufficient in effectively integrating multiple data sources such as policies and economic indicators in predicting public revenue data, and the accuracy of their prediction results is often low, showing a significant difference from the actual data, and the results usually show a specific value rather than an interval range.
[0004] In view of this, the present invention proposes a public revenue prediction system and method based on a hybrid time series to solve the above problems. Summary of the Invention
[0005] 1. A public revenue prediction system based on a hybrid time series, characterized by comprising: a data acquisition module for collecting historical public revenue data and policy texts in chronological order;
[0006] A data preprocessing module for preprocessing the historical public revenue data and policy texts to obtain time series data and policy text data;
[0007] A feature fusion module, which is used to extract the features of time series data to obtain the time-domain features of the time series data; extract the features of policy text data to obtain text features; fuse the obtained time-domain features and text features to obtain a weighted fusion feature vector;
[0008] A dynamic weight calculation module, which constructs a time series model and uses the time series model to extract the linear features in the time series data; constructs a long short-term memory network and uses the long short-term memory network to extract the non-linear features in the weighted fusion feature vector; calculates the dynamic weight of the time series model based on the obtained linear features and non-linear features;
[0009] A prediction reconciliation module, which constructs a hierarchical tree and hierarchically decomposes the time series data based on the hierarchical tree to obtain subsequences at each level; predicts each level of subsequences respectively to obtain prediction results at each level; performs reconciliation prediction on the prediction results at each level based on the MinT prediction reconciliation method to obtain a reconciliation prediction result;
[0010] A risk assessment module, which calculates the confidence intervals of the time series model, the long short-term memory network and the prediction reconciliation module respectively; fuses the corresponding confidence intervals based on the dynamic weight of the time series model and the preset weight of the prediction reconciliation module to obtain the prediction range of public revenue and the corresponding risk level;
[0011] A result output module, which sends the obtained prediction range of public revenue and risk level to the revenue prediction terminal.
[0012] 2. The public revenue prediction system of a hybrid time series according to claim 1, wherein the historical public revenue data includes tax revenue, social insurance fund revenue, donation revenue, fine revenue, land transfer fees, lottery public welfare funds, profits remitted by state-owned enterprises and state-owned equity transfer revenue; after filling in the missing values in the historical public revenue data, perform standardization processing to obtain time series data; extract the time-domain features obtained by extracting the features of the time series data;
[0013] The policy text includes macroeconomic analysis, fiscal policy and tax policy, public revenue structure analysis, economic structure adjustment and industrial policy, budget management and fiscal reform, and social and economic factors; remove the noise, perform word segmentation and remove stop words from the obtained policy text to obtain policy text data; extract the text features obtained by extracting the features of the policy text data; fuse the time-domain features and text features to obtain a weighted fusion feature vector.
[0014] 3. The public revenue prediction system of a hybrid time series according to claim 2, wherein the method of fusing the time-domain features and text features includes:
[0015] Extract statistical features, time-domain features, frequency-domain features, and non-linear features from time series data through Python; construct time-domain features based on statistical features, time-domain features, frequency-domain features, and non-linear features, and vectorize the time-domain features to obtain time-domain feature vectors
[0016] Use a pre-trained BERT model to extract all keywords and their sentiment tendencies from policy text data, and then obtain a set of keyword features and sentiment features, and construct a vocabulary based on the keyword features and sentiment features; input the policy text data and the vocabulary into the bag-of-words model to obtain the embedding vectors of each text; calculate the mean of the embedding vectors of each text to obtain the text feature vectors
[0017] Perform calculations based on a pre-constructed fully connected layer to obtain the attention weight α t ; According to the calculated attention weight α t Perform weighted summation on the time-domain feature vectors and text feature vectors to obtain weighted fusion feature vectors; calculate the dynamic weights of the time series model based on the weighted fusion feature vectors and time series data.
[0018] 4. A public revenue prediction system for a hybrid time series according to claim 3, wherein the method for calculating the dynamic weights of the time series model includes:
[0019] Construct a time series model, and optimize and determine the model parameters through ADF test and Bayesian information criterion; use the time series model to predict historical public revenues, and use the obtained historical public revenue prediction values as the linear features of the time series data, and then construct the feature vectors of the linear features
[0020] Use the pre-constructed Sequential model of Keras to construct a long short-term memory network; use the weighted fusion feature vectors to train different levels of the long short-term memory network, use the trained long short-term memory network to predict historical public revenues, and use the predicted historical public revenue values as the non-linear features of the long short-term memory network, and then construct the feature vectors of the non-linear features
[0021] The feature vectors of the linear features and the feature vectors of the non-linear features Are concatenated at time point t to obtain an input vector; use a pre-constructed fully connected layer to calculate the input vector to obtain the dynamic weight w of the time series model t; Use the constructed time series model and long short-term memory network to predict each level of the hierarchical tree respectively, and reconcile the predictions of each level to obtain the reconciled prediction result.
[0022] 5. A public revenue prediction system with a hybrid time series according to claim 4, wherein the method of reconciling the predictions of each level to obtain the reconciled prediction result includes:
[0023] Define the hierarchical structure; administrative dimension: country-province-city; business dimension: country-total tax-direct tax and indirect tax; decompose the time series data according to the administrative dimension and business dimension to obtain subsequences of each level;
[0024] Construct a hierarchical tree based on the administrative dimension and business dimension. Each node of the hierarchical tree represents a level, the leaf node represents the underlying time series, and the non-leaf node represents the summary layer; construct an S matrix based on the hierarchical tree, and initialize the S matrix as a matrix of all zeros; fill the S matrix based on the constructed hierarchical tree, mark the column corresponding to the leaf node as 1, and mark the row corresponding to the non-leaf node as 1 for all columns of its child nodes.
[0025] Use STL decomposition to extract features from the subsequence data of each level. If the subsequence data has a linear trend, stationarity, or a sequence that can be made stationary by differencing, construct an ARIAM model for prediction to obtain the subsequence prediction result; if the sequence has high noise, non-linear relationships, or long-term dependencies, construct a long short-term memory network for prediction to obtain the subsequence prediction result;
[0026] Based on the relationship between the underlying sequence and the upper-level summary sequence in the S matrix, use the MinT prediction reconciliation method to reconcile and predict the multi-level time series data to obtain the reconciled prediction result Fuse the corresponding confidence intervals based on the dynamic weight of the time series model and the weight of the preset prediction reconciliation module to obtain the prediction range of the public revenue and the corresponding risk level.
[0027] 6. A public revenue prediction system with a hybrid time series according to claim 5, wherein the method of obtaining the prediction range of the public revenue and the corresponding risk level includes:
[0028] Take the predicted value of the future public revenue by the time series model as the prediction mean of the time series model Use the time series model to predict the historical public revenue to obtain the prediction result; calculate the prediction variance by the prediction result of the historical public revenue Based on the prediction mean of the time series model and the prediction variance to obtain the prediction confidence interval of the time series model;
[0029] Add a Dropout layer to each layer in the constructed long short - term memory network to obtain a D - type long short - term memory network; use the D - type long short - term memory network to make M predictions on the public revenue to obtain M predicted values Calculate the M predicted values obtained to get the predicted mean of the D - type long short - term memory network And the predicted variance Take the predicted mean of the D - type long short - term memory network And the predicted variance As the predicted mean of the long short - term memory network And the predicted variance Based on the predicted mean of the long short - term memory network And the predicted variance Obtain the confidence interval of the long short - term memory network;
[0030] Take the predicted value As the predicted mean of the prediction harmonic model Use the MinT prediction harmonic module to predict the past public revenue and calculate the variance of the prediction result
[0031] Perform weighted calculation on the predicted means and variances of the time - series model, long short - term memory network, and prediction harmonic module to obtain the final predicted mean μ t And the final predicted variance σ 2 t ; Based on the predicted mean μ t And the predicted variance σ 2 t of the public revenue, calculate the confidence interval and the width width of the confidence interval of the predicted value of the public revenue;
[0032] Preset the risk assessment range [T1, T2]; where T1 is the lower limit of the confidence interval of the calculated historical public revenue, and T2 is the upper limit of the confidence interval range of the calculated historical public revenue;
[0033] If the width of the confidence interval is less than the lower limit of the confidence interval of the historical public revenue, it is a low - risk; if the width of the confidence interval is greater than the upper limit of the confidence interval of the historical public revenue, it is a high - risk; if the width of the confidence interval is between the upper and lower limits of the confidence interval of the historical public revenue, it is a medium - risk.
[0034] 7. A public revenue prediction method for a hybrid time series, which is implemented based on the public revenue prediction system for a hybrid time series according to any one of claims 1 to 6, and is characterized by including:
[0035] Step 1: Collect historical public revenue data and policy texts in chronological order;
[0036] Step 2: Preprocess the historical public revenue data and policy texts to obtain time series data and policy text data;
[0037] Step 3: Extract the features of the time series data to obtain the time domain features of the time series data; extract the features of the policy text data to obtain text features; fuse the obtained time domain features and text features to obtain a weighted fusion feature vector;
[0038] Step 4: Construct a time series model and use the time series model to extract the linear features in the time series data; construct a long short-term memory network and use the long short-term memory network to extract the non-linear features in the weighted fusion feature vector; calculate the dynamic weights of the time series model based on the obtained linear features and non-linear features;
[0039] Step 5: Construct a hierarchical tree, hierarchically decompose the time series data based on the hierarchical tree to obtain subsequences at each level; predict each level of subsequences respectively to obtain prediction results at each level; perform harmonic prediction on the prediction results at each level based on the MinT prediction harmonic method to obtain harmonic prediction results;
[0040] Step 6: Calculate the confidence intervals of the time series model, the long short-term memory network, and the prediction harmonic module respectively; fuse the corresponding confidence intervals based on the dynamic weights of the time series model and the preset weights of the prediction harmonic module to obtain the prediction range of the public revenue and the corresponding risk level;
[0041] Step 7: Send the obtained prediction range and risk level of the public revenue to the revenue prediction terminal.
[0042] Technical effects and advantages of a public revenue prediction system and method based on a hybrid time series of the present invention:
[0043] By constructing a feature fusion model, the present invention organically combines text features and time series features, fully considers the impact of policies on public revenue, and thus significantly improves the prediction accuracy of the system; by constructing an adaptive dynamic weight prediction model, the present invention realizes the dynamic allocation of model weights, effectively optimizes the collaborative efficiency between models, enables the model to dynamically adjust the contribution weights of each model according to different prediction scenarios, makes the prediction results more in line with the actual situation, and thus enhances the robustness of the system; by constructing a risk assessment module, the present invention conducts credibility assessment and risk warning on the predicted public revenue range, reminds decision-makers to pay attention to potential risks that may exist in the future period, and requires various pre-plans to be prepared in advance, further improving the practicability and reliability of the system. Description of the Drawings
[0044] Figure 1Schematic diagram of a public revenue prediction system based on a hybrid time series according to the present invention;
[0045] Figure 2 Schematic diagram of a public revenue prediction method based on a hybrid time series according to the present invention; Detailed implementation manners
[0046] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0047] Embodiment 1
[0048] Please refer to Figure 1 As shown, a public revenue prediction system and method based on a hybrid time series in this embodiment include:
[0049] A data acquisition module for collecting historical public revenue data and policy texts in chronological order;
[0050] A data preprocessing module for preprocessing the historical public revenue data and policy texts to obtain time series data and policy text data;
[0051] A feature fusion module for extracting the features of the time series data to obtain the time domain features of the time series data; extracting the features of the policy text data to obtain text features; and fusing the obtained time domain features and text features to obtain a weighted fusion feature vector;
[0052] A dynamic weight calculation module for constructing a time series model and using the time series model to extract the linear features in the time series data; constructing a long short-term memory network and using the long short-term memory network to extract the non-linear features in the weighted fusion feature vector; and calculating the dynamic weight of the time series model based on the obtained linear features and non-linear features;
[0053] A prediction reconciliation module for constructing a hierarchical tree, hierarchically decomposing the time series data based on the hierarchical tree to obtain subsequences at each level; respectively predicting the subsequences at each level to obtain prediction results at each level; and performing reconciliation prediction on the prediction results at each level based on the MinT prediction reconciliation method to obtain a reconciliation prediction result;
[0054] The risk assessment module calculates the confidence intervals of the time series model, the long short-term memory network, and the prediction reconciliation module respectively; based on the dynamic weights of the time series model and the preset weights of the prediction reconciliation module, the corresponding confidence intervals are fused to obtain the prediction range of the public revenue and the corresponding risk level;
[0055] The result output module sends the obtained prediction range of the public revenue and the risk level to the revenue prediction terminal.
[0056] Historical public revenue data includes tax revenue, non-tax revenue, social insurance fund revenue, donation revenue, interest revenue, and fine revenue, land transfer fees, lottery public welfare funds, profits remitted by state-owned enterprises, and revenue from the transfer of state-owned equity; among them, tax revenue includes value-added tax, enterprise income tax, individual income tax, consumption tax, and customs duties; non-tax revenue includes administrative fees, fines, and state-owned asset proceeds; social insurance fund revenue includes old-age insurance fund revenue, medical insurance fund revenue, and unemployment insurance fund revenue;
[0057] Policy texts include macroeconomic analysis, fiscal policies and tax policies, public revenue structure analysis, economic structure adjustment and industrial policies, budget management and fiscal reforms, and social and economic factors.
[0058] It should be clearly pointed out that the historical public revenue data is sourced from the fiscal revenue and expenditure data officially released by the National Bureau of Statistics and the Ministry of Finance, and the policy texts are from the China Government Network. Both are publicly available data.
[0059] The ways to preprocess the historical public revenue data and policy texts include:
[0060] Use linear interpolation to fill in the missing values in the historical public revenue data, fill the annual data and quarterly data into monthly data to align different-frequency data to the same time axis, and obtain the time series data of the historical public revenue; perform standardization processing on the obtained time series data to get the standardized time series data; use the first 80% of the standardized time series data as the training set and the last 20% as the test set;
[0061] Process the collected policy texts, remove the noise, perform word segmentation, and remove the stop words to obtain the policy text data.
[0062] The ways to obtain the weighted fusion feature vector include:
[0063] Extract the statistical features, time domain features, frequency domain features, and non-linear features from the time series data through pandas, scipy, and nolds in the Python library; construct time domain features based on the statistical features, time domain features, frequency domain features, and non-linear features, and vectorize the time domain features to obtain the time domain feature vector
[0064] Using the pre-trained BERT model, extract all keywords and their sentiment tendencies from the policy text data, so as to obtain a set of keyword features and sentiment features, and construct a vocabulary based on the keyword features and sentiment features;
[0065] Input the policy text data and the vocabulary into the pre-trained bag-of-words model to obtain the embedded vector representation of each text; calculate the mean value of the embedded vectors of each text to obtain the text feature vector
[0066] Calculate the attention weight α based on the pre-constructed fully connected layer t ; The calculation formula is: where, W α is the weight matrix, used to learn the relationship between features; b α is the bias vector; represents concatenating the time-domain features and text features at time point t; f() normalizes the output value into a probability distribution to ensure that the weight α t is between 0 and 1; α t reflects the importance of the time-domain features and text features at the current time point t;
[0067] According to the calculated attention weight α t perform weighted summation on the time-domain features and text features to obtain the weighted fusion feature vector. The weighted summation formula is: where α t represents the weight of the time-domain features, and 1 - α t represents the weight of the text features; for example, if the calculated value of α t is 0.54, it means that the weight of the start time-domain feature size is 0.54, and the weight of the text feature size is 0.46; sort the obtained weighted fusion feature vectors by year, and take the first 80% as training set 2 and the last 20% as test set 2.
[0068] The ways to construct the time series model include:
[0069] Construct an ARIMA model. The ARIMA model is usually expressed as ARIMA(p, d, q) and is applicable to sequences with linear trends, stationarity, or sequences that can be made stationary by differencing; where p refers to the order of the autoregressive (AR) part, representing the linear relationship between the current value and several past values; d refers to the order of differencing (I), representing the number of times the time series is differenced to make the sequence stationary (the mean and variance do not change with time); q refers to the order of the moving average (MA) part, representing the linear relationship between the current value and several past error terms.
[0070] Use the ADF test to perform stationarity analysis on the time series data to obtain the p-value of the ADF test; determine whether this p-value is lower than the established significance level (usually set at 0.05); if the p-value is lower than this significance level, it indicates that the time series data is already stationary and no further processing is required. At this time, the differencing order d can be set to 0; conversely, if the p-value is higher than the significance level, it means that the time series data has not reached a stationary state. At this time, the data must be processed accordingly to ensure that it meets the stationarity conditions;
[0071] The corresponding processing includes:
[0072] Use the differencing method on the time series data to obtain the offspring time series data, ΔX t =X t -X t-1 ; where X t represents the data at the current time point, and X t-1 represents the data at the previous time point of the current time point;
[0073] Use the ADF test to analyze the offspring time series data to obtain the p-value; if this value is lower than the significance level of 0.05, it indicates that the data is stationary and no further processing is required. At this time, the d value is 1; conversely, if the p-value is higher than 0.05, the data is not stationary and the offspring time series data needs to be differenced again; the number of times the differencing method is used in the process of making the time series data reach a stationary state is the value of d.
[0074] Use the Bayesian Information Criterion (BIC) to perform parameter screening on the p-value and q-value of the ARIMA model to determine the optimal parameter combination. The methods of parameter screening include:
[0075] Based on the obtained stationary time series data, draw the autocorrelation function graph and the partial autocorrelation function graph to observe the autocorrelation and partial autocorrelation of the time series, so as to determine the value range of the parameters p and q;
[0076] Generate all (p, d, q) combinations based on the obtained value ranges of p and q and the d value; it should be noted that the values of p and q in the combination are both integers;
[0077] Use the generated (p, d, q) parameters to fit the ARIMA model and calculate the corresponding BIC value; record the BIC value corresponding to each group of parameters, and select the parameter combination with the smallest BIC value as the optimal parameter.
[0078] It should be clearly pointed out that the BIC calculation formula is: BIC = ln(n)·k - 2ln(L); where n is the sample size, k is the number of free parameters in the model, k = p + q + 1; L is the maximum likelihood estimate of the ARIMA model; ln(n)·k represents the penalty term for model complexity, and as the number of model parameters k increases, the penalty term also increases, thus suppressing the model complexity; -2ln(L) represents the goodness of fit of the model; the larger the maximum likelihood estimate L (i.e., the better the model fits the data), the larger ln(L) is, and thus -2ln(L) is smaller; the model with the smallest BIC value is the theoretically optimal model.
[0079] Construct an ARIMA model using the best parameters (p, d, q); train the constructed ARIMA model on the training set and evaluate the trained ARIMA model using the validation set data; if the model shows excellent performance (small prediction deviation) on the training set but performs poorly (large prediction deviation) on the validation set, it indicates that the model may have overfitting; at this time, the model parameters (p, d, q) need to be readjusted until a set of parameters (p, d, q) is found such that the ARIMA model can maintain good performance on both the training set and the validation set; this set of parameters (p, d, q) is the optimal parameters of the required ARIMA model; after verification, the value of (p, d, q) is (3, 1, 2), that is, the constructed model is ARIAM(3, 1, 2);
[0080] Use the trained ARIMA model to predict the historical public revenue and obtain the predicted values where i is the index of the years included in the time series data.
[0081] The ways to construct a long short-term memory network include:
[0082] Use the pre-constructed Sequential model of Keras to construct a long short-term memory network;
[0083] Use the weighted fusion feature vector to train the long short-term memory network, and through step-by-step experiments and verifications, find the layer configuration that best suits the current needs (usually no more than 5 layers);
[0084] It should be clearly pointed out that the ways to determine the best number of layers of the long short-term memory network include:
[0085] First, start training from a single-layer long short-term memory network, use the training set 2 to train the model successively, and then use the trained model to predict the test set 2;
[0086] Calculate the mean squared error (MSE) and mean absolute error (MAE) of each layer of the long short-term memory network, and record the obtained MSE values and MAE values in detail; the calculation formula is: where y i is the actual value of historical public revenue; is the predicted value of historical public revenue by the long short-term memory network, n is the number of years included in the time series data, and i is the index of the years included in the time series data;
[0087] Conduct statistical analysis on the MSE values and MAE values of long short-term memory networks with different numbers of layers, and select the number of network layers corresponding to the minimized MSE values and MAE values to construct the long short-term memory network; after comparison, select the long short-term memory network with 2 network layers. At this time, the MSE value is 0.1868 and the MAE value is 0.2056.
[0088] Use the constructed long short-term memory network to predict historical public revenue to obtain the prediction results
[0089] The methods for calculating the dynamic weights of the time series model include:
[0090] Take the predicted value of the historical public revenue of the time series model as the linear feature of the time series data to construct a feature vector Take the predicted value of the historical public revenue of the long short-term memory network as the non-linear feature of the long short-term memory network to construct a vector Concatenate the feature vector of the time series model and the feature vector of the long short-term memory network at time point t to obtain an input vector
[0091] Based on the pre-constructed fully connected layer and calculate the dynamic weight w t , where W is the weight matrix of the fully connected layer, b is the bias vector of the fully connected layer; σ is the Sigmoid activation function, which restricts the weight between 0 and 1.
[0092] The methods for the prediction reconciliation method to conduct harmonic prediction on multi-level time series data include:
[0093] Define the hierarchical structure; administrative dimension: country-province-city; business dimension: country-total tax-direct tax and indirect tax, where direct tax includes value-added tax and enterprise income tax, and indirect tax includes consumption tax and customs duties; decompose the time series data according to the administrative dimension and business dimension to obtain subsequences at each level;
[0094] Construct a hierarchical tree based on the constructed hierarchical structure. Each node in the hierarchical tree represents a level, the leaf nodes represent the underlying time series, and the non-leaf nodes represent the aggregation layers; construct an S matrix based on the hierarchical tree. The size of the S matrix is m*n; where the number of rows m is the number of all hierarchical nodes, and the number of columns n is the number of underlying time series; initialize the S matrix as a matrix of all zeros; fill the S matrix based on the constructed hierarchical tree. The columns corresponding to the leaf nodes (underlying sequences) are marked as 1, and for the non-leaf nodes (aggregation layers), the rows corresponding to all their child nodes' columns are marked as 1.
[0095] Use STL decomposition to extract features from the subsequence data at each level. If the subsequence is a sequence with a linear trend, stationarity, or can be made stationary through differencing, then use the constructed ARIAM model for prediction to obtain the subsequence prediction result; if the subsequence is a sequence with high noise, non-linear relationships, or long-term dependencies, then use the constructed long short-term memory network for prediction to obtain the subsequence prediction result.
[0096] Based on the relationship between the underlying sequence and the upper-level aggregated sequence in the S matrix, use the MinT prediction harmonic method to perform harmonic prediction on the multi-level time series data to obtain the harmonic prediction result. The MinT prediction harmonic method is a prior art and will not be elaborated here.
[0097] The ways of fusing the confidence intervals include:
[0098] Take the predicted value of the future by the time series model As the predicted mean of the time series model Use the time series model to predict the past public revenue and calculate the variance of the prediction result. Among them, y i Is the actual value of the historical public revenue; Is the predicted value of the historical public revenue by the time series model, n is the number of years included in the time series data, and i is the index of the years included in the time series data;
[0099] Add a Dropout layer to each layer of the constructed long short-term memory network to obtain a D-type long short-term memory network; use the D-type long short-term memory network to make M predictions on the public revenue to obtain M predicted values. Calculate the predicted mean for the obtained M predicted values. And the predicted variance. Among them, j represents the index for obtaining M predicted values; Represents the prediction result obtained in the j-th run of the model.
[0100] Take the predicted value The predicted mean of the predictive harmonic model Use the MinT predictive harmonic module to predict past public revenues and calculate the variance of the prediction results
[0101] Perform weighted calculations on the predicted means and variances of the time series model, long short-term memory network, and predictive harmonic module to obtain the final predicted mean μ t And the final predicted variance σ 2 t ;
[0102]
[0103] Where θ1, θ1 are preset weights; θ1 + θ2 = 1;
[0104] Based on the average value μ of the final prediction t And the variance σ of the final prediction 2 t , calculate the confidence interval and the confidence interval width width of the predicted value of public revenue; Among them, z is the quantile corresponding to the confidence level; for example, the obtained final predicted mean μ t And the final predicted variance σ 2 t Are 100 and 25 respectively. When the confidence level is set to 95%, at this time z is 1.96, then the confidence interval is [90.2, 109.8] and the width is 19.6.
[0105] It should be clearly pointed out that the ways to obtain the corresponding risk levels include:
[0106] Repeat the above steps to obtain the predicted historical public revenue range, calculate the width of the obtained historical public revenue range, and obtain the mean μ h And variance σ 2 h ; Preset the risk assessment range [T1, T2], set the threshold T1 to μ h -σ 2 h , set the threshold T2 to μ h +σ 2 h ;
[0107] If the width of the confidence interval is less than the lower limit of the historical public revenue confidence interval, it is a low risk; if the width of the confidence interval is greater than the upper limit of the historical public revenue confidence interval, it is a high risk; if the width of the confidence interval is between the upper and lower limits of the historical public revenue confidence interval, it is a medium risk.
[0108] It should be clearly pointed out that the confidence interval represents the range of public revenue predicted by the system; the wider the confidence interval, the greater the range of fluctuations in the public revenue predicted by the system, indicating that the revenue is unstable in the future for a period of time and there are more potential risk factors in society; the narrower the confidence interval, the smaller the range of fluctuations in the public revenue predicted by the system, indicating that the revenue is stable in the future for a period of time and there are fewer potential risk factors in society.
[0109] In this embodiment, by constructing a feature fusion model, the text features and time series features are organically combined, fully considering the impact of policies on public revenue, thus significantly improving the prediction accuracy of the system; by constructing an adaptive dynamic weight prediction model, the dynamic allocation of model weights is realized, effectively optimizing the collaborative efficiency between models, enabling the model to dynamically adjust the contribution weights of each model according to different prediction scenarios, making the prediction results more in line with the actual situation, thereby enhancing the robustness of the system; by constructing a risk assessment module, the predicted values of different models are weighted and calculated to obtain the final predicted range of public revenue; by evaluating the credibility and risk warning of the final predicted range of public revenue, reminding decision-makers to pay attention to potential risks that may exist in the future for a period of time and to prepare various plans in advance, further improving the practicality and reliability of the system.
[0110] Embodiment 2
[0111] Please refer to Figure 2 As shown, for the parts not described in detail in this embodiment, refer to the description content of Embodiment 1. A public revenue prediction method based on a hybrid time series is provided, including:
[0112] Step 1: Collect historical public revenue data and policy texts in chronological order;
[0113] Step 2: Preprocess the historical public revenue data and policy texts to obtain time series data and policy text data;
[0114] Step 3: Extract the features of the time series data to obtain the time domain features of the time series data; extract the features of the policy text data to obtain text features; fuse the obtained time domain features and text features to obtain a weighted fusion feature vector;
[0115] Step 4: Construct a time series model to extract the linear features in the time series data using the time series model; construct a long short-term memory network to extract the non-linear features in the weighted fusion feature vector using the long short-term memory network; calculate the dynamic weights of the time series model based on the obtained linear features and non-linear features;
[0116] Step 5: Construct a hierarchical tree, hierarchically decompose the time series data based on the hierarchical tree to obtain subsequences at each level; predict each level of subsequences respectively to obtain prediction results at each level; perform harmonic prediction on the prediction results at each level based on the MinT prediction harmonic method to obtain a harmonic prediction result;
[0117] Step 6: Calculate the confidence intervals of the time series model, long short-term memory network, and prediction harmonic module respectively; fuse the corresponding confidence intervals based on the dynamic weight of the time series model and the preset weight of the prediction harmonic module to obtain the prediction range of the public revenue and the corresponding risk level;
[0118] Step 7: Send the obtained prediction range and risk level of the public revenue to the revenue prediction terminal.
[0119] Embodiment 3
[0120] This embodiment discloses an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the operation mode of the above-provided public revenue prediction system and method based on a hybrid time series.
[0121] Since the electronic device introduced in this embodiment is the electronic device adopted for implementing a public revenue prediction system and method based on a hybrid time series in an embodiment of the present application, based on the public revenue prediction system and method based on a hybrid time series introduced in an embodiment of the present application, those skilled in the art can understand the specific implementation manners and various variations of the electronic device in this embodiment. Therefore, the specific implementation of how this electronic device implements the method in an embodiment of the present application will not be described in detail here. As long as those skilled in the art implement the electronic device adopted for a public revenue prediction system and method based on a hybrid time series in an embodiment of the present application, it falls within the scope of protection of the present application.
[0122] The above formulas are all dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data for software simulation to obtain a formula that is closest to the actual situation. The preset parameters and threshold selection in the formulas are set by those skilled in the art according to the actual situation.
[0123] The above are only the preferred implementation manners of the present invention. The protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the idea of the present invention belong to the protection scope of the present invention. It should be noted that for ordinary technical users in the technical field, several improvements and refinements made without departing from the principle of the present invention should also be regarded as within the protection scope of the present invention.
Claims
1. A public revenue forecasting system based on mixed time series, characterized in that: include: A data collection module for collecting historical public revenue data and policy texts in chronological order; The data preprocessing module preprocesses historical public revenue data and policy texts to obtain time series data and policy text data; The feature fusion module is used to extract the features of the time series data to obtain the time domain features of the time series data; extract the features of the policy text data to obtain the text features; perform feature fusion on the obtained time domain features and text features to obtain a weighted fusion feature vector; Dynamic weight calculation module, build time series model, use time series model to extract linear features in time series data; build long short-term memory network, use long short-term memory network to extract nonlinear features in weighted fusion feature vector; Based on the obtained linear and nonlinear features, the dynamic weights of the time series model are calculated; The prediction and reconciliation module builds a hierarchical tree and decomposes the time series data hierarchically based on the hierarchical tree to obtain subsequences at each level; Predict the subsequences of each level separately to obtain the prediction results of each level; perform reconciliation prediction on the prediction results of each level based on the MinT prediction reconciliation method to obtain the reconciliation prediction results; The risk assessment module calculates the confidence intervals of the time series model, the long short-term memory network, and the forecast reconciliation module respectively; The corresponding confidence intervals are integrated based on the dynamic weights of the time series model and the weights of the preset forecast reconciliation module to obtain the forecast range of public revenue and the corresponding risk level; The result output module sends the obtained forecast range and risk level of public revenue to the revenue forecast terminal.
2. A hybrid time series public revenue forecasting system according to claim 1, characterized in that: The historical public revenue data include tax revenue, social insurance fund revenue, donation revenue, fine revenue, land transfer fee, lottery public welfare fund, profit remitted by state-owned enterprises and state-owned equity transfer revenue; after filling the missing values in the historical public revenue data, standardization processing is performed to obtain time series data; the time domain features are obtained by extracting the features of the time series data; The policy text includes macroeconomic analysis, fiscal policy and tax policy, public revenue structure analysis, economic structure adjustment and industrial policy, budget management and fiscal reform and social and economic factors; the obtained policy text is subjected to noise removal, word segmentation and stop words elimination to obtain policy text data; features in the policy text data are extracted to obtain text features; time domain features and text features are subjected to feature fusion to obtain a weighted fusion feature vector.
3. A hybrid time series public revenue forecasting system according to claim 2, characterized in that: The method of fusing the time domain features and the text features includes: Use Python to extract statistical features, time domain features, frequency domain features and nonlinear features from time series data; construct time domain features based on statistical features, time domain features, frequency domain features and nonlinear features, and quantize the time domain features to obtain the time domain feature vector Use the pre-trained BERT model to extract all keywords and their sentiment tendencies from the policy text data, and then obtain a set of keyword features and sentiment features, and build a vocabulary based on the keyword features and sentiment features; input the policy text data and vocabulary into the bag-of-words model to obtain the embedding vector of each text; calculate the mean value of the embedding vector of each text to obtain the text feature vector Based on the pre-built fully connected layer, the attention weight α is calculated t ; According to the calculated attention weight α t The time domain feature vector and the text feature vector are weightedly summed to obtain a weighted fusion feature vector; the dynamic weight of the time series model is calculated based on the weighted fusion feature vector and the time series data.
4. A hybrid time series public revenue forecasting system according to claim 3, characterized in that: The method of calculating the dynamic weight of the time series model includes: Construct a time series model, and determine the model parameters through ADF test and Bayesian information criterion optimization; use the time series model to predict historical public revenue, and use the predicted historical public revenue as the linear feature of the time series data, and then construct the feature vector of the linear feature. Use the pre-built Keras Sequential model to build a long short-term memory network; use weighted fusion feature vectors to train long short-term memory networks at different levels, use the trained long short-term memory network to predict historical public revenue, and use the predicted historical public revenue value as the nonlinear feature of the long short-term memory network, and then construct the feature vector of the nonlinear feature. The feature vector of the linear feature and the eigenvector of the nonlinear feature Concatenate at time point t to get the input vector; use the pre-built fully connected layer to calculate the input vector to get the dynamic weight w of the time series model t ; Use the constructed time series model and long short-term memory network to predict each level of the hierarchical tree separately, and reconcile the predictions of each level to obtain the harmonized prediction results.
5. A hybrid time series public revenue forecasting system according to claim 4, characterized in that: The method of reconciling the predictions of each level to obtain the reconciled prediction result includes: Define the hierarchical structure; administrative dimension: country-province-city; business dimension: country-total tax-direct tax and indirect tax; decompose the time series data according to the administrative dimension and business dimension to obtain subsequences at each level; A hierarchical tree is constructed based on the administrative dimension and the business dimension. Each node of the hierarchical tree represents a level. Leaf nodes represent the underlying time series, and non-leaf nodes represent the summary layer. An S matrix is constructed based on the hierarchical tree and initialized to an all-zero matrix. The S matrix is filled based on the constructed hierarchical tree. The columns corresponding to leaf nodes are marked as 1, and the columns corresponding to all child nodes of non-leaf nodes are marked as 1. Use STL decomposition to extract features from subsequence data at each level. If the subsequence data has a linear trend, stability, or a sequence that can be stabilized by difference, then build an ARIAM model for prediction to obtain the subsequence prediction result; if it has a sequence with high noise, nonlinear relationship, or long-term dependence, then build a long short-term memory network for prediction to obtain the subsequence prediction result; Based on the relationship between the bottom sequence and the upper summary sequence in the S matrix, the MinT forecasting reconciliation method is used to perform reconciliation forecasting on multi-level time series data to obtain the reconciliation forecasting results. The corresponding confidence intervals are integrated based on the dynamic weights of the time series model and the weights of the preset forecast reconciliation module to obtain the forecast range of public revenue and the corresponding risk level.
6. A hybrid time series public revenue forecasting system according to claim 5, characterized in that: The method of obtaining the forecast range of public revenue and the corresponding risk level includes: The predicted value of future public revenue from the time series model As the predicted mean of a time series model The prediction results are obtained by using the time series model to predict the historical public revenue; the prediction variance is obtained by calculating the prediction results of the historical public revenue Forecasting Mean Based on Time Series Model and the prediction variance Get the forecast confidence interval of the time series model; Add a Dropout layer to each layer of the constructed LSTM network to obtain a D-type LSTM network; use the D-type LSTM network to make M predictions on public revenue and obtain M prediction values Calculate the M predicted values to get the predicted mean of the D-type long short-term memory network and the prediction variance The predicted mean of the D-type long short-term memory network and the prediction variance As the predicted mean of the LSTM network and the prediction variance Prediction mean based on long short-term memory network and the prediction variance Get confidence intervals for LSTM networks; The predicted value As the forecast mean of the forecast harmonic model Use MinT forecast reconciliation module to forecast past public revenue and calculate the variance of the forecast results The prediction mean and variance of the time series model, long short-term memory network and prediction reconciliation module are weighted to obtain the final prediction mean μ t And the final prediction variance σ 2 t ; Predicted mean μ based on public revenue t and the prediction variance σ 2 t , calculate the confidence interval of the predicted value of public revenue and the width of the confidence interval; The preset risk assessment range is [T1, T2]; where T1 is the lower limit of the calculated historical public revenue confidence interval, and T2 is the upper limit of the calculated historical public revenue confidence interval; If the width of the confidence interval is less than the lower limit of the historical public revenue confidence interval, it is low risk; if the width of the confidence interval is greater than the upper limit of the historical public revenue confidence interval, it is high risk; if the width of the confidence interval is between the upper and lower limits of the historical public revenue confidence interval, it is medium risk.
7. A method for public revenue forecasting of a mixed time series, which is implemented based on a mixed time series public revenue forecasting system according to any one of claims 1 to 6, characterized in that: include: Step 1: Collect historical public revenue data and policy texts in chronological order; Step 2: Preprocess historical public revenue data and policy texts to obtain time series data and policy text data; Step 3: Extract the features of the time series data to obtain the time domain features of the time series data; extract the features of the policy text data to obtain the text features; perform feature fusion on the obtained time domain features and text features to obtain a weighted fusion feature vector; Step 4: Build a time series model and use it to extract linear features from time series data; build a long short-term memory network and use it to extract nonlinear features from weighted fusion feature vectors; Based on the obtained linear and nonlinear features, the dynamic weights of the time series model are calculated; Step 5: Construct a hierarchical tree, and decompose the time series data hierarchically based on the hierarchical tree to obtain subsequences at each level; Predict the subsequences of each level separately to obtain the prediction results of each level; perform reconciliation prediction on the prediction results of each level based on the MinT prediction reconciliation method to obtain the reconciliation prediction results; Step 6: Calculate the confidence intervals of the time series model, the long short-term memory network and the forecast reconciliation module respectively; merge the corresponding confidence intervals based on the dynamic weight of the time series model and the preset weight of the forecast reconciliation module to obtain the forecast range of public revenue and the corresponding risk level; Step 7: Send the obtained forecast range and risk level of public revenue to the revenue forecast terminal.
Citation Information
Patent Citations
Bank income prediction method and related device
CN117035968A