Convolution gating circulation unit stock prediction method based on limit gradient lifting feature selection and secondary decomposition
Through the convolutional gating loop unit method based on extreme gradient enhancement and quadratic decomposition, combined with the XGBoost algorithm to optimize feature selection and model parameters, a neural network framework for genetic algorithm optimization is built, which solves the noise, overfitting and interpretability problems of deep learning in stock price prediction, and improves prediction accuracy and interpretability.
Patent Information
- Application Number
- CN202510392433.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-08
AI Technical Summary
In stock price prediction, deep learning faces problems such as high noise impact, difficulty in processing missing values, many model parameters and easy to overfit, and poor interpretability. Especially when there is less sample data, it is difficult to effectively regularize the model and select hyperparameters.
The convolutional gating cyclic unit method based on extreme gradient enhancement feature selection and quadratic decomposition is adopted, and the data is decomposed into multiple IMFs through EMD, and the stock price prediction is performed using the GA-CGRU model. Combined with the XGBoost algorithm to optimize feature selection and model parameters, a convolutional gating cyclic unit neural network framework GA-XGBoost-CGRU optimized by genetic algorithm is constructed.
It improves the accuracy of stock price prediction and the interpretability of the model, reduces the overfitting phenomenon, and enhances the prediction ability under limited data.
Smart Images

Figure CN120278756A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of stock price prediction, and particularly relates to a stock prediction method based on extreme gradient boosting feature selection and quadratic decomposition for convolutional gated recurrent units. Background Art
[0002] The core of deep learning lies in neural networks, especially deep neural networks (DNNs) and convolutional neural networks. DNNs can effectively capture the non-linear features of stock prices through multi-layer network structures. For example, some studies have achieved good prediction results by using DNN models combined with historical price data and technical indicators. RNNs and their variants (such as long short-term memory networks, LSTMs) perform excellently in processing time series data. Since stock prices are time-dependent, LSTMs can effectively capture long-term dependencies, and many studies have used LSTM models for stock price prediction and achieved remarkable results. Although deep learning has made some progress in stock price prediction, it still faces some challenges: stock market data is often affected by noise and missing values, and how to improve data quality and obtain reliable data remains an important issue. Deep learning models have numerous parameters and are prone to overfitting, especially when the sample data is scarce. Therefore, how to effectively regularize the model and select appropriate hyperparameters is a key challenge. Deep learning models are usually regarded as "black boxes" and lack interpretability, which is particularly important in the financial field. Researchers need to explore how to improve the interpretability of the model in order to better understand the factors affecting stock prices. Summary of the Invention
[0003] To solve these problems, this application proposes a stock prediction method based on extreme gradient boosting feature selection and quadratic decomposition for convolutional gated recurrent units. Select a leading stock in the industry, which usually has strong competitive advantages and market share. Even with the influence of external investors, its stock price is not easily affected by obvious fluctuations. Decompose the unstable data into multiple IMFs through quadratic decomposition, and use the GA-CGRU model for stock price prediction. In this way, the problem of stock price prediction accuracy is solved.
[0004] The present invention is implemented through the following technical solutions: A stock prediction method based on extreme gradient boosting feature selection and quadratic decomposition for convolutional gated recurrent units:
[0005] Step 1) Collect historical market data of A-share stocks and perform preprocessing, extract the closing price from it, and serialize it;
[0006] Perform EMD decomposition on the time series data, determine how many IMF components to decompose the original time series data into according to the sample entropy value of the IMF, and perform secondary decomposition on the high-frequency IMF through VMD; convert the one-dimensional time series data into a two-dimensional phase space and reconstruct the phase space.
[0007] 1.1 Stock historical market data collection: Use the AKShare library to obtain the historical market data of a specific stock.
[0008] 1.2 Stock data processing:
[0009] Regarding the stock historical market data, extract the closing price from it, serialize it, and perform a stability test on the serialized data;
[0010] Extract the closing price from the stock historical data, serialize the obtained closing price and use the Adfuller method, that is, the Augmented Dickey-Fuller test for stationarity testing. The ADF test examines the stationarity of the sequence by constructing an autoregressive model. The form of the model is:
[0011]
[0012] Where X t is the value of the time series at time t, σ, β i are the regression coefficients, ∈ t is the error term, i.e., white noise;
[0013] Through the ADF test, confirm the applicability of the data before modeling, that is, proceed to the next step;
[0014] Perform secondary decomposition on the historical market data that passes the test: First, obtain all local extrema of the sequence s(t), then connect all the upper and lower extreme points respectively to obtain the envelope line. The average value between the extreme envelopes is m1(t), and h1(t) is obtained by subtracting the average value m1(t) from the sequence s(t);
[0015] h1(t) = s(t) - m1(t)
[0016] Perform a screening operation, using h1(t) as the input and repeating the above steps;
[0017] h 11 (t) = h1(t) - m 11 (t)
[0018] The screening process will be repeated i times until h 1i (t) becomes a true IMF;
[0019] h 1i (t) = h 1(i-1) (t) - m1i (t)
[0020] The stopping condition is the Cauchy convergence criterion. Specifically, the test requirement is the normalized squared deviation SD of the results of two screening operations i less than;
[0021]
[0022] When SD i is less than 0.2, c1(t) = h 1i (t), where c1(t) is the first IMF, and it is difficult for SD i to be less than 0.2. To ensure that the frequency and amplitude of the IMF have sufficient physical significance, c j (t) represents the j-th IMF, and the equation is as follows:
[0023] c1(t) = h 1i (t)
[0024] c j (t) = h ji (t)
[0025] The recurrence formula for the residual is as follows:
[0026] r1(t) = x(t) - c1(t)
[0027] r j (t) = r j-1 (t) - c j (t) (j ≥ 2)
[0028] where r j (t) is the residual after the j-th step of decomposition.
[0029] The sequence s(t) is represented by the IMF C j (t) and the decomposed residual r n (t), and the equation is as follows:
[0030]
[0031] where n is the number of IMFs obtained by decomposition;
[0032] For variational mode decomposition (VMD), the original signal s(t) is decomposed into k components, ensuring that the decomposed sequence is a modal component with a finite bandwidth centered at the frequency, and at the same time, the sum of the estimated bandwidths of each mode is minimized. The constraint condition is that the sum of all modes is equal to the original signal. Then the VMD constrained variational model is as follows:
[0033]
[0034] where, u k={u1, u2, …, u k} are the modal functions, ω k ={ω1, ω2, …, ω k} are the modal center frequencies, and δ(t) is the Dirac function;
[0035] To solve the constrained optimization problem, transform the constrained variational problem into an unconstrained variational problem, and utilize the advantages of the quadratic penalty term and the Lagrange multiplier method to introduce the augmented Lagrangian function, as shown in the formula:
[0036]
[0037] where α is the penalty parameter, λ is the Lagrangian multiplier, the quadratic penalty factor α is used to reduce the interference of Gaussian noise, and the modal components and center frequencies are obtained through the iteration of the alternating direction method of multipliers (ADMM). According to the ADMM algorithm, the update formulas for u_k, ω_k, and λ are as follows:
[0038]
[0039] where n is the number of iterations, is the corresponding Fourier transform of f(t), u(t), and λ(t), and γ is the update parameter;
[0040] Calculate the sample entropy of multiple IMFs after decomposition, determine how many IMFs to decompose the original sequence data according to the sample entropy, and perform secondary decomposition on the decomposed high-frequency IMFs using VMD:
[0041] The original signal s is decomposed into k components, ensuring that the decomposition sequence is a modal component with a finite bandwidth having a center frequency, and at the same time, the sum of the estimated bandwidths of each mode is minimized, and the constraint condition is that the sum of all modes is equal to the original signal; solve the constrained optimization problem, transform the constrained variational problem into an unconstrained variational problem, and utilize the advantages of the quadratic penalty term and the Lagrange multiplier method to introduce the augmented Lagrangian function;
[0042] Perform phase space reconstruction on the obtained multiple IMFs. Assume that the state space of the system is M and its dimension is d, and the evolution of the system is described by a smooth mapping The state of the system is represented as x(t), and assume that only a single-variable time series {x(t)} can be observed. Construct a new vector through time-delay embedding:
[0043] y(t) = [x(t), x(t + τ), x(t + 2τ), …, x(t + (m - 1)τ)]
[0044] where τ is the time delay and m is the embedding dimension.
[0045] Step 2) Construct a convolutional gated recurrent unit neural network framework based on XGBoost to predict the next price of the stock;
[0046] Calculate the feature importance of the data after phase space reconstruction by the extreme gradient boosting method, and select the top three features with the highest feature importance as input features;
[0047] The convolutional gated recurrent unit neural network framework includes: using a convolutional neural network to learn each deep feature of the IMF and residual after decomposing the historical stock price data; obtaining the predicted values of the IMF and residual for the next step through the gated recurrent unit's learning of the deep features; after obtaining the predicted values of all IMFs and residuals, merging all the predicted values to obtain the predicted value of the stock price.
[0048] 2.1 Feature Selection Based on XGBoost
[0049] Use the XGBoost algorithm to calculate the importance of features in sequence data and perform feature selection based on feature importance; use the residuals of the previous model to train the subsequent model, and the subsequent model corrects the errors generated by the previous model. XGboost uses the second-order Taylor expansion to approximate the loss function and adds a regularization term to control the model complexity in the objective function to avoid overfitting; the XGBoost feature importance selection method is one of the embedding methods. During the training of XGBoost, the gain is used to determine the best split node, and the final feature importance score is calculated by the average gain; the average gain is the total gain of all trees divided by the total number of splits for each feature. The more times a feature is split, the greater the gain of the model; calculate the gain to output the importance of each feature. The higher the feature importance score, the greater the contribution of the feature in the model construction and training process. Usually, the top-ranked features are obtained according to the descending order of the feature importance scores. The gain calculation method:
[0050]
[0051] where g i and h i represent the first-order and second-order gradients. I L and I R represent the samples of the left and right nodes after splitting respectively, I = I L ∪I R , and λ, γ are penalty parameters;
[0052] 2.2 Gated Recurrent Unit
[0053] The gated recurrent unit (GRU) has two gates, the update gate and the reset gate. The update gate combines the input gate and the forget gate of the LSTM, and the reset gate can directly process the previous hidden state. While ensuring performance, the training speed of the GRU is faster than that of the LSTM. The calculation equations are as follows:
[0054] z t =σ(W z *[h t-1 ,x t +b z )
[0055] r t =σ(W r *[h t-1 ,x t +b r )
[0056]
[0057] Where h t is the hidden layer vector in the GRU, x t is the output vector in the GRU, z t is the update gate in the GRU, r t is the reset gate in the GRU, W z is the parameter matrix of the update gate in the GRU, W r is the parameter matrix of the reset gate in the GRU, W h is the parameter matrix of the hidden layer in the GRU, b z is the bias vector of the update gate in the GRU, b r is the bias vector of the reset gate in the GRU, b h is the bias vector of the hidden layer in the GRU, σ, tanh are activation functions;
[0058] 2.3 Convolutional Gated Recurrent Unit
[0059] The convolutional layer automatically extracts features through filters. For the same convolutional layer, there are no connections between neurons and weight sharing is adopted. In the forward propagation stage, each convolutional layer applies an activation function and a convolution operation to the output of the previous layer. This feature of the CNN helps to extract hidden information and is not affected by uncertainties. The definition is as shown in the equation:
[0060]
[0061] Where represents the output after convolution, f represents the activation function, W k represents the weight of the k-th convolutional kernel, x represents the input value, b kDenote the bias of the k-th convolution operation, and i and j represent the number of convolution kernel operations required in two directions. The overall framework consists of a convolutional layer, two GRU layers, and a fully connected layer.
[0062] Step 3) Use the genetic algorithm to optimize the parameters of the convolutional gated recurrent unit model. Find the optimal parameters of the model through multiple selections, crossovers, and mutations, and import the optimal parameters into the model for prediction.
[0063] First, initialize the population and randomly generate the first-generation individuals. Each individual represents a feasible solution. Then, train each individual and calculate its fitness according to the validation set. Then, through the roulette wheel selection method, preferentially select individuals with higher fitness for genetic operations. Finally, generate the next-generation individuals through chromosome crossover and a small probability of gene mutation. After multiple selections, crossovers, and mutations, excellent individuals are retained, and individuals with lower fitness are eliminated. Eventually, a population with gradually increasing fitness is formed. This process continues until the preset termination conditions are met.
[0064] Step 4) Evaluate the prediction effect of the model through the mean square error, mean absolute error, root mean square error, and mean relative error between the predicted value and the true value.
[0065] Calculation of mean square error (MSE): The mean square error is obtained by taking the average of the squares of the differences between each predicted value and the actual value. The formula is:
[0066]
[0067] where, y i is the true value, is the predicted value, and n is the number of samples.
[0068] Calculation of mean absolute error (MAE): The mean absolute error is obtained by taking the average of the absolute differences between each predicted value and the actual value. The formula is:
[0069]
[0070] Calculation of root mean square error (RMSE): The root mean square error is the square root of the mean square error. The formula is:
[0071]
[0072] Calculation of mean average percentage error (MAPE): The mean average percentage error is obtained by calculating the average of the ratios of the absolute error between each predicted value and the actual value to the actual value. The formula is:
[0073]
[0074] The beneficial effects of the present invention are as follows: A stock prediction method based on extreme gradient boosting feature selection and quadratic decomposition for convolutional gated recurrent units provided by this application proposes a new genetic algorithm-optimized convolutional gated recurrent unit neural network framework based on XGBoost, GA-XGBoost-CGRU (Genetic Algorithm Extreme Gradient Boosting Gated Recurrent Unit, GA-XGBoost-CGRU). This framework obtains high-order hidden features in stock data and constructs models according to different IMFs after decomposition to optimize model parameters. A large number of experiments on real-world stock historical market datasets show that GA-XGBoost-CGRU can improve the accuracy of stock data prediction. Description of the Drawings
[0075] Figure 1 It is a schematic diagram of the instance model framework GA-XGBoost-CGRU of this application;
[0076] Figure 2 It is an evaluation index diagram for evaluating prediction accuracy involved in Embodiment 1;
[0077] Figure 3 It is the prediction effect diagram of Embodiment 1. Detailed Implementation Manner
[0078] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments of this application and the corresponding drawings. It should be understood that the specific embodiments described here are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0079] This application provides a new genetic algorithm-optimized convolutional gated recurrent unit neural network framework based on XGBoost for stock prediction, GA-XGBoost-CGRU (Genetic Algorithm Extreme Gradient Boosting Gated Recurrent Unit, GA-XGBoost-CGRU), as Figure 1 shown, which includes three modules, namely, a stock historical market data collection and data preprocessing module, a convolutional gated recurrent unit neural network framework module, and a genetic algorithm to optimize model parameters module. The detailed information is described as follows:
[0080] 1. Stock historical market data collection and data preprocessing module
[0081] 1.1 Stock historical market data collection
[0082] Use the AKShare library to obtain the historical market data of specific stocks.
[0083] 1.2 Stock data processing
[0084] Regarding the stock historical market data, extract the closing price from it and serialize it.
[0085] Conduct a stability test on the serialized data.
[0086] Extract the closing price from the stock historical data, serialize the obtained closing price and use the Adfuller method, i.e., the Augmented Dickey-Fuller (ADF) test for stationarity testing. The ADF test examines the stationarity of the sequence by constructing an autoregressive model. The form of the model is usually:
[0087]
[0088] where X t is the value of the time series at time t, σ, β i are the regression coefficients, ∈ t is the error term (white noise).
[0089] Many time series models (such as ARIMA) require the data to be stationary. If the data is not stationary, the predictive ability of the model may be affected. Through the ADF test, the applicability of the data can be confirmed before modeling, and thus the next step can be carried out.
[0090] Perform a second decomposition on the historical market data that passes the test.
[0091] The specific method of the second decomposition is as follows: First, obtain all the local extrema of the sequence s(t), and then connect all the upper and lower extreme points respectively to obtain the envelope line. The average value between the extreme envelopes is m1(t), and h1(t) is obtained by subtracting the average value m1(t) from the sequence s(t).
[0092] h1(t) = s(t) - m1(t)
[0093] Generally, h1(t) does not meet the definition of IMF. Therefore, a screening operation is required, using h1(t) as the input and repeating the above steps.
[0094] h 11 (t) = h1(t) - m 11 (t)
[0095] The screening process will be repeated i times until h 1i (t) becomes a true IMF.
[0096] h 1i h(t) = h 1(i-1) h(t) - m 1i h(t)
[0097] The stopping condition is the Cauchy convergence criterion. Specifically, the test requirement is the normalized squared deviation SD of the results of two screening operations i less than
[0098]
[0099] When SD i is less than 0.2, c1(t) = h 1i (t), where c1(t) is the first IMF. It is difficult for SD i to be less than 0.2, which can ensure that the frequency and amplitude of the IMF have sufficient physical significance. c j (t) represents the jth IMF. The equation is as follows:
[0100] c1(t) = h 1i (t)
[0101] c j c(t) = h ji (t)
[0102] The recurrence formula for the residual is as follows:
[0103] r1(t) = x(t) - c1(t)
[0104] r j r(t) = r j-1 (t) - c j (t) (j ≥ 2)
[0105] The sequence s(t) can be represented by the IMFC j (t) and the decomposed residual r n (t), and the equation formula is as follows:
[0106]
[0107] where n is the number of IMFs obtained by decomposition.
[0108] For variational mode decomposition (VMD), the original signal s(t) is decomposed into k components, ensuring that the decomposition sequence is a modal component with a finite bandwidth centered at the frequency, and at the same time, the sum of the estimated bandwidths of each mode is minimized. The constraint condition is that the sum of all modes is equal to the original signal. Then the VMD constrained variational model is as follows:
[0109]
[0110] where, uk = {u1, u2, …, u k} are the modal functions, ω k = {ω1, ω2, …, ω k} are the modal center frequencies, and δ(t) is the Dirac function.
[0111] To solve the constrained optimization problem, transform the constrained variational problem into an unconstrained variational problem, and utilize the advantages of the quadratic penalty term and the Lagrange multiplier method to introduce the augmented Lagrangian function, as shown in the formula:
[0112]
[0113] where α is the penalty parameter, λ is the Lagrangian multiplier, and the quadratic penalty factor α is used to reduce the interference of Gaussian noise. The modal components and center frequencies are obtained through the iteration of the Alternating Direction Method of Multipliers (ADMM). According to the ADMM algorithm, the update formulas for uk, ωk, and λ are as follows:
[0114]
[0115] where n is the number of iterations, is the corresponding Fourier transform of f(t), u(t), λ(t), and γ is the update parameter.
[0116] Calculate the sample entropy of multiple IMFs after decomposition, and determine how many IMFs to decompose the original sequence data according to the sample entropy. Perform secondary decomposition on the decomposed high-frequency IMFs using VMD. The specific decomposition method is as follows: The original signal s is decomposed into k components, ensuring that the decomposed sequence is a modal component with a finite bandwidth centered at the frequency, and at the same time, the sum of the estimated bandwidths of each mode is minimized, and the constraint condition is that the sum of all modes is equal to the original signal. Solve the constrained optimization problem, transform the constrained variational problem into an unconstrained variational problem, and utilize the advantages of the quadratic penalty term and the Lagrange multiplier method to introduce the augmented Lagrangian function.
[0117] Perform phase space reconstruction on the obtained multiple IMFs. Assume that the state space of the system is M and the dimension is d, and the evolution of the system is described by a smooth mapping . The state of the system can be expressed as x(t), and assume that we can only observe a single-variable time series {x(t)}. Then, a new vector can be constructed through time-delay embedding:
[0118] y(t) = [x(t), x(t + τ), x(t + 2τ), …, x(t + (m - 1)τ)]
[0119] where τ is the time delay and m is the embedding dimension.
[0120] 2. Convolutional Gated Recurrent Unit Neural Network Framework Based on XGBoost
[0121] 2.1 Feature Selection Based on XGBoost
[0122] The XGBoost algorithm used in this application is used to calculate the importance of features in sequence data and perform feature selection based on feature importance. The residual of the previous model is used to train the subsequent model, that is, the subsequent model can correct the errors generated by the previous model. XGboost uses second-order Taylor expansion to approximate the loss function to ensure the accuracy of the model, and obtains better generalization ability by adding a regularization term that controls the model complexity to the objective function, thus avoiding overfitting. The XGBoost feature importance selection method is one of the embedding methods. During the training of XGBoost, we use the gain to determine the best splitting node. The final feature importance score is calculated by averaging the gains. The average gain is the total gain of all trees divided by the total number of splits for each feature. The more times a feature is split, the greater the gain of the model. It outputs the importance of each feature by calculating the gain. The higher the feature importance score, the greater the contribution of the feature in the model construction and training process. We usually obtain the top-ranked features according to the descending order of the feature importance scores. Gain calculation method:
[0123]
[0124] where g i and h i represent the first-order and second-order gradients. I L and I R represent the samples of the left and right nodes after splitting respectively. I = I L ∪I R . λ and γ are penalty parameters.
[0125] 2.2 Gated Recurrent Unit
[0126] The Gated Recurrent Unit (GRU) has two gates, the update gate and the reset gate. The update gate combines the input gate and the forget gate of the LSTM, while the reset gate can directly process the previous hidden state. The training speed of the GRU is faster than that of the LSTM while ensuring performance. The calculation equations are as follows:
[0127] z t = σ(W z *[h t-1 , x t +b z )
[0128] r t = σ(W r *[ht-1 , x t + b r )
[0129]
[0130] where h t is the hidden layer vector in the GRU, x t is the output vector in the GRU, z t is the update gate in the GRU, r t is the reset gate in the GRU, W z is the parameter matrix of the update gate in the GRU, W r is the parameter matrix of the reset gate in the GRU, W h is the parameter matrix of the hidden layer in the GRU, b z is the bias vector of the update gate in the GRU, b r is the bias vector of the reset gate in the GRU, b h is the bias vector of the hidden layer in the GRU, σ, tanh are activation functions.
[0131] 2.3 Convolutional Gated Recurrent Unit
[0132] The convolutional layer can automatically extract features through filters (convolution kernels). For the same convolutional layer, there are no connections between neurons, and weights can be shared. In the forward propagation stage, each convolutional layer applies an activation function and a convolution operation to the output of the previous layer. This property of CNN can help extract hidden information without being affected by its uncertainty. It is defined as shown in the equation:
[0133]
[0134] where represents the output after convolution, f represents the activation function, W k represents the weight of the k-th convolution kernel, x represents the input value, b k represents the bias of the k-th convolution operation, i and j represent the number of convolution kernel operations required in two directions. The overall framework consists of a convolutional layer, two GRU layers, and a fully connected layer.
[0135] 3. Optimizing Model Parameters with Genetic Algorithm
[0136] First, initialize the population and randomly generate the first-generation individuals. Each individual (chromosome) represents a feasible solution. Then, train each individual and calculate its fitness based on the validation set. Next, through the roulette wheel selection method, preferentially select individuals with higher fitness for genetic operations. Finally, generate the next-generation individuals through chromosome crossover and low-probability gene mutation. After multiple selections, crossovers, and mutations, excellent individuals are gradually retained, while individuals with lower fitness are eliminated, and ultimately a population with gradually increasing fitness is formed. This process continues until the preset termination conditions are met.
[0137] 4. Evaluate the prediction effect of the model
[0138] The model evaluates the prediction effect of the model through the mean squared error, mean absolute error, root mean squared error, and mean relative error. The specific implementation method is as follows:
[0139] Calculation of mean squared error (MSE): The mean squared error is obtained by taking the average of the squares of the differences between each predicted value and the actual value. The formula is:
[0140]
[0141] where, y i is the true value, is the predicted value, and n is the number of samples.
[0142] Calculation of mean absolute error (MAE): The mean absolute error is obtained by taking the average of the absolute differences between each predicted value and the actual value. The formula is:
[0143]
[0144] Calculation of root mean squared error (RMSE): The root mean squared error is the square root of the mean squared error. The formula is:
[0145]
[0146] Calculation of mean average percentage error (MAPE): The mean average percentage error is obtained by calculating the average of the ratios of the absolute error between each predicted value and the actual value to the actual value. The formula is:
[0147]
[0148] Example 1: As shown in the appendix of the specification Figures 2 - 3 as shown. Figure 3Shows the prediction effect of the stock with the stock code 600030. The stock data is selected from March 1, 2018 to May 28, 2024. The first 80% is the training set, and the last 20% is the test set. The figure only shows the prediction effect of the test set data. The blue data is the actual data, and the red data is the model prediction data.
Claims
1. A convolutional gated recurrent unit stock prediction method based on extreme gradient boosting feature selection and quadratic decomposition, characterized in that: Step 1) Collect the historical market data of A-share stocks and perform preprocessing. Extract the closing price from it, and serialize it; Perform EMD decomposition on the time series data. Determine how many IMF components to decompose the original time series data according to the sample entropy value of the IMF. Perform secondary decomposition on the high-frequency IMF through VMD; Convert the one-dimensional time series data into a two-dimensional phase space and reconstruct the phase space; Step 2) Build a convolutional gated recurrent unit neural network framework based on XGBoost to predict the next price of the stock; Calculate the feature importance of the data after phase space reconstruction through the extreme gradient boosting method, and select the top three features with the highest feature importance rankings as input features; The convolutional gated recurrent unit neural network framework includes: using a convolutional neural network to learn each deep feature of the IMF and the residual after decomposing the historical stock price data; obtaining the predicted values of the IMF and the residual for the next step through the gated recurrent unit's learning of the deep features; after obtaining the predicted values of all IMFs and residuals, merge all the predicted values to obtain the predicted value of the stock price; Step 3) Use the genetic algorithm to optimize the parameters of the convolutional gated recurrent unit model. Find the optimal parameters of the model through multiple selections, crossovers, and mutations, and import the optimal parameters into the model for prediction; Step 4) Evaluate the prediction effect of the model through the mean square error, mean absolute error, root mean square error, and mean relative error between the predicted value and the true value.
2. The convolutional gated recurrent unit stock prediction method based on extreme gradient boosting feature selection and quadratic decomposition according to claim 1, characterized in that In the said Step 1), the specific method is as follows: 1.1 Collection of historical market data of stocks: Use the AKShare library to obtain the historical market data of specific stocks. 1.2 Processing of stock data: Regarding the historical market data of stocks, extract the closing price from it, serialize it, and perform a stability test on the serialized data; Extract the closing price from the historical stock data, serialize the obtained closing price, and use the Adfuller method, that is, the Augmented Dickey-Fuller test, to perform a stationarity test. The ADF test constructs an autoregressive model to test the stationarity of the sequence. The form of the model is: where X t is the value of the time series at time t, σ, β i are regression coefficients, ∈ t is the error term, i.e., white noise; Through the ADF test, confirm the applicability of the data before modeling, that is, perform the next operation; Perform secondary decomposition on the historical market data that has passed the test: First, obtain all the local extreme values of the sequence s(t), and then connect all the upper and lower extreme points respectively to obtain the envelope line. The average value between the extreme envelopes is m1(t), and h1(t) is obtained by subtracting the average value m1(t) from the sequence s(t); h1(t) = s(t) - m1(t) Perform a screening operation, using h1(t) as the input, and repeat the above steps; h 11 h(t) = h1(t) - m 11 (t) The screening process will be repeated i times until h 1i (t) becomes a true IMF; h 1i h(t) = 1(i-1) h(t) - m 1i (t) The stopping condition is the Cauchy convergence criterion. Specifically, the test requirement is the normalized squared deviation SD of the results of two screening operations i less than; When SD i is less than 0.2, c1(t) = h 1i (t), where c1(t) is the first IMF, and it is difficult for SD i to be less than 0.2 to ensure that the frequency and amplitude of the IMF have sufficient physical significance. c k (t) represents the j-th IMF, and the equation is as follows: c1(t) = h 1i (t) c j (t) = h ji (t) The recurrence formula for the residual is as follows: r1(t) = x(t) - c1(t) r j r(t) = j-1 r(t) - c j r(t) (j ≥ 2) where r j (t) is the residual after the j-th step of decomposition. The sequence s(t) is represented by the IMF C j (t) and the decomposed residual r n (t), and the equation is as follows: where n is the number of IMFs obtained by decomposition; For variational mode decomposition VMD, the original signal s(t) is decomposed into k components, ensuring that the decomposed sequences are modal components with finite bandwidths having central frequencies, and at the same time, the sum of the estimated bandwidths of each mode is minimized. The constraint condition is that the sum of all modes is equal to the original signal. Then the VMD constrained variational model is as follows: Among them, u k ={u1, u2, …, u k} are the modal functions, ω k ={ω1, ω2, …, ω k} are the modal center frequencies, and δ(t) is the Dirac function; To solve the constrained optimization problem, the constrained variational problem is transformed into an unconstrained variational problem. By taking advantage of the quadratic penalty term and the Lagrange multiplier method, the augmented Lagrangian function is introduced, as shown in the formula: Among them, α is the penalty parameter, λ is the Lagrangian multiplier, and the quadratic penalty factor α is used to reduce the interference of Gaussian noise. The modal components and center frequencies are obtained by iterating the alternating direction multiplier method ADMM. According to the ADMM algorithm, the update formulas of u_k, ω_k and λ are as follows: where n is the number of iterations, is the corresponding Fourier transform of f(t), u(t), λ(t), and γ is the update parameter; Calculate the sample entropy of multiple IMFs after decomposition, determine how many IMFs to decompose the original sequence data according to the sample entropy, and use VMD to perform secondary decomposition on the decomposed high-frequency IMFs: The original signal s is decomposed into k components, ensuring that the decomposition sequence is a modal component with a limited bandwidth and a center frequency, and at the same time the sum of the estimated bandwidths of each mode is minimized, with the constraint that the sum of all modes is equal to the original signal; solving the constrained optimization problem, transforming the constrained variational problem into an unconstrained variational problem, and introducing the augmented Lagrangian function by taking advantage of the quadratic penalty term and the Lagrangian multiplier method; Perform phase space reconstruction on the obtained multiple IMFs. Assume that the state space of the system is M with dimension d, and the evolution of the system is described by a smooth mapping Let the state of the system be represented as x(t), and assume that only a single-variable time series {x(t)} can be observed. Construct a new vector through time-delay embedding: y(t)=[x(t),x(t+τ),x(t+2τ),…,x(t+(m-1)τ)] Where τ is the time delay and m is the embedding dimension.
3. The convolutional gated recurrent unit stock prediction method based on extreme gradient boosting feature selection and quadratic decomposition according to claim 1, wherein In the step 2), the specific method is: 2.1 Feature Selection Based on XGBoost The XGBoost algorithm is used to calculate the importance of features in sequence data and select features based on feature importance. The residual of the previous model is used to train the subsequent model, and the subsequent model corrects the errors generated by the previous model. XGboost uses a second-order Taylor expansion to approximate the loss function, and avoids overfitting by adding a regularization term to the objective function to control the complexity of the model. The XGBoost feature importance selection method is one of the embedding methods. During the training of XGBoost, the gain is used to determine the best split node, and the final feature importance score is calculated by the average gain. The average gain is the total gain of all trees divided by the total number of divisions of each feature. The more times a feature is split, the greater the gain of the model. The importance of each feature is output by calculating the gain. The higher the feature importance score, the greater the contribution of the feature in the model construction and training process. Usually, the top-ranked features are obtained by arranging the feature importance scores in descending order. The gain calculation method is: where g i and h i represent the first- and second-order gradients. I L and I R represent the samples of the left and right nodes after segmentation respectively, I = I L ∪I R , and λ, γ are penalty parameters; 2.2 Gated Recurrent Unit The gated recurrent unit GRU has two gates, the update gate and the reset gate. The update gate combines the input gate and the forget gate of the LSTM. The reset gate can directly process the previous hidden state. While ensuring performance, the training speed of GRU is faster than that of LSTM. The calculation equation is as follows: z t = σ(W z * [h t-1 , x t + b z ) r t = σ(W r * [h t-1 , x t + b r ) where h t is the hidden layer vector in the GRU, x t is the output vector in the GRU, z t is the update gate in the GRU, r t is the reset gate in the GRU, W z is the parameter matrix of the update gate in the GRU, W r is the parameter matrix of the reset gate in the GRU, W h is the parameter matrix of the hidden layer in the GRU, b z is the bias vector of the update gate in the GRU, b r is the bias vector of the reset gate in the GRU, b h is the bias vector of the hidden layer in the GRU, σ, tanh are activation functions; 2.3 Convolutional Gated Recurrent Unit The convolution layer automatically extracts features through filters. For the same convolution layer, there is no connection between neurons, and the weights are shared; in the forward propagation stage, each convolution layer uses activation functions and convolution operations on the output of the previous layer. This feature of CNN helps extract hidden information without being affected by uncertainty. The definition is shown in the equation: Among them represents the output after convolution, f represents the activation function, W k represents the weight of the k-th convolutional kernel, x represents the input value, b k represents the bias of the k-th convolution operation, and i and j represent the number of convolution kernel operations required in two directions. The overall framework consists of a convolutional layer, two GRU layers, and a fully connected layer.
4. The convolutional gated recurrent unit stock prediction method based on extreme gradient boosting feature selection and quadratic decomposition according to claim 1, characterized in that, In the said step 3), the specific method is as follows: First, initialize the population and randomly generate the individuals of the first generation, where each individual represents a feasible solution; then, train each individual and calculate its fitness according to the validation set; Then, through the roulette wheel selection method, preferentially select the individuals with higher fitness for genetic operations; finally, generate the individuals of the next generation through chromosome crossover and small-probability gene mutation; after multiple selections, crossovers and mutations, the excellent individuals are retained and the individuals with lower fitness are eliminated, and finally a population with gradually increasing fitness is formed. This process continues until the preset termination condition is met.
5. The convolutional gated recurrent unit stock prediction method based on extreme gradient boosting feature selection and quadratic decomposition according to claim 1, wherein In the said step 4), the specific method is as follows: Mean Squared Error (MSE) calculation: The mean squared error is obtained by taking the average of the squares of the differences between each predicted value and the actual value. The formula is: Among them, y i is the true value, is the predicted value, and n is the number of samples. Mean Absolute Error (MAE) calculation: The mean absolute error is obtained by taking the average of the absolute differences between each predicted value and the actual value. The formula is: Root Mean Squared Error (RMSE) calculation: The root mean squared error is the square root of the mean squared error. The formula is: Mean Average Percentage Error (MAPE) calculation: The mean average percentage error is obtained by calculating the average of the ratios of the absolute errors between each predicted value and the actual value to the actual value. The formula is: