Crude oil price prediction scheme based on deep learning and multi-feature fusion

By combining multi-feature fusion methods with principal component analysis, variational modal decomposition, sample entropy, whale optimization algorithm, Bayesian optimization, LSTM and GRU, the problem that the existing technology is difficult to capture long-term dependency information and process nonlinear and non-stationary characteristics when predicting crude oil prices, significantly improving prediction accuracy and robustness.

CN119991190APending Publication Date: 2025-05-13ZHEJIANG GONGSHANG UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510068686.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

When predicting crude oil prices, the prior art is difficult to effectively capture long-term dependency information and process nonlinear and non-stationary characteristics, resulting in insufficient prediction accuracy.

Method used

A hybrid prediction model based on deep learning and multi-feature fusion is proposed, combining technologies such as principal component analysis (PCA), variational modal decomposition (VMD), sample entropy (SE), whale optimization algorithm (WOA), Bayesian optimization (BO), LSTM and GRU to improve the accuracy of crude oil price prediction through multi-level feature fusion.

Benefits of technology

The accuracy and robustness of crude oil price prediction are significantly improved, and the model's fitting ability to complex nonlinear data is enhanced through multi-feature fusion and hyperparameter optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991190A_ABST
    Figure CN119991190A_ABST
Patent Text Reader

Abstract

Aiming at the complex characteristics of the international crude oil price, the invention provides a hybrid prediction model to improve the prediction precision of the crude oil closing price. Firstly, historical data and influence factors thereof are subjected to dimension reduction processing through principal component analysis (PCA), and the multi-collinearity problem is reduced. Meanwhile, news text sentiment analysis is introduced to capture the influence of market emotion. Thirdly, the number of modes and penalty factors in variational mode decomposition (VMD) are optimized through a whale optimization algorithm (WOA), multiple subsequences are obtained through decomposition, and fluctuation terms and trend terms are reconstructed through sample entropy (SE); and carrying out iterative optimization on key parameters of the LSTM and the GRU by adopting a Bayesian optimization (BO) algorithm. And finally, predicting a fluctuation item by adopting an LSTM model, predicting a trend item by adopting a GRU model, and superposing predicted values of the subsequences to obtain a final result. Experimental results show that RMSE, MAE, MAP and R2 of the model on a test set are 0.792, 0.769, 0.860 and 0.966 respectively and are superior to those of other comparison models, and the effectiveness of the model is verified.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a crude oil price prediction method based on deep learning technology and multi-feature fusion, which belongs to the technical category of time series prediction and energy market analysis. Technical Background

[0002] Crude oil is an important strategic resource of the country and is known as the blood of industry. With the advancement of crude oil financialization, crude oil, as a key derivative product, has played an important role in price discovery and risk aversion, and has therefore attracted widespread attention from investors. However, the crude oil market is affected by a variety of factors, including changes in supply and demand, geopolitical conflicts, public health emergencies, and speculative sentiment, and its price fluctuates violently.

[0003] In the field of financial market analysis and forecasting, crude oil prices are an important indicator of the global energy market, and its forecasting problem has always been a hot topic. Some researchers use traditional time series methods to forecast crude oil prices, such as using the autoregressive integrated moving average (ARIMA) model to analyze price changes. These econometric models are good at capturing linear features in the series, but have certain limitations when dealing with nonlinear and non-stationary characteristics.

[0004] Compared with traditional methods, machine learning models have shown stronger capabilities in processing complex nonlinear and non-stationary data. Some researchers have verified the effectiveness of short-term predictions of energy prices using artificial neural networks (ANNs); others have used extreme gradient boosting trees (XGBoost) to predict natural gas prices, or analyzed electricity market prices using random forests (RFs). Although these models have certain advantages in predicting financial time series, they are still insufficient in capturing information about long-term dependencies. Improved recurrent neural networks (RNNs), such as long short-term memory networks (LSTMs) and gated recurrent units (GRUs), can effectively alleviate the problem of vanishing gradients and are good at capturing long-term dependency characteristics in sequences. Therefore, they have been widely used in time series prediction and have demonstrated excellent performance.

[0005] In addition, since crude oil price series usually exhibit non-stationarity, high noise and long memory characteristics, some researchers have tried to use signal decomposition methods such as empirical mode decomposition (EMD) and ensemble empirical mode decomposition (EEMD) to extract sequence features. However, these decomposition methods have certain shortcomings due to problems such as boundary effects and sensitivity to noise. Variational mode decomposition (VMD) is a non-recursive decomposition technique that decomposes the time series into several subsequences with specific frequency characteristics by optimizing in the frequency domain, thereby better capturing the nonlinear and non-stationary characteristics in the data.

[0006] In order to improve the prediction accuracy, some researchers have tried to combine signal decomposition with prediction models. For example, some researchers have proposed a hybrid method of deep learning and signal decomposition to decompose and predict crude oil price data. The results show that the hybrid model is better than a single model. In addition, some researchers have combined VMD and optimization algorithms to model the decomposed subsequences separately, and obtained the final prediction results by superposition, which significantly improved the accuracy and robustness of the prediction.

[0007] In order to further optimize the effect of crude oil price forecasting, this paper proposes a hybrid forecasting model based on multi-technology fusion, whose main contributions include:

[0008] (1) Hybrid model construction: This paper proposes an innovative hybrid model to predict crude oil prices, which integrates principal component analysis (PCA), variational mode decomposition (VMD), sample entropy (SE), whale optimization algorithm (WOA), Bayesian optimization (BO), LSTM and GRU, and improves the accuracy of crude oil price prediction at multiple levels.

[0009] (2) Dimensionality reduction of principal component analysis: Considering that crude oil prices are affected by multiple factors, some researchers have shown that principal component analysis can effectively reduce redundant information in multidimensional features. This study uses PCA to reduce the dimensionality of input variables, reduce multicollinearity problems and retain key influencing factors.

[0010] (3) Application of variational mode decomposition in time series decomposition: In view of the shortcomings of traditional decomposition methods, this paper uses VMD technology to decompose the original data to extract features of different frequencies, thereby improving the stability of the decomposition and avoiding mode aliasing problems.

[0011] (4) Sample entropy is used for feature division: In order to further improve the prediction accuracy of the model, the present invention introduces sample entropy (SE) to quantify the complexity of time series, and divides the subsequences after VMD decomposition into fluctuation terms and trend terms, thereby achieving targeted modeling.

[0012] (5) Combined application of hyperparameter optimization: To ensure the optimal configuration of model performance, this paper adopts the whale optimization algorithm (WOA) to optimize the decomposition parameters of VMD, and combines iteratively with Bayesian optimization (BO) to optimize the hyperparameters of LSTM and GRU, such as the number of hidden layer nodes and learning rate.

[0013] (6) Innovative application of sentiment analysis: Some researchers have pointed out that sentiment factors play an important role in energy market forecasting. To this end, this paper combines news text sentiment analysis tools to extract daily sentiment scores and introduces them into the crude oil price forecasting model to supplement the insufficiency of price data and capture the market's response to macroeconomic and political events. Summary of the invention

[0014] This paper proposes a crude oil price combination prediction model based on deep learning and multi-feature fusion, aiming to improve the prediction accuracy of crude oil closing prices. The model combines principal component analysis, variational mode decomposition, sample entropy, whale optimization algorithm, Bayesian optimization, long short-term memory network and gated recurrent unit, and improves the prediction accuracy of crude oil prices from multiple angles by multi-feature fusion. Figure 1 The representation is the framework of the proposed hybrid model. The specific scheme includes the following steps:

[0015] Step 1, data collection;

[0016] Step 2, principal component analysis (PCA) dimensionality reduction;

[0017] Step 3: sentiment analysis of news text;

[0018] Step 4, variational mode decomposition (VMD) and parameter optimization;

[0019] Step 5, sample entropy (SE) partition;

[0020] Step 6: LSTM and GRU model prediction and hyperparameter optimization;

[0021] Step 7: Prediction result superposition and evaluation;

[0022] Furthermore, in step 1, the price of WTI crude oil is affected by many factors, including external environmental factors and internal market factors. External environmental factors are affected by the macro-economy, international situation, related product markets, etc., while internal market factors are affected by supply and demand, market size, etc. The closing price, opening price, highest price, lowest price and trading volume of crude oil were obtained from the Investing.com website. The US industrial production index, US unemployment rate, US overall inflation rate, US crude oil inventory, US crude oil production, US crude oil consumption and US liquefied natural gas inventory were obtained from the Tonghuashun database. The Google Index was obtained from Google. News text was obtained from the International Petroleum Network. The trend chart of the WTI crude oil closing price is shown below: Figure 2 shown.

[0023] Furthermore, in step 2, principal component analysis is a method widely used in the field of data statistics. It can extract the most critical information from the original data, simplify the complexity of the data by reducing the data dimension, and maintain the core characteristics of the original data to the greatest extent, so as to facilitate subsequent analysis. Crude oil prices are affected by many factors, and there may be a high correlation between these factors, resulting in multicollinearity problems during model training. PCA is an effective dimensionality reduction method that can reduce the dimension of the data while retaining the most important feature information. The main steps of PCA include:

[0024] (1) Standardize the original data;

[0025] (2) Conduct correlation analysis on the standardized matrix and construct a correlation coefficient matrix;

[0026] (3) According to the correlation coefficient matrix, calculate the eigenvalue and the contribution rate and cumulative contribution rate of each principal component;

[0027] (4) The number of principal components is selected based on the principle that the cumulative contribution rate exceeds 99%. In the present invention, three principal components are finally selected.

[0028] Let F i is the main component; X i is the standardized matrix; u ij The principal component coefficient is the formula for principal component analysis as follows:

[0029]

[0030] In step 3, sentiment analysis is performed on the extracted news text data. SnowNLP tools are used to perform word segmentation and sentiment analysis on the text, and finally generate a daily sentiment score. If the score is greater than 0.5, it means that the news sentiment is positive, otherwise it is negative. The sentiment score of the news text is used as an important feature affecting the fluctuation of crude oil prices and integrated into the input data of the model to further capture the impact of market sentiment on crude oil prices.

[0031] In step 4, the crude oil price series has nonlinear and non-stationary characteristics, and direct prediction model training may ignore these characteristics. VMD is an adaptive signal decomposition method that can decompose complex time series signals into multiple intrinsic mode functions (IMFs) with different center frequencies. By decomposing the crude oil price series into multiple time series, each corresponding to different market dynamics, it provides richer information for subsequent prediction models.

[0032] The decomposition process of VMD can be regarded as the process of finding the optimal solution to the variational problem, which involves the construction and solution of the variational problem. The variational model under constraints can be expressed as:

[0033]

[0034] The constraints are:

[0035]

[0036] Among them, u k =(u k1 ,...,u kn ) are the modal functions, w k =(w k1 ,...,w kn) is the center frequency of each mode. By combining the advantages of the quadratic penalty term and the Lagrangian multiplier method, the augmented Lagrangian function is introduced to transform the above constrained variational problem into an unconstrained variational problem, which is expressed as follows:

[0037]

[0038] Among them, α is a quadratic penalty factor, which is used to reduce the interference of Gaussian noise. The enhanced Lagrangian L is determined in equation (3), and its saddle point can be found by ADMM (alternating direction method of multipliers). According to the ADMM method, u k and w k The VMD analysis process can be implemented in two directions. Iteratively search for the optimal solution u k ,ω k The expressions for and λ are as follows:

[0039]

[0040] In the above formula, n is the number of iterations, and Corresponding to u i The Fourier transform of f(t), f(t) and λ(t).

[0041] Since the selection of relevant parameters in the VMD algorithm has a significant impact on the prediction results, the whale optimization algorithm is used to optimize the VMD parameters to determine the optimal decomposition mode number K and penalty factor α to complete the decomposition of the sequence.

[0042] The Whale Optimization Algorithm (WOA) is inspired by the behavior of whale groups and aims to solve optimization problems, especially in continuous optimization problems. It is an evolutionary algorithm that simulates the social behavior and migration patterns of whale groups and mainly consists of the following three parts.

[0043] (1) Surrounding prey. When whales are surrounding their prey, they will choose the best path for themselves.

[0044] D=|CX * (t)-X(t)|

[0045] X(t+1)=X * (t)-AD

[0046] Where: D is the distance between the prey and the best search agent; X * (t) is the position of the prey; X(t) is the position of the whale; t is the number of iterations of the whale optimization algorithm; C and A represent coefficient vectors.

[0047] (2) Bubble net attack. This strategy simulates two modes of humpback whale hunting behavior, namely, contraction and spiral update. Both modes allow whale individuals to move toward the target prey and decide which mode to use for position update based on randomly generated probabilities.

[0048] X(t+1)=D'e bl cos(2π)+X * (t)

[0049] D'=|X * -X(t)|

[0050] Where b controls the shape of the spiral; l is a random number used to introduce randomness so that individual whales explore the search space in different ways.

[0051] (3) Random search. During the search phase, the whale updates its position through random search, adjusting the whale's displacement according to the position of the nearest search agent.

[0052] D=|CX rand (t)-X(t)|

[0053] X(t+1)=X rand (t)-AD

[0054] Where: X rand (t) represents the random position of prey.

[0055] In the WOA algorithm, a random number q∈[0,1] is also generated to determine the next behavior strategy: if the generated random number q satisfies q>0.5, the algorithm enters the bubble net attack mode, in which individual whales may adopt different behavior strategies to increase the diversity of the search; if q<0.5, the algorithm will judge the coefficient vector A, when |A|≥1, the individual whale chooses to search for prey; when A<1, the individual whale chooses to surround the prey. As the number of iterations of the WOA algorithm increases, the algorithm gradually transitions from the prey search stage to the prey surround stage.

[0056] In step 5, this method evaluates the complexity of each intrinsic mode function (IMF) after variational mode decomposition through sample entropy. The sample entropy value of each IMF is calculated and used as an indicator to measure the complexity of IMF. According to the high and low sample entropy values, IMFs are reclassified into volatility terms and trend terms. This method not only helps to distinguish different components in the signal more accurately, but also has important application value for time series data such as crude oil prices that are affected by multiple factors. The specific calculation method is as follows.

[0057] (1) Construct the time series X into an m-dimensional vector. Where: i = 1, 2, ..., N-m + 1

[0058] X(i)={x(i),x(i+1),…,x(i+m-1)}

[0059] (2) Define the distance between X(i) and X(j) as d[X(i),X(j)] (i and j are not equal), which is the one with the largest difference between the corresponding elements of the two.

[0060]

[0061] (3) Define a threshold r (r>0), count the number of d[X(i),X(j)]<r and calculate its ratio to the total number of vectors.

[0062]

[0063] (4) Average all the results obtained by formula (4).

[0064]

[0065] (5) Repeat the above steps with dimension m+1. Theoretically, the sample entropy of this sequence is:

[0066]

[0067] (6) However, in practice, N cannot be infinite but a finite value. Then the estimated value of sample entropy is:

[0068]

[0069] In step 6, LSTM is used to process the volatility of crude oil prices and capture the dynamic characteristics of price fluctuations. GRU is used to predict the trend of crude oil prices. The Bayesian optimization algorithm is used to optimize the hyperparameters of the LSTM and GRU models, such as the number of hidden layer nodes, learning rate, and maximum number of training times, to improve the prediction performance of the models.

[0070] In step 7, the prediction results of the GRU model for the trend term and the LSTM model for the volatility term are weighted and superimposed to form the final crude oil price forecast value. The root mean square error (RMSE), mean absolute error (MAE), mean absolute percentage error (MAPE) and determination coefficient (R 2 ) and other evaluation indicators are used to comprehensively evaluate the model to verify the effectiveness of the hybrid model for crude oil price prediction proposed in the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] Figure 1 :The process of hybrid model prediction method based on multi-feature fusion

[0072] Figure 2:WTI crude oil closing price time series chart

[0073] Figure 3 :Decomposition results of PCA1

[0074] Figure 4 :Decomposition results of PCA2

[0075] Figure 5 :Decomposition results of PCA3

[0076] Figure 6 :Decomposition results of closing price

[0077] Figure 7 : Decomposition results of sentiment scores

[0078] Figure 8 :Decomposition and reconstruction results of PCA1

[0079] Fig. 9 :Decomposition and reconstruction results of PCA2

[0080] Fig.10 :Decomposition and reconstruction results of PCA3

[0081] Fig.11 : Decomposition and reconstruction results of closing price

[0082] Fig.12 : Decomposition and reconstruction results of sentiment scores

[0083] Fig.13 :WTI crude oil price model fitting curve

[0084] Fig.14 :Comparison results of different evaluation indicators of each model DETAILED DESCRIPTION

[0085] The hybrid model proposed in the present invention is used to predict the closing price of crude oil. First, the original data is standardized to eliminate dimensional differences, and the interpolation method is used to fill missing values ​​to ensure the integrity and continuity of the data. The data content includes crude oil prices, trading volumes, macroeconomic indicators, and sentiment analysis scores of news texts. Then, principal component analysis (PCA) is used to reduce the dimensionality of multidimensional data, extract key factors, reduce redundant features, speed up model training, and improve computational efficiency and prediction accuracy. Then, sentiment analysis is performed on the news text to generate daily sentiment scores, which are integrated into the model input data as important features affecting crude oil price fluctuations to capture the impact of market sentiment on prices. Then, variational mode decomposition (VMD) is used to decompose the time series data into multiple intrinsic mode functions (IMFs) to separate high-frequency terms and trend terms, thereby improving prediction accuracy. The whale optimization algorithm (WOA) is introduced to optimize the parameters of VMD to avoid modal aliasing and noise interference. The sample entropy (SE) of each IMF is calculated, and the IMFs are divided into fluctuation items and trend items according to the sample entropy value. The GRU model and LSTM model are used for prediction respectively, and the model hyperparameters are optimized using Bayesian optimization (BO) to improve the model performance. Finally, the prediction results of the GRU and LSTM models are weighted and superimposed to obtain the final prediction value, and the root mean square error (RMSE), mean absolute error (MAE), mean absolute percentage error (MAPE) and determination coefficient (R 2 ) and other indicators to comprehensively evaluate the effectiveness of the model to ensure its reliability and accuracy in practical applications.

[0086] A. Data Collection and Preprocessing

[0087] Crude oil prices are affected by many factors, including basic trading factors (such as closing price, opening price, highest price, lowest price, trading volume, etc.), macroeconomic factors (such as industrial production index, unemployment rate, inflation rate, etc.), supply and demand (such as crude oil production, inventory, consumption, etc.), alternative energy (such as natural gas prices, etc.), search index (such as Google index), and social events and international situations (such as exchange rate fluctuations, changes in political situation, etc.). In order to eliminate dimensional differences, the original data is standardized and the missing values ​​are filled by interpolation to ensure data integrity and continuity. The 11 influencing factors are reduced in dimension by principal component analysis (PCA), and the KMO test and Bartlett sphericity test are used to confirm that the data is suitable for dimensionality reduction. Finally, 3 principal components are selected, and the cumulative contribution rate reaches 99.9%, which effectively reduces the data dimension, retains key information, and avoids multicollinearity problems.

[0088] At the same time, news texts are collected from the International Petroleum Network and saved in Excel files in chronological order. The news titles are cleaned to remove irrelevant characters and stop words, and the Jieba word segmentation tool is used to segment words to reduce context dependence and word order. The SnowNLP library is used to perform sentiment analysis on the cleaned news titles and calculate the mean of the daily news sentiment score. If the score is greater than 0.5, it means that the news sentiment is positive, and if it is less than 0.5, it is negative. The sentiment score is taken as an important feature affecting the fluctuation of crude oil prices and integrated into the input data of the model to capture the impact of market sentiment on crude oil prices. Combined prediction and result comparison analysis.

[0089] B. Decomposition and reconstruction of crude oil price related data

[0090] In order to improve the performance of variational mode decomposition (VMD), the whale optimization algorithm (WOA) is used to optimize the decomposition mode number K and penalty factor α of VMD. The number of whales in the WOA algorithm is set to 10, the maximum number of iterations is set to 30, the optimized parameters include the decomposition mode number K (ranging from 3 to 10) and the penalty factor α (ranging from 100 to 2500), and the optimization is performed with minimizing the envelope entropy as the objective function. Finally, WOA optimization results in the best decomposition mode number K of 4 and the best penalty factor α of 100. Then, the historical crude oil prices, the first three principal components of related influencing factors, and the sentiment analysis results of news text are used as input for VMD decomposition.

[0091] Figures 3 to 7 The decomposition results are shown, including VMD decomposition diagrams of PCA1, PCA2, PCA3, closing price and sentiment score. Each diagram shows multiple intrinsic mode functions (IMFs), each of which corresponds to components of different frequencies. These IMFs are reclassified by sample entropy. The IMF with the highest sample entropy value is identified as a trend term, and the other IMFs are classified as fluctuation terms. Figures 8 to 12 The decomposition and reconstruction results of these signals are shown, where IMFs are divided into trend terms and fluctuation terms so that different frequency characteristics can be processed separately in the subsequent modeling process. Combined prediction and result comparison analysis.

[0092] C. Combined prediction and comparative analysis of results

[0093] In order to evaluate the performance of the proposed hybrid model, a systematic experimental scheme was designed, and the contribution of each submodule to the final effect of the model was analyzed through ablation experiments. The experiment included a comparison of 9 models, including separate GRU (M1) and LSTM (M2) models, as well as hybrid models (M3-M9) that combined VMD, SE, WOA, BO and other technologies. For models that did not use optimization algorithms, the parameters were determined through pre-experiments and tuning; while for models that used optimization algorithms, the parameters were optimized through an iterative algorithm.

[0094] The performance of the model was evaluated using multiple metrics, including mean absolute error (MAE), root mean square error (RMSE), mean absolute percentage error (MAPE), and coefficient of determination (R 2 ), and a comprehensive evaluation of the prediction results of each model was performed. Fig.13 The fitting curve of the WTI crude oil price forecast results is displayed, and the fit between the forecast results of different models and the actual data is intuitively compared. Fig.14 The MAE, MAPE, RMSE and R 2 Comparison results on indicators.

[0095] The experimental results show that with the superposition of methods and the introduction of optimization algorithms, the prediction error of the model gradually decreases. When the GRU and LSTM models are used alone, the prediction effect is average; but after the introduction of VMD and SE technology, the prediction error of the model is significantly reduced. After further adding WOA and BO optimization, the performance of the model is significantly improved. Finally, the WOA-VMD-SE-BO-GRU-LSTM (M9) model performs well in all evaluation indicators, verifying the effectiveness of the combination of hybrid models and optimization algorithms, and significantly enhancing the model's ability to fit complex nonlinear data.

Claims

1. The present invention proposes a crude oil price combined prediction model based on deep learning and multi-feature fusion, aiming to improve the prediction accuracy of crude oil closing prices. The model combines principal component analysis, variational mode decomposition, sample entropy, whale optimization algorithm, Bayesian optimization, long short-term memory network and gated recurrent unit and other technologies, and improves the prediction accuracy of crude oil prices from multiple angles by means of multi-feature fusion. Figure 1 shows the framework of the proposed hybrid model. The specific scheme includes the following steps: Step 1, data collection; Step 2, principal component analysis (PCA) dimensionality reduction; Step 3: sentiment analysis of news text; Step 4, variational mode decomposition (VMD) and parameter optimization; Step 5, sample entropy (SE) partition; Step 6: LSTM and GRU model prediction and hyperparameter optimization; Step 7: Prediction result superposition and evaluation; Furthermore, in step 1, the price of WTI crude oil is affected by many factors, including external environmental factors and internal market factors. External environmental factors are affected by macroeconomics, international situation, related product markets, etc., while internal market factors are affected by supply and demand, market size, etc. The closing price, opening price, highest price, lowest price and trading volume of crude oil are obtained from the Yingwei Finance website. The US industrial production index, US unemployment rate, US overall inflation rate, US crude oil inventory, US crude oil production, US crude oil consumption and US liquefied natural gas inventory are obtained from the Tonghuashun database. The Google Index is obtained from Google. The news text is obtained from the International Petroleum Network. The trend chart of the closing price of WTI crude oil is shown in Figure 2. Furthermore, in step 2, principal component analysis is a method widely used in the field of data statistics. It can extract the most critical information from the original data, simplify the complexity of the data by reducing the data dimension, and maintain the core characteristics of the original data to the greatest extent, so as to facilitate subsequent analysis. Crude oil prices are affected by many factors, and these factors may be highly correlated, resulting in multicollinearity problems during model training. PCA is an effective dimensionality reduction method that can reduce the dimension of data while retaining the most important feature information. The main steps of PCA include: (1) Standardize the original data; (2) Conduct correlation analysis on the standardized matrix and construct a correlation coefficient matrix; (3) According to the correlation coefficient matrix, calculate the eigenvalue and the contribution rate and cumulative contribution rate of each principal component; (4) The number of principal components is selected based on the principle that the cumulative contribution rate exceeds 99%. In the present invention, three principal components are finally selected. Let F i is the main component; X i is the standardized matrix; u ij The principal component coefficient is the formula for principal component analysis as follows: In step 3, sentiment analysis is performed on the extracted news text data. SnowNLP tools are used to perform word segmentation and sentiment analysis on the text, and finally generate a daily sentiment score. If the score is greater than 0.5, it means that the news sentiment is positive, otherwise it is negative. The sentiment score of the news text is used as an important feature affecting the fluctuation of crude oil prices and integrated into the input data of the model to further capture the impact of market sentiment on crude oil prices. In step 4, the crude oil price series has nonlinear and non-stationary characteristics, and direct prediction model training may ignore these characteristics. VMD is an adaptive signal decomposition method that can decompose complex time series signals into multiple intrinsic mode functions (IMFs) with different center frequencies. By decomposing the crude oil price series into multiple time series, each corresponding to different market dynamics, it provides richer information for subsequent prediction models. The decomposition process of VMD can be regarded as the process of finding the optimal solution to the variational problem, which involves the construction and solution of the variational problem. The variational model under constraints can be expressed as: The constraints are: Among them, u k =(u k1 ,...,u kn ) are the modal functions, w k =(w k1 ,...,w kn ) is the center frequency of each mode. By combining the advantages of the quadratic penalty term and the Lagrangian multiplier method, the augmented Lagrangian function is introduced to transform the above constrained variational problem into an unconstrained variational problem, which is expressed as follows: Among them, α is a quadratic penalty factor, which is used to reduce the interference of Gaussian noise. The enhanced Lagrangian L is determined in equation (3), and its saddle point can be found by ADMM (alternating direction method of multipliers). According to the ADMM method, u k and w k The VMD analysis process can be implemented in two directions. Iteratively search for the optimal solution u k ,ω k The expressions for and λ are as follows: In the above formula, n is the number of iterations, and Corresponding to u i The Fourier transform of f(t), f(t) and λ(t). Since the selection of relevant parameters in the VMD algorithm has a significant impact on the prediction results, the whale optimization algorithm is used to optimize the VMD parameters to determine the optimal decomposition mode number K and penalty factor α to complete the decomposition of the sequence. The Whale Optimization Algorithm (WOA) is inspired by the behavior of whale groups and aims to solve optimization problems, especially in continuous optimization problems. It is an evolutionary algorithm that simulates the social behavior and migration patterns of whale groups and mainly consists of the following three parts. (1) Surrounding prey. When whales are surrounding their prey, they will choose the best path for themselves. D=|CX * (t)-X(t)| X(t+1)=X * (t)-AD Where: D is the distance between the prey and the best search agent; X * (t) is the position of the prey; X(t) is the position of the whale; t is the number of iterations of the whale optimization algorithm; C and A represent coefficient vectors. (2) Bubble net attack. This strategy simulates two modes of humpback whale hunting behavior, namely, contraction and spiral update. Both modes allow whale individuals to move toward the target prey and decide which mode to use for position update based on randomly generated probabilities. X(t+1)=D'e bl cos(2π)+X * (t) D'=|X * -X(t)| Where b controls the shape of the spiral; l is a random number used to introduce randomness so that individual whales explore the search space in different ways. (3) Random search. During the search phase, the whale updates its position through random search, adjusting the whale's displacement according to the position of the nearest search agent. D=|CX rand (t)-X(t)| X(t+1)=X rand (t)-AD Where: X rand (t) represents the random position of prey. In the WOA algorithm, a random number q∈[0,1] is also generated to determine the next behavior strategy: if the generated random number q satisfies q>0.5, the algorithm enters the bubble net attack mode, in which individual whales may adopt different behavior strategies to increase the diversity of the search; if q<0.5, the algorithm will judge the coefficient vector A, when |A|≥1, the individual whale chooses to search for prey; when A<1, the individual whale chooses to surround the prey. As the number of iterations of the WOA algorithm increases, the algorithm gradually transitions from the prey search stage to the prey surround stage. In step 5, the method evaluates the complexity of each intrinsic mode function (IMF) after variational mode decomposition through sample entropy. The sample entropy value of each IMF is calculated and used as an indicator to measure the complexity of IMF. According to the sample entropy value, IMFs are reclassified into volatility items and trend items. This method not only helps to distinguish the different components of the signal more accurately, but also has important application value, especially for time series data such as crude oil prices that are affected by multiple factors. The specific calculation method is as follows. (1) Construct the time series X into an m-dimensional vector. Where: i = 1, 2, ..., N-m + 1 X(i)={x(i),x(i+1),…,x(i+m-1)} (2) Define the distance between X(i) and X(j) as d[X(i),X(j)] (i and j are not equal), which is the one with the largest difference between the corresponding elements of the two. (3) Define a threshold r (r>0), count the number of d[X(i),X(j)]<r and calculate its ratio to the total number of vectors. (4) Average all the results obtained by formula (4). (5) Repeat the above steps with dimension m+1. Theoretically, the sample entropy of this sequence is: (6) However, in reality, N cannot be infinite, but a finite value. Then the estimated value of sample entropy is: In step 6, LSTM is used to process the volatility of crude oil prices and capture the dynamic characteristics of price fluctuations. GRU is used to predict the trend of crude oil prices. The Bayesian optimization algorithm is used to optimize the hyperparameters of the LSTM and GRU models, such as the number of hidden layer nodes, learning rate, and maximum number of training times, to improve the prediction performance of the models. In step 7, the prediction results of the GRU model for the trend term and the LSTM model for the volatility term are weighted and superimposed to form the final crude oil price forecast value. The root mean square error (RMSE), mean absolute error (MAE), mean absolute percentage error (MAPE) and determination coefficient (R 2 ) and other evaluation indicators are used to comprehensively evaluate the model to verify the effectiveness of the hybrid model for crude oil price prediction proposed in the present invention.