Iron ore price machine learning prediction method and system based on multi-scale decomposition
Through multi-scale decomposition and Bi-LSTM model combined with knowledge distillation technology, the problem of insufficient synergies of accuracy and model scale in iron ore price prediction is solved, and high-precision prediction of lightweight models is achieved.
Patent Information
- Application Number
- CN202510824120.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-07-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Conventional machine learning methods in the prior art are difficult to achieve high-performance synergy between iron ore price prediction accuracy and model structure scale, and traditional methods have shortcomings in dealing with complex market dynamics and nonlinear relationships.
The multi-scale decomposition method is used to screen feature subsequences through discrete wavelet transformation and correlation analysis, and combine Bi-LSTM model and knowledge distillation technology to build a lightweight integrated model for iron ore price prediction.
It realizes high-performance coordination between lightweight model structure and prediction accuracy, improves the accuracy and efficiency of iron ore price prediction, and is suitable for multi-factor data processing.
Smart Images

Figure CN120338855A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of mining data analysis, and particularly to a machine learning prediction method and system for iron ore prices based on multi-scale decomposition. Background Art
[0002] Due to numerous factors affecting iron ore prices and obvious price fluctuations in recent years, it is difficult to accurately predict them. Traditional methods mainly include economic methods, statistical methods, time series analysis methods, econometrics, and financial engineering methods. Economic methods focus on studying macroeconomic factors and ignore complex market dynamics and non-linear relationships. Statistical methods can capture the trends of time series, but have poor fitting ability for non-linear relationships. Financial engineering methods introduce complex economic models and can better explain market dynamics, but have limited large-scale data processing capabilities. Therefore, with the development of machine learning methods, iron ore price prediction based on machine learning methods has become a hot topic of concern for many scholars. Since there are many factors affecting iron ore prices and there are certain non-linear relationships, it is difficult for conventional machine learning methods to achieve high-performance coordination between model prediction accuracy and model structure scale. For example, high prediction accuracy requires a large model structure scale, and a complex structure will affect large-scale data processing capabilities. Correspondingly, a small model structure scale will reduce prediction accuracy. Summary of the Invention
[0003] The purpose of the present invention is to provide a machine learning prediction method and system for iron ore prices based on multi-scale decomposition to solve the technical problem that it is difficult for conventional machine learning methods to achieve high-performance coordination between model prediction accuracy and model structure scale in the prior art.
[0004] To solve the above technical problem, the present invention specifically provides the following technical solutions: A machine learning prediction method for iron ore prices based on multi-scale decomposition, comprising the following steps: Determine the first characteristic parameters for iron ore price prediction according to the factors affecting iron ore prices, obtain the historical data of the first characteristic parameters and iron ore prices, and perform data preprocessing on the historical data; Perform multi-scale sequence decomposition on the historical data of the first characteristic parameters through discrete wavelet transform to obtain multiple characteristic subsequences with different frequencies, determine the correlation between each characteristic subsequence and iron ore prices through the correlation analysis method, and screen out the second characteristic parameters for iron ore price prediction from each characteristic subsequence based on the correlation; Using the Bi-LSTM model structure, multiple deep learning models for predicting iron ore prices are established respectively with each feature subsequence in the second feature parameter and the historical data of iron ore prices. Through knowledge distillation technology, the prediction performance of the deep learning model corresponding to the feature subsequence with the highest correlation with iron ore prices is integrated and optimized among multiple deep learning models, and a lightweight integrated model for predicting iron ore prices is obtained; Using the lightweight integrated model to predict the iron ore price within the period to be predicted, and obtaining the iron ore price of the period to be predicted.
[0005] As a preferred embodiment of the present invention, the influencing factors of iron ore prices include the US dollar index, BDI shipping price, steel price index, iron ore inventory, exchange rate, interest rate, and news sentiment analysis; The method for determining the first feature parameter includes: Extracting five factors, namely the US dollar index, BDI shipping price, steel price index, iron ore inventory, and news sentiment analysis, from the influencing factors of iron ore prices to form the first feature parameter; Converting the news sentiment analysis in the first feature parameter from text data into time series data composed of three major categories of data: positive, negative, and neutral, through the Naive Bayes method.
[0006] As a preferred embodiment of the present invention, the sequence decomposition method of discrete wavelet transform includes: Setting the wavelet function in the discrete wavelet transform to a 3-layer Daubechies function; Performing multi-scale sequence decomposition on the historical data of the first feature parameter through discrete wavelet transform to obtain multiple feature subsequences {A1, A2,..., AN} with different frequencies. In the formula, A1, A2, AN are respectively the identifiers of the feature subsequences, and N is the total number of feature subsequences obtained by discrete wavelet transform; The expression of the discrete wavelet transform is: ; In the formula, is the wavelet coefficient of the discrete wavelet transform at the parameter , is the scale parameter, are respectively the scaling and translation parameters, is the complex conjugate function, is the time series composed of the historical data of the first feature parameter.
[0007] As a preferred embodiment of the present invention, the screening method of the second feature parameter includes: Calculate the Pearson correlation coefficient between each eigen-subsequence obtained by discrete wavelet transform and the time series composed of historical data of iron ore prices respectively; Select the eigen-subsequences with non-zero Pearson correlation coefficients in each eigen-subsequence obtained by discrete wavelet transform as the second characteristic parameters. Each eigen-subsequence in the second characteristic parameters is: {B1, B2, …, BM}, where B1, B2, BM are the identifiers of the eigen-subsequences respectively, and M is the total number of eigen-subsequences in the second characteristic parameters; Among them, in the second characteristic parameters, the eigen-subsequence with the largest absolute value of the Pearson correlation coefficient corresponds to the eigen-subsequence with the highest correlation with the iron ore price, and the eigen-subsequence with the largest absolute value of the Pearson correlation coefficient is marked as B1.
[0008] As a preferred solution of the present invention, the construction method of the deep learning model includes: Establish M deep learning models {Bi-LSTM-B1, Bi-LSTM-B2, …, Bi-LSTM-BM} that output iron ore prediction results according to the eigen-subsequences with the eigen-subsequences {B1, B2, …, BM} and the historical data of iron ore prices in the Bi-LSTM model structure. Among them, Bi-LSTM-B1, Bi-LSTM-B2, Bi-LSTM-BM are the identifiers of the deep learning models respectively, M is the total number of eigen-subsequences in the second characteristic parameters, and the deep learning model corresponding to the eigen-subsequence with the highest correlation with the iron ore price is Bi-LSTM-B1; Among them, the Bi-LSTM model structure is composed of a forward LSTM model and a backward LSTM model. The forward LSTM model is used to obtain the output of the iron ore prediction results in the forward order at each moment of the eigen-subsequence, and the backward LSTM is used to obtain the output of the iron ore prediction results in the reverse order at each moment of the eigen-subsequence. The output results of the iron ore prediction results in the two directions are spliced to form the output result as the final iron ore prediction result output at each moment; The expression of the Bi-LSTM model structure is: ; In the formula, , , , are the forward propagation hidden layer state, backward propagation hidden layer state, input eigen-subsequence, and iron ore prediction result output by the hidden layer at time t respectively, is the activation function of the hidden layer, , , , are respectively In 、 In 、 、 weight matrices, 、 are the forward propagation hidden layer state and the backward propagation hidden layer state at time t-1 respectively, 、 are the biases of the forward propagation hidden layer and the backward propagation hidden layer respectively, is the vector concatenation symbol; The parameters of the Bi-LSTM model structure include the number of hidden layer neurons, the weight matrices of different states, the learning rate, and the biases of different hidden layers; The parameters of the Bi-LSTM model structure are optimized by the particle swarm optimization algorithm.
[0009] As a preferred embodiment of the present invention, the construction method of the lightweight integrated model includes: Among the M deep learning models {Bi-LSTM-B1, Bi-LSTM-B2,…, Bi-LSTM-BM}, the deep learning model Bi-LSTM-B1 corresponding to the feature subsequence with the highest correlation with the iron ore price is marked as the lightweight model; Train the remaining M-1 deep learning models {Bi-LSTM-B2,…, Bi-LSTM-BM}, and use each trained deep learning model in the remaining M-1 deep learning models {Bi-LSTM-B2,…, Bi-LSTM-BM} as the teacher model one by one to perform knowledge distillation training on the lightweight model as the student model, and use the trained lightweight model as the lightweight integrated model; The training loss function of the lightweight model is: ; In the formula, is the training loss of the lightweight model, is the iron ore price prediction result output by the lightweight model, is the iron ore price prediction result output by the i-th deep learning model in {Bi-LSTM-B2,…, Bi-LSTM-BM}, is the absolute value of the Pearson correlation coefficient corresponding to the i-th deep learning model in {Bi-LSTM-B2,…, Bi-LSTM-BM}, is the iron ore price prediction result obtained by the M deep learning models {Bi-LSTM-B1, Bi-LSTM-B2,…, Bi-LSTM-BM} through the Stacking integration algorithm, is the KL divergence calculation formula; The training losses of each deep learning model in {Bi-LSTM-B2, …, Bi-LSTM-BM} are: ; In the formula, is the training loss of each deep learning model in {Bi-LSTM-B2, …, Bi-LSTM-BM}, is the true value of the iron ore price.
[0010] As a preferred solution of the present invention, the method for determining by the Stacking integration algorithm includes: Among the M deep learning models {Bi-LSTM-B1, Bi-LSTM-B2, …, Bi-LSTM-BM}, each deep learning model is used as the base model in the Stacking integration algorithm, and the XGBoost model is used as the meta-model in the Stacking integration algorithm; Taking the iron ore price prediction result output by the base model as the input of the meta-model, and taking as the output of the meta-model.
[0011] As a preferred solution of the present invention, the prediction method for the iron ore price in the to-be-predicted time period includes: Obtaining the data sequence of the first feature parameter in the to-be-predicted time period and performing data preprocessing; Performing multi-scale sequence decomposition on the data sequence of the first feature parameter through discrete wavelet transform, and screening out the second feature parameter in the to-be-predicted time period by using correlation analysis; Inputting the second feature parameter in the to-be-predicted time period into the lightweight integration model to obtain the iron ore price prediction result in the to-be-predicted time period.
[0012] As a preferred solution of the present invention, the data preprocessing method includes missing value processing and outlier cleaning. The missing values are processed by the interpolation method of summing and averaging the data of the two moments before and the two moments after the current missing value. The expression of the interpolation method is: ; In the formula, , , , , are the data values at time t, time t-2, time t-1, time t+1, and time t+2 respectively; The outlier is determined based on the difference between the data value at time t and the data values at adjacent times being higher than a preset threshold, and the outlier is assigned the average value of the historical data.
[0013] As a preferred embodiment of the present invention, the present invention provides an iron ore price prediction system based on multi-scale decomposition and machine learning algorithms, which is applied to an iron ore price machine learning prediction method based on multi-scale decomposition. The system includes: A data acquisition unit, configured to determine first feature parameters for iron ore price prediction according to iron ore price influencing factors, and obtain historical data of the first feature parameters and iron ore prices; A data preprocessing unit, configured to perform data preprocessing on the historical data; A feature determination unit, configured to perform multi-scale sequence decomposition on the historical data of the first feature parameters through discrete wavelet transform to obtain multiple feature subsequences with different frequencies, determine the correlation between each feature subsequence and the iron ore price through a correlation analysis method, and screen out second feature parameters for iron ore price prediction from each feature subsequence based on the correlation; A model construction unit, configured to establish multiple deep learning models for iron ore price prediction respectively with each feature subsequence in the second feature parameters and the historical data of the iron ore price through a Bi-LSTM model structure, and perform integrated optimization of the prediction performance of the deep learning model corresponding to the feature subsequence with the highest correlation with the iron ore price among the multiple deep learning models through knowledge distillation technology to obtain a lightweight integrated model for iron ore price prediction; A price prediction unit, configured to predict the iron ore price within a to-be-predicted time period through the lightweight integrated model to obtain the iron ore price of the to-be-predicted time period.
[0014] The present invention has the following beneficial effects compared with the prior art: The present invention combines discrete wavelet transform, a bidirectional long short-term memory artificial neural network Bi-LSTM, and knowledge distillation technology to construct a lightweight model for predicting iron ore prices, and integrates the prediction results of multiple complex Bi-LSTM model structures to obtain high-precision prediction performance, realizing the coordination of lightweight model structure and high-performance prediction accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only exemplary, and for those of ordinary skill in the art, without creative efforts, other implementation drawings can be obtained according to the provided drawings.
[0016] Figure 1 Flow chart of the machine learning prediction method for iron ore price based on multi-scale decomposition provided by the embodiment of the present invention; Figure 2 Block diagram of the iron ore price prediction system based on multi-scale decomposition and machine learning algorithm provided by the embodiment of the present invention; Figure 3 Structural diagram of the Bi-LSTM model provided by the embodiment of the present invention; Figure 4 Schematic diagram of the construction process of the Bi-LSTM deep learning model based on the discrete wavelet decomposition method provided by the embodiment of the present invention. Detailed implementation manners
[0017] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0018] As Figures 1 to 4 shown, the present invention provides a machine learning prediction method for iron ore price based on multi-scale decomposition, including the following steps: Determine the first characteristic parameters for iron ore price prediction according to the influencing factors of iron ore price, obtain the historical data of the first characteristic parameters and iron ore price, and perform data preprocessing on the historical data; Perform multi-scale sequence decomposition on the historical data of the first characteristic parameters through discrete wavelet transform to obtain multiple characteristic subsequences with different frequencies, determine the correlation between each characteristic subsequence and the iron ore price through the correlation analysis method, and screen out the second characteristic parameters for iron ore price prediction from each characteristic subsequence based on the correlation; Establish multiple deep learning models for iron ore price prediction through the Bi-LSTM model structure respectively with each characteristic subsequence in the second characteristic parameters and the historical data of the iron ore price, and perform integrated optimization of the prediction performance of the deep learning model corresponding to the characteristic subsequence with the highest correlation with the iron ore price among the multiple deep learning models through the knowledge distillation technology to obtain a lightweight integrated model for iron ore price prediction; The iron ore price in the forecast time period is predicted by a lightweight integrated model to obtain the iron ore price in the forecast time period. The present invention not only considers the time series information of the iron ore price itself, but also considers that the iron ore price is affected by many factors, integrates relevant information such as market supply and demand, macroeconomics (exchange rate, interest rate, etc.) and industry policy (news sentiment analysis data), establishes a machine learning model for predicting the iron ore price, predicts the iron ore price from a more comprehensive perspective, and provides a richer basis for decision-making.
[0019] The present invention takes into account the existence of text data (i.e., news sentiment analysis data) in the factors affecting iron ore prices, and uses the naive Bayes method to convert the text data, thereby converting the text data into time series data alone.
[0020] The present invention applies discrete wavelet transform to time series data, decomposing signal components of different frequencies in the time series data, such as the approximate part (low-frequency part) and the detail part (high-frequency part), so as to deeply mine the characteristic parameters in the time series. The mined characteristic parameters can be better used for iron ore price parameter prediction and improve prediction accuracy.
[0021] The present invention screens characteristic parameters with different frequencies (i.e. {A1, A2, ..., AN}) mined out by discrete wavelet transform through correlation analysis, selects characteristic parameters related to iron ore price (i.e. {B1, B2, ..., BM}), eliminates redundant characteristic parameters, and adopts relevant characteristic parameters to better predict iron ore price parameters, thereby improving prediction accuracy.
[0022] The present invention adopts a bidirectional long short-term memory artificial neural network method, namely Bi-LSTM (Bidirectional Long Short-Term Memory), which solves the limitations of traditional methods in processing nonlinear and multi-factor data. The method can not only consider the data before the current moment, but also fully consider the data after the current moment, effectively capture the context information, and is particularly suitable for the type of feature parameters with text data.
[0023] The present invention combines discrete wavelet transform and Bi-LSTM model structure, and uses the Bi-LSTM model structure to perform deep learning modeling on all feature subsequences {B1, B2, ..., BM} obtained by correlation analysis after discrete wavelet transform, so as to obtain M deep learning models {Bi-LSTM-B1, Bi-LSTM-B2, ..., Bi-LSTM-BM} for predicting iron ore prices according to B1, B2, ..., BM inputs, wherein the deep learning model constructed with the feature subsequence B1 with the highest correlation is Bi-LSTM-B1.
[0024] By using the Stacking integration method for these M deep learning models, multi-model integration can be achieved, and thus the integrated result of the iron ore prediction results of these M deep learning models can be obtained, resulting in a final high-precision price prediction result. However, the iron ore prediction model obtained by integrating M deep learning models using the Stacking integration method includes the M deep learning models as base models and also a meta-model for integrating the M prediction results. The model structure is complex, the data processing volume is large, the big data processing efficiency is limited, and the requirements for the model deployment environment are high.
[0025] In order to lightweight the model structure, the present invention adopts the knowledge distillation technology (a model compression technology that transfers knowledge through a teacher-student model). The deep learning model constructed with the most relevant feature subsequence B1, namely Bi-LSTM-B1, is used as the base model to learn the iron ore price prediction performance of the remaining Bi-LSTM-B2, …, Bi-LSTM-BM models. Since Bi-LSTM-B1 uses the most relevant feature subsequence B1 to predict the iron ore price, it can obtain a relatively high-confidence iron ore price prediction result. However, due to only analyzing B1, the single feature leads to the confidence not being optimized. Therefore, the prediction capabilities of Bi-LSTM-B2, …, Bi-LSTM-BM are transferred to Bi-LSTM-B1 through the knowledge distillation method of the teacher-student model structure, enabling Bi-LSTM-B1 to have the prediction performance of predicting the iron ore price based on B2, …, BM. Coupled with the original prediction performance of predicting the iron ore price based on B1, Bi-LSTM-B1 obtains the prediction performance equivalent to that of an integrated model composed of M base models and 1 meta-model. That is, Bi-LSTM-B1 can achieve the integrated result of the iron ore prediction results of M deep learning models by only inputting B2. The Bi-LSTM-B1 trained by the knowledge distillation method (i.e., the lightweight integrated model) is used as the final iron ore price prediction model, which can achieve high-precision performance while lightweighting the model, obtain faster training and inference speeds, be able to be deployed and run on smaller devices, and save system resources more.
[0026] The influencing factors of iron ore price include the US dollar index, BDI shipping price, steel price index, iron ore inventory, exchange rate, interest rate, and news sentiment analysis; The method for determining the first feature parameter includes: Extract the five factors of the US dollar index, BDI shipping price, steel price index, iron ore inventory, and news sentiment analysis from the influencing factors of iron ore price to form the first feature parameter; Convert the news sentiment analysis in the first feature parameter from text data into time series data consisting of three major categories of data: positive, negative, and neutral, through the Naive Bayes method.
[0027] The sequence decomposition method of discrete wavelet transform includes: Set the wavelet function in the discrete wavelet transform to the 3-layer Daubechies function; Perform multi-scale sequence decomposition on the historical data of the first feature parameter through discrete wavelet transform to obtain multiple feature subsequences {A1, A2, …, AN} with different frequencies. In the formula, A1, A2, AN are the identifiers of the feature subsequences respectively, and N is the total number of feature subsequences obtained by discrete wavelet transform; The expression of discrete wavelet transform is: ; In the formula, is the wavelet coefficient of the discrete wavelet transform at parameter ; is the scale parameter; are the scaling and translation parameters respectively; is the complex conjugate function; is the time series composed of the historical data of the first feature parameter.
[0028] The screening method of the second feature parameter includes: Calculate the Pearson correlation coefficient between each feature subsequence obtained by discrete wavelet transform and the time series composed of the historical data of iron ore prices respectively; Screen the feature subsequences with non-zero Pearson correlation coefficients in each feature subsequence obtained by discrete wavelet transform as the second feature parameter. Each feature subsequence in the second feature parameter is: {B1, B2, …, BM}. In the formula, B1, B2, BM are the identifiers of the feature subsequences respectively, and M is the total number of feature subsequences in the second feature parameter; Among them, in the second feature parameter, the feature subsequence with the largest absolute value of the Pearson correlation coefficient corresponds to the feature subsequence with the highest correlation with iron ore prices, and the feature subsequence with the largest absolute value of the Pearson correlation coefficient is marked as B1.
[0029] Scale the second feature parameter to the interval [0, 1] using the linear normalization method.
[0030] The construction method of the deep learning model includes: The Bi-LSTM model structure is used to establish M deep learning models {Bi-LSTM-B1, Bi-LSTM-B2, …, Bi-LSTM-BM} that output iron ore prediction results based on each feature subsequence in {B1, B2, …, BM} and the historical data of iron ore prices. In the formula, Bi-LSTM-B1, Bi-LSTM-B2, and Bi-LSTM-BM are the identifiers of the deep learning models respectively, M is the total number of feature subsequences in the second feature parameter, and the deep learning model corresponding to the feature subsequence with the highest correlation with the iron ore price is Bi-LSTM-B1; Among them, the Bi-LSTM model structure is composed of a forward LSTM model and a backward LSTM model. The forward LSTM model is used to obtain the output of the iron ore prediction results in the forward order at each moment of the feature subsequence, and the backward LSTM is used to obtain the output of the iron ore prediction results in the reverse order at each moment of the feature subsequence. The output results of the iron ore prediction results in the two directions are concatenated to form the final iron ore prediction result output at each moment; As Figure 3 shown, the expression of the Bi-LSTM model structure is: ; In the formula, , , , are the forward propagation hidden layer state, backward propagation hidden layer state, input feature subsequence, and iron ore prediction result output by the hidden layer at time t respectively, is the activation function of the hidden layer, , , , are respectively in , in , , weight matrices, , are the forward propagation hidden layer state and backward propagation hidden layer state at time t-1 respectively, , are the biases of the forward propagation hidden layer and backward propagation hidden layer respectively, is the vector concatenation symbol; The parameters of the Bi-LSTM model structure include the number of neurons in the hidden layer, weight matrices of different states, learning rate, and biases of different hidden layers.
[0031] The construction method of the lightweight integrated model includes: Among the M deep learning models {Bi-LSTM-B1, Bi-LSTM-B2, …, Bi-LSTM-BM}, the deep learning model Bi-LSTM-B1 corresponding to the feature subsequence with the highest correlation with the iron ore price is marked as the lightweight model; Train the remaining M - 1 deep learning models {Bi-LSTM-B2, …, Bi-LSTM-BM}, and use each trained deep learning model in the remaining M - 1 deep learning models {Bi-LSTM-B2, …, Bi-LSTM-BM} as the teacher model one by one to perform knowledge distillation training on the lightweight model as the student model, and use the trained lightweight model as the lightweight integrated model; The training loss function of the lightweight model is: ; In the formula, is the training loss of the lightweight model, is the predicted iron ore price result output by the lightweight model, is the predicted iron ore price result output by the i-th deep learning model in {Bi-LSTM-B2, …, Bi-LSTM-BM}, is the absolute value of the Pearson correlation coefficient corresponding to the i-th deep learning model in {Bi-LSTM-B2, …, Bi-LSTM-BM}, is the predicted iron ore price result obtained by the M deep learning models {Bi-LSTM-B1, Bi-LSTM-B2, …, Bi-LSTM-BM} through the Stacking integration algorithm, is the KL divergence operation formula, which can also be replaced by the mean square error; The training losses of the deep learning models in {Bi-LSTM-B2, …, Bi-LSTM-BM} are: ; In the formula, is the training loss of each deep learning model in {Bi-LSTM-B2, …, Bi-LSTM-BM}, is the true value of the iron ore price.
[0032] In the method of training Bi-LSTM-B1 by knowledge distillation in the present invention, the training loss of Bi-LSTM-B1 is set as Ls, which includes two aspects. One is , making the prediction results of Bi-LSTM-B1 approach those of Bi-LSTM-B2, …, Bi-LSTM-BM respectively. That is, the prediction ability of Bi-LSTM-B1 learns from the prediction abilities of Bi-LSTM-B2, …, Bi-LSTM-BM respectively. At the same time, the degree to which the prediction ability of Bi-LSTM-B1 learns from the prediction abilities of Bi-LSTM-B2, …, Bi-LSTM-BM is related to the correlation between B2, …, BM and the iron ore price prediction. Among them, the stronger the correlation between B2, …, BM and the iron ore price prediction, the higher the confidence level of the iron ore price prediction of the corresponding deep learning model of the feature subsequence, or in other words, the stronger the prediction ability for the iron ore price, and the higher the benefit for Bi-LSTM-B1 to learn from it. Therefore, a higher weight is set for it. , making the learning priority of this deep learning model higher, so as to achieve a deeper learning degree of the prediction ability of Bi-LSTM-B1 from it and obtain a higher performance improvement. This part enables Bi-LSTM-B1 to independently learn from the prediction abilities of Bi-LSTM-B2, …, Bi-LSTM-BM one by one, which is equivalent to independently developing the local prediction performances that Bi-LSTM-B1 has on the remaining deep learning models respectively, that is, equivalent to achieving the improvement of each local performance of Bi-LSTM-B1.
[0033] Another is , making the prediction result of Bi-LSTM-B1 approach the integrated result of the iron ore prediction results of the M deep learning models. The prediction ability of Bi-LSTM-B1 learns from the prediction ability of the iron ore prediction model obtained by integrating the M deep learning models with the Stacking integration method, which consists of M deep learning models and a meta-model for integrating the M prediction results. This part enables Bi-LSTM-B1 to learn from the prediction abilities of Bi-LSTM-B1, Bi-LSTM-B2, …, Bi-LSTM-BM as a whole, which is equivalent to developing the global prediction performance that Bi-LSTM-B1 has after integrating the M deep learning models as a whole, that is, equivalent to achieving the improvement of the global performance of Bi-LSTM-B1.
[0034] In summary, the dual improvement of local performance and global performance optimizes the prediction ability of Bi-LSTM-B1, and the most accurate iron ore prediction result can be achieved by only inputting B2.
[0035] In the training of each deep learning model in {Bi-LSTM-B2, …, Bi-LSTM-BM}, the common loss function is used for construction. , that is, the difference between the outputs of Bi-LSTM-B2, …, Bi-LSTM-BM and the true values, ensuring the lowest axis difference after training and enabling the optimal prediction performance of each of Bi-LSTM-B2, …, Bi-LSTM-BM, can be used for subsequent knowledge distillation.
[0036] Determined by the Stacking ensemble algorithm The method includes: Among the M deep learning models {Bi-LSTM-B1, Bi-LSTM-B2, …, Bi-LSTM-BM}, each deep learning model is used as the base model in the Stacking ensemble algorithm, and the XGBoost model is used as the meta-model in the Stacking ensemble algorithm; Taking the iron ore price prediction results output by the base model as the input of the meta-model, and As the output of the meta-model.
[0037] The Stacking ensemble algorithm is an ensemble learning technique that combines the prediction results of multiple base learners and uses a meta-model for secondary training to improve the generalization performance of the overall model. Stacking (stacking method) integrates the advantages of different algorithm types by constructing a multi-level prediction system, breaking through the limitations of a single model.
[0038] In the present invention, the M deep learning models can be integrated through the Stacking integration method, that is, the integration result of the iron ore prediction results of these M deep learning models is obtained, and the final high-precision price prediction result is obtained.
[0039] The prediction method for the iron ore price in the period to be predicted includes: Obtain the data sequence of the first characteristic parameter in the period to be predicted and perform data preprocessing; Perform multi-scale sequence decomposition on the data sequence of the first characteristic parameter through discrete wavelet transform, and use correlation analysis to screen out the second characteristic parameter in the period to be predicted; Input the second characteristic parameter in the period to be predicted into the lightweight integration model to obtain the iron ore price prediction result in the period to be predicted. In order to continuously update the historical data, the iron ore price prediction result in the period to be predicted is added back to the historical data to improve the accuracy of the iron ore price prediction model.
[0040] The data preprocessing method includes missing value processing and outlier cleaning. The missing values are processed by the interpolation method of summing and averaging the data of the two moments before and the two moments after the current missing value. The expression of the interpolation method is: ; In the formula, , , , , are the data values at time t, time t-2, time t-1, time t+1, and time t+2 respectively; Outliers are determined based on the difference between the data value at time t and the data values at adjacent times being higher than a pre-set threshold, and the outliers are assigned the average value of historical data.
[0041] The parameters of the Bi-LSTM model structure are optimized by the particle swarm optimization algorithm.
[0042] As Figure 2 shown, the present invention provides an iron ore price prediction system based on multi-scale decomposition and machine learning algorithms, which is applied to an iron ore price machine learning prediction method based on multi-scale decomposition. The system includes: A data acquisition unit, configured to determine first feature parameters for iron ore price prediction according to iron ore price influencing factors, and obtain the first feature parameters and historical data of iron ore prices; A data preprocessing unit, configured to perform data preprocessing on historical data; A feature determination unit, configured to perform multi-scale sequence decomposition on the historical data of the first feature parameters through discrete wavelet transform to obtain multiple feature subsequences with different frequencies, determine the correlation between each feature subsequence and the iron ore price through a correlation analysis method, and screen out second feature parameters for iron ore price prediction from each feature subsequence based on the correlation; A model construction unit, configured to establish multiple deep learning models for iron ore price prediction respectively with each feature subsequence in the second feature parameters and the historical data of iron ore prices through a Bi-LSTM model structure, and perform integrated optimization of the prediction performance of the deep learning model corresponding to the feature subsequence with the highest correlation with the iron ore price among the multiple deep learning models through knowledge distillation technology to obtain a lightweight integrated model for iron ore price prediction; A price prediction unit, configured to predict the iron ore price within a to-be-predicted time period through the lightweight integrated model to obtain the iron ore price of the to-be-predicted time period.
[0043] The present invention combines discrete wavelet transform, a bidirectional long short-term memory artificial neural network Bi-LSTM, and knowledge distillation technology to construct a lightweight model for predicting iron ore prices, and integrates the prediction results of the complex structures of multiple Bi-LSTM models to obtain high-precision prediction performance, realizing the coordination of lightweight model structure and high-performance prediction accuracy.
[0044] The above embodiments are only exemplary embodiments of the present application and are not intended to limit the present application. The protection scope of the present application is defined by the claims. Those skilled in the art can make various modifications or equivalent replacements within the essence and protection scope of the present application, and such modifications or equivalent replacements should also be regarded as falling within the protection scope of the present application.
Claims
1. A machine learning prediction method for iron ore prices based on multi-scale decomposition, characterized in that, It includes the following steps: Determine the first characteristic parameter for iron ore price prediction according to the influencing factors of iron ore price, obtain the historical data of the first characteristic parameter and iron ore price, and perform data preprocessing on the historical data; Perform multi-scale sequence decomposition on the historical data of the first characteristic parameter through discrete wavelet transform to obtain multiple characteristic subsequences with different frequencies, determine the correlation between each characteristic subsequence and iron ore price through the correlation analysis method, and screen out the second characteristic parameter for iron ore price prediction from each characteristic subsequence based on the correlation; Establish multiple deep learning models for iron ore price prediction respectively with each characteristic subsequence in the second characteristic parameter and the historical data of iron ore price through the Bi-LSTM model structure, and perform integrated optimization of the prediction performance of the deep learning model corresponding to the characteristic subsequence with the highest correlation with iron ore price among multiple deep learning models through knowledge distillation technology to obtain a lightweight integrated model for iron ore price prediction; Predict the iron ore price in the period to be predicted through the lightweight integrated model to obtain the iron ore price in the period to be predicted.
2. The machine learning prediction method for iron ore prices based on multi-scale decomposition according to claim 1, wherein The influencing factors of the iron ore price include the US dollar index, BDI shipping price, steel price index, iron ore inventory, exchange rate, interest rate, and news sentiment analysis; The method for determining the first characteristic parameter includes: Extract 5 factors including the US dollar index, BDI shipping price, steel price index, iron ore inventory, and news sentiment analysis from the influencing factors of iron ore price to form the first characteristic parameter; Convert the news sentiment analysis in the first characteristic parameter from text data into time series data composed of three major categories of data: positive, negative, and neutral through the Naive Bayes method.
3. The machine learning prediction method for iron ore price based on multi-scale decomposition according to claim 2, wherein, The sequence decomposition method of the discrete wavelet transform includes: Set the wavelet function in the discrete wavelet transform to the 3-layer Daubechies function; Perform multi-scale sequence decomposition on the historical data of the first characteristic parameter through discrete wavelet transform to obtain multiple characteristic subsequences {A1, A2, …, AN} with different frequencies, where A1, A2, AN are respectively the identifiers of the characteristic subsequences, and N is the total number of characteristic subsequences obtained by discrete wavelet transform; The expression of the discrete wavelet transform is: ; wherein, is the wavelet coefficient of the discrete wavelet transform at the parameter , is the scale parameter, are the scaling and translation parameters respectively, is the complex conjugate function, is the time series composed of the historical data of the first characteristic parameter.
4. A machine learning prediction method for iron ore prices based on multi-scale decomposition according to claim 3, characterized in that The screening method of the second characteristic parameter includes: Calculate the Pearson correlation coefficient between each characteristic subsequence obtained by discrete wavelet transform and the time series composed of the historical data of iron ore price respectively; Screen out the characteristic subsequences with non-zero Pearson correlation coefficients in each characteristic subsequence obtained by discrete wavelet transform as the second characteristic parameter. Each characteristic subsequence in the second characteristic parameter is: {B1, B2, …, BM}, where B1, B2, BM are respectively the identifiers of the characteristic subsequences, and M is the total number of characteristic subsequences in the second characteristic parameter; Among them, in the second characteristic parameter, the characteristic subsequence with the largest absolute value of the Pearson correlation coefficient corresponds to the characteristic subsequence with the highest correlation with iron ore price, and the characteristic subsequence with the largest absolute value of the Pearson correlation coefficient is marked as B1.
5. A machine learning prediction method for iron ore prices based on multi-scale decomposition according to claim 4, characterized in that The construction method of the deep learning model includes: Using the Bi-LSTM model structure to establish M deep learning models {Bi-LSTM-B1, Bi-LSTM-B2, …, Bi-LSTM-BM} that output iron ore prediction results based on each feature subsequence in {B1, B2, …, BM} and the historical data of iron ore prices. In the formula, Bi-LSTM-B1, Bi-LSTM-B2, and Bi-LSTM-BM are respectively the identifiers of the deep learning models, M is the total number of feature subsequences in the second feature parameter, and the deep learning model corresponding to the feature subsequence with the highest correlation with iron ore prices is Bi-LSTM-B1; Among them, the Bi-LSTM model structure is composed of a forward LSTM model and a backward LSTM model. The forward LSTM model is used to obtain the output of the iron ore prediction results in the forward order at each moment of the feature subsequence, and the backward LSTM is used to obtain the output of the iron ore prediction results in the reverse order at each moment of the feature subsequence. The output results of the iron ore prediction results in the two directions are concatenated and used as the final iron ore prediction result output at each moment; The expression of the Bi-LSTM model structure is: ; In the formula, , , , are respectively the forward propagation hidden layer state, backward propagation hidden layer state, input feature subsequence, and iron ore prediction result output by the hidden layer at time t, is the activation function of the hidden layer, , , , are respectively in , in , , weight matrices, , are respectively the forward propagation hidden layer state and backward propagation hidden layer state at time t-1, , are respectively the biases of the forward propagation hidden layer and backward propagation hidden layer, is the vector concatenation symbol; The parameters of the Bi-LSTM model structure include the number of neurons in the hidden layer, weight matrices in different states, learning rate, and biases in different hidden layers; The parameters of the Bi-LSTM model structure are optimized by the particle swarm optimization algorithm.
6. The machine learning prediction method for iron ore price based on multi-scale decomposition according to claim 5, wherein, The construction method of the lightweight integrated model includes: Among the M deep learning models {Bi-LSTM-B1, Bi-LSTM-B2, …, Bi-LSTM-BM}, the deep learning model Bi-LSTM-B1 corresponding to the feature subsequence with the highest correlation with iron ore prices is marked as the lightweight model; Train the remaining M - 1 deep learning models {Bi-LSTM-B2, …, Bi-LSTM-BM}, and use each trained deep learning model in the remaining M - 1 deep learning models {Bi-LSTM-B2, …, Bi-LSTM-BM} as the teacher model one by one to perform knowledge distillation training on the lightweight model as the student model, and use the trained lightweight model as the lightweight integrated model; The training loss function of the lightweight model is: ; Wherein, is the training loss of the lightweight model, is the iron ore price prediction result output by the lightweight model, is the iron ore price prediction result output by the i-th deep learning model in {Bi-LSTM-B2, …, Bi-LSTM-BM}, is the absolute value of the Pearson correlation coefficient corresponding to the i-th deep learning model in {Bi-LSTM-B2, …, Bi-LSTM-BM}, is the iron ore price prediction result obtained by the M deep learning models {Bi-LSTM-B1, Bi-LSTM-B2, …, Bi-LSTM-BM} through the Stacking integration algorithm, is the KL divergence operation formula; The training losses of each deep learning model in {Bi-LSTM-B2, …, Bi-LSTM-BM} are: ; Wherein, is the training loss of each deep learning model in {Bi-LSTM-B2, …, Bi-LSTM-BM}, is the true value of the iron ore price.
7. A machine learning prediction method for iron ore prices based on multi-scale decomposition according to claim 6, characterized in that Determined by the Stacking integration algorithm The method includes: Among the M deep learning models {Bi-LSTM-B1, Bi-LSTM-B2, …, Bi-LSTM-BM}, each deep learning model is used as the base model in the Stacking integration algorithm, and the XGBoost model is used as the meta model in the Stacking integration algorithm; The iron ore price prediction result output by the base model is used as the input of the meta-model, and is used as the output of the meta-model.
8. A machine learning prediction method for iron ore prices based on multi-scale decomposition according to claim 7, characterized in that The prediction method for the iron ore price in the to-be-predicted time period includes: Obtain the data sequence of the first feature parameter within the to-be-predicted time period and perform data preprocessing; Perform multi-scale sequence decomposition on the data sequence of the first characteristic parameter through discrete wavelet transform, and screen out the second characteristic parameter within the time period to be predicted by using correlation analysis; Input the second characteristic parameter within the time period to be predicted into the lightweight integrated model to obtain the iron ore price prediction result within the time period to be predicted.
9. A machine learning prediction method for iron ore prices based on multi-scale decomposition according to claim 8, characterized in that, The data preprocessing method includes missing value processing and outlier cleaning. The missing values are processed by the interpolation method of summing and averaging the data of the two moments before and the two moments after the current missing value. The expression of the interpolation method is: ; Wherein, , , , , are the data values at time t, time t - 2, time t - 1, time t + 1, and time t + 2, respectively; The outliers are determined according to the difference between the data value at time t and the data values at adjacent times being higher than a preset threshold, and the outliers are assigned the average value of the historical data.
10. An iron ore price prediction system based on multi-scale decomposition and machine learning algorithms, characterized in that, Applied to a machine learning prediction method for iron ore prices based on multi-scale decomposition according to any one of claims 1-9, the system includes: A data acquisition unit for determining the first characteristic parameter for iron ore price prediction according to the influencing factors of iron ore prices, and obtaining the historical data of the first characteristic parameter and iron ore prices; A data preprocessing unit for preprocessing the historical data; A feature determination unit for performing multi-scale sequence decomposition on the historical data of the first characteristic parameter through discrete wavelet transform to obtain multiple characteristic subsequences with different frequencies, determining the correlation between each characteristic subsequence and iron ore prices through the correlation analysis method, and screening out the second characteristic parameter for iron ore price prediction from each characteristic subsequence based on the correlation; A model construction unit for establishing multiple deep learning models for iron ore price prediction respectively with each characteristic subsequence in the second characteristic parameter and the historical data of iron ore prices through the Bi-LSTM model structure, and performing integrated optimization of the prediction performance of the deep learning model corresponding to the characteristic subsequence with the highest correlation with iron ore prices among the multiple deep learning models through knowledge distillation technology to obtain a lightweight integrated model for iron ore price prediction; A price prediction unit for predicting the iron ore price within the time period to be predicted through the lightweight integrated model to obtain the iron ore price of the time period to be predicted.