A method for predicting moisture content of sintering mixtures based on multivariate interaction

Through the interactive correlation gate mechanism based on dual-stage attention and multi-scale feature mapping, the problem of ignoring the variable coupling relationship in the deep learning method in the sintering mixture moisture prediction is solved, high-precision moisture prediction and trend forecast are achieved, and the applicability and interpretability of the model are improved.

CN119851803BActive Publication Date: 2025-09-16CENT SOUTH UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411938473.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2025-09-16
Estimated Expiration
2044-12-26

AI Technical Summary

Technical Problem

Existing deep learning methods ignore the coupling relationship and interaction between variables in the prediction of sintering mixture moisture, resulting in low accuracy of time series prediction results, and low applicability and interpretability of time series prediction results to downstream tasks.

Method used

An interactive correlation gate mechanism based on dual-stage attention is adopted. Through sequence stabilization processing, multi-scale embedding feature mapping, autocorrelation gate and cross-correlation gate algorithms, a model loss function that integrates prediction value and trend guidance is designed. Combined with the sliding window method and least squares fitting processing, accurate prediction of the moisture content of sintering mixture is achieved.

Benefits of technology

The accuracy and interpretability of sintering mixture moisture prediction are improved, the model's ability to understand multivariate time series data is enhanced, and accurate classification and forecasting of future trends are achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119851803B_ABST
    Figure CN119851803B_ABST
Patent Text Reader

Abstract

A method for predicting the moisture content of sintering mixtures based on multivariate interaction. In view of the fact that existing time series prediction methods ignore the coupling relationship and historical autocorrelation between variables, the present invention proposes an autocorrelation gate and cross-correlation gate algorithm based on the interaction-correlation gate mechanism. By improving the traditional gated cyclic unit, the autocorrelation gate and cross-correlation gate are introduced to display the historical correlation of the mining sequence itself and the cross-correlation between process variables. In order to solve the problem that traditional time series prediction only uses the accuracy of the predicted value as an evaluation indicator, the present invention provides feedback guidance for the update of model parameters by designing a model loss function that integrates the predicted value and trend guidance, so that the model information is more suitable for time series prediction and downstream trend forecasting tasks. In summary, the method proposed in the present invention can accurately predict the moisture content of sintering mixtures and their trends, has the advantages of high credibility and high accuracy, and also provides an innovative idea for time series prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of sintering mixture moisture prediction, and specifically relates to a sintering mixture moisture prediction method based on multivariate interaction correlation. Background Art

[0002] Sintering is an important process for converting various powdered iron-containing raw materials into dense bodies to improve their metallurgical properties. In this process, the moisture content of the mixed materials is one of the most important parameters affecting the quality of the final product.

[0003] Appropriate moisture content can effectively improve sintering process stability, ensure production continuity, reduce energy consumption, and enhance product metallurgical properties. However, the sintering process is complex and dynamic, with numerous intertwined variables, making real-time monitoring and prediction of mix moisture content challenging. Therefore, effectively predicting changes in sinter mix moisture content is crucial in sintering engineering. This not only helps improve process stability, but also helps optimize production processes and improve product quality.

[0004] Currently, the most popular time series forecasting solutions include statistical models, machine learning models, and deep learning models. Statistical models, such as the autoregressive integrated moving average model and seasonal decomposition, rely on modeling the statistical patterns of historical data and are suitable for time series with clear statistical characteristics. Machine learning models, such as support vector machines and random forests, enhance their ability to handle nonlinear relationships by introducing more complex data feature representations and learning mechanisms. In the industrial sector, using deep learning models for time series forecasting also plays a vital role in improving production efficiency, optimizing resource allocation, and diagnosing and preventing faults.

[0005] A comprehensive comparative analysis of the advantages and limitations of the various time series forecasting methods described above reveals that deep learning-based methods have the widest scope of applicability and the best model performance. However, current deep learning methods, when applied to time series forecasting, often overlook the coupling and interactions between variables, resulting in low accuracy in time series forecast results. Furthermore, while the purpose of time series forecasting is to support subsequent trend classification tasks, such as trend classification forecasts for moisture and silicon content in sintering mixes, existing time series forecasting processes are often separated from subsequent downstream tasks, resulting in low applicability and interpretability of time series forecast results.

[0006] Considering the sensitivity of recurrent neural networks to sequence features when processing time series data and their advantages in capturing long-term and short-term dependencies, this paper innovatively proposes an interactive correlation gate mechanism based on dual-stage attention for time series prediction. This paper constructs a feature representation of time series data, designs a dual-stage attention module to enhance the model's focus on key information, develops an interactive correlation gate mechanism to capture the mutual correlation of process variables and the autocorrelation of the sequence, designs a model loss function guided by time series prediction and downstream trend forecasting tasks, and studies the interactive correlation mechanism of the gated recurrent unit. This overcomes the shortcomings of traditional time series prediction models in dealing with long-term dependencies and feature correlations, thereby achieving high-precision prediction of time series.

[0007] The present application differs from the prior art in the following ways:

[0008] Technical comparison with patent CN201910883548.8 "A time series prediction method based on GRU neural network"

[0009] This invention proposes a time series prediction method based on GRU neural network, by using GRU-

[0010] The SES model predicts time series data, performing a quadratic exponential smoothing process on the initial forecast data to obtain the final forecast data value. However, this invention fails to fully consider the complex relationships between different variables in the time series during the time series forecast process, resulting in the time series forecast results being unable to accurately predict the time series.

[0011] Technical comparison with patent CN201910013775.5 "A traffic time series prediction method based on gated network and gradient boosting regression"

[0012] This invention proposes a method for predicting traffic time series based on a gated neural network (GRU) and a gradient boosting regression model (GBR). By extracting data from multiple time dimensions, the GRU is used to mine short-term and long-term patterns in time series data. However, this invention only considers the time dimension and ignores the relationships between variables, resulting in low accuracy in time series registration and a failure to predict trends. Summary of the Invention

[0013] In response to the above problems, the present invention provides a method for predicting the moisture content of sintering mixtures based on multivariate interaction. The method proposed in the present invention can accurately predict the moisture content of sintering mixtures and its trend, has the advantages of high credibility and high accuracy, and also provides an innovative idea for time series prediction.

[0014] To achieve the above object, the technical solution adopted by the present invention is:

[0015] A method for predicting moisture content of sintering mixtures based on multivariate interaction correlation is characterized by the following specific process:

[0016] (1) Aiming at the non-stationary and long-term dependence characteristics of sintering mixture moisture time series data, a sequence stabilization processing method is designed to simplify the difficulty of subsequent prediction tasks;

[0017] (2) Aiming at the complex relationship between process variables in the sintering industry, an embedded feature mapping method is created to provide the relationship information between multiple variables for the time series prediction model, which facilitates the construction of subsequent models;

[0018] (3) The autocorrelation gate and cross-correlation gate algorithms based on the interactive correlation gate mechanism are proposed. By improving the traditional gated recurrent unit, the autocorrelation gate and cross-correlation gate are introduced to display the historical correlation of the mining sequence itself and the cross-correlation between process variables;

[0019] (4) Design a model loss function that integrates prediction value and trend guidance to provide feedback guidance for the update of model parameters, making the model information more suitable for time series prediction and downstream trend forecasting tasks;

[0020] (5) A trend curve fitting processing method based on least squares is proposed, which combines historical and predicted data to achieve accurate prediction of the mixture moisture trend classification results.

[0021] As a further improvement of the present invention, the step (1) is specifically as follows:

[0022] Step 1: In view of the non-stationary characteristics of the sintering mixture moisture time series data, the ADF test is first used to determine whether the moisture sequence has a unit root. If a unit root exists, the sequence is non-stationary, otherwise it is stationary. Formula (1) is the model selection formula;

[0023]

[0024] Among them, y t is the original moisture sequence, Δy t is the first-order difference, d is the lag order, βt is the linear trend term, ε t is the random disturbance term;

[0025] Then the ADF statistic was calculated and compared with the critical value corresponding to the 5% confidence level;

[0026] Step 2: After calculating the ADF statistic and comparing it with the critical value corresponding to the 5% confidence level, the difference between adjacent observations is calculated through the first-order difference operation, which significantly reduces the random fluctuations in the sequence, thereby enhancing the stationarity of the sequence and weakening the long-term dependence of the original sequence. The calculation formula is as follows:

[0027] Δyt =y t -y t-1 (2)

[0028] Where Δy t To represent the first-order difference sequence, y t is the sequence value at time t, y t-1 is the series value at the previous time point;

[0029] In addition, in order to verify whether the data after differential processing has weakened the long-term dependence of the original sequence, the ACF autocorrelation function is further used to analyze the correlation structure of the sequence. The calculation formula of ACF is as follows:

[0030]

[0031] Among them, k is the lag order, y t is the sequence value at time t, N is the sequence length, is the mean of the series.

[0032] As a further improvement of the present invention, the step (2) is specifically as follows:

[0033] Step 1: Embed the data points at the same time step into a vector. Suppose there are multivariate time series data X∈R B×T×D , B is the batch size; T is the length of the time series; D is the dimension of the variable, that is, the number of features. In order to capture the characteristic patterns at different time scales, we first select a set of block sizes {p1, p2, ..., p N}, p i represents the i-th block size and p i ≤T, and the block size p i It means that at this scale, each block contains time steps, for each block size p i , calculate the number of blocks at the current scale, the calculation formula is as follows:

[0034]

[0035] in, Indicates rounding down;

[0036] For each batch of samples X (b) ∈R T×D , b=1,2,...,B, according to the block size p i , divide it, and get

[0037] To block sequence The jth block is:

[0038]

[0039] Representation Block Contains sample X (b) From (j-1)p i +1 time step to jp i time steps of data, with dimension p i ×D, followed by each scale p i Define a linear map: Embed pi :R pi×D →R E , E is the embedding dimension, that is, the feature space dimension to which the block is mapped. Specifically for each block The mapping formula is:

[0040]

[0041] in, Indicates that the block matrix Expand into vectors by rows, is the scale p i The weight matrix of the embedding layer below transforms the input vector from dimension p i ×D maps to the embedding dimension E,b pi ∈R E is the bias vector of the embedding layer, is the embedding vector after embedding mapping; W e ∈R E×L is the weight matrix of the embedding layer, mapping the input from dimension L to the embedding dimension E; b e ∈R E is the bias vector of the embedding layer; e b,d,s ∈R E is the embedding vector after embedding mapping, representing the block Feature representation of

[0042] Through linear mapping, a set of multi-scale embedding vectors is obtained for all batches, all scales, and all blocks. These embedding vectors are stacked on the scale dimension to form a four-dimensional embedding tensor:

[0043]

[0044] Where N is the number of block sizes, n max =max{n i} is the maximum number of blocks at all scales;

[0045] Step 2: Then, when processing sequence data, considering the order of the sequence and the influence of position information on the model, it is necessary to add position information to the embedding vector so that the model can distinguish the features of different positions. For each segment s, a position encoding vector PE is generated.j ∈R E , whose elements are calculated according to the following formula:

[0046] i is an even number:

[0047]

[0048] i is an odd number:

[0049]

[0050] Where i∈{0,1,...,E-1} is each index of the embedding dimension;

[0051] Then for each embedding vector The position encoding vector PE j Add it to get the embedding that includes position information:

[0052]

[0053] in, The embedded vector after adding the position encoding contains the feature representation and position information of the block;

[0054] Step 3: In order to integrate the feature representations of different scales, it is necessary to fuse the multi-scale embedding vectors. There are many options for fusion methods. Here we use the average method along the scale dimension. For each batch b and block position j, the embedding vectors of different scales are averaged:

[0055]

[0056] in, represents the fused embedding vector at position j. By averaging the embedding vectors of different scales, the information of each scale can be integrated, so that the model can consider the features of different time scales in subsequent processing. b′=b×S+s represents the new sample index;

[0057] Combine the fused embedding vectors into a new embedding sequence to obtain a unified embedding tensor

[0058] Step 4: In order to capture the dependencies between different positions in the sequence, a multi-head self-attention mechanism is applied to process the fused embedded sequence. Assume that there are H attention heads in total, and the dimension of each head is d k = E / H, for the h-th attention head, define the projection matrix of query, key and value Then we have:

[0059]

[0060] in, Map it to E h The query, key, and value vector spaces of the dimension are dimensional. This process can be understood as follows: for each attention head, the input is decomposed into three different representations Q, K, and V for calculating the attention weight;

[0061] Then, the attention weight is calculated for the h-th head. The weighted similarity between the query and the key is calculated first, and then scaled and Softmax normalized to obtain the attention matrix:

[0062]

[0063] in, is the transpose of the key matrix, To avoid gradient instability caused by excessive inner product value in high-dimensional space, the Softmax function ensures the normalization of attention weights in each variable dimension. Each row of represents the degree of attention to each variable;

[0064] Apply the attention weights to the value matrix to get the output of each head:

[0065]

[0066] Concatenate the outputs of all heads along the last dimension to get the combined output matrix:

[0067]

[0068] Among them, Concat means splicing on the last dimension;

[0069] Finally, the original embedding dimension is restored through linear mapping and residual connection is performed:

[0070]

[0071] Among them, W O is the output mapping matrix, and the residual connection helps the stability of training and the propagation of gradients.

[0072] As a further improvement of the present invention, the step (3) is specifically as follows:

[0073] Step 1: In order to capture the causal relationship between the mixture moisture series and its own historical state, the autocorrelation estimate of the mixture moisture series is calculated. Specifically, for each time step t, the autocorrelation is calculated. The formula is as follows:

[0074]

[0075] in, b p are the autocorrelation weight and bias of the mixture moisture time series, is the autocorrelation weight of the mixture moisture and the hidden state at the previous moment;

[0076] Next, to capture the potential temporal correlations between process variables and achieve more efficient joint modeling between variables, the temporal data of sintering process variables are introduced into the analysis framework. The sintering process variable matrix u is normalized to eliminate the influence of dimensional differences. The weights are learned and the dimension is reduced through convolution operations to obtain the abstract feature representation u′ of the sintering-related process variables containing more critical information. The formula is as follows:

[0077]

[0078] Among them, Conv is the convolution operation, σ is the Sigmoid activation function, T u is the number of categories of the relevant process variables, T is the sequence length, and the cross-correlation between the mixture moisture and the relevant process variable characteristics is calculated. The formula is as follows:

[0079]

[0080] in, b q are the cross-correlation weights and biases between the mixture moisture and related process variable characteristics, W h q is the cross-correlation weight between the mixture moisture and the hidden state of each at the previous moment;

[0081] The calculation results of the interactive correlation gate are used to update the model hidden state, so that the above autocorrelation and cross-correlation are learned and transmitted through the hidden state weights, so that the model can remember the global state correlation information. The formula is as follows:

[0082]

[0083] Among them, b h1 、b h2 The autocorrelation bias and cross-correlation bias are used for updating the hidden state of the mixture moisture, ω represents the weight hyperparameter for the autocorrelation term and the cross-correlation term, tanh[·] is the hyperbolic tangent activation function, and ⊙ is the Hadamard product operation;

[0084] Step 2: In order to capture the importance of different time steps in the time series to the prediction results, for each time step t, calculate the attention score α t , the formula is as follows:

[0085]

[0086] Among them, h tis the hidden state at the tth time step, W attn is the time attention weight vector, and then the hidden state is weighted summed to calculate the context vector c t , the formula is as follows:

[0087]

[0088] Finally, using the context vector c t The final prediction result is generated through the output layer. The specific calculation formula is as follows:

[0089] y t =W fc c t +b fc (25)

[0090] Among them, W fc is the weight matrix of the output layer, b fc is the bias vector of the output layer.

[0091] As a further improvement of the present invention, the step (4) is specifically as follows:

[0092] In order to combine the time series prediction process with downstream tasks, a model loss function that integrates the predicted value and trend guidance is designed to provide feedback guidance for the update of model parameters, making the model information more suitable for time series prediction and downstream trend forecasting tasks;

[0093] The designed model loss function L consists of two parts

[0094] L=λ MSE MSE+λ trend ·L trend (26)

[0095] Among them, MSE is the mean square error loss term, which calculates the numerical error between the predicted value and the true value, L trend is the trend loss term, which measures the trend consistency between the predicted sequence and the true sequence, λ MSE and λ trend is the weight hyperparameter of the mean square error loss term and the trend loss term;

[0096] Trend loss term L trend The trend consistency between the predicted sequence and the true sequence is measured by calculating the mean square error between the first-order differences. The specific calculation is as follows:

[0097]

[0098] in, is the first-order difference of the forecast series.

[0099] As a further improvement of the present invention, the step (5) is specifically as follows:

[0100] Step 1: Determine the trend characterization window, that is, take the historical five-minute data as the observation benchmark, and combine it with the mixture moisture forecast value of the next three minutes, use the sliding window method to perform local analysis on the time series data, the size of the sliding window is T, and slides on the time series in turn. In each sliding window, the data point {(t i ,y i )} perform least squares polynomial fitting, using constant, linear polynomial, and quadratic polynomial models for fitting:

[0101] y=a0 (28)

[0102] y=a0+a1t (29)

[0103] y=a0+a1t+a2t 2 (30)

[0104] For each model, the coefficients {a i}, so that the error between the fitted curve and the data points is minimized. Specifically, minimize the residual sum of squares:

[0105]

[0106] Where y is the actual or predicted mixture moisture data, The fitting curve at t i The value at

[0107] Step 2: Calculate the root mean square error (RMSE) of each model and select the polynomial model with the smallest RMSE as the trend curve of the window. The calculation formula is as follows:

[0108]

[0109] Then, the selected best-fit polynomial is differentiated to obtain the first-order derivative y′ and the second-order derivative y″ for trend determination. The trend state is determined by sorting the values ​​of the first-order derivative and the second-order derivative.

[0110] The advantages of the present invention compared with the prior art are:

[0111] (1) Aiming at the non-stationary and long-term dependence characteristics of the sintering mixture moisture time series data, the present invention proposes a sequence stabilization processing method. By performing ADF test and first-order difference processing operations on the sequence, the long-term dependence of the sintering mixture moisture series is weakened, the stationarity of the moisture series data is enhanced, and the difficulty of subsequent prediction tasks is simplified.

[0112] (2) The present invention aims at the complex relationship between process variables in the sintering industry and creates an embedded feature mapping method. By dividing the input sequence according to different scales, a multi-scale input representation is formed, which helps the model capture patterns and dependencies at different time scales, thereby improving the prediction performance of subsequent models.

[0113] (3) The present invention proposes an autocorrelation gate and a cross-correlation gate algorithm based on the interactive correlation gate mechanism. By improving the traditional gated recurrent unit, the autocorrelation gate and the cross-correlation gate are introduced to show the historical correlation of the mining sequence itself and the cross-correlation between process variables, thereby enhancing the model's ability to understand multivariate time series data.

[0114] (4) The present invention designs a model loss function that integrates prediction value and trend guidance, optimizes the two target information at the same time, provides feedback guidance for the update of model parameters, and makes the model information more suitable for time series prediction and downstream trend forecasting tasks;

[0115] (5) The present invention proposes a trend curve fitting processing method based on the sliding window method and least squares, which realizes accurate classification and trend forecasting of future trends by combining historical and predicted data.

[0116] (6) The present invention uses a multi-scale feature mapping method for the first time to capture the patterns and dependencies of sintering moisture sequences at different time scales, and combines it with an interactive correlation gate mechanism based on dual-stage attention for optimization training, thereby achieving accurate prediction of sintering moisture. Through the sliding window method and least squares method, accurate forecasts of future trends are achieved, providing a new technical path for sintering moisture time series prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0117] Figure 1 is a flow chart of the method;

[0118] Figure 2 It is a steady sintering sequence;

[0119] Figure 3 It is the experimental test effect (prediction map);

[0120] Figure 4 This is the experimental test effect (classification chart). DETAILED DESCRIPTION

[0121] The following is a detailed description of the technical solution of the application in conjunction with the accompanying drawings. The described embodiments are only part of the embodiments involved in this patent. All non-innovative embodiments based on this embodiment by other researchers in this field fall within the scope of protection of this patent.

[0122] The present invention proposes a time series prediction method based on multivariate interaction correlation. Figure 1The present invention is an implementation flow chart of the method, comprising the following steps:

[0123] (1) Construction of a stable sintering mixture moisture sequence;

[0124] Due to factors such as fluctuations in temperature and humidity at the sintering site and in material ratios, time series data on the moisture content of mixed materials often exhibit non-stationary characteristics. Stationary time series are a prerequisite for most time series analysis and modeling methods, and they help simplify parameter estimation and improve prediction accuracy. This paper addresses the non-stationary and long-term dependency characteristics of time series data on sintering mixed material moisture content, proposing a sequence stabilization method that simplifies subsequent prediction tasks. This method specifically includes the following steps:

[0125] Step 1: In view of the non-stationary characteristics of the sintering mixture moisture time series data, the present invention first uses the ADF (Augmented Dickey-Fuller) test to determine whether the moisture series has a unit root. If a unit root exists, the series is non-stationary, otherwise it is stationary. Equation (1) shows the model selection formula.

[0126]

[0127] Among them, y t is the original moisture sequence, Δy t is the first-order difference, d is the lag order, βt is the linear trend term, ε t is a random disturbance term.

[0128] Then the ADF statistic was calculated and compared with the critical value corresponding to the 5% confidence level. It can be seen from the experimental results in Table 1 that the p-value is greater than the significance level, and the original hypothesis cannot be rejected. Therefore, it is determined that the mixture moisture series is non-stationary, that is, its mean and variance change over time and are not suitable for direct prediction.

[0129] Table 1 ADF test results before and after preprocessing

[0130]

[0131] Step 2: After calculating the ADF statistic and comparing it with the critical value corresponding to the 5% confidence level, the present invention calculates the difference between adjacent observations through a first-order difference operation, significantly reducing the random fluctuations in the sequence, thereby enhancing the stationarity of the sequence and also weakening the long-term dependence of the original sequence. The calculation formula is as follows:

[0132] Δy t =y t -y t-1 (2)

[0133] Where Δy t To represent the first-order difference sequence, y tis the sequence value at time t, y t-1 It is the series value at the previous time point.

[0134] In addition, to verify whether the differentially processed data has weakened the long-term dependence of the original sequence, the present invention further uses ACF (autocorrelation function) to analyze the correlation structure of the sequence. The calculation formula of ACF is as follows:

[0135]

[0136] Among them, k is the lag order, y t is the sequence value at time t, N is the sequence length, is the mean of the series.

[0137] By comparing the autocorrelations of the sequences before and after preprocessing, the present invention shows that the moisture series before differencing has a strong long-term dependence, as evidenced by significant autocorrelation even with a large lag k. However, after first-order differencing, the autocorrelation coefficient of the moisture series rapidly decays to zero and fluctuates within the confidence interval, indicating that the series has become stationary. This further validates the effectiveness of first-order differencing in weakening long-term dependence and enhancing stationarity.

[0138] In summary, a sequence stabilization processing method is proposed to address the non-stationary and long-term dependence characteristics of sintering mixture moisture time series data, which simplifies the difficulty of subsequent prediction tasks.

[0139] (2) Feature mapping methods for complex variables;

[0140] Based on the stable sintering moisture sequence after the above preprocessing, the present invention proposes a multi-scale embedding feature mapping method to capture multi-scale characteristic patterns in response to the complex relationships between variables in the sintering industrial process. This method introduces a multi-scale processing method before the embedding layer, dividing the input sequence according to different scales to form a multi-scale input representation. This helps the model capture patterns and dependencies at different time scales, thereby improving the predictive performance of subsequent models. Specifically, it includes the following steps:

[0141] Step 1: Embed the data points at the same time step into a vector. Suppose there are multivariate time series data X∈R B×T×D , B is the batch size; T is the length of the time series; D is the dimension of the variable (i.e., the number of features). In order to capture the characteristic patterns at different time scales, the present invention first selects a set of block sizes (scales) {p1, p2, ..., p N}, p i represents the i-th block size and p i ≤T, and the block size p iIt represents the time steps contained in each block at this scale. For each block size p i , calculate the number of blocks at the current scale, the calculation formula is as follows:

[0142]

[0143] in, Indicates rounding down.

[0144] For each batch of samples X (b) ∈R T×D , b=1,2,...,B, according to the block size p i , divide it, and get

[0145] To block sequence The jth block is:

[0146]

[0147] Representation Block Contains sample X (b) From (j-1)p i +1 time step to jp i time steps of data, with dimension p i ×D. Then the present invention is for each scale p i Define a linear mapping (embedding layer): Embed pi :R pi×D →R E , E is the embedding dimension, that is, the feature space dimension to which the block is mapped. The mapping formula is:

[0148]

[0149] in, Indicates that the block matrix Expand into vectors by rows, is the scale p i The weight matrix of the embedding layer below transforms the input vector from dimension p i ×D maps to the embedding dimension E,b pi ∈R E is the bias vector of the embedding layer. is the embedding vector after embedding mapping; W e ∈R E×L is the weight matrix of the embedding layer, mapping the input from dimension L to the embedding dimension E; b e ∈R E is the bias vector of the embedding layer; e b,d,s ∈R E is the embedding vector after embedding mapping, representing the block feature representation.

[0150] Through linear mapping, a set of multi-scale embedding vectors is obtained for all batches, all scales, and all blocks. These embedding vectors are stacked on the scale dimension to form a four-dimensional embedding tensor:

[0151]

[0152] Where N is the number of block sizes (scales), n max =max{n i} is the maximum number of blocks at all scales.

[0153] Step 2: Then, when processing sequence data, considering the order of the sequence and the influence of position information on the model, it is necessary to add position information to the embedding vector so that the model can distinguish features at different positions. For each segment s, a position encoding vector PE is generated. j ∈R E , whose elements are calculated according to the following formula:

[0154] i is an even number:

[0155]

[0156] i is an odd number:

[0157]

[0158] Where i∈{0,1,…,E-1} is each index of the embedding dimension.

[0159] Then for each embedding vector The position encoding vector PE j Add it to get the embedding that includes position information:

[0160]

[0161] in, The embedded vector after adding position encoding contains the feature representation and position information of the block.

[0162] Step 3: In order to integrate the feature representations of different scales, it is necessary to fuse the embedding vectors of multiple scales. There are many options for fusion methods. Here we use the method of averaging along the scale dimension. For each batch b and block position j, the embedding vectors of different scales are averaged:

[0163]

[0164] in, represents the fused embedding vector at position j. By averaging the embedding vectors at different scales, we can integrate information from each scale, allowing the model to simultaneously consider features at different time scales in subsequent processing. b′ = b × S + s represents the new sample index.

[0165] Combine the fused embedding vectors into a new embedding sequence to obtain a unified embedding tensor

[0166] Step 4: In order to capture the dependencies between different positions in the sequence, a multi-head self-attention mechanism is applied to process the fused embedded sequence. Assume that there are H attention heads in total, and the dimension of each head is d k =E / H, for the hth attention head, define the projection matrix of query, key and value Then we have:

[0167]

[0168]

[0169] in, Map it to E h This process can be understood as: for each attention head, the input is decomposed into three different representations (Q, K, V) for calculating the attention weight.

[0170] Then, the attention weight is calculated for the h-th head. First, the weighted similarity between the query and the key is calculated, and then scaled and Softmax normalized to obtain the attention matrix:

[0171]

[0172] in, is the transpose of the key matrix, The scaling factor is used to avoid excessive inner product values ​​in high-dimensional space, which may lead to unstable gradients. The Softmax function ensures the normalization of attention weights in each variable dimension. h (b) Each row of represents the degree of attention to each variable.

[0173] Apply the attention weights to the value matrix to get the output of each head:

[0174]

[0175] Concatenate the outputs of all heads along the last dimension to get the combined output matrix:

[0176]

[0177] Among them, Concat means splicing on the last dimension.

[0178] Finally, the original embedding dimension is restored through linear mapping and residual connection is performed:

[0179]

[0180] Among them, W O Is the output mapping matrix. Residual connection can help the stability of training and the propagation of gradients.

[0181] In summary, the embedded feature mapping method proposed in the present invention, through multi-scale mapping embedding and multi-head attention mechanism enhancement, can provide rich and multi-scale feature representation for subsequent time series prediction and analysis models, improve the accuracy of characterizing and analyzing complex relationships in multivariate time series, and thus improve the prediction performance and interpretability of industrial processes.

[0182] (3) Autocorrelation gate and cross-correlation gate algorithms that capture variable relationships;

[0183] Based on the multi-scale embedded feature mapping method, this paper proposes a self-correlation gate and a cross-correlation gate algorithm to further capture the complex relationships between different variables in time series. By introducing autocorrelation and cross-correlation gating mechanisms into the gated recurrent unit, this algorithm effectively models the interdependencies between variables and enhances the model's ability to understand multivariate time series data. The specific implementation steps are as follows:

[0184] Step 1: In order to capture the causal relationship between the mixture moisture series and its own historical state, the autocorrelation estimate of the mixture moisture series is calculated. Specifically, for each time step t, the autocorrelation is calculated. The formula is as follows:

[0185]

[0186] in, b p are the autocorrelation weight and bias of the mixture moisture time series, is the autocorrelation weight between the mixture moisture and the hidden state at the previous moment.

[0187] Next, to capture the potential temporal correlations between process variables and achieve more efficient joint modeling of these variables, the temporal data of the sintering process variables is introduced into the analysis framework. The sintering process variable matrix u is normalized to eliminate the effects of dimensional differences. Convolution operations are then used to learn weights and perform dimensionality reduction, yielding an abstract feature representation u′ of the sintering-related process variables that contains more critical information. The formula is as follows:

[0188]

[0189] Among them, Conv is the convolution operation, σ is the Sigmoid activation function, T u is the number of categories of the relevant process variables, and T is the sequence length. Calculate the cross-correlation between the mixture moisture and the relevant process variable characteristics The formula is as follows:

[0190]

[0191] in, b q are the cross-correlation weights and biases between the mixture moisture and related process variable characteristics, W h q is the cross-correlation weight between the mixture moisture and the hidden state of each at the previous moment.

[0192] The calculation results of the interactive correlation gate are used to update the model hidden state, so that the above autocorrelation and cross-correlation are learned and transmitted through the hidden state weights, so that the model can remember the global state correlation information. The formula is as follows:

[0193]

[0194] Among them, b h1 、b h2 are the autocorrelation bias and cross-correlation bias used for updating the hidden state of the mixture moisture, ω represents the weight hyperparameter for the autocorrelation term and the cross-correlation term, tanh[·] is the hyperbolic tangent activation function, and ⊙ is the Hadamard product operation.

[0195] Step 2: In order to capture the importance of different time steps in the time series to the prediction results, the present invention calculates the attention score α for each time step t t , the formula is as follows:

[0196]

[0197] Among them, h t is the hidden state at the tth time step, W attn is the time attention weight vector. Then the hidden state is weighted summed to calculate the context vector c t , the formula is as follows:

[0198]

[0199] Finally, using the context vector c t The final prediction result is generated through the output layer. The specific calculation formula is as follows:

[0200] y t =W fc c t +b fc (25)

[0201] Among them, W fc is the weight matrix of the output layer, b fc is the bias vector of the output layer.

[0202] In summary, the autocorrelation gate and cross-correlation gate algorithms based on the interactive correlation gate mechanism proposed in the present invention improve the traditional gated recurrent unit, introduce the autocorrelation gate and cross-correlation gate to display the historical correlation of the mining sequence itself and the cross-correlation between process variables, and effectively model the interdependence between variables.

[0203] (4) Designing a model loss function that integrates time series prediction and downstream trend forecasting tasks;

[0204] In order to combine the time series prediction process with downstream tasks, the present invention designs a model loss function that integrates the predicted value and trend guidance, provides feedback guidance for the update of model parameters, and makes the model information more suitable for time series prediction and downstream trend forecasting tasks.

[0205] The model loss function L designed by the present invention consists of two contents

[0206] L=λ MSE MSE+λ trend ·L trend (26)

[0207] Among them, MSE is the mean square error loss term, which calculates the numerical error between the predicted value and the true value, L trend is the trend loss term, which measures the trend consistency between the predicted sequence and the true sequence, λ MSE and λ trend is the weight hyperparameter of the mean square error loss term and the trend loss term.

[0208] Trend loss term L trend By calculating the mean square error between the first-order differences of the predicted sequence and the true sequence, the trend consistency between the two is measured. The specific calculation is as follows:

[0209]

[0210] in, is the first-order difference of the forecast series.

[0211] In summary, the present invention designs a model loss function that integrates prediction value and trend guidance, provides feedback guidance for the update of model parameters, and enables the model to optimize two target information simultaneously during the training process, which is more suitable for time series prediction and downstream trend forecasting tasks.

[0212] (5) Forecasting mixture moisture trends based on historical and forecast data;

[0213] Based on the multi-step prediction of sintering mix moisture content using the aforementioned model, in order to further analyze the trend changes of the prediction results and provide more intuitive decision support, this invention uses the sliding window method and least squares method based on historical data and forecast data to fit the time series data of the mix moisture content with a polynomial to classify and forecast future trends. The specific implementation steps are as follows:

[0214] Step 1: Determine the trend characterization window. That is, use the historical five-minute data as the observation benchmark and combine it with the mixture moisture forecast value for the next three minutes. Use the sliding window method to perform local analysis on the time series data. The size of the sliding window is T, which slides on the time series in turn. In each sliding window, for the data point {(t i ,y i )} perform least squares polynomial fitting, using constant, linear polynomial, and quadratic polynomial models for fitting:

[0215] y=a0 (28)

[0216] y=a0+a1t (29)

[0217] y=a0+a1t+a2t 2 (30)

[0218] For each model, the coefficients {a i}, so that the error between the fitted curve and the data points is minimized. Specifically, minimize the residual sum of squares:

[0219]

[0220] Where y is the actual or predicted mixture moisture data, The fitting curve at t i The value at .

[0221] Step 2: Calculate the root mean square error (RMSE) of each model and select the polynomial model with the smallest RMSE as the trend curve for the window. The calculation formula is as follows:

[0222]

[0223] Next, the selected best-fit polynomial is differentiated to obtain the first-order derivative y′ and the second-order derivative y″ for trend determination. By sorting out the values ​​of the first-order derivative and the second-order derivative, the trend status can be determined according to the rules in Table 2:

[0224] Table 2 Trend change description rules

[0225]

[0226] In summary, the present invention proposes a trend curve fitting processing method based on least squares, which combines historical and forecast data to analyze the trend changes of the forecast results and accurately classify and forecast future trends.

[0227] Example 1

[0228] This implementation is based on the sintering process of a domestic steel plant. Figure 1 First, for the moisture and variable sequences collected from the sintering site, the sequence stabilization method was used to obtain the following Figure 2 The stationary sequence shown in the figure is used to capture the patterns and dependencies at different time scales through the embedding feature mapping method, and then a sintering moisture prediction model is established based on the interactive correlation gate mechanism of the dual-stage attention.

[0229] In this implementation, 330 time series of moisture content of mixed materials were collected at the sintering site to evaluate the model. The training set and test set were divided in a ratio of approximately 2:1. The input window and output window lengths were 5 and 3, respectively. Through rigorous experimental verification, good experimental results were achieved. Figure 3 、 4 A graph comparing the predicted and true values ​​and a trend classification confusion matrix are presented. The model's performance evaluation metrics include a root mean square error of 0.0056, a mean absolute error of 0.044, a coefficient of determination of 0.736, and a trend classification accuracy of 85%. These experimental results demonstrate that the proposed method performs well in sintering moisture prediction, meeting the needs of field operations and providing reliable decision support for field personnel.

[0230] The above description is merely a preferred embodiment of the present invention and does not constitute any other form of limitation to the present invention. Any modification or equivalent variation based on the technical essence of the present invention shall still fall within the scope of protection claimed by the present invention.

Claims

1. A method for predicting moisture content of sintering mixtures based on multivariate interaction, characterized by: The specific process is as follows: (1) Aiming at the non-stationary and long-term dependence characteristics of the sintering mixture moisture time series data, a sequence stabilization processing method is designed to simplify the difficulty of subsequent prediction tasks; (2) Aiming at the complex relationship between process variables in the sintering industry, an embedded feature mapping method is created to provide relationship information between multiple variables for the time series prediction model, facilitating the construction of subsequent models; (3) The autocorrelation gate and cross-correlation gate algorithms based on the interactive correlation gate mechanism are proposed. By improving the traditional gated recurrent unit, the autocorrelation gate and cross-correlation gate are introduced to display the historical correlation of the mining sequence itself and the cross-correlation between process variables; The step (3) is specifically as follows: Step 1: In order to capture the causal relationship between the mixture moisture series and its own historical state, the autocorrelation estimate of the mixture moisture series is calculated. Specifically, for each time step , calculate the autocorrelation , the formula is as follows: (19); in, 、 are the autocorrelation weight and bias of the mixture moisture time series, is the autocorrelation weight of the mixture moisture and the hidden state at the previous moment, is the serial value at time t; Then, in order to capture the potential temporal correlation between process variables and realize more efficient joint modeling between variables, the temporal data of sintering process variables are introduced into the analysis framework. Normalization is performed to eliminate the influence of dimensional differences, and weights are learned and dimension reduction is performed through convolution operations to obtain abstract feature representations of sintering-related process variables containing more critical information. , the formula is as follows: (20); in, is the convolution operation, is the Sigmoid activation function, is the number of categories of the relevant process variable, For the sequence length, calculate the cross-correlation between the mixture moisture and the relevant process variable characteristics , the formula is as follows: (21); in, 、 are the cross-correlation weights and biases between the mixture moisture and the relevant process variable characteristics, is the cross-correlation weight between the mixture moisture and the hidden state of each at the previous moment; The calculation results of the interactive correlation gate are used to update the model hidden state, so that the autocorrelation and cross-correlation are learned and transmitted through the hidden state weights, so that the model can remember the global state correlation information. The formula is as follows: (22); in, 、 The autocorrelation bias and cross-correlation bias used for updating the hidden state of mixture moisture respectively, represents the weight hyperparameters for autocorrelation and cross-correlation terms, is the hyperbolic tangent activation function, is the Hadamard product operation; Step 2: In order to capture the importance of different time steps in the time series to the prediction results, for each time step , calculate the attention score , the formula is as follows: (23); in, For the The hidden state of time steps, is the time attention weight vector, and then the hidden state is weighted summed to calculate the context vector , the formula is as follows: (24); Finally, using the context vector The final prediction result is generated through the output layer. The specific calculation formula is as follows: (25); in, is the weight matrix of the output layer, is the bias vector of the output layer; (4) Design a model loss function that integrates prediction value and trend guidance to provide feedback guidance for the update of model parameters, making the model information more suitable for time series prediction and downstream trend forecasting tasks; (5) A trend curve fitting processing method based on least squares is proposed. By combining historical and predicted data, an accurate prediction of the classification results of the mixture moisture trend is achieved.

2. The method for predicting sintering mixture moisture based on multivariate interaction according to claim 1, characterized in that: The step (1) is specifically as follows: Step 1: In view of the non-stationary characteristics of the sintering mixture moisture time series data, the ADF test is first used to determine whether the moisture sequence has a unit root. If a unit root exists, the sequence is non-stationary, otherwise it is stationary. Formula (1) is the model selection formula; (1); in, is the original moisture sequence, is the first-order difference, is the lag order, is the linear trend term, is the random disturbance term; Then the ADF statistic is calculated and compared with the critical value corresponding to the 5% confidence level; Step 2: After calculating the ADF statistic and comparing it with the critical value corresponding to the 5% confidence level, the difference between adjacent observations is calculated through the first-order difference operation, which significantly reduces the random fluctuations in the sequence, thereby enhancing the stationarity of the sequence and weakening the long-term dependence of the original sequence. The calculation formula is as follows: (2); in, To represent the first-order difference sequence, is the serial value at time t, is the series value at the previous time point; In addition, in order to verify whether the data after differential processing has weakened the long-term dependence of the original sequence, the ACF autocorrelation function is further used to analyze the correlation structure of the sequence. The calculation formula of ACF is as follows: (3); Where k is the lag order, is the sequence value at time t, N is the sequence length, is the mean of the series.

3. The method for predicting sintering mixture moisture based on multivariate interaction according to claim 1, characterized in that: The step (2) is specifically as follows: Step 1: Embed the data points at the same time step into a vector. Suppose there are multivariate time series data. , B is the batch size; T is the length of the time series; D is the dimension of the variable, that is, the number of features. In order to capture the characteristic patterns at different time scales, we first select a set of block sizes , Indicates the Block size and , and the block size It means that at this scale, each block contains time steps, for each block size , calculate the number of blocks at the current scale, the calculation formula is as follows: (4); in, Indicates rounding down; For each batch of samples , , according to the block size , divide it, and get To block sequence , among which The blocks are: (5); Representation Block Includes samples Zhongcong time step to the time steps of data, with dimensions , followed by each scale Define a linear map: , is the embedding dimension, that is, the feature space dimension to which the block is mapped. For each block The mapping formula is: (6); in, Indicates that the block matrix Expand into vectors by rows, It's a scale The weight matrix of the embedding layer below transforms the input vector from dimension Mapping to embedding dimension , is the bias vector of the embedding layer, is the embedding vector after embedding mapping; is the weight matrix of the embedding layer, which transforms the input from dimension Mapping to embedding dimension ; is the bias vector of the embedding layer; is the embedding vector after embedding mapping, representing the block Feature representation of Through linear mapping, a set of multi-scale embedding vectors is obtained for all batches, all scales, and all blocks. These embedding vectors are stacked on the scale dimension to form a four-dimensional embedding tensor: (7); in, is the number of block sizes, is the maximum number of blocks at all scales; Step 2: Then, when processing sequence data, considering the order of the sequence and the influence of position information on the model, it is necessary to add position information to the embedding vector so that the model can distinguish the features of different positions. , generating a position encoding vector , whose elements are calculated according to the following formula: is an even number: (8); For odd numbers: (9); in, for each index of the embedding dimension; Then for each embedding vector , the position encoding vector Add it to get the embedding that includes position information: (10); in, The embedded vector after adding the position encoding contains the feature representation and position information of the block; Step 3: In order to integrate the feature representations of different scales, it is necessary to fuse the multi-scale embedding vectors. There are many options for fusion methods. Here we use the average method along the scale dimension. For each batch and block location , averaging the embedding vectors of different scales: (11); in, Indicates the location The fused embedding vectors can be averaged across different scales to synthesize information from each scale, allowing the model to simultaneously consider features from different time scales in subsequent processing. Represents the new sample index; Combine the fused embedding vectors into a new embedding sequence to obtain a unified embedding tensor ; Step 4: In order to capture the dependencies between different positions in the sequence, a multi-head self-attention mechanism is applied to process the fused embedding sequence. Assuming that there are attention heads, each with a dimension of , for the hth attention head, define the projection matrix of query, key and value , then: (12); (13); (14); in, Map it to The query, key, and value vector spaces of the dimension are dimensional. This process can be understood as follows: for each attention head, the input is decomposed into three different representations Q, K, and V for calculating the attention weight; Afterwards, Calculate the attention weights by first calculating the weighted similarity between the query and the key, and then scale and Normalize and get the attention matrix: (15); in, is the transpose of the key matrix, To scale the factor, avoid excessive inner product values ​​in high-dimensional space, which may lead to unstable gradients. The function ensures that the attention weights are normalized across the dimensions of each variable. Each row of represents the degree of attention to each variable; Apply the attention weights to the value matrix to get the output of each head: (16); Concatenate the outputs of all heads along the last dimension to get the combined output matrix: (17); in, Indicates splicing on the last dimension; Finally, the original embedding dimension is restored through linear mapping and residual connection is performed: (18); in, is the output mapping matrix, and the residual connection helps the stability of training and the propagation of gradients.

4. The method for predicting sintering mixture moisture based on multivariate interaction according to claim 1, characterized in that: The step (4) is specifically as follows: In order to combine the time series prediction process with downstream tasks, a model loss function that integrates the predicted value and trend guidance is designed to provide feedback guidance for the update of model parameters, making the model information more suitable for time series prediction and downstream trend forecasting tasks; Designed model loss function It consists of two parts; (26); in, is the mean square error loss term, which calculates the numerical error between the predicted value and the true value. is the trend loss term, which measures the trend consistency between the predicted sequence and the true sequence. and is the weight hyperparameter of the mean square error loss term and the trend loss term; Trend loss term The trend consistency between the predicted sequence and the true sequence is measured by calculating the mean square error between the first-order differences. The specific calculation is as follows: (27); in, is the first-order difference of the forecast series.

5. The method for predicting sintering mixture moisture based on multivariate interaction according to claim 1, characterized in that: The step (5) is specifically as follows; Step 1: Determine the trend characterization window, that is, take the historical five-minute data as the observation benchmark, and combine it with the mixture moisture forecast value of the next three minutes, use the sliding window method to perform local analysis on the time series data, the size of the sliding window is M, and it slides on the time series in turn. In each sliding window, the data points Perform least squares polynomial fitting, using constant, linear polynomial, and quadratic polynomial models for fitting: (28); (29); (30); For each model, the coefficients were estimated using the least squares method. , so that the error between the fitted curve and the data points is minimized, specifically, the residual sum of squares is minimized: (31); in, is the actual or predicted mixture moisture data, For the fitting curve The value at Step 2: Calculate the root mean square error of each model ,choose The minimum polynomial model is used as the trend curve of the window, and its calculation formula is as follows: (32); Then the selected best fitting polynomial is differentiated to obtain the first-order derivative and the second-order derivative Used for trend determination, the trend status is determined by sorting out the values ​​of the first-order derivative and the second-order derivative.

Citation Information

Patent Citations

  • A traffic time sequence prediction method based on a gating network and gradient lifting regression

    CN109886387A

  • Time sequence prediction method based on GRU neural network

    CN110647980A

  • Optimal target moisture prediction method for sintering mixture

    CN116312839A

  • Intelligent sensing method and system for sintering mixture moisture and change trend thereof

    CN117575992A