Flexible prediction method for silicon content of blast furnace molten iron based on data fusion and time embedding
Patent Information
- Application Number
- CN202411419154.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-11
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2044-10-11
AI Technical Summary
对铁水硅含量的插值处理会造成原始信息的丢失;由于高炉是大惯性工业过程,半个小时的输入数据对于硅含量预测不够
[0034]1.本发明在钢铁工业高炉铁水硅含量预测上有较高的应用价值。
Smart Images

Figure CN119226708B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of industrial production, and in particular to a flexible prediction method for silicon content in blast furnace molten iron based on data fusion and time embedding. Background Technology
[0002] The silicon content of molten iron is a crucial indicator of its quality and indirectly reflects the furnace's thermal state, thus playing a significant role in both molten iron quality control and blast furnace operation control. However, in industrial settings, the silicon content of molten iron at the tapping point cannot be obtained in real time; it requires sampling and manual analysis, which introduces a time lag. Furthermore, because the blast furnace is a large inertial system, the impact of production operation adjustments on quality indicators also has a considerable time lag. Therefore, the early prediction of molten iron silicon content is essential for optimizing blast furnace production.
[0003] Blast furnaces are typical high-inertia, high-dimensional, complex, and nonlinear industrial production systems. Their internal structure is relatively closed, allowing data collection only through sensors installed at the furnace top, surface, and bottom. Therefore, modeling blast furnaces from a physicochemical mechanism perspective is extremely difficult. Data-driven modeling avoids the study of complex mechanisms by using only observational data to build a black-box model, thereby predicting quality indicators. Among data modeling methods, neural network models can characterize the nonlinear relationships within the system, learning autonomously from data collected in industrial settings. In particular, sequence-to-sequence model architectures based on recurrent neural networks have strong modeling capabilities for the prevalent time-series data in industry, thus becoming one of the important models for predicting the silicon content in blast furnace molten iron.
[0004] There are two types of industrial data in the blast furnace production process. One type is data automatically collected by sensors, such as oxygen enrichment flow rate, furnace top pressure, hot blast temperature, and actual pulverized coal injection rate. This type of data is uniformly sampled with a sampling interval of 10 seconds. The second type is manually analyzed data, such as the silicon content of molten iron. The time label for this type of data is determined by manual sampling, with an interval of approximately 20–40 minutes. The heterogeneous industrial data and non-uniform time labels make it difficult for recurrent neural networks that calculate based on unit time steps to directly process the raw industrial data. Therefore, it is challenging to fully utilize raw industrial data to construct an accurate predictive model for the silicon content of molten iron.
[0005] Traditional methods involve mechanical preprocessing of the raw data, averaging or downsampling high-frequency data in the time domain to reduce the data volume, and mapping them one-to-one with low-frequency data on time labels. The mapped data is then processed sequentially according to time, ignoring the non-uniform sampling characteristics of the original data. Alternatively, the original non-uniform data can be further linearly interpolated in the time domain to transform it into uniform data before subsequent modeling studies. However, this approach fails to consider the fusion of heterogeneous industrial data and the processing of non-uniformly sampled data. Interpolation of silicon content in molten iron can lead to the loss of original information; and since blast furnace operations are high-inertia industrial processes, half an hour of input data is insufficient for predicting silicon content.
[0006] Therefore, those skilled in the art are dedicated to developing a flexible prediction method for silicon content in blast furnace hot metal based on data fusion and time embedding. This method models heterogeneous raw industrial data from both uniform and non-uniform sampling, fully utilizing the effective information contained in time stamps to improve the prediction accuracy of silicon content in hot metal. It outputs predicted silicon content data for hot metal at any future time point, enhancing the flexibility and usability of the prediction and providing greater reference value for on-site production. Summary of the Invention
[0007] In view of the above-mentioned deficiencies of the prior art, the technical problem to be solved by the present invention is the problem of fusion of heterogeneous industrial data and the problem of processing non-uniform sampling data, so as to improve the prediction accuracy of silicon content in molten iron.
[0008] To achieve the above objectives, this invention provides a flexible prediction method for silicon content in blast furnace hot metal based on data fusion and temporal embedding, comprising the following steps:
[0009] Step 1: Preprocess the raw data;
[0010] Step 2: Perform correlation analysis to obtain sensor data most relevant to the silicon content of molten iron;
[0011] Step 3: Construct a dataset, where each data point consists of a time series of iron-silicon content and sensor data for the corresponding time range;
[0012] Step 4: Through the encoder's data fusion layer, high-frequency sensor data is fused with low-frequency and non-uniform historical molten iron silicon content data, and the representation sequence corresponding to the historical molten iron silicon content input time step is extracted (representation sequence refers to the vector combination at each time step output by the data fusion layer).
[0013] Step 5: Using the encoder uniform mapping layer, the non-uniform time label information is used to map the non-uniform representation sequence to uniform time points, so as to obtain the high-level representation sequence with uniform time intervals (the high-level representation sequence refers to the vector combination at each time step output by the uniform mapping layer).
[0014] Step 6: Input the high-level representation sequence output by the uniform mapping layer and the decoder state vector (decoder state vector, which refers to the state of the hidden layer neurons of the decoder network) of the previous prediction time step into the attention mechanism layer, calculate the attention weight of each high-level representation (high-level representation, which refers to the vector output by the uniform mapping layer at a certain time step) of the current prediction time step, and further calculate the context feature vector (context feature vector, which refers to the vector output by the attention mechanism layer) of the current time step.
[0015] Step 7: Input the context feature vector of the current time step and the predicted value of molten iron and silicon content of the previous time step into the decoder to calculate the decoder state vector and the predicted value of molten iron and silicon content of the current time step.
[0016] Step 8: Calculate the loss function based on the predicted and actual values of silicon content in molten iron, and then perform model backpropagation and parameter optimization.
[0017] Furthermore, step 1, data preprocessing, includes anomaly processing, missing value imputation, filtering, and normalization.
[0018] Furthermore, in step 2, the high-frequency data is sampled based on the time stamp of the low-frequency data according to a certain lag time, so as to realize the correspondence of heterogeneous data under different lag time settings.
[0019] Furthermore, in step 2, correlation analysis is performed by combining Pearson correlation coefficient, Spearman correlation coefficient, and maximum mutual information coefficient.
[0020] Further, in step 3, a time series of molten iron silicon content of a specific length is selected, and the latest silicon content data in the series is removed, with the remaining data constituting a historical molten iron silicon content series; according to the time range corresponding to the historical molten iron silicon content series, a time series of input variables such as oxygen-enriched flow rate and top pressure of a certain length are selected; all the data selected above constitute a data point in the dataset; in each data point, the latest silicon content data is used as the prediction data for the model, and all other data are used as the input data for the model.
[0021] Furthermore, in step 3, the entire dataset is divided into a training set, a validation set, and a test set; the training set is used for model training and parameter optimization, the validation set is used for hyperparameter tuning, and the test set is used to test the model's true performance.
[0022] Furthermore, in step 4, the original high-frequency uniform sensor data and low-frequency non-uniform historical iron-silicon content data are directly processed to achieve data fusion.
[0023] Furthermore, step 4 indicates that the sequence includes sensor data and historical iron-silicon content information.
[0024] Furthermore, in step 5, the representation sequence of non-uniform time intervals is mapped to uniform time points using a time tag embedding mechanism.
[0025] Furthermore, in step 5, the time-label information of the data is used to mine the temporal relationship between the data through a learnable mapping method, and the sequence of non-uniform time intervals is mapped into a uniform sequence.
[0026] Further, in step 5, the uniform mapping layer uses a learnable temporal embedding mechanism to map the non-uniform time label and uniform reference time label sequences into two corresponding sets of temporal embedding vector sequences; with the help of a learnable parameter matrix, the similarity between the two sets of vectors is calculated to obtain the weight relationship matrix from each original time point to the reference time point; using the weight relationship matrix and the subsequent linear mapping layer, the non-uniform representation sequence is mapped into a high-level representation sequence with uniform time intervals.
[0027] Furthermore, in step 6, a complete encoder-decoder structure based on the attention mechanism is adopted to calculate the temporal relationship between the decoder state vector at each prediction time point and the encoder high-level representation at each historical time point.
[0028] Furthermore, in step 7, the iron-silicon content sequence of any predicted time length is output through the decoder structure.
[0029] Existing methods for processing non-uniform data using temporal linear interpolation cannot consider the dynamic changes in the time series before and after the interpolation point, thus limiting the reliability of the interpolated data. This invention utilizes the time-stamp information of the data and employs a learnable mapping method to mine the temporal relationships between data, mapping non-uniform time interval data to uniform data. In the fifth step, a uniform mapping layer is designed to map the non-uniform encoder representation sequence to a uniform high-level representation sequence. The uniform mapping layer first uses a learnable temporal embedding mechanism to map the non-uniform time stamps and the uniform reference time stamp sequences to corresponding temporal embedding vector sequences, forming two temporal embedding matrices. Similarity calculations are performed on the two temporal embedding matrices using learnable parameters to obtain the weight relationship matrix K from each original time point to the reference time point. Using K and the subsequent linear mapping layer, the non-uniform representation sequence is mapped to a high-level representation sequence with uniform time intervals. This invention can model non-uniformly sampled raw industrial data, fully utilizing the effective information contained in the time stamps to improve the prediction accuracy of iron-silicon content.
[0030] Existing iron-silicon content prediction models cannot directly integrate raw heterogeneous industrial data. Data preprocessing based on downsampling is prone to information loss, affecting the prediction accuracy. This invention, through neural network structure design, enables high-frequency uniformly sampled industrial data and low-frequency non-uniformly sampled data to be simultaneously input into the prediction model and fused, avoiding the loss of effective information caused by downsampling. In the fourth step: based on the original sequence-to-sequence architecture, a data fusion layer is added to the encoder. The data fusion layer first feeds high-frequency input data, such as oxygen-enriched flow rate and top pressure, into a single-layer LSTM network (Long Short-Term Memory network). At this time, the hidden state of each time step of the LSTM is the representation of the input data up to the current time step. According to the low-frequency historical iron-silicon content time tags, the LSTM hidden state of the corresponding time step is extracted. This achieves the transformation of high-frequency input data into low-frequency representation data (representation data refers to the vector sequence output by the first layer of the LSTM network in the data fusion layer). Furthermore, the extracted representation is concatenated with the corresponding historical molten iron silicon content, and the concatenated data is fed into the second-layer LSTM network to obtain the encoder representation H (the encoder representation refers to the vector at a certain time step output by the data fusion layer). This representation contains 22 sets of data, such as oxygen-enriched flow rate and top pressure, as well as information on historical molten iron silicon content, thus achieving information fusion. This invention can directly fuse and process the most original heterogeneous industrial data—sensor data and manual test data—preventing information loss caused by preprocessing operations, improving the prediction accuracy of the model, and the end-to-end model from raw data to predicted data is easier to implement than prediction methods that require multiple processing steps.
[0031] Existing technologies mostly rely on single-step prediction, which cannot flexibly handle different prediction time lengths, contradicting the uneven prediction time of molten iron and silicon content. This invention utilizes a complete sequence-to-sequence architecture based on an attention mechanism, enabling the prediction model to output molten iron and silicon content prediction data of arbitrary length, meeting the requirements of arbitrary prediction time. In step six: the attention mechanism is used to better characterize the temporal relationship between system input and output variables. The dependency relationship between output data and input data at different time steps is different, and the attention weights calculated by the attention mechanism can well characterize this relationship of varying strength. The encoder's high-level representation H contains information about the input variables up to a certain time step, and the decoder's state vector s contains information about the output variables up to a certain time step. Therefore, the attention mechanism takes the high-level representation sequence and s at different time steps as inputs to calculate the attention weights between the output data and input data at different time steps. First, a series of H are convolved in the temporal domain to obtain a series of key values k. Since the key values k are calculated from H of a certain time length (convolution kernel length), they contain richer temporal features than H. The dot product of k at all input time steps and s at the previous output time step is performed. All the calculated results are fed into the Softmax layer to obtain the attention weight w for the current time step, where w represents the attention level of the current time step to the data from different input time steps. The context feature vector z for the current output time step is obtained by weighted summing of H at all different input time steps using w. z is then fed into the encoder for calculating the current output time step. Step 7: Output a predicted output sequence of arbitrary length through the decoder structure. The decoder is a GRU (Gated Recurrent Unit) network. The network inputs are the context feature vector *z* of the current time step and the molten iron / silicon content of the previous step. The hidden state is represented as *s*. The output is the decoder state vector *s* of the current time step and the predicted molten iron / silicon content. With a complete attention-based encoder-decoder structure, the model can continue to calculate the molten iron / silicon content for the next time step. First, the decoder state vector *s* is returned to the attention mechanism layer, where it undergoes another temporal dependency calculation with the already calculated encoder high-level representation *H* to obtain *z* for the next time step. Then, based on the latest *s* and *z*, the predicted molten iron / silicon content for the next time step can be calculated. Through iterative calculation, the prediction model can output molten iron / silicon content prediction data for any time length, overcoming the problem of uneven prediction time length and the bottleneck of the classic molten iron / silicon content model, which can only perform single-point predictions. This invention enables the prediction model to output molten iron / silicon content prediction data for any future time point, improving the flexibility and usability of prediction, and providing more valuable reference for on-site production workers.
[0032] No complete encoder-decoder structure has been found for predicting the silicon content in molten iron. Variable duration prediction is not only related to the decoder but also to the use of a complete encoder-decoder structure. A complete structure includes an encoder and a decoder, with the key point being that the decoder returns a series of decoder state vectors s to the attention mechanism layer for calculating a series of context feature vectors z. Existing techniques, because they only predict a single time point, only calculate one z and one s, thus eliminating the need to return s to the attention mechanism layer to calculate the next z and further obtain the next s.
[0033] Compared with the prior art, the present invention has the following obvious substantive features and significant advantages:
[0034] 1. This invention has high application value in predicting the silicon content of blast furnace hot metal in the iron and steel industry.
[0035] 2. This invention can effectively integrate heterogeneous industrial data and mine time-stamp information, resulting in a high accuracy rate in predicting the silicon content of molten iron in the model. It can meet the prediction needs of actual industrial sites and has high reference value.
[0036] 3. This invention can directly process raw industrial sensor data and manual test data in an end-to-end manner, and predict the silicon content of molten iron, avoiding cumbersome data processing and making it highly practical.
[0037] 4. The prediction time of this invention is relatively flexible, which can meet the variable prediction needs of the production site and has high reference value.
[0038] The following will further explain the concept, specific structure, and technical effects of the present invention in conjunction with the accompanying drawings, so as to fully understand the purpose, features, and effects of the present invention. Attached Figure Description
[0039] Figure 1 This is an overall flowchart of a preferred embodiment of the method for predicting the silicon content in molten iron according to the present invention;
[0040] Figure 2 This is an encoder data fusion layer according to a preferred embodiment of the present invention;
[0041] Figure 3 This is a preferred embodiment of the encoder uniform mapping layer of the present invention;
[0042] Figure 4 This is an attention mechanism of a preferred embodiment of the present invention;
[0043] Figure 5 This is a preferred embodiment of the decoder structure of the present invention;
[0044] Figure 6This is a model prediction effect of a preferred embodiment of the present invention;
[0045] Figure 7 This is a comparison of the test set accuracy during the training process of different datasets according to a preferred embodiment of the present invention. Detailed Implementation
[0046] The following description, with reference to the accompanying drawings, illustrates several preferred embodiments of the present invention to make its technical content clearer and easier to understand. The present invention can be embodied in many different forms, and the scope of protection of the present invention is not limited to the embodiments mentioned herein.
[0047] In the accompanying drawings, components with the same structure are indicated by the same numerical designation, and components with similar structures or functions are indicated by similar numerical designations. The dimensions and thicknesses of each component shown in the drawings are arbitrary, and the present invention does not limit the dimensions and thicknesses of each component. To make the illustrations clearer, the thickness of some components has been appropriately exaggerated in the drawings.
[0048] In this embodiment, a 2650m steel plant in China is used as an example. 3 Take, for example, the actual blast furnace production data collected from January 1, 2021 to December 31, 2021 for a certain blast furnace.
[0049] A method for predicting the variable duration of silicon content in molten blast furnace iron based on heterogeneous data fusion and temporal information embedding, such as... Figure 1 As shown, it includes the following steps:
[0050] Step 1: Preprocess the raw data collected from the steel plant. First, outlier data (outliers and missing values) is removed and filled in. Due to factors such as sensor failure and blast furnace shutdown for maintenance, the raw data contains outliers and missing values. Outliers are identified using box plots. Shorter periods of outlier data are identified as sensor failures and filled in by averaging the normal data at both ends of the outlier time period. Longer periods of outlier data are identified as blast furnace shutdown for maintenance, and therefore all data from that period are discarded.
[0051] The imputed data is then digitally filtered to smooth out any abnormal spikes.
[0052] Except for the sampling interval of silicon content in molten iron, which is approximately 20 to 40 minutes, the sampling interval for all other data (oxygen-enriched flow rate, top pressure, etc.) is 1 minute. Therefore, the unit step of the recurrent neural network in processing the data corresponds to 1 minute of actual time.
[0053] Zero-mean normalization is performed on the data to avoid the impact of different data magnitudes on the model accuracy.
[0054] Step 2: The correlation between molten iron silicon content and other variables such as oxygen-enriched flow rate and top pressure was analyzed using three methods: Pearson correlation coefficient, Spearman correlation coefficient, and maximum mutual information coefficient. The input variables for the model were then selected based on the analysis results and expert experience. Since the sampling frequency of molten iron silicon content is lower than that of other variables, the time label [t1, t2, t3, ..., t] of the molten iron silicon content data was used in the correlation analysis. n Based on [t1-lag, t2-lag, t3-lag, ..., t], high-frequency data is sampled according to a certain lag time, with the sampling time points being [t1-lag, t2-lag, t3-lag, ..., t]. n -lag]. Finally, 22 variables were selected as inputs to the model (excluding historical molten iron silicon content), including: CO2, actual pulverized coal injection rate of the previous hour, cold air flow rate, actual wind speed, oxygen enrichment flow rate, oxygen enrichment rate, actual pulverized coal injection rate of the current hour, furnace gas index, furnace gas volume, hot air temperature, set pulverized coal injection rate, permeability index, resistance coefficient, No. 1 top pressure, No. 2 top pressure, top temperature downcomer, top temperature northeast, top temperature southeast, top temperature northwest, top temperature southwest, blast kinetic energy, and blast humidity.
[0055] Step 3: Construct the dataset. The earliest time label in all data corresponds to step 0, 1 minute later is step 1, and n minutes later corresponds to step n. After the first two steps, the smallest unit of all data event labels is 1 minute. Except for the time period corresponding to the discarded data, each time step corresponds to a set of 22 input data such as oxygen-enriched flow rate and top pressure, but does not necessarily correspond to the iron-silicon content data.
[0056] Select a time series of molten iron silicon content of length [y] n(1) ,y n(2) ,y n(3) ,…,y n(step) According to its corresponding time step [n(1),n(2),n(3),…,n(step)], the time series of 22 variables such as oxygen-enriched flow rate and top pressure at time step [n(1)-40,…,n(1)-1,n(1),n(1)+1,…,n(step-1)] are selected. n(1)-40, …,x n(1)-1 ,x n(1) ,x n(1)+1 ,…,x n(step-1) This constitutes "one data point" in the dataset. Because the sampling time intervals for molten iron silicon content are different, the step lengths corresponding to the time series of molten iron silicon content with different step lengths are also different, resulting in different lengths for each data point. Each data point contains time series corresponding to 22 variables such as oxygen-enriched flow rate and top pressure [x]. n(1)-40 ,…,x n(1)-1 ,x n(1) ,x n(1)+1 ,…,xn(step-1) ] and historical iron-silicon content sequence [y n(1) ,y n(2) ,y n(3) ,…,y n(step-1) ] will be used as the input to the model, y n(step) As the model's predicted output.
[0057] The entire dataset contains 12,508 data points, including 10,000 data points in the training set, 1,308 data points in the validation set, and 1,200 data points in the test set. The training set is used for model training and parameter optimization, the validation set is used for hyperparameter tuning, and the test set is used to test the model's true performance. Steps four through six describe the processing of a single data point in the dataset.
[0058] Step 4: Building upon the existing sequence-to-sequence architecture, a data fusion layer is added to the encoder. This layer directly utilizes 22 sets of high-frequency data, such as oxygen-enriched flow rate and top pressure, as well as lower-frequency and uneven historical iron-silicon content data, thus avoiding information loss caused by excessive processing of the original data. The structure of the data fusion layer is as follows: Figure 2 As shown, firstly, the 22 sets of input data [x] representing "one data point" in the dataset are... n(1)-40 ,…,x n(1)-1 ,x n(1) ,x n(1)+1 ,…,x n(step-1) The data is fed into a single-layer LSTM network to extract the hidden state [h] at time steps [n(1),n(2),n(3),…,n(step-1)]. n(1) ,h n(2) ,h n(3) ,…,h n(step-1) Since the hidden state of the LSTM depends on all the input data up to the current step, it can be assumed that [h] n(1) ,h n(2) ,h n(3) ,…,h n(step-1) [] represents the original input data at time steps [n(1),n(2),n(3),…,n(step-1)], containing all the information of the original data. This operation transforms high-frequency input data into low-frequency representation data and achieves alignment with historical molten iron silicon content at non-uniform time steps. Next, the aligned and concatenated data [] will be used. <h n(1) ,y n(1) >, <h n(2) ,y n(2) >, <h n(3) ,y n(3) >,…, <h n(step-1) ,y n(step-1)The sequence [H] is fed into the second-layer LSTM network to obtain the encoder representation sequence. n(1) H n(2) H n(3) ,…,H n(step-1) This sequence contains 22 sets of data, including oxygen-enriched flow rate and top pressure, as well as information on historical silicon content in molten iron, achieving information fusion.
[0059] Step 5: Use a time-stamp embedding mechanism to map the representation sequence of non-uniform time intervals to uniform time points. The structure of the uniform mapping layer is as follows: Figure 3 As shown, a uniformly distributed reference time step [n] is first given. ref (1),n ref (2),…,n ref [step-1], the original time step and the reference time step are transformed into two sets of time embedding vector sequences through the time embedding layer and the linear mapping layer, forming the time embedding matrices φ and φ ref Functions of the temporal embedding layer The calculation corresponding to the linear mapping layer is shown below, where ω and β are learnable parameters, and d determines the embedding vector. Dimensions.
[0060]
[0061] Furthermore, with the help of learnable parameter matrices W and V, the weight relationship matrix K between each original time step and the reference time step can be obtained according to the following similarity calculation formula.
[0062]
[0063] Matrix K can represent the original non-uniform sequence [H] n(1) H n(2) H n(3) ,…,H n(step-1) The sequence is mapped to a uniform high-level representation sequence at the reference time step as follows: [H] nref(1) H nref(2) H nref(3) ,…,H nref(step-1) ].
[0064]
[0065] The uniform mapping layer fully utilizes the original time stamps, embedding non-uniform temporal information into the representation sequence to model the non-uniform temporal relationships of the original data. While obtaining a high-level representation sequence with a uniform time distribution, it rationalizes downstream standard neural network modeling methods based on equal-time step sizes. Subsequent steps, lacking the relationship between n and n... ref The difference is that it is abbreviated as n.
[0066] Step 6: Utilize attention mechanisms to better characterize the time-domain influence relationship between system input and output variables. For example... Figure 4 As shown, this is the overall structure of the encoder, which includes a data fusion layer and an attention mechanism. [H] n(1) H n(2) H n(3) ,…,H n(step-1) Perform a 1D convolution in the temporal domain to obtain the key value [k] at the corresponding time step. n(1) ,k n(2) ,k n(3) ,…,k n(step-1) The query value is the decoder state vector s. Initial hidden state s n(step-1) It is obtained by performing a temporal 1D convolution on the encoder's high-level representation H in the last few time steps. The decoder state vector s at time step n (step-1) and any subsequent time step is then convolved with [k...]. n(1) ,k n(2) ,k n(3) ,…,k n(step-1) Perform a dot product, then pass it through a softmax layer to obtain the attention weights [w]. n(1) ,w n(2) ,w n(3) ,…,w n(step-1) The weights are used to perform a weighted summation of the encoder's high-level representations to obtain the context feature vector z corresponding to time step n (step-1) and subsequent time steps. Since the attention weights calculated from the decoder state vector s differ at different time steps, the corresponding context feature vector z also differs. [s] n(step-1) ,s n(step-1)+1 ,s n(step-1)+2 ,…,s n(step)-1 [z] can be calculated. n(step-1) ,z n(step-1)+1 ,z n(step-1)+2 ,…,z n(step)-1 ].
[0067] The weights calculated using the attention mechanism represent the time-domain dependencies between input and output variables. Key value k n(1) It represents all the information from input variable x up to step n(1), s n(step-1) This implies information about the output variables in n(step-1) steps. If w is calculated after the attention mechanism... n(1) If the value of w is large, it indicates that the input data in step n(1) has a strong relationship with the output data in step n(step-1); if w n(1)If the value is small, then the opposite is true. Each time step from n(step-1) to n(step) is different, and the output at different time steps naturally has different temporal dependencies with the input data. Therefore, attention mechanism calculations are performed at each time step, and the resulting attention weights effectively characterize the dependencies between the output and input variables at different time steps.
[0068] Step 7: Output a predicted output sequence of arbitrary length through the decoder structure. Figure 5 For the structure of the decoder, using s n(step-1) The decoder's hidden state is initialized by using the context feature vector z calculated at the current time step. n(step-1)+1 Compared with the predicted output y of the previous time step n(step-1) Together, they are fed into the GRU network as input to obtain the hidden state s. n(step-1)+1 The predicted output y is then obtained through a linear output layer. n(step-1)+1 . s n(step-1)+1 The data is then fed back into the attention mechanism described in step five to calculate the z-value required for the next time step. n(step-1)+2 Along with y n(step-1)+1 Together, they serve as the input for the next time step. After the above iterative calculations, the predicted output sequence [y] can be obtained. n(step-1)+1 ,y n(step-1)+2 ,y n(step-1)+3 ,…,y n(step) It's important to note here that the original molten iron silicon content only has true data at time steps n(step-1) and n(step), but no true data at time steps [n(step-1)+1, n(step-1)+2, ..., n(step)-1]. Therefore, when calculating the prediction accuracy, only the predicted value y is considered. n(step) Compared with the true value y n(step) The error between them. Furthermore, in the dataset, the length of [n(step-1), n(step)] varies for each data point, due to the uneven sampling of molten iron and silicon content. However, the recurrent neural network structure in the decoder can obtain a predicted output sequence of arbitrary length after n(step-1) through iterative calculation, naturally overcoming the problem of uneven prediction time length and achieving flexible prediction of molten iron and silicon content.
[0069] Step 8: Calculate the loss function based on the predicted and actual values of silicon content in molten iron, and then perform model backpropagation and parameter optimization. The training set is used for model parameter optimization, employing the Adams algorithm and the mean absolute value loss function. The validation set is used to adjust model hyperparameters, primarily including step, temporal convolution length, hidden layer variable dimension, key-value dimension, and data batch size. The test set is used to test the model's real-world performance. Experimental results are as follows: Figure 6 The figure shows a comparison between predicted and actual values of silicon content in molten iron, along with an error range. In the figure, solid square lines represent predicted values, solid triangle lines represent actual values, solid lines represent the error between the two, and the error between the two dashed lines is less than ±0.1.
[0070] There are three metrics for evaluating prediction models of silicon content in molten iron, one of which is the hit rate (HR):
[0071]
[0072] Where N is the number of predicted samples, and the error e k The following condition is satisfied, where y true This is the actual value for silicon content:
[0073]
[0074] The second type is the root mean square error (RMSE):
[0075]
[0076] The third type is Mean Absolute Percent Error (MAPE):
[0077]
[0078] The model achieved a prediction accuracy of 95.15% and a root mean square error of 0.0443. The model demonstrated excellent prediction performance. Furthermore, following traditional data preprocessing methods, high-frequency sensor data were averaged over time to ensure a one-to-one correspondence between the sensor data and the iron-silicon content analysis values. Linear interpolation was then performed to obtain uniform data, which was used to construct the dataset. Figure 7 Comparing the accuracy of models trained on datasets with the same data frequencies after downsampling and linear interpolation with those trained on datasets containing heterogeneous data categories with different frequencies, it can be seen that the model trained on the original dataset (containing heterogeneous industrial data with different frequencies) without excessive preprocessing has better performance metrics. Therefore, it can be concluded that the model designed in this invention can effectively preserve the information of the original data, prevent data preprocessing operations such as time-domain downsampling and linear interpolation from damaging the original data, and ensure the model's prediction accuracy.
[0079] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.
Claims
1. A flexible prediction method for silicon content in blast furnace hot metal based on data fusion and temporal embedding, characterized in that, Includes the following steps: Step 1: Preprocess the raw data; Step 2: Perform correlation analysis to obtain sensor data most relevant to the silicon content of molten iron; Step 3: Construct a dataset, where each data point consists of a time series of iron-silicon content and sensor data for the corresponding time range; Step 4: Through the encoder's data fusion layer, high-frequency sensor data is fused with low-frequency and non-uniform historical molten iron silicon content data, and the representation sequence corresponding to the historical molten iron silicon content input time step is extracted. Step 5: Through the encoder uniform mapping layer, using the non-uniform time label information, the non-uniform representation sequence is mapped to uniform time points to obtain a high-level representation sequence with uniform time intervals. Step 6: Input the high-level representation sequence output by the uniform mapping layer and the decoder state vector of the previous prediction time step into the attention mechanism layer, calculate the attention weight of each high-level representation at the current prediction time step, and further calculate the context feature vector of the current time step. Step 7: Input the context feature vector of the current time step and the predicted value of molten iron and silicon content of the previous time step into the decoder to calculate the decoder state vector and the predicted value of molten iron and silicon content of the current time step. Step 8: Calculate the loss function based on the predicted and actual values of silicon content in molten iron, and then perform model backpropagation and parameter optimization; In step 5, the time-label information of the data is used to mine the temporal relationship between the data through a learnable mapping method, and the sequence of non-uniform time intervals is mapped to a uniform sequence. In step 5, the uniform mapping layer uses a learnable temporal embedding mechanism to map the non-uniform time label and uniform reference time label sequences into two corresponding sets of temporal embedding vector sequences. By using a learnable parameter matrix, the similarity between two sets of vectors is calculated, and the weight relationship matrix from each original time point to the reference time point is obtained. Using the weight relationship matrix and the subsequent linear mapping layer, the non-uniform representation sequence is mapped into a high-level representation sequence with uniform time intervals.
2. The flexible prediction method for silicon content in blast furnace hot metal based on data fusion and temporal embedding as described in claim 1, characterized in that, Step 1, data preprocessing, includes outlier processing, missing value imputation, filtering, and normalization.
3. The flexible prediction method for silicon content in blast furnace hot metal based on data fusion and temporal embedding as described in claim 1, characterized in that, Step 2 involves using the time stamp of low-frequency data as a reference to sample high-frequency data, thereby achieving correspondence between heterogeneous data under different lag time settings.
4. The flexible prediction method for silicon content in blast furnace hot metal based on data fusion and temporal embedding as described in claim 1, characterized in that, Step 2 involves performing a correlation analysis by combining the Pearson correlation coefficient, Spearman correlation coefficient, and maximum mutual information coefficient.
5. The flexible prediction method for silicon content in blast furnace hot metal based on data fusion and temporal embedding as described in claim 1, characterized in that, Step 4 involves directly processing the original high-frequency uniform sensor data and low-frequency non-uniform historical molten iron silicon content data to obtain a representation sequence.
6. The flexible prediction method for silicon content in blast furnace hot metal based on data fusion and temporal embedding as described in claim 1, characterized in that, In step 5, the time tag embedding mechanism is used to map the representation sequence of non-uniform time intervals to uniform time points.
7. The flexible prediction method for silicon content in blast furnace hot metal based on data fusion and temporal embedding as described in claim 1, characterized in that, In step 6, a complete encoder-decoder structure based on the attention mechanism is adopted to calculate the temporal relationship between the decoder state vector at each prediction time point and the encoder high-level representation at each historical time point.
8. The flexible prediction method for silicon content in blast furnace hot metal based on data fusion and temporal embedding as described in claim 1, characterized in that, In step 7, the iron-silicon content sequence of any predicted time length is output through the decoder structure.