Performance index time series prediction method based on double-flow attention and sequence structure constraint

CN122692402APending Publication Date: 2026-09-04CENT SOUTH UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202610769894.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-01
Publication Date
2026-09-04

AI Technical Summary

Technical Problem

此外,针对传统逐点损失缺乏结构约束的问题,引入序列结构损失并与点级误差协同优化,以提升预测结果在点值精度、局部形态与整体趋势上的一致性

Benefits of technology

[0103] This invention focuses on complex industrial processes and proposes a time-series prediction method for performance indicators based on two-stream attention and sequence structure constraints. This method is used to predict the future trends of key performance indicators in industrial production processes and provides technical support for process monitoring, status assessment, anomaly early warning, process optimization, and control decision-making. This invention fully considers the heterogeneity of process and target variables in their time-series attributes. It extracts the dynamic features of both types of variables separately through independent two-stream modeling, avoiding information interference caused by direct mixed modeling and improving the ability to characterize the endogenous evolution of target variables and the external driving forces of process variables. Simultaneously, this invention achieves effective information transfer and collaborative modeling between process and target variables through a bidirectional intervariate interaction mechanism, enhancing the model's ability to represent the multivariate coupling relationships in complex industrial processes. Furthermore, this invention introduces sequence structure constraints, combining local point value error optimization with the overall time-series structure, which effectively improves the numerical accuracy, local dynamic changes, and global trend consistency of the prediction results. This invention can provide more accurate, stable and reliable prediction results of key performance indicators for industrial sites, which helps to improve the predictability, stability and controllability of production process operation, reduce reliance on human experience, and promote the development of industrial process management towards intelligence and refinement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122692402A_ABST
    Figure CN122692402A_ABST
Patent Text Reader

Abstract

The present application aims to provide a performance index time series prediction method based on double-flow attention and sequence structure constraint. In view of the heterogeneity of process variables and target variables in dynamic characteristics, the time series representation of the two types of variables is learned through double-flow independent coding, so as to weaken the representation interference caused by shared coding. Further, a bidirectional cross-variable interaction mechanism is designed to depict the driving effect of process variables on target variables and the reverse modulation of target variable history dynamics on process information, so as to realize effective fusion of heterogeneous information. In addition, in view of the problem that the traditional point-by-point loss lacks structure constraint, a sequence structure loss is introduced and optimized cooperatively with the point-level error to improve the consistency of the prediction result in point value accuracy, local morphology and overall trend. The present application can provide more accurate, stable and reliable key performance index time series prediction results for the industrial field, reduce the dependence on artificial experience, and promote the development of industrial process management towards intelligence and refinement.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of artificial intelligence, big data and digital information processing technology, and specifically relates to a time series prediction method for performance indicators based on dual-stream attention and sequence structure constraints. Background Technology

[0002] Processes in metallurgical, chemical, and energy industries are generally characterized by strong coupling, strong nonlinearity, large time delays, and time-varying operating conditions, making the time-series evolution of key performance indicators (KPIs) highly complex and uncertain. Against this backdrop, industrial process control relies not only on current observations but also on the accurate characterization of the future evolution trends of KPIs to support process perception, early warning decision-making, and forward-looking control. However, due to factors such as complex industrial environments, limited detection conditions, and long analysis cycles, high-frequency, continuous online observation of KPIs is often difficult. Even single-point time values ​​are difficult to obtain stably, and their continuous dynamic sequences are even more difficult to observe directly, thus limiting timely perception of process dynamics and forward-looking judgments of future states. This makes industrial operations more reliant on lag corrections rather than proactive optimization. Therefore, research on time-series prediction of KPIs is of great significance for improving the operational stability, control foresight, and production efficiency of complex industrial systems.

[0003] Modern industrial sites deploy numerous sensors that can collect high-frequency, continuous data on process variables closely related to production processes, covering multi-dimensional operational status information such as temperature, pressure, flow rate, and composition, resulting in a massive accumulation of historical data. These process variable data have deep physicochemical correlations with key performance indicators (KPIs), containing rich information on process dynamics and mappings. In recent years, the rapid development of deep time series modeling methods has provided an effective technical approach for mining complex nonlinear mapping relationships from high-dimensional process variable sequences and predicting the future evolution sequences of KPIs. This provides a solid data foundation and methodological support for constructing data-driven prediction models that translate high-dimensional process variables into future sequences of KPIs. However, existing deep time series prediction methods typically encode process variables and target variables together, ignoring their significant heterogeneity in frequency characteristics, amplitude scales, and dependency patterns. Furthermore, they generally employ pointwise loss functions for optimization, lacking explicit constraints on the local structure and trend evolution of the predicted sequence, leading to systematic biases in the prediction results at the sequence structure level. Therefore, this invention proposes a temporal prediction framework that combines two-stream attention encoding and sequence structure loss to uniformly model the internal dependencies, cross-variable interactions, and predictive sequence structure consistency of heterogeneous temporal variables.

[0004] Patent application CN121615062A, entitled "An Online Soft Measurement Method and Device for Industrial Processes Based on Multimodal Data", discloses an online soft measurement method and device for industrial processes based on multimodal data. It integrates process variable data from industrial sensors and video data from cameras, extracts multimodal time-series features through bidirectional cross-attention and dual-channel autoencoders, and uses an integrated local weighted partial least squares model and incremental update measurement to predict key performance indicators in real time.

[0005] However, this patent uses the current single-time point value as the prediction target and does not model the future evolution sequence of key performance indicators, so it cannot provide multi-step trend information for forward-looking process regulation.

[0006] Patent application CN121834206A, entitled "A Time Series Prediction Method Based on Deep Learning", discloses a time series prediction method based on deep learning. It constructs a scale set for the observed sequence through a multi-scale sliding window, extracts seasonal and trend features at each scale using a decomposition module, and obtains the prediction result through bidirectional fusion and forward propagation, thereby realizing multi-granular dynamic modeling of non-stationary time series.

[0007] However, this patent is geared towards the prediction task of a single homogeneous sequence and does not take into account the significant heterogeneity between the process variable sequence and the historical sequence of the target variable in industrial scenarios, making it difficult to fully utilize the autocorrelation structural information contained in the historical evolution of the target variable itself.

[0008] Patent application CN121117620B, entitled "A Few-Sample Irregular Time Series Prediction Method Based on a Meta-Learning Framework," discloses a few-sample irregular time series prediction method based on a meta-learning framework. It constructs a time series prediction model based on a bidirectional recurrent neural network, optimizes the model weights through a model-independent dual-loop meta-learning strategy to improve the few-sample generalization ability, and combines random position mask self-supervised training to enhance the model's adaptability to irregular sampled data.

[0009] However, this patent uses a pointwise loss function in the design of the optimization target, with the training orientation of minimizing the numerical prediction error at each time step. It lacks explicit constraints on the trend evolution and local dynamic structure of the predicted sequence, making it difficult to guarantee the rationality and consistency of the predicted sequence at the overall structural level.

[0010] In summary, existing methods for time-series prediction of key industrial performance indicators model a one-way mapping from process variables to target variables. This results in insufficient modeling of the autocorrelation evolution within the historical sequence of the target variable, which is crucial for capturing the long-term temporal characteristics of output dynamics. Furthermore, existing methods generally employ pointwise loss functions, with the primary optimization objective being to minimize the numerical deviation between the predicted and actual sequences at each time step. However, this type of loss essentially reduces the time series to a collection of independent samples, lacking explicit constraints on temporal dependencies and local structures between adjacent moments. Consequently, it fails to adequately characterize complex patterns such as periodic fluctuations, trend evolution, and local abrupt changes within the time series. Summary of the Invention

[0011] This invention aims to propose a time-series prediction method for performance indicators based on two-stream attention and sequence structure constraints. Addressing the heterogeneity of process and target variables in their dynamic characteristics, the method employs independent two-stream encoding to learn the time-series representations of both types of variables separately, thereby mitigating representational interference caused by shared encoding. Furthermore, a bidirectional intervariate interaction mechanism is designed to characterize the driving effect of process variables on target variables and the reverse modulation of process information by the historical dynamics of target variables, achieving effective fusion of heterogeneous information. In addition, to address the lack of structural constraints in traditional pointwise loss, a sequence structure loss is introduced and co-optimized with point-level error to improve the consistency of prediction results in point value accuracy, local morphology, and overall trend.

[0012] The purpose of this invention is to provide an accurate and reliable method for time-series prediction of performance indicators, providing effective data support and decision-making basis for process perception, early warning decision-making and forward-looking control in industrial sites.

[0013] This invention provides a time series prediction method for performance metrics based on two-stream attention and sequence structure constraints, specifically including the following steps:

[0014] (1) Data preprocessing and dataset construction: Obtain multi-sampling rate time series data from industrial sites, and perform time alignment, outlier correction, missing value interpolation, and zero mean normalization in sequence. Based on the preprocessed data, construct a sample set containing historical process variable sequences, historical target variable sequences, and future target variable sequences.

[0015] (2) Block-based input sequence tokenization method: The historical process variable sequence and the target variable sequence are respectively divided into sliding window block slices and independent linear projections, and the learnable position code is superimposed to generate a dual-path token sequence as the input of the subsequent encoder;

[0016] (3) Learning the temporal representation of variables based on self-attention: Independent multi-head self-attention is used to learn the intra-variable temporal features of the dual-path token sequences to obtain the temporal representations of the process variables and the target variables, which provides a basis for subsequent cross-variable interactions;

[0017] (4) Cross-variable interactive learning based on bidirectional cross attention: bidirectional cross attention is introduced to perform cross-variable interactive fusion of intravariable representations of dual-path variables, and the driving effect of process variables on target variables and the reverse modulation of process information by target variables are modeled respectively, so as to obtain enhanced representations of process variables and target variables that fuse heterogeneous information.

[0018] (5) Joint supervised optimization based on sequence structure constraints: The enhanced representation is compressed into a global representation by end-query attention pooling, and the dual-path global representation is spliced ​​to predict the target sequence. Sequence structure loss and weighted MSE loss are introduced for joint optimization, and the time series prediction model of key industrial performance indicators is trained.

[0019] As a further improvement of the present invention, the specific implementation scheme of the data preprocessing and dataset construction in step (1) is as follows:

[0020] The sampling frequencies of various data acquisition systems in industrial processes vary significantly, ranging from minutes to days. To construct a complete minute-level dataset, the key performance indicator sequences sampled from irregular events are first reconstructed into minute-level continuous sequences using linear interpolation. Subsequently, time matching can be directly performed for minute-level process variables, while hourly and daily process variables are time-matched using same-day data expansion.

[0021] Because the raw data collected from the industrial site contained some anomalies, a moving average method was used to correct them, replacing the original data with the average of the observations within its neighborhood window.

[0022]

[0023] in Indicates the half width of the sliding window. Indicates the first The variables are in position The observed values. For a small number of missing values, linear interpolation is used to complete them:

[0024]

[0025] in Indicates the first Dimensional process variables in The observed value at time, and These represent the nearest non-missing observations before and after the missing point. To eliminate dimensional differences between variables, both process and target variables are normalized to zero mean, i.e.:

[0026]

[0027] in and They represent the first The mean and standard deviation of the process variables. , and They represent the first The observed values, mean, and standard deviation of the target variable are recorded. After preprocessing, the timestamps are recorded on a uniform minute-level time axis. The process variable observation vector and the target variable observation vector are respectively Then the dataset can be represented as .for The construction window length is at any time. Historical process variable sequence and historical target variable sequence The corresponding sequence of future target variables is defined as follows: ,in and These are the process variable and the target variable dimensions, respectively. It is the prediction interval of the target variable sequence. This indicates the prediction step size.

[0028] As a further improvement of the present invention, the specific implementation scheme of the block-based input sequence tokenization method in step (2) is as follows:

[0029] Process variables and target variables exhibit significant heterogeneity in semantic attributes, temporal dynamics, and distribution characteristics. Unified modeling may lead to heterogeneous information coupling and representation degradation. A sequence tokenization method is used to analyze the historical process variable sequences separately. With historical target variable sequence For process variables, use a length of [length missing]. With step size The sliding window generates the first A segment:

[0030]

[0031] Then it is mapped to the embedding space through a separate linear projection layer:

[0032] ;

[0033] in and For learnable parameters, This indicates a vectorization operation, which expands the window matrix in chronological order into a form of length [length missing]. A one-dimensional vector; for historical target variables Use the same window parameters Maintain time alignment:

[0034] ;

[0035] ;

[0036] in , This indicates a vectorization operation, which expands the window matrix in chronological order into a form of length [length missing]. A one-dimensional vector; to inject temporal sequence information, independent learnable positional codes are added to the two token sequences. ,make The encoder input is:

[0037]

[0038] By introducing non-shared tokens for the two types of variables, the model can learn a type-aware projection space, thereby better adapting to different distribution characteristics and temporal patterns.

[0039] As a further improvement of the present invention, the specific implementation scheme of the variable temporal representation learning based on self-attention in step (3) is as follows:

[0040] A two-stream independent modeling strategy is employed for the two types of variables to learn their respective dependency patterns, providing a fully extracted temporal representation for subsequent cross-variable interactions. For example, the query, key, and value matrix can be obtained through learnable linear projection:

[0041]

[0042] in These are learnable parameters. Based on scaled dot-product attention, the output of intra-sequence attention is defined as follows:

[0043]

[0044] in Indicates the single-head feature dimension. Given the number of attention heads, the output of multi-head attention is further defined as follows:

[0045]

[0046] in The output projection matrix is ​​given, and each attention head satisfies the following:

[0047]

[0048] To construct a complete variable encoder layer, a standard Transformer block is composed of residual connections, layer normalization (LayerNorm), and a position-wise feedforward network (FFN). Specifically:

[0049]

[0050]

[0051] Indicates will The target variable is fed into a position-by-position feedforward network. The intra-sequence encoding process is the same as described above, but uses an independent parameter set. Thus we get:

[0052]

[0053] and As a temporal representation of the two variables, it is input into the subsequent intervariate interaction module to achieve information fusion between process variables and target variables at the temporal semantic level.

[0054] As a further improvement of the present invention, the specific implementation scheme of the cross-variable interactive learning based on dual-stream cross-attention in step (4) is as follows:

[0055] For explicit modeling of process variables and target variable time series representation and The cross-variable coupling relationship is addressed by introducing bidirectional cross-attention to achieve two-way information exchange and fusion. Direction, using process variables as queries and target variables as keys and values, constructs linear projections of cross-attention respectively:

[0056]

[0057] in For learnable parameters, scaled dot product attention is used to compute attention for different sequences:

[0058]

[0059] in Indicates the single-head feature dimension. Given the number of attention heads, the output of multi-head attention is further defined as follows:

[0060]

[0061] in The output projection matrix is ​​given, and each attention head satisfies the following:

[0062]

[0063] Subsequently, residual connections and layer normalization are introduced to obtain the intermediate representation after cross-variable interaction:

[0064]

[0065] The feature update is then performed using a position-wise feedforward network (FFN) to obtain the fused process variable representation:

[0066]

[0067] While preserving the inherent temporal patterns of process variables, it integrates information related to the target variable, enhancing the ability of process variable representations to perceive the target variable. The direction is determined by using the target variable representation as the query and the process variable representation as the key and value. The target variable representation is updated through independent parameters to obtain the fused target variable representation. .

[0068] As a further improvement of the present invention, the specific implementation scheme of the joint supervised optimization based on sequence structure constraints shown in step (5) is as follows:

[0069] Attention pooling with end-query is used to process the sequence Compressed into a global representation for target variable prediction. For example, construct the query using the last time step, and construct the key and value from the entire sequence:

[0070]

[0071] in These are learnable parameters. The attention weights are normalized over time:

[0072]

[0073] And the global representation is obtained by weighted summation:

[0074]

[0075] right Performing the same operation yields The two global representations are then concatenated, fused, and used for prediction.

[0076]

[0077]

[0078] in These are the weight parameters in the two-layer linear prediction mapping. These are the learnable bias parameters in the two-layer linear prediction mapping; the difference between the predicted target sequence and the true target sequence is measured by the weighted mean square error.

[0079]

[0080] in It is the first The weights of each prediction step are used to characterize the importance of different time steps. (This relies solely on...) While it can improve point accuracy, it lacks sufficient constraint on the overall structure of the predicted sequence. Therefore, a sequence structure loss is introduced to regularize the predicted sequence from three complementary perspectives: trend consistency, fluctuation pattern, and mean level. For each output dimension... definition:

[0081]

[0082] And record its mean and standard deviation as . and .

[0083] Correlation loss: To align the trend consistency between the predicted and actual sequences within the prediction window, the Pearson correlation coefficient is used to measure consistency, resulting in correlation loss.

[0084]

[0085] Variance Loss: To align the relative fluctuations of the predicted and actual sequences within the prediction window, the deviation in each dimension is calculated and transformed into a probability distribution using softmax. The Kullback-Leibler (KL) divergence measure is then used to measure the similarity between the two distributions, resulting in the variance loss.

[0086]

[0087] in This represents a probabilistic mapping function used to map the deviation of the target variable sequence from its mean into a probability distribution, thereby satisfying the requirements for calculating the KL divergence.

[0088] Mean loss is introduced to further constrain the overall level of the predicted sequence and avoid systematic mean shifts.

[0089]

[0090] In summary, the sequence structure loss is defined as:

[0091]

[0092] in These are the weighting coefficients. The ultimate training objective is the joint optimization of sequence MSE error and structural error:

[0093]

[0094] in The contribution of the structural loss is controlled. This joint loss minimizes pointwise errors and aligns with the sequence structure within a unified optimization framework, ensuring that the predicted results are consistent with the true sequence in both local numerical values ​​and overall dynamic morphology. During the training phase, the preprocessed sample set is proportionally divided into training and testing sets. The Adam optimizer is used to iteratively optimize the joint loss through backpropagation. By minimizing the pointwise and structural errors of the predicted sequence, the optimal model parameters are trained. In the prediction phase, given the historical process variable sequence and the historical target variable sequence at any given time, the data is sequentially processed through block tokenization, dual-stream independent encoding, and bidirectional intervariate interaction.

[0095] After being integrated with attention pooling, it directly outputs the future. Prediction results of the target variable sequence.

[0096] The key point of this invention is:

[0097] (1) A time series prediction framework integrating dual-stream cross attention and structural loss is proposed to address the core challenges of industrial time series prediction from two dimensions: heterogeneous time series dependence and sequence structure constraints.

[0098] (2) A dual-stream independent coding and bidirectional cross-variable interaction mechanism was constructed to characterize the heterogeneous temporal dependency of process variables and target variables and their bidirectional coupling relationship;

[0099] (3) A sequence structure consistency loss is proposed, which makes up for the limitations of traditional pointwise loss in modeling time sequence structure by explicitly constraining the overall structure of the target sequence;

[0100] (4) A joint optimization strategy for prediction error and sequence structure consistency was designed to improve the overall trend fitting ability while ensuring local numerical accuracy and enhance the stability of the model in time series prediction.

[0101] (5) The proposed key performance indicator time series prediction method can effectively improve the ability to perceive the dynamic changes of complex industrial processes and provide decision-making basis for industrial process monitoring, anomaly early warning and optimization control.

[0102] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0103] This invention focuses on complex industrial processes and proposes a time-series prediction method for performance indicators based on two-stream attention and sequence structure constraints. This method is used to predict the future trends of key performance indicators in industrial production processes and provides technical support for process monitoring, status assessment, anomaly early warning, process optimization, and control decision-making. This invention fully considers the heterogeneity of process and target variables in their time-series attributes. It extracts the dynamic features of both types of variables separately through independent two-stream modeling, avoiding information interference caused by direct mixed modeling and improving the ability to characterize the endogenous evolution of target variables and the external driving forces of process variables. Simultaneously, this invention achieves effective information transfer and collaborative modeling between process and target variables through a bidirectional intervariate interaction mechanism, enhancing the model's ability to represent the multivariate coupling relationships in complex industrial processes. Furthermore, this invention introduces sequence structure constraints, combining local point value error optimization with the overall time-series structure, which effectively improves the numerical accuracy, local dynamic changes, and global trend consistency of the prediction results. This invention can provide more accurate, stable and reliable prediction results of key performance indicators for industrial sites, which helps to improve the predictability, stability and controllability of production process operation, reduce reliance on human experience, and promote the development of industrial process management towards intelligence and refinement. Attached Figure Description

[0104] Figure 1 This is a schematic diagram illustrating the overall performance metrics of time series prediction for dual-stream attention and sequence structure constraints.

[0105] Figure 2 This is a schematic diagram of a sequence of independently segmented tokens.

[0106] Figure 3 This is a schematic diagram of time series representation learning for two-stream variables.

[0107] Figure 4 This is a diagram illustrating bidirectional cross-variable interaction.

[0108] Figure 5 This is a schematic diagram of joint supervised optimization with sequence structure constraints.

[0109] Figure 6 This is a graph showing the actual and predicted sequence of furnace top pressure. Detailed Implementation

[0110] The following examples are used to illustrate the present invention, but are not intended to limit the scope of the invention.

[0111] Example 1

[0112] This implementation case uses a 2650m³ iron smelter in a certain ironmaking plant as an example.3 The large blast furnace was used for verification.

[0113] A time-series prediction method based on two-stream attention and sequence structure constraints. A general schematic diagram is shown below. Figure 1 As shown, the specific steps include the following:

[0114] (1) Data preprocessing and dataset construction: The data collected from the blast furnace monitoring device were processed to improve the data quality, including time alignment, outlier correction, missing value interpolation, and zero mean normalization. The detailed information of the dataset after preprocessing is shown in Table 1, which provides a reliable data foundation for subsequent blast furnace top pressure sequence prediction.

[0115] Table 1. Detailed information on the furnace top pressure dataset.

[0116]

[0117] (2) Block-based input sequence tokenization method: The historical process variable sequence and the furnace top pressure sequence are respectively divided into sliding window blocks and independent linear projections to generate dual-path token sequences as inputs to the subsequent encoder, such as... Figure 2 As shown, the historical input window has a length of 48 time steps. The two sequences are divided into blocks using a sliding window of length 8 and a step size of 4, generating 11 segments in each block. Each segment is mapped to a 128-dimensional embedding space through its own independent linear projection layer to obtain the corresponding token vector. Subsequently, independent learnable positional codes are superimposed on the two token sequences to inject temporal order information, ultimately resulting in the dual-encoder input sequence.

[0118] (3) Self-attention-based learning of variable temporal representations: Intra-variable temporal features are learned by using multi-head self-attention Transformer blocks with independent parameters for the two-channel token sequences. The two sequences are each processed by... A feedforward network with head self-attention, residual connections, layer normalization, and a hidden layer dimension of 512 is used to progressively deepen the temporal dependency representations of each layer. After stacking two layers, intra-variable temporal representations of process variables and furnace top pressure are obtained, such as... Figure 3 As shown.

[0119] (4) Cross-variable interactive learning based on bidirectional cross-attention: Subsequently, a bidirectional cross-attention mechanism is introduced to perform cross-variable interactive fusion of the two representations. In the direction from process variable to furnace top pressure, the process variable representation is used as the query and the furnace top pressure representation is used as the key and value calculation. By employing cross-attention, the same operation is symmetrically performed in the opposite direction. After updating via residual connections and a feedforward network, process variables and furnace top pressure representations incorporating heterogeneous information are obtained, achieving explicit bidirectional modeling of the process variable driving effect and the inverse modulation of furnace top pressure. Figure 4 As shown.

[0120] (5) Joint supervised optimization based on sequence structure constraints: End-query attention pooling is used to compress the dual-path representation into a global representation. The attention weight is calculated using the token of the last time step as the query and the token of the whole sequence as the key and value. After weighted summation, the two global representations are concatenated and the future is output through a linear prediction head. The system predicts the furnace top pressure sequence at a prediction interval of Δt = 1 min. In the loss function design, the weighted MSE loss assigns decreasing weights to each prediction step to highlight the importance of recent predictions. The sequence structure loss is a weighted combination of correlation loss, variance loss, and mean loss, with the two components weighted... The proportional joint optimization synergistically improves both the accuracy of prediction points and the consistency of sequence structure, such as... Figure 5 As shown. During the training phase, the dataset was divided into training, validation, and test sets in an 8:1:1 ratio. The Adam optimizer was used with a learning rate of 0.0001 and a batch size of 128. The training lasted for 300 epochs, resulting in the final time-series prediction model for the top pressure of the blast furnace smelting process.

[0121] To verify the effectiveness of the proposed method, comparative experiments were conducted on a blast furnace top pressure dataset, comparing it with several mainstream time series prediction methods. The comparison methods included LSTM, DLinear, PatchTST, Transformer, and Autoformer, covering representative methods based on recursive networks, linear mapping, and attention mechanisms. The evaluation metrics used were mean squared error (MSE), mean absolute error (MAE), and hit rate (HR), measuring prediction performance from two dimensions: point value accuracy and interval prediction reliability, respectively. The experimental results are shown in Table 2. Overall, the proposed method achieved the best MSE and MAE, and obtained the highest hit rate across all seven prediction steps. Although the hit rate of each method decreased with increasing prediction step length, the proposed method consistently maintained its lead, with a smaller decrease, demonstrating stronger multi-step prediction stability and error control capabilities. Autoformer showed strong competitiveness in this task, especially in short-step prediction, but its overall performance was still lower than the proposed method. The performance decline of LSTM and DLinear was relatively gradual, but their overall prediction accuracy remained low. PatchTST and Transformer have certain advantages in short-to-medium step prediction, but their performance degrades significantly as the prediction step size increases.

[0122] Table 2. Time-series prediction results of furnace top pressure based on different models

[0123]

[0124] Figure 6 The figure shows a comparison between the predicted and actual values ​​of the blast furnace top pressure at 140 randomly selected consecutive time steps. The alternating gray and white backgrounds in the figure indicate the starting position of each sliding prediction window. Figure 6 As can be seen, the model proposed in this invention can track the rapid fluctuations of furnace top pressure well, maintaining a consistent direction of change with the true value in most abrupt intervals, and responding promptly to local peaks and troughs. In contrast, LSTM and DLinear show significant deviations in some intervals, while PatchTST, Transformer, and Autoformer, although able to reflect the overall fluctuation range, exhibit amplitude deviations and phase shifts near several extreme values, making it difficult to stably characterize the transient changes of furnace top pressure. The above results demonstrate that the method proposed in this invention can better characterize the continuous fluctuation characteristics of the furnace top pressure sequence and exhibits stronger comprehensive advantages in prediction accuracy and long-step stability.

Claims

1. A time-series prediction method for performance metrics based on two-stream attention and sequence structure constraints, characterized in that, Specifically, the following steps are included: (1) Data preprocessing and dataset construction: Obtain multi-sampling rate time series data from industrial sites, and perform time alignment, outlier correction, missing value interpolation, and zero mean normalization in sequence. Based on the preprocessed data, construct a sample set containing historical process variable sequences, historical target variable sequences, and future target variable sequences. (2) Block-based input sequence tokenization method: The historical process variable sequence and the target variable sequence are respectively divided into sliding window block slices and independent linear projections, and the learnable position code is superimposed to generate a dual-path token sequence as the input of the subsequent encoder; (3) Learning the temporal representation of variables based on self-attention: Independent multi-head self-attention is used to learn the intra-variable temporal features of the dual-path token sequences to obtain the temporal representations of the process variables and the target variables, which provides a basis for subsequent cross-variable interactions; (4) Cross-variable interactive learning based on bidirectional cross attention: bidirectional cross attention is introduced to perform cross-variable interactive fusion of intravariable representations of dual-path variables, and the driving effect of process variables on target variables and the reverse modulation of process information by target variables are modeled respectively, so as to obtain enhanced representations of process variables and target variables that fuse heterogeneous information. (5) Joint supervised optimization based on sequence structure constraints: The enhanced representation is compressed into a global representation by end-query attention pooling, and the dual-path global representation is spliced ​​to predict the target sequence. Sequence structure loss and weighted MSE loss are introduced for joint optimization, and the time series prediction model of key industrial performance indicators is trained.

2. The time series prediction method for performance indicators based on two-stream attention and sequence structure constraints according to claim 1, characterized in that, The specific implementation scheme for data preprocessing and dataset construction in step (1) is as follows: The sampling frequencies of various data acquisition systems in industrial processes vary significantly, ranging from minutes to days. To construct a complete minute-level dataset, the key performance indicator sequences sampled from irregular events are first reconstructed into minute-level continuous sequences using linear interpolation. Subsequently, time matching can be directly performed for minute-level process variables, while time matching for hourly and daily process variables is completed using data expansion on the same day. Because the raw data collected from the industrial site contained some anomalies, a moving average method was used to correct them, replacing the original data with the average of the observations within its neighborhood window. ; in Indicates the half width of the sliding window. Indicates the first The variables are in position The observed values; for a small number of missing values, linear interpolation is used to complete them: ; in Indicates the first Dimensional process variables in The observed value at time, and These represent the nearest non-missing observations before and after the missing point, respectively. To eliminate dimensional differences between different variables, both process and target variables are normalized to zero mean, i.e.: ; in and They represent the first The mean and standard deviation of the process variables. , and They represent the first The observed values, mean, and standard deviation of the target variable; After preprocessing, record the time on a unified minute-level timeline. The process variable observation vector and the target variable observation vector are respectively Then the dataset can be represented as ;for The window length at any given time is Historical process variable sequence and historical target variable sequence The corresponding sequence of future target variables is defined as follows: ,in and These are the process variable and the target variable dimensions, respectively. It is the prediction interval of the target variable sequence. This indicates the prediction step size.

3. The time series prediction method for performance indicators based on two-stream attention and sequence structure constraints according to claim 1, characterized in that, The specific implementation scheme of the block-based input sequence tokenization method in step (2) is as follows: Process variables and target variables exhibit significant heterogeneity in semantic attributes, temporal dynamics, and distribution characteristics. Unified modeling may lead to heterogeneous information coupling and representation degradation. A sequence tokenization method is used to separately analyze the historical process variable sequences. With historical target variable sequence For process variables, use a length of [length missing]. With step size The sliding window generates the first A segment: ; Then it is mapped to the embedding space through a separate linear projection layer: ; in and For learnable parameters, This indicates a vectorization operation, which expands the window matrix in chronological order into a form of length [length missing]. A one-dimensional vector; for historical target variables Use the same window parameters Maintain time alignment: ; ; in , This indicates a vectorization operation, which expands the window matrix in chronological order into a form of length [length missing]. A one-dimensional vector; to inject temporal sequence information, independent learnable positional codes are added to the two token sequences. ,make The encoder input is: ; By introducing non-shared tokens for the two types of variables, the model can learn a type-aware projection space, thereby better adapting to different distribution characteristics and temporal patterns.

4. The time series prediction method for performance indicators based on two-stream attention and sequence structure constraints according to claim 1, characterized in that, The specific implementation scheme of the self-attention-based variable temporal representation learning in step (3) is as follows: A two-stream independent modeling strategy is employed for the two types of variables to learn their respective dependency patterns, providing a fully extracted temporal representation for subsequent cross-variable interactions; For example, the query, key, and value matrix can be obtained through learnable linear projection: ; in These are learnable parameters; based on scaled dot product attention, the output of intra-sequence attention is defined as: ; in Indicates the single-head feature dimension. Given the number of attention heads, the output of multi-head attention is further defined as follows: ; in The output projection matrix is ​​given, and each attention head satisfies the following: ; To construct a complete variable encoder layer, a standard Transformer block is composed of residual connections, a layer normalization layer (LayerNorm), and a position-wise feedforward network (FFN); specifically: ; ; Indicates will The target variable is fed into a position-by-position feedforward network. The intra-sequence encoding process is the same as described above, but uses an independent parameter set. Thus we get: ; and As a temporal representation of the two variables, it is input into the subsequent intervariate interaction module to achieve information fusion between process variables and target variables at the temporal semantic level.

5. The time series prediction method for performance indicators based on two-stream attention and sequence structure constraints according to claim 1, characterized in that, The specific implementation scheme of the cross-variable interactive learning based on dual-stream cross-attention in step (4) is as follows: For explicit modeling of process variables and target variable time series representation and The cross-variable coupling relationship is addressed by introducing bidirectional cross-attention to achieve two-way information exchange and fusion; Direction, using process variables as queries and target variables as keys and values, constructs linear projections of cross-attention respectively: ; in For learnable parameters, scaled dot product attention is used to compute attention for different sequences: ; in Indicates the single-head feature dimension. Given the number of attention heads, the output of multi-head attention is further defined as follows: ; in The output projection matrix is ​​given, and each attention head satisfies the following: ; Subsequently, residual connections and layer normalization are introduced to obtain the intermediate representation after cross-variable interaction: ; The feature update is then performed using a position-wise feedforward network (FFN) to obtain the fused process variable representation: ; While preserving the temporal patterns of the process variables themselves, it integrates information related to the target variable, enhancing the ability of the process variable representation to perceive the target variable; The direction is determined by using the target variable representation as the query and the process variable representation as the key and value. The target variable representation is updated through independent parameters to obtain the fused target variable representation. .

6. The time series prediction method for performance indicators based on two-stream attention and sequence structure constraints according to claim 1, characterized in that, The specific implementation scheme of the joint supervised optimization based on sequence structure constraints shown in step (5) is as follows: Attention pooling with end-query is used to process the sequence Compression into a global representation for target variable prediction; For example, construct the query using the last time step, and construct the key and value from the entire sequence: ; in These are learnable parameters; the attention weights are normalized over time. ; And the global representation is obtained by weighted summation: ; right Performing the same operation yields Then, the two global representations are spliced ​​together and fused for prediction. ; ; in These are the weight parameters in the two-layer linear prediction mapping. These are the learnable bias parameters in the two-layer linear prediction mapping; the difference between the predicted target sequence and the true target sequence is measured by the weighted mean square error. ; in It is the first The weights of each prediction step are used to characterize the importance of different time steps; relying solely on While it can improve the accuracy of data points, it lacks sufficient constraint on the overall structure of the predicted sequence. Therefore, a sequence structure loss is introduced to regularize the predicted sequence from three complementary perspectives: trend consistency, fluctuation pattern, and mean level. For each output dimension... definition: ; And record its mean and standard deviation as . and ; Correlation loss: To align the trend consistency between the predicted and actual sequences within the prediction window, the Pearson correlation coefficient is used to measure consistency, resulting in the correlation loss. ; Variance Loss: To align the relative fluctuations of the predicted and actual sequences within the prediction window, the deviation in each dimension is calculated and transformed into a probability distribution using softmax. The Kullback-Leibler divergence is then used to measure the similarity between the two distributions, yielding the variance loss. ; in This represents a probabilistic mapping function used to map the deviation of the target variable sequence from the mean into a probability distribution, thereby satisfying the calculation requirements of KL divergence. Mean loss is introduced to further constrain the overall level of the predicted sequence and avoid systematic mean shifts. ; In summary, the sequence structure loss is defined as: ; in These are the weighting coefficients; the final training objective is the joint optimization of sequence MSE error and structural error. ; in The contribution of the structural loss is controlled by the joint loss, which minimizes pointwise errors and aligns with the sequence structure within a unified optimization framework, ensuring that the predicted results are consistent with the true sequence in both local numerical values ​​and overall dynamic morphology. During training, the preprocessed sample set is proportionally divided into training and testing sets. The Adam optimizer is used to iteratively optimize the joint loss through backpropagation, minimizing the pointwise and structural errors of the predicted sequence to obtain the optimal model parameters. In the prediction phase, given the historical process variable sequence and the historical target variable sequence at any given time, the sequence undergoes block tokenization, dual-stream independent encoding, bidirectional cross-variable interaction, and attention pooling fusion, directly outputting the future prediction. Prediction results of the target variable sequence.

Citation Information

Patent Citations

  • Few-shot Irregular Time Series Prediction Method Based on Meta-learning Framework

    CN121117620B

  • Industrial process online soft measurement method and device based on multi-modal data

    CN121615062A

  • Time sequence prediction method based on deep learning

    CN121834206A