Method for predicting operation and maintenance time sequence data based on sequence information perception
By sharding and processing the operation and maintenance time series data with an attention mechanism, and dynamically integrating long-term and short-term information, the problem of capturing global change patterns and long-range dependencies in the prediction of operation and maintenance time series data is solved, achieving more efficient prediction results.
Patent Information
- Application Number
- CN202510842606.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-10-03
AI Technical Summary
Existing technologies find it difficult to effectively capture the global dynamic change patterns and long-range dependencies in operation and maintenance time series data, resulting in insufficient prediction accuracy.
By sharding the long-term operation and maintenance time series data and mapping it into a high-dimensional embedding representation, combining the global attention and self-attention values to perform element-by-element multiplication, accumulation and operation, the long-term and short-term sequential information is dynamically integrated to generate prediction results.
It improves the ability to capture the long-term cumulative effects and trends of operation and maintenance data, reduces computational complexity, enhances information integration effects, adapts to data changes, and improves prediction accuracy.
Smart Images

Figure CN120744447A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of operation and maintenance data prediction, and in particular to a method for predicting operation and maintenance time series data based on sequential information perception. Background Art
[0002] Long-sequence prediction technology for operation and maintenance time series data plays a crucial role in numerous operation and maintenance scenarios, including enterprise information system maintenance, network equipment monitoring, cloud computing platform resource scheduling, and industrial equipment health management. The massive amount of operation and maintenance monitoring data generated by modern network infrastructure and industrial systems exhibits distinct time series characteristics. Accurately predicting this operation and maintenance time series data is crucial for fault early warning, performance optimization, resource planning, and cost control. Predicting operation and maintenance time series data requires models that can deeply analyze historical monitoring data, capturing complex dependencies across different time scales and the dynamic evolution of system states, to accurately predict future system behavior and potential anomalies.
[0003] In recent years, the emergence of deep learning technology, particularly the Transformer architecture, has revolutionized the way time series data is processed, leveraging its powerful self-attention mechanism. This has effectively improved the ability to capture long-range dependencies within sequences and facilitated the efficient implementation of parallel computing. Recurrent neural networks (RNNs) and their variants have also excelled in modeling sequential patterns, capturing the dynamics of time series data.
[0004] However, traditional time series prediction models often have limitations in capturing long-range dependencies, handling sequential change patterns, or efficient parallel computing when processing long-series data commonly found in operations and maintenance scenarios, making it difficult to meet the high-precision prediction needs of modern complex operations and maintenance environments. The Transformer architecture lacks understanding of the global change patterns in operations and maintenance data sequences and is insufficiently capable of modeling sequential patterns during the dynamic evolution of system states. These limitations make it difficult for the model to accurately capture equipment load fluctuations, resource utilization trends, and the periodic characteristics of system performance indicators, thus affecting prediction accuracy. On the other hand, while recurrent neural networks (RNNs) excel at modeling sequential patterns, they are severely constrained by low computational efficiency and vanishing gradients when processing long-series data commonly found in the operations and maintenance field. Summary of the Invention
[0005] The present invention proposes a prediction method for operation and maintenance time series data based on sequential information perception, which solves the problem that the existing technology is difficult to effectively capture the global dynamic change patterns and long-range dependencies in long-sequence data.
[0006] To solve the above technical problems, the present invention provides a method for predicting operation and maintenance time series data based on sequential information perception, comprising the following steps:
[0007] Step S1: Slice the long operation and maintenance time series data and map the sliced long operation and maintenance time series data into a high-dimensional embedding representation through linear transformation and position encoding;
[0008] Step S2: Perform a linear transformation on the high-dimensional embedding representation to obtain a query matrix, a key matrix, a value matrix, and a global attention matrix representing the weight of each time position's contribution to the entire sequence. Calculate the self-attention value of each time position based on the query matrix, key matrix, and value matrix.
[0009] Step S3: performing element-wise multiplication of the global attention and self-attention values, accumulating all items before each time position in the product of the element-wise multiplication to obtain long-term sequence information; performing weighted combination of the self-attention values at the current time position and the two time positions before the current time position to obtain short-term sequence information;
[0010] Step S4: Dynamically fuse long-term and short-term sequential information and self-attention, add the fused features to the high-dimensional embedding representation to obtain the fused feature tensor, and map the fused feature tensor to the prediction result.
[0011] Preferably, the sharding of the long operation and maintenance time series data in step S1 includes the following steps:
[0012] Step S11: normalize the long operation and maintenance time series data, and transpose the normalized long operation and maintenance time series data into ;
[0013] Step S12: Set the length to , the step length is ,right Fragment the data obtained after sharding. Reshape the fragments into a 2D tensor , the two-dimensional tensor The expression is:
[0014] ;
[0015] Where, Indicates sharding operation; Represents a reshape operation; Indicates the expansion operation; Represents a fill operation.
[0016] Preferably, the expression of the high-dimensional embedding representation in step S1 is:
[0017] ;
[0018] ;
[0019] ;
[0020] In the above formula, Long operation and maintenance time series data after embedding position encoding; is the weight matrix; Encode for position; is the position index; is the dimension index; is the embedding dimension.
[0021] Preferably, the expression of the self-attention value in step S2 is:
[0022] ;
[0023] ;
[0024] ;
[0025] ;
[0026] ;
[0027] ;
[0028] In the above formula, For self-attention; is the activation function; is the query matrix; is the bond matrix; is the value matrix; is the embedding dimension; To generate queries and key Basic characteristics of To generate values and global attention weight Basic characteristics of Represents matrix dot product; for Activation function; is the scaling factor; 、 is the bias term; It is the long operation and maintenance time series data after embedding position encoding.
[0029] Preferably, the expression of the global attention in step S2 is:
[0030] ;
[0031] ;
[0032] In the above formula, For global attention; is the activation function; To generate values and global attention weight Basic characteristics of is the scaling factor; is the bias term; Represents matrix dot product; for Activation function; is the weight matrix; is the bias term.
[0033] Preferably, the expression of the long-term sequence information in step S3 is:
[0034] ;
[0035] Where, is long-term sequential information; Indicates accumulation and operation; For self-attention; Represents matrix dot product; For global attention.
[0036] Preferably, the expression of the short-term sequence information in step S3 is:
[0037] ;
[0038] Where, is short-term sequential information; For self-attention; represents an all-zero vector; Indicates that all vectors are merged in the last dimension.
[0039] Preferably, the expression of the fused feature tensor in step S4 is:
[0040] ;
[0041] ;
[0042] Where, is the fused feature tensor; Represents normalization operation; Long operation and maintenance time series data after embedding position encoding; is a learnable gating parameter; is long-term sequential information; is short-term sequential information; For self-attention; 、 is the weight matrix; 、 is the bias term; for Activation function.
[0043] Preferably, the expression for mapping the fused feature tensor into the prediction result in step S4 is:
[0044] ;
[0045] Where, To predict the results; Represents a dimensionality reduction operation; Represents a reshape operation; is the fused feature tensor; is the weight matrix; is the position deviation.
[0046] Preferably, after obtaining the prediction result in step S4, the prediction loss is calculated using the mean square error, and the prediction loss is optimized using the Adam algorithm. The expression of the optimization process is:
[0047] ;
[0048] ;
[0049] In the above formula, The gradient of the optimization for round t; Indicates finding the gradient; is the loss function; 、 are the estimates of the first and second moments of the gradient, respectively; 、 For control and exponential decay rate; Represents matrix multiplication; is the bias-corrected first-order moment estimate; yes t to the power of Bias-corrected second-order moment estimates; yes t to the power of are the model parameters at training step t; It is a hyperparameter that controls the step size of parameter update; A small value added to prevent the denominator from being zero; is the mean square error loss function; is the number of variables; is the sequence length of the prediction results; For the The variable in The predicted value of the step length; For the The variable in The true value of the step size.
[0050] The benefits of the present invention include at least:
[0051] 1. By introducing global attention to weight the self-attention values and using a cumulative summation operation to calculate long-term sequential information, the model captures the long-term cumulative effects and trends in the operation and maintenance data. Short-term sequential information is obtained by combining the self-attention values and the values of their neighboring historical time steps to focus on short-term fluctuations and local patterns in the data. A gating network is designed to dynamically fuse long-term sequential information, short-term sequential information, and the original self-attention values, enabling the model to adaptively integrate features of different time scales based on data characteristics and enhance the effect of information integration. This strategy of combining self-attention with long-term and short-term sequential information enables the model to analyze data from multiple perspectives and generate richer and more temporally dynamic feature representations.
[0052] 2. By sharding long operational time series data and mapping it into high-dimensional embedding representations, it can effectively decompose long sequences into several shorter subsequences, helping to reduce the complexity of subsequent self-attention calculations and, to a certain extent, alleviate the excessive memory consumption that may occur when directly processing extremely long sequences. By dividing data into units that are easier for the model to process, it improves overall processing efficiency and the scalability of the model for long sequences.
[0053] 3. By dynamically fusing long-term and short-term sequential information and self-attention to generate feature tensors and map them into prediction results, it can dynamically update features at each time step and capture data changes in a timely manner. Compared with static feature extraction methods, it can better adapt to changes that may occur at any time in long-term operation and maintenance time series data. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 Schematic diagram of a method flow in an embodiment of the present invention;
[0055] Figure 2 Schematic diagram of the system structure of an embodiment of the present invention. DETAILED DESCRIPTION
[0056] The following is a clear and complete description of the technical solutions in the embodiments of the present invention, in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts are within the scope of protection of the present invention.
[0057] like Figure 1 As shown, an embodiment of the present invention provides a method for predicting operation and maintenance time series data based on sequential information perception, comprising the following steps:
[0058] Step S1: Slice the long operation and maintenance time series data, and map the sliced long operation and maintenance time series data into a high-dimensional embedding representation through linear transformation and position encoding.
[0059] Specifically, historical operation and maintenance time series data are collected, a long-term operation and maintenance time series data set is constructed, the long-term operation and maintenance time series data set is divided into data sets, and the initial time series data set is Perform normalization preprocessing to obtain time series operation and maintenance data ,in is the look-back window length, is the number of variables.
[0060] Transpose the normalized operation and maintenance time series data to obtain . Use patching technology to Perform slicing and cut the operation and maintenance time series data into , the step length is The fragments obtained by segmentation Reshape the fragments into a 2D tensor for further processing:
[0061] ;
[0062] ;
[0063] ;
[0064] ;
[0065] In the above formula, This is the operation and maintenance time series data after block division; Represents a sharding operation; represents a reshape operation; Indicates the expansion operation; Represents a filling operation; is the weight matrix; This is the operation and maintenance time series data after embedding position encoding; Encode for position; is the position index; is the dimension index; is the embedding dimension.
[0066] Step S2: Perform a linear transformation on the high-dimensional embedding representation to obtain the query matrix, key matrix, value matrix, and global attention representing the weight of each time position's contribution to the entire sequence. Calculate the self-attention value of each time position based on the query matrix, key matrix, and value matrix.
[0067] Step S3: Perform element-wise multiplication of the global attention and self-attention values, accumulate all items before each time position in the product of the element-wise multiplication, and obtain long-term sequential information; perform weighted combination of the self-attention values of the current time position and the two time positions before the current time position to obtain short-term sequential information.
[0068] Specifically, the weight of each time position relative to the entire sequence is calculated and multiplied by the corresponding long-range dependency, thereby accumulating and summing the dynamic changes of the previous time steps to perceive the sequential information. By taking advantage of the attention mechanism and the accumulation and operation to maintain the global view of the sequence while capturing the cumulative effect of time changes, the complexity and inefficiency of the traditional RNN architecture are avoided. Specifically, by embedding the high-dimensional representation Perform linear transformation and learnable scaling and translation to map to the query matrix , key matrix , value matrix and global attention ,in Used to represent the contribution weight of each time series position to the entire sequence, 、 、 and The calculation expression is:
[0069] ;
[0070] ;
[0071] ;
[0072] ;
[0073] ;
[0074] ;
[0075] In the above formula, is the weight matrix; , is the scaling factor; , is the bias term; for Activation function; is the activation function, which is performed on the last dimension; To generate queries and key Basic characteristics of To generate values and global attention weight basic features.
[0076] Calculate the self-attention value based on the above results :
[0077] .
[0078] The self-attention mechanism can calculate the association between the feature representation of each position and the feature representation of all other positions and generate corresponding attention weights.
[0079] Based on self-attention value and global attention , two different dynamic sequence information are obtained through long-term cumulative sum and short-term cumulative sum respectively:
[0080] ;
[0081] ;
[0082] In the above formula, is long-term sequential information; Represents the state of each time step accumulated and all previous time steps; It is short-term sequential information, which means that only the information of the previous two time steps is accumulated for each time step; is an all-zero vector; Indicates that all vectors are merged in the last dimension by multiplying Scaling is done to keep the values within a more stable range.
[0083] Step S4: Dynamically fuse long-term and short-term sequential information and self-attention, add the fused features to the high-dimensional embedding representation to obtain the fused feature tensor, and map the fused feature tensor to the prediction result.
[0084] Specifically, in order to control the fusion of long-term and short-term sequential information and the original self-attention calculation results, a gating network is designed, and residual connections and layer normalization are applied to enhance the training stability and expressiveness of the model:
[0085] ;
[0086] ;
[0087] Where, is the fused feature tensor; Represents normalization operation; is a learnable gating parameter; for Activation function; 、 is the weight matrix; 、 is the bias term.
[0088] After layer normalization and residual connection, it can be sent to the next layer for processing. An additional reshaping operation is performed when outputting the prediction results:
[0089] ;
[0090] Where, To predict the results; Represents a dimensionality reduction operation; Represents a reshape operation; is the fused feature tensor; is the weight matrix; is the position deviation; is the sequence length of the prediction results.
[0091] The mean square error (MSE) is used to calculate the model loss, and the Adam algorithm is used to iteratively optimize the model. The expression is:
[0092] ;
[0093] Where, is the number of variables; For the The variable in The predicted value of the step length; For the The variable in The true value of the step size.
[0094] The expression for iterative optimization of the model using the Adam algorithm is:
[0095] ;
[0096] Where, The gradient of the optimization for round t; Indicates finding the gradient; is the loss function; 、 are the estimates of the first and second moments of the gradient, respectively; 、 For control and The exponential decay rate is set to 0.9 and 0.999 in the embodiment of the present invention; Represents matrix multiplication; is the bias-corrected first-order moment estimate; yes t to the power of Bias-corrected second-order moment estimates; yes t to the power of are the model parameters at training step t; It is a hyperparameter that controls the step size of parameter update; A very small value is added to prevent the denominator from being zero, and is set to 1e-8 in the embodiment of the present invention.
[0097] like Figure 2 As shown, an embodiment of the present invention also provides a prediction system for operation and maintenance time series data based on sequential information perception, which is implemented based on the above-mentioned prediction method for operation and maintenance time series data based on sequential information perception, and includes: a data preprocessing module, a multivariate time series embedding module, a sequential information perception module and a prediction output module.
[0098] Data preprocessing module: used to collect historical time series operation and maintenance data, build long-term time series data sets, and divide and normalize the data sets to eliminate differences between operation and maintenance indicators of different dimensions and improve the stability of model training.
[0099] Multivariate Time Series Embedding Module: This module is responsible for transposing the normalized operation and maintenance data, implementing data sharding through patching technology, and combining position encoding to generate an embedded representation containing time series position information.
[0100] Sequential Information Perception Module: Utilizes the self-attention mechanism to establish long-range dependencies between time points within a sequence and weights them through global attention. On this basis, the module captures global evolution trends through long-term accumulation and focuses on local dynamic changes through short-term accumulation and context aggregation. By integrating the three mechanisms of self-attention, long-term sequential information, and short-term sequential information, the module is able to capture features at different time scales, thereby achieving a comprehensive understanding of the dynamic evolution of operation and maintenance data. At the same time, a gating mechanism is used to adaptively fuse long-term and short-term sequential information with the original self-attention calculation results. Residual connections and layer normalization techniques are used to enhance the model's training stability and expressiveness, generating a fused feature representation with rich temporal semantics.
[0101] Prediction output module: reshapes the fused features, generates the target prediction sequence through fully connected layer mapping, and efficiently updates the model parameters based on the mean square error loss function and Adam optimization algorithm to output the final operation and maintenance data prediction results.
[0102] The present invention is verified to be effective through the following experiments:
[0103] 1. Data Description
[0104] To verify the effectiveness of the method in this embodiment of the present invention for timing prediction, a statistical dataset of network traffic time series from a real commercial LTE network was used. This dataset captures the Quality of Service (QoS) and Quality of Experience (QoE) characteristics of three commercial user equipments (UEs) interacting with three edge applications. Specifically, it includes the throughput and jitter of each UE-application, as well as the channel quality indicator (CQI) for each UE. The throughput is measured in bytes, and the jitter is measured in seconds.
[0105] This dataset is stored in a CSV file format, recording operational monitoring metrics once per second. Each row of data represents the value of each metric at a specific timestamp. The dataset columns include: the date column, which represents the timestamp accurate to the second; columns such as UE1: web-rtc, UE1: sipp, and UE1: web-server, which represent the throughput data for a specific UE interacting with different applications, such as web-rtc, sipp, and web-server, respectively. Generally, higher throughput values indicate higher data transmission efficiency. Columns such as UE1-Jitter, UE2-Jitter, and UE3-Jitter represent the jitter experienced by each UE. Jitter is the variation in packet arrival delay; lower values indicate better network stability. Columns such as UE1-CQI, UE2-CQI, and UE3-CQI represent the channel quality indicator for each UE, indicating the quality of the wireless channel. Generally, higher values indicate better channel quality and support for higher data rates. Table 1 shows a partial data structure.
[0106] Table 1 Data structure table
[0107]
[0108] In the above examples, some throughputs are zero, indicating that there is no traffic or no data generated in the corresponding scenario at that moment. After the data is loaded, in order to effectively utilize the data set for model training and evaluation, the embodiment of the present invention first performs a series of preprocessing and partitioning operations on the data. In terms of feature selection, the embodiment of the present invention can process univariate time series data, that is, by selecting a single operation and maintenance monitoring indicator specified by the target parameter, such as UE1: web-rtc; it can also process multivariate time series data, that is, considering all features except the date column. After the data is loaded, the target column will be moved to the last column, and the remaining columns will be used as input features.
[0109] The dataset is then strictly divided into training, validation, and test sets in chronological order to ensure fair model evaluation and prevent future data leakage into the training process. The specific division ratio is as follows: the training set accounts for 70% of the total data volume and is used for model learning and parameter optimization; the test set accounts for 20% of the total data volume and is used to evaluate the model's generalization ability and final performance on unseen data; the remaining 10% of the data constitutes the validation set, which is used during training to adjust model hyperparameters and perform early stopping, effectively preventing overfitting.
[0110] Finally, for time series prediction tasks, this embodiment of the present invention uses a sliding window approach to construct input-output sequence pairs. Through the detailed data processing and partitioning steps described above, this embodiment of the present invention can prepare a high-quality, structured dataset, laying a solid foundation for the subsequent training and evaluation of deep learning time series prediction models.
[0111] 2. Description of Benchmark Methodology
[0112] To verify the effectiveness of the model, this embodiment of the present invention selected a series of state-of-the-art time series prediction models as baselines for comparison, including Transformer-based models: iTransformer (ICLR 2024), PatchTST (ICLR 2023), Autoformer (NeurIPS 2021), and non-Transformer models: TimesNet (ICLR2023), DLinear (AAAI 2023), and SCINet (NeurIPS 2022).
[0113] 3. Experimental Setup
[0114] To fairly compare the model in this embodiment with the baseline model, all experiments were conducted on two NVIDIA Tesla T4 16GB GPUs and implemented using the PyTorch framework. Adam was selected as the optimizer, with 10 training epochs, a batch size of 32, and an early stopping mechanism of 3. The sequential information perception module in this embodiment consists of four layers, with all hidden layers set to 512 dimensions and a dropout value of 0.1.
[0115] 4. Results of Algorithm Performance Comparison
[0116] To systematically evaluate the effectiveness of the methods described in this embodiment of the present invention in the field of operational time series data forecasting, we designed a comparative experiment. We uniformly set the input sequence length to 96 time steps and conducted comprehensive tests for prediction lengths of 96, 192, 336, and 720 time steps. The experiments used mean squared error (MSE) and mean absolute error (MAE) as evaluation metrics. The prediction results are shown in Table 2.
[0117] Table 2 Comparison of prediction results
[0118]
[0119] The experimental results in Table 1 show that the time series prediction method proposed in this embodiment of the present invention achieves optimal performance in all test scenarios. Specifically, compared to the state-of-the-art iTransformer model, the method in this embodiment of the present invention achieves an average MSE reduction of 14.1%, demonstrating its significant advantage in prediction accuracy. This outstanding performance is primarily attributed to the sequential information-aware attention mechanism designed in this embodiment of the present invention, which effectively captures long-term correlations in time series data and accurately perceives sequential dependencies in time series data through cumulative and dynamic time step information. Notably, the global attention mechanism introduced in this embodiment of the present invention not only effectively captures long-term temporal dependencies but also successfully avoids technical challenges such as computational inefficiency and vanishing gradients that are common in traditional recurrent neural network (RNN) architectures. Comprehensive experimental results demonstrate that the method in this embodiment of the present invention exhibits remarkable robustness and adaptability in time series modeling capabilities, particularly in the complex application scenario of network operation and maintenance, providing reliable technical support for practical applications in related fields.
[0120] In summary, this paper addresses the shortcomings of existing network operation and maintenance time series data prediction technologies and proposes a method and system for network operation and maintenance time series prediction with enhanced sequential information perception. This method improves the Transformer architecture and designs an innovative sequential information perception enhancement strategy. This strategy aims to simultaneously capture the relative importance of each time point in the operation and maintenance data sequence within the overall sequence and its dynamic changes over time through an optimized attention mechanism, effectively constructing long-term dependencies and thus enhancing the Transformer's sequential understanding and long-term prediction capabilities when processing operation and maintenance time series data.
[0121] The Transformer, enhanced with sequential information awareness, significantly improves its ability to model long-range dependencies in operational data, capturing both long-term and short-term sequential information while avoiding the complexity and inefficiency of RNN-like architectures in capturing sequential information. The improved attention mechanism effectively captures the global view, overcoming the limitations of the traditional Transformer and adapting to diverse prediction needs. The designed gating network aggregates time-step information, requiring only a single attention head and eliminating the need for a feedforward layer. This maintains efficient parallel computation and adapts to the real-time demands of large-scale operational environments.
[0122] The method and system of the present invention are particularly suitable for time series data prediction in complex operation and maintenance environments such as IT infrastructure, network equipment, cloud computing platforms and industrial equipment. They can achieve high-precision prediction of key operation and maintenance indicators such as system performance, resource utilization, and load changes, providing reliable support for operation and maintenance decision-making and abnormal warning.
[0123] The technical features of the above embodiments may be combined in any manner. To simplify the description, not all possible combinations of the technical features in the above embodiments are described. Only preferred embodiments of the present invention are presented. While the description is relatively specific and detailed, it should not be construed as limiting the scope of the present invention. As long as there are no conflicts in the combination of these technical features, they should be considered to be within the scope of this specification.
[0124] It should be noted that, for those skilled in the art, various modifications and improvements can be made without departing from the scope of the present invention, and these modifications and improvements fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.
Claims
1. A prediction method for operation and maintenance time series data based on sequential information perception, characterized in that: The following steps are involved: Step S1: Slice the long operation and maintenance time series data and map the sliced long operation and maintenance time series data into a high-dimensional embedding representation through linear transformation and position encoding; Step S2: Perform a linear transformation on the high-dimensional embedding representation to obtain a query matrix, a key matrix, a value matrix, and a global attention matrix representing the weight of each time position's contribution to the entire sequence. Calculate the self-attention value of each time position based on the query matrix, key matrix, and value matrix. Step S3: performing element-wise multiplication of the global attention and self-attention values, accumulating all items before each time position in the product of the element-wise multiplication to obtain long-term sequence information; performing weighted combination of the self-attention values at the current time position and the two time positions before the current time position to obtain short-term sequence information; Step S4: Dynamically fuse long-term and short-term sequential information and self-attention, add the fused features to the high-dimensional embedding representation to obtain the fused feature tensor, and map the fused feature tensor to the prediction result.
2. The method for predicting operation and maintenance time series data based on sequential information perception according to claim 1, characterized in that: The sharding of the long operation and maintenance time series data in step S1 includes the following steps: Step S11: normalize the long operation and maintenance time series data, and transpose the normalized long operation and maintenance time series data into ; Step S12: Set the length to , the step length is ,right Fragment the data obtained after sharding. Reshape the fragments into a 2D tensor , the two-dimensional tensor The expression is: ; Where, Indicates sharding operation; Represents a reshape operation; Indicates the expansion operation; Represents a fill operation.
3. The method for predicting operation and maintenance time series data based on sequential information perception according to claim 2, characterized in that: The expression of the high-dimensional embedding representation in step S1 is: ; ; ; In the above formula, Long operation and maintenance time series data after embedding position encoding; is the weight matrix; Encode for position; is the position index; is the dimension index; is the embedding dimension.
4. The method for predicting operation and maintenance time series data based on sequential information perception according to claim 1 is characterized by: The expression of the self-attention value in step S2 is: ; ; ; ; ; ; In the above formula, For self-attention; is the activation function; is the query matrix; is the bond matrix; is the value matrix; is the embedding dimension; To generate queries and key Basic characteristics of To generate values and global attention weight Basic characteristics of Represents matrix dot product; for Activation function; is the scaling factor; 、 is the bias term; It is the long operation and maintenance time series data after embedding position encoding.
5. The method for predicting operation and maintenance time series data based on sequential information perception according to claim 1 is characterized by: The expression of the global attention in step S2 is: ; ; In the above formula, For global attention; is the activation function; To generate values and global attention weight Basic characteristics of is the scaling factor; is the bias term; Represents matrix dot product; for Activation function; is the weight matrix; is the bias term.
6. The method for predicting operation and maintenance time series data based on sequential information perception according to claim 1, characterized in that: The expression of the long-term sequence information in step S3 is: ; Where, is long-term sequential information; Indicates accumulation and operation; For self-attention; Represents matrix dot product; For global attention.
7. The method for predicting operation and maintenance time series data based on sequential information perception according to claim 1, characterized in that: The expression of the short-term sequence information in step S3 is: ; Where, is short-term sequential information; For self-attention; represents an all-zero vector; Indicates that all vectors are merged in the last dimension.
8. The method for predicting operation and maintenance time series data based on sequential information perception according to claim 1, characterized in that: The expression of the fused feature tensor in step S4 is: ; ; Where, is the fused feature tensor; Represents normalization operation; Long operation and maintenance time series data after embedding position encoding; is a learnable gating parameter; is long-term sequential information; is short-term sequential information; For self-attention; 、 is the weight matrix; 、 is the bias term; for Activation function.
9. The method for predicting operation and maintenance time series data based on sequential information perception according to claim 1, characterized in that: The expression for mapping the fused feature tensor to the prediction result in step S4 is: ; Where, To predict the results; Represents a dimensionality reduction operation; Represents a reshape operation; is the fused feature tensor; is the weight matrix; is the position deviation.
10. The method for predicting operation and maintenance time series data based on sequential information perception according to claim 1, characterized in that: After the prediction result is obtained in step S4, the prediction loss is calculated using the mean square error, and the prediction loss is optimized using the Adam algorithm. The expression of the optimization process is: ; ; In the above formula, The gradient of the optimization for round t; Indicates finding the gradient; is the loss function; 、 are the estimates of the first and second moments of the gradient, respectively; 、 For control and exponential decay rate; Represents matrix multiplication; is the bias-corrected first-order moment estimate; yes t to the power of Bias-corrected second-order moment estimates; yes t to the power of are the model parameters at training step t; It is a hyperparameter that controls the step size of parameter update; A small value added to prevent the denominator from being zero; is the mean square error loss function; is the number of variables; is the sequence length of the prediction results; For the The variable in The predicted value of the step length; For the The variable in The true value of the step size.