Robust time sequence mask and reconstruction method based on information bottleneck and optimal transmission
By generating a mask sequence through an adaptive detector based on information bottleneck and using the optimal transmission strategy for interpolation and reconstruction, the problem of irregular subsequences interfering with time series prediction in the existing technology is solved, and higher prediction accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510718660.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-09-16
AI Technical Summary
Existing robust time series forecasting methods have difficulty in effectively handling irregular subsequences, resulting in significant forecasting errors. In addition, existing data cleaning strategies fail to fully consider the specific requirements of the forecasting task and are unable to accurately capture irregular subsequences that are unfavorable to the forecasting task.
An adaptive detector based on information bottleneck is used to generate mask sequences, which are then interpolated and reconstructed using the optimal transmission strategy. The parameters of the adaptive detector are optimized through mask loss and prediction loss to generate mask sequences that shield irregular subsequences, and the optimal transmission strategy is used to reduce the reconstruction cost.
It improves the prediction accuracy of time series data, can effectively resist the influence of irregular subsequences, reduces reconstruction costs, and improves the robustness of the model.
Smart Images

Figure CN120653957A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of time series prediction, and more particularly to a robust time series masking and reconstruction method based on information bottleneck and optimal transmission. Background Art
[0002] Real-world time series data often contains irregular subsequences that deviate from the regular pattern of the entire series, posing a challenge to time series forecasting. Irregular subsequences can be caused by a variety of reasons, such as sensor failure, transmission interference, or malicious attacks. These irregular subsequences can interfere with the model's ability to correctly interpret the time series pattern, leading to significant forecasting errors.
[0003] However, most existing robust time series forecasting methods focus solely on addressing point anomalies or distribution shifts. They address point anomalies by using robust loss functions and sample selection strategies tailored to specific types of anomalies, or by incorporating a self-adaptation phase before prediction to address distribution shifts. However, compared to point anomalies, irregular subsequences are more complex because they exhibit diverse lengths or patterns, and they may share the same distribution as normal sequences. Therefore, existing methods struggle to withstand the interference caused by irregular subsequences.
[0004] Another intuitive strategy is to perform additional data cleaning before prediction to filter out irregular subsequences in the data. However, this strategy separates the data cleaning and prediction tasks into independent models, failing to fully consider the specific requirements of the prediction task during data cleaning. Consequently, it is difficult to accurately capture irregular subsequences that are detrimental to the prediction task and may introduce additional noise into the prediction. Therefore, addressing the prediction challenge of irregular subsequence data is crucial. Summary of the Invention
[0005] In view of this, the present invention provides a robust timing masking and reconstruction method based on information bottleneck and optimal transmission, which is used to at least solve some of the technical problems in the background technology.
[0006] In order to achieve the above object, the present invention adopts the following technical solutions:
[0007] A robust timing masking and reconstruction method based on information bottleneck and optimal transmission includes the following steps:
[0008] Obtain time series data and perform normalization;
[0009] The normalized time series data is detected using an adaptive detector based on the information bottleneck principle to generate a mask sequence that can shield irregular subsequences.
[0010] The mask sequence is interpolated and reconstructed using an optimal transmission strategy to generate an interpolated reconstructed time series sequence.
[0011] Furthermore, the above method also includes performing time series prediction using the reconstructed time series sequence after interpolation.
[0012] Furthermore, the time series data includes load time series data of the power transformer.
[0013] Furthermore, an adaptive detector based on the information bottleneck principle is used to detect the normalized time series data to generate an optimal mask sequence that can shield irregular subsequences. Specifically, the following steps are performed:
[0014] Map the normalized time series x to the latent variable matrix z through linear transformation;
[0015] Use the self-attention mechanism to process the latent variable matrix z to obtain the self-attention matrix A and the hidden feature matrix E;
[0016] Calculate the similarity between the hidden variable z and the hidden feature matrix E through the cross attention mechanism;
[0017] The similarity results are transformed linearly to generate the mask probability matrix λ of the irregular pattern;
[0018] Generate the mask matrix M according to the mask probability matrix λ and use the Gumbel-Softmax algorithm;
[0019] Generate mask sequence x based on mask matrix M m .
[0020] Furthermore, the similarity between the latent variable z and the hidden feature matrix E is calculated through the cross attention mechanism, specifically including:
[0021] The hidden feature matrix E is used as the query matrix Q′ in the cross attention mechanism, and the latent variable matrix z is used as the key matrix K′ and value matrix V′ in the cross attention mechanism;
[0022] The similarity score between the query matrix Q′ and the key matrix K′ is calculated by dot product:
[0023]
[0024] Where: T represents the matrix transpose symbol; represents the scaling factor; d k represents the dimension of the key matrix K'; softmax represents the softmax function used to normalize the similarity scores into attention weights.
[0025] Furthermore, a mask sequence x is generated based on the mask matrix M m, specifically including:
[0026] Multiply the mask matrix M by the normalized time series x bit by bit to obtain the mask sequence x m .
[0027] Furthermore, the mask sequence is interpolated and reconstructed using the optimal transmission strategy, specifically including:
[0028] The mask sequence x is obtained by connecting the Transformer encoder and the linear transformation module. m Perform preliminary transformation to obtain the transformation sequence
[0029] Perform a linear transformation on the self-attention matrix A generated by the self-attention mechanism to generate the transmission strategy matrix P;
[0030] Apply the transmission strategy matrix P to the transformed sequence The reconstructed mask sequence after interpolation is obtained, and the interpolation process is optimized using the optimal transmission loss. The reconstructed mask sequence after interpolation is obtained by the following expression:
[0031]
[0032] Where x' represents the reconstructed mask sequence after interpolation.
[0033] Furthermore, it also includes optimizing the adaptive detector based on the information bottleneck principle by using mask loss and prediction loss.
[0034] Furthermore, the mask loss is used to optimize the adaptive detector based on the information bottleneck principle, specifically including:
[0035] The mask loss is formulated to optimize and minimize the mutual information between the mask sequence and the input time series data, where the mask loss specifically includes the following function:
[0036]
[0037] λ is the mask probability matrix (mentioned above); τ is a hyperparameter; L is the sequence length of the input time series data.
[0038] Furthermore, the prediction loss is used to optimize the adaptive detector based on the information bottleneck principle, specifically including:
[0039] MSE is formulated as the prediction loss for predicting the mask to optimize and maximize the mutual information between the mask sequence and the true value of the predicted sequence.
[0040] It can be seen from the above technical solutions that, compared with the prior art, the present invention discloses a robust timing masking and reconstruction method based on information bottleneck and optimal transmission, which has the following beneficial effects:
[0041] The present invention establishes an adaptive detector based on information bottleneck, takes the original sequence as input, learns to generate a mask matrix to mask the original sequence, and optimizes the parameters of the adaptive detector through mask loss and prediction loss, so as to generate a mask sequence that shields irregular subsequences, thereby improving the accuracy of subsequent applications of time series data such as load time series data.
[0042] The present invention utilizes an optimal transmission strategy to reconstruct the mask sequence, thereby reducing reconstruction costs.
[0043] The present invention utilizes the generated mask sequence for prediction, which can resist the influence of irregular subsequences and improve the accuracy of prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0045] Figure 1 The overall flow chart of the method provided for the embodiment of the invention.
[0046] Figure 2 A schematic diagram of the overall architecture of the method provided in an embodiment of the present invention.
[0047] Figure 3 A schematic diagram of the optimal transmission strategy provided by an embodiment of the present invention.
[0048] Figure 4 This is a schematic diagram comparing the effects of the methods provided in the embodiments of the present invention. DETAILED DESCRIPTION
[0049] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0050] The technical solution proposed by the present invention is shown in the figure above. The present invention adopts a detection-interpolation-prediction workflow. First, an information bottleneck-based adaptive detector (IB-based Detector) is used to identify irregular subsequences in the time series, and then the time series is masked according to the detection results. Next, the masked sequence is interpolated using the optimal transmission-based reconstruction module (OTReconstruction Module) to obtain a reconstructed sequence. Finally, the reconstructed sequence is input into the prediction module (Prediction Module) to generate the final prediction result. The entire workflow integrates the irregular subsequence detection, interpolation, and prediction tasks into a unified optimization goal.
[0051] The robust timing masking and reconstruction method based on information bottleneck and optimal transmission disclosed in an embodiment of the present invention includes the following steps:
[0052] Obtain time series data and perform normalization;
[0053] The normalized time series data is detected using an adaptive detector based on the information bottleneck principle to generate a mask sequence that can shield irregular subsequences.
[0054] The mask sequence is interpolated and reconstructed using an optimal transmission strategy to generate an interpolated reconstructed time series sequence.
[0055] Reference below Figure 1 、 Figure 2 、 Figure 3 The inventive principle and specific implementation steps of the present invention are further explained.
[0056] First, raw time series data is acquired and normalized. Time series data in this paper refers to related data that can form a time series, including load time series data of power transformers. These raw time series data often contain irregular subsequences. These irregular subsequences exhibit different lengths and patterns and may share the same distribution as regular sequences, making them difficult to detect directly. These irregular subsequences affect the analysis of complex patterns and dependencies in time series, thereby limiting subsequent applications and related performance of time series data, such as prediction.
[0057] To eliminate the impact of irregular subsequences on prediction, we propose an information bottleneck based detector to accurately detect these irregular subsequences.
[0058] Specifically, we need to obtain an optimal mask sequence (intermediate representation) that contains as little information as possible about the original input x, while retaining the relevant information needed to predict the future sequence y. Based on this concept, we set up a learnable mask matrix M, trying to find the optimal mask sequence representation x*m by masking irrelevant or irregular subsequences in the original sequence x:
[0059]
[0060] Among them, α is a hyperparameter used to balance the two terms.
[0061] In practice, we first map the input sequence x to a latent variable Z through a linear transformation. Then, z is fed into the self-attention mechanism to capture long-term dependencies, resulting in the attention matrix A and the generated features E. z is a matrix of size [sequence length × hidden dimension], and x is obtained by a linear transformation (mapping 1 to the hidden dimension).
[0062]
[0063] Next, the similarity between the hidden feature E and the input Z is calculated through cross attention, where the hidden feature E is used as the query and the input Z is used as the key and value of the cross attention. The result is then transformed into a mask probability matrix λ of the irregular pattern through a linear transformation.
[0064]
[0065] To generate the one-hot mask matrix from the mask probability matrix λ, we apply the Gumbel-Softmax algorithm to generate the mask matrix M.
[0066] M = GumbelSoftmax(λ).
[0067] Applying the mask matrix M to the original sequence x can generate the mask sequence x m Since the mask matrix is one-hot encoded, the places where the value is 0 mean they are masked, and the places where the value is not 0 mean they are retained. Therefore, the mask matrix M is multiplied bit by bit by the time sequence x to obtain the mask sequence x. m .
[0068] The mask loss is formulated to optimize the minimization of the mutual information between the mask matrix (intermediate representation) and the input sequence:
[0069]
[0070] λ i is the probability that the i-th position in the time series is masked, τ is a hyperparameter; L is the sequence length of the input time series data.
[0071] The mask loss function includes λ i , which is the mask probability, can be obtained from the mask probability matrix. By optimizing the mask loss function to increase the mask probability, more positions can be masked, thereby indirectly minimizing the mutual information between the mask sequence and the input time sequence.
[0072] The first item is used to achieve as much masking as possible, and the second item is a continuity item, which is used to maintain the continuity of the mask so that it can completely mask the entire irregular subsequence.
[0073] The purpose of the above steps is to mask out (set to 0) the regions of the irregular subsequence and retain only the normal regions. In other words, the final masked sequence does not contain the masked subsequence. To achieve this, we establish an adaptive detector that uses the original sequence as input to learn to generate a mask matrix to mask the original sequence. The parameters of the adaptive detector are optimized through masking loss and prediction loss to accurately mask out the irregular subsequence regions. The masking loss is used to achieve as much masking as possible and maintain the continuity of the mask, so that it can completely mask out the entire irregular subsequence; the prediction loss is used to optimize the model to retain prediction-related information in the mask sequence, thus preventing normal regions from being masked.
[0074] After obtaining the mask sequence, the optimal transmission strategy is used to interpolate the mask sequence, and finally input it into the predictor to generate the predicted value In the present invention, the predictor can adopt the PatchTST model structure, but can also be adapted to the Dlinear and iTransformer structures, with the goal of ultimately generating a predicted value.
[0075] MSE is formulated as the prediction loss to optimize the mutual information between the intermediate representation and the true value of the predicted sequence:
[0076]
[0077] y is the real future sequence value, The predicted value output by the model.
[0078] We formulate the optimization process of mask sequence interpolation as an optimal transmission problem. Figure 3 As shown, first, in order to restore the continuity of the masked sequence, we perform a preliminary transformation on the masked sequence through a network combining the Transformer encoder and the linear head (one-time linear transformation module), and obtain Will The distribution of is taken as the source distribution, and the distribution of the original sequence x is taken as the target distribution. Our goal is to obtain a transmission strategy P that transforms the source distribution into the target distribution while making the corresponding migration cost as small as possible:
[0079]
[0080] Where C is the transfer cost. We hope that the masked area can be reconstructed back to a normal pattern rather than the original irregular pattern. Therefore, we set a larger transfer cost for transferring to the masked area:
[0081]
[0082] P is the transfer strategy. We perform a linear transformation on the attention matrix A generated by the self-attention mechanism to generate P. Apply P to The overall process of obtaining the interpolated sequence x' is as follows Figure 3 shown.
[0083]
[0084] Optimize the interpolation process by setting the optimal transmission loss:
[0085]
[0086] β is a hyperparameter used to balance the two terms.
[0087] The overall optimization loss consists of three parts: mask loss, prediction loss, and transmission loss:
[0088]
[0089] α is a hyperparameter, as mentioned above. By optimizing the model parameters through the overall loss, the three steps of detection, interpolation, and prediction are optimized end-to-end.
[0090] The present invention (ours) can make the model surpass the existing time series prediction methods under real data sets, using mean square error (MSE) and mean absolute error (MAE) as evaluation indicators.
[0091]
[0092] To further verify the robustness of the present invention against irregular subsequences in historical data, we artificially injected five types of irregular subsequences into the original data to simulate a scenario where the proportion of irregular subsequences is more significant. The present invention demonstrated a significant performance improvement in the presence of irregular subsequences, further demonstrating its effectiveness in combating the effects of irregular subsequences.
[0093]
[0094] Pre-preparation: time series dataset (with irregular subsequences).
[0095] Specific implementation steps:
[0096] 1. Normalize the time series data. (Input: original time series data Output: normalized time series data)
[0097] 2. Use adaptive detectors to generate mask probabilities and attention matrices from time series data. (Input: normalized time series data Output: mask matrix and attention matrix)
[0098] 3. Calculate the mask loss. (Input: mask probability Output: mask loss)
[0099] 4. Use Gumbel-Softmax to generate a mask matrix. (Input: mask probability Output: mask matrix)
[0100] 5. Use the mask matrix to mask the time series data. (Input: mask matrix and normalized time series data Output: masked time series)
[0101] 6. Input the mask sequence into the reconstruction module based on optimal transmission for interpolation. (Input: mask sequence and attention matrix Output: reconstructed sequence)
[0102] 7. Calculate the optimal transmission loss. (Input: mask timing, reconstruction timing. Output: optimal transmission loss.)
[0103] 8. Use the reconstructed time series to make predictions. (Input: reconstructed time series Output: predicted value)
[0104] 9. Calculate the prediction loss. (Input: predicted value and true value Output: prediction loss)
[0105] 10. Calculate the overall loss. (Input mask loss, optimal transfer loss and prediction loss output: overall loss)
[0106] 11. Use the overall loss to optimize the above process and use gradient descent to update the parameters.
[0107] Take the energy data application scenario as an example:
[0108] This scenario requires predicting the load value of the power transformer. Due to occasional sensor failures and other reasons, the signal will fluctuate, so the time series sampled historically will produce irregular subsequences. For example, Figure 4 Because the sensor stops working temporarily, the sampled data may remain unchanged for a short period of time (as shown by the black broken line in the first column of the figure). Sensor abnormalities may cause the collected sequence to contain spikes (as shown by the black broken line in the fourth column of the figure).
[0109] Conventional time series forecasting methods (the second row) are susceptible to interference from irregular subsequences within these historical sequences, resulting in poor prediction performance. Furthermore, because irregular subsequences are diverse in type, length, and randomness, conventional data cleaning methods struggle to accurately identify and remove them.
[0110] Our method (first row) can resist the interference of irregular subsequences by executing the "irregular subsequence detection -> interpolation -> prediction" process, achieving more accurate load value sequence prediction. Specifically, we normalize the original historical load value sequence (including irregular subsequences) and then input it into the model: First, an adaptive detector generates a mask matrix to mask the load value sequence, aiming to mask out the irregular subsequence regions; then, an optimal transmission reconstruction module is used to reconstruct the masked regions to obtain a normal load value sequence; finally, the reconstructed load value sequence is input into the predictor to obtain the predicted future load value. The model is trained and updated using the masking loss, optimal transmission loss, and prediction loss.
[0111] like Figure 4 As shown, the present invention can effectively detect irregular patterns (purple areas) in the original load value sequence and convert them into normal patterns, ultimately generating more accurate predictions. This invention effectively improves the model's robustness when dealing with this type of energy data. Visualization also reveals that the prediction curve (orange) of the present invention is closer to the true value (black), making it suitable for time series prediction of this type of energy data.
[0112] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.
[0113] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A robust timing masking and reconstruction method based on information bottleneck and optimal transmission, characterized in that: The following steps are involved: Obtain time series data and perform normalization; The normalized time series data is detected using an adaptive detector based on the information bottleneck principle to generate a mask sequence that can shield irregular subsequences. The mask sequence is interpolated and reconstructed using an optimal transmission strategy to generate an interpolated reconstructed time series sequence.
2. The robust timing masking and reconstruction method based on information bottleneck and optimal transmission according to claim 1, characterized in that: The method also includes performing time series prediction using the reconstructed time series sequence after interpolation.
3. The robust timing masking and reconstruction method based on information bottleneck and optimal transmission according to claim 1, characterized in that: The time series data includes load time series data of the power transformer.
4. The method of robust timing masking and reconstruction based on information bottleneck and optimal transmission according to claim 1, characterized in that: The normalized time series data is detected using an adaptive detector based on the information bottleneck principle to generate an optimal mask sequence that can shield irregular subsequences. Specifically, the following steps are performed: Map the normalized time series x to the latent variable matrix z through linear transformation; Use the self-attention mechanism to process the latent variable matrix z to obtain the self-attention matrix A and the hidden feature matrix E; Calculate the similarity between the hidden variable z and the hidden feature matrix E through the cross attention mechanism; The similarity results are transformed linearly to generate the mask probability matrix λ of the irregular pattern; Generate the mask matrix M according to the mask probability matrix λ and use the Gumbel-Softmax algorithm; Generate mask sequence x based on mask matrix M m .
5. The method of robust timing masking and reconstruction based on information bottleneck and optimal transmission according to claim 4, characterized in that: The similarity between the hidden variable z and the hidden feature matrix E is calculated through the cross attention mechanism, specifically including: The hidden feature matrix E is used as the query matrix Q′ in the cross attention mechanism, and the latent variable matrix z is used as the key matrix K′ and value matrix V′ in the cross attention mechanism; The similarity score between the query matrix Q′ and the key matrix K′ is calculated by dot product: Where: T represents the matrix transpose symbol; represents the scaling factor; d k represents the dimension of the key matrix K'; softmax represents the softmax function used to normalize the similarity scores into attention weights.
6. The method of robust timing masking and reconstruction based on information bottleneck and optimal transmission according to claim 4, characterized in that: Generate mask sequence x based on mask matrix M m , specifically including: Multiply the mask matrix M by the normalized time series x bit by bit to obtain the mask sequence x m .
7. The method of robust timing masking and reconstruction based on information bottleneck and optimal transmission according to claim 4, characterized in that: The mask sequence is interpolated and reconstructed using the optimal transmission strategy, including: The mask sequence x is obtained by connecting the Transformer encoder and the linear transformation module. m Perform preliminary transformation to obtain the transformation sequence Perform a linear transformation on the self-attention matrix A generated by the self-attention mechanism to generate the transmission strategy matrix P; Apply the transmission strategy matrix P to the transformed sequence The reconstructed mask sequence after interpolation is obtained, and the interpolation process is optimized using the optimal transmission loss. The reconstructed mask sequence after interpolation is obtained by the following expression: Where x' represents the reconstructed mask sequence after interpolation.
8. The method of robust timing masking and reconstruction based on information bottleneck and optimal transmission according to claim 1, characterized in that: It also includes optimizing the adaptive detector based on the information bottleneck principle using mask loss and prediction loss.
9. The method of robust timing masking and reconstruction based on information bottleneck and optimal transmission according to claim 8, characterized in that: The adaptive detector based on the information bottleneck principle is optimized using mask loss, specifically including: The mask loss is formulated to optimize and minimize the mutual information between the mask sequence and the input time series data, where the mask loss specifically includes the following function: λ i is the probability that the i-th position in the time series is masked; τ is a hyperparameter; L is the sequence length of the input time series data.
10. The method of robust timing masking and reconstruction based on information bottleneck and optimal transmission according to claim 8, characterized in that: The prediction loss is used to optimize the adaptive detector based on the information bottleneck principle, including: MSE is formulated as the prediction loss for predicting the mask to optimize and maximize the mutual information between the mask sequence and the true value of the predicted sequence.
Citation Information
Patent Citations
Mask time sequence model training method and device, mask time sequence model prediction method and device, equipment and medium
CN117556882A
Long-period multivariate time sequence prediction method based on multivariate information interaction
CN117786602A
Adaptive timestamp coding enhanced complex equipment incomplete state monitoring data time sequence interpolation method
CN119646555A
Forecasting in multivariate irregularly sampled time series with missing values
US20220058465A1