Industrial time series prediction method based on out-of-distribution representation learning
Through the quantization of timing similarity and the time series distribution external representation learning method DIVERSIFY combined with spatial attention, the prediction performance degradation of traditional methods under the timing drift phenomenon is solved, and the accurate prediction of industrial time series is achieved.
Patent Information
- Application Number
- CN202410012888.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-04
- Publication Date
- 2025-07-04
AI Technical Summary
Traditional time series prediction methods have limitations in dealing with timing drift phenomena, resulting in severe decline in prediction performance of industrial time series prediction models in long-term predictions.
The distribution characteristics of industrial time series are described using time-sequence similarity quantification technology, and the potential distribution of sequences is explored through the time convolution network and time-sequence distribution external representation learning method DIVERSIFY, and predictions are made in combination with spatial attention.
It effectively solves the problem of prediction performance degradation caused by timing drift phenomenon, and improves the accuracy and stability of industrial time series prediction.
Smart Images

Figure CN120257225A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of industrial time series prediction, and specifically relates to an industrial time series prediction method based on out-of-distribution representation learning. Background Art
[0002] With the rapid development of modern industrial information technology, a large amount of real-time data is generated in industrial processes. These data contain rich information related to production efficiency and product quality, such as the flame temperature in the kiln sintering process, the molten iron quality in the blast furnace ironmaking process, the heating temperature of the coal-fired boiler, etc. To a certain extent, these industrial time series data reveal the characteristics and laws of the actual production condition changes. Therefore, industrial time series modeling and prediction can effectively explore the laws of industrial production condition changes, which has also become an important part of improving production efficiency, promoting the stable development of industry, and driving industrial intelligence.
[0003] Traditional time series prediction methods have been widely applied to industrial time series prediction tasks, but they have certain limitations in dealing with the phenomenon of temporal drift. The temporal drift phenomenon mainly describes the phenomenon that the data distribution at future moments will show a large difference compared with historical data, resulting in a serious decline in the prediction performance of the prediction model established on historical data during the prediction stage. The existence of the temporal drift phenomenon causes traditional time series prediction methods to fail in long-term prediction because they cannot adapt to the change of data distribution. In industrial processes, due to the long duration, complex mechanism, and vulnerability to external environmental interference of industrial processes, the temporal drift phenomenon is often widespread and cannot be ignored. Summary of the Invention
[0004] The present invention provides an industrial time series prediction method based on out-of-distribution representation learning. This method first applies temporal similarity quantification technology to describe the distribution characteristics of industrial time series, thereby obtaining time series segments with different distribution characteristics. Subsequently, these segments are input into a temporal convolutional network, and the out-of-distribution representation learning method for time series (DIVERSIFY) is used to explore the potential distribution of the sequences. Finally, TCN is fused with spatial attention to obtain accurate prediction results.
[0005] The technical solution adopted by the present invention to achieve the above object is:
[0006] An industrial time series prediction method based on out-of-distribution representation learning, comprising: using a time series similarity quantification technique to describe the distribution characteristics of industrial time series to obtain time series segments with different distribution characteristics; inputting these segments into a temporal convolutional network (TCN), and exploring the potential distribution of the sequences with the help of a time series out-of-distribution representation learning method DIVERSIFY; combining TCN with spatial attention to obtain a final prediction model for predicting industrial time series data.
[0007] The method specifically includes the following steps:
[0008] Step 1: Collect industrial time series data and perform preprocessing;
[0009] Step 2: Input the industrial time series data into a temporal distribution similarity quantification module to obtain K most dissimilar subsequences and assign corresponding labels;
[0010] Step 3: Divide the obtained sample data set into a training data set I, a validation data set I, and a prediction data set I according to a ratio, and then perform a normalization operation;
[0011] Step 4: Establish an out-of-distribution representation learning model DIVERSIFY for time series, which models the internal potential distribution of time series with dynamic distributions and learns the temporal differences and potential distribution characteristics between different distributions;
[0012] Step 5: Use the training data set I to train DIVERSIFY to obtain corresponding weights and biases, bring the obtained weights and biases into the validation data set I, then calculate the prediction error of the validation data set I, and save the weights and biases that minimize the prediction error of the validation set;
[0013] Step 6: Perform input processing on the industrial time series obtained in Step 1: Combine the values of each variable at m historical moments and the predicted target value after t seconds to form a sample data; where m is the number of input historical time points and t is the prediction time point; and divide the sample data set into a training data set II, a validation data set II, and a prediction data set II according to a ratio, and then perform a normalization operation;
[0014] Step 7: Establish a prediction model based on spatial attention and temporal convolutional network. This model uses the feature extraction layer in the model DIVERSIFY trained in Steps 4 and 5 as the temporal convolutional network here, and at the same time uses spatial attention to mine the multi-variable coupling characteristics of industrial time series, and finally combines the two parts of features through a gated fusion mechanism to achieve prediction;
[0015] Step 8: Use the training dataset II to train the prediction model to obtain the corresponding weights and biases. Substitute the obtained weights and biases into the validation dataset II, then calculate the prediction error of the validation dataset II, and save the weights and biases that minimize the prediction error of the validation set. Select the AdamW optimizer, use the mean squared error as the loss function, and adjust the parameters of the prediction model through backpropagation to improve the prediction accuracy of the model;
[0016] Step 9: Substitute the weights and biases that minimize the prediction error of the validation set in the previous step into the prediction model, and then apply it to the industrial time series dataset to calculate the predicted values of the prediction model on the dataset.
[0017] The preprocessing includes the following steps:
[0018] Step 1-1: Remove outliers;
[0019] Step 1-2: Filter the industrial time series.
[0020] The time series distribution similarity quantification module can maximize the use of the information contained in the continuous time series by finding the time series periods that are least similar to each other, mainly including the following steps:
[0021] Average the time series into n parts, where each part is the smallest unit period;
[0022] Define the set {q1,..., q K}, given the K value, let it iterate from the initial value to the given value, and in each iteration round, further select each time series period of the input sequence based on the greedy strategy; the further selection of each time series period of the input sequence is as follows: between the start and end points of the time series, select 1 split point from the candidate split points to refine the data segment by maximizing the distribution distance, and then further select split points for the current smallest data segment until all K data segments are selected, which are used to represent the start and end points of the time subseries with the least similar distribution;
[0023] In the set, the K value that maximizes the average distribution distance of the K most dissimilar subsequences is the optimal K value, and output it.
[0024] The out-of-distribution representation learning model DIVERSIFY structure of the time series includes:
[0025] DIVERSIFY consists of a min-max adversarial game, aiming to characterize the latent distribution to cope with the distribution dynamics in time series data, mainly divided into the following three steps:
[0026] Step 4-1: Fine-grained feature update:
[0027] First, use the pseudo-domain class label as the label of the classifier to update the feature extractor, so that the feature extractor can better capture the feature information within the domain; the pseudo-domain class label regards each category of each domain as a new class, thereby attaching more fine-grained category and domain information, and its calculation formula is as follows:
[0028] s = d′ × C + y
[0029] where s ∈ {1, 2, …, S}, S = K × C, K is the predefined number of latent distributions, and C is the number of initial categories; d′ is the domain label, and in the first iteration, all samples are initialized, that is, d′ = 0;
[0030] Secondly, use the pseudo-domain class label for supervised learning, and the loss function is as follows:
[0031]
[0032] where, h f , respectively represent the feature extractor, the Bottleneck layer and the classifier layer of this part, and L represents the cross-entropy loss function;
[0033] Step 4-2, Latent distribution representation: By maximizing the difference between different latent distributions, identify the domain label of each sample to obtain latent distribution information and expand the diversity of data distribution;
[0034] First, obtain the centroid of each domain with intra-class features, and the calculation formula is as follows:
[0035]
[0036] where, respectively represent the Bottleneck layer and the classifier layer of this part. is the initial centroid of the k-th latent domain, and δ k is the k-th element of the softmax output;
[0037] Secondly, use the distance function D to obtain the pseudo-domain label through the nearest centroid classifier, and the calculation formula is as follows:
[0038]
[0039] Thirdly, calculate the centroid and obtain the updated pseudo-domain label:
[0040]
[0041]
[0042] Among them, Ι(a) is 1 only when a is true, otherwise it is 0; finally, the loss function of this step can be obtained:
[0043]
[0044] Among them is the discriminator of this part, which includes multiple linear layers and a classification layer; is the gradient reversal layer with hyperparameter λ1;
[0045] Step 4-3, Domain-invariant Representation Learning: By using the pseudo-domain labels in the previous step to learn domain-invariant representations, which are used for the model to learn general features to handle data in different sub-domains; specifically, adversarial training is directly used to update the classification loss and the domain classifier loss as follows:
[0046]
[0047] The establishment of the prediction model based on spatial attention and temporal convolutional network includes the following steps:
[0048] Step 7-1, Determine the model input dimension: batchsize×m×n features , where batchsize is the batch size set during batch training of the model, m is the number of input historical time points, and n features is the number of features at each moment of the input;
[0049] Step 7-2, According to the model input dimension in Step 4-1, determine the structural parameters of the prediction model.
[0050] The temporal convolutional module TCN effectively mines the temporal dynamic features and non-linear features of the variable to be predicted by using causal convolution.
[0051] The spatial attention module includes the following steps:
[0052] Step 7-3, Generate a spatial attention map by using the spatial relationship of features: Apply average pooling along the channel axis direction, which represents the average aggregated features of the entire channel; then, through a standard convolutional layer and residual connection, generate a two-dimensional spatial attention map, and the calculation is as follows:
[0053] SAM(x) = σ(f(x))·x + x
[0054] where f represents the convolution operation and σ is the activation function.
[0055] The training and calculation process of the prediction model includes the following steps:
[0056] Step 8-1: Determine the hyperparameters for predicting model training: learning rate lr, maximum number of iterations I MAX ; Randomly initialize the weight matrix w and bias β of each network layer, and set the initial number of iterations I = 0;
[0057] Step 8-2: Input the input two-dimensional matrix X = [X(t), …, X(t - m)] into the temporal convolutional network and the spatial attention module:
[0058] h T = TCN(X)
[0059] h S = SAM(X)
[0060] Step 8-3: Input the temporal features and spatial features of the latent distribution obtained in the above steps into the Gated Fusion module, and the calculation formula is as shown in the following formula:
[0061] h = h S ·σ(h S + h T ) + h T ·(1 - σ(h S + h T ))
[0062] Step 8-4: Obtain the fused feature h through the above steps, and finally use the fully connected layer to weight the fused feature to obtain the final output;
[0063] y = ReLU(w a h + β)
[0064] y represents the final predicted value of the model; w a represents the transformation matrix; β represents the bias value; ReLU is a non-linear activation function.
[0065] The industrial time series prediction device based on out-of-distribution representation learning includes a front-end interface and a background. The background is provided with a memory and a processor. A program is stored in the processor. When the processor loads the program, it executes the above method steps to obtain a trained and optimized prediction model, so that the model predicts the industrial time series data and obtains the predicted values of the corresponding industrial data, and visually displays them to the front-end interface of the industrial scenario for users.
[0066] The present invention has the following beneficial effects and advantages:
[0067] Industrial Time Series Prediction Method Based on Out-of-Distribution Representation Learning. To address the prevalent temporal drift phenomenon in industrial time series, this paper proposes a method called AdaTCN. This method focuses on out-of-distribution representation learning of industrial process time series. First, it applies temporal similarity quantification technology to describe the distribution characteristics of industrial time series, thereby obtaining time series segments with different distribution characteristics. Subsequently, these segments are input into a temporal convolutional network, and the out-of-distribution representation learning method for time series (DIVERSIFY) is used to explore the potential distribution of the sequence. Through this process, a TCN based on dynamic time series is successfully constructed, and the potential distribution of the time series is effectively learned. Finally, the TCN is fused with spatial attention to obtain accurate prediction results. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] Figure 1 Schematic diagram of an industrial example of the method of the present invention.
[0069] Figure 2 Schematic diagram of the method of the present invention.
[0070] Figure 3 Comparison chart of prediction errors of different models on the test set.
[0071] Figure 4 Comparison chart of prediction results between the method of the present invention and other models. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0072] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following detailed description of the specific implementation method of the present invention will be given in conjunction with the accompanying drawings. Many specific details are set forth in the following description in order to fully understand the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the connotation of the invention. Therefore, the present invention is not limited by the specific implementations disclosed below.
[0073] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art belonging to the technical field of the present invention. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention.
[0074] As Figure 1-2 shown, it is a schematic diagram of an industrial example and a schematic diagram of the method of the present invention.
[0075] Industrial Time Series Prediction Method Based on Out-of-Distribution Representation Learning. This method focuses on the out-of-distribution representation learning of industrial process time series. First, it applies a time series similarity quantification technique to describe the distribution characteristics of industrial time series, thereby obtaining time series segments with different distribution characteristics. Subsequently, these segments are input into a temporal convolutional network, and an out-of-distribution representation learning method for time series (DIVERSIFY) is used to explore the potential distribution of the sequence. Finally, TCN is fused with spatial attention to obtain accurate prediction results. The programming languages used for the program execution steps of the present invention are not limited to MATLAB, Python, etc.
[0076] The specific steps of the present invention are as follows:
[0077] Step 1: Collect industrial time series. Taking the actual loose rewetting industrial process as an example, the material is conveyed to the drum along with the conveyor belt, and the humidifying water is atomized by high-temperature steam in the drum, enabling the material to fully absorb moisture. The operator then adjusts the amount of humidifying water based on the real-time quantities of each sensor during this process to control the outlet moisture to meet the indicators. The key parameters in this process mainly include 24 variables such as the inlet moisture of the material at the inlet of the drum, the outlet moisture of the material at the outlet of the drum, the actual value of the added water involved in the process, the water valve opening, the actual outlet temperature, the set value of the added water, the actual temperature of bucket 1, the actual temperature of bucket 2, and the water addition coefficient (where the buckets are used to hold auxiliary materials, and the number of buckets is set according to the actual situation). The key parameter to be predicted is the outlet moisture value; and preprocess the above variable data collected.
[0078] Step 1-1: Remove outliers according to the set error range.
[0079] Step 1-2: Filter the industrial time series; for example, the wavelet threshold denoising method is used for the filtering method.
[0080] Step 2: Input the industrial time series data into the temporal distribution similarity quantification module to obtain K most dissimilar subsequences and assign corresponding labels.
[0081] Step 2-1: To effectively calculate and avoid trivial solutions, first evenly divide the time series into n parts, where each part is regarded as the smallest unit period that cannot be further divided. Here, n = 10 is taken as an example.
[0082] Step 2-2: Randomly search for the value of K in the set {2, 3, 4, 5, 6, 7, 8, 9, 10}.
[0083] Step 2-3: Given K, select each period of the input sequence based on the greedy strategy. First, consider K = 2, and use A and B to represent the start and end points of the time series respectively. By maximizing the distribution distance d(S AC , S CB)Select one split point (denoted as C) from 9 candidate split points. Among them, d is a distance metric function, such as Euclidean or edit distance, and S AC and S CB represent the period from A to C and the period from C to B respectively.
[0084] Step 2-4: After determining C, consider K = 3, and use the same strategy to select another point D. Applying a similar strategy to different values of K can obtain the starting and ending points of the K subsequences with the least similar distribution.
[0085] Step 2-5: The K value in the set {2, 3, 4, 5, 6, 7, 8, 9, 10} that maximizes the average distribution distance of the K subsequences with the least similar distribution is the optimal K value.
[0086] Step 3: Divide the obtained sample data set into a 70% training data set I, a 10% validation data set I, and a 20% prediction data set I according to a ratio, and then perform a standardization operation. The formula is as follows:
[0087]
[0088] where mean(x i ) and std(x i ) represent the mean and variance of the variable x i respectively.
[0089] Step 4: Establish an out-of-distribution representation learning model DIVERSIFY for time series. This model models the internal latent distribution of time series with dynamic distributions and learns the temporal differences and latent distribution characteristics between different distributions;
[0090] Step 4-1: Fine-grained feature update: In this step, the pseudo-domain class label is used as the label of the classifier to update the feature extractor, which helps the feature extractor better capture the feature information within the domain. The pseudo-domain class label here treats each category of each domain as a new class, thereby attaching more fine-grained category and domain information. The calculation formula is as follows:
[0091] s = d′ × C + y (2)
[0092] where s ∈ {1, 2, …, S}, S = K × C, K is the predefined number of latent distributions, and C is the number of initial categories; d′ is the domain label. In the first iteration, all samples are initialized, i.e., d′ = 0.
[0093] Then, supervised learning is carried out using the pseudo-domain class label, and the loss function is as shown in the following formula:
[0094]
[0095] Among them, h f , respectively represent the feature extractor, the Bottleneck layer and the classifier layer of this part, and L represents the cross-entropy loss function.
[0096] Step 4-2, Latent distribution representation: The goal of this step is to identify the domain label of each sample to obtain the latent distribution information. The main idea is to maximize the difference between different latent distributions to expand the diversity of the data distribution. First, obtain the centroid of each domain with intra-class features, and the calculation formula is shown as follows:
[0097]
[0098] Among them, respectively represent the Bottleneck layer and the classifier layer of this part. is the initial centroid of the k-th latent domain, and δ k is the k-th element of the softmax output.
[0099] Then, we use the distance function D to obtain the pseudo-domain label through the nearest centroid classifier, and the calculation formula is shown as follows:
[0100]
[0101] Next, calculate the centroid and obtain the updated pseudo-domain label:
[0102]
[0103]
[0104] Among them, Ι(a) is 1 only when a is true, otherwise it is 0. Finally, the loss function of this step can be obtained:
[0105]
[0106] Among them is the discriminator of this part, which contains multiple linear layers and a classification layer. is the gradient reversal layer with hyperparameter λ1.
[0107] Step 4-3, Domain-invariant representation learning: This step learns the domain-invariant representation by using the pseudo-domain label in the previous step, which will help the model learn general features to handle data in different sub-domains. The specific idea is to directly use adversarial training to update the classification loss and the domain classifier loss as shown in the following formula:
[0108]
[0109] Step 5: Use the training dataset I to train DIVERSIFY to obtain the corresponding weights and biases. Substitute the obtained weights and biases into the validation dataset I, then calculate the prediction error of the validation dataset I, and save the weights and biases that minimize the prediction error of the validation set.
[0110] Step 6: Process the input of the industrial time series obtained in Step 1: Combine the variable values at m historical moments and the predicted target value after t seconds to form a sample data; where m is the number of input historical time points and t is the prediction time point; and divide the sample data set into a training dataset II, a validation dataset II, and a prediction dataset II according to a ratio, and then perform a standardization operation.
[0111] Step 7: Establish a prediction model based on spatial attention and temporal convolutional network. This model uses the feature extraction layer in the model DIVERSIFY trained in Steps 4 and 5 as the temporal convolutional network here, and at the same time uses spatial attention to mine the multi-variable coupling features of the industrial time series, and finally combines the two parts of features through a gated fusion mechanism to achieve prediction.
[0112] Step 7-1: Determine the model input dimension: batchsize×m×n features , where batchsize is the batch size set during the batch training of the model, m is the number of input historical time points, and n features is the number of features at each moment of the input.
[0113] Step 7-2: Determine the structural parameters of the prediction model according to the model input dimension in Step 4-1.
[0114] Step 7-3: Generate a spatial attention map by utilizing the spatial relationship of features: Apply average pooling along the channel axis direction to represent the average aggregated features of the entire channel; then generate a two-dimensional spatial attention map through a standard convolutional layer and a residual connection, and the calculation is as shown in the following formula:
[0115] SAM(x) = σ(f(x))·x + x (10)
[0116] where f represents the convolutional operation and σ is the activation function.
[0117] Step 8: Use the training dataset II to train the deep neural network to obtain the corresponding weights and biases. Substitute the obtained weights and biases into the validation dataset II, then calculate the prediction error of the validation dataset II, and save the weights and biases that minimize the prediction error of the validation set; select the AdamW optimizer, use the mean squared error as the loss function, and adjust the prediction model parameters through backpropagation to improve the prediction accuracy of the model.
[0118] Step 8-1: Determine the hyperparameters for predicting model training: learning rate lr, maximum number of iterations I MAX ; Randomly initialize the weight matrix w and bias β of each network layer, and set the initial iteration number I = 0;
[0119] Step 8-2: Input the input two-dimensional matrix X = [X(t), …, Z(t-m)] into the temporal convolutional network and spatial attention module:
[0120] h T = TCN(X) (11)
[0121] h S = SAM(X) (12)
[0122] Step 8-3: Input the temporal features and spatial features of the latent distribution obtained in the above steps into the Gated Fusion module, and the calculation formula is shown as follows:
[0123] h = h S ·σ(h S + h T ) + h T ·(1 - σ(h S + h T )) (13)
[0124] Step 8-4: Obtain the fused feature h through the above steps, and finally use the fully connected layer to weight and fuse the features to obtain the final output;
[0125] y = ReLU(w a h + β) (14)
[0126] y represents the final predicted value of the model; w a represents the transformation matrix; β represents the bias value; ReLU is a non-linear activation function.
[0127] Step 9: Substitute the weights and biases that minimize the prediction error of the validation set in the previous step into the prediction model, and then apply it to the industrial time series dataset to calculate the predicted values of the prediction model on the dataset.
[0128] The results of the above method are as Figure 3 、 4 shown. Figure 3 It is a comparison chart of the mean absolute error of predicting the outlet moisture of the next 60, 90, and 120 s using 120 s of historical data of the industrial loose rewetting process. Figure 4 (a) is a comparison chart of the predicted value of the outlet moisture of 90 s of each model in this process with the true result, Figure 4(b) is a comparison chart of the absolute value of the error between the predicted value and the true value at some time points during this process. It can be seen that the proposed method can effectively complete the time series prediction of the loose moisture regain process, improve the prediction accuracy of the model. At the same time, the experimental results can clearly demonstrate the superiority of the proposed method in industrial time series prediction tasks, as well as its effective capture of time series characteristics and the ability to handle time series drift phenomena.
[0129] In summary, the present invention proposes an industrial time series prediction method AdaTCN based on out-of-distribution representation learning. This method describes the distribution characteristics of industrial time series through time series similarity quantification technology, so as to obtain time series segments with different distribution characteristics. Subsequently, these segments are input into a temporal convolutional network, and the out-of-distribution representation learning method for time series (DIVERSIFY) is used to explore the potential distribution of the sequence. Through this process, a TCN based on dynamic time series is successfully constructed, and the potential distribution of the time series is effectively learned. Finally, by fusing TCN with spatial attention, accurate prediction results are obtained to complete the industrial time series prediction task. The present invention effectively solves the problem that the prediction performance of the prediction model will seriously decline during the prediction stage due to the time series offset phenomenon commonly existing in industrial time series, and has theoretical and practical significance for the field of industrial time series prediction.
[0130] The present invention also provides an industrial time series prediction device based on out-of-distribution representation learning, including a front-end interface and a background. The background is provided with a memory and a processor. A program is stored in the processor. When the processor loads the program, it executes the above-mentioned method steps to obtain a trained and optimized prediction model, so that the model predicts industrial time series data to obtain the predicted values of the corresponding industrial data, and visually displays them to the front-end user interface of the industrial scenario.
[0131] The embodiments described above will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several transformations and improvements can be made. These all belong to the protection scope of the present invention.
Claims
1. An industrial time series prediction method based on out-of-distribution representation learning, characterized in that, Including: Using temporal similarity quantification technology to describe the distribution characteristics of industrial time series, so as to obtain time series segments with different distribution characteristics; inputting these segments into the Temporal Convolutional Network (TCN), and exploring the potential distribution of the sequence with the help of the out-of-distribution representation learning method DIVERSIFY for time series; combining TCN with spatial attention to obtain the final prediction model for predicting industrial time series data.
2. The industrial time series prediction method based on out-of-distribution representation learning according to claim 1, wherein The method specifically includes the following steps: Step 1: Collect industrial time series data and perform preprocessing. Step 2: Input the industrial time series data into the temporal distribution similarity quantification module to obtain K most dissimilar subsequences and assign corresponding labels. Step 3: Divide the obtained sample data set into training data set I, validation data set I, and prediction data set I according to a certain proportion, and then perform standardization operations. Step 4: Establish an out-of-distribution representation learning model DIVERSIFY for time series. This model models the internal potential distribution of time series with dynamic distributions, and learns the temporal differences and potential distribution characteristics between different distributions. Step 5: Use training data set I to train DIVERSIFY to obtain the corresponding weights and biases, bring the obtained weights and biases into validation data set I, then calculate the prediction error of validation data set I, and save the weights and biases that minimize the prediction error of the validation set. Step 6: Input and process the industrial time series obtained in Step 1: Combine the variable values at m historical moments and the predicted target value after t seconds to form a sample data; where m is the number of input historical time points and t is the prediction time point; and divide the sample data set into training data set II, validation data set II, and prediction data set II according to a certain proportion, and then perform standardization operations. Step 7: Establish a prediction model based on spatial attention and temporal convolutional network. This model uses the feature extraction layer in the model DIVERSIFY trained in Steps 4 and 5 as the temporal convolutional network here, and at the same time uses spatial attention to mine the multi-variable coupling characteristics of industrial time series, and finally combines the two parts of features through a gating fusion mechanism to achieve prediction. Step 8: Use training data set II to train the prediction model to obtain the corresponding weights and biases, bring the obtained weights and biases into validation data set II, then calculate the prediction error of validation data set II, and save the weights and biases that minimize the prediction error of the validation set; select the AdamW optimizer, use the mean squared error as the loss function, and adjust the parameters of the prediction model through backpropagation to improve the prediction accuracy of the model. Step 9: Substitute the weights and biases that minimize the prediction error of the validation set in the previous step into the prediction model, and then apply it to the industrial time series data set to calculate the predicted values of the prediction model on the data set.
3. The industrial time series prediction method based on out-of-distribution representation learning according to claim 1, wherein The preprocessing includes the following steps: Step 1-1: Remove outliers. Step 1-2: Filter the industrial time series.
4. The industrial time series prediction method based on out-of-distribution representation learning according to claim 1, wherein The temporal distribution similarity quantification module can maximize the use of the information contained in continuous time series by finding the most dissimilar time series periods, mainly including the following steps: The time series is evenly divided into n parts, where each part is the smallest unit period; Define the set {q1, ……, q K}, given a K value to iterate from an initial value to a given value, and in each iteration round, further select each time series period of the input sequence based on a greedy strategy; the further selection of each time series period of the input sequence is as follows: between the start point and the end point of the time series, select 1 split point from the candidate split points to subdivide the data segment by maximizing the distribution distance, and then further select split points for the current smallest data segment until K data segments are all selected, which are used to represent the start point and the end point of the time sub-sequences with the least similar distributions; In the set, the K value that maximizes the average distribution distance of the K most dissimilar subsequences is the optimal K value, and it is output.
5. The industrial time series prediction method based on out-of-distribution representation learning according to claim 1, wherein The out-of-distribution representation learning model DIVERSIFY structure of the time series described above includes: DIVERSIFY consists of a min-max adversarial game, aiming to characterize the latent distribution to cope with the distribution dynamics in time series data, which is mainly divided into the following three steps: Step 4-1, Fine-grained feature update: First, use the pseudo-domain class label as the label of the classifier to update the feature extractor, so that the feature extractor can better capture the in-domain feature information; the pseudo-domain class label regards each category of each domain as a new class, so as to attach more fine-grained category and domain information, and its calculation formula is as follows: s = d′×C + y where s ∈ {1, 2, …, S}, s refers to the value range of the pseudo-domain class label, S refers to the maximum number of this value range, that is, the number of classes of the pseudo-domain class label, S = K×C, K is the predefined number of latent distributions, and C is the number of initial classes; d′ is the domain label, and in the first iteration, all samples are initialized, that is, d′ = 0; Secondly, use the pseudo-domain class label for supervised learning, and the loss function is as follows: where h f , represent the feature extractor, the Bottleneck layer and the classifier layer of this part respectively, and L represents the cross-entropy loss function; Step 4-2, Latent distribution representation: By maximizing the difference between different latent distributions, identify the domain label of each sample to obtain latent distribution information and expand the diversity of data distribution; First, obtain the centroid of each domain with intra-class features, and the calculation formula is as follows: Among them, respectively represent the Bottleneck layer and the classifier layer of this part. is the initial centroid of the k-th latent domain, and δ k is the k-th element of the softmax output; Secondly, use the distance function D to obtain the pseudo-domain label through the nearest centroid classifier, and the calculation formula is as follows: Thirdly, calculate the centroid and obtain the updated pseudo-domain label: where Ι(a) is 1 only when a is true, otherwise it is 0; finally, the loss function of this step can be obtained: Among them is the discriminator for this part, which includes multiple linear layers and a classification layer; is the gradient reversal layer with hyperparameter λ1; Step 4-3, Domain-Invariant Representation Learning: By using the pseudo-domain labels in the previous step to learn domain-invariant representations, the model can learn general features to handle data from different sub-domains; specifically, adversarial training is directly used to update the classification loss and the domain classifier loss as follows:
6. The industrial time series prediction method based on out-of-distribution representation learning according to claim 1, wherein The establishment of the prediction model based on spatial attention and temporal convolutional network includes the following steps: Step 7-1: Determine the model input dimension: batchsize×m×n features , where batchsize is the batch size set during the batch training of the model, m is the number of input historical time points, and n features is the number of features at each input moment; Step 7-2, Determine the structural parameters of the prediction model according to the model input dimension in Step 4-1.
7. The industrial time series prediction method based on out-of-distribution representation learning according to claim 1, wherein The temporal convolutional module TCN effectively mines the temporal dynamic features and non-linear features of the variable to be predicted by using causal convolution.
8. The industrial time series prediction method based on out-of-distribution representation learning according to claim 1, wherein The spatial attention module includes the following steps: Step 7-3, Generate a spatial attention map by using the spatial relationship of features: Apply average pooling along the channel axis direction, which represents the average aggregated feature of the entire channel; then generate a two-dimensional spatial attention map through a standard convolutional layer and a residual connection, and the calculation is as follows: SAM(x) = σ(f(x))·x + x where f represents the convolution operation and σ is the activation function.
9. The industrial time series prediction method based on out-of-distribution representation learning according to claim 1, wherein The prediction model training and calculation process includes the following steps: Step 8-1: Determine the hyperparameters for predicting model training: learning rate lr, maximum number of iterations I MAX ; Randomly initialize the weight matrix w and bias β of each network layer, and set the initial number of iterations I = 0; Step 8-2, Input the input two-dimensional matrix X = [X(t), …, X(t - m)] into the temporal convolutional network and the spatial attention module: h T = TCN(X) h S = SAM(X) Step 8-3, Input the temporal features and spatial features of the latent distribution obtained in the above steps into the gated fusion module Gated Fusion, and the calculation formula is as shown below: h = h S ·σ(h S +h T )+h T ·(1 - σ(h S +h T )) Step 8-4, Obtain the fused feature h through the above steps, and finally use the fully connected layer to weight the fused feature to obtain the final output; y = ReLU(w a h + β) y represents the final predicted value of the model; w a represents the conversion matrix; β represents the bias value; ReLU is the non-linear activation function.
10. An industrial time series prediction device based on out-of-distribution representation learning, characterized in that, It includes a front end of the interface and a back end. The back end is provided with a memory and a processor. A program is stored in the processor. When the processor loads the program, it executes the method steps described in any one of claims 1-9 to obtain a trained and optimized prediction model, so that the model predicts industrial time series data, obtains prediction values of corresponding industrial data, and visually displays them to the user front end of the industrial scenario.