Time Series Prediction Model Training Method, Time Series Prediction Method and Device

By disaggregating the time series samples and generating the attention weight matrix, the problem of difficult and low prediction accuracy in the time series prediction model in the prior art is solved, and efficient and interpretable long-term time series prediction is achieved.

CN119025924BActive Publication Date: 2025-06-27INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411498056.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-25
Publication Date
2025-06-27
Estimated Expiration
2044-10-25

AI Technical Summary

Technical Problem

Existing time series prediction models have problems with difficult interpretation and low accuracy in prediction results, especially when predicting long-term time series.

Method used

By discretizing the time series samples, the characteristics and information of the sequence fragment samples are obtained, the attentional computing mechanism is used to generate an attention weight matrix, and the importance of each sequence fragment to the prediction results are quantified, so as to achieve the interpretability of the model and more accurate prediction.

Benefits of technology

The explanatory nature of the time series prediction model is achieved, the accuracy of long-term time series prediction is improved, the calculation cost is reduced, and the length of the input data is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119025924B_ABST
    Figure CN119025924B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method for training a time series prediction model, a time series prediction method and a device, which are applied to the technical field of data processing. The method includes: obtaining a time series sample, and performing discretization processing on the time series sample to obtain a sequence segment sample; obtaining the segment sample features and segment sample information of the sequence segment sample; obtaining an attention weight matrix of the time series sample according to the attention operation mechanism, segment sample features, and segment sample information of the time series prediction model; obtaining a time series sample prediction result according to the attention weight matrix; and training the time series prediction model through the time series sample prediction result to obtain a trained time series prediction model. The time series prediction model obtained by the present invention has interpretability and high prediction result accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of data processing, and particularly to a method for training a time series prediction model, a time series prediction method, and a device. Background Art

[0002] Time series data is a data sequence including multiple data arranged in chronological order. Time series prediction can be applied to various application scenarios. For example, weather information within a historical time period can be used to predict weather information within a future time period. Another example is that traffic flow information within a historical time period can be used to predict traffic flow within a future time period.

[0003] In related technologies, a time series prediction model (Sample Convolution and Interaction Network, SCINet) can be used for time series prediction, or a deep learning model for long-term time series prediction (Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series Forecasting, AuToFormer) can be used for time series prediction.

[0004] However, the time series prediction models in related technologies have the disadvantages of being difficult to interpret, and these time series prediction models have problems with low accuracy of prediction results when predicting long-period time series data. Summary of the Invention

[0005] Embodiments of the present disclosure provide a method for training a time series prediction model, a time series prediction method, and a device, so as to solve the problems that the time series prediction models in related technologies are difficult to interpret and have low accuracy of prediction results.

[0006] In a first aspect, embodiments of the present disclosure provide a method for training a time series prediction model, including:

[0007] Obtain a time series sample, and perform discretization processing on the time series sample to obtain a sequence segment sample;

[0008] Obtain the segment sample features and segment sample information of the sequence segment sample;

[0009] According to the attention operation mechanism of the time series prediction model, the segment sample features, and the segment sample information, obtain the attention weight matrix of the time series sample;

[0010] According to the attention weight matrix, obtain the time series sample prediction result of the time series sample;

[0011] Train a time series prediction model with the time series prediction results of the time series samples to obtain a trained time series sample prediction model.

[0012] In a second aspect, an embodiment of the present disclosure provides a time series prediction method, including:

[0013] Obtain a target time series to be predicted;

[0014] Perform discretization processing on the target time series to obtain a plurality of target sequence segments;

[0015] Input the target sequence segments into the trained time prediction model to obtain the time series prediction result of the target time series; wherein, the trained time prediction model is obtained by the time series prediction model acquisition method described in the first aspect.

[0016] In a third aspect, an embodiment of the present disclosure provides a time series prediction model training device, including:

[0017] A first acquisition module, configured to acquire a time series sample and perform discretization processing on the time series sample to obtain a sequence segment sample;

[0018] A second acquisition module, configured to acquire the segment sample features of the sequence segment sample; the segment sample features include the sample information of the sequence segment sample corresponding to the segment sample features;

[0019] A third acquisition module, configured to obtain an attention weight matrix according to the attention operation mechanism of the time series prediction model and the segment sample features;

[0020] A fourth acquisition module, configured to obtain a time series sample prediction result according to the attention weight matrix;

[0021] A fifth acquisition module, configured to train a time series prediction model with the time series sample prediction result to obtain a time series prediction model.

[0022] In a fourth aspect, an embodiment of the present disclosure provides a time series prediction device, including:

[0023] A sixth acquisition module, configured to acquire a target time series to be predicted;

[0024] A seventh acquisition module, configured to perform discretization processing on the target time series to obtain a plurality of target sequence segments;

[0025] An eighth acquisition module, configured to input the target sequence segments into the trained time prediction model to obtain the time series prediction result of the target time series;

[0026] Among them, the time prediction model is obtained by the method for obtaining a time series prediction model described in the first aspect.

[0027] In a fifth aspect, an embodiment of the present disclosure provides an electronic device, including:

[0028] A memory for storing a computer program;

[0029] A processor for executing the computer program to implement the method described in the first aspect or the second aspect.

[0030] In a sixth aspect, an embodiment of the present disclosure provides a computer-readable storage medium for storing a computer program; wherein when the computer program is executed by a processor, the method described in the first aspect or the second aspect is implemented.

[0031] In a seventh aspect, an embodiment of the present disclosure provides a computer program product, including computer programs / instructions, and when the computer programs / instructions are executed by a processor, the steps of the method described in any one of the first aspect or the second aspect are implemented.

[0032] Regarding the prior art, the present disclosure has the following advantages:

[0033] In the embodiment of the present disclosure, in this embodiment, the time series sample is discretized to obtain a sequence segment sample, the segment sample feature of the sequence segment sample and the sample information of the sequence segment sample are obtained, and according to the attention operation mechanism of the time series prediction model and the segment sample feature, an attention weight matrix is obtained. The obtained attention weight matrix includes the sample information of the segment sample feature used to calculate the attention weight. Based on the attention weights in the attention weight matrix and the sample information of the segment sample feature used to calculate the attention weight, the importance of each sequence segment sample input into the time series prediction model for the prediction result can be quantified, and the importance degree of the sequence segment sample corresponding to the segment sample information for the time series prediction result can be determined according to the sample information of the attention weight matrix. In other words, based on the attention weight matrix including the sample information of the segment sample feature used to calculate the attention weight, the model interpretation of the time prediction model is realized, and the time series prediction model obtained based on the attention weight matrix can predict long-term time series in an interpretable manner.

[0034] The above description is only an overview of the technical solution of the present disclosure. In order to be able to understand the technical means of the present disclosure more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present disclosure more obvious and understandable, the specific embodiments of the present disclosure are specifically exemplified below. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] To more clearly illustrate the technical solutions in the embodiments of the present disclosure or related technologies, the following will briefly introduce the accompanying drawings required for the description of the embodiments.

[0036] Figure 1 It is a step diagram of a time series prediction model training method provided by an embodiment of the present disclosure;

[0037] Figure 2 It is a schematic diagram of the result of discretization processing of a time series sample provided by an embodiment of the present disclosure;

[0038] Figure 3 It is a schematic diagram of the structure of a time series prediction model provided by an embodiment of the present disclosure;

[0039] Figure 4 It is a step diagram of another time series prediction model training method provided by an embodiment of the present disclosure;

[0040] Figure 5 It is a step diagram of yet another time series prediction model training method provided by an embodiment of the present disclosure;

[0041] Figure 6 It is a schematic diagram of the result of discretization processing of another time series sample provided by an embodiment of the present disclosure;

[0042] Figure 7 It is a step diagram of a time series prediction method provided by an embodiment of the present disclosure;

[0043] Figure 8(a) is a schematic diagram of the attention weight matrix corresponding to Parameter 1 provided by an embodiment of the present disclosure;

[0044] Figure 8(b) is a schematic diagram of the attention weight matrix corresponding to Parameter 2 provided by an embodiment of the present disclosure;

[0045] Figure 8(c) is a schematic diagram of the attention weight matrix corresponding to Parameter 3 provided by an embodiment of the present disclosure;

[0046] Figure 8(d) is a schematic diagram of the attention weight matrix corresponding to Parameter 4 provided by an embodiment of the present disclosure;

[0047] Figure 8(e) is a schematic diagram of the attention weight matrix corresponding to Parameter 5 provided by an embodiment of the present disclosure;

[0048] Figure 8(f) is a schematic diagram of the attention weight matrix corresponding to Parameter 6 provided by an embodiment of the present disclosure;

[0049] Figure 8(g) is a schematic diagram of the attention weight matrix corresponding to Parameter 7 provided by an embodiment of the present disclosure;

[0050] Figure 9(a) is a schematic diagram of the attention weight matrix corresponding to Single Prediction Parameter 1 provided by an embodiment of the present disclosure;

[0051] Figure 9(b) is a schematic diagram of the attention weight matrix corresponding to the single prediction parameter 2 provided by the embodiment of the present disclosure;

[0052] Figure 9(c) is a schematic diagram of the attention weight matrix corresponding to the single prediction parameter 3 provided by the embodiment of the present disclosure;

[0053] Figure 9(d) is a schematic diagram of the attention weight matrix corresponding to the single prediction parameter 4 provided by the embodiment of the present disclosure;

[0054] Figure 9(e) is a schematic diagram of the attention weight matrix corresponding to the single prediction parameter 5 provided by the embodiment of the present disclosure;

[0055] Figure 9(f) is a schematic diagram of the attention weight matrix corresponding to the single prediction parameter 6 provided by the embodiment of the present disclosure;

[0056] Figure 9(g) is a schematic diagram of the attention weight matrix corresponding to the single prediction parameter 7 provided by the embodiment of the present disclosure;

[0057] Figure 10 is a schematic diagram of the true value and predicted value of OT provided by the embodiment of the present disclosure;

[0058] Figure 11(a) is a schematic diagram of the global statistical attention weight matrix corresponding to parameter 1 provided by the embodiment of the present disclosure;

[0059] Figure 11(b) is a schematic diagram of the global statistical attention weight matrix corresponding to parameter 2 provided by the embodiment of the present disclosure;

[0060] Figure 11(c) is a schematic diagram of the global statistical attention weight matrix corresponding to parameter 3 provided by the embodiment of the present disclosure;

[0061] Figure 11(d) is a schematic diagram of the global statistical attention weight matrix corresponding to parameter 4 provided by the embodiment of the present disclosure;

[0062] Figure 11(e) is a schematic diagram of the global statistical attention weight matrix corresponding to parameter 5 provided by the embodiment of the present disclosure;

[0063] Figure 11(f) is a schematic diagram of the global statistical attention weight matrix corresponding to parameter 6 provided by the embodiment of the present disclosure;

[0064] Figure 11(g) is a schematic diagram of the global statistical attention weight matrix corresponding to parameter 7 provided by the embodiment of the present disclosure;

[0065] Figure 12(a) is a schematic diagram of the influence degree of each parameter on the prediction result of parameter 1 provided by the embodiment of the present disclosure;

[0066] Figure 12(b) is a schematic diagram showing the influence degree of each parameter on the prediction result of parameter 2 provided by the embodiment of the present disclosure;

[0067] Figure 12(c) is a schematic diagram showing the influence degree of each parameter on the prediction result of parameter 3 provided by the embodiment of the present disclosure;

[0068] Figure 12(d) is a schematic diagram showing the influence degree of each parameter on the prediction result of parameter 4 provided by the embodiment of the present disclosure;

[0069] Figure 12(e) is a schematic diagram showing the influence degree of each parameter on the prediction result of parameter 5 provided by the embodiment of the present disclosure;

[0070] Figure 12(f) is a schematic diagram showing the influence degree of each parameter on the prediction result of parameter 6 provided by the embodiment of the present disclosure;

[0071] Figure 12(g) is a schematic diagram showing the influence degree of each parameter on the prediction result of parameter 7 provided by the embodiment of the present disclosure;

[0072] Figure 13 is a screening matrix of an ETTH2 dataset provided by the embodiment of the present disclosure;

[0073] Figure 14 is a schematic diagram of a time series prediction model training device provided by the embodiment of the present disclosure;

[0074] Figure 15 is a schematic diagram of a time series prediction device provided by the embodiment of the present disclosure;

[0075] Figure 16 is a block diagram of an electronic device provided by the embodiment of the present disclosure. Detailed implementation manners

[0076] Hereinafter, exemplary embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be completely conveyed to those skilled in the art.

[0077] Some nouns or terms appearing in this application are applicable to the following explanations:

[0078] Electricity Transformer Temperature (ETT) dataset, which includes time series of transformer power load, oil temperature, and the corresponding moments of power load and oil temperature within different time periods.

[0079] It should be noted that the data samples involved in this application (including but not limited to the data for analysis, stored data, displayed data, etc.) are all information and data authorized by users or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or reject.

[0080] Figure 1 is a step diagram of a time series prediction model training method provided by an embodiment of the present disclosure. Referring to Figure 1 , the time series prediction model training method includes:

[0081] Step 101, obtain a time series sample, and perform discretization processing on the time series sample to obtain multiple sequence segment samples.

[0082] The time series sample is obtained by sorting multiple data samples according to the acquisition time of the data samples, and the sequence segment sample includes at least one data sample.

[0083] Furthermore, obtain data samples at multiple preset times, and arrange the multiple data samples in the order of the preset times to obtain a time series sample.

[0084] The data in the time series sample can be weather data, traffic flow data, power load data of equipment in the power system, or other data. This embodiment does not limit the type of data.

[0085] Exemplarily, the time series sample can be discretized according to a preset sequence length to obtain multiple sequence segment samples. Among them, the sequence length reflects the number of data in the sequence.

[0086] For example, referring to Figure 2 , the time series sample includes six data samples from data a1 to data a6. Among them, each data sample has a corresponding acquisition time, and data a1 to data a6 are arranged in the order of the acquisition time. For example, the acquisition time of data a1 is earlier than that of other data, and the acquisition time of data a2 is earlier than that of data a3 to data a6.

[0087] Exemplarily, the preset sequence length is 3, and according to the preset sequence length of 3, the Figure 2 time series sample is discretized to obtain 3 sequence segment samples. Each sequence segment sample includes 3 data samples.

[0088] To preserve the context information of sequence segment samples, the data of two adjacent sequence segment samples need to be connected end to end. Specifically, in two adjacent sequence segment samples, the data at the end of the first sequence segment sample is the same as the data at the beginning of the second sequence segment sample. Refer to Figure 2 , the first sequence segment sample includes data a1, data a2, and data a3, and the second sequence segment sample includes data a3, data a4, and data a5. The last data a3 in the first sequence segment sample is also the first data in the second sequence segment sample, which is data a3.

[0089] Step 102: Obtain the segment sample features and segment sample information of the sequence segment sample.

[0090] By obtaining the segment sample features and sample information of the sequence segment sample, it is possible to ensure that in subsequent processing, the correspondence between the sequence segment sample input to the time series prediction model and the segment sample features is retained.

[0091] Exemplarily, the segment sample information may include the identifier of the sequence segment sample and the moment corresponding to the sequence segment sample. Among them, the identifier of the sequence segment sample can be the data identifier of the sequence segment sample, and the data identifier can be the name of the data type, the data type code, or other data identifiers. For example, if the sequence segment sample includes the power load data of device A in the power system, the segment sample information may include the code for identifying the power load of device A and the acquisition moment of each power load.

[0092] Exemplarily, through a one-dimensional convolutional layer, the sequence segment sample is convolved to obtain the segment sample features of the sequence segment sample.

[0093] Step 103: According to the attention operation mechanism, segment sample features, and segment sample information of the time series prediction model, obtain the attention weight matrix of the time series sample.

[0094] Among them, the attention weight matrix includes the segment sample information.

[0095] Furthermore, the attention weight matrix includes: attention weights, and the segment sample information corresponding to the segment sample features used to calculate the attention weights.

[0096] Exemplarily, refer to Figure 3 , the segment sample features are convolved to obtain the query (Query) vector, key (Key) vector, and value (Value) vector of the sample features, and the attention weight matrix of the segment sample features is obtained according to the query connection and the key vector.

[0097] Step 104: Obtain the time series sample prediction result according to the attention weight matrix.

[0098] Exemplarily, a dot product attention operation is performed on the attention weight matrix and the value vector to obtain attention features, and the attention features are input into the forward channel of the time series prediction model to obtain the time series sample prediction result. Among them, the forward channel may include multiple fully connected layers. The forward channel composed of multiple fully connected layers serves as a predictor in the time series prediction model.

[0099] Step 105: Train the time series prediction model with the time series sample prediction result to obtain the trained time series prediction model.

[0100] Exemplarily, the training of the time series prediction model includes multiple training rounds. In each training round, an attention weight matrix is obtained according to the segment sample features, segment sample information of the sequence segment sample, and the attention operation mechanism of the time series prediction model, and the time series sample prediction result is obtained according to the time series prediction model and the attention weight matrix.

[0101] When the current round does not meet the training stop condition, optimize the model parameters of the time series prediction model according to the time series sample prediction result and the time series result as the label, and continue to train the time series prediction model until the model training meets the training stop condition, and then determine the time series prediction model obtained in the last round as the time series prediction model.

[0102] For example, when the number of rounds reaches the preset round number threshold, it is determined that the model training meets the training stop condition, and the time series prediction model obtained in the last training is determined as the time series model.

[0103] For another example, during the training process of each round, obtain the loss function value, and optimize the model parameters in the time series prediction model according to the loss function value to complete one training of the time series prediction model. Further, in each training round, obtain the time series sample prediction result output by the time series prediction model, and obtain the loss function value according to the time series prediction sample result and the time series data as the label. When the loss function value meets the preset loss function value requirement, it is determined that the model training meets the training stop condition, and the time series prediction model obtained in the last training is determined as the time series model.

[0104] Furthermore, the loss function values of each training round in the training process of the most recent preset number of rounds can be obtained. When the loss function value in the preset round does not decrease as the number of rounds increases, it is determined that the loss function value meets the preset loss function requirement.

[0105] In the related art, common time series prediction models include: Long Short-Term Memory (LSTM) model, Recurrent Neural Network (RNN) model, Transformer structure model, etc. There are two problems to be solved in time series prediction models. One is how to improve the accuracy of long sequence time series prediction (LSTF), and the other is how to explain the reasons for the prediction results obtained by deep neural network models. Before the Transformer structure was proposed, the Recurrent Neural Network (RNN) model and the Long Short-Term Memory (LSTM) model dominated the time series prediction field. Although the RRN model and the LSTM model can achieve time series prediction, their cyclic structures lead to low computational efficiency and difficulties in feature extraction when dealing with long sequences, which limits their applications in time series prediction. Among them, the time series prediction effect of the Transformer structure model and its branches is better than that of other models.

[0106] However, for the Transformer model in the related art, its computational cost is proportional to the square of the length of the input and output data. The length of the input and output data of long-period time series is large. When using the Transformer model in the related art for long-period time series prediction, there are problems of high computational cost, and its attention value cannot directly represent the importance degree of variables, resulting in the problem that the Transformer model in the related art cannot be explained.

[0107] In the related art, a long-term time series prediction (AutoFormer) model can also be used for long-period time series prediction. However, the AutoFormer model relies on the assumption of time series periodicity, and its computational complexity is very high, and there are also unexplainable problems. Models based on the Transformer structure also include the LogSparseTransformer model, the Longformer model for processing long texts, and the Beyond Efficient Transformer for Long equence Time-Series Forecasting (Informer) model, etc. These models also have the problem of being difficult to explain. Based on these models, it is impossible to clarify the importance of parameters for the prediction results. The deep learning multi-level time series prediction (Temporal Fusion Transformer, TFT) model in the related art can explain the data basis of the model, but it uses attention to explain the importance of variables without considering the influence of the model structure on the attention value, so its interpretability is inaccurate. That is, the time series prediction models in the related art at least have the problem of lacking interpretability.

[0108] In this embodiment, the time series sample is discretized to obtain a sequence segment sample, and the segment sample feature and sample information of the sequence segment sample are obtained. According to the attention operation mechanism of the time series prediction model and the segment sample feature, an attention weight matrix is obtained. The obtained attention weight matrix includes the sample information of the segment sample feature used to calculate the attention weight. Based on the attention weights in the attention weight matrix and the sample information of the segment sample feature used to calculate the attention weights, the importance of each sequence segment sample input into the time series prediction model for the prediction result can be quantified, and the importance of the sequence segment sample corresponding to the sample information for the time series sample prediction result can be clearly obtained, thereby realizing the interpretability of the time series prediction model. That is, in this embodiment, based on the attention weight matrix including the sample information of the segment sample feature used to calculate the attention weights, the model interpretation of the time prediction model is realized. In other words, the time series prediction model obtained based on the attention weight matrix can predict long-term time series in an interpretable manner.

[0109] In addition, the time series samples are discretized. The sequence segment samples obtained based on the discretization are also discrete, and the segment sample features obtained from the sequence segment samples are also discrete. Based on the discrete segment sample features, the influence degree of the sequence segment samples on the attention weights in the attention weight matrix can be obtained. Based on the discrete segment sample features, when a screening matrix is obtained from the attention weights obtained from the subsequent segment sample features and a part of the sequence segment samples are screened out according to the screening matrix, the influence on the prediction result of the time series samples can be minimized when shielding unnecessary sequence segment samples. In other words, based on the sequence segment samples obtained by discretization, the computational cost of the time series prediction model for predicting long-period time series can be reduced, the prediction accuracy can be improved, and the interpretability of the time series prediction model can be realized.

[0110] Figure 4 is another time series model training method provided by an embodiment of the present application. Referring to Figure 4 , the method may include the following steps:

[0111] Step 201: Discretize the time series samples according to a preset sequence length to obtain sequence segment samples.

[0112] Among them, in two adjacent sequence segment samples, the first quantity of data samples at the end of the first sequence segment sample is the same as the first quantity of data samples at the beginning of the second sequence segment sample.

[0113] Among them, the first quantity can be a preset value and can be set according to user needs.

[0114] Among them, the sequence segment samples include a plurality of data arranged in chronological order, and the preset sequence length is equal to the number of data samples in the sequence segment samples.

[0115] For example, referring to Figure 2 , if the preset sequence length is 3, then each obtained sequence segment sample includes 3 data samples, and the data sample at the end of the first sequence segment sample is the same as one data sample at the beginning of the second sequence segment sample, and both are a3.

[0116] Step 202: Obtain the segment sample features and segment sample information of the sequence segment samples.

[0117] The segment sample information includes at least one of the following: the sample identifier of the sequence segment sample and the moment information of the sequence segment sample.

[0118] Among them, the sample identifier can be the name, number, or other identification information of the time series sample to which the sequence segment sample belongs.

[0119] The time information of the sequence segment sample can be the acquisition time information of the data samples in the sequence segment sample or the time period information of the sequence segment sample.

[0120] In the case where there are multiple sequence segment samples, the sequence segment samples and the segment sample features correspond one by one, and the segment sample features and the segment sample information correspond one by one. For example, if the segment sample information includes a sample identifier and time information, then the segment sample features have corresponding sample identifiers and time information.

[0121] Exemplarily, referring to Figure 3 the structural schematic diagram of the time series prediction model shown, through the interpretable discrete feature mapping layer in the model, the time series sample is segmented into discrete sequence segment samples. Through the interpretable discrete feature mapping process, the sequence segment samples are mapped into segment sample features.

[0122] Exemplarily, through the one-dimensional convolutional layer in the discrete feature mapping layer, the segment sample features of the sequence segment samples are extracted from the small-scale sequence segment samples obtained after discrete processing. Among them, feature extraction is performed through the one-dimensional convolutional layer to effectively retain the segment sample information corresponding to the extracted segment sample features.

[0123] Exemplarily, the width and stride of the kernel function in the one-dimensional convolutional layer can be set according to user requirements. Preferably, the width of the kernel function in the one-dimensional convolutional layer is 1 and the stride is 1. Based on the kernel function with a width of 1, it can be ensured that the segment sample features of one sequence segment sample are extracted at a time, ensuring that the segment sample features and the sequence segment samples can correspond one by one. Setting the kernel function stride to 1 can ensure that feature extraction is performed on each sequence segment sample to avoid data omission.

[0124] Specifically, the time series prediction module can include a feature mapping module (Feature Encoding, FE) and a positional encoding mapping module (Positional Encoding, PE). The segment sample features of the sequence segment samples can be obtained through the feature mapping module FE and the positional encoding mapping module PE, and the segment sample information tags are added to the segment sample features. Among them, in the time series prediction model, the positional encoding mapping module PE is used to mark the sequence order of the sequence variables. Specifically, the positional encoding mapping module PE is used to mark the position of the sequence segment sample in the time series sample. By retaining the positional encoding for the sequence segment sample, the segment sample features of each sequence segment sample can be tagged.

[0125] Step 203, perform convolutional processing on the segment sample features to obtain a query vector and a key vector.

[0126] Exemplarily, referring to Figure 3, the fragment sample features are convolved through a one-dimensional convolutional layer to obtain query vectors and key vectors. Among them, the fragment sample features correspond one-to-one with the sequence fragment samples. There are multiple sequence fragment samples, and there are also multiple fragment sample features. As Figure 3 shown, the sequence fragment sample h is convolved to obtain the query vector Qh, the key vector Kh, and the value vector Vh of the sequence fragment sample h.

[0127] Step 204: Obtain the attention weight matrix of the time series sample according to the query vector, the key vector, and the fragment sample information.

[0128] The attention weight matrix includes multiple attention weights, and each attention weight corresponds to the sample information of the fragment sample feature used to calculate the attention weight.

[0129] Specifically, the query vector and the key vector are multiplied to obtain an initial attention weight matrix, and each attention weight in the attention weight matrix is labeled with the fragment sample information to obtain an attention weight matrix containing the fragment sample information.

[0130] Among them, the attention weight matrix includes the fragment sample information of the sequence fragment sample. According to the attention weights in the attention weight matrix and their corresponding fragment sample information, it is possible to know the importance of the sequence fragment sample corresponding to the fragment sample information to the prediction result of the subsequent obtained time series sample. Thus, the interpretability of the time series prediction model is realized.

[0131] Step 205: Obtain the target attention weights greater than or equal to the preset weight threshold from the attention weight matrix.

[0132] The preset weight threshold can be set according to user needs. For example, when the requirement for the importance of the sequence fragment sample to the prediction result of the time series sample is relatively high, the preset weight threshold can be set larger. Conversely, the preset weight threshold can be set smaller.

[0133] Step 206: Obtain the target fragment sample features corresponding to the target attention weights from multiple fragment sample features.

[0134] Among them, one fragment sample feature corresponds to one sequence fragment sample.

[0135] Exemplarily, a screening matrix is constructed according to the target attention weights, and according to the screening matrix, the target fragment sample features corresponding to the target attention weights are screened out from multiple fragment sample features.

[0136] Exemplarily, Step 206 may include the following sub-steps (sub-step 2061 to sub-step 2063):

[0137] Sub-step 2061: Configure a first weight for the target attention weight and a second weight for other attention weights except the target attention weight.

[0138] Among them, the first weight corresponds to the segment sample features used to calculate the target attention weight, and the second weight corresponds to the segment sample features used to calculate other attention weights.

[0139] Exemplarily, the first weight is 1 and the second weight is 0. Specifically, the first weight corresponding to the sequence segment sample used to calculate the target attention weight is 1, and the second weight corresponding to the sequence segment sample used to calculate other attention weights is 0.

[0140] Sub-step 2062: Construct a screening matrix according to the first weight, the segment sample information of the segment sample features corresponding to the first weight, the second weight, and the segment sample information of the segment sample features corresponding to the second weight.

[0141] Exemplarily, if the first weight is 1 and the second weight is 0, the constructed screening matrix is a matrix with matrix elements being 1 or 0.

[0142] Sub-step 2063: Screen out the target segment sample features from multiple segment sample features according to the screening matrix.

[0143] Exemplarily, perform a multiplication process on the screening matrix and the matrix composed of multiple segment sample features to obtain an evaluation value matrix. Among them, the elements in the evaluation value matrix are evaluation values, and the evaluation values correspond one-to-one with the segment sample features, indicating a positive correlation with the probability of the segment sample features being selected.

[0144] Obtain the target evaluation values in the evaluation value matrix that are greater than or equal to a preset threshold, and determine the segment sample features corresponding to the target evaluation values as the target segment sample features.

[0145] Exemplarily, if the first weight is 1 and the second weight is 0, and the screening matrix is a matrix with matrix elements being 1 or 0. Perform a multiplication process on the screening matrix and the matrix composed of multiple segment sample features to obtain an evaluation value matrix, and the evaluation values in the evaluation value matrix are equal to 1 or 0. Specifically, the evaluation value corresponding to the segment sample features of the first weight is 1, and the evaluation value corresponding to the segment sample features of the second weight is 0. Obtain the segment sample features with an evaluation value equal to 1 and determine them as the target segment sample features.

[0146] Based on the screening matrix with matrix elements being 1 or 0, the target segment sample features can be quickly obtained from the segment sample features.

[0147] Step 207: Input the target segment sample features into the time series prediction model to obtain the time series sample prediction result.

[0148] Specifically, input the target segment sample features into the time series prediction model. Through the one-dimensional convolutional layer in the time series prediction model, obtain the query vector, key vector, and value vector corresponding to the target sample features. Obtain the attention weight matrix through the query vector and the key vector, perform dot product attention operation on the attention weight matrix and the value vector to obtain the attention features, and input the attention features into the forward channel to obtain the time series sample prediction result.

[0149] Step 208: Train the time series prediction model with the time series sample prediction result to obtain the trained time series prediction model.

[0150] The method in this step has been described in the aforementioned step 105 and will not be elaborated here.

[0151] Exemplarily, the attention operation mechanism in the Transformer structure can be used to obtain the attention weight matrix, and an interpretable time series prediction model can be obtained based on the attention weight matrix. The time series prediction model has the advantages of high prediction accuracy and interpretability. The interpretable time series prediction model based on the Transformer structure can, while accurately predicting long-period time series, explain the time series prediction model.

[0152] In this embodiment, the attention weight matrix contains the sequence segment information of the segment sample features. Based on the attention weight matrix, the interpretability of the time series prediction model is realized. An extraction matrix is obtained based on the interpretable attention weight matrix. According to the extraction matrix, the key target segment sample features can be obtained from multiple segment sample features. The time series prediction model is trained with the target segment sample features. When using this time series prediction model to predict the time series, the prediction result can be obtained through the key time series segments, thereby reducing the amount of data used for prediction, reducing the prediction time, lowering the prediction cost, and improving the prediction efficiency. The method of this embodiment is applicable to time series prediction, especially long-period time series prediction. It solves the problem in the related technology that when performing long-period time series prediction, the amount of data input into the time series prediction model is large, resulting in high prediction cost.

[0153] In one embodiment, there is at least one time series sample. Referring to Figure 5 , the time series prediction model training method may include the following steps:

[0154] Step 301: Discretize each time series sample separately to obtain at least one set of sequence segment samples.

[0155] Among them, one time series sample corresponds to one set of sequence segment samples.

[0156] For example, referring to Figure 6 , the time series sample includes a first time series sample and a second time series sample. The first time series sample includes 6 data samples from data sample a1 to data sample a6, and the second time series sample includes 6 data samples from data sample b1 to data sample b6.

[0157] The first time series sample is discretized to obtain three groups of sequence segment samples. These three groups of sequence segment samples are respectively: the sequence segment sample containing data samples a1 to a3, the sequence segment sample containing data samples a3 to a5, and the sequence segment sample containing data samples a5 to a6.

[0158] The second time series sample is discretized to obtain three groups of sequence segment samples. These three groups of sequence segment samples are respectively: the sequence segment sample containing data samples b1 to b3, the sequence segment sample containing data samples b3 to b5, and the sequence segment sample containing data samples b5 to b6.

[0159] Exemplarily, one time series sample corresponds to at least one parameter, and the time series sample is used to represent the data of the corresponding parameter at different times.

[0160] For example, the time series prediction model is used to predict the weather. There are multiple time series samples used to train the time series prediction model, specifically including: the first time series sample containing humidity data, and the second time series sample including temperature data. The first time series sample is discretized to obtain multiple sequence segment samples containing humidity data, where each sequence segment sample includes multiple humidity sample data arranged according to the acquisition time.

[0161] Step 302: Extract features from at least one group of sequence segment samples to obtain the segment sample features of the sequence segment samples.

[0162] Among them, one time series sample corresponds to one dimension of segment sample features. One dimension of segment sample features includes all the sequence segment sample features of the time series sample. Specifically, according to Figure 3 the interpretable discrete feature mapping layer in

[0163] extracts features from at least one group of sequence segment samples to obtain at least one dimension of segment sample features. Among them, the interpretable discrete feature mapping layer includes a feature mapping module FE and a position encoding mapping module PE.

[0164]

[0165] Among them, represents the number of dimensions, represents the time length, R represents the real number space, X1, X i , X T respectively represent the data of the sequence segment sample of the first time length, the data of the sequence segment sample of the i-th time length, and the data of the sequence segment sample of the T-th time length. represents d x dimensional real number space.

[0166] Among them, the calculation method of the position encoding mapping module PE is shown in the following formula.

[0167]

[0168]

[0169] Among them, represents the position encoding result obtained by encoding the 2j-th data of the t-th time length using the position encoder function, represents the position encoding result obtained by encoding the 2j + 1-th data of the t-th time length using the position encoder function, .

[0170] Furthermore, the feature mapping module FE and the position encoding mapping module PE constitute an interpretable discrete feature mapping layer EF. The mapping layer EF discretizes the input time series sample to obtain a discrete sequence segment sample, and maps the sequence segment sample to the corresponding segment sample feature. The mapping formula is as follows:

[0171]

[0172] Among them, represents the segment sample feature obtained by encoding the time series sample X according to the feature encoder; represents the position encoding of the time series sample X. represents the mapping result of mapping the segment sample feature obtained by encoding the time series sample X and the time series sample X.

[0173] Based on the segment sample features extracted by the interpretable discrete feature mapping layer, they correspond one-to-one with the input sequence segment samples. Therefore, it can ensure the intrinsic interpretability of the time series model.

[0174] Furthermore, referring to Figure 3 , after the discretized sequence segment sample passes through the interpretable discrete feature mapping layer, it is marked with a position label and then input into the dot product attention structure.

[0175] By extracting features from sequence segment samples of the same time series sample, feature vectors of segment samples in one dimension are obtained. Among the feature vectors of segment samples belonging to the same dimension, the feature vectors of segment samples and the sequence segment samples correspond one by one. Based on the feature vectors of segment samples corresponding one by one to the sequence segment samples, when obtaining the attention weight matrix according to the feature vectors of segment samples subsequently, the attention weight matrix can present the importance of sequence segment samples to the prediction result, realizing the interpretability of the time series model.

[0176] Furthermore, one time series sample corresponds to one dimension. By discretizing each time series sample respectively, isolation according to different dimensions and different time periods can be achieved. By performing isolation-based discretization processing on multi-dimensional time series samples, discretized sequence segment samples corresponding to the time series samples can be obtained, and features can be extracted from the set of discretized sequence segment samples. Thus, during the feature extraction process, interaction between different data can be strictly prevented to ensure one-to-one correspondence between sequence segment samples and feature vectors of segment samples, making the data flow clear and simple, and clearly realizing the interpretability of the time series prediction model.

[0177] Exemplarily, there is at least one time series sample. For example, in the application scenario of weather prediction, the time series samples can include time series samples of humidity, time series samples of temperature, and other time series samples related to weather prediction.

[0178] Step 303: Obtain the attention weight matrix of the time series sample according to the attention operation mechanism, feature vectors of segment samples, and segment sample information of the time series prediction model.

[0179] The method of this step has been described in the aforementioned step 103 and will not be elaborated here.

[0180] Step 304: Obtain the head attention feature of the time series sample according to the attention weight matrix of the time series sample.

[0181] Exemplarily, perform dot product attention calculation on the attention weight matrix of the time series sample and the value vector of the feature vector of the segment sample to obtain the head attention feature of the time series sample.

[0182] Exemplarily, refer to Figure 3 , and obtain the weight attention matrix corresponding to the feature vector of the segment sample based on the interpretable independent multi-head dot product attention layer. The interpretable independent multi-head dot product attention layer can effectively fuse the input feature vectors of segment samples, generate the attention weight matrix corresponding to the feature vectors of segment samples, and then independently output each head attention according to the attention weight matrix.

[0183] The multi-head attention is interpretable. The head attention features of the interpretable independent multi-head attention are positively correlated with the importance of the corresponding segment sample features, and can be used to characterize the relative importance of the segment sample features.

[0184] Exemplarily, the head attention is obtained through the attention operation mechanism based on the Transformer structure. Among them, compared with the Transformer structure in the related art, the present embodiment makes the following improvements to the Transformer structure: abandoning the residual connection structure.

[0185] In order to accurately calculate the importance of variables, it is necessary to adjust the weights of the attention values according to the gradient of the attention matrix and the model structure. The Transformer model in the related art has a residual connection layer. Due to the stacking of the residual connection and the attention layer, it is very difficult to calculate the weights, making it difficult for the method based on attention to interpret the neural network. In the present embodiment, the interpretable independent multi-head dot product attention mechanism abandons the residual connection, ensuring that the input data with low attention values cannot propagate forward through the residual connection, making the data screening based on the attention mechanism more reliable. That is, after abandoning the residual connection structure, the output of the attention mechanism is no longer added to the original input. Thus, it can be ensured that only the fused features adjusted by the attention weights can participate in the subsequent calculation process of the network model.

[0186] By abandoning the residual connection through the interpretable independent multi-head dot product attention mechanism, the time series prediction model of the present embodiment can easily filter out unnecessary sequence segments, fundamentally reducing the length of the input sequence and significantly reducing the computational cost of long-term prediction. In addition, the time series prediction model of the present embodiment can avoid the errors caused by loop calculations and improve the calculation speed.

[0187] Exemplarily, referring to Figure 3 , the discrete sequence segment samples are subjected to feature extraction through a one-dimensional convolutional layer to obtain query vectors, value vectors, and key vectors.

[0188] Preferably, the kernel function width of the one-dimensional convolutional layer is 1 and the stride is 1. Based on the kernel function with a width of 1 and a stride of 1, the query vectors, value vectors, and key vectors obtained by the one-dimensional convolutional layer can correspond one-to-one with the mathematical meanings of the segment sample features obtained through the interpretable discrete feature mapping layer. Based on the query vectors, value vectors, and key vectors obtained by the one-dimensional convolutional layer, they correspond one-to-one with the sequence segment samples input to the time series prediction model.

[0189] Step 305, according to the head attention features, obtain the first time series sample prediction result of the time series sample.

[0190] Exemplarily, the head attention feature input and the predictor corresponding to the time series sample are used to obtain the first time series sample prediction result corresponding to the segment sample feature.

[0191] Exemplarily, the predictor consists of two fully connected layers. Figure 3 An independent forward channel layer is shown. Figure 3 The forward channels in it constitute the predictor in this step. Among them, each forward channel predicts the segment sample feature of a single dimension according to the input head attention feature, and obtains the first time series sample prediction result corresponding to the segment sample feature of the single dimension.

[0192] Among them, the segment sample features of all sequence segment samples of a time series sample correspond to the same dimension. In other words, a segment sample feature of one dimension includes the segment sample features of all sequence segment samples belonging to the same time series sample.

[0193] For example, if the time series sample is the humidity data sample within 24 hours, and the collection time interval of the humidity data sample is 4 hours, then there are 6 humidity data samples, and the sequence length of the time series sample is 6. Refer to Figure 2 , and discretize the time series sample with a sequence length of 6 according to a sequence length equal to 3, 3 sequence segment samples can be obtained. The segment sample features of these 3 sequence segment samples correspond to one dimension, and this dimension is humidity.

[0194] Furthermore, refer to Figure 3 , Figure 3 The forward channels in it are predictors for predicting the time series prediction result. The independent forward channel layer consists of multiple independent forward channels. The number of forward channels is the same as the number of head attention features, and the forward channels and the head attention features correspond one by one. The number of forward channels is the same as the dimension of the segment sample feature, and the forward channels and the dimension of the segment sample feature correspond one by one. Exemplarily, each independent forward channel consists of two fully connected layers.

[0195] The independent prediction channels correspond one by one to the head attention features and are used to predict the first time series sample prediction result of the time series sample of a single dimension according to the head attention features.

[0196] Step 306, splice the first time series sample prediction results of at least one time series sample to obtain the time series sample prediction result.

[0197] Based on the segment sample features of one dimension, obtain the attention weight matrix of the time series sample corresponding to the segment sample features of this dimension. According to the attention weight matrix, obtain the head attention features of the time series sample. Input the head attention features independently output from the segment sample features of one dimension into the predictor (forward channel) to obtain the first time series sample prediction result of the time series sample. In this embodiment, instead of coupling the output results of all head attention features into a single output result and then making a prediction, for the segment sample features of one dimension, its attention weight matrix is obtained separately, and then according to the attention weight matrix and the value vector, a head attention feature of the segment sample features is obtained. Thus, the importance degree of the segment sample features in the value vector to the prediction result can be clarified.

[0198] Furthermore, according to the method of this embodiment, when predicting multi-dimensional time series samples, for each dimension, an attention weight matrix can be obtained. The attention weight matrix can directly represent the importance degree of the corresponding segment sample features and the sequence segment samples corresponding to the segment sample features to predicting the segment sample features of this dimension. Based on the method of this embodiment, the attention weight matrices of all head attention can be obtained, and through one prediction, the importance degree of all sequence segment samples to the sequence segment samples of any one dimension can be obtained.

[0199] Refer to Figure 3 , after obtaining the first time series sample prediction results of the segment sample features of at least one dimension, splice and fuse the first time series sample prediction results of all dimensions to finally obtain the multi-dimensional time series sample prediction result.

[0200] For example, the time series sample includes multiple dimensions. The head attention feature obtained from the time series sample of the 1st dimension is A1, the head attention feature obtained from the time series sample of the hth dimension is Ah, and the head attention feature obtained from the time series sample of the Hth dimension is AH. Input the head attention features A1 to AH into their respective corresponding forward channels to obtain their respective first time series sample prediction results, and splice all the first time series sample prediction results to obtain the predicted multi-dimensional time series.

[0201] Step 307, train the time series prediction model with the time series sample prediction result to obtain the trained time series prediction model.

[0202] The method of this step has been described in the foregoing step 105 and will not be elaborated here.

[0203] According to this embodiment, and Figure 3The schematic diagram of the model structure shown can extract the segment sample features of multi-dimensional time series samples through an interpretable discrete feature mapping layer. Through the segment sample features of multi-dimensional time series samples, the inherent interpretability of the time series prediction model can be achieved.

[0204] Referring to Figure 3 , the interpretable discrete feature mapping layer first divides the single-dimensional or multi-dimensional long-term time series samples into discrete sequence segment samples. The basic unit of the discrete sequence segment samples is the data samples within the specified length in a single dimension.

[0205] To preserve the context information, the data in the discretized sequence segment samples is connected end to end. The discretization process is as Figure 2 shown. The discretization process has been described in the foregoing embodiments and will not be elaborated here.

[0206] Based on the discrete sequence segment samples, segment sample features and segment sample information are obtained. Based on the segment sample features and segment sample information, a weight attention matrix containing segment sample information is obtained. According to the weight attention matrix, the interpretability of the model is achieved. The interpretability of the time series prediction model in this embodiment is realized through the attention weight matrix generated when predicting columns, which can measure the importance of all input time series segment samples. In the attention weight matrix, the attention weight is positively correlated with the importance of the input sequence segment samples.

[0207] Based on the structure of the inherent interpretability of the interpretable time series prediction model, during the process of using the model for time series prediction, the data flow of the time series prediction model is clear and definite, and the importance of the input sequence segments can be directly measured according to the dot product attention.

[0208] The time series prediction model of this embodiment has an independent multi-head attention structure. Based on the independent multi-head attention structure, each head attention can be decoupled. Thus, during the model training and prediction processes, without performing gradient correction on the time series model, the importance of all input sequence segment samples can be measured when predicting any dimension of the time series samples of the time series prediction model. The time series prediction model proposed in this embodiment can obtain the importance of all input sequence segment samples when predicting any dimension of the time series samples in a single prediction, and then explain the prediction of any dimension.

[0209] Furthermore, in this embodiment, the fragment sample features with clear mathematical meanings are processed by an interpretable independent multi-head dot product attention mechanism. Based on the independent multi-head attention mechanism, while effectively fusing the input fragment sample features, an attention weight matrix can be generated to represent the importance degree of each fragment sample feature in the network model. The interpretable independent multi-head dot product attention mechanism independently inputs the fused features adjusted by the attention weights represented by each head attention into the forward channels corresponding to the time series samples one by one, so as to predict the time series samples of a single dimension through each forward channel respectively. Finally, all the prediction results of the single dimension (the first time series sample prediction results) are concatenated into the final multi-dimensional prediction results (the time series sample prediction results).

[0210] Based on the time series prediction model of the independent multi-head attention mechanism in this embodiment, data can be obtained clearly and simply. Based on the time series prediction model with this structure, the obtained attention values can directly reflect the importance of variables without complex correction.

[0211] The embodiment of the present application also provides a time series prediction method. Refer to Figure 7 , the method may include the following steps:

[0212] Step 401, obtain the target time series to be predicted.

[0213] Among them, the target time series is obtained by sorting multiple data according to the data collection time.

[0214] Data collected at different times can be obtained, and multiple data are arranged in chronological order to form the target time series.

[0215] Step 402, perform discretization processing on the target time series to obtain multiple target sequence segments;

[0216] Among them, the target sequence segment includes at least one data.

[0217] For the method of performing discretization processing on the target time series to obtain multiple target sequence segments, reference may be made to the method description of performing discretization processing on the time series samples in the foregoing step 101, which will not be elaborated here.

[0218] Step 403, input the target sequence segment into the trained time prediction model to obtain the time series sample prediction result.

[0219] Among them, the trained time prediction model is obtained by the acquisition method of the time series prediction model in any of the foregoing embodiments.

[0220] Exemplarily, step 403 may include the following steps:

[0221] Step 4031: Obtain the segment features of the target sequence segment.

[0222] Exemplarily, through a one-dimensional convolutional layer, obtain the segment features of the target sequence segment.

[0223] Sub-step 4032: According to the screening matrix, screen out the target segment features from multiple segment features.

[0224] Exemplarily, the screening matrix is a matrix with elements of 0 or 1. Perform a multiplication operation on the screening matrix and the matrix of segment features to obtain the target segment features.

[0225] Sub-step 4033: According to the target segment features and the trained time series prediction model, obtain the time series sample prediction result for the target time series.

[0226] The screening matrix is obtained through the screening matrix acquisition method of the foregoing embodiment. According to the attention mechanism of the time series prediction model, obtain the time series sample prediction result for the target time series.

[0227] Screen out the target segment features from multiple segment features through the screening matrix, and based on the target segment features and the time series prediction model, obtain the time series sample prediction result for the target time series, reducing the data processing volume and improving the prediction efficiency of the model for long time series.

[0228] In this embodiment, the target time series is discretized to obtain target sequence segments, and the target sequence segments are input into the trained time prediction model to obtain the time series sample prediction result for the target time series. The time series prediction model obtained according to the method of the present application has the characteristics of interpretability and high prediction result accuracy. Therefore, based on the time prediction result obtained in this embodiment, it has the characteristic of high accuracy.

[0229] Next, taking the time series composed of ETT data as the time series sample as an example, the time series prediction model of the embodiment of the present application is further exemplarily described. In this embodiment, the public ETT dataset is used to verify the accuracy and interpretability of the interpretable time series prediction model proposed by the present application when predicting long-term time series. Among them, the ETT dataset records the power system data collected from two counties within two years. The dataset includes a 1-hour-level sub-dataset: the sub-dataset ETTH1 reflecting the power system data of the first county, and the sub-dataset ETTH2 reflecting the power system data of the second county. The dataset also includes a 15-minute-level sub-dataset ETTM1, and ETTM1 is used to reflect the power system data of one of the counties.

[0230] The ETT dataset is a seven-dimensional time series, where each dimension of the time series corresponds to a type of data, and each type of data has a corresponding acquisition time. For each preset time, seven types of data are collected. For each type of data, the data is arranged in chronological order to obtain a one-dimensional time series.

[0231] Specifically, each data point has six power load characteristics and a target value. The six power loads are the high-side useful load (HUFL), the high-side useless load (HULL), the medium-side useful load (MUFL), the medium-side useless load (MULL), the low-side useful load (LUFL), and the low-side useless load (LULL), and the target value is the oil temperature (OT). The time series samples are divided into a training set, a validation set, and a test set. Among them, the training set, the validation set, and the test set contain ETT data within 12 months, ETT data within 4 months, and ETT data within 4 months, respectively. When training the time series model, the Adam optimizer is used to optimize the interpretable time series prediction model. Preferably, the learning rate of the Adam optimizer is set to 0.001, and the number of samples processed at one time is 64.

[0232] When performing time series prediction through the forward channel, two evaluation metrics are used for each prediction window, namely the mean squared error (MSE) and the mean absolute error (MAE). Further, when predicting a multi-dimensional time series, the time series prediction results for each dimension are obtained respectively, and the MSE and MAE of the time series prediction results are obtained. The MSE of the time series prediction results for multiple dimensions is averaged to obtain the average MSE of all dimensions, and the MAE of the time series prediction results for multiple dimensions is averaged to obtain the average MAE of all dimensions, and all predictions are tested by rolling with a step size of 1. When analyzing the accuracy of the prediction results, the average MSE of all dimensions and the average MAE of all dimensions are used for analysis.

[0233] Based on the time series prediction model training method in the foregoing embodiments, after obtaining an interpretable time series prediction model, the interpretable time series prediction model is used to perform single-dimensional and multi-dimensional long-period time series predictions on the ETT dataset respectively, and compare them with the following time series prediction algorithms: Autoregressive Integrated Moving Average Model (ARIMA), Prophet prediction model, Probabilistic Forecasting with Autoregressive Recurrent (DeepAR) structure, Neural Basis Expansion Analysis for Interpretable Time Series Forecasting (N-Beats) model, Informer, LogSparse Transformer (LogTrans), Reversible Transformer (Reformer) model, Long Short Term Memory-a (LSTMa), Long and Short-term Time-series Networks (LSTNet) model, Temporal Convolutional Network (TCN) model, Pyraformer (a network model based on PyTorch), AuToFormer, and SCINet. Among them, based on each prediction model, the prediction errors of single-dimensional time series predictions for different prediction periods in the ETTH1 and ETTH2 datasets are shown in Table 1:

[0234]

[0235] The prediction errors of single-dimensional time series predictions for different prediction periods in part of the ETTH2 dataset and the ETTM1 dataset are shown in Table 2:

[0236]

[0237] Among them, based on each prediction model, the prediction errors of multi-dimensional time series predictions for different prediction periods in the ETTH1 and ETTH2 datasets are shown in Table 3:

[0238]

[0239] The prediction errors for the multi-dimensional sequences in the ETTH2 dataset and the ETTM1 dataset under different prediction periods are shown in Table 4 as follows:

[0240]

[0241] According to the error results shown in Tables 1 to 4, among all the models, the total error of the time series prediction model obtained through this embodiment is the smallest, and the number of times of obtaining the minimum error is the most. In other words, compared with other time series prediction models in the related art, the time series prediction model obtained through the method of this embodiment has a higher prediction accuracy on the basis of realizing interpretability.

[0242] In this embodiment, one prediction result is selected from the ETT dataset for specific analysis. Among them, the ETT dataset is a seven-dimensional time series, and the length of the time series for each dimension is 256, that is, the time series for each dimension includes 256 data, and each data corresponds to a collection moment.

[0243] The time series data with a sequence length of 256 is input into the interpretable time series prediction model to predict the time series with a sequence length of 720. That is, in this embodiment, based on the data at 256 moments in the historical time period, the data at 720 moments in the future time is predicted.

[0244] Exemplarily, the multi-dimensional time series data is discretized into a set of multiple single-dimensional sequence segments with a sequence length of 8. And the time series prediction results of the time series data for each dimension, as well as the attention weight matrix of the time series data for each dimension, are obtained respectively.

[0245] According to the time series data of the seven prediction dimensions, seven attention weight matrices can be obtained. The size of each attention weight matrix is 7 rows and 32 columns, and the attention weight matrix is used to measure the importance of 32 sequence segments of the seven input dimensions. Exemplarily, the attention weight matrices for predicting the time series data of the seven dimensions are shown in Figure 8. Among them, Figure 8(a) is the attention weight matrix of parameter 1, Figure 8(b) is the attention weight matrix of parameter 2, Figure 8(c) is the attention weight matrix of parameter 3, Figure 8(d) is the attention weight matrix of parameter 4, Figure 8(e) is the attention weight matrix of parameter 5, Figure 8(f) is the attention weight matrix of parameter 6, and Figure 8(g) is the attention weight matrix of parameter 7. Among them, in Figures 8(a) to 8(g), the horizontal axis is the preset moment, and the vertical axis is the heat value of the attention weight.

[0246] Among them, parameters 1 to 7 are respectively: HUFL, HULL, MUFL, MULL, LUFL, LULL, OT.

[0247] In FIGS. 8(a) to 8(g), the higher the brightness of the color block, the greater the attention weight corresponding to the color block, and the higher the importance of the parameter at the preset moment corresponding to the color block to the prediction result. For example, referring to FIG. 8(a), the brightness of the color block corresponding to parameter 6 at the 6th preset moment is the highest, indicating that the data sample of parameter 6 at the 6th preset moment has the highest importance to the prediction result of parameter 1.

[0248] Furthermore, the attention weight matrix when predicting the OT sequence of the prediction variable is corresponded to the time series data of seven dimensions, and a schematic diagram of the interpretability analysis result of the input time series data shown in FIG. 9 is obtained, as well as Figure 10 a comparison chart of the predicted curve and the target curve of OT shown in FIG. 5. Among them, FIG. 9(a) is a schematic diagram of the interpretability analysis result of parameter 1, FIG. 9(b) is a schematic diagram of the interpretability analysis result of parameter 2, FIG. 9(c) is a schematic diagram of the interpretability analysis result of parameter 3, FIG. 9(d) is a schematic diagram of the interpretability analysis result of parameter 4, FIG. 9(e) is a schematic diagram of the interpretability analysis result of parameter 5, FIG. 9(f) is a schematic diagram of the interpretability analysis result of parameter 6, and FIG. 9(g) is a schematic diagram of the interpretability analysis result of parameter 7.

[0249] Among them, FIG. 9 can reflect the time series segments that play a key role and the time series segments that are irrelevant or have little impact on the prediction result in the input time series segment when the model predicts the specified time series data. Based on FIG. 9, the interpretation result of one prediction using the time series model can be obtained. Based on the interpretation result presented in FIG. 9, the operation logic and data basis of the time series prediction model can be analyzed. Among them, the horizontal axis in FIG. 9 represents the moment, and the vertical axis represents the heat value corresponding to the attention weight. The brighter the color block, the greater the attention weight, and the higher the importance of the parameter to the prediction result.

[0250] According to FIG. 9, the time series prediction model assigns higher attention weights to several obvious segment features. The analysis of the time series segments corresponding to these segment features is as follows: referring to the trough interval of the input parameter 1 (HUFL) sequence in FIG. 9(a); referring to FIG. 9(b), the two segments where the input parameter 2 (HULL) sequence first fluctuates upward and downward; referring to FIG. 9(d), the segment where the input parameter 4 (MULL) sequence suddenly drops; referring to FIG. 9(g), the fluctuation period of the input parameter 7 (OT) sequence. Based on the inherent interpretability of the interpretable time series prediction model, data analysis can obtain the potential data rules in the process of processing data by the time series prediction model. The heat map shown in FIG. 9 indicates the inherent interpretability of the interpretable time series prediction model, which has the characteristics of stability and reliability.

[0251] Based on Figure 10It can be seen that the error between the predicted curve and the true curve obtained by the time series prediction model of this embodiment is small, indicating that the prediction result accuracy of the time series prediction model of this embodiment is high.

[0252] In this embodiment, the intrinsic interpretability of the time series prediction model is exploited during single prediction, and the analysis of the overall data can also be achieved by statistically analyzing the attention weight matrices in all prediction experiments.

[0253] Specifically, the attention weight matrices obtained when predicting the entire dataset are accumulated to obtain the globally statistically attention matrix shown in FIG. 11.

[0254] FIG. 11 shows 7 attention weight matrices. Among them, FIG. 11(a) is the attention weight matrix of parameter 1, FIG. 11(b) is the attention weight matrix of parameter 2, FIG. 11(c) is the attention weight matrix of parameter 3, FIG. 11(d) is the attention weight matrix of parameter 4, FIG. 11(e) is the attention weight matrix of parameter 5, FIG. 11(f) is the attention weight matrix of parameter 6, and FIG. 11(g) is the attention weight matrix of parameter 7. In FIGS. 11(a) to 11(g), the horizontal axis represents the time moment, and the vertical axis represents the heat value of the attention weight. Among them, the higher the brightness of the color block, the greater the attention weight corresponding to the color block, and the higher the importance of the parameter at the preset time moment corresponding to the color block to the prediction result.

[0255] The size of each attention weight matrix is 7 rows and 32 columns, and the attention weight matrix is used to measure the importance of 32 segments of 7 dimensions of the input. According to the globally statistically attention weight matrix shown in FIG. 11, the time series prediction model extracts features from equally spaced time series, and this characteristic conforms to the characteristics of time series. For example, referring to FIG. 11(a), the segments with relatively high importance of parameter 3 to the prediction result of parameter 1 are the segments with relatively high brightness of the color blocks in the figure. These segments are parameter 3 at the 4th moment, the 15th moment, and the 23rd moment, approximately showing a periodic change in importance, and the change period is 9, indicating that the change of the attention weight is related to the extraction moment of the segment sample features and shows periodicity as the extraction moment changes.

[0256] To facilitate the evaluation of the relationships between various data types, the total attention weights of each data type are counted one by one when predicting a specified dimension. The statistical results are shown in Figure 12. Among them, Figure 12(a) shows the importance of each parameter to Parameter 1, Figure 12(b) shows the importance of each parameter to Parameter 2, Figure 12(c) shows the importance of each parameter to Parameter 3, Figure 12(d) shows the importance of each parameter to Parameter 4, Figure 12(e) shows the importance of each parameter to Parameter 5, Figure 12(f) shows the importance of each parameter to Parameter 6, and Figure 12(g) shows the importance of each parameter to Parameter 7. In Figures 12(a) to 12(g), the horizontal axis represents time, and the vertical axis represents the heat value of the attention weight.

[0257] Based on the statistical results shown in Figure 12, the roles of various data in predicting a certain data sequence can be specifically analyzed. For example, referring to Figure 12(g), the heat values of the attention weights corresponding to Parameter 3 (MUFL), Parameter 4 (MULL), and Parameter 7 (OT) are the largest. Therefore, according to Figure 12(g), it can be known that the variables MUFL, MULL, and OT are the most important for the prediction of OT. In other words, the prediction of the variable OT mainly depends on the variables MUFL, MULL, and OT. The statistical results shown in Figure 12 provide an important reference for further exploring the internal relationships of the data set.

[0258] Based on the inherent interpretability of the time series prediction model, the relationships between data types can be analyzed qualitatively and quantitatively. Among them, in addition to providing a reference for data analysis, the inherent interpretability can also effectively screen data.

[0259] Specifically, from the set of discretized time series segment sets, a subset of input sequence segments with high attention weights is screened. By performing time series prediction based on the subset of input sequence segments with high attention weights, the length of the input time series can be fundamentally reduced, thereby significantly reducing the computational cost of long-term prediction. Based on the inherent data screening function of the time series prediction model, the time series prediction model can reduce the computational cost when predicting long-period sequences. In addition, based on the internal discretized data structure of the time series prediction model, the negative impact of masked data can be minimized.

[0260] For example, in the global statistical attention weight matrix for predicting each dimension, the positions of the most critical specified number of sequence segments are obtained, and the weights of these positions are set to 1, while the weights of the remaining positions are set to 0 to obtain a screening matrix. When screening data according to the screening matrix, the sequence segments at the positions with a weight of 1 are retained, and the segments at the positions with a weight of 0 are no longer input into the prediction model.

[0261] By screening data, retraining, and testing the prediction model, the final time series prediction model is obtained. Compared with the original time series prediction model, the computational cost of the prediction model based on the screened data is significantly reduced. Figure 13 The screening matrix is shown when screening the original 7×32 sequence segments of the dataset ETTH2. The horizontal axis represents the time, and the vertical axis represents the heat value of the attention weights. Among them, the color blocks with high brightness correspond to large weights, and the color blocks with low brightness correspond to small weights. For example, the weight corresponding to the color block with high brightness is 1, and the weight corresponding to the color block with low brightness is 0.

[0262] According to Figure 13 It can be seen that when predicting the 7 dimensions corresponding to the 7 parameters from parameter 1 to parameter 7, parameter 1 has the greatest importance at the 7th, 9th, 12th, 18th, 20th, 23rd, and 29th to 31st moments. Then, when performing time series prediction, the segment sample features of these moments of parameter 1 are screened out, the time series prediction samples are trained, and in the subsequent process of model inference, the time series data of these moments are used for time series prediction.

[0263] In this embodiment, the interpretable time series prediction model only relies on the screened data for long-term sequence prediction, and it is marked as the interpretable time series prediction model (screening). The interpretable time series prediction model relies on 24-length multi-dimensional data for long-term sequence prediction, and it is marked as the interpretable time series prediction model (short input). The prediction errors of different time series prediction models are obtained, and the prediction errors are shown in Table 3:

[0264]

[0265] According to Table 5, it can be seen that the prediction error of the time series prediction model based on the screened data is very close to the prediction error of the original time series prediction model, and is much smaller than the prediction error when only inputting 24-length data. However, the computational cost of the prediction model based on the screened data is smaller than that of the model that only inputs 24-length data.

[0266] According to the prediction error comparison results shown in Table 5, by screening the prediction model data, the accuracy of the prediction results can be improved while ensuring the interpretability of the time series model. In other words, the interpretable time series prediction model obtained in this embodiment can use the data screening function to predict long-period time series at low cost and with high accuracy of the prediction results.

[0267] In the related art, the SCINet model uses a fully convolutional network without a dot product attention structure and has a low computational cost. However, the SCINet model is not interpretable and has a large prediction error when performing long-term sequence prediction. The AuToFormer model uses a dot product attention structure, but the computational consumption of the dot product attention structure is proportional to the square of the input data length.

[0268] The maximum data length of the time series prediction model obtained in this application is , and the maximum data length of the input dot product attention structure of the AuToFormer model .

[0269] Among them, represents the time length of the input data, represents the prediction time length, represents the dimension of the original data, represents the unit time length of the proposed interpretable time series prediction model.

[0270] The maximum data length of the time series prediction model obtained in this application is smaller than the maximum data length of the input dot product attention structure of the AuToFormer model. Therefore, compared with the AuToFormer model, the time series prediction model in this embodiment has a small computational consumption.

[0271] The time series prediction model of this application has a data screening function, which can shorten the length of the input data to an extremely short value, greatly reducing the computational cost. Therefore, when predicting long-term sequences, the computational amount of the time series prediction model of this application is significantly lower than that of the AuToFormer model, and data screening can even be achieved according to requirements, extremely compressing the computational cost.

[0272] In the embodiment of this application, a new interpretable long-term time series prediction model is built through discrete feature mapping, interpretable independent multi-head dot product attention, and independent forward channels; the proposed interpretable long-term time series prediction model surpasses many current excellent sequence prediction algorithms in terms of time series prediction accuracy; the proposed interpretable long-term time series prediction model utilizes its interpretable advantage to perform importance analysis on the input data and reveals the reasons for the model output prediction results; according to the importance analysis results of the input data by the proposed interpretable long-term time series prediction model, screening of the original data is realized, and a method for fundamentally reducing the computational complexity of the time series prediction algorithm is proposed.

[0273] The present invention is an exploration and improvement of the Transformer structure for time series prediction, which can provide an important basic technology for the improvement of large language models in the future

[0274] Exemplarily, the improvement of the Transformer structure in the time series prediction task in the present invention can be applied to a large language model to implement an interpretable large language model with improved performance.

[0275] Figure 14 FIG. is a schematic diagram of a time series prediction model training device provided by an embodiment of the present disclosure. Referring to Figure 14 , the device 50 includes:

[0276] A first acquisition module 501, configured to perform discretization processing on time series samples to obtain a plurality of sequence segment samples;

[0277] A second acquisition module 502, configured to acquire segment sample features and segment sample information of the sequence segment samples;

[0278] A third acquisition module 503, configured to obtain an attention weight matrix of the time series samples according to the attention operation mechanism of the time series prediction model, segment sample features, and segment sample information;

[0279] A fourth acquisition module 504, configured to obtain a time series sample prediction result according to the attention weight matrix;

[0280] A fifth acquisition module 505, configured to train the time series prediction model through the time series sample prediction result to obtain a trained time series prediction model.

[0281] Optionally, the fourth acquisition module 504 may include:

[0282] A first acquisition sub-module, configured to acquire target attention weights greater than or equal to a preset weight threshold from the attention weight matrix;

[0283] A second acquisition sub-module, configured to acquire target sequence segment sample features corresponding to the target attention weights from a plurality of segment sample features;

[0284] A third acquisition sub-module, configured to input the target sequence segment sample features into the time series prediction model to obtain a time series sample prediction result.

[0285] Optionally, the second acquisition sub-module includes:

[0286] A first configuration unit, configured to configure a first weight for the target attention weights and a second weight for other attention weights except the target attention weights;

[0287] A first construction unit, configured to construct a screening matrix according to the first weight, segment sample information corresponding to the segment sample features corresponding to the first weight, the second weight, and segment sample information of segment sample features for calculating other attention weights;

[0288] The first screening unit is configured to screen out the target sequence fragment sample features from multiple fragment sample features according to a screening matrix.

[0289] Optionally, there is at least one time series sample; the first acquisition module 501 may include:

[0290] The fourth acquisition sub-module is configured to perform discretization processing on each time series sample respectively to obtain at least one set of sequence fragment samples;

[0291] The second acquisition module 502 includes:

[0292] The fifth acquisition sub-module is configured to perform feature extraction on at least one set of sequence fragment samples to obtain the fragment sample features of the sequence fragment samples.

[0293] Optionally, the fourth acquisition module 504 includes:

[0294] The sixth acquisition sub-module is configured to obtain the head attention features of the time series sample according to the attention weight matrix of the time series sample;

[0295] The seventh acquisition sub-module is configured to obtain the first time series sample prediction result of the time series sample according to the head attention features;

[0296] The eighth acquisition sub-module is configured to splice the first time series sample prediction results of at least one time series sample to obtain the time series sample prediction result.

[0297] Optionally, the sixth acquisition sub-module includes:

[0298] The first acquisition unit is configured to perform dot product attention calculation on the attention weight matrix of the time series sample and the value vector of the fragment sample features to obtain the head attention features of the time series sample; wherein, the value vector is obtained by performing convolution processing on the fragment sample features.

[0299] Optionally, the seventh acquisition sub-module includes:

[0300] The second acquisition unit is configured to input the head attention features into the predictor corresponding to the time series sample to obtain the first time series sample prediction result corresponding to the fragment sample features.

[0301] Optionally, the first acquisition module 501 includes:

[0302] The ninth acquisition sub-module is configured to perform discretization processing on the time series sample according to a preset sequence length to obtain sequence fragment samples, wherein, among two adjacent sequence fragment samples, the first quantity of data samples at the end of the first sequence fragment sample is the same as the first quantity of data samples at the beginning of the second sequence fragment sample.

[0303] Optionally, the sample information includes at least one of the following: the sample identifier of the sequence fragment sample, and the time information of the sequence fragment sample.

[0304] Optionally, the third acquisition module 503 includes:

[0305] The tenth acquisition sub-module is configured to perform convolution processing on the fragment sample features to obtain a query vector and a key vector;

[0306] The tenth acquisition sub-module is configured to obtain an attention weight matrix according to the query vector, the key vector, and the fragment sample information.

[0307] In summary, in this embodiment, based on the attention weights in the attention weight matrix and the sample information of the fragment sample features used to calculate the attention weights, the importance of each sequence fragment sample in the input time series prediction model can be quantified, and the importance degree of the sequence fragment sample corresponding to the sample information can be clearly obtained for the time series prediction result. In other words, based on the attention weight matrix containing the sample information of the fragment sample features used to calculate the attention weights, the model interpretation of the time prediction model is realized. The trained time series prediction model obtained based on the attention weight matrix has the predictability for predicting long-term time series.

[0308] Figure 15 is a schematic diagram of a time series prediction device provided by an embodiment of the present disclosure. Refer to Figure 15 , the device 60 includes:

[0309] The sixth acquisition module 601 is configured to acquire a target time series to be predicted;

[0310] The seventh acquisition module 602 is configured to perform discretization processing on the target time series to obtain a plurality of target sequence fragments;

[0311] The eighth acquisition module 603 is configured to input the target sequence fragment into the trained time prediction model to obtain a time series prediction result; wherein, the trained time prediction model is obtained by the time series prediction model acquisition method in the foregoing embodiment.

[0312] Exemplarily, the eighth acquisition module 603 includes:

[0313] The twelfth acquisition sub-module is configured to acquire the fragment features of the target sequence fragment;

[0314] The thirteenth acquisition sub-module is configured to screen out the target fragment features from the plurality of fragment features according to the screening matrix;

[0315] A fourteenth acquisition sub-module, configured to obtain a time series prediction result of a target time series according to the target segment feature and the trained time series prediction model;

[0316] The screening matrix is obtained by the screening matrix obtaining method in the foregoing embodiment.

[0317] In this embodiment, the target time series is discretized to obtain a target sequence segment, and the target sequence segment is input into the trained time prediction model to obtain a time series prediction result of the target time series. The time series prediction model obtained according to the method of the present application has the characteristics of interpretability and high prediction result accuracy. Therefore, the time prediction result obtained based on this embodiment has the characteristic of high accuracy.

[0318] This disclosure embodiment also provides an electronic device, as Figure 16 shown, including a processor 1101, a communication interface 1102, a memory 1103, and a communication bus 1104. Among them, the processor 1101, the communication interface 1102, and the memory 1103 communicate with each other through the communication bus 1104.

[0319] The memory 1103 is used to store a computer program.

[0320] When the processor 1101 is used to execute the program stored in the memory 1103, the following steps are implemented: obtaining a time series sample, and discretizing the time series sample to obtain a sequence segment sample; obtaining a segment sample feature and segment sample information of the sequence segment sample; according to the attention operation mechanism of the time series prediction model, the segment sample feature, and the segment sample information, obtaining an attention weight matrix of the time series sample; the attention weight matrix includes segment sample information; obtaining a time series sample prediction result according to the attention weight matrix; training the time series prediction model through the time series sample prediction result to obtain a trained time series prediction model.

[0321] Among them, the processor 1101 can also implement other steps in the above time series prediction model training method or time series prediction method, which will not be elaborated here.

[0322] The communication bus mentioned in the above electronic device may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. The communication bus can be divided into an address bus, a data bus, and a control bus. For the sake of simplicity, only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.

[0323] The communication interface is used for communication between the above electronic device and other devices.

[0324] The memory may include a Random Access Memory (RAM), and may also include a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.

[0325] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU) and a Network Processor (NP); it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0326] In another embodiment provided by the present disclosure, a computer-readable storage medium is also provided. Instructions are stored in the computer-readable storage medium. When it runs on a computer, it causes the computer to execute the time series prediction model training method or the time series prediction method in the above embodiments.

[0327] In another embodiment provided by the present disclosure, a computer program product containing instructions is also provided. When it runs on a computer, it causes the computer to execute the time series prediction model training method or the time series prediction method in the above embodiments.

[0328] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present disclosure are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center via wired (such as coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a server, data center, or other data storage device that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).

[0329] It should be noted that, in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0330] Each embodiment in this specification is described in a related manner. The same or similar parts among the embodiments can be referred to each other, and the differences between each embodiment and other embodiments are emphasized. For the embodiments of the device, electronic device, computer-readable storage medium, and computer program product including instructions, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiments.

[0331] The above are only the preferred embodiments of the present disclosure and are not intended to limit the protection scope of the present disclosure. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present disclosure are included in the protection scope of the present disclosure.

Claims

1. A time series prediction model training method, characterized in that: include: Acquire time series samples, and discretize the time series samples to obtain sequence fragment samples; the time series samples include power load data of equipment in the power system, and the power load data includes at least one of the following: effective load on the high voltage side, invalid load on the high voltage side, effective load on the medium voltage side, invalid load on the medium voltage side, effective load on the low voltage side, and invalid load on the low voltage side; Acquire fragment sample features and fragment sample information of the sequence fragment samples; According to the attention operation mechanism of the time series prediction model, the characteristics of the fragment samples, and the fragment sample information, an attention weight matrix of the time series samples is obtained; the fragment sample information includes at least one of the following: a sample identifier of the sequence fragment sample, and the moment information of the sequence fragment sample; the attention weight matrix includes the fragment sample information; Obtaining a time series sample prediction result according to the attention weight matrix; The time series prediction model is trained by using the time series sample prediction results to obtain a trained time series prediction model.

2. The method according to claim 1, characterized in that According to the attention weight matrix, the time series sample prediction result is obtained, including: Obtaining a target attention weight greater than or equal to a preset weight threshold from the attention weight matrix; Obtaining a target segment sample feature corresponding to the target attention weight from a plurality of segment sample features; The target segment sample features are input into the time series prediction model to obtain a time series sample prediction result.

3. The method according to claim 2, characterized in that Obtaining a target segment sample feature corresponding to the target attention weight from the plurality of segment sample features, including: Configuring a first weight for the target attention weight, and configuring a second weight for other attention weights except the target attention weight; constructing a screening matrix according to the first weight, the fragment sample information of the fragment sample feature corresponding to the first weight, the second weight, and the fragment sample information of the fragment sample feature used to calculate other attention weights; The target segment sample feature is screened out from a plurality of segment sample features according to the screening matrix.

4. The method according to claim 1, characterized in that The time series sample has at least one; the time series sample is discretized to obtain a sequence fragment sample, including: Discretizing each of the time series samples to obtain at least one group of sequence fragment samples; The obtaining of the fragment sample features of the sequence fragment samples comprises: Feature extraction is performed on at least one group of the sequence fragment samples to obtain fragment sample features of the sequence fragment samples.

5. The method according to claim 1, characterized in that The time series sample has at least one, and obtaining a prediction result of the time series sample according to the attention weight matrix includes: According to the attention weight matrix of the time series sample, a head attention feature of the time series sample is obtained; Obtaining a first time series sample prediction result of the time series sample according to the head attention feature; The first time series sample prediction result of at least one of the time series samples is concatenated to obtain the time series sample prediction result.

6. The method according to claim 5, characterized in that According to the attention weight matrix of the time series sample, the head attention feature of the time series sample is obtained, including: Performing dot product attention calculation on the attention weight matrix of the time series sample and the value vector of the segment sample feature to obtain the head attention feature of the time series sample; The value vector is obtained by performing convolution processing on the features of the fragment samples.

7. The method according to claim 5, characterized in that Obtaining a first time series sample prediction result of the time series sample according to the head attention feature, including: The head attention feature is input into a predictor corresponding to the time series sample to obtain a first time series sample prediction result corresponding to the segment sample feature.

8. The method according to claim 1, characterized in that Discretization processing is performed on the time series samples to obtain sequence fragment samples, including: According to a preset sequence length, the time series samples are discretized to obtain sequence fragment samples; Among two adjacent sequence segment samples, a first number of data samples located at the end of a first sequence segment sample is the same as a first number of data samples located at the beginning of a second sequence segment sample.

9. The method according to claim 1, characterized in that: According to the attention operation mechanism of the time series prediction model, the characteristics of the fragment samples, and the fragment sample information, an attention weight matrix of the time series samples is obtained, including: Performing convolution processing on the features of the fragment samples to obtain a query vector and a key vector; An attention weight matrix of the time series sample is obtained according to the query vector, the key vector, and the fragment sample information.

10. A time series prediction method, characterized in that: include: Acquire a target time series to be predicted; the target time series to be predicted includes power load data of equipment in the power system, the power load data includes at least one of the following: a high-voltage side effective load, a high-voltage side invalid load, a medium-voltage side effective load, a medium-voltage side invalid load, a low-voltage side effective load, and a low-voltage side invalid load; Discretize the target time series to obtain target sequence fragments; The target sequence segment is input into a trained time prediction model to obtain a time series prediction result; wherein the trained time prediction model is obtained by the time series prediction model training method according to any one of claims 1 to 9.

11. The method according to claim 10, characterized in that Input the target sequence fragment into the trained time prediction model to obtain the time series prediction result, including: Acquiring fragment features of the target sequence fragment; Screening out a target fragment feature from the plurality of fragment features according to a screening matrix; Obtaining a time series prediction result of a target time series according to the target segment features and the trained time series prediction model; Wherein, the screening matrix is ​​obtained by the method of claim 3.

12. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the method according to any one of claims 1 to 11.

13. A computer-readable storage medium, characterized in that: Used to store a computer program; wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 11 is implemented.

14. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the method according to any one of claims 1 to 11 is implemented.

Citation Information

Patent Citations

  • Time sequence prediction method based on A2form

    CN117574958A