A time series prediction method and system based on exception processing

By detecting abnormal fluctuations and adjusting their weights in time series prediction, combined with the ConvLSTM network and the dual attention mechanism, the problem of existing methods ignoring abnormal fluctuations and spatial relationships is solved, which significantly improves the accuracy of time series trend prediction.

CN119557802BActive Publication Date: 2025-06-06SHANDONG UNIV OF FINANCE & ECONOMICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411611098.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-12
Publication Date
2025-06-06
Estimated Expiration
2044-11-12

AI Technical Summary

Technical Problem

Existing time series prediction methods ignore the spatial relationship between abnormal fluctuations in time series data and multiple variables, resulting in low prediction accuracy.

Method used

The time series prediction method based on exception processing is adopted to detect abnormal fluctuations through the fluctuation point detection network, calculate its weight and adjust the weight of other data points. Then, feature extraction is performed using the ConvLSTM network and the influence weights of different time series data are adaptively adjusted through a dual attention mechanism.

Benefits of technology

Effectively weaken the impact of abnormal fluctuations on feature extraction, improve the comprehensiveness and accuracy of feature extraction of time series data, and thus improve the accuracy of time series trend prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119557802B_ABST
    Figure CN119557802B_ABST
Patent Text Reader

Abstract

The present invention proposes a time series prediction method and system based on exception processing, which uses abnormal fluctuation points to detect abnormal fluctuation points, calculates the weight of the abnormal fluctuation points according to the distance between the abnormal fluctuation points and other data points, and adjusts the weights of other data points according to the abnormal fluctuation point weights, which can effectively reduce the influence of abnormal fluctuation points on subsequent data feature extraction, thereby reducing feature extraction errors and improving overall prediction accuracy; uses a ConvLSTM network for feature extraction, and the ConvLSTM network can analyze the spatial relationship between multivariate variables on the basis of extracting the time relationship of time series data, which can significantly improve the problem of incomplete data feature information extracted only from the time relationship level, and improves prediction accuracy; proposes a dual attention network, which uses a self-attention mechanism to adaptively learn and weigh and calculate the influence weights of different time series, thereby improving prediction accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field related to time series prediction, and in particular relates to a time series prediction method and system based on exception processing. Background Art

[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.

[0003] Time series analysis has been widely used in practical problems in many fields such as electricity, transportation, environmental monitoring, weather forecasting, etc. It is mainly divided into accurate point prediction and trend prediction of time series data. Since accurate point numerical prediction is difficult, recent research focuses on the future trend of time series. Time series trend is to predict the future direction of time series data, rising, stable or falling.

[0004] Time series data is a data sequence with fixed time intervals and arranged in chronological order. Time series prediction can mine potential patterns from time series data and provide guidance for future decision-making. Existing time series prediction methods can be roughly divided into three categories: traditional methods, machine learning-based methods, and deep learning-based methods. The traditional time series prediction method is to model the time series data, and fit it based on historical time series data through autoregressive models (AR), auto-regressive moving average models (ARIMA) and other variants. Although traditional methods have certain prediction effects, they all assume that the data is linear and stable, which is contrary to the actual situation of high volatility, so they have great limitations and poor stability. Later, machine learning-based methods such as support vector machines (SVM) and logistic regression were applied to time series prediction, but the processing effect of high-dimensional and complex time series data was poor. In order to better consider the nonlinearity, non-stationarity, high-dimensional complexity and other characteristics of time series data, researchers began to apply deep learning methods to time series prediction. Among them, recurrent neural network (RNN), convolutional neural network (CNN), etc. can capture the potential characteristics of time series and perform well in time series prediction. Long short-term memory network (LSTM) is widely used as a special RNN. LSTM can learn and store the context information of time series and solve the problems of gradient disappearance and gradient explosion in long time series training. However, LSTM has spatial redundant data and does not consider spatial correlation. Similarly, CNN only obtains time features and ignores the spatial features between multivariate variables of time series data. Moreover, these methods ignore the mutual influence between data at different time points of time series data. In addition, these methods pay less attention to abnormal fluctuation points in time series data. Here, abnormal fluctuation points refer to points that rise or fall sharply in the process of changing data points related to the prediction task, that is, obviously abnormal change points. Obviously, the unified processing of abnormal fluctuation points and normal points will have a greater impact on the accuracy of time series prediction.

[0005] In summary, existing methods often ignore abnormal fluctuation points in time series data, ignore the different impacts of time series data at different time points, which affects the prediction accuracy, and only use the time relationship of time series data to ignore the spatial relationship between multivariate variables, which leads to problems such as incomplete feature extraction of time series data and easy information loss. Therefore, how to provide a time series trend prediction solution that can improve the accuracy of time series trend prediction is a technical problem that needs to be solved in this field. Summary of the invention

[0006] In order to overcome the above-mentioned deficiencies of the prior art, the present invention provides a time series prediction method and system based on exception processing, which improves the accuracy of time series prediction results.

[0007] In order to achieve the above object, the present invention adopts the following technical solution:

[0008] In a first aspect, the present invention provides a time series prediction method based on exception processing, comprising:

[0009] Obtain the historical time series data corresponding to the target prediction time;

[0010] The historical time series data corresponding to the target prediction time is input into the trained prediction model. When the prediction model processes the historical time series data corresponding to the target prediction time, the fluctuation point detection network of the prediction model is used to detect abnormal fluctuation points of the historical time series data, and the weight of the abnormal fluctuation point is calculated according to the distance between the abnormal fluctuation point and other data points, and the weight of other data points is adjusted according to the abnormal fluctuation point weight;

[0011] The ConvLSTM network of the prediction model is used to extract features from the adjusted historical time series data, and the dual attention mechanism is used to adaptively adjust the influence weights of different historical time series data to obtain the features of the time series data after weight adjustment.

[0012] Based on the time series data characteristics after weight adjustment, the prediction model is used to predict the time series prediction result corresponding to the target prediction time.

[0013] In a second aspect, the present invention provides a time series prediction system based on exception processing, comprising:

[0014] An acquisition module is configured to: acquire historical time series data corresponding to a target prediction time;

[0015] An anomaly detection module is configured to: input the historical time series data corresponding to the target prediction time into a trained prediction model, when the prediction model processes the historical time series data corresponding to the target prediction time, use the fluctuation point detection network of the prediction model to detect abnormal fluctuation points of the historical time series data, calculate the weight of the abnormal fluctuation point according to the distance between the abnormal fluctuation point and other data points, and adjust the weight of other data points according to the abnormal fluctuation point weight;

[0016] An adjustment module is configured to: extract features of the adjusted historical time series data using the ConvLSTM network of the prediction model, and adaptively adjust the influence weights of different historical time series data using a dual attention mechanism to obtain features of the time series data after weight adjustment;

[0017] The time series prediction module is configured to: based on the time series data characteristics after weight adjustment, use the prediction model to predict the time series prediction result corresponding to the target prediction time.

[0018] In a third aspect, the present invention provides an electronic device comprising a memory and a processor, and computer instructions stored in the memory and executed on the processor, wherein when the computer instructions are executed by the processor, the method described in the first aspect is performed.

[0019] In a fourth aspect, the present invention provides a computer-readable storage medium for storing computer instructions, wherein when the computer instructions are executed by a processor, the method described in the first aspect is performed.

[0020] In a fifth aspect, the present invention provides a computer program product, comprising a computer program, which, when executed by a processor, implements the method described in the first aspect.

[0021] One or more of the above technical solutions have the following beneficial effects:

[0022] In the present invention, abnormal fluctuation points are detected by using abnormal fluctuation points, and the weight of the abnormal fluctuation point is calculated according to the distance between the abnormal fluctuation point and other data points. The weight of other data points is adjusted according to the abnormal fluctuation point weight, which can effectively reduce the impact of abnormal fluctuation points on subsequent data feature extraction, thereby reducing feature extraction errors and improving overall prediction accuracy.

[0023] In the present invention, the ConvLSTM network is used for feature extraction. The ConvLSTM network can analyze the spatial relationship between multivariate variables based on the time relationship of time series data, which can significantly improve the problem of incomplete data feature information extracted only from the time relationship level and improve the prediction accuracy.

[0024] In the present invention, a dual attention network is proposed, which uses a self-attention mechanism to adaptively learn and weigh and calculate the influence weights of different time series to improve prediction accuracy.

[0025] Advantages of additional aspects of the present invention will be given in part in the following description, and in part will become obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.

[0027] Figure 1 This is a diagram of the overall network structure of the prediction model in Embodiment 1 of the present invention;

[0028] Figure 2 This is a graph showing the influence of different values ​​of ω on the accuracy of electricity price trend prediction in the first embodiment of the present invention;

[0029] Figure 3 This is a diagram showing the influence of different values ​​of θ on the accuracy of electricity price trend prediction in the first embodiment of the present invention. DETAILED DESCRIPTION

[0030] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.

[0031] It should be noted that the terms used herein are for describing specific embodiments only and are not intended to be limiting of exemplary embodiments according to the present invention.

[0032] In the absence of conflict, the embodiments of the present invention and the features of the embodiments may be combined with each other.

[0033] The time series prediction method in the embodiments of this specification can be applied to various fields, such as electricity, transportation, environmental monitoring, weather forecast, various Internet products and many other fields, such as: it can be used to predict indicators such as sales, number of users, order volume on the advertising platform, for example, by analyzing the impact of extreme events such as promotional activities and holidays on the advertising platform on sales, thereby predicting future sales. It can be used to predict traffic flow, for example, by analyzing the impact of daily traffic flow conditions on holidays and daily traffic flow conditions on non-holidays on traffic flow, thereby predicting future traffic flow. It can be used to predict indicators such as wind power and load, for example, by analyzing the impact of different weather, seasons, etc. on electricity consumption, thereby predicting future electricity load, etc. Of course, according to actual needs, it can also be used in other application scenarios.

[0034] Embodiment 1

[0035] This embodiment discloses a time series prediction method based on exception processing, including:

[0036] Obtain the historical time series data corresponding to the target prediction time;

[0037] The historical time series data corresponding to the target prediction time is input into the trained prediction model. When the prediction model processes the historical time series data corresponding to the target prediction time, the fluctuation point detection network of the prediction model is used to detect abnormal fluctuation points of the historical time series data, and the weight of the abnormal fluctuation point is calculated according to the distance between the abnormal fluctuation point and other data points, and the weight of other data points is adjusted according to the abnormal fluctuation point weight;

[0038] The ConvLSTM network of the prediction model is used to extract features from the adjusted historical time series data, and the dual attention mechanism is used to adaptively adjust the influence weights of different historical time series data to obtain the features of the time series data after weight adjustment.

[0039] Based on the time series data characteristics after weight adjustment, the prediction model is used to predict the time series prediction result corresponding to the target prediction time.

[0040] The prediction model of this embodiment consists of three parts: a time series data processing module, a time series data feature extraction module, and a time series trend prediction module. In the time series data processing module, the time series data is standardized and input into the abnormal fluctuation point detection network, and the influence of the abnormal fluctuation point on other data points is quantified and weighed through the abnormal fluctuation point weight; in the time series data feature extraction module, the time series data is extracted from the time relationship level and the spatial relationship level through the ConvLSTM network, and then the time dependency of the time series data is refined through the dual attention network, and the influence weights of different time series are adaptively learned and adjusted, and the weight-adjusted time series data features are output; in the time series trend prediction module, the time series features are input into the fully connected layer, and the predicted time series trend is output.

[0041] Combine the following Figure 1 The time series trend prediction method based on exception processing proposed in this embodiment is described in detail.

[0042] Step 1: Obtain the historical time series data corresponding to the target prediction time and perform preprocessing, use the fluctuation point detection network of the prediction model to detect the abnormal fluctuation points of the historical time series data, calculate the weight of the abnormal fluctuation point according to the distance between the abnormal fluctuation point and other data points, and adjust the weight of other data points according to the abnormal fluctuation point weight.

[0043] The target prediction time can be understood as the time when time series prediction is required. The target prediction time can be a moment or a time range, which can be determined according to actual needs. For example, if you need to predict the traffic flow in a certain place from October 1st to 7th, 2023, then October 1st to 7th, 2023 is the target prediction time. If you need to predict the weather conditions in a certain place on October 1st, 2023, then October 1st, 2023 is the target prediction time. Generally, the target prediction time refers to a time that has not yet occurred, that is, time series prediction is generally a prediction of data at a certain time in the future. Time series data is data collected at different times and is used to describe the situation where the phenomenon described changes over time. This type of data reflects the state or degree of change of a certain thing, phenomenon, etc. over time. Historical time series data can be understood as time series data between the times that need to be predicted. For example, if you need to predict the sales of a certain product in October 2023, then the historical time series data can be understood as the data related to the product before October 2023, such as: promotional activities of the product before October 2023, sales in October of previous years, sales of products of the same type as the product, etc.

[0044] The historical time series data corresponding to the target prediction time is standardized and preprocessed, and then the data is processed initially. Then, the abnormal fluctuation points in the historical time series data are detected through the abnormal fluctuation point detection network, and the abnormal fluctuation point weight matrix is ​​constructed using the distance between the abnormal fluctuation point and the prediction point. The farther the transaction point is from the abnormal fluctuation point, the less it is affected by the abnormal fluctuation point.

[0045] By constructing a weight matrix, the influence of abnormal fluctuation points on the feature extraction of time series data can be effectively weakened, the feature extraction error can be reduced, and the prediction accuracy can be improved.

[0046] Step 11: Standardization preprocessing.

[0047] Each attribute of all historical time series data is subjected to a separate min-max normalization calculation using the following formula, and the normalized data is then used to extract the features of the time series data.

[0048]

[0049] Among them, x t represents the data at time t, represents the data at time t after standardization, x min Indicates the minimum value of the data in all time series, x max Indicates the maximum value of the data in all time series.

[0050] Step 12: Abnormal fluctuation point detection network.

[0051] Historical time series data may experience abnormal fluctuations in the short term due to special circumstances. Such abnormal fluctuations are occasional special circumstances and are obviously not suitable for predicting the future trend of time series. Through a large number of analyses of historical time series data, it is found that historical time series data usually have multiple abnormal fluctuation points in the historical time, and the rise or fall of the time series data at this point is obviously beyond the normal range. In order to reduce the negative impact of abnormal fluctuation points on time series prediction, this embodiment proposes an abnormal fluctuation point detection network.

[0052] Specifically, if the fluctuation gap of the historical time series data at time t is t If the following formula is satisfied, the time point t is regarded as an abnormal fluctuation point, that is, the detection criteria of the abnormal fluctuation point are defined as follows:

[0053]

[0054] Among them, data t Represents the time series data at time t, data t-1 represents the time series data at time t-1, represents the average value of multiple data in the sliding window m, represents the standard deviation of multiple data in the sliding window m, and ω is the experimental coefficient that controls the detection threshold of abnormal fluctuation points.

[0055] After using the abnormal fluctuation point detection network to find abnormal fluctuation points in historical time series data, the abnormal fluctuation point weight is output to quantify and weigh the impact of the abnormal fluctuation point on other data points. The weight is calculated by the distance between the abnormal fluctuation point and other data points in the sliding window at the current time point. According to the principle that the closer the point is to the abnormal fluctuation point, the greater the impact, and the farther away from the abnormal fluctuation point, the smaller the impact. The quantitative calculation expression is:

[0056]

[0057] Among them, distance t,τ It represents the time difference between the data time node t and the time node τ. For example, if the collected power data is collected every hour, then this distance is the hour difference between t and τ. θ represents the adjustment coefficient of the influence of abnormal fluctuation points on other points. m is the size of the sliding window. Represents the standardized data point at time t.

[0058] It is understandable that through Calculate all data points adjusted by the abnormal fluctuation point detection network. For example, if the sliding window size is 5, then the combined impact of the first point, the second point, the third point, the fourth point, and the fifth point will be calculated. If the second point, the third point, the fourth point, and the fifth point are all normal points, then g t,τ All are 1, g t Also 1, constant.

[0059] Step 2: Use the ConvLSTM network of the prediction model to extract features from the adjusted historical time series data, and use the dual attention mechanism to adaptively adjust the influence weights of different historical time series data to obtain the features of the time series data after weight adjustment.

[0060] In the time series data feature extraction module, the ConvLSTM network is used to extract time series features from the temporal relationship level and the spatial relationship level. Then, the dual attention network is used to refine the time dependency of the time series data, and the influence weights of different time series are adaptively learned and adjusted to output the time series data features after weight adjustment.

[0061] The ConvLSTM network is used to extract time series data features. ConvLSTM combines convolutional neural networks (CNN) and recurrent neural networks (RNN), uses convolutional neural networks to replace the fully connected layer in the LSTM network, and extracts basic spatial features through convolution to obtain the spatiotemporal correlation of time series data. The biggest feature of time series data is the time label. Data with different time labels have different effects on the final prediction task. This embodiment proposes a dual attention network, which innovatively combines the hierarchical idea with the self-attention network. The time dependency of time series data is analyzed layer by layer in a layered manner. The self-attention mechanism is used to adaptively learn and adjust the influence weights of different time series, further improving the accuracy of time series prediction tasks.

[0062] Step 21: The ConvLSTM network extracts time series data features from the temporal relationship level and the spatial relationship level.

[0063] ConvLSTM combines the architecture of convolutional neural network (CNN) and long short-term memory network (LSTM). It combines convolution and recurrent structures to capture spatiotemporal information at the same time, and controls the flow of information through a gating mechanism. Convolution operations are added to each time step to capture the spatial information of time series multivariate data. The convolution kernel is shared at each time step, which helps to extract similar features. The information flow of memory units is controlled by forget gates and update gates, which is more suitable for processing complex time series data. Before extracting time series features, the ConvLSTM network needs to perform network parameter training and optimization to achieve the best effect. The calculation formula for a single training is:

[0064]

[0065] Among them, * represents convolution, Representing Hadamard, Represents input, C t represents the output unit state, h t represents the hidden state, h t-1 represents the hidden state of all information before time t-1, i t represents the input gate, f t represents the forget gate, o t Represents the output gate, W h , W x , W c represents the parameters to be learned, b i 、b f 、b c 、b o represents bias, σ(·) represents the sigmoid activation function, and tanh represents the tanh activation function.

[0066] All the time series data after data processing at T time points Input the ConvLSTM network and after calculation, we get the time series features [h 1 h 2 ……h T ].

[0067] Step 22: Dual Attention Network.

[0068] The biggest feature of time series data is the time label. Data with different time labels have different effects on the final prediction task. This embodiment proposes a dual attention network, which innovatively combines the hierarchical idea with the self-attention network. The time dependency of time series data is analyzed layer by layer in a hierarchical manner. The influence weights of different time series are adaptively learned and adjusted through the self-attention mechanism, thereby further improving the accuracy of the time series prediction task.

[0069] Specifically, the time series data features extracted by the ConvLSTM network are input into the inner self-attention network, and adaptive learning is performed at the fine-grained level to adjust the influence weight of the time series. The specific calculation process is as follows:

[0070]

[0071] Among them, s(h t ,h t-τ+1 ) represents the approximate function, Q, K, V are learnable parameter matrices, d is the spatial dimension, m is the sliding window size, and h is converted to t ,h t-τ+1 Convert it into latent space and calculate the inner product in the latent space to get the similarity β between the two hidden spaces τ , It is the time series data feature at time t output after the internal self-attention calculation transformation at the fine-grained level.

[0072] Then, adaptive learning is performed through the outer self-attention network to further adjust the influence weight of the time series. The specific calculation process is as follows:

[0073]

[0074] in, represents the approximate function, is a learnable parameter matrix, d is the spatial dimension, m is the sliding window size, and through the parameter matrix Will Convert it into latent space and calculate the inner product in the latent space to get the similarity of the two hidden spaces It is the time series data feature at time t output after the outer self-attention calculation transformation.

[0075] Step 3: Based on the time series data characteristics after weight adjustment, use the prediction model to predict the time series prediction results corresponding to the target prediction time.

[0076] The time series data features processed by the time series data processing module and the time series data feature extraction module in sequence, that is, the time series data features after integrating the spatiotemporal relationship of the time series data and adjusting the weights Input to the fully connected layer and calculate the time series trend through the softmax activation function.

[0077] Specifically, the time series trend is obtained by calculating the change of time series data. In this paper, the time series trend is divided into three categories: rising (+1), stable (0), and falling (-1). The specific calculation formula for the time series data trend is as follows:

[0078]

[0079] Among them, data t+1 is the time series data at time t+1, data t is the time series data at time t, ε rising is the rising threshold of 0.55%, ε falling is the decline threshold -0.50%.

[0080] This embodiment uses data sets and performance evaluation indicators to conduct comparative experiments and ablation experiments to analyze the proposed prediction model.

[0081] This embodiment selects real data from three fields, namely power, traffic, and environmental detection, to conduct a large number of experiments, namely, power data set, traffic data set, and air quality data set.

[0082] Electricity data set: hourly data from the PJM electricity market in the United States from January 1, 2017 to January 1, 2021. The data includes basic electricity information such as historical electricity prices, loads, energy prices, installed capacity, and multivariate variable information related to electricity information such as temperature, wind speed, and holidays. The prediction target is "electricity price".

[0083] Traffic Dataset: Contains traffic data measured between Minneapolis and St. Paul, Minnesota, USA from October 2, 2012 to September 30, 2018. The dataset has hourly numerical features such as traffic volume, speed, holidays, and weather, and the prediction target is "traffic volume".

[0084] Air quality dataset: Contains statistics of daily PM2.5, PM10, SO2, CO, NO2, O3 concentrations and AQI features from December 2, 2013 to October 31, 2018. The prediction target is "PM2.5 concentration".

[0085] The time series data is divided into a training set, a validation set, and a test set in a 7:2:1 manner in chronological order. This embodiment uses python to implement the model, and the time series data multivariate variables in the input data set of the model are output as predicted time series trends. The parameters are trained by the training set, the parameters are optimized by the validation set, and the model performance is tested by the test set. Two hidden layers are set in the ConvLSTM convolutional neural network of the model, with 32 units in each layer, and the sliding window is set to 16. The training of the first 16 time points predicts the next time point. The parameter optimization uses the Adam optimizer, and the learning rate decay scheduler is adopted, in which the initial learning rate is set to 0.0001, the decay coefficient is 0.5, the maximum number of training epochs is set to 300, and the early stopping strategy is adopted to prevent the model from overfitting.

[0086] The accuracy and F1 score are used as evaluation indicators. The calculation formula of the accuracy is as follows:

[0087]

[0088] Among them, t p is the number of correct predictions made in the positive class, t n is the number of correct predictions made in the negative class, f p is the number of wrong predictions made in the positive class, f n is the number of wrong predictions made in the negative class.

[0089] F 1 The calculation formula is as follows:

[0090]

[0091] The method of this embodiment is compared with various time series prediction methods to verify the effectiveness of the model of this embodiment. The comparison methods are trained and learned according to the parameter settings in their papers, and the parameters are optimized through the validation set. ARIMA: Moving average autoregressive model, a traditional method; SVM: support vector machine, a machine learning method; LSTM: a deep learning method, which models time series data and predicts time series trends; GRU: optimizes LSTM, with a simple architecture, faster computing efficiency and training speed; BiLSTM: a bidirectional LSTM model, which can make up for the defect that LSTM only captures positive dependency information and better captures bidirectional dependency information. The classification index evaluation results of the method of this embodiment and the above-mentioned method are shown in Table 1.

[0092] Table 1 Performance of different methods for time series forecasting (bold indicates the best result in each indicator, underline indicates the suboptimal result)

[0093]

[0094] As can be seen from Table 1, the prediction effect of the machine learning method SVM is greatly improved compared with the traditional method ARIMA. This is because the traditional method assumes that the data is linear and stable, which is contrary to the actual situation of high fluctuations, resulting in great limitations. Deep learning methods such as LSTM have further improved the prediction effect of the machine learning method SVM. This is because machine learning has poor processing effect on high-dimensional and complex data, and deep learning methods can process data while considering the nonlinearity, non-stationarity and high-dimensional complexity of the data. In the deep learning method, the accuracy and F1 value of the method of this embodiment are significantly improved compared with the method using only LSTM. Among them, in the power data set, the accuracy of the method of this embodiment is 7.1% higher than that of the LSTM method, and the F1 value is increased by 7.2%. The comparison results on the traffic and air quality data sets are similar to those on the power data set. The experimental results fully demonstrate the effectiveness of the model proposed in this embodiment for time series trend prediction.

[0095] This example uses ablation experiments to verify the importance and benefits of three key parts in the model in this paper to the final results. They are: 1. Abnormal fluctuation point detection network; 2. ConvLSTM network; 3. Dual attention network.

[0096] The model of this embodiment introduces an abnormal fluctuation point detection network to find out the points of abnormal numerical fluctuation in the time series data, and quantifies the influence of abnormal fluctuation points on the time series feature extraction through the abnormal fluctuation point weight matrix, thereby reducing the feature extraction error and improving the overall prediction accuracy. This embodiment compares the time series prediction results of the model with the abnormal fluctuation point detection network removed, denoted as OURS-noGAP, and the model with the abnormal fluctuation point detection network not removed, denoted as OURS, and the various indicators are shown in Table 2.

[0097] Table 2 Performance comparison of models with and without abnormal fluctuation point detection network (bold indicates the best result in each indicator)

[0098]

[0099] As can be seen from Table 2, for the three data sets of electricity, traffic and air quality, the index values ​​obtained by the model with the abnormal fluctuation point detection network are all optimal. This shows that the abnormal fluctuation point weight matrix calculated by the network can effectively reduce the impact of abnormal fluctuation points on time series prediction and effectively improve the accuracy of the prediction results.

[0100] This embodiment also studies the influence of different values ​​of the two experimental parameters ω (formula (3)) and θ (formula (4)) in the abnormal fluctuation point detection network on the final prediction results. Taking the power data set as an example, the specific results are as follows: Figure 2 and Figure 3 shown.

[0101] Figure 3 At the beginning, the experimental coefficient θ was very small, which resulted in the abnormal fluctuation point detection network obtaining too small an influence weight on other points, and the accuracy of electricity price trend prediction was very poor. As the value of θ increases, the influence weight of abnormal fluctuation points on other points gradually increases, and the improvement effect on the accuracy of electricity price trend prediction is obvious. When θ reaches 25, the accuracy is the highest. Then, as θ continues to increase, the influence weight of abnormal fluctuation points on other points in the abnormal fluctuation point detection network is too large, and the role of the abnormal fluctuation point detection network gradually disappears. The negative impact of abnormal fluctuation points on electricity price feature extraction gradually increases, and the improvement effect on the accuracy of electricity price trend prediction is no longer obvious. Therefore, when using the power data set for experiments, the θ experimental coefficient is set to 25, and the θ experimental coefficient setting operation is the same in the traffic data set and the air quality data set.

[0102] This embodiment selects the ConvLSTM network to extract features from time series data. The ConvLSTM network can analyze the spatial relationship between multivariate variables based on the time relationship extracted from the time series data, which can significantly improve the problem of incomplete feature information of time series data extracted only from the time relationship level and improve prediction accuracy.

[0103] In order to verify the improvement of ConvLSTM network on the model time series prediction performance, we designed the model variant OURS-L that only uses LSTM network to extract time relationship, the model variant OURS-G that only uses GRU network to extract time relationship, and the model variant OURS-B that only uses BiLSTM network to extract time relationship. OURS is used to represent the complete model that uses ConvLSTM network to extract time series data features from both time relationship and spatial relationship levels. The prediction results of the complete model and the three model variants are compared in terms of indicators, as shown in Table 3. It can be seen that compared with OURS-L, all indicators of OURS-G and OURS-B have been significantly improved, and compared with OURS-G and OURS-B, all indicators of OURS have been improved more significantly. This shows that extracting time series data features from the spatial relationship level can enrich the extraction of time series data feature information on the basis of extracting features from time relationship. Therefore, after the ConvLSTM network fuses the time series data features extracted from time relationship and spatial relationship, it has a better improvement on the time series prediction effect.

[0104] Table 3 Performance comparison of time series feature extraction model variants using different networks (bold indicates the best result in each indicator)

[0105]

[0106]

[0107] In the model of this embodiment, a dual attention network is introduced to capture the time dependency of time series data, hierarchically refine the time dependency of time series data, adaptively learn to weigh and calculate the influence weights of different time series, and improve the prediction accuracy. To verify the role of the dual attention network, this paper designs three model variants: OURS-HT, OURS-AT, and OURS. Among them, OURS-HT means removing dual attention, OURS-AT means single attention, and OURS means the model of this paper. The performance index comparison of the experimental results of the model variants and the complete model is shown in Table 4:

[0108] Table 4 Performance comparison of various variants of dual attention network (bold indicates the best result in each indicator)

[0109]

[0110] As can be seen from Table 4, the variant without dual attention network performs the worst, the variant with single attention network has some improvement, and the performance of the model in this paper is most significantly improved by adding dual attention network, which shows the gain effect of dual attention module on feature extraction of time series data. The dual attention network is introduced to adaptively learn and weigh and calculate the influence weights of different time series to improve the prediction accuracy.

[0111] This embodiment proposes a time series prediction method and system based on exception handling, which can perform time series prediction by learning the temporal and spatial relationships between time series data. The proposed abnormal fluctuation point detection network significantly reduces the negative impact of abnormal fluctuation points on the accuracy of time series prediction. On the basis of extracting the temporal relationship features of time series data, the ConvLSTM network introduces spatial relationship feature extraction, and finally fuses the spatiotemporal feature information of time series data, enriching and improving the extraction of feature information of time series data. The proposed dual attention network refines and adaptively encodes the weights of time series signals, selects appropriate weights for data at different time points, and improves the accuracy of time series prediction. In subsequent work, the spatial relationship of time series data can be further analyzed by combining graph neural networks with convolutional neural networks, so as to further improve the accuracy of time series prediction.

[0112] Embodiment 2

[0113] The purpose of this embodiment is to provide a time series prediction system based on exception processing, including:

[0114] An acquisition module is configured to: acquire historical time series data corresponding to a target prediction time;

[0115] An anomaly detection module is configured to: input the historical time series data corresponding to the target prediction time into a trained prediction model, when the prediction model processes the historical time series data corresponding to the target prediction time, use the fluctuation point detection network of the prediction model to detect abnormal fluctuation points of the historical time series data, calculate the weight of the abnormal fluctuation point according to the distance between the abnormal fluctuation point and other data points, and adjust the weight of other data points according to the abnormal fluctuation point weight;

[0116] An adjustment module is configured to: extract features of the adjusted historical time series data using the ConvLSTM network of the prediction model, and adaptively adjust the influence weights of different historical time series data using a dual attention mechanism to obtain features of the time series data after weight adjustment;

[0117] The time series prediction module is configured to: based on the time series data characteristics after weight adjustment, use the prediction model to predict the time series prediction result corresponding to the target prediction time.

[0118] In further embodiments, there is also provided:

[0119] An electronic device includes a memory and a processor, and computer instructions stored in the memory and executed on the processor, wherein when the computer instructions are executed by the processor, the method described in Embodiment 1 is performed. For the sake of brevity, no further description is given here.

[0120] It should be understood that in this embodiment, the processor may be a central processing unit CPU, and the processor may also be other general-purpose processors, digital signal processors DSP, application-specific integrated circuits ASIC, off-the-shelf programmable gate arrays FPGA or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0121] The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A portion of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.

[0122] A computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by a processor, the method described in embodiment 1 is completed.

[0123] The method in the first embodiment can be directly embodied as a hardware processor, or a combination of hardware and software modules in the processor. The software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware. To avoid repetition, it will not be described in detail here.

[0124] A computer program product includes a computer program, and when the computer program is executed by a processor, the method described in the first embodiment is implemented.

[0125] The present invention also provides at least one computer program product tangibly stored on a non-transitory computer-readable storage medium. The computer program product includes computer executable instructions, such as instructions included in a program module, which are executed in a device on a real or virtual processor of the target to perform the process / method as described above. Typically, program modules include routines, programs, libraries, objects, classes, components, data structures, etc. that perform specific tasks or implement specific abstract data types. In various embodiments, the functions of program modules can be combined or divided between program modules as needed. Machine executable instructions for program modules can be executed in local or distributed devices. In distributed devices, program modules can be located in local and remote storage media.

[0126] The computer program code for implementing the method of the present invention can be written in one or more programming languages. These computer program codes can be provided to the processor of a general-purpose computer, a special-purpose computer or other programmable data processing device, so that the program code, when executed by the computer or other programmable data processing device, causes the function / operation specified in the flow chart and / or block diagram to be implemented. The program code can be executed completely on a computer, partially on a computer, as an independent software package, partially on a computer and partially on a remote computer or completely on a remote computer or server.

[0127] In the context of the present invention, computer program codes or related data may be carried by any appropriate carrier to enable a device, apparatus or processor to perform the various processes and operations described above. Examples of carriers include signals, computer readable media, etc. Examples of signals may include electrical, optical, radio, acoustic or other forms of propagation signals, such as carrier waves, infrared signals, etc.

[0128] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with this embodiment can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0129] Although the above describes the specific implementation mode of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without creative work are still within the scope of protection of the present invention.

Claims

1. A time series prediction method based on exception processing, characterized in that: include: Obtain the historical time series data corresponding to the target prediction time; The historical time series data corresponding to the target prediction time is input into the trained prediction model. When the prediction model processes the historical time series data corresponding to the target prediction time, the fluctuation point detection network of the prediction model is used to detect the abnormal fluctuation points of the historical time series data. The weight of the abnormal fluctuation point is calculated according to the distance between the abnormal fluctuation point and other data points. The weight of other data points is adjusted according to the abnormal fluctuation point weight, specifically: Among them, distance t,τ represents the time difference between time node t and time node τ, θ represents the adjustment coefficient of the influence of abnormal fluctuation point on other points, m is the sliding window size, represents the standardized data point at time t; The abnormal fluctuation point meets the following requirements: Among them, data t Represents the historical time series data at time t, data t-1 Represents the historical time series data at time t-1, represents the average value of multiple data points in the sliding window m, represents the standard deviation of multiple data points in the sliding window m, ω is the experimental coefficient for controlling the threshold of abnormal fluctuation point detection, Gap t It is an abnormal fluctuation point; The ConvLSTM network of the prediction model is used to extract features from the adjusted historical time series data, and the dual attention mechanism is used to adaptively adjust the influence weights of different historical time series data to obtain the features of the time series data after weight adjustment. Based on the time series data characteristics after weight adjustment, the prediction model is used to predict the time series prediction result corresponding to the target prediction time.

2. A time series prediction method based on exception processing as claimed in claim 1, characterized in that: The dual attention mechanism is used to adaptively adjust the influence weights of different historical time series data to obtain the characteristics of the time series data after weight adjustment, specifically: Calculate the similarity between the time series features corresponding to the current moment and the time series features corresponding to the preset time in the hidden space, and obtain the time series data features corresponding to the current moment after the first weight adjustment through the internal self-attention layer based on the calculated similarity; Through the similarity in the hidden space between the time series data features corresponding to the current moment and the time series data features corresponding to the preset time, the final time series data features corresponding to the current moment after the second weight adjustment are calculated through the external self-attention layer according to the calculated similarity.

3. A time series prediction method based on exception processing as claimed in claim 1, characterized in that: Based on the time series data characteristics after weight adjustment, the prediction model is used to predict the time series prediction result corresponding to the target prediction time, specifically: The fully connected layer is used to process the time series data features after weight adjustment, and the time series trend is calculated through the activation function.

4. A time series prediction method based on exception processing as claimed in claim 1, characterized in that: Before the historical time series data corresponding to the target prediction time is input into the trained prediction model, it also includes standardization preprocessing of the acquired historical time series data corresponding to the target prediction time.

5. A time series prediction system based on exception processing, characterized in that: include: An acquisition module is configured to: acquire historical time series data corresponding to a target prediction time; The anomaly detection module is configured to: input the historical time series data corresponding to the target prediction time into the trained prediction model; when the prediction model processes the historical time series data corresponding to the target prediction time, the fluctuation point detection network of the prediction model is used to detect the abnormal fluctuation point of the historical time series data; the weight of the abnormal fluctuation point is calculated according to the distance between the abnormal fluctuation point and other data points; and the weight of other data points is adjusted according to the abnormal fluctuation point weight, specifically: Among them, distance t,τ represents the time difference between time node t and time node τ, θ represents the adjustment coefficient of the influence of abnormal fluctuation point on other points, m is the sliding window size, represents the standardized data point at time t; The abnormal fluctuation point meets the following requirements: Among them, data t Represents the historical time series data at time t, data t-1 Represents the historical time series data at time t-1, represents the average value of multiple data points in the sliding window m, represents the standard deviation of multiple data points in the sliding window m, ω is the experimental coefficient for controlling the threshold of abnormal fluctuation point detection, Gap t It is an abnormal fluctuation point; An adjustment module is configured to: extract features of the adjusted historical time series data using the ConvLSTM network of the prediction model, and adaptively adjust the influence weights of different historical time series data using a dual attention mechanism to obtain features of the time series data after weight adjustment; The time series prediction module is configured to: based on the time series data characteristics after weight adjustment, use the prediction model to predict the time series prediction result corresponding to the target prediction time.

6. An electronic device, characterized in that: The method comprises a memory and a processor and computer instructions stored in the memory and executed on the processor, wherein when the computer instructions are executed by the processor, the method according to any one of claims 1 to 4 is completed.

7. A computer-readable storage medium, characterized in that: Used to store computer instructions, which, when executed by a processor, complete the method described in any one of claims 1 to 4.

8. A computer program product, characterized in that The invention comprises a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Time series data anomaly detection method based on channel fusion self-attention mechanism

    CN117034175A

  • Online evaluation method for household energy storage battery

    CN117368744A