A Time Series Prediction Method and System Based on Hierarchical Self-Attention Mechanism
The hierarchical self-attention mechanism addresses the challenges of spatial correlations and local-global feature handling in time series prediction by clustering and using enhanced attention layers, resulting in improved long-term prediction accuracy and efficiency.
Patent Information
- Application Number
- CN202510495706.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-04-21
AI Technical Summary
Existing time series prediction methods are difficult to effectively capture the dynamic changes of time series in long-term prediction, especially ignoring the impact of hierarchical structure and local emergencies, resulting in unsatisfactory prediction results.
The time series prediction method based on the hierarchical self-attention mechanism is adopted, and the historical time series is clustered and analyzed, and the time dependence and spatial relationship are captured by the intra-class attention layer and the inter-class attention layer, and prediction is carried out in combination with the decoder.
It improves the accuracy and efficiency of time series prediction, especially the ability to effectively handle the propagation of local emergencies over a long period of time, reduces the time complexity of the model, and improves the modeling ability of long-term dependence.
Smart Images

Figure CN120011838B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of time series prediction, and particularly to a time series prediction method and system based on a hierarchical self-attention mechanism. Background Art
[0002] With the acceleration of the global digitalization process, the construction of intelligent systems has gradually become an important goal for the development of various industries. As a core component of it, the intelligent data management system aims to improve data processing efficiency and optimize the allocation and utilization of resources through information technology and data analysis means. As a key link in the intelligent data management system, time series prediction can help the management department optimize decisions, plan resource allocation, and respond to emergencies by predicting data trends in advance, thereby improving the overall operation efficiency of the system. However, due to its complex spatio-temporal dynamic changes, especially in long-term prediction, time series prediction still faces many challenges and it is difficult to obtain ideal prediction results. Summary of the Invention
[0003] The purpose of this application is to provide a time series prediction method and system based on a hierarchical self-attention mechanism, which can improve the prediction accuracy and efficiency of time series.
[0004] To achieve the above purpose, this application provides the following solutions:
[0005] In the first aspect, this application provides a time series prediction method based on a hierarchical self-attention mechanism, including:
[0006] Performing clustering analysis on the historical time series, dividing the historical time series into multiple categories, and obtaining multiple sub-matrices; each sub-matrix includes time series of the same category;
[0007] According to the multiple sub-matrices, using a pre-trained time series prediction model to perform time series prediction to obtain a future time series; the time series prediction model includes an intra-class attention layer, an inter-class attention layer, and a decoder;
[0008] Among them, the intra-class attention layer is used to block each sub-matrix respectively, and uses the self-attention mechanism to learn the time dependence relationship between each time series block within each sub-matrix to obtain a first time series representation;
[0009] The inter-class attention layer is used to block the first time series representation, uses the mapping attention mechanism to capture the original spatial relationship between each time series block within the first time series representation, and uses the enhanced attention mechanism to enhance each time series block within the first time series representation to obtain a second time series representation;
[0010] The decoder is used to map the second time series representation into a future time series.
[0011] In a second aspect, the present application provides a time series prediction system based on a hierarchical self-attention mechanism, including:
[0012] A classification module for performing clustering analysis on historical time series, dividing the historical time series into multiple categories to obtain multiple sub-matrices; each sub-matrix includes time series of the same category;
[0013] A prediction module for performing time series prediction on the basis of multiple sub-matrices by using a pre-trained time series prediction model to obtain future time series; the time series prediction model includes an intra-class attention layer, an inter-class attention layer, and a decoder;
[0014] Wherein, the intra-class attention layer is used to block each sub-matrix respectively, and learn the temporal dependence relationship between each temporal block in each sub-matrix by using a self-attention mechanism to obtain a first temporal representation;
[0015] The inter-class attention layer is used to block the first temporal representation, capture the original spatial relationship between each temporal block in the first temporal representation by using a mapping attention mechanism, and enhance each temporal block in the first temporal representation by using an enhanced attention mechanism to obtain a second temporal representation;
[0016] The decoder is used to map the second temporal representation to future time series.
[0017] According to the specific embodiments provided by the present application, the present application has the following technical effects:
[0018] The present application provides a time series prediction method and system based on a hierarchical self-attention mechanism. By dividing historical time series into multiple categories to form a hierarchical spatio-temporal structure, the time complexity of the time series prediction model is reduced, the modeling ability for long-term dependencies is improved, and thus the efficiency of time series prediction is enhanced. The intra-class attention layer is used to learn the temporal dependence relationship between temporal blocks, which can effectively capture the temporal features within a class. The mapping attention mechanism is used to capture the original spatial relationship between each temporal block in the first temporal representation, and the enhanced attention mechanism is used to enhance each temporal block in the first temporal representation, which can effectively handle the propagation of local emergencies, thereby significantly improving the accuracy of long-term time series prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0020] Figure 1This is an application environment diagram of a time series prediction method based on a hierarchical self-attention mechanism in an embodiment of the present application.
[0021] Figure 2 This is a schematic flowchart of a time series prediction method based on a hierarchical self-attention mechanism provided in an embodiment of the present application.
[0022] Figure 3 This is an overall framework diagram of a time series prediction model provided in an embodiment of the present application.
[0023] Figure 4 This is a framework diagram of an intra-class block attention mechanism provided in an embodiment of the present application.
[0024] Figure 5 This is a model framework diagram of an inter-class mapping attention mechanism provided in an embodiment of the present application.
[0025] Figure 6 This is a model framework diagram of an inter-class enhanced attention mechanism provided in an embodiment of the present application.
[0026] Figure 7 This is a schematic diagram of the functional modules of a time series prediction system based on a hierarchical self-attention mechanism provided in an embodiment of the present application. Detailed implementation manners
[0027] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0028] Most early time series predictions adopted traditional time series analysis methods. For example, classical statistical models such as autoregressive integrated moving average models and support vector regression. These models assume that the law of change of time series over time is linear. However, the change of time series data is not only affected by complex non-linear dynamic systems, but also restricted by various external factors (such as weather changes, market fluctuations, holidays, etc.), which makes it difficult for traditional methods based on linear assumptions to obtain ideal prediction effects.
[0029] With the rapid development of deep learning, especially the introduction of Recurrent Neural Network (RNN) and Long Short-Term Memory (LSTM), the modeling ability of sequence data has been significantly improved. Such models can capture non-linear features and achieve relatively good prediction results. However, models such as RNN and LSTM mainly focus on modeling time dependencies and ignore the spatial correlations between different time series. Time series data not only changes in the time dimension but also includes the mutual influence between individual time series. Due to factors such as geographical location and system structure, there are significant spatial correlations between different data points, and these models fail to effectively capture and utilize this feature.
[0030] To address the limitations of traditional methods, in recent years, many researchers have started applying graph neural networks to time series prediction. A typical approach is to pre-define a graph structure, use time series as nodes, and the relationships between time series as edges, and capture spatial dependencies through graph convolutional networks. Combining with time series models, these methods can model the spatio-temporal dynamic relationships of time series simultaneously. Representative works include those based on Multivariate Time series Graph Neural Networks (MTGNN), etc. By integrating spatial and temporal features into a unified framework, these methods have greatly improved the prediction accuracy. However, these methods still have deficiencies in modeling long-term dependencies and complex hierarchical structures.
[0031] With the success of the self-attention mechanism in fields such as natural language processing, researchers have also begun to try introducing the attention mechanism into time series prediction. The attention mechanism has excellent sequence modeling capabilities and can flexibly handle long-distance dependencies in sequence data. These characteristics make it very suitable for application to the problem of time series prediction with strong spatio-temporal correlations. However, models based on the attention mechanism still face challenges in the following aspects: one is the neglect of hierarchical relationships. Hierarchical structures are prevalent in time series, and different types of time series have different characteristics, but related methods do not take this into account. The other is the insufficient combination of local and global features. In the changes of time series, local sudden events will significantly affect the overall time series, but related methods are difficult to simultaneously capture the dynamic changes of local mutations and global trends.
[0032] This application proposes a time series prediction method and system based on a hierarchical self-attention mechanism for long-term prediction problems to solve the problems existing in the above-mentioned time series prediction methods.
[0033] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0034] The time series prediction method based on the hierarchical self-attention mechanism provided by the embodiments of the present application can be applied to, for example, Figure 1 the application environment shown in the figure. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be set separately, integrated on the server 104, placed on the cloud or other servers. The terminal 102 can send the historical time series to the server 104. After receiving the historical time series, the server 104 performs clustering analysis on the historical time series and uses a pre-trained time series prediction model to perform time series prediction to obtain the future time series. The server 104 can feedback the future time series to the terminal 102. In addition, in some embodiments, the time series prediction method based on the hierarchical self-attention mechanism can also be implemented separately by the server 104 or the terminal 102.
[0035] Among them, the terminal 102 can be, but is not limited to, various desktop computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers, and can also be a cloud server.
[0036] In an exemplary embodiment, as Figure 2 shown in the figure, a time series prediction method based on the hierarchical self-attention mechanism is provided. This method is executed by a computer device, and can be specifically executed by a computer device such as a terminal or a server alone, or jointly executed by a terminal and a server. In the embodiments of the present application, taking this method applied to Figure 1 the server 104 in the figure as an example for illustration, it includes the following steps 201 to step 202.
[0037] Step 201, perform clustering analysis on the historical time series, divide the historical time series into multiple categories, and obtain multiple sub-matrices. Each sub-matrix includes time series of the same category.
[0038] In an exemplary embodiment, a hierarchical segmentation algorithm is used to perform clustering analysis on the historical time series, divide the historical time series into multiple categories, and obtain multiple sub-matrices.
[0039] In complex time series prediction, there are significant correlations between different time series, especially between similar time series. Therefore, this application first uses a hierarchical segmentation algorithm (METIS) to cluster time series, dividing the time series into several categories, which facilitates the time series prediction model to learn the spatial relationships between time series. The core of the METIS algorithm is to gradually reduce the scale of the original graph by sparsifying nodes and edges, and its principle is as follows: ; where is the original input time series, is the number of categories, is the predefined graph structure, is the classified cluster, is the M th time series of a category.
[0040] The METIS algorithm optimally partitions the predefined graph structure to divide the time series into multiple highly correlated categories, laying a foundation for subsequent intra-class and inter-class modeling.
[0041] Step 202: Based on multiple sub-matrices, use a pre-trained time series prediction model to predict the time series and obtain the future time series. The time series prediction model includes an intra-class attention layer, an inter-class attention layer, and a decoder.
[0042] Among them, the intra-class attention layer is used to block each sub-matrix respectively, and uses the self-attention mechanism to learn the temporal dependencies between each time series block within each sub-matrix, obtaining the first time series representation.
[0043] The inter-class attention layer is used to block the first time series representation, uses the mapping attention mechanism to capture the original spatial relationships between each time series block within the first time series representation, and uses the enhanced attention mechanism to enhance each time series block within the first time series representation, obtaining the second time series representation.
[0044] The decoder is used to map the second time series representation to the future time series.
[0045] The time series prediction model provided by this application realizes the accurate modeling of spatio-temporal correlations in complex time series through category division and time chunk processing, and the mutual cooperation of the intra-class attention layer and the inter-class attention layer, especially effectively capturing and dealing with emergencies and data fluctuations, improving the accuracy and efficiency of time series prediction, especially for emergencies in time series.
[0046] In an exemplary embodiment, such as Figure 3As shown, first, a custom graph structure is read, then category division and time chunking are performed. After that, intra-class attention calculation is carried out, and then inter-class projection attention calculation and inter-class enhancement attention calculation are respectively performed. The results of the inter-class projection attention calculation and the inter-class enhancement attention calculation are added and normalized, and a multi-layer perceptron is used to output the prediction result.
[0047] The following specifically introduces the detailed process of the time series prediction model for time series prediction.
[0048] (1) Intra-class attention layer.
[0049] In this application, time series blocks are used as the input of the attention mechanism instead of the feature values of each time step. Specifically, the time series of the same class are used as the input to model the time series within the same class. Within each class, the time series blocks are mapped and the query matrix, key matrix, and value matrix are calculated respectively. Subsequently, the attention calculation between different time series blocks is completed through the self-attention mechanism, thereby obtaining the final representation of the class, which can effectively learn the time series features within the class and help decouple the complex correlations between different classes. In addition, sudden events may occur in the time series, resulting in sudden changes in the time series. The block-based analysis method can help the time series prediction model better model these sudden changes.
[0050] In the intra-class attention layer, the time series data within each class is divided into several time series blocks, and the time dependence between the time series within each class is learned through the self-attention mechanism. This layer effectively captures the time series features within the class through the attention calculation of the time series blocks.
[0051] In an exemplary embodiment, as Figure 4 shown, the intra-class attention layer first performs time chunking, then determines the query weight matrix, key weight matrix, and value weight matrix through a linear layer, calculates the attention score matrix according to the query weight matrix and the key weight matrix, weights the value weight matrix according to the attention score matrix by the attention score, then performs a residual connection between the weighted result and the input time series, and then performs normalization, fully connected feed-forward neural network, and normalization in sequence. The result obtained is then subjected to a residual connection with the result of the first normalization to obtain the first time series representation.
[0052] The intra-class attention layer divides each sub-matrix into blocks respectively, and uses the self-attention mechanism to learn the time dependence between the time series blocks within each sub-matrix. Obtaining the first time series representation specifically includes the following steps 301 to 305.
[0053] Step 301, for any sub-matrix, perform an expansion process on the time series of the sub-matrix to obtain the expanded time series of the sub-matrix.
[0054] In this application, time series in each category are segmented in the time dimension. The segmentation process fills zero values at the front end of the time series to ensure that the length of each segment is the same, guaranteeing the continuity of the segmentation process.
[0055] Step 302: Divide the extended time series of the sub-matrix into multiple blocks to obtain multiple time series blocks of the sub-matrix.
[0056] After completing the category division, sub-matrices composed of different time series are obtained. For each sub-matrix, this application provides a segmentation method to divide it according to a fixed size Before segmentation, this application performs an extension process on the time series, that is, fills zero values at the front end of the time series. Specifically, time steps are filled, where is the length of the sub-matrix. Then, according to the size of the time series block, the time series is divided into time series blocks. The segmentation process is a continuous operation, that is, the starting position of each time series block follows the ending position of the previous time series block.
[0057] The specific operation is as follows: ; where is the j th sub-matrix, is the segmentation operation, is the time series, is the number of time series blocks of each sub-matrix, is the j th sub-matrix's S th time series block, is the j th sub-matrix's number of time series.
[0058] Through the above segmentation method, the length of the input time series is shortened from to . Under the same memory usage conditions, longer sequences can be processed. In addition, the shortening of the input length effectively reduces the time complexity of the time series prediction model.
[0059] Step 303: According to the query weight matrix, key weight matrix, and value weight matrix of the sub-matrix, map each time series block of the sub-matrix from the time length to a multi-dimensional space to obtain the query matrix, key matrix, and value matrix of each time series block of the sub-matrix.
[0060] The query weight matrix, key weight matrix, and value weight matrix of each sub - matrix are pre - trained. In the intra - class attention layer of this application, a multi - head attention mechanism is adopted, and different attention mechanisms are used for different classes, which means that the parameters between classes are not shared, and the weight matrices of queries, keys, and values for each class are different.
[0061] Specifically, the following formula is used to map the j -th time series block of the s -th sub - matrix from the time length to the d -dimensional space to obtain the query matrix, key matrix, and value matrix of the j -th time series block of the s -th sub - matrix.
[0062] .
[0063] .
[0064] .
[0065] Where is the query matrix of the j -th time series block of the s -th sub - matrix, , is the key matrix of the j -th time series block of the s -th sub - matrix, is the value matrix of the j -th time series block of the s -th sub - matrix, is the j -th time series block of the s -th sub - matrix, is the query weight matrix of the j -th sub - matrix, is the key weight matrix of the j -th sub - matrix, is the value weight matrix of the j -th sub - matrix.
[0066] Step 304: Perform attention calculation between blocks according to the query matrix, key matrix, and value matrix of each time series block of the sub - matrix to obtain the time series representation of the sub - matrix.
[0067] Specifically, the following formula is used to obtain the time series representation of the j -th sub - matrix.
[0068] .
[0069] Where is the jTemporal representation of sub - matrices , is the temporal representation of the j th temporal block of the S th sub - matrix, is the normalized exponential function, is the query matrix of the j th sub - matrix, , is the key matrix of the j th sub - matrix, , is the value matrix of the j th sub - matrix, , d is the spatial dimension.
[0070] Step 305: Concatenate the temporal representations of each sub - matrix to obtain the first temporal representation: , where N is the number of original time series.
[0071] (2) Inter - class attention layer.
[0072] In the intra - class attention layer, information is aggregated in the time dimension by block division. In time series prediction tasks, there are mutual influence relationships between different time series. Therefore, the purpose of the inter - class attention layer is to model the spatial relationships between different time series.
[0073] In the inter - class attention layer, this application models the spatial relationships between different time series through two different attention mechanisms. Mapping attention is used to capture the original spatial relationships between time series, while enhanced attention is used to improve the sensitivity of the time series prediction model to time series mutations. Enhanced attention enhances the capture of mutation information through pooling operations, and then performs standard attention calculations to ensure that in the face of emergencies, important information can be transmitted in a timely manner, thereby making more accurate predictions for time series. Finally, the results of mapping attention and enhanced attention are fused to generate the final inter - class representation, that is, the second temporal representation.
[0074] In an exemplary embodiment, as Figure 5 shown, the mapping attention mechanism first determines the query weight matrix through a linear layer, determines the key weight matrix and value weight matrix through two multi - layer perceptrons respectively, then calculates the attention score matrix according to the query weight matrix and the key weight matrix, weights the value weight matrix by the attention score matrix, then performs a residual connection between the weighted result and the input time series, and then performs normalization, fully - connected feed - forward neural network, and normalization in sequence. After that, a residual connection is made between the obtained result and the result of the first normalization.
[0075] AsFigure 6 As shown, to enhance attention, first, the query weight matrix, key weight matrix, and value weight matrix are determined successively through the pooling layer and the linear layer. Then, the attention score matrix is calculated based on the query weight matrix and the key weight matrix. The value weight matrix is weighted by the attention score matrix according to the attention scores. After that, the weighted result is subjected to residual connection with the input time series. Then, normalization, fully connected feed-forward neural network, and normalization are performed successively. Finally, the obtained result is subjected to residual connection with the result of the first normalization.
[0076] The inter-class attention layer divides the first time series representation into blocks, adopts the mapping attention mechanism to capture the original spatial relationship between the time series blocks within the first time series representation, and adopts the enhanced attention mechanism to enhance each time series block within the first time series representation to obtain the second time series representation, which specifically includes the following steps 401 to 406.
[0077] Step 401: Divide the first time series representation into blocks to obtain a plurality of representation blocks, that is ; where is the R th representation block, , R is the number of representation blocks, which is the same as the number of time series blocks S of each sub-matrix.
[0078] Step 402: For any representation block, according to the query weight matrix, key weight matrix, and value weight matrix of the representation block, map the representation block from the time length to the multi-dimensional space to obtain the query matrix, key matrix, and value matrix of the representation block. The query weight matrix, key weight matrix, and value weight matrix of each representation block are pre-trained.
[0079] The process of mapping the representation block from the time length to the multi-dimensional space is the same as the process of mapping the time series block from the time length to the multi-dimensional space in the intra-class attention layer, which will not be elaborated here.
[0080] Step 403: According to the query matrix, key matrix, and value matrix of the representation block, adopt the multi-layer perceptron and the multi-head attention mechanism to capture the original spatial relationship between the time series to obtain the mapping attention representation of the representation block.
[0081] Specifically, the following formula is used to obtain the mapping attention representation of the r th representation block.
[0082] .
[0083] Where is the mapping attention representation of the r th representation block, is the normalization exponential function, is the query matrix for the r th representation block, is the key matrix for the r th representation block, is the value matrix for the r th representation block, and are both pre-trained multi-layer perceptrons.
[0084] Specifically, using the multi-layer perceptron for the and mapping operations helps the time series prediction model better learn the spatial relationships.
[0085] Step 404, perform a max pooling operation on the representation block to obtain the enhanced attention representation of the representation block.
[0086] The enhanced attention is used to improve the attention of the time series prediction model to the fluctuating data. When a sudden change occurs in a certain time series, it may affect other time series. Therefore, this application propagates the mutation information to other time series through enhanced attention. Specifically, a pooling operation is performed on the mapped attention representation obtained in the intra-class attention layer to highlight the parts with higher attention scores, and these parts represent the changes in the time series. That is, use the formula to determine the enhanced attention representation of the r th representation block. Among them, is the enhanced attention representation of the r th representation block.
[0087] Step 405, combine the mapped attention representation and the enhanced attention representation of the representation block to obtain the inter-class attention representation of the representation block.
[0088] Specifically, use the following formula to obtain the inter-class attention representation of the r th representation block.
[0089] .
[0090] Among them, is the inter-class attention representation of the r th representation block, is the mapped attention representation of the r th representation block, is the normalization process.
[0091] Step 406, concatenate the inter-class attention representations of multiple representation blocks to obtain the second time series representation: . Among them, is the second time series representation.
[0092] (3) Decoder.
[0093] In the decoding stage, the present application designs a simple multi-layer perceptron to map the generated second temporal representation to the final predicted value. The temporal prediction model processes the different chunks of each time series as a whole, thereby reducing the dimension of the representation from to . The loss function uses the mean squared error to measure the difference between the predicted value and the true value, thereby guiding the training of the temporal prediction model.
[0094] That is, the decoder uses the formula to determine the future time series . Among them, is the multi-layer perceptron.
[0095] During the training process of the temporal prediction model, the mean squared error is used as the loss function . Among them, is the loss function value, is the future time series output by the temporal prediction model, is the true time series.
[0096] The present application can simplify the decoding process through the multi-layer perceptron, and at the same time, by minimizing the mean squared error, effectively improve the accuracy of time series prediction.
[0097] The present application proposes a novel solution based on two types of hierarchical characteristics in the actual scenario: First, the spatial dependence between data points is hierarchical. Due to the different characteristics of different data, the data can be naturally divided into multiple categories. The data within the same category has high similarity or related behaviors. Second, the temporal dependence of the time series is hierarchical. Sudden events often occur suddenly, resulting in a sharp change in the time series in a short period of time, but this drastic change may have little impact on the long-term trend. The above two hierarchies have great potential to revolutionize the existing time series prediction methods.
[0098] Therefore, the present application simultaneously captures the local features and global features in the time series through a hierarchical attention mechanism. First, based on a predefined graph structure, the time series is divided into several categories and temporal chunks, thereby forming a hierarchical spatio-temporal structure, reducing the time complexity of the temporal prediction model, enhancing the modeling ability for long-term dependencies, and proposing an intra-class and inter-class attention mechanism to capture and fuse global and local spatial dependencies. Finally, an enhanced attention layer is designed to capture the dynamic changes of time series mutations, effectively handling the propagation of local emergencies, thereby significantly improving the accuracy of long-time period time series prediction.
[0099] This application not only effectively reduces the time complexity of the time series prediction model, but also improves the processing ability for long time series. Compared with the time complexity of the traditional self-attention mechanism, this application reduces the time complexity to by introducing a time chunking mechanism, significantly improving the applicability and operation efficiency of the time series prediction model in long time series.
[0100] This application has conducted a large number of experiments on multiple real-world datasets and achieved excellent prediction results in tests with different prediction time lengths (1 hour, 3 hours, and 6 hours), verifying the effectiveness and generality of the time series prediction model.
[0101] This application also provides an application scenario that applies the above-mentioned time series prediction method based on the hierarchical self-attention mechanism. Specifically: The time series prediction method based on the hierarchical self-attention mechanism provided in this embodiment can be applied to the traffic flow prediction scenario. In the traffic flow prediction scenario, the historical time series of the time series prediction method based on the hierarchical self-attention mechanism is the historical traffic flow monitoring data, and the future time series is the future traffic flow monitoring data.
[0102] Based on the same inventive concept, an embodiment of this application also provides a system for implementing the above-mentioned time series prediction method based on the hierarchical self-attention mechanism. The implementation solutions provided by this system to solve problems are similar to the implementation solutions described in the above method. Therefore, the specific limitations in one or more of the following embodiments of the time series prediction system based on the hierarchical self-attention mechanism can refer to the limitations on the time series prediction method based on the hierarchical self-attention mechanism in the above text and will not be repeated here.
[0103] In an exemplary embodiment, as Figure 7 shown, a time series prediction system based on the hierarchical self-attention mechanism is provided, including: a classification module 701 and a prediction module 702.
[0104] The classification module 701 is used to perform clustering analysis on the historical time series, divide the historical time series into multiple categories, and obtain multiple sub-matrices. Each sub-matrix includes time series of the same category.
[0105] The prediction module 702 is used to perform time series prediction on the basis of multiple sub-matrices by using a pre-trained time series prediction model to obtain a future time series; the time series prediction model includes an intra-class attention layer, an inter-class attention layer, and a decoder.
[0106] Among them, the intra-class attention layer is used to respectively chunk each sub-matrix, and use the self-attention mechanism to learn the temporal dependence relationship between each temporal chunk in each sub-matrix to obtain a first temporal representation.
[0107] The inter-class attention layer is used to block the first time series representation, capture the original spatial relationship between each time series block in the first time series representation by using a mapping attention mechanism, and enhance each time series block in the first time series representation by using an enhanced attention mechanism to obtain a second time series representation.
[0108] The decoder is used to map the second time series representation to a future time series.
[0109] In an exemplary embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.
[0110] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0111] In an exemplary embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0112] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0113] In this application, all actions of obtaining signals, information, or data are carried out on the premise of complying with the corresponding data protection regulations and policies of the country where the device is located and obtaining the authorization given by the owner of the corresponding device.
[0114] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the various embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tapes, floppy disks, flash memories, optical memories, high-density embedded non-volatile memories, resistive random access memories (ReRAM), magnetoresistive random access memories (MRAM), ferroelectric random access memories (FRAM), phase change memories (PCM), graphene memories, etc. Volatile memories can include random access memory (RAM) or external cache memories, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0115] The databases involved in the various embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the various embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logics, data processing logics based on quantum computing, etc., without limitation.
[0116] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0117] Specific examples are used in this article to elaborate on the principles and implementation manners of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A time series prediction method based on a hierarchical self-attention mechanism, characterized in that, The time series prediction method based on the hierarchical self-attention mechanism includes: Performing clustering analysis on the historical time series, dividing the historical time series into multiple categories to obtain multiple sub-matrices; each sub-matrix includes time series of the same category; According to the multiple sub-matrices, using a pre-trained time series prediction model to perform time series prediction to obtain the future time series; the time series prediction model includes an intra-class attention layer, an inter-class attention layer, and a decoder; Among them, the intra-class attention layer is used to block each sub-matrix respectively, and uses the self-attention mechanism to learn the time dependence relationship between each time series block in each sub-matrix to obtain the first time series representation; The inter-class attention layer is used to block the first time series representation, uses the mapping attention mechanism to capture the original spatial relationship between each time series block in the first time series representation, and uses the enhanced attention mechanism to enhance each time series block in the first time series representation to obtain the second time series representation; The decoder is used to map the second time series representation to the future time series; the historical time series is historical traffic flow monitoring data, and the future time series is future traffic flow monitoring data.
2. The time series prediction method based on the hierarchical self-attention mechanism according to claim 1, wherein Using a hierarchical segmentation algorithm to perform clustering analysis on the historical time series, dividing the historical time series into multiple categories to obtain multiple sub-matrices.
3. The time series prediction method based on the hierarchical self-attention mechanism according to claim 1, characterized in that The intra-class attention layer blocks each sub-matrix respectively, and uses the self-attention mechanism to learn the time dependence relationship between each time series block in each sub-matrix to obtain the first time series representation, specifically including: For any sub-matrix, performing expansion processing on the time series of the sub-matrix to obtain the expanded time series of the sub-matrix; Dividing the expanded time series of the sub-matrix into multiple blocks to obtain multiple time series blocks of the sub-matrix; According to the query weight matrix, key weight matrix, and value weight matrix of the sub-matrix, mapping each time series block of the sub-matrix from the time length to a multi-dimensional space to obtain the query matrix, key matrix, and value matrix of each time series block of the sub-matrix; the query weight matrix, key weight matrix, and value weight matrix of each sub-matrix are pre-trained; Performing attention calculation between blocks according to the query matrix, key matrix, and value matrix of each time series block of the sub-matrix to obtain the time series representation of the sub-matrix; Concatenating the time series representations of each sub-matrix to obtain the first time series representation.
4. The time series prediction method based on the hierarchical self-attention mechanism according to claim 3, wherein The timing representation of the j th sub-matrix is obtained using the following formula: ; Among them, is the timing representation of the j th sub - matrix, , is the timing representation of the j th timing block of the S th sub - matrix, S is the number of timing blocks for each sub - matrix, is the normalization exponential function, is the query matrix of the j th sub - matrix, , is the query matrix of the j th timing block of the S th sub - matrix, is the key matrix of the j th sub - matrix, , is the key matrix of the j th timing block of the S th sub - matrix, is the value matrix of the j th sub - matrix, , the j th timing block of the S th sub - matrix's value matrix, is the spatial dimension.
5. The time series prediction method based on the hierarchical self-attention mechanism according to claim 1, wherein The inter-class attention layer blocks the first time series representation, uses the mapping attention mechanism to capture the original spatial relationship between each time series block in the first time series representation, and uses the enhanced attention mechanism to enhance each time series block in the first time series representation to obtain the second time series representation, specifically including: Blocking the first time series representation to obtain multiple representation blocks; For any representation block, according to the query weight matrix, key weight matrix, and value weight matrix of the representation block, mapping the representation block from the time length to a multi-dimensional space to obtain the query matrix, key matrix, and value matrix of the representation block; the query weight matrix, key weight matrix, and value weight matrix of each representation block are pre-trained; According to the query matrix, key matrix, and value matrix representing the blocks, a multi-layer perceptron and a multi-head attention mechanism are adopted to capture the original spatial relationships between time series, and a mapped attention representation of the said block is obtained. A max pooling operation is performed on the said block to obtain an enhanced attention representation of the said block. The mapped attention representation and the enhanced attention representation of the said block are combined to obtain an inter-class attention representation of the said block. The inter-class attention representations of multiple blocks are concatenated to obtain a second time series representation.
6. The time series prediction method based on the hierarchical self-attention mechanism according to claim 5, wherein The mapping attention representation of the r th representation block is obtained by using the following formula: ; Among them, is the r -th mapping attention representation of the representation block, is the normalized exponential function, is the r -th query matrix of the representation block, is the r -th key matrix of the representation block, is the r -th value matrix of the representation block, and are both pre-trained multi-layer perceptrons, d is the spatial dimension.
7. The time series prediction method based on the hierarchical self-attention mechanism according to claim 5, characterized in that The following formula is used to obtain the inter-class attention representation of the r th representation block: ; Among them, is the r inter-class attention representation of the th representation block, r is the mapping attention representation of the th representation block, r is the enhanced attention representation of the th representation block, is the normalization process.
8. The time series prediction method based on the hierarchical self-attention mechanism according to claim 1, characterized in that The decoder is a multi-layer perceptron.
9. A time series prediction system based on a hierarchical self-attention mechanism, which is applied to the time series prediction method based on the hierarchical self-attention mechanism according to any one of claims 1-8, and is characterized in that, The time series prediction system based on the hierarchical self-attention mechanism includes: A classification module for performing clustering analysis on historical time series, dividing the historical time series into multiple categories, and obtaining multiple sub-matrices; each sub-matrix includes time series of the same category. A prediction module for performing time series prediction on the basis of multiple sub-matrices by using a pre-trained time series prediction model to obtain future time series; the time series prediction model includes an intra-class attention layer, an inter-class attention layer, and a decoder. Among them, the intra-class attention layer is used to divide each sub-matrix into blocks respectively, and a self-attention mechanism is adopted to learn the time dependence relationships between the time series blocks within each sub-matrix to obtain a first time series representation. The inter-class attention layer is used to divide the first time series representation into blocks, adopt a mapped attention mechanism to capture the original spatial relationships between the time series blocks within the first time series representation, and adopt an enhanced attention mechanism to enhance the time series blocks within the first time series representation to obtain a second time series representation. The decoder is used to map the second time series representation into future time series; the historical time series is historical traffic flow monitoring data, and the future time series is future traffic flow monitoring data.
Citation Information
Patent Citations
Short-term time sequence prediction method and system based on time and space attention
CN115081586A
Prediction method, device and system based on multivariable time series
CN115345220A