Time sequence prediction method and system based on hierarchical self-attention mechanism
By introducing a hierarchical self-attention mechanism in time series prediction, dividing time series categories and combining attention layer and decoder, the problem of spatiotemporal correlation and long-term dependence in time series prediction is solved, and the accuracy and efficiency of prediction are significantly improved.
Patent Information
- Application Number
- CN202510495706.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-21
AI Technical Summary
Time series prediction is difficult to obtain ideal prediction effects in long-term prediction, mainly due to complex spatial and temporal dynamic changes and failure to effectively capture the spatial correlation between time series.
The time series prediction method based on the hierarchical self-attention mechanism is adopted, and the historical time series is clustered and analyzed, and they are divided into multiple categories to form a hierarchical spatio-temporal structure. The time series is then learned and predicted using a combined model of the intra-class attention layer, the inter-class attention layer, and the decoder.
The accuracy and efficiency of time series prediction are improved, especially in long-term prediction, which can effectively capture the spatiotemporal correlation of time series and the propagation of emergencies.
Smart Images

Figure CN120011838A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of time series prediction, and in particular to a time series prediction method and system based on a hierarchical self-attention mechanism. Background Art
[0002] With the acceleration of the global digitalization process, the construction of intelligent systems has gradually become an important goal for the development of various industries. As its core component, the intelligent data management system aims to improve data processing efficiency and optimize resource allocation and utilization through information technology and data analysis. As a key link in the intelligent data management system, time series prediction can help management departments optimize decision-making, plan resource allocation and respond to emergencies by predicting data trends in advance, thereby improving the operating efficiency of the overall system. However, due to its complex spatiotemporal dynamic changes, time series prediction still faces many challenges, especially in long-term predictions, and it is difficult to obtain ideal prediction results. Summary of the invention
[0003] The purpose of this application is to provide a time series prediction method and system based on a hierarchical self-attention mechanism, which can improve the prediction accuracy and efficiency of time series.
[0004] To achieve the above objectives, this application provides the following solutions: In a first aspect, the present application provides a time series prediction method based on a hierarchical self-attention mechanism, comprising: Performing cluster analysis on the historical time series, dividing the historical time series into multiple categories, and obtaining multiple sub-matrices; each sub-matrix includes the time series of the same category; According to the multiple sub-matrices, a pre-trained time series prediction model is used to perform time series prediction to obtain a future time series; the time series prediction model includes an intra-class attention layer, an inter-class attention layer and a decoder; The intra-class attention layer is used to divide each sub-matrix into blocks respectively, and adopts a self-attention mechanism to learn the time dependency between each time series block in each sub-matrix to obtain a first time series representation; The inter-class attention layer is used to divide the first time series representation into blocks, use a mapping attention mechanism to capture the original spatial relationship between each time series block in the first time series representation, and use an enhanced attention mechanism to enhance each time series block in the first time series representation to obtain a second time series representation; The decoder is used to map the second time series representation into a future time series.
[0005] In the second aspect, the present application provides a time series prediction system based on a hierarchical self-attention mechanism, comprising: A classification module is used to perform cluster analysis on the historical time series, divide the historical time series into multiple categories, and obtain multiple sub-matrices; each sub-matrix includes the time series of the same category; A prediction module, used to perform time series prediction based on multiple sub-matrices using a pre-trained time series prediction model to obtain a future time series; the time series prediction model includes an intra-class attention layer, an inter-class attention layer and a decoder; The intra-class attention layer is used to divide each sub-matrix into blocks respectively, and adopts a self-attention mechanism to learn the time dependency between each time series block in each sub-matrix to obtain a first time series representation; The inter-class attention layer is used to divide the first time series representation into blocks, use a mapping attention mechanism to capture the original spatial relationship between each time series block in the first time series representation, and use an enhanced attention mechanism to enhance each time series block in the first time series representation to obtain a second time series representation; The decoder is used to map the second time series representation into a future time series.
[0006] According to the specific embodiments provided in this application, this application has the following technical effects: The present application provides a time series prediction method and system based on a hierarchical self-attention mechanism, which divides the historical time series into multiple categories to form a hierarchical spatiotemporal structure, reduces the time complexity of the time series prediction model, improves the modeling ability of long-term dependencies, and thus improves the efficiency of time series prediction. The intra-class attention layer is used to learn the time dependency relationship between time series blocks, which can effectively capture the time series characteristics within the class, and the mapping attention mechanism is used to capture the original spatial relationship between each time series block in the first time series representation. The enhanced attention mechanism is used to enhance each time series block in the first time series representation, which can effectively handle the propagation of local emergencies, thereby significantly improving the accuracy of long-term time series prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0008] Figure 1 This is an application environment diagram of a time series prediction method based on a hierarchical self-attention mechanism in one embodiment of the present application.
[0009] Figure 2 A flowchart of a time series prediction method based on a hierarchical self-attention mechanism is provided for one embodiment of the present application.
[0010] Figure 3 An overall framework diagram of a time series prediction model provided in one embodiment of the present application.
[0011] Figure 4 A framework diagram of the intra-class block attention mechanism provided in one embodiment of the present application.
[0012] Figure 5 A model framework diagram of the inter-class mapping attention mechanism provided in one embodiment of the present application.
[0013] Figure 6 A model framework diagram of the inter-class enhanced attention mechanism provided in one embodiment of the present application.
[0014] Figure 7 A schematic diagram of the functional modules of a time series prediction system based on a hierarchical self-attention mechanism provided in one embodiment of the present application. DETAILED DESCRIPTION
[0015] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0016] Early time series forecasting mostly used traditional time series analysis methods. For example, classic statistical models such as the autoregressive integrated moving average model and support vector regression assume that the law of time series change over time is linear. However, the change of time series data is not only affected by complex nonlinear dynamic systems, but also subject to various external factors (such as weather changes, market fluctuations, holidays, etc.), which makes it difficult for traditional methods based on linear assumptions to obtain ideal prediction results.
[0017] With the rapid development of deep learning, especially the introduction of recurrent neural networks (RNN) and long short-term memory networks (LSTM), the modeling ability of sequence data has been significantly improved. Such models can capture nonlinear features and achieve relatively good prediction results. However, models such as RNN and LSTM mainly focus on modeling time dependencies, while ignoring the spatial correlation between different time series. Time series data is not only about changes in the time dimension, but also includes the mutual influence between each time series. There is a significant spatial correlation between different data points due to factors such as geographical location and system structure, and these models fail to effectively capture and utilize this feature.
[0018] In order to address the limitations of traditional methods, in recent years, many researchers have begun to apply graph neural networks to time series prediction. The typical approach is to predefine a graph structure, use time series as nodes, and the relationships between time series as edges, and capture spatial dependencies through graph convolutional networks. Combined with time series models, these methods can simultaneously model the spatiotemporal dynamic relationships of time series. Representative works include multivariate time series graph neural networks (MTGNN), which greatly improve the accuracy of predictions by integrating spatial and temporal features into a unified framework. However, such methods still have shortcomings in modeling long-term dependencies and complex hierarchical structures.
[0019] With the success of the self-attention mechanism in fields such as natural language processing, researchers have also begun to try to introduce the attention mechanism into time series prediction. The attention mechanism has excellent sequence modeling capabilities and can flexibly handle long-distance dependencies in sequence data. These characteristics make it very suitable for time series prediction, a problem with strong spatiotemporal correlation. However, models based on the attention mechanism still have challenges in the following aspects: First, the neglect of hierarchical relationships. There is a common hierarchical structure in time series, and different types of time series have different characteristics, which the relevant methods do not take into account. Second, the combination of local and global features is insufficient. In the change of time series, local emergencies will significantly affect the overall time series, but the relevant methods find it difficult to capture the dynamic changes of local mutations and global trends at the same time.
[0020] This application proposes a time series prediction method and system based on a hierarchical self-attention mechanism for long-term prediction problems, in order to solve the problems existing in the above-mentioned time series prediction methods.
[0021] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0022] The time series prediction method based on the hierarchical self-attention mechanism provided in the embodiment of the present application can be applied to Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be set up separately, or it can be integrated on the server 104, or it can be placed on the cloud or other servers. The terminal 102 can send historical time series to the server 104. After receiving the historical time series, the server 104 performs cluster analysis on the historical time series, and uses a pre-trained time series prediction model to predict the time series to obtain the future time series. The server 104 can feedback the future time series to the terminal 102. In addition, in some embodiments, the time series prediction method based on the hierarchical self-attention mechanism can also be implemented by the server 104 or the terminal 102 alone.
[0023] The terminal 102 may be, but is not limited to, various desktop computers, laptop computers, smart phones, tablet computers, IoT devices, and portable wearable devices. The IoT devices may be smart speakers, smart TVs, smart air conditioners, smart vehicle-mounted devices, etc. The portable wearable devices may be smart watches, smart bracelets, head-mounted devices, etc. The server 104 may be implemented as an independent server or a server cluster consisting of multiple servers, or may be a cloud server.
[0024] In an exemplary embodiment, Figure 2 As shown, a time series prediction method based on a hierarchical self-attention mechanism is provided. The method is executed by a computer device, and can be executed by a computer device such as a terminal or a server alone, or by a terminal and a server together. In the embodiment of the present application, the method is applied to Figure 1 The server 104 in the example is used for explanation, and the steps include the following steps 201 to 202.
[0025] Step 201, cluster analysis is performed on the historical time series, and the historical time series is divided into multiple categories to obtain multiple sub-matrices. Each sub-matrix includes the time series of the same category.
[0026] In an exemplary embodiment, a hierarchical segmentation algorithm is used to perform cluster analysis on the historical time series, and the historical time series is divided into multiple categories to obtain multiple sub-matrices.
[0027] In complex time series prediction, there are significant correlations between different time series, especially between similar time series. Therefore, this application first uses a hierarchical segmentation algorithm (METIS) to cluster time series and divide the time series into several categories, so that the time series prediction model can learn the spatial relationship between time series. The core of the METIS algorithm is to gradually reduce the size of the original graph by sparse nodes and edges. The principle is: ;in, is the original input time series, is the number of categories, is a predefined graph structure, is the classified cluster, For the M time series of categories.
[0028] The METIS algorithm divides the time series into multiple highly correlated classes by optimizing the segmentation of the predefined graph structure, thus laying the foundation for subsequent intra-class and inter-class modeling.
[0029] Step 202: Based on the multiple sub-matrices, a pre-trained time series prediction model is used to perform time series prediction to obtain future time series. The time series prediction model includes an intra-class attention layer, an inter-class attention layer, and a decoder.
[0030] The intra-class attention layer is used to divide each sub-matrix into blocks respectively, and adopts the self-attention mechanism to learn the time dependency between each time series block in each sub-matrix to obtain the first time series representation.
[0031] The inter-class attention layer is used to block the first time series representation, use a mapping attention mechanism to capture the original spatial relationship between each time series block in the first time series representation, and use an enhanced attention mechanism to enhance each time series block in the first time series representation to obtain a second time series representation.
[0032] The decoder is used to map the second time series representation into a future time series.
[0033] The time series prediction model provided in this application achieves accurate modeling of spatiotemporal correlations in complex time series through category division and time block processing, and the coordination of intra-class attention layer and inter-class attention layer, especially the effective capture and response to sudden events and data fluctuations, thereby improving the accuracy and efficiency of time series prediction, especially for sudden situations in time series.
[0034] In an exemplary embodiment, Figure 3 As shown, the custom graph structure is first read, and then the category division and time block division are performed. After that, the intra-class attention calculation is performed, and then the inter-class projection attention calculation and the inter-class enhanced attention calculation are performed respectively. The results of the inter-class projection attention calculation and the inter-class enhanced attention calculation are added and normalized, and the multi-layer perceptron is used to output the prediction result.
[0035] The following is a detailed introduction to the time series prediction process of the time series prediction model.
[0036] (1) Intra-class attention layer.
[0037] This application uses time series blocks as the input of the attention mechanism instead of the feature values of each time step. Specifically, the time series of the same class is used as input to model the time series within the same class. In each class, the time series blocks are mapped and the query matrix, key matrix and value matrix are calculated separately. Then, the attention calculation between different time series blocks is completed through the self-attention mechanism to obtain the final representation of the class, which can effectively learn the time series features within the class and help decouple the complex correlations between different classes. In addition, sudden events may occur in the time series, resulting in sudden changes in the time series. The block-based analysis method can help the time series prediction model better model these sudden changes.
[0038] In the intra-class attention layer, the time series data in each class is divided into several time series blocks, and the time dependency between the time series in the class is learned through the self-attention mechanism. This layer effectively captures the time series characteristics within the class by calculating the attention of the time series blocks.
[0039] In an exemplary embodiment, Figure 4 As shown, the intra-class attention layer first performs time blocking, and then determines the query weight matrix, key weight matrix and value weight matrix through the linear layer, calculates the attention score matrix according to the query weight matrix and the key weight matrix, and uses the attention score matrix to weight the value weight matrix according to the attention score. After that, the weighted result is residually connected with the input time series, and then normalized, fully connected feedforward neural network and normalized are performed in sequence, and the obtained result is residually connected with the result of the first normalization to obtain the first time series representation.
[0040] The intra-class attention layer divides each sub-matrix into blocks respectively, and adopts the self-attention mechanism to learn the time dependency between each time series block in each sub-matrix, so as to obtain the first time series representation, which specifically includes the following steps 301 to 305.
[0041] Step 301: for any sub-matrix, perform expansion processing on the time series of the sub-matrix to obtain an expanded time series of the sub-matrix.
[0042] This application processes the time series in each category in the time dimension in blocks. The block process fills the front end of the time series with zero values to ensure that the length of each block is the same, thus ensuring the continuity of the block processing.
[0043] Step 302: Divide the extended time series of the submatrix into multiple blocks to obtain multiple time series blocks of the submatrix.
[0044] After completing the classification, we get For each sub-matrix, this application provides a block method to divide it into fixed sizes. In order to improve the prediction effect of the time series prediction model for the nearest time step, this application expands the time series before segmentation, that is, fills the front end of the time series with zero values, specifically filling time steps, where is the length of the submatrix. Then, according to the size of the timing block Divide the time series into The block division process is a continuous operation, that is, the starting position of each timing block follows the ending position of the previous timing block.
[0045] The specific operations are: ;in, For the j sub-matrices, For block operation, is a time series, is the number of timing blocks for each submatrix, For the j The sub-matrix S A timing block, , For the j The number of time series in the submatrix.
[0046] Through the above-mentioned block method, the length of the input time series is changed from Shorten to , under the same memory usage conditions, it can process longer sequences. In addition, the shortening of input length effectively reduces the time complexity of the time series prediction model.
[0047] Step 303, according to the query weight matrix, key weight matrix and value weight matrix of the submatrix, each time block of the submatrix is mapped from the time length to the multidimensional space to obtain the query matrix, key matrix and value matrix of each time block of the submatrix.
[0048] The query weight matrix, key weight matrix and value weight matrix of each submatrix are pre-trained. This application adopts a multi-head attention mechanism in the intra-class attention layer, and uses different attention mechanisms for different classes, which means that the parameters between classes are not shared, and the query, key, and value weight matrices of each class are different.
[0049] Specifically, the following formula is used to convert the j The sub-matrix s The time blocks are mapped from time length to d dimensional space, to obtain the j The sub-matrix s The query matrix, key matrix, and value matrix of each time series block.
[0050] .
[0051] .
[0052] .
[0053] in, For the j The sub-matrix s The query matrix of time series blocks, , For the j The sub-matrix s The key matrix of the timing blocks, For the j The sub-matrix s The value matrix of a time series block, For the j The sub-matrix s A timing block, For the j The query weight matrix of the sub-matrices, For the j The key weight matrix of the sub-matrices, For the j The value weight matrix of the sub-matrices.
[0054] Step 304, performing attention calculation between blocks according to the query matrix, key matrix and value matrix of each time-series block of the submatrix to obtain a time-series representation of the submatrix.
[0055] Specifically, the following formula is used to obtain j The time series representation of the sub-matrices.
[0056] .
[0057] in, For the j The time series representation of the sub-matrices is: , For the j The sub-matrix S The timing representation of a timing block, is the normalized exponential function, For the j The query matrix of sub-matrices, , For the j The key matrix of the submatrices, , For the j The value matrix of the submatrices, , d is the spatial dimension.
[0058] Step 305, concatenate the time series representation of each sub-matrix to obtain a first time series representation: ,in, N is the number of original time series.
[0059] (2) Inter-class attention layer.
[0060] In the intra-class attention layer, information is aggregated in the time dimension in a block-by-block manner. In the time series prediction task, different time series have mutual influence relationships, so the purpose of the inter-class attention layer is to model the spatial relationship between different time series.
[0061] In the inter-class attention layer, this application models the spatial relationship between different time series through two different attention mechanisms. Mapping attention is used to capture the original spatial relationship between time series, while enhanced attention is used to improve the sensitivity of the time series prediction model to time series mutations. Enhanced attention enhances the capture of mutation information through pooling operations, and then performs standard attention calculations to ensure that important information can be transmitted in a timely manner when faced with emergencies, thereby making more accurate predictions of time series. Finally, the results of mapping attention and enhanced attention are fused to generate the final inter-class representation, i.e., the second time series representation.
[0062] In an exemplary embodiment, Figure 5 As shown in the figure, the mapping attention mechanism first determines the query weight matrix through a linear layer, determines the key weight matrix and the value weight matrix through two multi-layer perceptrons respectively, and then calculates the attention score matrix according to the query weight matrix and the key weight matrix, and weights the value weight matrix according to the attention score matrix. After that, the weighted result is residually connected with the input time series, and then normalized, fully connected feedforward neural network and normalized are performed in sequence, and the result is residually connected with the result of the first normalization.
[0063] like Figure 6 As shown in the figure, the enhanced attention first determines the query weight matrix, key weight matrix and value weight matrix through the pooling layer and the linear layer in sequence, and then calculates the attention score matrix based on the query weight matrix and the key weight matrix, and weights the value weight matrix according to the attention score matrix. After that, the weighted result is residually connected with the input time series, and then normalized, fully connected feedforward neural network and normalized are performed in sequence, and the result is residually connected with the result of the first normalization.
[0064] The inter-class attention layer divides the first time series representation into blocks, adopts the mapping attention mechanism to capture the original spatial relationship between the time series blocks in the first time series representation, and adopts the enhanced attention mechanism to enhance the time series blocks in the first time series representation to obtain the second time series representation, which specifically includes the following steps 401 to 406.
[0065] Step 401, divide the first time series representation into blocks to obtain multiple representation blocks, namely ;in, For the R A representation block, , R is the number of representation blocks, which is related to the number of sequential blocks in each submatrix S same.
[0066] Step 402: for any representation block, the representation block is mapped from the time length to the multidimensional space according to the query weight matrix, key weight matrix and value weight matrix of the representation block to obtain the query matrix, key matrix and value matrix of the representation block. The query weight matrix, key weight matrix and value weight matrix of each representation block are pre-trained.
[0067] The process of mapping the representation block from time length to multi-dimensional space is the same as the process of mapping the temporal block from time length to multi-dimensional space in the intra-class attention layer, and will not be repeated here.
[0068] Step 403, according to the query matrix, key matrix and value matrix of the representation block, a multi-layer perceptron and a multi-head attention mechanism are used to capture the original spatial relationship between time series to obtain the mapping attention representation of the representation block.
[0069] Specifically, the following formula is used to obtain r The mapped attention representation of each representation block.
[0070] .
[0071] in, For the r The mapping attention representation of the representation block, is the normalized exponential function, For the r A query matrix representing a block, For the r A key matrix representing a block, For the r A matrix of values representing blocks, and All are pre-trained multi-layer perceptrons.
[0072] In particular, a multi-layer perceptron is used to and The mapping operation helps the time series prediction model to better learn spatial relationships.
[0073] Step 404: perform a maximum pooling operation on the representation block to obtain an enhanced attention representation of the representation block.
[0074] Enhanced attention is used to improve the attention of the time series prediction model to fluctuating data. When a time series changes suddenly, it may affect other time series. Therefore, this application propagates the mutation information to other time series by enhancing attention. Specifically, the mapping attention representation obtained in the intra-class attention layer is pooled to highlight the parts with higher attention scores, which represent the changes in the time series. That is, the formula is used Determine r The enhanced attention representation of the representation block. Among them, For the r Enhanced attention representation of representation blocks.
[0075] Step 405: Combine the mapped attention representation and the enhanced attention representation of the representation block to obtain the inter-class attention representation of the representation block.
[0076] Specifically, the following formula is used to obtain r The inter-class attention representation of the representation block.
[0077] .
[0078] in, For the r The inter-class attention representation of the representation block, For the r The mapping attention representation of the representation block, For standardized processing.
[0079] Step 406, concatenate the inter-class attention representations of the multiple representation blocks to obtain a second temporal representation: .in, It is the second timing representation.
[0080] (3) Decoder.
[0081] In the decoding stage, the present application designs a simple multi-layer perceptron to map the generated second time series representation to the final prediction value. The time series prediction model processes the different blocks of each time series as a whole, thereby transforming the dimension of the representation from Convert to The loss function uses mean square error to measure the difference between the predicted value and the true value, thereby guiding the training of the time series prediction model.
[0082] That is, the decoder uses the formula Determine future time series .in, A multi-layer perceptron.
[0083] In the training process of the time series prediction model, the mean square error is used as the loss function .in, is the loss function value, is the future time series output by the time series forecasting model. is a real time series.
[0084] This application can simplify the decoding process through a multi-layer perceptron and effectively improve the accuracy of time series prediction by minimizing the mean square error.
[0085] This application proposes a novel solution based on two types of hierarchical characteristics in actual scenarios: First, the spatial dependency between data points is hierarchical. Due to the different characteristics of different data, the data can be naturally divided into multiple categories. Data within the same category have high similarity or related behaviors. Second, the temporal dependency of time series is hierarchical. Sudden events often occur suddenly, causing time series to change dramatically in a short period of time, but such drastic changes may have little impact on long-term trends. The above two layers have great potential to innovate existing time series prediction methods.
[0086] Therefore, this application captures both local and global features in time series through a hierarchical attention mechanism. First, the time series is divided into several categories and time series blocks based on a predefined graph structure to form a hierarchical spatiotemporal structure, which reduces the time complexity of the time series prediction model and improves the modeling ability of long-term dependencies. In addition, the intra-category and inter-category attention mechanisms are proposed to capture and fuse global and local spatial dependencies. Finally, an enhanced attention layer is designed to capture the dynamic changes of time series mutations and effectively handle the propagation of local emergencies, thereby significantly improving the accuracy of long-term time series prediction.
[0087] This application not only effectively reduces the time complexity of the time series prediction model, but also improves the processing ability of long time series. The time complexity of this application is reduced to , significantly improving the applicability and computational efficiency of time series prediction models in long time series.
[0088] This application conducted a large number of experiments on multiple real-world datasets and achieved excellent prediction results in tests with different prediction time lengths (1 hour, 3 hours, and 6 hours), verifying the effectiveness and versatility of the time series prediction model.
[0089] The present application also provides an application scenario, which applies the above-mentioned time series prediction method based on the hierarchical self-attention mechanism. Specifically: The time series prediction method based on the hierarchical self-attention mechanism provided in this embodiment can be applied in the traffic flow prediction scenario. In the traffic flow prediction scenario, the historical time series of the time series prediction method based on the hierarchical self-attention mechanism is the historical traffic flow monitoring data, and the future time series is the future traffic flow monitoring data.
[0090] Based on the same inventive concept, the embodiment of the present application also provides a system for implementing the above-mentioned method for time series prediction based on the hierarchical self-attention mechanism. The implementation scheme for solving the problem provided by the system is similar to the implementation scheme recorded in the above-mentioned method, so the specific limitations in one or more embodiments of the time series prediction system based on the hierarchical self-attention mechanism provided below can be referred to the limitations of the time series prediction method based on the hierarchical self-attention mechanism above, and will not be repeated here.
[0091] In an exemplary embodiment, Figure 7 As shown, a time series prediction system based on a hierarchical self-attention mechanism is provided, including: a classification module 701 and a prediction module 702.
[0092] The classification module 701 is used to perform cluster analysis on the historical time series, and divide the historical time series into multiple categories to obtain multiple sub-matrices. Each sub-matrix includes the time series of the same category.
[0093] The prediction module 702 is used to perform time series prediction based on multiple sub-matrices using a pre-trained time series prediction model to obtain future time series; the time series prediction model includes an intra-class attention layer, an inter-class attention layer and a decoder.
[0094] The intra-class attention layer is used to divide each sub-matrix into blocks respectively, and adopts the self-attention mechanism to learn the time dependency between each time series block in each sub-matrix to obtain the first time series representation.
[0095] The inter-class attention layer is used to block the first time series representation, use a mapping attention mechanism to capture the original spatial relationship between each time series block in the first time series representation, and use an enhanced attention mechanism to enhance each time series block in the first time series representation to obtain a second time series representation.
[0096] The decoder is used to map the second time series representation into a future time series.
[0097] In an exemplary embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.
[0098] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0099] In an exemplary embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0100] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0101] In this application, all actions to obtain signals, information or data are carried out in compliance with the relevant data protection laws and policies of the country where they are located and with the authorization given by the owner of the corresponding device.
[0102] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).
[0103] The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. The non-relational database may include a distributed database based on blockchain, etc., but is not limited thereto. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but is not limited thereto.
[0104] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0105] This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application. At the same time, for those skilled in the art, according to the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.
Claims
1. A time series prediction method based on hierarchical self-attention mechanism, characterized in that: The time series prediction method based on the hierarchical self-attention mechanism includes: Performing cluster analysis on the historical time series, dividing the historical time series into multiple categories, and obtaining multiple sub-matrices; each sub-matrix includes the time series of the same category; According to the multiple sub-matrices, a pre-trained time series prediction model is used to perform time series prediction to obtain a future time series; the time series prediction model includes an intra-class attention layer, an inter-class attention layer and a decoder; The intra-class attention layer is used to divide each sub-matrix into blocks respectively, and adopts a self-attention mechanism to learn the time dependency between each time series block in each sub-matrix to obtain a first time series representation; The inter-class attention layer is used to divide the first time series representation into blocks, use a mapping attention mechanism to capture the original spatial relationship between each time series block in the first time series representation, and use an enhanced attention mechanism to enhance each time series block in the first time series representation to obtain a second time series representation; The decoder is used to map the second time series representation into a future time series.
2. The time series prediction method based on hierarchical self-attention mechanism according to claim 1 is characterized in that: A hierarchical segmentation algorithm is used to perform cluster analysis on the historical time series, and the historical time series is divided into multiple categories to obtain multiple sub-matrices.
3. The time series prediction method based on hierarchical self-attention mechanism according to claim 1 is characterized in that: The intra-class attention layer divides each sub-matrix into blocks, and uses the self-attention mechanism to learn the temporal dependency between the time series blocks in each sub-matrix to obtain the first temporal representation, which specifically includes: For any sub-matrix, performing expansion processing on the time series of the sub-matrix to obtain an expanded time series of the sub-matrix; Dividing the extended time series of the submatrix into a plurality of blocks to obtain a plurality of time series blocks of the submatrix; According to the query weight matrix, key weight matrix and value weight matrix of the submatrix, each time series block of the submatrix is mapped from the time length to the multidimensional space to obtain the query matrix, key matrix and value matrix of each time series block of the submatrix; the query weight matrix, key weight matrix and value weight matrix of each submatrix are pre-trained; Performing attention calculation between blocks according to the query matrix, the key matrix and the value matrix of each time-series block of the submatrix to obtain a time-series representation of the submatrix; The time series representation of each sub-matrix is concatenated to obtain a first time series representation.
4. The time series prediction method based on hierarchical self-attention mechanism according to claim 3 is characterized in that: Use the following formula to get j The time series representation of the sub-matrices is: ; in, For the j The time series representation of the sub-matrices is: , For the j The sub-matrix S The timing representation of a timing block, S is the number of timing blocks for each sub-matrix, is the normalized exponential function, For the j The query matrix of sub-matrices, , For the j The sub-matrix S The query matrix of time series blocks, For the j The key matrix of the submatrices, , For the j The sub-matrix S The key matrix of the timing blocks, For the j The value matrix of the submatrices, , No. j The sub-matrix S The value matrix of a time series block, is the spatial dimension.
5. The time series prediction method based on hierarchical self-attention mechanism according to claim 1 is characterized in that: The inter-class attention layer divides the first temporal representation into blocks, uses the mapping attention mechanism to capture the original spatial relationship between the temporal blocks in the first temporal representation, and uses the enhanced attention mechanism to enhance the temporal blocks in the first temporal representation to obtain the second temporal representation, which specifically includes: Dividing the first time series representation into blocks to obtain a plurality of representation blocks; For any representation block, according to the query weight matrix, key weight matrix and value weight matrix of the representation block, the representation block is mapped from the time length to the multidimensional space to obtain the query matrix, key matrix and value matrix of the representation block; the query weight matrix, key weight matrix and value weight matrix of each representation block are pre-trained; According to the query matrix, key matrix and value matrix of the representation block, a multi-layer perceptron and a multi-head attention mechanism are used to capture the original spatial relationship between time series to obtain a mapping attention representation of the representation block; Performing a maximum pooling operation on the representation block to obtain an enhanced attention representation of the representation block; Combining the mapped attention representation and the enhanced attention representation of the representation block to obtain an inter-class attention representation of the representation block; The inter-class attention representations of multiple representation blocks are concatenated to obtain the second temporal representation.
6. The time series prediction method based on hierarchical self-attention mechanism according to claim 5 is characterized in that: Use the following formula to get r The mapping attention representation of the representation block is: ; in, For the r The mapping attention representation of the representation block, is the normalized exponential function, For the r A query matrix representing a block, For the r A key matrix representing a block, For the r A matrix of values representing blocks, and All are pre-trained multi-layer perceptrons. d is the spatial dimension.
7. The time series prediction method based on hierarchical self-attention mechanism according to claim 5 is characterized in that: Use the following formula to get r Inter-class attention representation of representation blocks: ; in, For the r The inter-class attention representation of the representation block, For the r The mapping attention representation of the representation block, For the r The enhanced attention representation of the representation block, For standardized processing.
8. The time series prediction method based on hierarchical self-attention mechanism according to claim 1, characterized in that: The decoder is a multi-layer perceptron.
9. The time series prediction method based on hierarchical self-attention mechanism according to claim 1, characterized in that: The historical time series is historical traffic flow monitoring data; the future time series is future traffic flow monitoring data.
10. A time series prediction system based on a hierarchical self-attention mechanism, applied to the time series prediction method based on a hierarchical self-attention mechanism according to any one of claims 1 to 9, characterized in that: The time series prediction system based on the hierarchical self-attention mechanism includes: A classification module is used to perform cluster analysis on the historical time series, divide the historical time series into multiple categories, and obtain multiple sub-matrices; each sub-matrix includes the time series of the same category; A prediction module, used to perform time series prediction based on multiple sub-matrices using a pre-trained time series prediction model to obtain a future time series; the time series prediction model includes an intra-class attention layer, an inter-class attention layer and a decoder; The intra-class attention layer is used to divide each sub-matrix into blocks respectively, and adopts a self-attention mechanism to learn the time dependency between each time series block in each sub-matrix to obtain a first time series representation; The inter-class attention layer is used to divide the first time series representation into blocks, use a mapping attention mechanism to capture the original spatial relationship between each time series block in the first time series representation, and use an enhanced attention mechanism to enhance each time series block in the first time series representation to obtain a second time series representation; The decoder is used to map the second time series representation into a future time series.
Citation Information
Patent Citations
Short-term time sequence prediction method and system based on time and space attention
CN115081586A
Prediction method, device and system based on multivariable time series
CN115345220A
Short-term commodity demand prediction method based on cross-sequence
CN117217809A
Deep learning work order quantity prediction method based on wavelet multi-resolution decomposition
CN118657255A
Elevator risk early warning method based on local patch and enhanced position coding
CN119503573A