Multimodal Spatiotemporal Sequence Prediction Method and Device for Visualization of Multi-Scale Time-Series Data
Through the fusion of multi-scale spatiotemporal data imagery and multimodal feature, the problems of insufficient dynamic spatiotemporal mode capture and high computational complexity in the long-term span of the existing technology are solved, and more efficient multivariate spatiotemporal sequence prediction is achieved.
Patent Information
- Application Number
- CN202510578525.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-05-07
AI Technical Summary
The existing multivariate spatiotemporal sequence prediction methods fail to fully capture dynamic spatiotemporal patterns within a long-term span, have high computational complexity, and lack the effective fusion ability of multimodal data, resulting in limited prediction performance.
By dividing single-scale spatiotemporal sequence data into multi-scale long and short-term data, it is converted into fixed-resolution image data, combined with convolutional neural networks and spatiotemporal bottleneck attention networks, adaptively weighted fusion multimodal features are used for prediction.
It improves prediction performance and convergence speed, enhances the learning effect of spatiotemporal features, solves the problem of high computational complexity, and is suitable for multivariate spatiotemporal sequence prediction tasks such as traffic flow and stock prices.
Smart Images

Figure CN120105347B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of time series prediction, and in particular, to a multi-modal spatio-temporal sequence prediction method and device for visualizing multi-scale time series data. Background Art
[0002] Multivariate spatio-temporal sequence prediction has important application values in fields such as modern intelligent transportation, environmental monitoring, and financial analysis. Typical multivariate spatio-temporal sequence data includes traffic flow, stock prices, electroencephalogram signals, and temperature distributions, etc. These data usually contain multi-dimensional variable information such as location and timestamp. The core of the multivariate spatio-temporal sequence prediction problem lies in mining the complex, heterogeneous, and dynamic spatio-temporal feature distribution laws to achieve accurate prediction of different nodes at future time steps.
[0003] Existing multivariate spatio-temporal sequence prediction methods mainly rely on models such as graph convolutional neural networks and recurrent neural networks. For example, the spatial dimension correlation and scale dependence are extracted through graph convolutional neural networks, or the temporal dimension temporal correlation, periodicity, and trend features are extracted through recurrent neural networks and their variants (such as long short-term memory networks and gated recurrent units) and Transformer models. However, although the existing technologies have made remarkable progress in spatio-temporal feature extraction, there are still some key problems to be solved.
[0004] First of all, most existing studies construct prediction models based on short-term spatio-temporal data, usually using historical data with no more than 12 time steps as input. This method only focuses on learning spatio-temporal features from short-term data and ignores the continuous and dynamic spatio-temporal features hidden in longer historical data. Although some studies attempt to introduce data with a longer time span (such as daily or weekly data), these methods often only focus on the spatio-temporal features within the same period and fail to fully capture the dynamically changing spatio-temporal patterns within a long time span. This limitation results in the model being unable to deeply mine the potential deep features in spatio-temporal data, thus restricting the further improvement of prediction performance. Secondly, some studies have begun to explore using longer historical spatio-temporal data for modeling, but these methods usually rely on single-scale spatio-temporal data and are difficult to comprehensively express deep spatio-temporal features. In addition, due to the increase in the amount of input data, the convergence speed of the model is slow, affecting the efficiency in practical applications. Finally, in some time series prediction tasks, most existing models are constructed based on single-modal spatio-temporal data sources and lack the ability to effectively fuse multi-modal data. Although some studies attempt to introduce external data (such as weather data or point of interest information) to enhance the expression ability of time series data, they still face great challenges in feature selection and data denoising, restricting the robustness and adaptability of the model.
[0005] In view of this, the applicant has specifically proposed this application after studying the existing technologies. Summary of the Invention
[0006] The present invention aims to provide a multi-modal spatio-temporal sequence prediction method and device for visualizing multi-scale time-series data, so as to solve the problems of high computational complexity and slow convergence rate in the prior art.
[0007] To solve the above technical problems, the present invention is realized through the following technical solutions:
[0008] A multi-modal spatio-temporal sequence prediction method for visualizing multi-scale time-series data, comprising:
[0009] S1. Obtain single-scale spatio-temporal sequence data;
[0010] S2. Combine a sliding window with a preset length to divide the single-scale spatio-temporal sequence data to obtain single-scale short-term spatio-temporal data for each sliding window;
[0011] S3. Perform step expansion on the single-scale short-term spatio-temporal data to generate multiple single-scale long-term spatio-temporal data with different scales and fuse them to obtain a multi-scale long-short-term spatio-temporal data set;
[0012] S4. Convert the multi-scale long-short-term spatio-temporal data set into image data with a fixed resolution to obtain multi-scale long-term image data;
[0013] S5. Use a convolutional neural network to extract spatio-temporal features of the multi-scale long-term image data;
[0014] S6. Use a spatio-temporal bottleneck attention network to extract spatio-temporal features of the single-scale short-term spatio-temporal data;
[0015] S7. Perform adaptive weighted fusion on the spatio-temporal features respectively extracted from the single-scale short-term spatio-temporal data and the multi-scale long-term image data to obtain a prediction result of the spatio-temporal sequence.
[0016] Preferably, the single-scale spatio-temporal sequence data includes traffic flow data, stock price data, electroencephalogram data, humidity data, and temperature data of a single time scale; specifically, S2 is as follows:
[0017] Suppose the obtained single-scale spatio-temporal sequence data is , where is the dimension representation symbol of X , C represents the number of nodes, and any node ; T represents the total number of time steps, time , Q represents the step to be predicted;
[0018] Suppose the length of the sliding window isP , each sliding window is represented as , then the spatio-temporal sequence of each sliding window, represented as :
[0019] ;
[0020] Among them, is the single-scale short-term spatio-temporal data; represents the node c at a certain time point t of P spatio-temporal data.
[0021] Preferably, define the time scale corresponding to the single-scale short-term spatio-temporal data as the first scale, and its time interval is , which contains P spatio-temporal data;
[0022] By performing forward step expansion on the data of the first scale, multi-scale long-time series spatio-temporal data is generated, and the generation method is:
[0023] Add adjacent two data in the of P data to obtain a new data set, and its time interval is , and at this time the number of the new data set is half of the number P , that is, generate the spatio-temporal data in the second half of the second scale;
[0024] For forward expand a sliding window to obtain the corresponding single-scale short-term spatio-temporal data, that is, the spatio-temporal data of 1 historical P steps , add adjacent two data in to obtain the spatio-temporal data in the first half of the second scale;
[0025] Then, add of P consecutive s data in the s scale to generate the P / s new data set in the second half of the scale ; for forward expand s -1 sliding window to obtain the single-scale short-term spatio-temporal data within the corresponding sliding window, that is, the steps of single-scale long-term spatio-temporal data , add Add two adjacent data to obtain the s remaining data at the scale, and ;
[0026] Stitch and fuse the generated single-scale spatio-temporal data at multiple different scales to obtain a multi-scale long-term and short-term spatio-temporal data set. The expression is:
[0027] ;
[0028] Among them, represents the P short-term spatio-temporal data at the first scale, represents the s generated P long-term spatio-temporal data at the scale, is the multi-scale long-term and short-term spatio-temporal data set corresponding to the node c , d represents the number of scales for stitching.
[0029] Preferably, the S4 is specifically:
[0030] Plot the multi-scale long-term and short-term spatio-temporal data set c corresponding to any node as d scatter sub-plots;
[0031] Stack the scatter sub-plots vertically into a column to generate an RGB image with three channels; among them, , respectively represent the height and width of the image;
[0032] By fine-tuning d generate an RGB image containing d spatio-temporal data corresponding to different time scales. Each RGB image contains d spatio-temporal data corresponding to the scatter sub-plots at different time scales;
[0033] Then, C nodes will generate C RGB images ;
[0034] Convert C RGB images into a grayscale image with C channels to better extract the spatial features of multiple nodes;
[0035] Convert the grayscale image After dimensional conversion to , that is, using a grayscale image containing C channels to represent C nodes at a certain time point t of the multi-scale long short-term spatio-temporal data set to reduce the computational complexity.
[0036] Preferably, during the process of expanding and padding the new data set:
[0037] If t ∈ P , s × P -1], then the remaining P - P / s data are filled with zeros;
[0038] If t ∈ s × P , T - Q , then the corresponding spatio-temporal sequence data at the first scale are introduced s × P to generate the spatio-temporal data of the s th scale, and the expression is: P ;
[0039] ;
[0040] ;
[0041] Among them, represents the aggregated data containing the s th scale after expansion and padding. P
[0042] Preferably, the convolutional neural network is a lightweight ConvNeXt-T, including a convolutional unit, a pooling layer, a layer normalization and a linear transformation layer;
[0043] The multi-scale long-term image data undergoes multiple-stage convolutional downsampling operations through the convolutional unit, and then the output results of the convolutional unit are successively subjected to global average pooling operations of the pooling layer, normalization operations of the layer normalization, and channel dimension transformation operations of the linear transformation layer to output the learned spatio-temporal features.
[0044] Preferably, the convolutional unit contains a pre-convolutional module and four cascaded stage structures, and each of the four stages contains several ConvNeXt blocks;
[0045] Among them, the pre-convolution module consists of a convolutional layer with a kernel size of 4×4 and a stride of 4, which is used to perform preliminary downsampling and output the input features of the first stage;
[0046] The first stage contains 3 ConvNeXt blocks, each of which is composed of a depthwise convolutional layer with a kernel size of 7×7 and two pointwise convolutional layers with a kernel size of 1×1 stacked in sequence, which is used to enhance the local receptive field and improve the feature expression ability. At the end of it, a convolutional layer with a kernel size of 2×2 and a stride of 2 is set, which is used to further downsample and obtain the output features of the first stage, that is, the input features of the second stage;
[0047] The second stage contains 3 ConvNeXt blocks, and at the end of it, a convolutional layer with a kernel size of 2×2 and a stride of 2 is set to obtain the output features of the second stage, that is, the input features of the third stage;
[0048] The third stage contains 9 ConvNeXt blocks, and a convolutional layer with a kernel size of 2×2 and a stride of 2 is also set to obtain the output features of the third stage, that is, the input features of the fourth stage;
[0049] The fourth stage contains 3 ConvNeXt blocks, without downsampling operation, and finally outputs the high-order expression features of the convolutional unit.
[0050] Preferably, the spatio-temporal bottleneck attention network includes a spatio-temporal encoder, a transform attention block, and a spatio-temporal prediction decoder;
[0051] Among them, the spatio-temporal encoder consists of 3 sequentially cascaded spatio-temporal bottleneck attention blocks with residual connections, which is used to model the spatial dependence between nodes and the temporal correlation in the sequence dimension, so as to extract single-scale short-term spatio-temporal features and generate the spatio-temporal feature tensor corresponding to the single-scale spatio-temporal sequence data;
[0052] The transform attention block further models the spatio-temporal feature tensor, enhances the expression ability of key features by suppressing irrelevant information, and outputs global attention features;
[0053] The spatio-temporal prediction decoder consists of 3 sequentially cascaded spatio-temporal bottleneck attention blocks with residual connections, which is used to map the attention features into a single-scale future spatio-temporal representation and realize the prediction of the target sequence.
[0054] Preferably, the S7 is specifically:
[0055] Initialize two learnable parameters and , and generate the corresponding weights through softmax function for normalization and , the expression is:
[0056] ;
[0057] ;
[0058] in, e Represents a natural constant, used to calculate exponents;
[0059] pass and The single-scale short-term spatiotemporal data are Extracted spatiotemporal features With the multi-scale long-term image data Extracted spatiotemporal features Perform weighted fusion to obtain the final prediction result, which is expressed as:
[0060]
[0061] in, The final prediction result; is the spatiotemporal features extracted from the single-scale short-term spatiotemporal data; is the spatiotemporal features extracted from the multi-scale long-term image data, t Indicates a point in time.
[0062] The present invention also provides a multi-modal spatiotemporal sequence prediction device for multi-scale time series data visualization, comprising:
[0063] An acquisition unit, used to acquire single-scale spatiotemporal series data;
[0064] A sliding window division unit, used to divide the single-scale spatiotemporal sequence data into single-scale short-term spatiotemporal data of each sliding window in combination with a sliding window of a preset length;
[0065] An aggregation unit, used for performing step expansion on the single-scale short-term spatiotemporal data, generating a plurality of single-scale long-term spatiotemporal data of different scales and fusing them to obtain a multi-scale long-term and short-term spatiotemporal data set;
[0066] An imaging unit, used for converting the multi-scale long-term and short-term spatiotemporal data set into image data with a fixed resolution to obtain multi-scale long-term image data;
[0067] A multi-scale long-term image data feature extraction unit, used to extract the spatiotemporal features of the multi-scale long-term image data using a convolutional neural network;
[0068] A single-scale short-term spatiotemporal data feature extraction unit, used to extract the spatiotemporal features of the single-scale short-term spatiotemporal data using a spatiotemporal bottleneck attention network;
[0069] The adaptive fusion prediction unit is used to adaptively weight the spatiotemporal features extracted from the single-scale short-term spatiotemporal data and the multi-scale long-term image data to obtain the prediction result of the spatiotemporal sequence.
[0070] The present invention also provides a spatiotemporal series prediction device based on multi-scale time series data visualization, comprising a processor and a memory, wherein the memory stores a computer program, and the computer program can be executed by the processor to implement a spatiotemporal series prediction method based on multi-scale time series data visualization as described above.
[0071] The present invention also provides a computer-readable storage medium, on which computer-readable instructions are stored. When the computer-readable instructions are executed by a processor of a device where the computer-readable storage medium is located, a multimodal spatiotemporal series prediction method for visualizing multi-scale time series data as described above is implemented.
[0072] In summary, compared with the prior art, the present invention has the following beneficial effects:
[0073] (1) The present invention adopts a multi-scale time series data visualization method to convert multiple spatiotemporal series data of different single time scales into an image with a fixed resolution, so that a single image contains multiple long-term time series data of any combination at different time scales, thereby enhancing the feature expression capability of the input data.
[0074] (2) The present invention adaptively weightedly fuses the spatiotemporal features extracted from single-scale short-term spatiotemporal series data and the spatiotemporal features extracted from multi-scale long-term image data, i.e., features of two different modalities, so that the model can learn more comprehensive and diverse spatiotemporal dependencies, thereby effectively improving the prediction performance of the model.
[0075] (3) The present invention transforms the spatiotemporal prediction problem based on spatiotemporal feature learning into a spatiotemporal prediction problem based on the parallel learning of image features and spatiotemporal feature values. By leveraging the rich and effective feature extraction model of computer vision, the present invention deeply explores and learns the spatiotemporal features of multivariate spatiotemporal data from a new perspective, thereby enhancing the effect of spatiotemporal feature learning.
[0076] (4) Through the multi-scale spatio-temporal data visualization and multi-modal feature fusion design, the present invention reduces the computational complexity, solves the deficiencies in spatio-temporal feature learning of the prior art, and significantly improves the prediction performance and convergence speed. Verified by embodiments, this method is particularly applicable to multi-variable spatio-temporal sequence prediction tasks such as traffic flow prediction and stock price prediction, and has important practical application value. In addition, the present invention can avoid problems such as the misalignment and insufficient denoising of multi-modal data from different sources, and expands the research ideas of existing spatio-temporal sequence prediction methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0077] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as a limitation of the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0078] Figure 1 Schematic diagram of a multi-modal spatio-temporal sequence prediction method for multi-scale time series data visualization provided for Embodiment 1.
[0079] Figure 2 Overall framework schematic diagram of a multi-modal spatio-temporal sequence prediction method for multi-scale time series data visualization provided for Embodiment 1.
[0080] Figure 3 Schematic diagram of the process of aggregating single-scale short-term spatio-temporal data into multi-scale long-term spatio-temporal data provided for Embodiment 1.
[0081] Figure 4 Schematic diagram of the process of visualizing multi-scale long-term spatio-temporal passenger flow data of a certain city subway station provided for Embodiment 1.
[0082] Figure 5 Schematic diagram of the feature learning process based on multi-scale long-term image data provided for Embodiment 1.
[0083] Figure 6(a) is a comparison diagram of the convergence speeds of the STID, STAEformer, and MSTSI models on the BJMetro dataset provided for Embodiment 1.
[0084] Figure 6(b) is a comparison diagram of the convergence speeds of the PDFormer, DDGCRN, and MSTSI models on the BJMetro dataset provided for Embodiment 1.
[0085] Figure 7 Schematic diagram of a multi-modal spatio-temporal sequence prediction device for multi-scale time series data visualization provided for Embodiment 2.
[0086] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. Specific Embodiment
[0087] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the present invention to be protected, but merely represents the selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0088] Embodiment 1
[0089] Embodiment 1 of the present invention provides a multi-modal spatio-temporal sequence prediction method for visualizing multi-scale time-series data, which can be implemented by a multi-modal spatio-temporal sequence prediction device for visualizing multi-scale time-series data (hereinafter referred to as the prediction device), and in particular, is executed by one or more processors in the prediction device.
[0090] In this embodiment, the prediction device may be an electronic device equipped with a processor, and the processor has a computer program for the multi-modal spatio-temporal sequence prediction method for visualizing multi-scale time-series data and the computer program can be executed, such as a computer, a smart phone, a smart tablet, a workstation, etc., which is not limited herein.
[0091] As Figure 1 - Figure 2 shown, a multi-modal spatio-temporal sequence prediction method for visualizing multi-scale time-series data includes steps S1 to S7.
[0092] S1. Obtain single-scale spatio-temporal sequence data.
[0093] In this embodiment, the single-scale spatio-temporal sequence data contains information data of multi-dimensional variables such as position and time stamp, and can be observation values from different sensors and different time points. Typical multi-variable spatio-temporal sequence data includes: traffic flow data, stock price data, electroencephalogram data, temperature / humidity data, etc. The multi-variable spatio-temporal sequence prediction problem focuses on mining the complex, heterogeneous and dynamic spatio-temporal feature distribution laws in spatio-temporal data to achieve accurate prediction of different nodes at different time points.
[0094] For example, based on sensors at 500 intersections in a city, the single-scale traffic flow data is counted every 5 minutes. The spatiotemporal series data includes:
[0095] Location: 116.4°E, 39.9°N (Xidan intersection);
[0096] Timestamp: 2025-04-16 08:00:00;
[0097] Multivariate: Predict traffic flow for 500 road segments in the next half hour.
[0098] For example, the single-scale weather data sent back every hour by 100 meteorological stations on the Qinghai-Tibet Plateau, such as ultraviolet intensity, PM2.5 concentration, wind speed, etc., are as follows:
[0099] Spatial dimension: Nagqu Station at an altitude of 4,700 meters;
[0100] Time dimension: 14:00 every day during the rainy season in 2025;
[0101] Multivariate: Forecast weather conditions for 100 weather stations for the next hour.
[0102] By capturing the spatiotemporal coupling effect of altitude and meteorological parameters, severe convective weather can be warned in advance.
[0103] For example, the smart bracelet data of multiple patients for 30 consecutive days, such as heart rate, blood oxygen saturation, etc., are as follows:
[0104] Spatial anchor point: bedroom (relative position when GPS is in sleep state);
[0105] Time granularity: heart rate fluctuation per minute (from 72 to 58 beats / minute during sleep);
[0106] Multivariate: Predicting the risk of apnea at 3 a.m. in multiple patients.
[0107] Specifically, Figure 2 As shown in the figure, the model MSTSI (Multimodal Spatio-Temporal Time Series Prediction Model Based on Multi-Scale Time Series Data Imaging) corresponding to the multimodal spatio-temporal series prediction method based on multi-scale time series data imaging proposed in the present invention is mainly composed of Figure 2It consists of four parts: (a) a feature learning module based on single-scale spatio-temporal data, (b) a conversion module for visualizing multi-scale spatio-temporal data, (c) a feature learning module based on multi-scale image data, and (d) an adaptive fusion and prediction module for multi-modal features.
[0108] Let the original input single-scale spatio-temporal sequence data be in numerical form. Among them, is the dimension representation symbol of X , C represents the number of nodes in the entire network, T represents the total number of time steps. represents any node c at any time point t of the spatio-temporal data, , , Q represents the step length to be predicted.
[0109] S2. Combine a sliding window with a preset length to divide the single-scale spatio-temporal sequence data to obtain single-scale short-term spatio-temporal data for each sliding window.
[0110] Let the length of the sliding window be P , for example, let P = 12, and each sliding window is represented as .
[0111] As shown in the following formula, the single-scale spatio-temporal sequence data P with a sliding window of length , that is, the single-scale short-term spatio-temporal data, is represented as , and the specific expression is:
[0112] .
[0113] As Figure 2 shown, the spatio-temporal sequence data C of t nodes within a sliding window range at any time point is represented as:
[0114] ;
[0115] Among them, represents the single-scale short-term spatio-temporal data of node C at any time point t , that is, the spatio-temporal data within a sliding window.
[0116] S3. Expand the step length of the single-scale short-term spatio-temporal data to generate multiple single-scale long-term spatio-temporal data with different scales and fuse them to obtain a multi-scale long-short-term spatio-temporal data set.
[0117] As Figure 3 shown, it details the process of aggregating into multi-scale spatio-temporal data.
[0118] Specifically, the time scale of single-scale short-term spatio-temporal data is defined as the first scale (i.e., the scale 1 in Figure 3 , the spatio-temporal data within a sliding window), and its time interval is minutes. Then, the values of two consecutive data in are added together to generate a new data, and the scale of the corresponding new data set is defined as the second scale (i.e., the scale 2 in Figure 3 ). For example, taking 4 data as an example, it is expressed as:
[0119] .
[0120] Therefore, P data on scale 1 spanning the same time span can generate data on scale 2. The time interval of the data on scale 2 is minutes. To keep the number of data on scale 2 equal to that on scale 1, both being P , another P steps of P data from the forward extended history on scale 1 are introduced to fill the other data on scale 2. For example, filling the
[0121] data according to the above 4-data example is expressed as:
[0122] Through such data extension, the number of data on scale 2 also reaches P ones, the same as the number of data on scale 1. Following this idea, the above extension and filling generation process is defined as:
[0123] ;
[0124] ;
[0125] where represents the P aggregated data on scale 2 after extension and filling.
[0126] It should be noted that if t ∈ P , 2 P -1], then those used to generate in scale 2 PThe data corresponding to Scale 1 is missing to varying degrees. Therefore, zero-padding needs to be performed on the missing data on Scale 1. The following formula defines in detail the three zero-padding strategies for generating t the data on Scale 2 from the missing data on Scale 1 at different times: P
[0127] ;
[0128] After the same processing, we can generate the s -scale P aggregated data. The time interval of the data on Scale s is minutes. Such a multi-scale aggregation strategy for any scale s can be expressed as:
[0129] ;
[0130] ;
[0131] where represents the s -th scale P aggregated spatio-temporal data, and .
[0132] Specifically, the P data on Scale 1 will generate the s -scale P / s data without any padding; the remaining data needs to be padded. If , then the data will be padded with zeros; if , where Q represents the step size to be predicted, then based on the s × P actual spatio-temporal sequence data on Scale 1, the s -scale P data is generated.
[0133] Finally, the d single-scale P data generated are concatenated to obtain the multi-scale long short-term spatio-temporal data set shown in the following formula .
[0134] ;
[0135] Then, the set of multi-scale long short-term spatio-temporal data containing C nodes It can be expressed by the following formula:
[0136] Wherein, ; d represents the number of splicing scales; represents any node c corresponding multi-scale long-term and short-term spatio-temporal data set; represents the node C corresponding multi-scale long-term and short-term spatio-temporal data set.
[0137] In this embodiment, the short-term spatio-temporal data refers to the spatio-temporal data within a sliding window, and the long-term spatio-temporal data refers to the spatio-temporal data exceeding one sliding window.
[0138] S4. Convert the multi-scale long-term and short-term spatio-temporal data set into image data with a fixed resolution to obtain multi-scale long-term image data.
[0139] Since contains spatio-temporal data of multiple scales, as the scale increases from 1 to d , the step size of the corresponding spatio-temporal data will also increase. This makes the calculation cost of training become higher. To overcome this problem, the present invention specifically designs a conversion module for visualizing multi-scale spatio-temporal data, which is used to convert into C pictures .
[0140] Specifically, first, plot the multi-scale long-term and short-term spatio-temporal data set corresponding to any node c as scatter subgraphs. d
[0141] Then, stack these subgraphs vertically into a column to generate an RGB picture , which has three channels and the resolution is strictly limited to a fixed size, such as H × W fixed at 128×128 pixels.
[0142] Next, by fine-tuning the number of scales d , a picture containing spatio-temporal data corresponding to d different time scales can be flexibly generated. Each picture contains spatio-temporal data corresponding to the different time scales of d subgraphs. Through such conversion, the multi-scale spatio-temporal data in the form of pictures Spatio-temporal data with any number of different time scale combinations can be ingeniously included. At the same time, ensure that the computational complexity of the image is independent of the length of the input spatio-temporal data and the number of time scales, so as to solve the problem of the high computational cost of multi-scale spatio-temporal data during training due to multiple scale combinations.
[0143] Generally, at any point in time t , including C a set of multi-scale long-short-term spatio-temporal data with nodes can be converted into C RGB images with three channels . The process of visualizing multi-scale long-term spatio-temporal data is expressed as:
[0144] ;
[0145] where represents the visualization function.
[0146] To reduce the computational complexity, continue to convert C images into a grayscale image with C channels . Then, through dimensional transformation, is converted into . In this way, a grayscale image containing C channels can be ingeniously used to represent, at any point in time t , the original C set of multi-scale long-short-term spatio-temporal data with nodes, which helps to better extract the spatial features between multiple nodes.
[0147] As Figure 4 shown, in a specific application, a process of visualizing multi-scale long-term spatio-temporal passenger flow data of a certain city's subway station is provided. The time interval of the original passenger flow input data is 5 minutes, corresponding to scale 1. We define the time intervals of scales 2, 6, and 12 as 10, 30, and 60 minutes respectively. Therefore, at scales 1, 2, 6, and 12, P = 12 time steps of passenger flow corresponding to time spans of 17:00 to 18:00, 16:00 to 18:00, 12:00 to 18:00, and 6:00 to 18:00 respectively, corresponding to time spans of 60, 120, 360, and 720 minutes. In this example, d = 4, P = 12, indicating that 4 single-scale spatio-temporal data are combined together, and each scale uses 12 time steps of data, generating an image with a resolution ofH × W The picture of W . Obviously, the larger the single-scale time span is, the longer the step length of the input historical spatio-temporal data is, so that the generated picture contains richer long-term historical passenger flow data.
[0148] Therefore, for P to T - Q within the time span of C nodes, we can obtain a total of M pictures , where M = T - Q - P +1. These pictures are used as the input of the feature learning module based on multi-scale image data.
[0149] S6. Use a convolutional neural network to extract the spatio-temporal features of the multi-scale long-term image data.
[0150] As Figure 5 shown in the feature learning process based on image data. In this embodiment, a lightweight convolutional neural network ConvNeXt-T is used, including a convolutional unit (ConvNeXt-T), a global average pooling layer (Global AveragePooling, GAP), a layer normalization (Layer Normalization, LN), and a linear transformation layer (Linear), to extract features from the picture .
[0151] Specifically, the multi-scale long-term image data undergoes downsampling operations in multiple stages through the convolutional unit, and then the output result of the convolutional unit passes through global average pooling of the pooling layer, normalization of the layer normalization, and channel dimension transformation of the linear transformation layer in sequence to obtain spatio-temporal features .
[0152] In this embodiment, the feature learning process based on the convolutional unit is divided into four stages, and each stage contains several ConvNeXt blocks. For example, the number of channels in each stage can be set to =(96, 192, 384, 768), and the number of ConvNeXt blocks in each stage is =(3, 3, 9, 3), and the quantity can also be set according to actual needs, which is not limited here.
[0153] As Figure 5 shown in the picture at any time point t Taking as an example, the feature extraction process in the feature learning module based on image data is described. First, It is downsampled through a convolutional layer with a convolutional kernel size of 4×4 and a stride of 4 in the convolutional unit to obtain a feature map , which is the input of the first stage. Next, the first stage passes through 3 ConvNeXt blocks, each ConvNeXt block is composed of a 7×7 depth convolutional layer and two 1×1 pointwise convolutional layers stacked together, and a convolutional layer with a convolutional kernel size of 2×2 and a stride of 2 for downsampling, obtaining the output of the first stage, that is, the input of the second stage . The second stage contains 3 ConvNeXt blocks, and a convolutional layer with a convolutional kernel size of 2×2 and a stride of 2 is set at its end to obtain the output features of the second stage, that is, the input features of the third stage. The third stage contains 9 ConvNeXt blocks, and a convolutional layer with a convolutional kernel size of 2×2 and a stride of 2 is also set to obtain the output features of the third stage, that is, the input features of the fourth stage. The fourth stage contains 3 ConvNeXt blocks, without downsampling operation, and finally outputs the high-order expression features of the convolutional unit.
[0154] Following this process, we can obtain the final output after four stages . Then, the global average pooling GAP is used to calculate the average value channel by channel . After that, layer normalization LN is used to standardize it to obtain . Finally, through a Linear linear transformation, is transformed to , and the dimension of is transformed into . The above process is defined as the function corresponding to the following formula :
[0155] ;
[0156] In addition, in terms of spatial feature learning, by utilizing the natural property of channel interaction in convolutional operations, features can be effectively extracted from , and the spatial dependence relationships between all nodes are equivalently learned. In terms of temporal feature learning, extracting features from each channel of can equivalently learn the short-term and long-term spatio-temporal patterns of each node.
[0157] It should be noted that when dividing the dataset for all images, they are strictly arranged in the chronological order of the original data without random rearrangement. Therefore, for all P from T to Q within the time span of Cnodes, the features learned from the image data are .
[0158] S6. Extract the spatio-temporal features of the single-scale short-term spatio-temporal data by using the spatio-temporal bottleneck attention network.
[0159] In this step, for all P to T - Q within the time span of C nodes, the spatio-temporal data with historical P time steps is the input of the feature learning module based on the single-scale spatio-temporal data.
[0160] Select a self-supervised spatio-temporal bottleneck attention network (Self-Supervised Spatial-Temporal Bottleneck Attentive Network, SSTBAN) as the baseline method, and remove the self-supervised branch structure of SSTBAN, that is, only use the spatio-temporal bottleneck attention network (Spatial-Temporal Bottleneck Attention Network, STBAN) to learn the features based on the spatio-temporal data.
[0161] In this embodiment, STBAN is a deep learning model that combines spatio-temporal feature modeling and attention mechanism, aiming to efficiently process spatio-temporal sequence data (such as traffic flow prediction, stock prediction, video analysis, etc.). By introducing the bottleneck attention mechanism, this network can effectively capture spatio-temporal dependencies while reducing the computational complexity, thereby improving the model performance.
[0162] Taking the spatio-temporal data at any time point t as an example, describe the feature extraction process in the feature learning module based on the single-scale spatio-temporal data. First, is input into the STBAN composed of a spatio-temporal encoder, a transform attention block, and a spatio-temporal prediction decoder. In this embodiment, the spatio-temporal encoder and the spatio-temporal prediction decoder are respectively stacked with = 3 and = 3 spatio-temporal bottleneck attention blocks (Spatial-Temporal Bottleneck Attention, STBA) with residual connections to learn . Specifically, . Capture the spatial dependencies between nodes and the temporal correlations in the time series through a spatio-temporal encoder, learn single-scale short-term spatio-temporal features, extract the spatio-temporal feature tensor of single-scale short-term spatio-temporal data, and provide the input for subsequent attention calculation. Subsequently, capture the dependencies between the spatio-temporal feature tensors at different spatial positions and different time steps through a transform attention block, further enhance the feature expression ability, and reduce the dimension of attention calculation through a dimensionality reduction operation (such as linear projection), thereby reducing the computational complexity and obtaining attention features. Perform the final decoding prediction task in the spatio-temporal prediction decoder to generate the final single-scale short-term spatio-temporal features 。
[0163] The above process is represented by the spatio-temporal bottleneck attention network function as:
[0164] ;
[0165] Therefore, for all P to T - Q nodes within the time span from C the single-scale short-term spatio-temporal features learned from the single-scale short-term spatio-temporal data are , where 。
[0166] S7 adaptively weights and fuses the spatio-temporal features extracted from the single-scale short-term spatio-temporal data and the multi-scale long-term image data respectively to obtain the prediction result of the spatio-temporal sequence.
[0167] To further improve the prediction performance of MSTSI, the spatio-temporal features learned in the feature learning module based on multi-scale image data and the spatio-temporal features learned in the feature learning module based on single-scale spatio-temporal data are adaptively fused for multi-modal features.
[0168] Specifically, initialize two learnable parameters and , and generate two weights softmax and and through normalization by the softmax function. The process of normalization using the
[0169] ;
[0170] where e represents the natural constant used to calculate the exponent.
[0171] Then, pass through and respectively for and Perform weighted fusion to obtain the final prediction result , and its process is defined as follows:
[0172] ;
[0173] Finally, for the P to T - Q time span of C nodes, the predicted result can be obtained, such as traffic flow data for the future Q time step lengths. In this way, more comprehensive and diverse spatio-temporal dependency relationships can be learned, and these dependency relationships embed short-term and long-term patterns, enhancing the expression and learning of spatio-temporal data.
[0174] In summary, the method of the present invention can be specifically summarized as follows:
[0175] Assume that the entire network contains C nodes. First, at any time point , select historical C nodes, spanning T single-scale spatio-temporal sequence data ( X ) on P time step lengths of short-term spatio-temporal data, and represent it as , that is, single-scale short-term spatio-temporal data.
[0176] Then, input into the feature learning module based on single-scale short-term spatio-temporal data to learn spatio-temporal patterns .
[0177] Next, propose a conversion module for visualizing multi-scale spatio-temporal data. By extending the time step lengths to introduce more historical data, and aggregate it into a multi-scale long-short-term spatio-temporal data set .
[0178] After that, convert into C images with a fixed resolution size , and pack all C images into . Through this special processing, the deep learning ability of MSTSI for multi-node spatial features can be improved.
[0179] Furthermore, a convolutional neural network based on ConvNeXt-T is introduced to design a feature extraction module based on multi-scale image data for learning the patterns of multi-scale long-term spatio-temporal data 。
[0180] Finally, adaptive weighted fusion and is performed to learn more comprehensive and diverse spatio-temporal dependencies embedded in short-term and long-term spatio-temporal data.
[0181] Through the learning of such multi-modal features, more accurate predictions can be achieved for future Q time steps. It is worth mentioning that the learning of multi-scale long-term spatio-temporal patterns is achieved by relying on image data with a fixed resolution, and this process is completed in parallel with the learning of single-scale short-term spatio-temporal patterns. Such a design not only ensures the accuracy of the prediction performance but also greatly improves the convergence speed of the method.
[0182] In another preferred embodiment, in order to conduct experimental tests on the method of the present invention, the following data is selected as the research object:
[0183] (1) Passenger flow data of 44 stations of the Bus Rapid Transit (BRT) in City A from March 4, 2019
[0184] to March 29, 2019, namely the XMBRT dataset;
[0185] (2) Passenger flow
[0186] data of 80 stations of the subway in City B from January 1, 2019 to January 26, 2019 is used as the research object, namely the HZMetro dataset;
[0187] (3) Passenger flow data of 276 stations of the subway in City C from February 29, 2016 to April 1, 2016, namely the BJMetro dataset.
[0188] Then, evaluation is carried out using three evaluation indicators: Mean Square Error (MAE), Root Mean Square Error (RMSE), and Mean Absolute Percentage Error (MAPE).
[0189] All comparative experiments are completed under the same hardware environment and the same hyperparameter settings. The AdamW optimizer with a learning rate of 0.001 is used for training, the loss function is the Huber loss, the batch size is 16, the maximum number of training epochs is 100, and an early stopping condition with a threshold of 10 epochs is added (that is, when the loss optimization curve tends to converge and there is no change for 10 consecutive epochs, the training stops) to verify its effectiveness.
[0190] Twelve baseline models were used for the experimental comparison. Most of their hyperparameters were set according to the original text, and only the hyperparameters of some models were fine-tuned to achieve their best prediction performance as much as possible. The comparison models are as follows:
[0191] (1) Attention-Based Spatial–Temporal Graph Convolutional Network (ASTGCN-r), which is a model combining attention mechanism and spatio-temporal graph convolution, used to capture complex spatio-temporal dependencies in the traffic network.
[0192] (2) Graph WaveNet, whose core idea is to capture long-range temporal dependencies and hidden spatial dependencies through an adaptive adjacency matrix and dilated convolution, and is used for traffic speed prediction and time series modeling.
[0193] (3) Graph Multi-Attention Network (GMAN), whose core idea is to simultaneously model dynamic spatial correlations and non-linear temporal correlations through the multi-head attention mechanism, to capture spatial dependencies between nodes and non-linear correlations between different time steps, and to adaptively fuse spatial and temporal representations for traffic prediction.
[0194] (4) Spatio-Temporal Synchronous Graph Convolutional Networks (STSGCN), whose core idea is to synchronously capture complex spatio-temporal correlations through local spatio-temporal graph convolution modules, and is used for spatio-temporal network data prediction.
[0195] (5) Spectral Temporal Graph Neural Network (StemGNN), whose core idea is to simultaneously model the correlations within and between time series in the spectral domain, and is used for multivariate time series prediction.
[0196] (6) Spatio-Temporal Fusion Graph Neural Network (STFGNN), whose core idea is to learn hidden spatio-temporal dependencies through spatio-temporal fusion graphs and gated convolution modules, and is used for traffic flow prediction.
[0197] (7) Adaptive Spatial Temporal Graph Neural Network (ASTGNN). The core idea is to capture dynamic spatial dependencies through adaptive graph convolutional layers for optimizing spatio-temporal graph models.
[0198] (8) Spatio-Temporal Embedding Prediction (STEP). The core idea is to encode time and space information as low-dimensional vectors into the model to enhance prediction ability.
[0199] (9) Spatial and Temporal IDentity information Network (STID). The core idea is to improve the model's ability to distinguish different samples through spatial and temporal identity information.
[0200] (10) Propagation Delay-aware Dynamic Long-range Transformer for Traffic Flow Prediction (PDFormer). The core idea is to capture spatio-temporal dependencies in traffic flow through propagation delay awareness and dynamic long-range attention.
[0201] (11) Decomposition Dynamic Graph Convolutional Recurrent Network (DDGCRN). The core idea is to process spatio-temporal sequence data by decomposing dynamic graph convolution and recurrent networks.
[0202] (12) Spatial-Temporal Adaptive Embedding Transformer (STAEformer). The core idea is to enhance the spatio-temporal modeling ability of the Transformer model through spatio-temporal adaptive embedding.
[0203] The experimental results are shown in Table 1. It can be seen from Table 1 that the MSTSI model proposed in the present invention achieves the best performance in most metrics, especially showing superiority on the complex BJMetro dataset. This result indicates that the method of the present invention can parallelly extract single-scale short-term spatio-temporal patterns and multi-scale long-term spatio-temporal patterns and has significant prediction performance.
[0204] Table 1 Experimental Results
[0205]
[0206] In addition, it was also found that the convergence rate of the MSTSI model of the present invention is significantly faster than that of other models. To present the results more clearly, according to the number of training epochs of the experiment, the loss values of the five models on the validation set are plotted in two subgraphs, as shown in FIGS. 6(a) and 6(b), and the epochs corresponding to the lowest loss values obtained by them are highlighted with asterisks. As shown in FIGS. 6(a) and 6(b), the optimal epochs of the five models, namely the Spatio-Temporal Identity Model (STID), Spatio-Temporal Adaptive Embedding Transformer (STAEformer), Propagation Delay-Aware Dynamic Long-Range Transformer (PDFormer), Decomposed Dynamic Graph Convolutional Recurrent Network (DDGCRN), and the model of the method of the present invention (MSTSI), are 52, 42, 295, 283, and 33 respectively. Obviously, MSTSI has the fastest convergence rate and better efficiency than other models.
[0207] In summary, compared with the prior art, the present invention has the following beneficial effects:
[0208] By converting multiple single-scale spatio-temporal data into an image with a resolution limited to a fixed size, the present invention enables a single image to contain multiple long-term time-series data with arbitrary combinations on different time scales, enhancing the feature expression ability of the input data.
[0209] By adaptively fusing the features of two modalities, namely the short-term single-scale spatio-temporal features and the spatio-temporal features of the image data, the model can learn more comprehensive and diverse spatio-temporal dependencies, thereby effectively improving the prediction performance of the model.
[0210] The present invention converts the spatio-temporal prediction problem based on spatio-temporal feature learning into a spatio-temporal prediction problem based on image feature learning, and deeply excavates and learns the spatio-temporal features of multi-variable spatio-temporal data from a new perspective by means of rich and effective feature extraction models in computer vision, improving the effect of spatio-temporal feature learning.
[0211] The present invention can freely and parallelly embed the feature learning module based on image data into the existing baseline models for time series prediction. By adaptively fusing the spatio-temporal features extracted from two different modalities of the same source, the prediction ability of the model is effectively improved. This method can reduce problems such as the difficulty in aligning and denoising multi-modal data from different sources, expanding the research ideas of existing spatio-temporal prediction methods.
[0212] Embodiment 2
[0213] Such as Figure 7As shown in the figure, the second embodiment of the present invention further provides a multi-modal spatio-temporal sequence prediction device for visualizing multi-scale time-series data, including:
[0214] An acquisition unit for acquiring single-scale spatio-temporal sequence data;
[0215] A sliding window partitioning unit for partitioning the single-scale spatio-temporal sequence data by combining a sliding window with a preset length to obtain single-scale short-term spatio-temporal data for each sliding window;
[0216] An aggregation unit for performing step expansion on the single-scale short-term spatio-temporal data to generate multiple single-scale long-term spatio-temporal data with different scales and fusing them to obtain a multi-scale long-short-term spatio-temporal data set;
[0217] An imaging unit for converting the multi-scale long-short-term spatio-temporal data set into image data with a fixed resolution to obtain multi-scale long-term image data;
[0218] A multi-scale long-term image data feature extraction unit for extracting spatio-temporal features of the multi-scale long-term image data using a convolutional neural network;
[0219] A single-scale short-term spatio-temporal data feature extraction unit for extracting spatio-temporal features of the single-scale short-term spatio-temporal data using a spatio-temporal bottleneck attention network;
[0220] An adaptive fusion prediction unit for adaptively weighted fusing the spatio-temporal features extracted from the single-scale short-term spatio-temporal data and the multi-scale long-term image data respectively to obtain a prediction result of the spatio-temporal sequence.
[0221] Embodiment III
[0222] The third embodiment of the present invention further provides a multi-modal spatio-temporal sequence prediction device for visualizing multi-scale time-series data, which includes a memory and a processor. A computer program is stored in the memory, and the computer program can be executed by the processor to implement the multi-modal spatio-temporal sequence prediction method for visualizing multi-scale time-series data as described above.
[0223] Embodiment IV
[0224] The fourth embodiment of the present invention further provides a computer-readable storage medium. Computer-readable instructions are stored on the computer-readable storage medium. When the computer-readable instructions are executed by a processor of the device where the computer-readable storage medium is located, the multi-modal spatio-temporal sequence prediction method for visualizing multi-scale time-series data as described above is implemented.
[0225] In several embodiments provided by the embodiments of the present invention, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device and method embodiments described above are merely illustrative. For example, the flowcharts in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0226] In addition, in each embodiment of the present invention, the various functional modules can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part.
[0227] If the described functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, an electronic device, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs. It should be noted that in this article, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to this process, method, article, or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article, or device including the said element.
[0228] The terms used in the embodiments of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The singular forms "a", "said" and "the" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise.
[0229] It should be understood that the term "and / or" used herein is merely a description of an association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B may represent: the sole existence of A, the simultaneous existence of A and B, and the sole existence of B. Additionally, the character " / " herein generally represents an "or" relationship between the associated objects before and after.
[0230] Depending on the context, the word "if" as used herein can be interpreted as "when" or "while" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detected (stated condition or event)" can be interpreted as "when determined" or "in response to determining" or "when detected (stated condition or event)" or "in response to detecting (stated condition or event)".
[0231] The "first / second" mentioned in the embodiments is merely to distinguish similar objects and does not represent a specific order for the objects. It can be understood that the "first / second" can be interchanged in the specific order or sequence when permitted. It should be understood that the objects distinguished by the "first / second" can be interchanged appropriately so that the embodiments described herein can be implemented in an order other than those illustrated or described herein.
[0232] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, various changes and modifications can be made to the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A multi-modal spatio-temporal sequence prediction method for visualizing multi-scale time series data, characterized in that Including: S1. Obtain single-scale spatio-temporal sequence data; S2. Combine a sliding window with a preset length to divide the single-scale spatio-temporal sequence data to obtain single-scale short-term spatio-temporal data for each sliding window; S3. Perform step expansion on the single-scale short-term spatio-temporal data to generate multiple single-scale long-term spatio-temporal data with different scales and fuse them to obtain a multi-scale long-short-term spatio-temporal data set; S4. Convert the multi-scale long-short-term spatio-temporal data set into image data with a fixed resolution to obtain multi-scale long-term image data; S5. Use a convolutional neural network to extract spatio-temporal features of the multi-scale long-term image data; the convolutional neural network is a lightweight ConvNeXt-T, including a convolutional unit, a pooling layer, a layer normalization layer, and a linear transformation layer; The multi-scale long-term image data Perform convolutional downsampling operations in multiple stages through the convolutional unit, and then perform global average pooling operations of the pooling layer, normalization operations of layer normalization, and channel dimension transformation operations of the linear transformation layer on the output results of the convolutional unit in sequence to output the learned spatio-temporal features; S6. Use a spatio-temporal bottleneck attention network to extract spatio-temporal features of the single-scale short-term spatio-temporal data; the spatio-temporal bottleneck attention network includes a spatio-temporal encoder, a transform attention block, and a spatio-temporal prediction decoder; Wherein, the spatio-temporal encoder is composed of 3 sequentially cascaded spatio-temporal bottleneck attention blocks with residual connections, and is used to model the spatial dependence between nodes and the temporal correlation in the sequence dimension, so as to extract single-scale short-term spatio-temporal features and generate a spatio-temporal feature tensor corresponding to the single-scale spatio-temporal sequence data; The transform attention block models the spatio-temporal feature tensor, enhances the expression ability of key features by suppressing irrelevant information, and outputs global attention features; The spatio-temporal prediction decoder is composed of 3 sequentially cascaded spatio-temporal bottleneck attention blocks with residual connections, and is used to map the attention features into a single-scale future spatio-temporal representation to realize the prediction of the target sequence; S7. Perform adaptive weighted fusion on the spatio-temporal features respectively extracted from the single-scale short-term spatio-temporal data and the multi-scale long-term image data to obtain a prediction result of the spatio-temporal sequence; Specifically, S7 is: Initialize two learnable parameters and , and generate corresponding weights through softmax function for normalization. The expressions are as follows: and , the expression is: ; ; Among them, e represents the natural constant, which is used to calculate the exponent; pass and The single-scale short-term spatiotemporal data are Extracted spatiotemporal features With the multi-scale long-term image data Extracted spatiotemporal features Perform weighted fusion to obtain the final prediction result, which is expressed as: Among them, is the final prediction result; is the spatio-temporal feature extracted from the single-scale short-term spatio-temporal data; is the spatio-temporal feature extracted from the multi-scale long-term image data, t represents a certain time point.
2. A multi-modal spatio-temporal sequence prediction method for visualizing multi-scale time-series data according to claim 1, characterized in that , the single-scale spatio-temporal sequence data includes traffic flow data, stock price data, brain wave data, humidity data, or temperature data of a single time scale; Specifically, S2 is: Suppose the obtained single-scale spatio-temporal sequence data is , where is X the dimension representation symbol of C denotes the total number of nodes, and any node ; T denotes the total number of time steps, and the time , Q denotes the step length to be predicted; Let the length of the sliding window be P , and each sliding window is represented as . Then the spatio-temporal sequence of each sliding window is expressed as: ; Among them, is the single-scale short-term spatio-temporal data; represents the node c at a certain time point t of P spatio-temporal data.
3. A multi-modal spatio-temporal sequence prediction method for visualizing multi-scale time-series data according to claim 2, characterized in that , specifically, S3 is: Define the time scale corresponding to the single-scale short-term spatio-temporal data as the first scale, and its time interval is , including P spatio-temporal data; By performing forward step expansion on the data at the first scale to generate multi-scale long time series spatio-temporal data, the generation method is as follows: Add of P adjacent two data in the data to obtain a new data set, and the time interval is . At this time, the number of the new data set is quantity P half, that is, generate the second scale space-time data for the second half; Pair Expand a sliding window forward to obtain single-scale short-term spatio-temporal data within the corresponding sliding window, that is, spatio-temporal data for 1 historical P step , and add two adjacent data in to obtain spatio-temporal data for the first half of the second scale; Then, add consecutive P data among the s data to generate a scale s for the latter segment of P / s new data sets, with a time interval of ; For extend s -1 sliding windows forward to obtain the single-scale short-term spatio-temporal data within the corresponding sliding windows, that is, steps of single-scale long-term spatio-temporal data , add two adjacent data in to obtain the remaining data at the s-th scale, and ; Stitch and fuse the generated single-scale spatio-temporal data with multiple different scales to obtain a multi-scale long-short-term spatio-temporal data set, and the expression is: ; Among them, represents the P short-term spatio-temporal data of the first scale, represents the s th P long-term spatio-temporal data of the generated scale, represents the c corresponding multi-scale long-short-term spatio-temporal data set of the node, d represents the number of scales for splicing.
4. A multi-modal spatio-temporal sequence prediction method for visualizing multi-scale time series data according to claim 3, characterized in that , specifically, S4 is: Take any one node c The corresponding multi-scale long short-term spatio-temporal data set Plot it as d a scatter subplot; Vertically stack the scatter plots into a column to generate an RGB image with three channels ; where , respectively represent the height and width of the image ; By fine-tuning d Generate an RGB image corresponding to spatio-temporal data with d different time scales. Each RGB image contains d sub-plots of scatter points corresponding to spatio-temporal data with different time scales; Then, C nodes will generate C RGB pictures ; Convert C multiple RGB images into a grayscale image with C channels, in order to better extract the spatial features of multiple nodes; Convert the grayscale image through dimensional conversion to , that is, use a grayscale image containing C channels to correspondingly represent the multi-scale long short-term spatio-temporal data set of C nodes at a certain time point t to reduce the computational complexity.
5. A multi-modal spatio-temporal sequence prediction method for visualizing multi-scale time series data according to claim 3, characterized in that , during the process of expanding the new data set: If t ∈ P , s × P -1], then fill the remaining P - P / s data with zeros; If t ∈ s × P , T - Q , then introduce the corresponding s × P spatio-temporal sequence data at the first scale to generate the s -th spatio-temporal data at the P -th scale. The expression is: ; ; Among them, represents the s -th P aggregated data after extended filling.
6. A multi-modal spatio-temporal sequence prediction method for visualizing multi-scale time-series data according to claim 1, characterized in that , the convolutional unit includes a pre-convolutional module and four cascaded stage structures, and each of the four stages includes several ConvNeXt blocks; Among them, the pre-convolution module consists of a convolutional layer with a kernel size of 4×4 and a stride of 4, which is used to perform preliminary downsampling and output the input features of the first stage; The first stage includes 3 ConvNeXt blocks, and each ConvNeXt block is sequentially stacked by a 7×7 depth convolutional layer and two 1×1 pointwise convolutional layers, which is used to enhance the local receptive field and improve the feature expression ability, and a convolutional layer with a kernel size of 2×2 and a stride of 2 is set at the end thereof for further downsampling to obtain the output features of the first stage, that is, the input features of the second stage; The second stage contains 3 ConvNeXt blocks, and a convolutional layer with a kernel size of 2×2 and a stride of 2 is set at its end to obtain the output features of the second stage, which are the input features of the third stage; The third stage contains 9 ConvNeXt blocks, and a convolutional layer with a kernel size of 2×2 and a stride of 2 is also set to obtain the output features of the third stage, which are the input features of the fourth stage; The fourth stage contains 3 ConvNeXt blocks, without downsampling operation, and finally outputs the high-order expression features of the convolutional unit.
7. A multi-modal spatio-temporal sequence prediction device for visualizing multi-scale time series data, which is used to implement the multi-modal spatio-temporal sequence prediction method for visualizing multi-scale time series data according to any one of claims 1-6, characterized in that, Including: An acquisition unit for acquiring single-scale spatio-temporal sequence data; A sliding window partitioning unit for partitioning the single-scale spatio-temporal sequence data in combination with a sliding window of a preset length to obtain the single-scale short-term spatio-temporal data of each sliding window; An aggregation unit for performing stride expansion on the single-scale short-term spatio-temporal data, generating multiple single-scale long-term spatio-temporal data of different scales and fusing them to obtain a multi-scale long-short-term spatio-temporal data set; An imaging unit for converting the multi-scale long-short-term spatio-temporal data set into image data with a fixed resolution to obtain multi-scale long-term image data; A multi-scale long-term image data feature extraction unit for extracting spatio-temporal features of the multi-scale long-term image data by using a convolutional neural network; A single-scale short-term spatio-temporal data feature extraction unit for extracting spatio-temporal features of the single-scale short-term spatio-temporal data by using a spatio-temporal bottleneck attention network; An adaptive fusion prediction unit for adaptively weighted fusion of the spatio-temporal features respectively extracted from the single-scale short-term spatio-temporal data and the multi-scale long-term image data to obtain the prediction result of the spatio-temporal sequence.
Citation Information
Patent Citations
Short temporary rainfall prediction method and device based on multi-scale space-time consistency
CN116152620A
Multi-scale convolution cycle unit space-time sequence prediction method fusing spatial local correlation
CN117671444A