Multi-modal space-time sequence prediction method and device for multi-scale time sequence data imaging
By converting multivariate spatiotemporal sequence data into multi-scale image data and combining convolutional neural networks and spatiotemporal bottleneck attention networks, the problems of high computational complexity and difficulty in fusion of multimodal data in the prior art are solved, and more efficient spatiotemporal sequence prediction is achieved.
Patent Information
- Application Number
- CN202510578525.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-05-07
AI Technical Summary
The existing multivariate spatiotemporal sequence prediction method has high computational complexity, slow convergence speed when processing long-term spatiotemporal data, and it is difficult to effectively fusion of multimodal data, which limits the robustness and adaptability of the model.
By converting single-scale spatiotemporal sequence data into a multi-scale long and short-term spatiotemporal data set and converting it into fixed-resolution image data, the spatiotemporal features are extracted using convolutional neural networks and spatiotemporal bottleneck attention networks to perform adaptive weighted fusion to achieve prediction.
It improves the prediction performance and convergence speed of the model, enhances the expression ability of spatiotemporal characteristics, and can more effectively integrate multimodal data, which is suitable for multivariate spatiotemporal sequence prediction tasks such as traffic flow and stock prices.
Smart Images

Figure CN120105347A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of time series prediction, and in particular to a multi-modal spatiotemporal series prediction method and device for multi-scale time series data visualization. Background Art
[0002] Multivariate spatiotemporal series prediction has important application value in modern intelligent transportation, environmental monitoring, financial analysis and other fields. Typical multivariate spatiotemporal series data include traffic flow, stock prices, brain wave signals and temperature distribution, which usually contain multi-dimensional variable information such as location and timestamp. The core of the multivariate spatiotemporal series prediction problem lies in mining the complex, heterogeneous and dynamic spatiotemporal feature distribution laws to achieve accurate prediction of different nodes in the future time step.
[0003] Existing multivariate spatiotemporal series prediction methods mainly rely on models such as graph convolutional neural networks and recurrent neural networks. For example, graph convolutional neural networks are used to extract the correlation and scale dependence of spatial dimensions, or recurrent neural networks and their variants (such as long short-term memory networks and gated recurrent units) and Transformer models are used to extract temporal correlation, periodicity, and trend characteristics of the time dimension. However, although the existing technology has made significant progress in spatiotemporal feature extraction, there are still some key issues that need to be solved.
[0004] First, most existing studies build prediction models based on short-term spatiotemporal data, usually using historical data of no more than 12 time steps as input. This approach only focuses on learning spatiotemporal features from short-term data, while ignoring the continuous and dynamic spatiotemporal features hidden in longer historical data. Although some studies have tried to introduce data with longer time spans (such as daily or weekly data), these methods often only focus on spatiotemporal features within the same period and fail to fully capture the spatiotemporal patterns of dynamic changes within the long span. This limitation makes it impossible for the model to deeply explore the potential deep-level features in spatiotemporal data, thereby limiting the further improvement of prediction performance. Secondly, some studies have begun to explore the use of longer historical spatiotemporal data for modeling, but these methods usually rely on spatiotemporal data of a single scale and are difficult to fully express deep-level spatiotemporal features. In addition, due to the increase in the amount of input data, the convergence speed of the model is slow, which affects the efficiency in practical applications. Finally, in some time series prediction tasks, most existing models are built based on single-modal spatiotemporal data sources and lack the ability to effectively fuse multimodal data. Although some studies have attempted to introduce external data (such as weather data or point of interest information) to enhance the expressiveness of time series data, they still face great challenges in feature selection and data denoising, which limits the robustness and adaptability of the model.
[0005] In view of this, the applicant filed this application after studying the existing technology. Summary of the invention
[0006] The present invention aims to provide a multi-modal spatiotemporal sequence prediction method and device for multi-scale time series data visualization, so as to solve the problems of high computational complexity and slow convergence speed existing in the prior art.
[0007] In order to solve the above technical problems, the present invention is implemented through the following technical solutions: A multi-modal spatiotemporal series prediction method for multi-scale time series data visualization, comprising: S1, obtain single-scale spatiotemporal series data; S2, combining a sliding window of a preset length, dividing the single-scale spatiotemporal sequence data to obtain single-scale short-term spatiotemporal data of each sliding window; S3, performing step expansion on the single-scale short-term spatiotemporal data to generate a plurality of single-scale long-term spatiotemporal data of different scales and fusing them to obtain a multi-scale long-term and short-term spatiotemporal data set; S4, converting the multi-scale long-term and short-term spatiotemporal data set into image data with a fixed resolution to obtain multi-scale long-term image data; S5, extracting spatiotemporal features of the multi-scale long-term image data using a convolutional neural network; S6, extracting the spatiotemporal features of the single-scale short-term spatiotemporal data using a spatiotemporal bottleneck attention network; S7, adaptively weighted fusion of the spatiotemporal features extracted from the single-scale short-term spatiotemporal data and the multi-scale long-term image data to obtain a prediction result of the spatiotemporal sequence.
[0008] Preferably, the single-scale spatiotemporal series data includes traffic flow data, stock price data, brain wave data, humidity data and temperature data at a single time scale; S2 is specifically: Assume that the acquired single-scale spatiotemporal series data is ,in, for X The dimension represents the symbol, C Indicates the number of nodes. Any node ; T represents the total time step, time , Q Indicates the step size that needs to be predicted; Assume the length of the sliding window is P , each sliding window is represented as , then the spatiotemporal sequence of each sliding window is expressed as : ; in, is the single-scale short-term spatiotemporal data; Representation Node c At a certain point in time t of P spatiotemporal data.
[0009] Preferably, the time scale corresponding to the single-scale short-term spatiotemporal data is defined as the first scale, and its time interval is , including P spatiotemporal data; By analyzing the data of the first scale Perform forward step expansion to generate multi-scale long-time series spatiotemporal data. The generation method is: Will of P The two adjacent data in the data are added to get a new data set, and its time interval is , the number of new data sets is quantity P half of the second scale, that is, to generate the second half of the spatiotemporal data; right Expand a sliding window forward to obtain the corresponding single-scale short-term spatiotemporal data, that is, the spatiotemporal data of 1 historical P steps ,Will Add the two adjacent data in the first half of the second scale to get spatiotemporal data; Then, of P Continuous s Add the data to generate the scale s The latter P / s New data sets, whose time interval is ;right Scaling forward s -1 The sliding window obtains the single-scale short-term spatiotemporal data within the corresponding sliding window, that is, Single-scale long-term spatiotemporal data ,Will Add the two adjacent data in the s The remaining scale data, and ; The generated single-scale spatiotemporal data of multiple different scales are spliced and fused to obtain a multi-scale long-term and short-term spatiotemporal data set, which is expressed as: ; in, Represents the first scale P Short-term spatiotemporal data, Indicates the generateds Scale P Long-term spatiotemporal data, For Node c The corresponding multi-scale long-term and short-term spatiotemporal data sets, d Indicates the number of scales to be spliced.
[0010] Preferably, the S4 is specifically: Any node c Corresponding multi-scale long-term and short-term spatiotemporal data sets Draw d Scatter plots; Stack the scatter plots vertically into a column to generate an RGB image with three channels ;in, , Respectively represent pictures The height and width of By fine-tuning d Generate a table containing d RGB images corresponding to spatiotemporal data of different time scales, each RGB image contains d Spatiotemporal data of different time scales corresponding to the scatter plots; but, C The nodes will generate C RGB images ; Will C RGB images Convert to a C Grayscale image with 1 channel , to better extract the spatial features of multiple nodes; Grayscale image Converted to , that is, use a sheet containing C Grayscale image with 1 channel express C A node at a certain point in time t Multi-scale long-term and short-term spatiotemporal data sets are used to reduce computational complexity.
[0011] Preferably, in the process of expanding and filling the new data set: like t ∈[ P , s × P -1], then the remaining P - P / s The data is filled with zero; like t ∈[ s × P ,T - Q ], then the corresponding s × P The spatiotemporal series data generates s Scale P spatiotemporal data, the expression is: ; ; in, Indicates that after expansion and filling, it contains s Scale P Aggregate data.
[0012] Preferably, the convolutional neural network is a lightweight version of ConvNeXt-T, including a convolution unit, a pooling layer, a layer normalization and a linear transformation layer; The multi-scale long-term image data The convolution unit performs multiple stages of convolution downsampling operations, and then the output result of the convolution unit is sequentially subjected to the global average pooling operation of the pooling layer, the standardization operation of the layer normalization, and the channel dimension transformation operation of the linear transformation layer to output the learned spatiotemporal features.
[0013] Preferably, the convolution unit comprises a pre-convolution module and four cascaded stage structures, and the four stages respectively comprise a number of ConvNeXt blocks; The pre-convolution module consists of a convolution layer with a convolution kernel size of 4×4 and a step size of 4, which is used to Perform preliminary downsampling and output the input features of the first stage; The first stage contains three ConvNeXt blocks, each of which is composed of a 7×7 deep convolution layer and two 1×1 point-by-point convolution layers stacked in sequence to enhance the local receptive field and improve the feature expression ability. At the end of the ConvNeXt block, a convolution layer with a convolution kernel size of 2×2 and a stride of 2 is set for further downsampling to obtain the output features of the first stage, which are the input features of the second stage. The second stage contains 3 ConvNeXt blocks, and a convolution layer with a convolution kernel size of 2×2 and a stride of 2 is set at the end to obtain the output features of the second stage, which are the input features of the third stage; The third stage includes 9 ConvNeXt blocks, and also sets a convolution layer with a convolution kernel size of 2×2 and a stride of 2 to obtain the output features of the third stage, which are the input features of the fourth stage; The fourth stage includes 3 ConvNeXt blocks, without downsampling operation, and finally outputs the high-order expression features of the convolution unit.
[0014] Preferably, the spatiotemporal bottleneck attention network comprises a spatiotemporal encoder, a transform attention block and a spatiotemporal prediction decoder; The spatiotemporal encoder is composed of three sequentially cascaded spatiotemporal bottleneck attention blocks with residual connections, which are used to model the spatial dependency between nodes and the temporal correlation in the sequence dimension, so as to extract single-scale short-term spatiotemporal features and generate a spatiotemporal feature tensor corresponding to the single-scale spatiotemporal sequence data; The transformation attention block further models the spatiotemporal feature tensor, suppresses irrelevant information, enhances the expressiveness of key features, and outputs global attention features; The spatiotemporal prediction decoder consists of three sequentially cascaded spatiotemporal bottleneck attention blocks with residual connections, which are used to map the attention features into a single-scale future spatiotemporal representation to achieve prediction of the target sequence.
[0015] Preferably, the S7 is specifically: Initialize two learnable parameters and ,pass softmax The function is normalized to generate the corresponding weight and , the expression is: ; ; in, e Represents a natural constant, used to calculate exponents; pass and The single-scale short-term spatiotemporal data are Extracted spatiotemporal features With the multi-scale long-term image data Extracted spatiotemporal features Perform weighted fusion to obtain the final prediction result, which is expressed as:
[0016] in, The final prediction result; is the spatiotemporal features extracted from the single-scale short-term spatiotemporal data; is the spatiotemporal features extracted from the multi-scale long-term image data, t Indicates a point in time.
[0017] The present invention also provides a multi-modal spatiotemporal sequence prediction device for multi-scale time series data visualization, comprising: An acquisition unit, used to acquire single-scale spatiotemporal series data; A sliding window division unit, used to divide the single-scale spatiotemporal sequence data into single-scale short-term spatiotemporal data of each sliding window in combination with a sliding window of a preset length; An aggregation unit, used for performing step expansion on the single-scale short-term spatiotemporal data, generating a plurality of single-scale long-term spatiotemporal data of different scales and fusing them to obtain a multi-scale long-term and short-term spatiotemporal data set; An imaging unit, used for converting the multi-scale long-term and short-term spatiotemporal data set into image data with a fixed resolution to obtain multi-scale long-term image data; A multi-scale long-term image data feature extraction unit, used to extract the spatiotemporal features of the multi-scale long-term image data using a convolutional neural network; A single-scale short-term spatiotemporal data feature extraction unit, used to extract the spatiotemporal features of the single-scale short-term spatiotemporal data using a spatiotemporal bottleneck attention network; The adaptive fusion prediction unit is used to adaptively weight the spatiotemporal features extracted from the single-scale short-term spatiotemporal data and the multi-scale long-term image data to obtain the prediction result of the spatiotemporal sequence.
[0018] The present invention also provides a spatiotemporal series prediction device based on multi-scale time series data visualization, comprising a processor and a memory, wherein the memory stores a computer program, and the computer program can be executed by the processor to implement a spatiotemporal series prediction method based on multi-scale time series data visualization as described above.
[0019] The present invention also provides a computer-readable storage medium, on which computer-readable instructions are stored. When the computer-readable instructions are executed by a processor of a device where the computer-readable storage medium is located, a multimodal spatiotemporal series prediction method for visualizing multi-scale time series data as described above is implemented.
[0020] In summary, compared with the prior art, the present invention has the following beneficial effects: (1) The present invention adopts a multi-scale time series data visualization method to convert multiple spatiotemporal series data of different single time scales into an image with a fixed resolution, so that a single image contains multiple long-term time series data of any combination at different time scales, thereby enhancing the feature expression capability of the input data.
[0021] (2) The present invention adaptively weightedly fuses the spatiotemporal features extracted from single-scale short-term spatiotemporal series data and the spatiotemporal features extracted from multi-scale long-term image data, i.e., features of two different modalities, so that the model can learn more comprehensive and diverse spatiotemporal dependencies, thereby effectively improving the prediction performance of the model.
[0022] (3) The present invention transforms the spatiotemporal prediction problem based on spatiotemporal feature learning into a spatiotemporal prediction problem based on the parallel learning of image features and spatiotemporal feature values. By leveraging the rich and effective feature extraction model of computer vision, the present invention deeply explores and learns the spatiotemporal features of multivariate spatiotemporal data from a new perspective, thereby enhancing the effect of spatiotemporal feature learning.
[0023] (4) The present invention reduces computational complexity through multi-scale spatiotemporal data visualization and multi-modal feature fusion design, solves the shortcomings of the existing technology in spatiotemporal feature learning, and significantly improves prediction performance and convergence speed. The examples show that this method is particularly suitable for multivariate spatiotemporal series prediction tasks such as traffic flow prediction and stock price prediction, and has important practical application value. In addition, the present invention can avoid the problems of difficult alignment of multimodal data from different sources and insufficient denoising, and expands the research ideas of existing spatiotemporal series prediction methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.
[0025] Figure 1 A schematic diagram of a multimodal spatiotemporal series prediction method for visualizing multi-scale time series data provided in Example 1.
[0026] Figure 2 A schematic diagram of the overall framework of a multimodal spatiotemporal series prediction method for multi-scale time series data visualization provided in Example 1.
[0027] Figure 3 A schematic diagram of the process of aggregating single-scale short-term spatiotemporal data into multi-scale long-term spatiotemporal data provided in Example 1.
[0028] Figure 4 A schematic diagram of the process of visualizing multi-scale long-term spatiotemporal passenger flow data of a certain city subway station provided in Example 1.
[0029] Figure 5 A schematic diagram of a feature learning process based on multi-scale long-term image data provided in Example 1.
[0030] FIG6( a ) is a comparison chart of the convergence speeds of the STID, STAEformer, and MSTSI models provided in Example 1 on the BJMetro dataset.
[0031] FIG6( b ) is a comparison chart of the convergence speeds of the PDFormer, DDGCRN, and MSTSI models provided in Example 1 on the BJMetro dataset.
[0032] Figure 7 A schematic diagram of a multimodal spatiotemporal sequence prediction device for visualizing multi-scale time series data provided in Example 2.
[0033] The present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. DETAILED DESCRIPTION
[0034] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the invention claimed for protection, but merely represents the selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0035] Embodiment 1 Embodiment 1 of the present invention provides a multimodal spatiotemporal series prediction method for multi-scale time series data visualization, which can be implemented by a multimodal spatiotemporal series prediction device for multi-scale time series data visualization (hereinafter referred to as the prediction device), and in particular, executed by one or more processors in the prediction device.
[0036] In this embodiment, the prediction device may be an electronic device equipped with a processor, the processor having a computer program of the multimodal spatiotemporal series prediction method for multi-scale time series data visualization and the computer program can be executed, such as a computer, a smart phone, a smart tablet, a workstation, etc., which is not limited here.
[0037] like Figure 1-Figure 2 As shown, a multimodal spatiotemporal sequence prediction method for multi-scale time series data visualization includes steps S1 to S7.
[0038] S1, obtain single-scale spatiotemporal series data.
[0039] In this embodiment, the single-scale spatiotemporal series data includes information data of multi-dimensional variables such as location and timestamp, which can come from observations of different sensors and different time points. Typical multivariate spatiotemporal series data include: traffic flow data, stock price data, brain wave data, temperature / humidity data, etc. The multivariate spatiotemporal series prediction problem focuses on mining the complex, heterogeneous and dynamic spatiotemporal feature distribution laws in spatiotemporal data to achieve accurate prediction of different nodes at different time points.
[0040] For example, based on sensors at 500 intersections in a city, the single-scale traffic flow data is counted every 5 minutes. The spatiotemporal series data includes: Location: 116.4°E, 39.9°N (Xidan intersection); Timestamp: 2025-04-16 08:00:00; Multivariate: Predict traffic flow for 500 road segments in the next half hour.
[0041] For example, the single-scale weather data sent back every hour by 100 meteorological stations on the Qinghai-Tibet Plateau, such as ultraviolet intensity, PM2.5 concentration, wind speed, etc., are as follows: Spatial dimension: Nagqu Station at an altitude of 4,700 meters; Time dimension: 14:00 every day during the rainy season in 2025; Multivariate: Forecast weather conditions for 100 weather stations for the next hour.
[0042] By capturing the spatiotemporal coupling effect of altitude and meteorological parameters, severe convective weather can be warned in advance.
[0043] For example, the smart bracelet data of multiple patients for 30 consecutive days, such as heart rate, blood oxygen saturation, etc., are as follows: Spatial anchor point: bedroom (relative position when GPS is in sleep state); Time granularity: heart rate fluctuation per minute (from 72 to 58 beats / minute during sleep); Multivariate: Predicting the risk of apnea at 3 a.m. in multiple patients.
[0044] Specifically, Figure 2 As shown in the figure, the model MSTSI (Multimodal Spatio-Temporal Time Series Prediction Model Based on Multi-Scale Time Series Data Imaging) corresponding to the multimodal spatio-temporal series prediction method based on multi-scale time series data imaging proposed in the present invention is mainly composed of Figure 2The system consists of four parts: (a) a feature learning module based on single-scale spatiotemporal data, (b) a conversion module for multi-scale spatiotemporal data visualization, (c) a feature learning module based on multi-scale image data, and (d) a multi-modal feature adaptive fusion and prediction module.
[0045] Assume that the original input single-scale spatiotemporal series data is in numerical form. for X The dimension represents the symbol, C Represents the number of nodes in the entire network, T Represents the total time step. Represents any node c At any point in time t spatiotemporal data, , , Q Indicates the step size required for prediction.
[0046] S2, combining with a sliding window of preset length, dividing the single-scale spatiotemporal series data to obtain single-scale short-term spatiotemporal data of each sliding window.
[0047] Assume the length of the sliding window is P , if P = 12, each sliding window is represented by .
[0048] As shown in the following formula, P Single-scale spatiotemporal series data with sliding windows of length , that is, single-scale short-term spatiotemporal data, expressed as , specifically expressed as: .
[0049] like Figure 2 As shown, within a sliding window C A node at any point in time t Spatiotemporal series data , expressed as: ; in, Representation Node C At any point in time t Single-scale short-term spatiotemporal data, that is, spatiotemporal data within a sliding window.
[0050] S3, performing step expansion on the single-scale short-term spatiotemporal data to generate multiple single-scale long-term spatiotemporal data of different scales and fusing them to obtain a multi-scale long-term and short-term spatiotemporal data set.
[0051] like Figure 3The following table details how to The process of aggregating into multi-scale spatiotemporal data.
[0052] Specifically, the single-scale short-term spatiotemporal data The time scale is defined as the first scale (i.e. Figure 3 Scale 1 in the sliding window, spatiotemporal data), whose time interval is minutes. Then, The values of two consecutive data in are added to generate a new data set, and the scale of the corresponding new data set is defined as the second scale (i.e. Figure 3 Scale 2), for example, taking 4 data as an example, it is expressed as: .
[0053] Therefore, across the same time span scale 1 P The data can generate scale 2 The time interval of the data on scale 2 is In order to keep the number of data on scale 2 equal to that on scale 1, both P , and then introduce the forward expansion history on scale 1 P Step P Data To fill the other For example, according to the above 4 data examples, The data is filled in, expressed as: ; Through this data expansion, the amount of data on scale 2 has reached P The number of data is the same as the number of data on scale 1. According to this idea, the above expansion filling generation process is defined as: ; ; in, Represents the scale 2 after expansion and padding P Aggregate data.
[0054] It is worth noting that if t∈[ P , 2 P -1], it is used to generate scale 2 P The data corresponding to scale 1 will be missing to varying degrees. Therefore, it is necessary to perform zero filling on the missing data at scale 1. The following formula defines in detail the t Zero padding is performed on the missing data at scale 1 to generate P Three filling strategies for data of scale 2: ; After the same process, we can generate the scale s on P Aggregate data. Scale s The time interval of the above data is Minutes. Any scale s This multi-scale aggregation strategy on can be expressed as: ; ; in, Indicates the generated s Scale P Aggregated spatiotemporal data, and .
[0055] Specifically, on scale 1 P The data will generate the scale s on P / s data without any padding; the remaining data needs to be filled. If , then The data is filled with zero; if ,in Q Indicates the step size that needs to be predicted, based on scale 1 s × P The actual spatiotemporal series data generation scale s of P data.
[0056] Finally, the generated d A single scale P The data are spliced together to obtain a multi-scale long-term and short-term spatiotemporal data set as shown in the following formula: .
[0057] ; Then, containing C A collection of multi-scale long-term and short-term spatiotemporal data of nodes It can be expressed as follows:
[0058] in, ; d Indicates the number of scales for splicing; Represents any node c The corresponding multi-scale long-term and short-term spatiotemporal data sets; Representation Node C A collection of corresponding multi-scale long-term and short-term spatiotemporal data.
[0059] In this embodiment, short-term spatiotemporal data refers to spatiotemporal data within a sliding window, and long-term spatiotemporal data refers to spatiotemporal data exceeding one sliding window.
[0060] S4, converting the multi-scale long-term and short-term spatiotemporal data set into image data with a fixed resolution to obtain multi-scale long-term image data.
[0061] because Contains spatiotemporal data at multiple scales, with the scale ranging from 1 to d As increases, the step length of the corresponding spatiotemporal data will also increase. This makes training To overcome this problem, the present invention specifically designs a multi-scale spatiotemporal data image conversion module for converting Convert to C Pictures .
[0062] Specifically, first, any node c Corresponding multi-scale long-term and short-term spatiotemporal data sets Draw d A scatter plot.
[0063] Then, these sub-images are stacked vertically into a column to generate an RGB image. , the image has three channels and the resolution is strictly limited to a fixed size, such as H × W Fixed to 128×128 pixels.
[0064] Next, by fine-tuning the number of scales d , you can flexibly generate a d The pictures corresponding to the spatiotemporal data at different time scales, each picture contains d Through such conversion, the multi-scale spatiotemporal data in the form of images can be It can cleverly include any number of spatiotemporal data of different time scale combinations. At the same time, it ensures that the computational complexity of the image is independent of the length of the input spatiotemporal data and the number of time scales, so as to solve the problem of high computational cost of multi-scale spatiotemporal data during training due to multiple scale combinations.
[0065] Usually, at any point in time t ,Include C A collection of multi-scale long-term and short-term spatiotemporal data of nodes Can be converted to C An RGB image with three channels . For multi-scale long-term spatiotemporal data The process of visualization is expressed as: ; in, Represents a graphical function.
[0066] In order to reduce the computational complexity, C Pictures Convert to a C Grayscale image with 1 channel Then, we transform Convert to In this way, you can cleverly use a sheet containing C Grayscale image with 1 channel To indicate that at any point in time t , original C A collection of multi-scale long-term and short-term spatiotemporal data of nodes , which helps to better extract spatial features between multiple nodes.
[0067] like Figure 4 As shown in the figure, in a specific application, the process of visualizing the multi-scale long-term spatiotemporal passenger flow data of a certain city subway station is provided. The time interval of the original passenger flow input data is 5 minutes, corresponding to scale 1. We define the time intervals of scales 2, 6 and 12 as 10, 30 and 60 minutes respectively. Therefore, at scales 1, 2, 6 and 12, P =12 time steps of passenger flow correspond to the time spans of 17:00 to 18:00, 16:00 to 18:00, 12:00 to 18:00 and 6:00 to 18:00, and the corresponding time spans are 60, 120, 360 and 720 minutes respectively. In this example, d =4, P =12, which means that four single-scale spatiotemporal data are combined together, and each scale uses 12 time steps of data to generate a resolution of H × W Obviously, the larger the time span of a single scale, the longer the step length of the corresponding input historical spatiotemporal data, so that the generated picture contains richer long-term historical passenger flow data.
[0068] Therefore, for P arrive T - Q within the time span C nodes, we can get a total of M Pictures , where M= T - Q - P +1 for these pictures It is used as input to the feature learning module based on multi-scale image data.
[0069] S6, using a convolutional neural network to extract the spatiotemporal features of the multi-scale long-term image data.
[0070] like Figure 5 The feature learning process based on image data is shown in FIG. This embodiment uses a lightweight convolutional neural network ConvNeXt-T, including a convolution unit (ConvNeXt-T), a global average pooling layer (Global AveragePooling, GAP), a layer normalization (Layer Normalization, LN) and a linear transformation layer (Linear) to learn from the image. Extract features from .
[0071] Specifically, multi-scale long-term image data The convolution unit performs multiple stages of downsampling operations, and then the output of the convolution unit is sequentially processed through global average pooling of the pooling layer, normalization of the layer normalization, and channel dimension transformation of the linear transformation layer to obtain spatiotemporal features. .
[0072] In this embodiment, the feature learning process based on the convolution unit is divided into four stages, each of which contains several ConvNeXt blocks. For example, the number of channels in each stage can be set to =(96, 192, 384, 768), the number of ConvNeXt blocks are =(3, 3, 9, 3). The quantity can also be set according to actual needs and is not limited here.
[0073] like Figure 5 At any time point t Pictures Taking image data as an example, the feature extraction process in the feature learning module is described. First, After downsampling through a convolutional layer with a convolution kernel size of 4×4 and a step size of 4, the feature map is obtained. , which is the input of the first stage. Then, the first stage passes through 3 ConvNeXt blocks, each of which is composed of a 7×7 depth convolution layer and two 1×1 point-by-point convolution layers stacked together, and a convolution layer with a convolution kernel size of 2×2 and a stride of 2 for downsampling, to obtain the output of the first stage, which is the input of the second stage. . The second stage contains 3 ConvNeXt blocks, and a convolution layer with a convolution kernel size of 2×2 and a stride of 2 is set at the end to obtain the output features of the second stage, that is, the input features of the third stage. The third stage contains 9 ConvNeXt blocks, and a convolution layer with a convolution kernel size of 2×2 and a stride of 2 is also set to obtain the output features of the third stage, that is, the input features of the fourth stage. The fourth stage contains 3 ConvNeXt blocks, without downsampling operations, and finally outputs the high-order expression features of the convolution unit.
[0074] Following this process, we can get the final output after four stages Then, global average pooling GAP is used to calculate channel by channel Then, we use layer normalization LN to normalize it and get Finally, through Linear transformation, Transform to , and The dimension is transformed into The above process is defined as the function corresponding to the following formula : ; In addition, in terms of spatial feature learning, the natural characteristics of channel interaction in convolution operations can be used to In terms of temporal feature learning, we can effectively extract features from Extracting features from each channel can equivalently learn the short-term and long-term spatiotemporal patterns of each node.
[0075] It is worth noting that when dividing the data set, all images are strictly arranged in the time order of the original data without random rearrangement. P arrive T - Q All within the time span C nodes, the features learned from the image data are .
[0076] S6, using the spatiotemporal bottleneck attention network to extract the spatiotemporal features of the single-scale short-term spatiotemporal data.
[0077] In this step, for P arrive T - Q All within the time span C nodes, with history P Spatiotemporal data with time steps It is the input of the feature learning module based on single-scale spatiotemporal data.
[0078] A Self-Supervised Spatial-Temporal Bottleneck Attentive Network (SSTBAN) was selected as the baseline method, and the self-supervised branch structure of SSTBAN was removed, that is, only the Spatial-Temporal Bottleneck Attention Network (STBAN) was used to learn features based on spatiotemporal data.
[0079] In this embodiment, STBAN is a deep learning model that combines spatiotemporal feature modeling with an attention mechanism, and is designed to efficiently process spatiotemporal sequence data (such as traffic flow prediction, stock prediction, video analysis, etc.). By introducing a bottleneck attention mechanism, the network effectively captures spatiotemporal dependencies while reducing computational complexity, thereby improving model performance.
[0080] At any time t Spatiotemporal data Taking the feature extraction process in the feature learning module based on single-scale spatiotemporal data as an example, we describe the feature extraction process in the feature learning module based on single-scale spatiotemporal data. First, is input into the STBAN consisting of a spatiotemporal encoder, a transform attention block, and a spatiotemporal prediction decoder. In this embodiment, the spatiotemporal encoder and the spatiotemporal prediction decoder are stacked =3 and = 3 Spatial-Temporal Bottleneck Attention (STBA) blocks with residual connections to learn Specifically, The spatial dependencies between nodes and the temporal correlations in the time series are captured through the spatiotemporal encoder, and the single-scale short-term spatiotemporal features are learned to extract the spatiotemporal feature tensor of the single-scale short-term spatiotemporal data to provide input for the subsequent attention calculation. Subsequently, the dependency between the spatiotemporal feature tensor at different spatial positions and different time steps is captured by transforming the attention block to further enhance the feature expression ability, and the dimension of the attention calculation is reduced through dimensionality reduction operations (such as linear projection), thereby reducing the computational complexity and obtaining the attention feature. The final decoding prediction task is performed in the spatiotemporal prediction decoder to generate the final single-scale short-term spatiotemporal features. .
[0081] The above process is achieved through the spatiotemporal bottleneck attention network function Indicates that the expression is: ; Therefore, for P arrive TQ All within the time span C nodes, the single-scale short-term spatiotemporal features learned from the single-scale short-term spatiotemporal data are ,in, .
[0082] S7, adaptively weighted fusion of the spatiotemporal features extracted from the single-scale short-term spatiotemporal data and the multi-scale long-term image data to obtain a prediction result of the spatiotemporal sequence.
[0083] In order to further improve the prediction performance of MSTSI, the spatiotemporal features learned in the feature learning module based on multi-scale image data are adaptively fused with the spatiotemporal features learned in the feature learning module based on single-scale spatiotemporal data.
[0084] Specifically, initialize two learnable parameters and and through softmax Function normalization generates two weights and .use softmax The process of normalizing the function is: ; in, e Represents a natural constant and is used to calculate exponents.
[0085] Then, by and Respectively and Perform weighted fusion to obtain the final prediction result , whose process is defined as: ; Finally, for P arrive T - Q within the time span C nodes, we can get the predicted results , as in the future Q In this way, more comprehensive and diverse spatiotemporal dependencies can be learned, which embed short-term and long-term patterns and enhance the representation and learning of spatiotemporal data.
[0086] In summary, the method of the present invention can be specifically summarized as follows: Assume that the entire network contains C nodes. First, at any point in time ,from C nodes, spanning TSingle-scale spatiotemporal series data at time steps ( X ) Select History P time steps, and express it as , that is, single-scale short-term spatiotemporal data.
[0087] Next, Input into the feature learning module based on single-scale short-term spatiotemporal data to learn spatiotemporal patterns .
[0088] Then, a multi-scale spatiotemporal data visualization conversion module is proposed by extending time step to introduce more historical data and aggregate them into a multi-scale long-term and short-term spatiotemporal data set .
[0089] Afterwards, Convert to C Fixed-resolution image size , and all C images Package as Through this special processing, the deep learning ability of MSTSI for multi-node spatial features is improved.
[0090] Secondly, a convolutional neural network based on ConvNeXt-T is introduced to design a feature extraction module based on multi-scale image data to learn the patterns of multi-scale long-term spatiotemporal data. .
[0091] Finally, adaptive weighted fusion and , to learn more comprehensive and diverse spatiotemporal dependencies embedded in both short-term and long-term spatiotemporal data.
[0092] Through the learning of this multimodal feature, future Q More accurate predictions at each time step. It is worth mentioning that the learning of multi-scale long-term spatiotemporal patterns relies on fixed-resolution image data, and this process is completed in parallel with the learning of single-scale short-term spatiotemporal patterns. This design not only ensures the accuracy of the prediction performance, but also greatly improves the convergence speed of the method.
[0093] In another preferred embodiment, in order to conduct experimental tests on the method of the present invention, the following data are selected as research objects: (1) City A’s Bus Rapid Transit (BRT) system was completed on March 4, 2019. - Passenger flow data of 44 stations on March 29, 2019, namely the XMBRT dataset; (2) Passenger flow at 80 subway stations in City B from January 1, 2019 to January 26, 2019 The data is used as the research object, namely the HZMetro dataset; (3) The passenger flow data of 276 subway stations in City C from February 29, 2016 to April 1, 2016, namely the BJMetro dataset.
[0094] Then, three evaluation indicators, namely Mean Square Error (MAE), Root Mean Square Error (RMSE) and Mean Absolute Percentage Error (MAPE), were used for evaluation.
[0095] All comparative experiments were completed under the same hardware environment and the same hyperparameter settings. The AdamW optimizer with a learning rate of 0.001 was used for training, the loss function was Huber loss, the batch size was 16, the maximum number of training rounds was 100, and an early termination condition with a threshold of 10 rounds was added (that is, when the loss optimization curve tends to converge and there is no change for 10 consecutive rounds, the training is terminated) to verify its effectiveness.
[0096] The experimental comparison uses 12 baseline models, and most of their hyperparameters are set according to the original text. Only the hyperparameters of some models are fine-tuned to achieve the best prediction performance as much as possible. The comparison models are as follows: (1) Attention-Based Spatial–Temporal Graph Convolutional Network (ASTGCN-r) is a model that combines the attention mechanism and spatiotemporal graph convolution to capture the complex spatiotemporal dependencies in transportation networks.
[0097] (2) Graph WaveNet: Its core idea is to capture long-range temporal dependencies and hidden spatial dependencies through adaptive adjacency matrices and dilated convolutions, which is used for traffic speed prediction and time series modeling.
[0098] (3) Graph Multi-Attention Network (GMAN). Its core idea is to simultaneously model dynamic spatial correlation and nonlinear temporal correlation through a multi-head attention mechanism to capture the spatial dependency between nodes and the nonlinear correlation between different time steps, and adaptively fuse spatial and temporal representations for traffic prediction.
[0099] (4) Spatio-Temporal Synchronous Graph Convolutional Networks (STSGCN), whose core idea is to synchronously capture complex spatiotemporal correlations through local spatiotemporal graph convolution modules for spatiotemporal network data prediction.
[0100] (5) Spectral Temporal Graph Neural Network (StemGNN), whose core idea is to simultaneously model the correlation within and between time series in the spectral domain for multivariate time series prediction.
[0101] (6) Spatio-Temporal Fusion Graph Neural Network (STFGNN). Its core idea is to learn hidden spatio-temporal dependencies through spatio-temporal fusion graph and gated convolution modules for traffic flow prediction.
[0102] (7) Adaptive Spatial Temporal Graph Neural Network (ASTGNN): Its core idea is to capture dynamic spatial dependencies through adaptive graph convolutional layers for optimization of spatiotemporal graph models.
[0103] (8) Spatio-Temporal Embedding Prediction (STEP): The core idea is to encode temporal and spatial information into low-dimensional vectors into the model to enhance the prediction ability.
[0104] (9) Spatial and Temporal IDentity information Network (STID): The core idea is to improve the model’s ability to distinguish different samples through spatial and temporal identity information.
[0105] (10) Propagation Delay-aware Dynamic Long-range Transformer for Traffic Flow Prediction (PDFormer): Its core idea is to capture the spatiotemporal dependencies in traffic flow through propagation delay awareness and dynamic long-range attention.
[0106] (11) Decomposition Dynamic Graph Convolutional Recurrent Network (DDGCRN): The core idea is to process spatiotemporal sequence data by decomposing dynamic graph convolutional and recurrent networks.
[0107] (12) Spatial-Temporal Adaptive Embedding Transformer (STAEformer): The core idea is to enhance the spatial-temporal modeling capability of the Transformer model through spatial-temporal adaptive embedding.
[0108] The experimental results are shown in Table 1. As can be seen from Table 1, the MSTSI model proposed in the present invention achieves the best performance in most indicators, especially on the complex BJMetro dataset. This result shows that the method of the present invention can extract single-scale short-term spatiotemporal patterns and multi-scale long-term spatiotemporal patterns in parallel, and has significant prediction performance.
[0109] Table 1 Experimental results
[0110] In addition, it is found that the convergence speed of the MSTSI model of the present invention is significantly faster than that of other models. In order to present the results more clearly, the loss values of the five models in the validation set are plotted into two sub-graphs according to the number of training rounds (epochs) of the experiment, as shown in Figures 6 (a) and 6 (b), and the number of training rounds corresponding to the lowest loss values obtained by them is highlighted with asterisks. As shown in Figures 6 (a) and 6 (b), the best epochs of the five models, spatiotemporal identity model (STID), spatiotemporal adaptive embedding transformer (STAEformer), propagation delay-aware dynamic long-distance transformer (PDFormer), decomposed dynamic graph convolutional recurrent network (DDGCRN) and the model of the present invention method (MSTSI) are 52, 42, 295, 283 and 33 respectively. Obviously, MSTSI converges the fastest and is more efficient than other models. In summary, compared with the prior art, the present invention has the following beneficial effects: The present invention converts multiple single-scale spatiotemporal data into an image with a resolution limited to a fixed size, so that a single image contains multiple long-term time series data in any combination at different time scales, thereby enhancing the feature expression capability of the input data.
[0111] The present invention adaptively fuses the features of these two modalities, namely, short-term single-scale spatiotemporal features and spatiotemporal features of image data, so that the model can learn more comprehensive and diverse spatiotemporal dependencies, thereby effectively improving the prediction performance of the model.
[0112] The present invention transforms the spatiotemporal prediction problem based on spatiotemporal feature learning into a spatiotemporal prediction problem based on image feature learning. By leveraging the rich and effective feature extraction model of computer vision, the spatiotemporal features of multivariate spatiotemporal data are deeply mined and learned from a new perspective, thereby improving the effect of spatiotemporal feature learning.
[0113] The present invention can freely and parallelly embed the feature learning module based on image data into the existing baseline model of time series prediction, and adaptively fuse the spatiotemporal features extracted from two different modalities of the same source to effectively improve the prediction ability of the model. This method can reduce the problems of difficult alignment and insufficient denoising of multimodal data from different sources, and expand the research ideas of existing spatiotemporal prediction methods.
[0114] Embodiment 2 like Figure 7 As shown, the second embodiment of the present invention further provides a multi-modal spatiotemporal sequence prediction device for multi-scale time series data visualization, comprising: An acquisition unit, used to acquire single-scale spatiotemporal series data; A sliding window division unit, used to divide the single-scale spatiotemporal sequence data into single-scale short-term spatiotemporal data of each sliding window in combination with a sliding window of a preset length; An aggregation unit, used for performing step expansion on the single-scale short-term spatiotemporal data, generating a plurality of single-scale long-term spatiotemporal data of different scales and fusing them to obtain a multi-scale long-term and short-term spatiotemporal data set; An imaging unit, used for converting the multi-scale long-term and short-term spatiotemporal data set into image data with a fixed resolution to obtain multi-scale long-term image data; A multi-scale long-term image data feature extraction unit, used to extract the spatiotemporal features of the multi-scale long-term image data using a convolutional neural network; A single-scale short-term spatiotemporal data feature extraction unit, used to extract the spatiotemporal features of the single-scale short-term spatiotemporal data using a spatiotemporal bottleneck attention network; The adaptive fusion prediction unit is used to adaptively weight the spatiotemporal features extracted from the single-scale short-term spatiotemporal data and the multi-scale long-term image data to obtain the prediction result of the spatiotemporal sequence.
[0115] Embodiment 3 The third embodiment of the present invention also provides a multimodal spatiotemporal sequence prediction device for multi-scale time series data visualization, which includes a memory and a processor, wherein the memory stores a computer program, and the computer program can be executed by the processor to implement the multimodal spatiotemporal sequence prediction method for multi-scale time series data visualization as described above.
[0116] Embodiment 4 The fourth embodiment of the present invention also provides a computer-readable storage medium, on which computer-readable instructions are stored. When the computer-readable instructions are executed by a processor of a device where the computer-readable storage medium is located, a multimodal spatiotemporal series prediction method for visualizing multi-scale time series data as described above is implemented.
[0117] In several embodiments provided in the embodiments of the present invention, it should be understood that the disclosed apparatus and method can also be implemented in other ways. The apparatus and method embodiments described above are merely schematic. For example, the flowcharts in the accompanying drawings show the possible architecture, functions and operations of the apparatus, method and computer program product according to multiple embodiments of the present invention. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of a code, and the module, program segment or a part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.
[0118] In addition, the functional modules in the various embodiments of the present invention may be integrated together to form an independent part, or each module may exist independently, or two or more modules may be integrated to form an independent part.
[0119] If the function is implemented in the form of a software function module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, an electronic device, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc. Various media that can store program codes. It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements that are not explicitly listed, or also includes elements inherent to such a process, method, article or device. Without more constraints, an element defined by the phrase "comprising a..." does not exclude the existence of other identical elements in the process, method, article or apparatus comprising the element.
[0120] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms "a", "said" and "the" used in the embodiments of the present invention and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings.
[0121] It should be understood that the term "and / or" used in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.
[0122] The word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)", depending on the context.
[0123] The "first\second" mentioned in the embodiments is only to distinguish similar objects, and does not represent a specific order for the objects. It is understandable that the "first\second" can be interchanged with the specific order or sequence where permitted. It should be understood that the objects distinguished by "first\second" can be interchanged where appropriate, so that the embodiments described herein can be implemented in an order other than those illustrated or described herein.
[0124] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A multi-modal spatiotemporal series prediction method for multi-scale time series data visualization, characterized in that: include: S1, obtain single-scale spatiotemporal series data; S2, combining a sliding window of a preset length, dividing the single-scale spatiotemporal sequence data to obtain single-scale short-term spatiotemporal data of each sliding window; S3, performing step expansion on the single-scale short-term spatiotemporal data to generate a plurality of single-scale long-term spatiotemporal data of different scales and fusing them to obtain a multi-scale long-term and short-term spatiotemporal data set; S4, converting the multi-scale long-term and short-term spatiotemporal data set into image data with a fixed resolution to obtain multi-scale long-term image data; S5, extracting spatiotemporal features of the multi-scale long-term image data using a convolutional neural network; S6, extracting the spatiotemporal features of the single-scale short-term spatiotemporal data using a spatiotemporal bottleneck attention network; S7, adaptively weighted fusion of the spatiotemporal features extracted from the single-scale short-term spatiotemporal data and the multi-scale long-term image data to obtain a prediction result of the spatiotemporal sequence.
2. According to the multi-scale time series data visualization multi-modal spatiotemporal series prediction method of claim 1, characterized in that ,The single-scale spatiotemporal series data include traffic flow data, stock price data, brain wave data, humidity data or temperature data at a single time scale; The S2 is specifically: Assume that the acquired single-scale spatiotemporal series data is ,in, for X The dimension represents the symbol, C Indicates the total number of nodes. Any node ; T represents the total time step, time , Q Indicates the step size that needs to be predicted; Assume the length of the sliding window is P , each sliding window is represented as , then the spatiotemporal sequence of each sliding window is expressed as: ; in, is the single-scale short-term spatiotemporal data; Representation Node c At a certain point in time t of P spatiotemporal data.
3. According to claim 2, a multi-modal spatiotemporal sequence prediction method for multi-scale time series data visualization is characterized in that , the S3 is specifically: Define the time scale corresponding to the single-scale short-term spatiotemporal data as the first scale, and its time interval is ,Include P spatiotemporal data; By analyzing the data of the first scale Perform forward step expansion to generate multi-scale long-time series spatiotemporal data. The generation method is: Will of P The two adjacent data in the data are added to get a new data set, and its time interval is , the number of new data sets is quantity P half of the second scale, that is, to generate the second half of the spatiotemporal data; right Expand a sliding window forward to obtain the single-scale short-term spatiotemporal data in the corresponding sliding window, that is, 1 history P Step-by-step spatiotemporal data ,Will Add the two adjacent data in the second scale to get the first half of the second scale. spatiotemporal data; Then, of P Continuous s Add the data to generate the scale s The latter P / s New data sets, whose time interval is ;right Scaling forward s -1 sliding window obtains the single-scale short-term spatiotemporal data within the corresponding sliding window, that is Single-scale long-term spatiotemporal data ,Will Add the two adjacent data in the sth scale to get the remaining data, and ; The generated single-scale spatiotemporal data of multiple different scales are spliced and fused to obtain a multi-scale long-term and short-term spatiotemporal data set, which is expressed as: ; in, Represents the first scale P Short-term spatiotemporal data, Indicates the generated s Scale P Long-term spatiotemporal data, Representation Node c The corresponding multi-scale long-term and short-term spatiotemporal data sets, d Indicates the number of scales to be spliced.
4. According to claim 3, a multi-modal spatiotemporal sequence prediction method for multi-scale time series data visualization is characterized in that , the S4 is specifically: Any node c Corresponding multi-scale long-term and short-term spatiotemporal data sets Draw d Scatter plots; Stack the scatter plots vertically into a column to generate an RGB image with three channels ;in, , Respectively represent pictures The height and width of By fine-tuning d Generate a table containing d RGB images corresponding to spatiotemporal data of different time scales, each RGB image contains d Spatiotemporal data of different time scales corresponding to the scatter plots; but, C The nodes will generate C RGB images ; Will C RGB images Convert to a C Grayscale image with 1 channel , to better extract the spatial features of multiple nodes; Grayscale image The dimension is converted to , that is, use a sheet containing C Grayscale image with 1 channel Corresponding representation C A node at a certain point in time t Multi-scale long-term and short-term spatiotemporal data sets are used to reduce computational complexity.
5. According to claim 3, a multi-modal spatiotemporal sequence prediction method for multi-scale time series data visualization is characterized in that ,In the process of expanding the new data set: like t ∈[ P , s × P -1], then the remaining P - P / s The data is filled with zero; like t ∈[ s × P , T - Q ], then the corresponding s × P The spatiotemporal series data generates s Scale P spatiotemporal data, the expression is: ; ; in, Indicates that after expansion and filling, it contains s Scale P Aggregate data.
6. The multi-modal spatiotemporal sequence prediction method for multi-scale time series data visualization according to claim 1 is characterized in that ,The convolutional neural network is a lightweight version of ConvNeXt-T, including convolution units, pooling layers, layer normalization and linear transformation layers; The multi-scale long-term image data The convolution unit performs multiple stages of convolution downsampling operations, and then the output result of the convolution unit is sequentially subjected to the global average pooling operation of the pooling layer, the standardization operation of the layer normalization, and the channel dimension transformation operation of the linear transformation layer to output the learned spatiotemporal features.
7. The multi-modal spatiotemporal sequence prediction method for multi-scale time series data visualization according to claim 6 is characterized in that ,The convolution unit includes a pre-convolution module and four cascaded stage structures, and the four stages respectively include several ConvNeXt blocks; The pre-convolution module consists of a convolution layer with a convolution kernel size of 4×4 and a step size of 4, which is used to Perform preliminary downsampling and output the input features of the first stage; The first stage contains three ConvNeXt blocks, each of which is composed of a 7×7 deep convolution layer and two 1×1 point-by-point convolution layers stacked in sequence to enhance the local receptive field and improve the feature expression ability. At the end of the ConvNeXt block, a convolution layer with a convolution kernel size of 2×2 and a stride of 2 is set for further downsampling to obtain the output features of the first stage, which are the input features of the second stage. The second stage contains 3 ConvNeXt blocks, and a convolution layer with a convolution kernel size of 2×2 and a stride of 2 is set at the end to obtain the output features of the second stage, which are the input features of the third stage; The third stage includes 9 ConvNeXt blocks, and also sets a convolution layer with a convolution kernel size of 2×2 and a stride of 2 to obtain the output features of the third stage, which are the input features of the fourth stage; The fourth stage includes 3 ConvNeXt blocks, without downsampling operation, and finally outputs the high-order expression features of the convolution unit.
8. The multi-modal spatiotemporal sequence prediction method for multi-scale time series data visualization according to claim 1 is characterized in that ,The spatiotemporal bottleneck attention network includes a spatiotemporal encoder, a transform attention block and a spatiotemporal prediction decoder; The spatiotemporal encoder is composed of three sequentially cascaded spatiotemporal bottleneck attention blocks with residual connections, which are used to model the spatial dependency between nodes and the temporal correlation in the sequence dimension, so as to extract single-scale short-term spatiotemporal features and generate a spatiotemporal feature tensor corresponding to the single-scale spatiotemporal sequence data; The transformation attention block models the spatiotemporal feature tensor, suppresses irrelevant information, enhances the expression ability of key features, and outputs global attention features; The spatiotemporal prediction decoder consists of three sequentially cascaded spatiotemporal bottleneck attention blocks with residual connections, which are used to map the attention features into a single-scale future spatiotemporal representation to achieve prediction of the target sequence.
9. The multi-modal spatiotemporal sequence prediction method for multi-scale time series data visualization according to claim 1 is characterized in that , the S7 is specifically: Initialize two learnable parameters and ,pass softmax The function is normalized to generate the corresponding weight and , the expression is: ; ; in, e Represents a natural constant, used to calculate exponents; pass and The single-scale short-term spatiotemporal data are Extracted spatiotemporal features With the multi-scale long-term image data Extracted spatiotemporal features Perform weighted fusion to obtain the final prediction result, which is expressed as: in, The final prediction result; is the spatiotemporal features extracted from the single-scale short-term spatiotemporal data; is the spatiotemporal features extracted from the multi-scale long-term image data, t Indicates a point in time.
10. A multi-modal spatiotemporal sequence prediction device for multi-scale time series data visualization, characterized in that: include: An acquisition unit, used to acquire single-scale spatiotemporal series data; A sliding window division unit, used to divide the single-scale spatiotemporal sequence data into single-scale short-term spatiotemporal data of each sliding window in combination with a sliding window of a preset length; An aggregation unit, used for performing step expansion on the single-scale short-term spatiotemporal data, generating a plurality of single-scale long-term spatiotemporal data of different scales and fusing them to obtain a multi-scale long-term and short-term spatiotemporal data set; An imaging unit, used for converting the multi-scale long-term and short-term spatiotemporal data set into image data with a fixed resolution to obtain multi-scale long-term image data; A multi-scale long-term image data feature extraction unit, used to extract the spatiotemporal features of the multi-scale long-term image data using a convolutional neural network; A single-scale short-term spatiotemporal data feature extraction unit, used to extract the spatiotemporal features of the single-scale short-term spatiotemporal data using a spatiotemporal bottleneck attention network; The adaptive fusion prediction unit is used to adaptively weight the spatiotemporal features extracted from the single-scale short-term spatiotemporal data and the multi-scale long-term image data to obtain the prediction result of the spatiotemporal sequence.
Citation Information
Patent Citations
Short temporary rainfall prediction method and device based on multi-scale space-time consistency
CN116152620A
Multi-scale convolution cycle unit space-time sequence prediction method fusing spatial local correlation
CN117671444A
Time series data classification prediction method, electronic equipment and storage medium
CN118606809A
Multivariate time sequence long-term prediction method based on multi-scale time sequence feature enhancement
CN118657253A
Road transportation capacity distribution prediction method integrating long-term and short-term fluctuation trends
CN119623763A