A multi-element time series analysis method based on fusion of multiple visibility graphs and transformers

By constructing the MVGFormer framework with multiple visibility graphs and visibility graph attention mechanism, the problems of high computational complexity and distracted attention in multivariate time series analysis are solved, and more efficient and accurate analysis is achieved, which is suitable for multivariate time series analysis scenarios.

CN120541496BActive Publication Date: 2025-10-14TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510996250.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-18
Publication Date
2025-10-14
Estimated Expiration
2045-07-18

AI Technical Summary

Technical Problem

Existing time series analysis methods have difficulty in effectively capturing the temporal dependencies in complex patterns when processing multivariate time series, resulting in distraction and information loss. Especially in large-scale data, the computational complexity is high, affecting the accuracy and efficiency of the analysis.

Method used

A sliding window mechanism is used to construct multiple visibility graphs, and the cross-channel consensus connection is extracted through the AND aggregation mechanism. Combined with the visibility graph attention mechanism VG-Attention, the computational complexity is reduced and global dependencies are strengthened. Batch normalization and denormalization operations are used to solve the dimensionality difference problem, forming the MVGFormer framework.

Benefits of technology

It significantly reduces computational complexity, improves the performance of multivariate time series analysis, and enhances the accuracy and efficiency of prediction, classification, filling, and anomaly detection. It can better capture global dependencies and is suitable for multiple scenarios such as medical diagnosis, meteorological data forecasting, and industrial process monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120541496B_ABST
    Figure CN120541496B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on multiple visibility map and the fusion of Transformer multi-element time series analysis method, comprising the following steps: S1, using sliding window mechanism converts each channel of multi-element time series into single layer visibility map, combined into multiple visibility map;S2, cross-channel consensus connection is extracted from multiple visibility map by AND aggregation mechanism, and consensus visibility map is generated;S3, with the adjacency relationship of consensus visibility map as constraint, the node embedding is iteratively updated using VG-Attention;S4, the final embedding is input into downstream module, and tasks such as prediction, classification are executed.The method can effectively capture global dependence, improve the performance of multi-element time series analysis.The method can effectively capture global dependence, improve the performance of multi-element time series analysis, and can be widely applied to physiological signal classification in medical diagnosis, weather data prediction, anomaly detection of industrial process monitoring and other multi-element time series analysis scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to time series analysis and deep learning technology, and in particular to a multivariate time series analysis method based on the fusion of multiple visibility graphs and Transformer. Background Art

[0002] In complex systems, multivariate time series (MTS) analysis provides a richer and more comprehensive perspective than univariate analysis, helping to better understand system behavior. Multivariate time series analysis has been widely applied and extensively studied in numerous practical scenarios, such as physiological signal classification for medical diagnosis, meteorological factor prediction for weather forecasting, and anomaly detection in industrial maintenance monitoring data. In recent years, Transformer-based methods have demonstrated outstanding performance in time series analysis. This success stems from the fact that both linguistic and temporal dependencies fundamentally explore the "relationships" between data points (words or time points). However, there is a key difference between the two: linguistic dependencies primarily focus on the order and grammatical structure of words in a sentence, such as subject-verb agreement or syntactic rules. Temporal dependencies in time series, on the other hand, are quite different, emphasizing time-varying patterns such as periodicity, trends, and seasonality. For example, a time series may exhibit a regular pattern of peaks and troughs (periodicity), or a data point may gradually rise or fall over time (trend). While both types of dependencies involve connections between data points, they differ significantly in their structure and evolution. Therefore, methods like full-attention or its variants that rely on pairwise correlations struggle to directly capture meaningful temporal structure from discrete time points and may even introduce noise and potential attention loss when applied to time series data. This noise hinders accurate understanding of temporal structure, leading to incomplete or inaccurate analysis of temporal patterns, especially when temporal dependencies are deeply embedded in complex patterns. Ultimately, this can produce meaningless attention maps and result in information loss.

[0003] In recent years, various deep learning models have been proposed for temporal modeling, such as those based on MLP, CNN, and Transformer. Many of these methods are designed for specific downstream tasks. For example, in forecasting tasks, models like Rlinear and DLinear use a single-layer fully connected neural network to model the relationship between past and future data points for multi-step forecasting. In classification tasks, methods like InceptionTime, Rocket, and TodyNet treat one-dimensional time series as a 1×N pixel matrix and use complex convolutional architectures to generate rich features for time series classification. Other models, such as the Anomaly Transformer, are specifically designed for anomaly detection.

[0004] Visibility graphs were first proposed in the Proceedings of the National Academy of Sciences (PANS) in 2008 and have since found effective applications in time series analysis, particularly in fields such as physiology, economics, and climate research. The core idea of ​​visibility graphs is to transform time series data into a graph, where each data point is represented as a node and edges are formed according to the geometric visibility criterion: two points are connected if they can "see" each other and no other points in between block the line of sight. This technique has been shown to capture the inherent structure of time series data and has been used in various fields to analyze underlying dynamics. In recent years, researchers have focused on combining visibility graphs with other advanced techniques to overcome their limitations and improve their applicability. For example, AVGNet integrates visibility graphs with graph neural networks (GNNs) to create a more adaptable graph structure for signal classification tasks. Furthermore, MAGNN combines multi-scale graph learning techniques with visibility graphs for multivariate time series prediction, preserving temporal dependencies across different scales. These methods aim to improve the flexibility and scalability of visibility graphs while maintaining their interpretability and simplicity. However, visibility graphs still face challenges, especially their high computational complexity, which limits their application to large-scale time series data.

[0005] It should be noted that the information disclosed in the above background technology section is only used to understand the background of this application, and therefore may include information that does not constitute prior art known to ordinary technicians in this field. Summary of the Invention

[0006] The main purpose of the present invention is to overcome the defects existing in the above-mentioned background technology and provide a multivariate time series analysis method based on the fusion of multiple visibility graphs and Transformer.

[0007] To achieve the above object, the present invention adopts the following technical solutions:

[0008] A multivariate time series analysis method based on the fusion of multiple visibility graphs and Transformer includes the following steps:

[0009] S1. Multiple visibility graph construction: A sliding window mechanism is used to convert the time series of each channel in the multivariate time series into a single-layer visibility graph, and then combine them to form a multiple visibility graph;

[0010] S2. Channel consensus information extraction: Based on the multi-visibility graph, the cross-channel consensus connection relationship is extracted through the AND aggregation mechanism to generate a consensus visibility graph that integrates global dependencies;

[0011] S3, visibility graph Transformer encoding: takes the adjacency relationship of the consensus visibility graph as a structural constraint and uses the visibility graph attention mechanism VG-Attention to iteratively update the node embedding representation;

[0012] S4. Task adaptation output: The final node embedding representation is input to the downstream task module to perform prediction, classification, filling or anomaly detection tasks.

[0013] A computer program product includes a computer program, which, when executed by a processor, implements the multivariate time series analysis method based on the fusion of multiple visibility graphs and Transformer.

[0014] The present invention has the following beneficial effects:

[0015] This paper proposes a multivariate time series analysis method based on the fusion of multiple visibility graphs and Transformer, and designs MVGFormer, which is a Transformer framework guided by visibility graphs and can better capture the structural characteristics and temporal changes of multivariate time series. In this paper, the multivariate time series is efficiently converted into a multiple visibility graph through a sliding window mechanism, significantly reducing the computational complexity; the AND aggregation mechanism is used to extract cross-channel consensus connections, strengthen global dependencies and filter noise; the adjacency relationship based on the consensus graph constrains the VG-Attention calculation range to implement a sparse attention mechanism, while retaining the time series structural characteristics (such as periodicity, trend and mutation points) while effectively suppressing attention distraction and reducing memory consumption; combining batch normalization and denormalization operations to solve the problem of multi-channel dimensional differences, so that the same framework can achieve performance improvements in the four major tasks of prediction, classification, filling and anomaly detection. Among them, cross-channel consensus extraction reduces prediction error and improves classification accuracy, ultimately forming a unified solution that takes into account efficiency, robustness and multi-task generalization. The method of the present invention can effectively capture global dependencies and improve the performance of multivariate time series analysis. It can be widely used in multivariate time series analysis scenarios such as physiological signal classification in medical diagnosis, meteorological data prediction, and anomaly detection in industrial process monitoring.

[0016] Other beneficial effects of the embodiments of the present invention will be further described below. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 Schematic diagram of a new perspective for analyzing time series using visibility graph criteria according to an embodiment of the present invention.

[0018] Figure 2 This is a framework diagram of MVGFormer according to an embodiment of the present invention.

[0019] Figure 3 This is a process diagram of mapping a time series into a graph using visibility graph criteria according to an embodiment of the present invention.

[0020] Figure 4 Schematic diagram of visibility graph criteria for periodic and non-periodic time series according to an embodiment of the present invention.

[0021] Figure 5 This is a sliding window visibility graph according to an embodiment of the present invention.

[0022] Figure 6 Layer aggregation mechanisms ((a) based on AND and (b) based on OR) for embodiments of the present invention.

[0023] Figure 7 Schematic diagram of the visibility graph attention mechanism according to an embodiment of the present invention.

[0024] Figure 8 2 is a performance comparison chart of an embodiment of the present invention and a Transformer-based model.

[0025] Figure 9 Comparison of time and memory consumption among full attention, ProbAttention, AutoCorrelation and VG-Attention.

[0026] Figure 10 Comparison chart of F1 scores for anomaly detection tasks.

[0027] Figure 11 Visualization of the ILI prediction results of different models under the "input 36 - prediction 24" setting.

[0028] Figure 12 Visualization of the filling results of ETTh2 under the 12.5% ​​mask ratio setting for different models.

[0029] Figure 13 This is a visualization of the temporal attention learned by an embodiment of the present invention.

[0030] Figure 14Figure 2 shows the impact of input time range on prediction results ((a) ETTh1 and (b) ILI datasets) according to an embodiment of the present invention.

[0031] Figure 15A This is a visualization diagram of multivariate time series embedding on different data sets in an embodiment of the present invention.

[0032] Figure 15B FIG2 is a diagram showing abnormal and multi-cycle patterns in an atrial fibrillation (AF) dataset according to an embodiment of the present invention.

[0033] Figure 16 This is an embedding visualization diagram using t-SNE on the UEA dataset in an embodiment of the present invention (showing the evolution of the learning process during model training).

[0034] Figure 17 This is an overall flow chart of the multivariate time series analysis method based on the fusion of multiple visibility graphs and Transformer in the present invention. DETAILED DESCRIPTION

[0035] The following is a detailed description of the embodiments of the present invention. It should be emphasized that the following description is only exemplary and is not intended to limit the scope of the present invention and its application.

[0036] See Figure 17 The embodiment of the present invention provides a multivariate time series analysis method based on the fusion of multiple visibility graphs and Transformer, comprising the following steps:

[0037] Step S1, multiple visibility graph construction: using a sliding window mechanism to convert the time series of each channel in the multivariate time series into a single-layer visibility graph, and combine them to form a multiple visibility graph.

[0038] In some embodiments, the sliding window mechanism described in step S1 includes: initializing the window: intercepting the starting segment of the time series to construct an initial visibility graph; sliding update: moving the window along the time axis, removing the earliest time point in the window and its associated edges, and adding new time points; connection relationship update: removing the associated edges of the old time points, and adding valid connections of the new time points based on the visibility criteria; loop iteration: repeating the update steps until the entire time series is covered, and outputting a single-layer visibility graph for each channel.

[0039] Step S2, channel consensus information extraction: Based on the multiple visibility graphs, the cross-channel consensus connection relationship is extracted through the AND aggregation mechanism to generate a consensus visibility graph that integrates global dependencies;

[0040] In some embodiments, the AND aggregation mechanism in step S2 includes: inputting the adjacency matrices of each layer of the multi-visibility graph; performing a cross-layer logical AND operation on each pair of nodes: retaining the edge if and only if all channels are connected; and outputting a binary consensus adjacency matrix to represent the strong dependencies shared across channels.

[0041] In some embodiments, the AND aggregation mechanism is configured to: replace the OR aggregation mechanism as the default cross-channel information fusion method; and enhance the model's ability to capture global dependencies by strengthening coexisting connections between channels.

[0042] Step S3, visibility graph Transformer encoding: The adjacency relationship of the consensus visibility graph is used as a structural constraint, and the visibility graph attention mechanism VG-Attention is used to iteratively update the node embedding representation.

[0043] In some embodiments, the VG-Attention mechanism described in step S3 includes: determining the neighbor set of each node based on the consensus visibility graph; calculating the query-key similarity only for the neighbor nodes; normalizing the attention weights of the neighbor nodes through the softmax function; aggregating the neighbor node value vector based on the attention weights, and updating the current node representation.

[0044] In some embodiments, the VG-Attention mechanism controls complexity by: limiting the scope of calculation to the edge connection set of the consensus graph; allocating attention weights only to node pairs with connections; and utilizing the sparsity of the consensus graph to avoid calculation of all node pairs.

[0045] In some embodiments, the node embedding preprocessing in step S3 includes: performing Z-score normalization on the input node features; projecting the normalized features into the latent space through a learnable parameter matrix; superimposing sine / cosine position encoding and injecting time point sequence information.

[0046] In some embodiments, the iterative learning layer of step S3 is specifically composed of the following sequential operations: VG-Attention layer: updating node representation based on consensus graph neighbor relationship; residual connection and batch normalization: adding VG-Attention output to input residual, and performing batch normalization along the feature dimension; feedforward network: extracting nonlinear features through two layers of linear transformation and ReLU activation function; quadratic residual connection and batch normalization: adding the feedforward network output to the batch normalization result residual of step b, and performing batch normalization on the added result.

[0047] In some embodiments, the final output layer of step S3 includes: projecting the node embedding to the original feature dimension through a learnable parameter matrix; and restoring statistical characteristics based on the mean and variance when the input is normalized.

[0048] Step S4, task adaptation output: The final node embedding representation is input to the downstream task module to perform prediction, classification, filling or anomaly detection tasks.

[0049] This multivariate time series analysis method, based on the fusion of multiple visibility graphs and the Transformer, constructs a multiple visibility graph through a sliding window mechanism, reducing time complexity from quadratic to linear, enabling efficient processing of large-scale data. An AND aggregation mechanism is employed to extract cross-channel consensus connections and generate a consensus visibility graph, integrating global dependencies. Based on the structured constraints of this graph, the VG-Attention mechanism is used to replace fully connected attention, alleviating distraction and reducing computational cost. As a general-purpose model, it demonstrates excellent performance in prediction (reaching SOTA in 40% of cases), classification (SOTA in 6 / 17 datasets), imputation, and anomaly detection (F1 surpasses baseline models), with the learned embeddings aligning with the inherent characteristics of the time series. This method effectively captures global dependencies and improves the performance of multivariate time series analysis. It is widely applicable to multivariate time series analysis scenarios, such as physiological signal classification in medical diagnosis, meteorological data forecasting, and anomaly detection in industrial process monitoring.

[0050] The following further describes specific embodiments of the present invention, its algorithm examples and experimental verification.

[0051] To effectively identify temporal patterns, this paper proposes a method inspired by the visibility graph criterion, which explicitly constructs temporal connections between time points based on sequence features, such as Figure 1 As shown in Figure 2. This transformation inherits several intrinsic structural properties of time series and provides a global perspective to enhance the memory capacity of sequences. Unlike current graph neural network (GNN4TS)-based methods that generate graphs through heuristics or learning from sequences, visibility graphs offer a new perspective on constructing connections between time points, characterized by versatility, theoretical validity, interpretability, simplicity, and effectiveness. Therefore, the goal of this paper is to address the limited understanding of time in traditional Transformers and leverage the advantages of visibility graphs to enhance the Transformer's ability to represent time series.

[0052] This paper aims to enhance the Transformer's understanding of temporal relationships by incorporating visibility graph principles. This approach provides a new perspective that helps the model capture the complex network connections inherent in time series and deepens its understanding of temporal dependencies.

[0053] Specifically, the method of the present invention consists of two main stages: constructing a visibility graph of temporal relationships and encoding them using a visibility graph transformer. In the first stage, to address the quadratic time complexity (O(N²)) of the visibility graph, the present invention proposes a sliding window method that avoids redundant computation of time point relationships and achieves linear time complexity. On this basis, the present invention constructs a consensus visibility graph for multivariate time series, which captures common patterns or interactions across channels and reveals the global dependencies of multivariate time series. In the second stage, the present invention adopts a method based on the visibility graph structure (VG-Attention) to improve the new attention similarity calculation, breaking away from the fully connected learning approach commonly used in the standard Transformer mechanism. This improvement ensures that the present invention's model focuses on the most important relationships between dispersed time points, reducing learning costs and alleviating the problem of scattered attention. Experimental results show that MVGFormer outperforms existing methods in four major time series analysis tasks, achieving state-of-the-art performance. The contributions of this invention are summarized as follows:

[0054] We design MVGFormer, a visibility graph-guided Transformer framework, to better capture the structural properties and temporal variations of multivariate time series.

[0055] This paper proposes a sliding window-based visibility graph (SVG), which reduces the computational complexity from O(N2) to O(N), making it more suitable for large-scale time series data.

[0056] To capture global dependencies in multivariate time series, this paper proposes a consensus visibility graph that integrates temporal dependencies and inter-channel dependencies based on consensus relations from graph theory.

[0057] Based on the consensus graph, this paper proposes a visibility graph-based attention (VG-Attention) mechanism, which focuses on learning key structural relationships in temporal patterns and effectively solves the problems of attention distraction and information loss.

[0058] As a general model, MVGFormer consistently achieves state-of-the-art performance on four mainstream time series analysis tasks, surpassing most current Transformer-based methods.

[0059] MVGFormer framework such as Figure 2 As shown in the figure, it consists of three stages: (1) multi-visibility graph: projecting multivariate time series into a multi-layer network; (2) channel-level consensus information extraction; (3) visibility graph Transformer model for encoding time series representation.

[0060] Visibility graph. For a univariate time series consisting of N real-valued data x(t)t=1N, according to the visibility algorithm, nodes usually represent specific time points, and edges represent connections between these time points that meet the visibility graph criteria:

[0061]

[0062] in and are any two data values, and there is another data point A straight line called a “visible line” connects the points in the sequence data and does not intersect any intermediate data height, such as Figure 3 shown. Figure 3 The process of mapping time series to graphs via visibility graph criteria is demonstrated.

[0063] Advantages of visibility maps. This transformation highlights the inherent ability of visibility maps to capture ordered and chaotic structures in time series. Two core advantages: (1) The visibility criterion quantifies the “receptive field” of each time point. Similar to the receptive field in convolutional neural networks (CNNs), the “visibility line” provides a global perspective, enabling the model to evaluate the impact of each time point. Figure 4 Visibility graph criteria for periodic and non-periodic time series are presented. Figure 4 As shown, panel (a) shows that the influence of the periodic peak is confined to each period, while in panel (b), the influence of the sharp drop extends to all subsequent points. (2) Visibility plots are applicable to various types of time series data, including periodic and non-periodic sequences. Specifically, ordered sequences are represented by regular plots (visualized in different colors for clarity), while random sequences correspond to exponential random plots.

[0064] Multiple visibility graph

[0065] This paper applies visibility graphs to multivariate time series. First, to address the high time complexity, this paper introduces an improved method, the sliding window visibility graph (SVG), which achieves linear complexity. Then, this paper constructs a visibility graph for each variable, thereby forming multiple visibility graphs for the multivariate time series, such as Figure 2 (b)

[0066] Sliding Window Visibility Graph

[0067] Traditional methods for constructing visibility graphs for time series data require a full traversal of the dataset to compute visibility relationships between all pairs of points, resulting in a time complexity of O(N2), where N is the length of the sequence. This approach is computationally expensive and inefficient for large-scale datasets. To overcome these limitations, the present invention introduces a sliding window visibility graph (SVG), which processes time series in fixed-size and overlapping windows, as shown in Figure 5 . Figure 5 A sliding window visibility graph is demonstrated: the time complexity is reduced from O(N2) to O(N). As the window slides over the time series, only newly added and removed time points are updated, significantly reducing the amount of recomputation. This results in an overall time complexity of O(N), a substantial improvement over traditional methods. By combining the sliding window technique with traditional visibility graph algorithms, the SVG ensures efficient real-time updates and linear time complexity, making it highly suitable for large-scale time series analysis and real-time data streaming.

[0068] Multiple visibility graph structure

[0069] Based on the proposed SVG, the present invention extends the visibility approach by introducing a novel structure, called a multiple visibility graph, denoted as M. This graph is constructed from M channel time series At any given time point t, each x(t) is a vector resulting from M different sensors. The structure of M is multi-layered, where each layer η contains a visibility graph corresponding to a specific variable In particular, the multiple visibility graph M is composed of a set of adjacency matrices, collectively denoted as:

[0070]

[0071] where each is an N x N adjacency matrix corresponding to layer α. In this multiple graph, each matrix represents connectivity: indicates the presence of a link between nodes i and j in layer η, while = 0 indicates the absence of a link, which applies to all node pairs i, j = 1, 2,..., N.

[0072] Consensus information extraction

[0073] The multi-visibility graph, with its hierarchical structure and shared nodes, captures inter-node relationships that reflect the correlations between different channels in a multivariate time series. To better understand these relationships from a global perspective, this paper leverages consensus relationships in the context of a multi-graph to explore common patterns or interactions between layers. An intuitive idea is to construct an aggregate graph (also called a consensus visibility graph) by merging adjacency matrices to capture consensus information from all views, thereby revealing the global dependencies of a multivariate time series.

[0074] Figure 6 This section shows the layer aggregation mechanisms (a) based on AND and (b) based on OR. AND-based aggregation requires edges to exist in every layer. OR-based aggregation considers edges that exist in any layer.

[0075] like Figure 6 As shown in (a), the present invention uses an AND-based aggregation mechanism (MVG-AND) to simplify the multi-graph M into a single-layer topology A, focusing on the commonalities between layers. This enables the model to highlight the strongly correlated regions between channels. To verify the effectiveness of this method, the present invention compares it with Figure 6 The paper compares the two approaches with the OR-based aggregation method (MVG-OR) shown in (b), which emphasizes the diversity of connections by retaining edges present in any layer. By comparing these two methods, we evaluate the impact of focusing on commonality (AND-based) versus diversity (OR-based) in capturing key relationships in multivariate time series.

[0076] To simplify the graph, both mechanisms operate by binarizing connections (0 and 1), eliminating weights and multiple edges. For the AND-based mechanism, the consensus visibility graph A is:

[0077]

[0078] For each pair of nodes i, j, we have:

[0079]

[0080] For an OR-based mechanism:

[0081]

[0082] in is a matrix whose elements are all 1 and:

[0083]

[0084] Based on this, we construct a consensus visibility graph A using multiple visibility graphs and consensus relationship extraction, which integrates time dependency and channel-level consensus information. The entire process is shown in the pseudocode of Algorithm 1 in Table 1 below.

[0085] Table 1

[0086]

[0087] Visibility Graph Transformer

[0088] The consensus visibility graph A combines sequence and graph-based features, promising focused attention and significant computational efficiency compared to traditional fully connected sequence methods (which have quadratic complexity). This example describes the learning mechanisms of the proposed visibility graph Transformer architecture: temporal embedding, visibility graph attention mechanism, batch normalization, and feedforward networks.

[0089] Temporal embedding. First, we prepare the input node embeddings to be passed to the visibility graph Transformer layer. For the consensus graph A, each node i has a node embedding , where M represents the number of variables. In order to deal with the inconsistency of units between variables and reduce the distribution differences between each input time series, the input node βi is Z-score normalized along the time dimension and then embedded into the d-dimensional hidden feature by linear projection , as shown below:

[0090]

[0091] in

[0092]

[0093] and , j=1,2,…,M is the dimension index of the multivariate time series, N is the input sequence length, , correction factor Set as , is the parameter of the linear projection layer. In order to capture the temporal dependency in the time series, the present invention uses the sinusoidal embedding method commonly used in natural language processing to embed the pre-computed timestamp position encoding.

[0094]

[0095] Where i represents the sequence position of the time point, j represents the dimension index, d represents the dimension of the time encoding, and k=0,1,2,…,d / 2−1. It is then added to the node feature h^i0 to obtain the initial node embedding

[0096]

[0097] Visibility graph attention mechanism. Figure 7 The visibility graph attention mechanism is demonstrated. Figure 7 As shown, the standard (dense) Transformer attention mechanism works well on language sequences because the relationships between tokens are difficult to determine in advance. In this case, it involves matrix multiplication ( ) to calculate the similarity between node pairs as follows:

[0098]

[0099] in, 、 and is a linear projection of the input embedding. The goal of this invention is to focus attention on the inherent characteristics of time series (such as periodicity, trend, fractal properties, etc.). Therefore, compared to the fully connected full attention mechanism, our method incorporates reliable connections between time points, thereby adding additional graph structure information to the input, as shown below:

[0100]

[0101] in

[0102]

[0103] Where (i, j) represents the edge between node i and node j in the consensus graph A, N(i) represents the set of adjacent nodes of node i, Qi and Kj represent the feature vectors of node i and node j respectively. This method will effectively reduce the amount of calculation and solve the problem of attention distraction caused by redundant relationships in the sequence. Specifically, node The update formula is:

[0104]

[0105] in k=1 to H represents the number of heads (i.e. the number of reduction layers), Indicates splicing, N i is the adjacent node of node i.

[0106] Batch Normalization and Feedforward Networks.

[0107] Attention Output is then fed into a feed-forward network (FFN) surrounded by residual connections and normalization layers. Instead of layer normalization, which is commonly used around the Transformer feed-forward layers, the present invention adopts batch normalization, as features from different channels are not comparable; for example, temperature, humidity, and air pressure variables in a multivariate weather dataset cannot be directly compared. Batch normalization enables the present invention to focus more effectively on the distribution of different samples within the same channel, and the ablation study results in Table 8 also support this decision. Then, the feed-forward network performs two linear transformations and incorporates a nonlinear activation function ReLU to extract deeper-level features, as follows:

[0108]

[0109] where W1, W2, and b1, b2 are weight matrices and bias vectors, respectively. denotes the intermediate representation. Finally, for the last layer of the graph Transformer model , a reverse normalization operation is performed to restore the output to its original statistical properties, as follows:

[0110]

[0111] where W1, W2, and b1, b2 are weight matrices and bias vectors, respectively. is used to project the learned node vector representation back to the original feature dimension M for use in downstream tasks.

[0112] Time complexity analysis. There is a significant difference in the time complexity of the full attention (Full-Attention) and visibility graph attention (VG-Attention) mechanisms. The time complexity of the full attention is O(N 2 d), where N is the number of nodes (or time points) and d is the dimension of the node embedding. This complexity is derived from the matrix multiplication , softmax operation, and final matrix multiplication with V, all of which have a quadratic relationship with N. In contrast, the time complexity of the visibility graph attention is , where E is the number of edges in the graph. This has a linear relationship with the number of edges, and when the graph is sparse (i.e. ), the computational efficiency is higher. Therefore, the visibility graph attention is often more efficient than the fully connected full attention mechanism, especially in sparse graphs.

[0113] Experiments

[0114] To verify the effectiveness of the MVGFormer, the present invention conducted a large number of comparative experiments according to the settings of TimesNet. Table 2 summarizes the experimental benchmarks.

[0115] Table 2 ​

[0116]

[0117] Dataset details. To evaluate the performance of the proposed MVGFormer, the present invention conducted prediction task experiments on 7 real-world datasets, including ETT (ETTh1, ETTh2, ETTm1, ETTm2), Exchange Rate, Weather, and ILI, as shown in Table 2. 17 UEA datasets were used on classification tasks, including gesture, action and audio recognition, medical diagnosis through heartbeat monitoring, and other practical field datasets. In addition, the present invention extended the method of the present invention to univariate time series (such as the M4 dataset) to demonstrate its effectiveness on all types of time series.

[0118] Baseline models. The present invention compared the MVGFormer with mature advanced models on all five tasks, including CNN-based models: TimesNet, TCN; MLP-based models: LightTS and DLinear; RNN-based models: LSSL; Transformer-based models: Informer, Pyraformer, Autoformer, FEDformer, Nonstationary Transformer, and ETSformer. In addition, the present invention also compared with the most advanced model for each specific task, such as Anomaly Transformer for anomaly detection, Rocket and TodyNet for classification. Overall, in order to make a comprehensive comparison, the present invention included more than 20 baseline models. All models were trained / tested on a single Nvidia V100-32G GPU.

[0119] Main results

[0120] As a general task model, the MVGFormer achieved consistent state-of-the-art performance on five mainstream analysis tasks. To demonstrate its advantages, the present invention mainly compared it with existing point-based Transformer models, such as FEDformer, Autoformer, Informer, etc., as shown in Figure 8 . Figure 8 Performance comparison with Transformer-based models is shown.

[0121] In addition, the sparse nature of the VG-Attention matrix means that fewer parameters need to be stored and calculated, thus improving memory efficiency. Therefore, as shown in Figure 9 As shown, VG-Attention shows significant advantages in both running time and memory consumption compared to the corresponding attention variants, such as Full-Attention, ProbAttention, and AutoCorrelation. In addition, it can be observed that the MVG-AND layer aggregation strategy outperforms the MVG-OR strategy in terms of MSE and training time, as it can better capture the consensus relationship between channels. Figure 9 The comparison of Full-Attention, ProbAttention, AutoCorrelation, and Visibility Graph Attention in terms of time and memory consumption is shown.

[0122] Therefore, the subsequent experiment adopts an aggregation strategy based on "and" (AND).

[0123] Long-term prediction

[0124] Settings. To evaluate the performance of the model in the prediction task, the present invention conducts experiments on five real-world benchmark datasets, including power datasets (ETTh1, ETTh2), weather datasets, exchange rate datasets, and disease datasets (ILI). Following the settings of TimesNet, the input sequence length of ILI is set to 36, and the input sequence length of other datasets is set to 96.

[0125] Results. As shown in Table 3, MVGFormer outperforms TimesNet and advanced Transformer-based models in more than 40% (16 / 40) of cases in the prediction task. It is worth noting that most of the sub-optimal predictions are very close to the prediction results of TimesNet. This excellent performance can be attributed to the global perspective provided by the visibility graph criterion, which integrates historical and future information. This criterion enhances the model's ability to extract periodic and trend information, helping to make more accurate predictions for remote future time points. In addition, although TimesNet performs outstandingly mainly on the ETTh2 dataset, MVGFormer performs stably on various datasets, indicating that it may be suitable for a wider range of scenarios and has better generalization ability.

[0126] Table 3 shows the long-term prediction task. All results come from 4 different prediction lengths, namely {24, 36, 48, 60} for ILI and {96, 192, 336, 720} for other datasets. The lower the MSE and MAE, the better the performance, and the best result is shown in bold, and the second best result is shown in underline.

[0127] Table 3

[0128]

[0129] Short-term forecasting

[0130] Setting. The proposed MVGFormer method is mainly designed for multivariate time series data. In fact, it can also be applied to univariate time series data without the layer aggregation step. Therefore, the present invention conducts comparative experiments using the univariate M4 dataset, which contains marketing data collected annually, quarterly, and monthly. For short-term forecasting indicators, the present invention adopts the Symmetric Mean Absolute Percentage Error (SMAPE), Mean Absolute Scaled Error (MASE), and Overall Weighted Average (OWA) as evaluation indicators, where OWA is a special indicator used in the M4 competition.

[0131] Results. The M4 dataset contains 100,000 time series from different sources and different frequencies, resulting in diverse temporal variations, making prediction more challenging. As shown in Table 4, MVGFormer consistently outperforms advanced Transformer-based and MLP-based models at all sampling frequencies. This is because the visibility graph based on time features is not affected by the sampling frequency. High-frequency data produces denser subgraphs, while low-frequency data presents sparser connections.

[0132] Table 4 shows the short-term forecasting task on the M4 dataset. The prediction length is in the range of [6, 48], and the result is the weighted average of multiple datasets at different sampling intervals.

[0133] Table 4

[0134]

[0135] Anomaly detection

[0136] Setting. Anomaly detection in industrial monitoring data is crucial for maintenance, but there are labeling challenges due to the subtlety of anomalies in large-scale datasets. The present invention compares models on five real-world datasets, including MSL, SMD, SWaT, SMAP, and PSM.

[0137] Results. Figure 10 Anomaly detection task is shown. F1 score (in %) as the evaluation indicator for each dataset, the higher the value, the better the performance. As shown in Table 5, the proposed MVGFormer model outperforms the state-of-the-art models on all five datasets. Figure 10As shown, MVGFormer performs well in anomaly detection, outperforming TimesNet and advanced Transformer-based models such as FEDformer, Pyraformer, and Autoformer. By leveraging the visibility criterion, MVGFormer can effectively identify anomalies as highly connected nodes in the visibility graph, thereby focusing attention on these nodes. In contrast, standard Transformer methods tend to spread attention across multiple edges due to pair-wise correlation-based computation, which can lead to weakened or overlooked anomaly detection. The sparse attention mechanism based on graphs in the present invention ensures more focused and accurate anomaly detection, thereby improving precision and reducing the likelihood of false positives or negatives.

[0138] Classification

[0139] Settings. To evaluate the model's ability in advanced representation learning, the present invention selected 17 multivariate time series datasets from UEA for sequence-level classification.

[0140] Results. As shown in Table 5, MVGFormer outperforms existing methods, including the classic method Rocket, the deep learning method InceptionTime, and the GNN-based method TodyNet. In particular, it shows significant effectiveness in handling physiological signals, such as on the AtrialFibrillation, StandWalkJump, and SelfRegulationSCP2 datasets.

[0141] Table 5 shows the classification task. The present invention reports the classification accuracy (%) as the result, with the best result in bold.

[0142] Table 5

[0143]

[0144] Imputation

[0145] Settings. System failures in real-world often lead to partial data missing in continuously collected time series, which poses challenges to downstream analysis. Therefore, in practical scenarios, imputing data becomes crucial. Following the settings used in TimesNet (random mask ratio {12.5%, 25%, 37.5%, 50%}), the present invention conducted experiments on the ETT (4 subsets) and Weather datasets, which often experience data missing.

[0146] Results. Table 6 shows that in 50% of the cases, the MVGFormer outperforms most Transformer-based methods when evaluated using the MSE and MAE metrics. This is due to the visibility graph's extraction of temporal trends and the reliable sparse VG-Attention mechanism's focus on key temporal structural dependencies. These factors give the inventive algorithm strong fitting capabilities at both local and global levels. This fitting capability is also proven to be advantageous in the prediction task.

[0147] Table 6 shows the imputation task. The performance of the models under different missingness levels (12.5%, 25%, 37.5%, 50%) is compared. The best result is shown in bold, and the second-best result is underlined.

[0148] Table 6

[0149]

[0150] Model Analysis

[0151] Efficiency Analysis

[0152] For a comprehensive comparison, the inventive shows the prediction task results as shown in Figure 11 and Figure 12 It can be observed that, thanks to the "receptive field" of the visibility graph, the MVGFormer exhibits stronger local fitting capabilities and trend judgment capabilities. Compared with Transformer-based methods such as Autoformer, Stationary, and Transformer, this enables it to better adapt to temporal pattern extraction. Figure 11 The prediction performance is shown: visualization of the prediction results of different models on ILI (Influenza-like Illness) under the "input 36 - predict 24" setting. The blue line represents the true value, and the orange line represents the predicted value. Figure 12 The imputation performance is shown: (a) input time series (b) MVGFormer (the invention) (c) Dlinear (d) Stationary (e) Transformer. Visualization of the ETTh2 imputation results of different models under the 12.5% mask ratio setting. The blue line represents the true value, and the orange line represents the predicted value.

[0153] Advantages of VG-Attention

[0154] From previous experimental results, we observed that MVGFormer exhibited excellent performance in both prediction and classification tasks. To further analyze the reasons for VG-Attention's effectiveness, we compared it with other attention mechanisms, such as full-attention and probabilistic attention. Figure 13 Shows the visualization of learned temporal attention. Among them: (a) input time series (b) VG-Attention (c) ProbAttention (d) Full-Attention. (b) VG-Attention from our invention; (c) from Informer, showing distraction; (d) from standard full attention, there is information loss problem. Figure 13 As shown, both VG-Attention and full attention focus on the peak points in the time series.

[0155] However, probabilistic attention fails to capture some key temporal information, exhibits distraction problems, and learns some incorrect details (marked as ?). In addition, compared to full attention, which requires pairwise learning between all time points, VG-Attention's learning is more targeted.

[0156] This enables VG-Attention to capture some local information and avoid information loss, such as key turning points (highlighted in the circled area). These results demonstrate the advantage of using visibility graphs to enhance the attention mechanism in multivariate time series analysis.

[0157] Longer historical length

[0158] Theoretically, due to its global perspective, visibility graphs are expected to be more suitable for long time series data than other Transformer-based methods. To verify this, we compared MVGFormer with FEDformer, Autoformer, Informer, and Transformer models to determine whether they can capture more historical information and achieve better forecasting performance as the length of the historical time series increases. Figure 14 The figure shows the impact of the input time range on the prediction results. Among them: (a) ETTh1 (b) disease (Illness). On the ETTh1 (T=96) and ILI (T=24) datasets, the MSE results of the model with different lookback lengths (X-axis) and a fixed prediction length T (Y-axis). Figure 14As shown, we found that only MVGFormer and Crossformer showed a gradual decrease in MSE as the input time series lengthened. Interestingly, the performance of the traditional Transformer model actually degraded with increasing input length. This may be because full attention tends to capture more noise as the number of time series tokens increases, leading to distraction.

[0159] Expressive ability

[0160] Figure 15A and Figure 15B Demonstrating representation capabilities: visualization of multivariate time series embeddings on different datasets (ac), and highlighting abnormal and multi-cycle patterns in an atrial fibrillation (AF) dataset (de).

[0161] Among them: (a) fill time series (b) MVGFormer (proposed by this invention) (c) DLinear (d) Stationary (e) Transformer. Figure 15A As shown in Figure 2, the present invention visualizes the embedding representation of multivariate time series. The embedding learned by MVGFormer is highly consistent with the characteristics of the time series. Specifically, the model is good at predicting the key segments of the sequence ( Figure 15A (a)-(b)), such as peaks and large fluctuations, give more prominent embedding. For stable and periodic time series ( Figure 15A (c)), which indicates a more uniform distribution. In addition, we observe that MVGFormer can effectively capture other complex temporal features such as anomalies and multi-periodicity ( Figure 15B (d)-(e)). This further explains its excellent performance in anomaly detection tasks. In addition, as Figure 16 As shown in the figure, MVGFormer shows strong representation ability in classification tasks and can effectively distinguish multivariate time series of different categories in the vector space. Figure 16 Demonstrating representation power (2): Embedding visualizations using t-SNE on PenDigits, NATOPS, and ArticularyWordRecognition (AWR) from the UEA dataset, showing the evolution of the learning process during model training. (a) PenDigits: Round = 0; (b) PenDigits: Round = 20; (c) PenDigits: Round = 53; (d) NATOPS: Round = 0; (e) NATOPS: Round = 30; (f) NATOPS: Round = 100; (g) AWR: Round = 0; (h) AWR: Round = 10; (i) AWR: Round = 40.

[0162] Ablation study

[0163] Impact of channel-level consensus

[0164] To verify the channel-level consensus information extraction method proposed in the present application, the present application conducts experiments to evaluate the impact of the MVG-AND layer aggregation mechanism (with or without aggregation) on the prediction and classification tasks, and the results are shown in Table 7. Layer aggregation is used for ablation study of channel consensus information extraction.

[0165] Table 7

[0166]

[0167] In the no aggregation (w / o_Aggregation) scenario, the multivariate time series is treated as a single variable sequence, each channel is processed independently (channel independent, CI) and encoded using visibility graph attention. Interestingly, after applying MVG-AND for layer aggregation to extract the consensus relationship between channels, both the experimental performance and the training speed are significantly improved: the mean square error is reduced by an average of 2.82%, the classification accuracy is increased by an average of 9.73%, and the training speed is increased by an average of 67.48%. This shows that strengthening the consensus relationship will result in a more sparse connection, but will not harm the time dependence, but rather will enhance the correlation between channels. In addition, as the number of channels increases, the performance also improves, which further emphasizes the importance and necessity of extracting consensus relationships to obtain the global dependence of multivariate time series.

[0168] Batch normalization and layer normalization

[0169] In the MVGFormer, each token (node) corresponds to the value of a different channel. Unlike traditional Transformers that use layer normalization (LayerNorm), different normalization methods are needed because the features in the multivariate time series data may have different scales (units). Batch normalization (BatchNorm) normalizes each feature (i.e., each channel) at all time steps in a batch, solving the problem of inconsistent feature scales and ensuring scale uniformity during training. In contrast, layer normalization normalizes all features at each time step, which cannot effectively solve the problem of inconsistent scales across time steps. As shown in Table 8, compared with layer normalization, batch normalization improves the mean square error (MSE) by 10.81% and the mean absolute error (MAE) by 4.74%. Therefore, batch normalization is more suitable for MVGFormer.

[0170] Table 8 performs ablation study on batch normalization and layer normalization when the prediction length is equal to 96.

[0171] Table 8

[0172]

[0173] In summary, the present invention proposes a multivariate time series analysis method based on the fusion of multiple visibility graphs and Transformer. The designed MVGFormer surpasses most of the current most advanced Transformer-based models in four major tasks, and even surpasses TimesNet. By integrating the temporal structure into the Transformer architecture and relying on the optimized visibility graph attention mechanism, this marks a major advancement in the field of time series analysis. The method of the present invention provides a new solution for attention guidance in multivariate time series models, and uses complex network theory to solve the inherent attention distraction problem in the standard Transformer architecture. The method of the present invention can effectively capture global dependencies and improve the performance of multivariate time series analysis. It can be widely used in multivariate time series analysis scenarios such as physiological signal classification in medical diagnosis, meteorological data forecasting, and anomaly detection in industrial process monitoring.

[0174] The above is a further detailed description of the present invention in conjunction with specific preferred embodiments, and the specific implementation of the present invention should not be considered to be limited to these descriptions. Those skilled in the art will recognize that, without departing from the scope of the present invention, several equivalent substitutions or obvious variations can be made, and the performance or use of the same should be considered to fall within the scope of protection of the present invention.

Claims

1. A multivariate time series analysis method based on the fusion of multiple visibility graphs and Transformer, characterized by: The following steps are involved: S1. Multiple visibility graph construction: A sliding window mechanism is used to convert the time series of each channel in the multivariate time series into a single-layer visibility graph, and then combine them to form a multiple visibility graph; S2. Channel consensus information extraction: Based on the multi-visibility graph, the cross-channel consensus connection relationship is extracted through the AND aggregation mechanism to generate a consensus visibility graph that integrates global dependencies; The AND aggregation mechanism in step S2 includes: Input the adjacency matrix of each layer of the multi-visibility graph; Perform a cross-layer logical AND operation on each pair of nodes: retain the edge if and only if all channels are connected; Output binary consensus adjacency matrix, representing the strong dependencies shared across channels; The AND aggregation mechanism is configured to: Replace the OR aggregation mechanism as the default cross-channel information fusion method; Improve the model's ability to capture global dependencies by strengthening coexisting connections between channels; S3, visibility graph Transformer encoding: takes the adjacency relationship of the consensus visibility graph as a structural constraint and uses the visibility graph attention mechanism VG-Attention to iteratively update the node embedding representation; The VG-Attention mechanism in step S3 includes: Determine the neighbor set of each node based on the consensus visibility graph; Calculate query-key similarity only for neighbor nodes; Normalize the attention weights of neighbor nodes through the softmax function; Aggregate neighbor node value vectors based on attention weights and update the current node representation; The VG-Attention mechanism controls complexity in the following ways: Limit the computation scope to the set of edge connections in the consensus graph; Attention weights are only assigned to connected node pairs; Exploit the sparsity of the consensus graph to avoid full-node pair computations; S4, task adaptation output: embed the final node representation into the downstream task module to perform prediction, classification, filling, or anomaly detection tasks.

2. The method according to claim 1, wherein The sliding window mechanism in step S1 includes: Initialization window: intercept the starting segment of the time series to construct the initial visibility graph; Sliding update: Move the window along the time axis, remove the earliest time point in the window and its associated edges, and add the new time point; Connection relationship update: remove the associated edges of the old time point and add the valid connections of the new time point based on the visibility criteria; Loop iteration: Repeat the update steps until the entire time series is covered, and output a single-layer visibility map for each channel.

3. The method according to any one of claims 1 to 2, characterized in that The node embedding preprocessing in step S3 includes: Perform Z-score normalization on input node features; Projecting normalized features into latent space via a learnable parameter matrix; Superimpose sine / cosine position codes to inject time point order information.

4. The method according to any one of claims 1 to 2, characterized in that The iterative learning layer of step S3 consists of the following sequential operations: VG-Attention layer: updates node representation based on consensus graph neighbor relationships; Residual connection and batch normalization: add the VG-Attention output to the input residual and perform batch normalization along the feature dimension; Feedforward network: extracts nonlinear features through two layers of linear transformation and ReLU activation function; Quadratic residual connection and batch normalization: Add the feedforward network output to the batch normalization result residual of step b, and perform batch normalization on the added result.

5. The method according to any one of claims 1 to 2, characterized in that The final output layer of step S3 includes: Project node embeddings to the original feature dimension through a learnable parameter matrix; Restores statistical properties based on the mean and variance when the input is normalized.

6. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the multivariate time series analysis method based on the fusion of multiple visibility graphs and Transformer is implemented as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Aero-engine residual life prediction method based on multilevel graph feature fusion

    CN117312770A