Time sequence prediction method and device, equipment and storage medium

By extracting data from different timescales from multivariate time series data and using graph neural network to interactively extract features, the problem of insufficient noise robustness in the prior art is solved, and more efficient multivariate time series prediction is achieved.

CN120124686APending Publication Date: 2025-06-10TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311679806.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-07
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The prior art is not robust to noise in multivariate time series prediction, especially in the interactive modeling of time dimensions and variable dimensions, and is susceptible to unexpected noise.

Method used

By extracting second time series data of different time scales from the first time series data, the graph neural network interacts with these data, extracting time node and variable node features, and predicting based on these features to improve the robustness to noise.

Benefits of technology

This method can effectively extract the characteristics of time nodes and variable nodes in data of different time scales, improve the robustness of multivariate time series prediction to noise and improve prediction performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120124686A_ABST
    Figure CN120124686A_ABST
Patent Text Reader

Abstract

The invention provides a time sequence prediction method and device, equipment and a storage medium, and relates to the field of artificial intelligence. The time series prediction method comprises the following steps: acquiring first time series data across at least two variables in a first time period; extracting at least two second time series data with different time scales from the first time series data; interacting the second time sequence data of the at least two different time scales to obtain at least two time node features; interacting the at least two time node features to obtain at least two variable node features; and according to the at least two variable node features, predicting third time series data across the at least two variables in the second time period. According to the embodiment of the invention, the robustness of MTS prediction performance to noise can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of artificial intelligence technology, and in particular, to a time series prediction method, apparatus, device, and storage medium. Background Art

[0002] Due to the emergence of deep neural networks, multivariate time series (MTS) prediction has made significant progress and has been widely applied in various fields (such as climate, transportation, energy, finance, etc.). Exemplarily, MTS prediction can be performed based on Convolutional Neural Networks (CNN), Recursive Neural Networks (RNN), Transformer, or Graph Neural Networks (GNN), etc. Methods based on Transformer and methods based on GNN have shown great potential due to their strong modeling ability for the interaction between the time dimension and the variable dimension.

[0003] In the related art, models based on Transformer have achieved great success in MTS prediction, which benefits from its attention mechanism that can model the long-term interaction between different time points (across time) of the sequence. GNN also shows good results in MTS prediction, which can extract predefined or adaptive interactions between different variables (across variables). However, the MTS prediction performance based on deep neural networks still needs to be further improved. Summary of the Invention

[0004] The present application provides a time series prediction method, apparatus, device, and storage medium, which is beneficial to improving the robustness of MTS prediction performance to noise.

[0005] In a first aspect, an embodiment of the present application provides a time series prediction method, including:

[0006] Obtain first time series data across at least two variables in a first time period;

[0007] Extract second time series data of at least two different time scales from the first time series data;

[0008] Interact the second time series data of the at least two different time scales to obtain at least two time node features; wherein, each time node feature is obtained according to the previous layer features and the previous layer neighbor time node features related to each time node;

[0009] Interact with the at least two time node features to obtain at least two variable node features; wherein, each of the variable node features is obtained based on the relevant previous layer features of each variable node and the previous layer neighbor variable node features;

[0010] Predict the third time series data across the at least two variables in the second time period based on the at least two variable node features.

[0011] In a second aspect, an embodiment of the present application provides a time series prediction device, including:

[0012] An acquisition unit, configured to acquire the first time series data across at least two variables in the first time period;

[0013] An extraction unit, configured to extract at least two second time series data with different time scales from the first time series data;

[0014] A first graph neural network, configured to interact with the at least two second time series data with different time scales to obtain at least two time node features; wherein, each of the time node features is obtained based on the relevant previous layer features of each time node and the previous layer neighbor time node features;

[0015] A second graph neural network, configured to interact with the at least two time node features to obtain at least two variable node features; wherein, each of the variable node features is obtained based on the relevant previous layer features of each variable node and the previous layer neighbor variable node features;

[0016] A prediction unit, configured to predict the third time series data across the at least two variables in the second time period based on the at least two variable node features.

[0017] In a third aspect, an embodiment of the present application provides an electronic device, including: a processor and a memory, the memory is used to store a computer program, and the processor is used to call and run the computer program stored in the memory to execute the method in the first aspect.

[0018] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, including instructions, which when running on a computer cause the computer to execute the method in the first aspect.

[0019] In a fifth aspect, an embodiment of the present application provides a computer program product, including computer program instructions, which cause the computer to execute the method in the first aspect.

[0020] In a sixth aspect, an embodiment of the present application provides a computer program, which causes the computer to execute the method in the first aspect.

[0021] In the above technical solution, at least two second time series data with different time scales are extracted from the first time series data, and at least one second time series data with different time scales from coarse to fine of the first time series data can be extracted. Furthermore, the at least one second time series is interacted, at least two time node features related in the at least one second time series data are extracted, the at least two time node features are interacted, at least two variable node features related in the at least two time node features are extracted, and further, through the at least two variable node features, the third time series data across the at least two variables in the second time period is predicted.

[0022] In the embodiment of the present application, by extracting at least two second time series data with different time scales from the first time series data, the second time series data with different time scales obtained have significantly different noise intensities. For example, a coarser scale exhibits a lower noise intensity. Further, by extracting at least two time node features related in the at least one second time series data, the dependence between time nodes at different time scales can be obtained, making the cross - time relationship robust to noise. At the same time, by extracting at least two variable node features from the at least two time node features, the invariant correlation between variables in the time process can be obtained, thereby improving the robustness of variables to noise. Therefore, the embodiment of the present application is beneficial to improving the robustness of the MTS prediction performance to noise. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 A schematic diagram of the prediction result of the Transformer - based model under unexpected noise;

[0024] Figure 2 A schematic diagram of the application scenario of the embodiment of the present application;

[0025] Figure 3 A schematic flowchart of a time series prediction method according to an embodiment of the present application;

[0026] Figure 4 A schematic diagram of the MTS prediction model according to an embodiment of the present application;

[0027] Figure 5 A schematic flowchart of another time series prediction method according to an embodiment of the present application;

[0028] Figure 6 A schematic diagram of the process of extracting multi - scale MTS according to an embodiment of the present application;

[0029] Figure 7 A schematic diagram related to the multi - scale MTS according to an embodiment of the present application;

[0030] Figure 8 Schematic flowchart of another time series prediction method according to an embodiment of the present application;

[0031] Figure 9 Schematic flowchart of another time series prediction method according to an embodiment of the present application;

[0032] Figure 10 Schematic diagram of cross-scale interaction according to an embodiment of the present application;

[0033] Figure 11 Schematic flowchart of another time series prediction method according to an embodiment of the present application;

[0034] Figure 12 Schematic diagram of cross-variable interaction according to an embodiment of the present application;

[0035] Figure 13 Schematic diagram of a decoder according to an embodiment of the present application;

[0036] Figure 14A Schematic diagram of noise signals at different levels in a multi-scale time series;

[0037] Figure 14B Schematic diagram of homogeneous and heterogeneous relationships between variables;

[0038] Figure 15 Example of quantitative results of MTS prediction using different methods;

[0039] Figure 16 MSE results of each model at different noise ratios on the ETTm2 dataset;

[0040] Figure 17 Schematic diagram of performance comparison of ablation variants;

[0041] Figure 18 MSE results of models with different lookback window sizes on four datasets;

[0042] Figure 19 MSE (left Y-axis) and MAE (right Y-axis) results of CrossGNN for traffic and weather;

[0043] Figure 20 Theoretical computational complexity of CrossGNN and Transformer-based methods;

[0044] Figure 21 Schematic diagram of time and memory consumption comparison of ETTh2;

[0045] Figure 22Schematic block diagram of a time series prediction device according to an embodiment of the present application;

[0046] Figure 23 Schematic block diagram of an electronic device according to an embodiment of the present application. Detailed implementation manners

[0047] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention.

[0048] It should be understood that in the embodiments of the present application, "B corresponding to A" means that B is associated with A. In one implementation, B can be determined according to A. However, it should also be understood that determining B according to A does not mean determining B only according to A, and B can also be determined according to A and / or other information.

[0049] In the description of the present application, unless otherwise specified, "at least one" means one or more, and "a plurality" means two or more than two. In addition, "and / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the preceding and following associated objects. "At least one (item)" or its similar expression refers to any combination of these items, including any combination of single item (item) or plural items (items). For example, at least one (item) of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.

[0050] It should also be understood that the first, second, etc. descriptions in the embodiments of the present application are only for schematic and differentiating the described objects, without an order, and do not represent a special limitation on the number of devices in the embodiments of the present application, and cannot constitute any limitation to the embodiments of the present application.

[0051] It should also be understood that the specific features, structures, or characteristics related to the embodiments in the specification are included in at least one embodiment of the present application. In addition, these specific features, structures, or characteristics can be combined in one or more embodiments in any suitable manner.

[0052] In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or server that includes a series of steps or units does not necessarily have to be limited to those clearly listed steps or units, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products, or devices.

[0053] The embodiments of the present application are applied to the field of artificial intelligence technology.

[0054] Artificial Intelligence (AI) uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, including theories, methods, technologies, and application systems for perceiving the environment, acquiring knowledge, and using knowledge to achieve optimal results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also involves researching the design principles and implementation methods of various intelligent machines to enable them to have functions of perception, reasoning, and decision-making. Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, pre-trained model technologies, operation / interaction systems, mechatronics, etc. Among them, pre-trained models, also known as large models or foundation models, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0055] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, robots, smart healthcare, smart customer service, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0056] The embodiments of this application may be related to Natural Language Processing (NLP) in artificial intelligence technology. NLP is an important direction in the field of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers using natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language people use in daily life, so it has a close connection with the research of linguistics. Natural language processing technology usually includes technologies such as text processing, semantic understanding, machine translation, robot question answering, and knowledge graphs.

[0057] Embodiments of this application may also be related to Machine Learning (ML) in artificial intelligence technology. ML is an interdisciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.

[0058] Embodiments of this application may also be related to Graph Neural Networks (GNN) in artificial intelligence technology. As a machine learning algorithm, GNN can extract important information from graphs and make useful predictions. As graphs become more common and information-rich, and artificial neural networks become more popular and powerful, GNN has become a powerful tool for many important applications. GNN can be created like any other neural network, using fully connected layers, convolutional layers, pooling layers, etc. The type and number of layers depend on the type and complexity of the graph data and the required output. GNN receives formatted graph data as input and generates a numerical vector representing relevant information about the nodes and their relationships. This vector representation is called a "graph embedding". Graph embeddings are usually used in machine learning to convert complex information into a distinguishable and learnable structure. Data is collected from the graph and aggregated with the values obtained from the previous layer. Finally, the output layer of GNN produces an embedding, which is a vector representation of the node data and its knowledge of other nodes in the graph.

[0059] Embodiments of this application may also be related to MTS prediction. MTS consists of time series of multiple variables. MTS prediction aims to estimate future time series values based on historical time series. Deep learning networks have proven to have excellent performance in MTS prediction. The main focus of research on neural network-based MTS prediction lies in designing the interaction between the time dimension (across time) and the variable dimension (across variables). Transformer-based methods and GNN-based methods have shown great potential due to their strong modeling capabilities for the interaction between the time dimension and the variable dimension.

[0060] Cross-temporal interaction modeling aims to capture the correlation between different time points. In related technologies, the CNN-based model TimesNet converts the time series into a two-dimensional matrix and uses a CNN-based backbone network for feature extraction. The RNN-based model LSTnet uses a long short-term memory network (LSTM) to model temporal dependencies, but it may be limited by the inherent problem of gradient vanishing / exploding in RNN. The Transformer-based model benefits from its self-attention mechanism, which enables it to capture long-term cross-temporal dependencies. AutoFormer combines a decomposition mechanism to divide the input sequence into trend and seasonality, and integrates the autocorrelation module into the Transformer to capture long-term cross-temporal dependencies. FedFormer uses a frequency-enhanced decomposition mechanism while incorporating additional frequency information.

[0061] Cross-variable interactions have been shown to be key to MTS prediction, and many works have adopted GNNs to capture cross-variable relationships. In related technologies, one solution uses GNNs to model cross-variable correlations in traffic predictions, which can effectively capture the correlations between different roads in a predefined topology map. Another solution extends the use of GNNs from spatiotemporal predictions to MTS predictions, and proposes a direct method for calculating adaptive cross-variable graphs. In addition, Transformer-based MTS predictions also recognize the potential of cross-variable interactions to improve prediction performance. However, the performance of MTS predictions based on deep neural networks needs to be further improved.

[0062] In the related art, the MTS prediction performance of simple linear neural network models is significantly better than many state-of-the-art complex models. By exploring the reasons why the existing cross-time and cross-variable interaction modeling methods based on deep learning models fail to improve the MTS prediction performance, the inventors found that the existence of some unexpected noise (caused by human error, sensor distortion, etc.) may be the cause of this situation. For example, Figure 1 As shown in (a), the Transformer-based model in the time dimension relies heavily on the input sequence to generate the attention map, and its prediction may be susceptible to occasional noise. Even some small fluctuations (i.e., noise) can easily lead to significant changes in time dependencies. Therefore, the self-attention mechanism tends to assign high scores to abnormal points in the time series, resulting in spurious cross-time correlations.

[0063] In addition, in the variable dimension, cross-variable correlations show complex and dynamic evolution over time. Although there is a potential causal relationship between variables, it is difficult to extract cross-variable interactions due to the influence of noise interference. In addition, Figure 1 As shown in (b), the inventors found that the unexpected noise detected by the outlier detection algorithm accounts for a large proportion in the time series.

[0064] By comprehensively analyzing the data in the real world, the inventors found that the related technologies cannot handle well the reduction in predictability caused by unexpected noise. Specifically, in the time dimension interaction modeling, although the Transformer-based methods have excellent performance, it is observed that their self-attention mechanisms are vulnerable to unexpected noise; in the variable dimension interaction modeling, the cross-variable relationships are dynamic and may be greatly affected by noise during the learning process.

[0065] In view of this, the embodiments of the present application provide a time series prediction method, device, equipment, and storage medium, which can help further improve the robustness of the MTS prediction performance to noise.

[0066] Specifically, first-time series data across at least two variables in a first time period can be obtained; second-time series data of at least two different time scales can be extracted from the first-time series data; a first graph neural network is used to interact with the second-time series data of at least two different time scales to obtain at least two time node features; wherein each time node feature is obtained according to the previous layer features related to each time node and the previous layer neighbor time node features; a second graph neural network is used to interact with the at least two time node features to obtain at least two variable node features; wherein each variable node feature is obtained according to the previous layer features related to each variable node and the previous layer neighbor variable node features; according to the at least two variable node features, third-time series data across at least two variables in a second time period is predicted.

[0067] Through the above technical solution, second-time series data of at least two different time scales is extracted from the first-time series data, and at least one second-time series data of different time scales from coarse to fine of the first-time series data can be extracted. Furthermore, the at least one second-time series is interacted to extract at least two time node features related in the at least one second-time series data, the at least two time node features are interacted to extract at least two variable node features related in the at least two time node features, and further through the at least two variable node features, third-time series data across the at least two variables in the second time period is predicted.

[0068] In the embodiments of the present application, by extracting second time series data of at least two different time scales from the first time series data, the second time series data of different time scales have significantly different noise intensities. For example, coarser scales exhibit lower noise intensities. Further, by extracting at least two time node features related to at least one second time series data, the dependencies between time nodes at different time scales can be obtained, making the cross-time relationship robust to noise. At the same time, by extracting at least two variable node features from at least two time node features, the invariant correlation between variables during the time process can be obtained, thereby improving the robustness of variables to noise. Therefore, the embodiments of the present application are beneficial to improving the robustness of MTS prediction performance to noise.

[0069] The embodiments of the present application can be used for MTS prediction in various fields such as climate, transportation, energy, and finance. Specifically, based on the time series data of multiple variables in a historical time period, the time series data of these multiple variables in a future time period can be predicted.

[0070] As an example, in weather prediction, multiple variables may include physical quantities such as temperature, humidity, precipitation, and wind speed. The embodiments of the present application can predict the predicted sequence values of physical quantities such as temperature, humidity, precipitation, and wind speed in a future period (such as consecutive days or hours) based on the sequence values of physical quantities such as temperature, humidity, precipitation, and wind speed on each day (or each hour) in historical time.

[0071] As another example, in traffic prediction, multiple variables may include physical quantities such as traffic flow, road capacity, vehicle speed, and traffic accidents. The embodiments of the present application can predict the predicted sequence values of physical quantities such as traffic flow, road capacity, vehicle speed, and traffic accidents in a future period (such as consecutive days or hours) based on the sequence values of physical quantities such as traffic flow, road capacity, vehicle speed, and traffic accidents on each day (or each hour) in historical time.

[0072] As another example, in energy prediction, multiple variables may include physical quantities such as power load, power demand, power price, fuel price, and renewable energy power generation. The embodiments of the present application can predict the predicted sequence values of physical quantities such as power load, power demand, power price, fuel price, and renewable energy power generation in a future period (such as consecutive days or hours) based on the sequence values of physical quantities such as power load, power demand, power price, fuel price, and renewable energy power generation on each day (or each hour) in historical time.

[0073] Figure 2 A schematic diagram showing an application scenario of the embodiments of the present application is shown.

[0074] As Figure 2As shown, this application scenario involves a terminal 102 and a server 104. The terminal 102 can communicate with the server 104 through a communication network. The server 104 can be the background server of the terminal 102.

[0075] Exemplarily, the terminal 102 can refer to a type of device with rich human-computer interaction methods, the ability to access the Internet, usually equipped with various operating systems, and having strong processing capabilities. The terminal device can be a smart phone, a tablet computer, a portable notebook computer, a desktop computer, a wearable device, a vehicle-mounted device, and other terminal devices, but not limited thereto.

[0076] The server 104 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The server can also become a node of the blockchain.

[0077] The server can be one or more. When there are multiple servers, there are at least two servers for providing different services, and / or there are at least two servers for providing the same service, such as providing the same service in a load balancing manner. The embodiments of the present application do not limit this.

[0078] The terminal device and the server can be directly or indirectly connected through wired or wireless communication. The present application does not limit this. The present application does not limit the number of servers or terminal devices. The solution provided by the present application can be completed independently by the terminal device, or independently by the server, or jointly completed by the terminal device and the server. The present application does not limit this.

[0079] Optionally, as Figure 2 shown, this application scenario can further include a data storage system 106. The data storage system 106 can store the data required by the server 104. The data storage system can be integrated on the server 104, or deployed on the cloud or other servers, without limitation.

[0080] It should be understood that Figure 2 is only an exemplary illustration and does not specifically limit the application scenario of the embodiments of the present application. For example, Figure 2 exemplarily shows one terminal device and one server. In fact, it can include other numbers of terminal devices and servers. The present application does not limit this.

[0081] The technical solutions of the embodiments of the present application will be described in detail below through some embodiments. These several embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0082] Figure 3 FIG. 4 is a schematic flowchart of a time series prediction method 300 according to an embodiment of the present application. The method 300 can be executed by any electronic device with data processing capabilities. For example, the electronic device can be implemented as a server or a terminal device, and the present application does not limit this. As Figure 3 shown, the time series prediction method 300 includes steps 310 to 350.

[0083] 310. Obtain first time series data of at least two variables in a first time period.

[0084] Among them, the at least two variables may include actual physical quantities in practical applications, and the first time series data may include time series values of at least two actual physical quantities in practical applications in the first time period. Exemplarily, the first time may include historical time, such as each day in the past month, or each hour in the previous day, etc., and the present application does not limit this.

[0085] For example, in the weather field, the at least two variables may be at least two of physical quantities such as temperature, humidity, precipitation, wind power, etc., and the first time series data may be time series values of at least two variables across physical quantities such as temperature, humidity, precipitation, wind power, etc. in a historical time period.

[0086] Again, for example, in the traffic field, the at least two variables may be at least two of physical quantities such as traffic flow, road capacity, vehicle speed, traffic accidents, etc., and the first time series data may be time series values of at least two variables across physical quantities such as traffic flow, road capacity, vehicle speed, traffic accidents, etc. in a historical time period.

[0087] Again, for example, in the energy field, the at least two variables may be at least two of physical quantities such as power load, power demand, power price, fuel price, renewable energy power generation, etc., and the first time series data may be time series values of at least two variables across physical quantities such as power load, power demand, power price, fuel price, renewable energy power generation, etc. in a historical time period.

[0088] In some embodiments, the first time series data contains noise, such as some accidental noises caused by humans, sensor distortion, etc.

[0089] In some embodiments, the first time series data includes MTS, which is the input for MTS prediction. Exemplarily, the first time series data may include a historical sequence of D variables, which is It is represented that where L represents the size of the retrospective window, represents the time series of the i-th variable at the t-th time step. The goal of MTS prediction is to predict the future time series based on χ by represented. Where T represents the predicted time step. Optionally, T >> 1.

[0090] After obtaining the first time series data in step 310, the first time series data can be input into the MTS prediction model. Figure 4 shows an example of the MTS prediction model 400, including an Adaptive Multi-Scale Identifier (AMSI) 410, a Cross-scale GNN 420, a Cross-variable GNN 430, and a forecasting layer 440. The MTS prediction model can also be called a crossGNN, which is a linear complex GNN model used to refine the cross-scale and cross-variable interactions of MTS prediction. Specifically, the AMSI 410 can be used to construct a multi-scale time series with noise reduction to handle unexpected noise in the time dimension. The cross-scale GNN 420 is used to extract scales with clearer trends and weaker noise. The cross-variable GNN 430 is used to extract invariant correlations between variables. The prediction layer 440 can predict the future time series according to the output of the GNN 420. Hereinafter, by combining Figure 4 the model architecture of, to describe the MTS prediction process of steps 320 to 350.

[0091] 320, extract second time series data of at least two different time scales from the first time series data.

[0092] Exemplarily, by using different time periods as time windows for data collection, time series data of different time scales can be obtained. For example, by collecting data with k different time periods, k data sequence data of different time scales can be obtained. Where k is a positive integer. As an example, the time scale can be minutes, hours, days, etc., without limitation.

[0093] Exemplarily, it can be achieved by adopting Figure 4 the AMSI 410 in to extract at least two second time series of different time scales from the first time series data, and construct a multi-scale time series data with noise reduction. Among them, the noise intensities of the second time series data of different time scales are different. Specifically, the AMSI 410 can capture different time scales of the MTS from coarse to fine and reduce unexpected noise at the coarse scale.

[0094] In some embodiments, refer toFigure 5 At least two second time series data with different time scales can be obtained through the following steps 311 to 312.

[0095] 311. Perform a Fourier transform on the first time series data to obtain time series data of the first time series data at at least two frequencies.

[0096] Exemplarily, the first time series data can be transformed into the frequency domain through a Fourier transform, such as a Fast Fourier Transform (FFT). By analyzing the first time series data in the frequency domain, the potential period of the first time series data can be obtained. Exemplarily, the time series data of the first time series data at at least two frequencies after FFT can be expressed as FFT(χ).

[0097] 312. Obtain the time series data of S frequencies with the largest amplitude values among the time series data of at least two frequencies. S is a positive integer greater than 1.

[0098] First, the amplitude of each time series data of at least two frequencies at each frequency can be calculated. Exemplarily, for the time series data of each frequency, the amplitudes of D variables can be averaged in the variable dimension to obtain the amplitude of the time series data of each frequency. For example, the amplitude can be shown as the following formula (1):

[0099]

[0100] where Amp(·) is the calculation of the amplitude, FFT(·) is the calculation of the FFT, and A ∈ R L represents the calculated amplitude of each frequency, which takes the average value from D variables through Avg(·).

[0101] Then, the time series data of S (Top-S) frequencies with the largest amplitudes can be selected from the time series data of at least two frequencies. Exemplarily, the frequencies corresponding to the Top-S amplitude values are shown as the following formula (2):

[0102] {f 1 ,..., f s} = argTop-S(A) (2)

[0103] where argTop-S(A) is to select S frequency values with the highest amplitudes from A.

[0104] 313. According to the frequencies of the time series data of S frequencies and the time window size of the first time series data, perform a downsampling operation on the first time series data to obtain the second time series data with different time scales corresponding to the S frequencies respectively.

[0105] Exemplarily, the downsampling operation may be an average pooling (AvgPool) operation. The second time series data corresponding to S frequencies respectively at different time scales, that is, the above-mentioned second time series data at at least two different time scales.

[0106] As an implementable manner, the period length of each of the S frequencies may be determined according to the frequencies of the time series data of the S frequencies and the time window size of the first time series data, and then the first time series data may be downsampled with the period length of each frequency as the kernel size and the stride to obtain the second time series data at the time scale corresponding to each frequency.

[0107] Exemplarily, according to the following formula (3), the period lengths {p 1 ,..., f s} may be determined by using the selected S frequencies {f 1 , p 2 ,..., p s}:

[0108]

[0109] Then, the average pooling operation is performed on the first time series data X with the period size p s of each frequency as the kernel size and the stride to obtain the time series data at the s-th scale. Exemplarily, it may be as shown in the following formula (4):

[0110]

[0111] where the length of the time series in the s-th scale is where is the operation of rounding down.

[0112] In some embodiments, the second time series data at at least two time scales obtained may also be concatenated in the time dimension to obtain a periodic multi-scale MTS, denoted as x′∈R L′×D , where is the sum of the lengths across all scales. Exemplarily, x′ may be expressed as the following formula (5):

[0113] χ′ = Concat(χ 1 , χ 2 ,..., χ S ) (5)

[0114] Figure 6 FIG. shows a schematic diagram of the process of AMSI extracting multi-scale MTS. As Figure 6As shown, for the input MTS χ ∈ R L×D , AMSI can obtain the period through FFT (Period Capture by FFT). Specifically, it can obtain the average amplitude of the time series data of multiple variables at each frequency, and then obtain S highest frequencies from the time series data of each frequency. Figure 7 Figure (a) in [reference] is a schematic diagram of S highest frequencies. Then, AMSI can obtain the scale through average pooling (Scale Capture by AvgPool). Specifically, it can mark different periods in the time domain according to S highest frequencies, and then obtain time series at different scales (time series at different Scales). Figure 7 Figure (b) in [reference] is a schematic diagram of the periods marked in the time domain, where the lower time series is the long period and the upper time series is the short period. Figure 7 Figure (c) in [reference] is a schematic diagram of time series at different scales, where the lower time series is the fine sequence and the upper time series is the coarse sequence. Then, AMSI concatenates the time series of all scales in the time dimension (concatenate all Scales in time dimension) to obtain the output MTS χ' ∈ R L′×D .

[0115] 330, interact with the second time series data of at least two different time scales to obtain at least two time node features; wherein, each time node feature is obtained according to the previous layer features related to each time node and the previous layer neighbor time node features.

[0116] Exemplarily, a first graph neural network can be used to interact with the second time series data of at least two different time scales to obtain at least two time node features. Exemplarily, each time node feature can be a time node feature of at least two variables (such as multi-variables). The first graph neural network can also be called a cross-scale GNN. The interaction process of the first graph neural network is the information propagation process. During the information propagation process, the previous layer of the current layer is the upper layer of the current layer. For example, during the information propagation process, the previous layer of the Nth layer is the (N - 1)th layer. N is a positive integer greater than 1.

[0117] Continue to refer to Figure 4 , the cross-scale GNN 420 can cross-scale interact with the second time series data of at least two different time scales (multi-scale MTS χ') to extract at least two time nodes that include scales with clearer associations and weaker noises.

[0118] In some embodiments, referring to Figure 8 , at least two time node features can be obtained according to the following steps 331 to 334.

[0119] 331. Construct a first graph neural network based on at least two second time series data; wherein, the first graph neural network includes time nodes of at least two different time scales and first correlation weights between each time node.

[0120] Specifically, the first graph neural network can be a cross-scale graph in the time dimension. Exemplarily, the first graph neural network can be represented as: G scale =(V scale , E scale ). Wherein, is the set of time nodes in all time scales, is the i-th time node. E scale ∈ R L′×L′ is the correlation weight assigned to between each time node, that is, an example of the first correlation weight. Each element in E scale refers to the correlation weight between two time nodes (inter-scale or intra-scale). The main purpose of the cross-scale GNN is to learn cross-scale time correlation weights E scale that are insensitive to noise interference.

[0121] Optionally, in order to weaken the influence of noise on the correlation weight, two learnable vectors and can be generated to initialize E scale , so as to maintain the independence of E scale .

[0122]

[0123] Wherein, ReLU(·) is the activation function of the regularization weight matrix, making each element positive; Softmax(·) is the operation to ensure that the sum of the weights of all nodes related to a specific time node is 1.

[0124] 332. Obtain at least one neighbor time node for each time node.

[0125] Wherein, at least one neighbor time node for each time node is a time node related to each time node. The neighbor time node can include at least one of time nodes of the same scale and cross-scale time nodes, and this application does not limit this. In some embodiments, as Figure 9 shown, at least one neighbor time node can be obtained according to the following steps 3321 to 3323.

[0126] 3321. Obtain at least one first neighbor time node at each scale of each time node according to the first correlation weight between each time node.

[0127] Specifically, step 3321 obtains at least one first neighbor time node through scale-sensitive restriction. Among them, the number of first neighbor time nodes at each scale is negatively correlated with the cycle length of each scale. Specifically, for each time node, the number of related time nodes at its fine scale should be greater than that at the coarse scale. For any time node where the number of related nodes at the s-th scale is restricted to where p s is the cycle length of the s-th scale, represents the ceiling function, and K is a constant. This ensures that finer time series contribute more time node associations.

[0128] Optionally, at least one node with the highest correlation weight at each scale of each time node can be obtained according to the first correlation weight between each time node, as at least one first neighbor time node.

[0129] Exemplarily, the neighbor time nodes set at the s-th scale of v i (i.e., the time nodes related at the s-th scale) are represented as as shown in the following formula (7):

[0130]

[0131] where argTop-k s (·) is the operation of extracting the s nodes with the highest correlation weight k , is the time node at the s-th scale 's correlation weight. In this way, the number of neighbor nodes at different scales can be restricted based on the first correlation weight matrix E scale .

[0132] 3322. Obtain at least one second neighbor time node that shares the same scale for each time node.

[0133] Specifically, step 3322 obtains at least one second neighbor time node through trend-aware selection. Specifically, to ensure that time trends can be captured, at least one second neighbor time node that shares the same scale for each time node can be obtained to retain the association between the time node and its previous and subsequent nodes. Exemplarily, Represented as time nodes The trend neighbor set of, which can be defined by the following formula (8):

[0134]

[0135] Where scale(·) provides the scale of the time nodes The trend neighbor set of consists of neighboring time nodes that share the same scale (i.e., |i - j| ≤ 1). This gives the ability to preserve the time trend in the cross-scale correlation graph

[0136] 3323. Determine at least one neighbor time node according to at least one first neighbor time node and at least one second neighbor time node

[0137] Specifically, at least one neighbor time node can be determined according to at least one of at least one first neighbor time node and at least one second neighbor time node. Exemplarily, the union of the first neighbor time node and the second neighbor time node can be used as at least one neighbor time node. Exemplarily, can be used As the time node The selected neighbor set of (i.e., at least one neighbor time node)

[0138] 333. Normalize the first correlation weight between each time node and at least one neighbor time node to obtain the first relative correlation weight between each time node and the at least one neighbor time node

[0139] Exemplarily, for the selected neighbor set Its correlation weight can be renormalized according to the following formula (9):

[0140]

[0141] Where E scale [i,j] is And The correlation weight between. By normalization, insignificant correlations can be filtered out and a restricted set of neighbor nodes can be retained for each node In addition, applying renormalization to the retained correlations can construct a cross-scale correlation graph between time nodes at different scales

[0142] 334. Obtain the time node feature of each time node according to the first relative correlation weight, the previous layer neighbor time node features associated with each time node, and the previous layer features associated with each time node

[0143] Specifically, after obtaining the cross-scale time correlation graph, the time node features of each time node can be obtained based on the cross-scale interaction in the time dimension of the GNN. Specifically, the cross-scale interaction in the time dimension of the GNN, that is, the process of using the GNN to interact with the second time series data of different time scales, can specifically refer to the information propagation process based on the GNN. Specifically, each information propagation process of the GNN is as follows: according to the first relative correlation weight, aggregate the neighbor time node features of the previous layer related to each time node to obtain the aggregated neighbor time node features of the current layer of each time node; according to the aggregated neighbor time node features of the current layer of each time node and the features of the previous layer related to each time node, obtain the time node features of each time node. Among them, the neighbor time node features of the previous layer refer to the neighbor time node features of the previous layer.

[0144] The information propagation process can be stacked for N layers. Among them, N is a positive integer. Among them, the previous layer of the Nth layer is the (N-1)th layer. Exemplarily, the information propagation process in the cross-scale GNN is shown in the following formula (10):

[0145]

[0146]

[0147]

[0148] Among them, σ(·) is the activation function, W is the learnable matrix, H time is the time node feature; NH time is the aggregation of neighbor time node features, that is, the aggregated neighbor time node features.

[0149] In formula (10), first obtain the aggregated neighbor time node features of the current layer (Nth layer) of each time node That is Aggregate the neighbor node features of the previous layer (N-1) layer related to the time nodes from Then, according to the aggregated time node features And the features of the previous layer Obtain the time node features of each time node That is From the aggregated time node features And the features of the previous layer Update. Finally, the normalized of the Nth layer is the output of the cross-scale GNN.

[0150] Figure 10 Shows a schematic diagram of cross-scale interaction. As Figure 10 Shown, it can be based on two learnable variables Initialize the GNN to obtain the initial graph (initializedgraph). For the input χ′ containing different variables (differentvariable), determine the neighbor time nodes of each time node of each variable in the cross-scale graph (Gross-scalegraph) through scale-sensitive constraints and trend-aware selection. The cross-scale graph can obtain the output H of the Nth layer through information propagation. time Is the output of the cross-scale GNN.

[0151] Therefore, in the time dimension, the embodiment of the present application models the dependencies between different scales through the cross-scale GNN, where the scale with clearer trends and weaker noise will be assigned more edge weights, so as to extract the time node features of the time scale of the scale with clearer trends and weaker noise.

[0152] 340, interact at least two time node features to obtain at least two variable node features; wherein, each variable node feature is obtained according to the relevant previous layer features of each variable node and the previous layer neighbor variable node features.

[0153] Exemplarily, a second graph neural network can be used to interact at least two time node features to obtain at least two variable node features. Exemplarily, the at least two variable node features are the node features corresponding to at least two variables respectively, that is, one variable can correspond to one variable node feature. The second graph neural network can also be called the cross-scale GNN. The interaction process of the second graph neural network is the information propagation process. In the information propagation process, the previous layer of the current layer is the upper layer of the current layer. For example, the previous layer of the Nth layer in the information propagation process is the (N - 1)th layer.

[0154] Exemplarily, the second graph neural network can also be called the cross-variable GNN. Continue to refer to Figure 4 The cross-variable GNN430 can extract the variable-variable correlation by utilizing the interaction of at least two time node features of each variable, such as the invariant correlation composed of homogeneous and heterogeneous relationships.

[0155] In some embodiments, refer to Figure 11 At least two variable node features can be obtained through the following steps 341 to 344.

[0156] 341, construct a second graph neural network according to at least two time node features; wherein, the second graph neural network includes at least two variable nodes and the second correlation weights between each variable node.

[0157] Specifically, the second graph neural network can be a cross-variable graph in the variable dimension. Exemplarily, the second graph neural network can be represented as G var =(Vvar , E var ). Among them, is the set of variable nodes in all variable scales, is the i-th variable node. E var is the correlation weight assigned between each variable node, that is, an example of the second correlation weight. E var Each element in refers to the correlation weight between two variable nodes. The main purpose of the cross-variable GNN is to learn the correlation weight E var .

[0158] Optionally, it can be initialized by generating two learnable vectors and to maintain the independence of E var , to maintain E var 's independence.

[0159]

[0160] Among them, ReLU(·) is the activation function of the regularization weight matrix, making each element positive; Softmax(·) is an operation to ensure that the sum of the weights of all nodes related to a specific variable node is 1.

[0161] 342, obtain at least one neighbor variable node of each variable node.

[0162] Among them, at least one neighbor variable node of each variable node is a variable node related to each variable node. In some embodiments, at least one neighbor variable node includes a homogeneous neighbor variable node and a heterogeneous neighbor variable node.

[0163] Optionally, at least one homogeneous neighbor variable node and at least one heterogeneous neighbor variable node of each variable node can be obtained according to the second correlation weight between each variable node.

[0164] Specifically, at least one (such as number) of variable nodes with the highest correlation weight can be selected as the positive neighbors of the homogeneous connection of each variable node, that is, the homogeneous neighbor variable nodes, and at least one (such as number) of variable nodes with the lowest correlation weight can be selected as the negative neighbors of the heterogeneous connection of each variable node, that is, the heterogeneous neighbor variable nodes. Exemplarily, is represented as the correlation weight of D variable nodes related to , then for variable v i , its two decoupled neighbor sets can be respectively represented as and

[0165] 343. Normalize the second correlation weight between each variable node and at least one neighboring variable node to obtain the second relative correlation weight between each variable node and at least one neighboring variable node.

[0166] Exemplarily, when at least one neighboring variable node includes at least one homogeneous neighboring variable node and at least one heterogeneous neighboring variable node, the second correlation weight between each variable node and at least one homogeneous neighboring variable node can be normalized to obtain the relative correlation weight between each variable node and the at least one homogeneous neighboring variable node, and the second correlation weight between each variable node and at least one heterogeneous neighboring variable node can be normalized to obtain the relative correlation weight between each variable node and at least one heterogeneous neighboring variable node. Then, the second relative correlation weight can be obtained according to the relative correlation weight between each variable node and at least one homogeneous neighboring variable node and the relative correlation weight between each variable node and at least one heterogeneous neighboring variable node.

[0167] Exemplarily, the relative correlation weight between each variable node and at least one homogeneous neighboring variable node and the relative correlation weight between each variable node and at least one heterogeneous neighboring variable node can be obtained according to the following formula (12):

[0168]

[0169] Through the above normalization process, other edges of the variable node except for the homogeneous and heterogeneous nodes can be filtered out. Among them, the weight of the edge between the variable node and the homogeneous neighboring variable node is positively correlated with its correlation score, while the weight of the edge between the variable node and the heterogeneous neighboring variable node is negatively correlated with its correlation score. In addition, in the embodiments of the present application, the weights of the edges between the variable node and the homogeneous neighboring variable node and the weights of the edges between the variable node and the heterogeneous neighboring variable node are separately normalized. Then, a cross-variable graph with disentangled homogeneous and heterogeneous correlations is constructed according to the normalized weights.

[0170] 344. Obtain the variable node feature of each variable node according to the second relative correlation weight, the feature of the previous layer neighboring variable node related to each variable node, and the feature of the previous layer related to each variable node.

[0171] Specifically, after obtaining the cross-variable graph, the variable node features of each variable node can be obtained based on the cross-variable interaction of the variable dimensions of the GNN. Specifically, based on the cross-scale interaction of the variable dimensions of the GNN, that is, the process of using the GNN to interact the variable node features of different time scales, it can specifically refer to the information propagation process based on the GNN. Specifically, each information propagation process of the GNN is as follows: according to the second relative correlation weight, aggregate the previous layer neighbor variable node features related to each variable node to obtain the neighbor variable node aggregation feature of each variable node at the current layer. Then, according to the neighbor variable node aggregation feature of each variable node at the current layer and the previous layer features related to each variable node, obtain the variable node features of each variable node. Among them, the previous layer neighbor variable node features refer to the neighbor variable node features of the previous layer.

[0172] Exemplarily, the information propagation process can be stacked for N layers. Where N is a positive integer. Among them, the previous layer of the Nth layer is the (N - 1)th layer. Exemplarily, the information propagation process in the cross-variable GNN is shown in the following formula (13):

[0173]

[0174]

[0175]

[0176] Among them, σ(·) is the activation function, W is the learnable matrix, H var is the variable node feature; NH var is the aggregation of the neighbor variable node features, that is, the neighbor variable node aggregation feature.

[0177] Figure 12 shows a schematic diagram of the cross-variable interaction. As Figure 12 shown, the GNN can be initialized according to two learnable variables to obtain the initialized graph. For the input H time containing different variables, determine the neighbor variable nodes of each variable in the cross-variable graph (Gross-variable graph) through the disentanglement of homogeneous and heterogeneous relationships. The cross-variable graph can obtain the output H vaa at the Nth layer through information propagation, which is the output of the cross-variable GNN.

[0178] Therefore, in the variable dimension, the embodiments of the present application introduce the heterogeneous interaction modeling between variables into the MTS prediction through the cross-variable GNN, so as to utilize the homogeneity and heterogeneity between different variables with positive and negative edge weights to learn the invariant associations including the homogeneous and heterogeneous relationships between variables.

[0179] 350. Predict third time series data across at least two variables in a second time period based on at least two variable node features.

[0180] Exemplarily, referring back to Figure 4 , the prediction layer 440 can predict time series data across at least two variables in a second time period based on at least two variable node features. Exemplarily, the second time can include future times, such as each day in the next month, or each hour on a certain future day, etc., and this application does not limit this.

[0181] Exemplarily, after obtaining at least two variable node features of the cross-variable GNN output, direct multi-step (DMS) prediction can be used to enable the decoder to predict multiple steps of MTS at once. As an example, as Figure 13 shown, two multi-layer perceptrons (MLPs) can be used as the decoder, where the first MLP C maps the time dimension of the variable node features from C to 1, and the second MLP T maps the time dimension from the historical input sequence L′ to the output sequence length. The final prediction result can be expressed as the following formula (14):

[0182]

[0183] Therefore, in the embodiments of this application, by extracting at least two second time series data of different time scales from the first time series data, at least one second time series data of different time scales from coarse to fine of the first time series data can be extracted. Furthermore, by interacting with the at least one second time series, at least two time node features related in the at least one second time series data are extracted, by interacting with the at least two time node features, at least two variable node features related in the at least two time node features are extracted, and further, based on the at least two variable node features, third time series data across the at least two variables in a second time period is predicted.

[0184] Figure 14A FIG. shows a schematic diagram of noise signals at different levels in a multi-scale time series. By performing multi-scale extraction on the time series, it can be observed that different scales have significantly different levels of noise intensity, and generally, coarser scales exhibit lower noise intensity. Therefore, by capturing dependencies at different scales, the cross-time relationship is made robust to noise. Figure 14BA schematic diagram showing homogeneous and heterogeneous relationships between variables is presented. Cross-variable interactions in real-world data have both homogeneity and heterogeneity. Both of these relationships can contribute to invariant associations over time. Therefore, learning invariant associations that include homogeneous and heterogeneous relationships between variables can improve their robustness to noise. Accordingly, embodiments of the present application can refine cross-time interactions and cross-variable interactions in noisy MTS and improve the robustness of MTS prediction performance to noise by capturing cross-scale interactions that are insensitive to unexpected input noise and extracting cross-variable relationships between heterogeneous variables.

[0185] Exemplarily, the interactive GNN disclosed in embodiments of the present application is used for MTS prediction to refine cross-time and cross-variable interactions. To handle unexpected noise in the time dimension, AMSI is first designed to construct multi-scale time series with different noise levels. In the time dimension, the dependencies between different scales are modeled by a cross-scale GNN (a temporal correlation graph), where scales with clearer trends and weaker noise will be assigned more edge weights. In the variable dimension, the heterogeneous interaction modeling between variables is introduced into MTS prediction, and a cross-variable GNN is proposed to utilize the homogeneity and heterogeneity between different variables with positive and negative edge weights. By focusing on edges with higher significance scores while constraining those with lower scores, a linear time and space complexity of O(L) for an input sequence length of L is achieved.

[0186] Extensive experiments conducted on 8 real-world MTS datasets demonstrate that CrossGNN proposed in embodiments of the present application is more effective than existing state-of-the-art SOTA methods. Specifically, compared with 7 state-of-the-art models with varying predicted lengths, CrossGNN achieves the top-ranked performance in 51 settings and the top-2 performance in 13 settings. At the same time, it can maintain linear memory occupancy and computational time as the input size increases.

[0187] First, the datasets and experimental settings are described.

[0188] Datasets: Extensive experiments are conducted on 8 real-world datasets, including weather, traffic, exchange rates, electricity, and 4 ETT datasets (ETTh1, ETTh2, ETTm1, and ETTm2). For the last 4 ETT datasets, the dataset is divided into a training set, a validation set, and a test set in a ratio of 6:2:2. For other datasets, the dataset is divided into a training set, a validation set, and a test set in a ratio of 7:1:2.

[0189] Baselines and settings: The embodiments of this application are compared with 7 state-of-the-art methods, including Times-Net; 4 Transformer-based methods: ETSformer, FEDformer, Pyraformer, Autoforme; the GNN-based method: MTGNN; and the simple yet powerful linear model Dlinear. The expected length of all datasets is T ∈ {96, 192, 336, 720}. The default backtracking window L = 96. The mean squared error (MSE) and mean absolute error (MAE) of MTS prediction are calculated as metrics.

[0190] Figure 15 An example of the quantitative results of MTS prediction using different methods is shown. CrossGNN achieves excellent performance on most datasets with various expected length settings, obtaining 51 first-place and 13 second-place rankings out of a total of 64 settings. Quantitatively, compared with the best results that Transformer-based methods can provide, CrossGNN achieves an overall reduction of 10.43% in MSE and 10.11% in MAE. Compared with the GNN-based method MTGNN, CrossGNN achieves a more significant reduction of 22.57% in MSE and 25.74% more significant reduction in MAE. Compared with other strong baselines such as TimesNet and Dlinear, CrossGNN is still generally superior. CrossGNN does not achieve the best performance on the electricity dataset. Further analysis reveals that the more severe out-of-distribution (OOD) problem in the electricity dataset results in lower generalization ability of the learned temporal graph relationships on the test set.

[0191] Noise robustness analysis: To evaluate the robustness of the model to noise, we add Gaussian white noise with different intensities to the original MTS and observe the performance changes of different methods. Figure 16 The MSE results of CrossGNN, ETSformer, and MTGNN at different noise ratios on the ETTm2 dataset are shown, and the input length setting is 96. As the noise ratio increases from 100db to 0db, the mean squared error (MSE) on Cross-GNN (0.177) increases more slowly compared with ETSformer (0.191) and MTGNN (0.205). The quantitative results prove that CrossGNN has good robustness to noisy data and has great advantages in handling unexpected fluctuations. This improvement benefits from the explicit modeling of the interactions at their respective scale levels and variable levels.

[0192] Ablation Study: The ablation study was conducted by removing the corresponding modules from CrossGNN on three datasets. C-AMSI removes the Adaptive Multi-Scale Identifier (AMSI) and directly divides the scale by a fixed length. C-CS removes the Cross-Scale GNN module. C-Hete removes the Cross-Variable GNN module and focuses on modeling the homogeneous correlations between different variables. Analysis Figure 17 of the results in shows that: 1) Removing the Cross-Scale GNN leads to the most significant drop in the predicted metrics, highlighting its strong ability to model the interactions between different scales and time points; 2) The Cross-Variable GNN also greatly improves the model performance, demonstrating the importance of modeling the complex dynamic interactions between different variables; 3) AMSI continuously improves the prediction accuracy, indicating that different scales of MTS contain rich interaction information.

[0193] Hyperparameter Sensitivity: Figure 18 Shows the MSE results of models with different lookback window sizes on four datasets. Among them, the MSE results (Y-axis) of models with different lookback window sizes (X-axis) on ETTh2, ETTm3, traffic, and weather, and the output length is set to 336. As the window size increases, the performance of the Transformer-based model fluctuates, while CrossGNN continuously improves. This indicates that the attention mechanism of the Transformer-based method may pay more attention to temporal noise, but the method provided in the embodiments of the present application can better extract the relationships between different time nodes through the Cross-Scale GNN.

[0194] By setting the number of scales to vary from 4 to 8 and reporting the MSE and MAE results on the weather and traffic datasets. As shown in Figure 19 (a) and Figure 19 (b), it can be seen that after a certain number of scales, the performance improvement becomes less significant, indicating that a certain scale size is sufficient to eliminate most of the effects of temporal noise. Limiting the number of neighbor nodes for each time node is mainly determined by the hyperparameter K. As shown in Figure 19 (c) and Figure 19 (d), we experimented with K values of 10, 15, 20, 25, and 30 and found that CrossGNN is insensitive to the number of K. This indicates that only by focusing on strongly correlated nodes can effective information aggregation be performed in temporal interactions. Among them, the left Y-axis is the MSE result, and the right Y-axis is the MAE result.

[0195] Complexity Analysis: Figure 20Shows the theoretical computational complexity of CrossGNN and existing Transformer-based methods. Among them, L represents the length of the input data. To verify that the time and space complexity of the method of the embodiments of the present application is indeed O(L), TVM is used to implement the GNN calculation part, and the calculation time and memory usage are compared with that of a fully connected graph during inference on ETTh2. The comparative experiment is implemented on an Intel(R) 8255C CPU@2.50 GHZ (CPU with 40GB of memory, centos7.8, and TVM 1.0.0). Figure 21 Shows a schematic diagram of the comparison of the time and memory consumption of ETTh2, and the method proposed by the embodiments of the present application is nearly linear with the input length. Among them, Figure (a) shows the memory occupancy, and Figure (b) shows the time spent on per-batch calculation.

[0196] The specific implementation manners of the present application have been described in detail above in conjunction with the accompanying drawings. However, the present application is not limited to the specific details in the above implementation manners. Within the scope of the technical concept of the present application, various simple modifications can be made to the technical solutions of the present application, and these simple modifications all belong to the protection scope of the present application. For example, the various specific technical features described in the above specific implementation manners can be combined in any suitable manner without contradiction. To avoid unnecessary repetition, the present application will not separately describe various possible combination manners. Again, for example, any combination can be made between the various different implementation manners of the present application, as long as it does not violate the idea of the present application, it should also be regarded as the content disclosed by the present application.

[0197] It should also be understood that in the various method embodiments of the present application, the magnitudes of the serial numbers of the above processes do not mean the order of execution is prior or subsequent. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application. It should be understood that these serial numbers can be interchanged under appropriate circumstances so that the embodiments of the present application described can be implemented in an order other than those shown or described.

[0198] The method embodiments of the present application have been described in detail above. Below, in conjunction with Figures 22 to 23 ,the device embodiments of the present application will be described in detail.

[0199] Figure 22 Is a schematic block diagram of a time series prediction device 10 according to an embodiment of the present application. As Figure 22 shown, the device 10 may include an acquisition unit 11, an extraction unit 12, a first graph neural network 13, a second graph neural network 14, and a prediction unit 15.

[0200] The acquisition unit 11 is configured to acquire first time series data of at least two variables in a first time period;

[0201] An extraction unit 12, configured to extract second time series data of at least two different time scales from the first time series data;

[0202] A first graph neural network 13, configured to interact with the second time series data of the at least two different time scales to obtain at least two time node features; wherein, each of the time node features is obtained according to the previous layer features related to each time node and the previous layer neighbor time node features;

[0203] A second graph neural network 14, configured to interact with the at least two time node features to obtain at least two variable node features; wherein, each of the variable node features is obtained according to the previous layer features related to each variable node and the previous layer neighbor variable node features;

[0204] A prediction unit 15, configured to predict third time series data across the at least two variables in a second time period according to the at least two variable node features.

[0205] In some embodiments, the first graph neural network 13 is specifically configured to:

[0206] Construct a first graph neural network according to at least two of the second time series data; wherein, the first graph neural network includes time nodes of at least two different time scales and first correlation weights between each time node;

[0207] Obtain at least one neighbor time node of each time node;

[0208] Normalize the first correlation weights between each time node and the at least one neighbor time node to obtain first relative correlation weights between each time node and the at least one neighbor time node;

[0209] Obtain the time node feature of each time node according to the first relative correlation weight, the previous layer neighbor time node features related to each time node, and the previous layer features related to each time node.

[0210] In some embodiments, the first graph neural network 13 is specifically configured to:

[0211] Aggregate the previous layer neighbor time node features related to each time node according to the first relative correlation weight to obtain neighbor time node aggregation features of the current layer of each time node;

[0212] Aggregate the neighbor time node features of the current layer for each time node and the features of the previous layer associated with each time node to obtain the time node features of each time node.

[0213] In some embodiments, the first graph neural network 13 is specifically configured to:

[0214] Obtain at least one first neighbor time node at each scale of each time node according to the first correlation weights between each time node;

[0215] Obtain at least one second neighbor time node sharing the same scale of each time node;

[0216] Determine the at least one neighbor time node according to the at least one first neighbor time node and the at least one second neighbor time node.

[0217] Optionally, the number of the first neighbor time nodes at each scale is negatively correlated with the period length at each scale.

[0218] In some embodiments, the second graph neural network 14 is specifically configured to:

[0219] Construct a second graph neural network according to the at least two time node features; wherein the second graph neural network includes at least two variable nodes and second correlation weights between each variable node;

[0220] Obtain at least one neighbor variable node of each variable node;

[0221] Normalize the second correlation weights between each variable node and the at least one neighbor variable node to obtain the second relative correlation weights between each variable node and the at least one neighbor variable node;

[0222] Obtain the variable node features of each variable node according to the second relative correlation weights, the features of the neighbor variable nodes of the previous layer associated with each variable node, and the features of the previous layer associated with each variable node.

[0223] In some embodiments, the second graph neural network 14 is specifically configured to:

[0224] Aggregate the features of the neighbor variable nodes of the previous layer associated with each variable node according to the second relative correlation weights to obtain the aggregated features of the neighbor variable nodes of the current layer of each variable node;

[0225] Aggregate the features of the neighbor variable nodes in the current layer of each variable node and the features of the previous layer related to each variable node to obtain the variable node features of each variable node.

[0226] In some embodiments, the second graph neural network 14 is specifically configured to:

[0227] Obtain at least one homogeneous neighbor variable node and at least one heterogeneous neighbor variable node of each variable node according to the second correlation weights between each variable node. Wherein, the at least one homogeneous neighbor variable node includes at least one variable node with the highest correlation weight, and the at least one heterogeneous neighbor variable node includes at least one variable node with the lowest correlation weight.

[0228] In some embodiments, the second graph neural network 14 is specifically configured to:

[0229] Normalize the second correlation weights between each variable node and the at least one homogeneous neighbor variable node to obtain the relative correlation weights between each variable node and the at least one homogeneous neighbor variable node;

[0230] Normalize the second correlation weights between each variable node and the at least one heterogeneous neighbor variable node to obtain the relative correlation weights between each variable node and the at least one heterogeneous neighbor variable node;

[0231] Obtain the second relative correlation weights according to the relative correlation weights between each variable node and the at least one homogeneous neighbor variable node, and the relative correlation weights between each variable node and the at least one heterogeneous neighbor variable node.

[0232] In some embodiments, the extraction unit 12 is specifically configured to:

[0233] Perform a Fourier transform on the first time series data to obtain the time series data of the first time series data at at least two frequencies;

[0234] Obtain the time series data of S frequencies with the largest amplitude values in the time series data of the at least two frequencies; S is a positive integer greater than 1;

[0235] Perform a downsampling operation on the first time series data according to the frequencies of the time series data of the S frequencies and the time window size of the first time series data to obtain the second time series data of different time scales corresponding to the S frequencies respectively.

[0236] In some embodiments, the extraction unit 12 is specifically configured to:

[0237] Determine the period length of each of the S frequencies according to the frequencies of the time series data of the S frequencies and the time window size of the first time series data;

[0238] Perform a downsampling operation on the first time series data with the period length of each frequency as the kernel size and stride to obtain the second time series data of the time scale corresponding to each frequency.

[0239] In some embodiments, the noise intensities of the second time series data of different time scales are different.

[0240] It should be understood that the apparatus embodiments and the method embodiments can correspond to each other, and similar descriptions can refer to the method embodiments. To avoid repetition, it will not be elaborated here. Specifically, Figure 9 The time unit prediction apparatus 10 shown can execute the above method embodiment, and the foregoing and other operations and / or functions of each module in the time series prediction apparatus 10 respectively correspond to the corresponding processes in the above method 300. For the sake of brevity, it will not be elaborated here.

[0241] In the foregoing, the apparatus of the embodiments of the present application has been described from the perspective of functional modules. It should be understood that the functional modules can be implemented in the form of hardware, or in the form of instructions of software, or in the form of a combination of hardware and software modules. Specifically, the steps of the method embodiments in the present application can be completed by the integrated logic circuit of the hardware in the processor and / or the instructions in the form of software. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by the combination of the hardware and software modules in the decoding processor. Optionally, the software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps in the above method embodiments.

[0242] Figure 23 It is a schematic block diagram of an electronic device provided by an embodiment of the present application.

[0243] As Figure 23 shown, the electronic device 30 may include:

[0244] A memory 33 and a processor 32. The memory 33 is used to store a computer program 34 and transmit the program code 34 to the processor 32. In other words, the processor 32 can call and run the computer program 34 from the memory 33 to implement the time series prediction method in the embodiments of the present application, including:

[0245] Obtain first time series data across at least two variables in a first time period;

[0246] Extract second time series data of at least two different time scales from the first time series data;

[0247] Interact with the second time series data of the at least two different time scales to obtain at least two time node features; wherein, each time node feature is obtained based on the previous layer features and previous layer neighbor time node features associated with each time node;

[0248] Use a second graph neural network to interact with the at least two time node features to obtain at least two variable node features; wherein, each variable node feature is obtained based on the previous layer features and previous layer neighbor variable node features associated with each variable node;

[0249] Predict third time series data across the at least two variables in a second time period based on the at least two variable node features..

[0250] For example, the processor 32 can be used to execute the steps in the above method 300 according to the instructions in the computer program 34.

[0251] In some embodiments of the present application, the processor 32 may include but is not limited to:

[0252] General-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and so on.

[0253] In some embodiments of the present application, the memory 33 includes but is not limited to:

[0254] Volatile memory and / or non-volatile memory. Among them, the non-volatile memory can be a Read-Only Memory (ROM), a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically Erasable PROM (EEPROM), or a flash memory. The volatile memory can be a Random Access Memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as Static RAM (SRAM), Dynamic RAM (DRAM), Synchronous DRAM (SDRAM), Double DataRate SDRAM (DDR SDRAM), Enhanced SDRAM (ESDRAM), synch link DRAM (SLDRAM), and Direct Rambus RAM (DR RAM).

[0255] In some embodiments of the present application, the computer program 34 can be divided into one or more units, and the one or more units are stored in the memory 33 and executed by the processor 32 to complete the method provided by the present application. The one or more units can be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program 34 in the electronic device 30.

[0256] Optionally, as Figure 23 shown, the electronic device 30 may further include:

[0257] A transceiver 33, which can be connected to the processor 32 or the memory 33.

[0258] Among them, the processor 32 can control the transceiver 33 to communicate with other devices. Specifically, it can send information or data to other devices, or receive information or data sent by other devices. The transceiver 33 can include a transmitter and a receiver. The transceiver 33 may further include an antenna, and the number of antennas can be one or more. It should be understood that the various components in the electronic device are connected through a bus system, where the bus system includes, in addition to the data bus, a power bus, a control bus, and a status signal bus.

[0259] The present application also provides a computer storage medium, on which a computer program is stored. When the computer program is executed by a computer, the computer is enabled to execute the methods in the above method embodiments. Or rather, the embodiments of the present application also provide a computer program product containing instructions, which, when executed by a computer, enable the computer to execute the methods in the above method embodiments.

[0260] When implemented using software, it may be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from a website, a computer, a server, or a data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium accessible by a computer or a data storage device such as a server or a data center integrating one or more available media. The available medium may be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a digital video disc (DVD)), or a semiconductor medium (such as a solid state disk (SSD)), etc.

[0261] It can be understood that in the specific implementation of the present application, when the above embodiments of the present application are applied to specific products or technologies and involve relevant data such as user information, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.

[0262] Those of ordinary skill in the art can realize that the modules and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0263] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or modules can be in electrical, mechanical, or other forms.

[0264] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules, that is, they can be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. For example, in each embodiment of the present application, the various functional modules can be integrated in a processing module, or each module can exist physically alone, or two or more modules can be integrated in one module.

[0265] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A time series prediction method, characterized in that, it includes: Obtain first time series data of at least two variables in a first time period; Extract second time series data of at least two different time scales from the first time series data; Interact with the second time series data of the at least two different time scales to obtain at least two time node features; wherein, each time node feature is obtained according to the previous layer features related to each time node and the previous layer neighbor time node features; Interact with the at least two time node features to obtain at least two variable node features; wherein, each variable node feature is obtained according to the previous layer features related to each variable node and the previous layer neighbor variable node features; Predict third time series data of the at least two variables in a second time period according to the at least two variable node features.

2. The method according to claim 1, characterized in that, The interacting with the second time series data of the at least two different time scales to obtain at least two time node features includes: Construct a first graph neural network according to at least two of the second time series data; wherein, the first graph neural network includes time nodes of at least two different time scales and first correlation weights between each time node; Obtain at least one neighbor time node of each time node; Normalize the first correlation weights between each time node and the at least one neighbor time node to obtain first relative correlation weights between each time node and the at least one neighbor time node; Obtain the time node feature of each time node according to the first relative correlation weight, the previous layer neighbor time node features related to each time node, and the previous layer features related to each time node.

3. The method according to claim 2, characterized in that, The obtaining the time node feature of each time node according to the first relative correlation weight, the previous layer neighbor time node features related to each time node, and the previous layer features related to each time node includes: Aggregate the previous layer neighbor time node features related to each time node according to the first relative correlation weight to obtain the neighbor time node aggregation feature of the current layer of each time node; Obtain the time node feature of each time node according to the neighbor time node aggregation feature of the current layer of each time node and the previous layer features related to each time node.

4. The method according to claim 2, characterized in that, The obtaining at least one neighbor time node of each time node includes: Obtain at least one first neighbor time node of each time node on each scale according to the first correlation weights between each time node; Obtain at least one second neighbor time node of each time node sharing the same scale; Determine the at least one neighbor time node according to the at least one first neighbor time node and the at least one second neighbor time node.

5. The method according to claim 4, wherein, the number of the first neighbor time nodes at each scale is negatively correlated with the period length of each scale.

6. The method according to claim 1, wherein, the interacting of the at least two time node features to obtain at least two variable node features includes: constructing a second graph neural network according to the at least two time node features; wherein, the second graph neural network includes at least two variable nodes and second correlation weights between each variable node; obtaining at least one neighbor variable node of each variable node; normalizing the second correlation weights between each variable node and the at least one neighbor variable node to obtain second relative correlation weights between each variable node and the at least one neighbor variable node; obtaining the variable node feature of each variable node according to the second relative correlation weights, the previous layer neighbor variable node features related to each variable node, and the previous layer features related to each variable node.

7. The method according to claim 6, wherein, the obtaining the variable node feature of each variable node according to the second relative correlation weights, the previous layer neighbor variable node features related to each variable node, and the previous layer features related to each variable node includes: aggregating the previous layer neighbor variable node features related to each variable node according to the second relative correlation weights to obtain the neighbor variable node aggregation feature of each variable node at the current layer; obtaining the variable node feature of each variable node according to the neighbor variable node aggregation feature of each variable node at the current layer and the previous layer features related to each variable node.

8. The method according to claim 6, wherein, the obtaining at least one neighbor variable node of each variable node includes: obtaining at least one homogeneous neighbor variable node and at least one heterogeneous neighbor variable node of each variable node according to the second correlation weights between each variable node; wherein, the at least one homogeneous neighbor variable node includes at least one variable node with the highest correlation weight, and the at least one heterogeneous neighbor variable node includes at least one variable node with the lowest correlation weight.

9. The method according to claim 8, wherein, the normalizing the second correlation weights between each variable node and the at least one neighbor variable node to obtain second relative correlation weights between each variable node and the at least one neighbor variable node includes: normalizing the second correlation weights between each variable node and the at least one homogeneous neighbor variable node to obtain relative correlation weights between each variable node and the at least one homogeneous neighbor variable node; Normalize the second correlation weight between each of the variable nodes and the at least one heterogeneous neighbor variable node to obtain the relative correlation weight between each of the variable nodes and the at least one heterogeneous neighbor variable node; Obtain the second relative correlation weight according to the relative correlation weight between each of the variable nodes and the at least one homogeneous neighbor variable node, and the relative correlation weight between each of the variable nodes and the at least one heterogeneous neighbor variable node.

10. The method according to claim 1, wherein, the extracting at least two second time series data with different time scales from the first time series data includes: Performing a Fourier transform on the first time series data to obtain time series data of the first time series data at at least two frequencies; Obtain the time series data of S frequencies with the largest amplitude values among the time series data of the at least two frequencies; S is a positive integer greater than 1; Perform a downsampling operation on the first time series data according to the frequencies of the time series data of the S frequencies and the time window size of the first time series data to obtain the second time series data with different time scales corresponding to the S frequencies respectively.

11. The method according to claim 10, wherein, the performing a downsampling operation on the first time series data according to the frequencies of the time series data of the S frequencies and the time window size of the first time series data to obtain the second time series data with different time scales corresponding to the S frequencies respectively includes: Determine the period length of each of the S frequencies according to the frequencies of the time series data of the S frequencies and the time window size of the first time series data; Perform a downsampling operation on the first time series data with the period length of each frequency as the kernel size and stride to obtain the second time series data with the time scale corresponding to each frequency.

12. The method according to any one of claims 1-11, wherein, the noise intensities of the second time series data with different time scales are different.

13. A time series prediction device, wherein, comprises: An acquisition unit for acquiring first time series data of at least two variables in a first time period; An extraction unit for extracting at least two second time series data with different time scales from the first time series data; A first graph neural network for interacting the at least two second time series data with different time scales to obtain at least two time node features; wherein, each of the time node features is obtained according to the previous layer features related to each time node and the previous layer neighbor time node features; A second graph neural network for interacting the at least two time node features to obtain at least two variable node features; wherein, each of the variable node features is obtained according to the previous layer features related to each variable node and the previous layer neighbor variable node features; A prediction unit, configured to predict third time series data across the at least two variables in a second time period according to the at least two variable node features.

14. An electronic device, characterized in that it includes a processor and a memory, and instructions are stored in the memory. When the processor executes the instructions, the processor executes the method according to any one of claims 1-12.

15. A computer storage medium, characterized in that it is used to store a computer program, and the computer program includes a method for executing any one of claims 1-12.

16. A computer program product, characterized in that it includes computer program code. When the computer program code is run by an electronic device, the electronic device executes the method according to any one of claims 1-12.

Citation Information

Cited By

  • Time sequence prediction method and system based on double-domain feature fusion

    CN121542613A