Time sequence data prediction method and system of time variable double-domain jump attention mechanism

Through the method of the time variable double-domain jump attention mechanism, the season and trend parts are decomposed and processed, and the two-way interaction between the time and channel domains is achieved, solving the problem of insufficient prediction accuracy in traditional methods, and improving the accuracy and adaptability of time series prediction.

CN120278332APending Publication Date: 2025-07-08XI AN JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510395369.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

When facing complex situations, traditional time series prediction methods fail to fully consider the complex interaction between time and channel domains, resulting in insufficient prediction accuracy and difficulty in dealing with seasonal fluctuations and sudden trend changes, affecting the accuracy of industry decisions.

Method used

The time variable double-domain jump attention mechanism is adopted, and the multi-grained feature refining, double-domain jump mixing and multi-predictor mixing methods are combined with multi-layer perceptron and self-attention mechanisms to decompose seasonal and trend parts, perform feature projection and prediction, and achieve two-way interaction between time and channel domains.

Benefits of technology

It significantly improves the accuracy and reliability of time series prediction, can better capture hidden patterns and dynamic changes in the data, adapt to the multi-dimensional timing prediction needs in complex real-life scenarios, and improves the coherence and accuracy of prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120278332A_ABST
    Figure CN120278332A_ABST
Patent Text Reader

Abstract

The invention discloses a time sequence data prediction method and system for a time variable double-domain jump attention mechanism, and the method comprises the steps: carrying out the processing of historical time sequence observation data through different kernel sizes through iteratively applying 1-D average pooling operation, obtaining a group of time sequences highlighting different particle size changes, and carrying out the prediction of the time sequence data of the time variable double-domain jump attention mechanism. Performing normalization and embedding operation on each time sequence; for the kth layer, decomposing the multi-resolution time sequence # imgabs0 # into a seasonal part and a trend part, and performing feature projection on the seasonal part by adopting a multi-layer sensor to obtain Skout; for the trend part, firstly dividing the trend part into a plurality of patches # imgabs1 #, then dividing the patches # imgabs1 # into sub-patches # imgabs2 #, grouping the patches # imgabs1 # and the sub-patches # imgabs2 # according to the same phase, regarding each group as a token, performing processing through a self-attention mechanism to obtain trend feature output Tkout, and adding seasons and trend features to obtain feature output Xkout; and based on the obtained feature output Xkout, adopting a corresponding predictor to predict to obtain a prediction result, and then accumulating to obtain a final future time sequence prediction value.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] In today's increasingly digital social environment, time series data emerges like a tide in various key fields and becomes an important basis for industry decision-making and operation management. In the field of traffic planning, the time series data of urban traffic flow records the real-time dynamic changes of vehicles on the road. Through the analysis and prediction of this data, traffic management departments can formulate scientific and reasonable traffic signal timing plans in advance, plan road construction and renovation projects, and optimize the allocation of public transport capacity, thereby effectively alleviating traffic congestion and improving the operation efficiency of urban traffic. For example, during the morning and evening rush hours, according to the time series change rules of historical traffic flow, accurately predict the traffic volume of each section, adjust the signal light duration in a timely manner, and guide vehicle diversion to avoid traffic paralysis.

[0003] In weather forecasting, the time series data of meteorological elements such as temperature, air pressure, humidity, and wind speed are the core basis for meteorologists to predict weather changes. Accurate weather prediction has immeasurable value in aspects such as agricultural production, aviation and navigation safety, and energy supply and demand balance. For example, in the agricultural field, accurate precipitation and temperature prediction can guide farmers to reasonably arrange farming activities, select appropriate sowing, irrigation, and harvesting times, reduce the impact of natural disasters on crops, and ensure agricultural harvests; in the aviation field, knowing the approach of bad weather in advance, airlines can adjust flight plans in a timely manner to ensure flight safety, reduce passenger delays and economic losses.

[0004] The energy consumption field also highly relies on time series prediction technology. Power companies need to predict the electricity demand in different regions, different seasons, and different time periods based on the time series data of electricity consumption, so as to reasonably arrange power generation plans, dispatch power resources, and maintain the stable operation of the power grid. Industrial enterprises can optimize production processes, reduce energy costs, improve energy utilization efficiency, and enhance their competitiveness by analyzing the time series of their own energy consumption.

[0005] However, traditional time series prediction methods have revealed many deficiencies when dealing with the complex situations of the real world. Many traditional methods only focus on analyzing the information in the time domain or channel domain separately, failing to fully consider the inconsistencies of time series characteristics within the sequence and the complex interaction relationships between the time and channel domains. For example, some methods based on simple statistical models often cannot accurately capture the hidden patterns and dynamic changes in the data when dealing with time series with seasonal fluctuations and sudden trend changes, resulting in large deviations in the prediction results. In practical applications, such as the electricity market affected by various factors such as seasonal factors, economic situations, and emergencies (such as a sharp increase in electricity demand caused by extreme weather), traditional methods are difficult to comprehensively consider these complex factors and cannot provide accurate electricity load forecasts for power companies, which may lead to problems of electricity supply shortages or surpluses, affecting the normal operation of society and the economic benefits of enterprises.

[0006] With the continuous growth of data volume and the intensification of data complexity, the limitations of traditional methods have become increasingly prominent. There is an urgent need for an innovative prediction method that can comprehensively integrate time and channel information to meet the urgent needs of various industries for high-precision time series prediction. Summary of the Invention

[0007] The technical problem to be solved by the present invention is to provide a time series data prediction method and system with a time-variable dual-domain jump attention mechanism in view of the above deficiencies in the prior art, so as to solve the technical problem of insufficient accuracy of traditional methods in the face of complex time series data and provide more reliable and accurate prediction technical support for many industries relying on time series prediction.

[0008] The present invention adopts the following technical solutions:

[0009] A time series data prediction method with a time-variable dual-domain jump attention mechanism includes the following steps:

[0010] Receive historical time series observation data containing T time steps and N variables;

[0011] By iteratively applying 1-D average pooling operations to process the historical time series observation data with different kernel sizes, a set of time series highlighting different granularity changes is obtained while keeping the data length unchanged, and then each time series is normalized and embedded.

[0012] For the k-th layer, the multi-resolution time series is decomposed into a seasonal part S k and a trend part T k , and the seasonal part is projected by a multi-layer perceptron to obtain S kout ; for the trend part, the trend part is first divided into multiple patches T kp , and then divided into sub-patches After grouping by the same phase, each group is regarded as a token and processed through the self-attention mechanism to obtain the trend feature output T kout , and add the seasonal and trend features to obtain the feature output X kout ;

[0013] Based on the obtained feature output X kout , use the corresponding predictor to make a prediction to obtain the prediction result, and then accumulate each prediction result to obtain the final future time series prediction value.

[0014] Preferably, the normalization adopts Z-score normalization or Min-Max normalization.

[0015] Preferably, by performing average pooling on X k , during the pooling process, sample and average the data, and fill it with the average value along the time direction after pooling.

[0016] Preferably, when filling, according to the time series characteristics and statistical laws of the data, choose simple arithmetic average or weighted average.

[0017] Preferably,

[0018] X k ' = AvgPool(X k , p k )

[0019] where X k is the input time series, AvgPool is average pooling, taking the average of the specified area, X k ' is the pooled sequence, and p k is the range of the specified area.

[0020] Preferably, the multi-layer perceptron has multiple hidden layers, the number of neurons in each layer gradually decreases and uses a non-linear activation function.

[0021] Preferably, after grouping by the same phase, each group is regarded as a token for self-attention mechanism calculation. During the calculation process, follow the standard attention calculation framework. For the query q i and key k j of each token, accurately calculate the pre-softmax score A i,j = (QK T / √d) i,j , d is the feature dimension; multiply by the value V after obtaining the attention weights through the softmax function to integrate the information of the grouped sub-patches.

[0022] Preferably, before accumulating the prediction results of each granularity, each prediction result is assigned a weight.

[0023] Preferably, the predictor Pred k adopts a multi-layer perceptron structure.

[0024] In a second aspect, an embodiment of the present invention provides a time-variable dual-domain jump attention mechanism-based time series data prediction system, including:

[0025] A data module that receives historical time series observation data containing T time steps and N variables;

[0026] A preprocessing module that processes the historical time series observation data with different kernel sizes by iteratively applying 1-D average pooling operations, and obtains a set of time series that highlight different granularity changes while keeping the data length unchanged. Subsequently, each time series is normalized and embedded;

[0027] A decomposition module. For the k-th layer, the multi-resolution time series is decomposed into a seasonal part S k and a trend part T k . The seasonal part is subjected to feature projection using a multi-layer perceptron to obtain S kout . For the trend part, the trend part is first divided into multiple patches T k p , and then divided into sub-patches . After grouping by the same phase, each group is regarded as a token and processed through a self-attention mechanism to obtain a trend feature output T kout , and the seasonal and trend features are added to obtain a feature output X kout ;

[0028] A prediction module that, based on the obtained feature output X kout , uses a corresponding predictor to make a prediction to obtain a prediction result, and then accumulates each prediction result to obtain a final predicted value of the future time series.

[0029] In a third aspect, a computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above time-variable dual-domain jump attention mechanism-based time series data prediction method are implemented.

[0030] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, including a computer program. When the computer program is executed by a processor, the steps of the above time-variable dual-domain jump attention mechanism-based time series data prediction method are implemented.

[0031] In a fifth aspect, a chip includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the time-variable dual-domain jumping attention mechanism-based time series prediction method are implemented.

[0032] In a sixth aspect, an embodiment of the present invention provides an electronic device including a computer program. When the computer program is executed by the electronic device, the steps of the time-variable dual-domain jumping attention mechanism-based time series prediction method are implemented.

[0033] Compared with the prior art, the present invention has at least the following beneficial effects:

[0034] A time-variable dual-domain jumping attention mechanism-based time series prediction method applies a cross-attention mechanism in the time-channel joint space, effectively realizing two-way interaction between the two domains and deeply mining the hidden associations across time and variables; HopMixer mainly includes three core modules: a multi-granularity feature refinement (MFR) block, a dual-domain jump mixing (DHM) block, and a multi-predictor mixing (MPM) block; the MFR block can highlight the changes under different granularities and horizons without changing the time series structure, effectively protecting the non-periodic details; the DHM block further applies the cross-attention mechanism to explore the hidden connections between subsequences of different times and variables; the MPM block integrates the results of multiple predictors to infer the future sequence; verified by experiments on a large number of public data sets, the performance of HopMixer is significantly better than the existing methods, providing a highly innovative solution for time series prediction tasks, strongly promoting the development of technologies in this field, and expected to greatly improve the prediction accuracy and reliability in many practical application scenarios, bringing significant economic and social benefits.

[0035] Furthermore, multi-variable historical time series data (T time steps, N variables) is received, and the original time structure and the association between variables are retained, providing a complete input for subsequent processing and adapting to the multi-dimensional time series prediction requirements in complex real-world scenarios.

[0036] Furthermore, multi-core average pooling is used to extract features of different granularities, enhancing the model's sensitivity to short-term fluctuations and long-term trends; normalization eliminates the dimensionality difference, and embedding mapping improves the feature expressiveness, while keeping the sequence length to avoid information loss, providing a multi-scale and robust input for downstream tasks.

[0037] Furthermore, seasonal-trend decomposition reduces the modeling complexity, and the MLP efficiently captures the periodicity of the seasonal component; after the trend is segmented, self-attention of sub-patches aggregates the global dependencies, taking into account both local details and long-range associations, and the dual-path feature fusion enhances the model's representation ability for complex patterns.

[0038] Furthermore, the hierarchical prediction results are accumulated and integrated to contribute multi-scale features, mitigating the single-layer prediction bias and enhancing the model generalization; the progressive fusion fully preserves the seasonal periodicity and trend evolution information, improving the coherence and accuracy of multi-step prediction.

[0039] It can be understood that the beneficial effects of the second to sixth aspects above can be referred to the relevant descriptions in the first aspect above, and will not be elaborated here.

[0040] In summary, the present invention proposes a brand-new time series prediction method, which can be widely applied to multiple fields such as transportation and energy. At the same time, the present invention also has a high prediction accuracy, achieving accurate prediction of complex time series in multiple fields.

[0041] The technical solutions of the present invention will be further described in detail below with reference to the drawings and embodiments. Description of the Drawings

[0042] To more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required to be used in the embodiments of the present application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.

[0043] Figure 1 Schematic diagram of the time series processed by the present invention;

[0044] Figure 2 Results graph of applying FFT (Fast Fourier Transform) to multiple time series;

[0045] Figure 3 Overall design diagram of the method of the present invention;

[0046] Figure 4 Implementation details diagram of jump attention;

[0047] Figure 5 Schematic diagram of a computer device provided by an embodiment of the present invention;

[0048] Figure 6 Block diagram of an electronic device provided by an embodiment of the present invention.

[0049] Among them, 60. computer device; 61. processor; 62. memory; 63. computer program; 600. electronic device; 610. processing unit; 620. storage unit; 6201. random access storage unit; 6202. cache storage unit; 6203. read-only storage unit; 6204. program / utilities; 6205. program module; 630. bus; 640. display unit; 650. input / output interface; 660. network adapter; 700. external device. Detailed implementation manners

[0050] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0051] In the description of the present invention, it should be understood that the terms "including" and "comprising" indicate the presence of the described features, wholes, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their combinations.

[0052] It should also be understood that the terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms.

[0053] It should be further understood that the term " / and" as used in the specification of the present invention and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations. For example, A and / or B can represent three cases: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in the present invention generally represents an "or" relationship between the contextually related objects.

[0054] It should be understood that although the terms first, second, third, etc. may be used in the embodiments of the present invention to describe preset ranges, etc., these preset ranges should not be limited to these terms. These terms are only used to distinguish the preset ranges from each other. For example, without departing from the scope of the embodiments of the present invention, the first preset range may also be referred to as the second preset range, and similarly, the second preset range may also be referred to as the first preset range.

[0055] Depending on the context, as used herein, the word "if" can be interpreted as "when" or "while" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detecting (stated condition or event)" can be interpreted as "when determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)".

[0056] Schematic diagrams of various structures according to the disclosed embodiments of the present invention are shown in the accompanying drawings. These figures are not drawn to scale, where for the purpose of clear illustration, some details are enlarged and some details may be omitted. The shapes of various regions and layers shown in the figures and their relative sizes and positional relationships are merely exemplary, and in practice, there may be deviations due to manufacturing tolerances or technical limitations, and those skilled in the art can design regions / layers with different shapes, sizes, and relative positions according to actual needs.

[0057] The present invention provides a method for predicting time series data with a time-varying dual-domain jump attention mechanism, which uses a multi-granularity feature refinement (MFR) block to extract time series at different granularities, and then applies a dual-domain jump mixing (DHM) block to process the trend components of sequences at different granularities. Finally, a multi-layer prediction mixing (MPM) block is used to obtain the final prediction result, achieving high-precision prediction of time series, which can be widely applied to multiple fields.

[0058] Embodiment 1

[0059] Please refer to Figure 3 , a method for predicting time series data with a time-varying dual-domain jump attention mechanism of the present invention, a brand-new time series prediction architecture - HopMixer, which mainly consists of three key parts: a multi-granularity feature refinement (MFR) block, a dual-domain jump mixing (DHM) block, and a multi-layer prediction mixing (MPM) block. Each part closely cooperates to jointly achieve efficient and accurate time series prediction, including the following steps:

[0060] S1. Data input;

[0061] Receive historical time series observation data containing T time steps and N variables Ensure the integrity and accuracy of the data. The data source can cover various channels such as sensor acquisition, database storage, or network transmission, and the data needs to be preliminarily cleaned and verified to remove obvious error data and outliers.

[0062] S2. Multi-granularity feature refinement;

[0063] By iteratively applying 1-D average pooling operations with different kernel sizes p kProcess the input data, where the selection of the kernel size is scientifically determined based on the time - period characteristics of the data, the change frequency of variables, and the required granularity level;

[0064] While keeping the data length unchanged, through precise calculations and data filling operations, obtain a set of time series {X0, …, X k}; The specific filling method is to fill along the time direction with the average value. The filling process must strictly follow the established algorithm logic to ensure the continuity and consistency of the data;

[0065] Subsequently, perform normalization and embedding operations on each time series. For normalization, advanced standardization methods such as Z - score normalization or Min - Max normalization are used to eliminate the dimensional differences and non - stationarity effects of the data;

[0066] For the embedding operation, utilize a pre - trained professional embedding model or an embedding layer customized according to the data characteristics to convert the data into a vector representation suitable for subsequent complex calculations The parameter settings of the embedding model need to go through a large number of experiments and optimizations to ensure the optimality of the embedding effect.

[0067] Through X k Perform average pooling operation. During the pooling process, precise sampling and average calculations of the data are required to ensure accurate data processing within each pooling window. And when filling with the average value along the time direction after pooling, according to the time - series characteristics and statistical laws of the data, select a suitable average - value calculation method, such as simple arithmetic average or weighted average, etc., to ensure the rationality and effectiveness of the filled data.

[0068] X k ' = AvgPool(X k , p k )

[0069] Where, X k is the input time series, AvgPool is the average pooling, performing pooling operation on the specified area, and p k is the range of the specified area.

[0070] The primary goal of the Multi - Granularity Feature Refinement (MFR) block is to extract rich feature information that can reflect different granularity changes without destroying the original structure of the time series, while effectively protecting the non - periodic details in the data, which often contain important prediction clues. To achieve this goal, the MFR block adopts iterative 1 - D average pooling operation, and the design of its kernel size p k is based on in - depth analysis of data characteristics and a large number of experimental verifications.

[0071] During the processing, for the input time series data Through precise calculations and data filling operations, a set of multi-granularity time series {X0, …, X k} is generated. For example, when processing energy consumption data, if the data exhibits obvious daily and weekly periodic characteristics, an appropriate kernel size can be selected so that the fluctuations within the daily cycle can be captured in a single pooling operation. After multiple pooling operations, the trend changes of the weekly cycle and even the monthly cycle can be further highlighted.

[0072] In the normalization operation, different standardization methods are selected according to the distribution characteristics and statistical properties of the data:

[0073] S101. Z-score standardization is applicable to the situation where the data distribution is relatively uniform and there are no obvious outliers. By converting the data into a standard normal distribution with a mean of 0 and a standard deviation of 1, the dimensional difference of the data and the influence of some non-stationarity are eliminated.

[0074] S102. Min-Max standardization is more suitable for scenarios where the data range is known and the data needs to be mapped to a specific interval (such as [0, 1]), and it can highlight the relative change relationship of the data within this interval.

[0075] The embedding operation uses advanced deep learning techniques, such as pre-trained word vector models or embedding layers specifically designed for time series data, to convert the time series data into a low-dimensional dense vector representation form. These vectors can better capture the semantic information and potential features of the data, providing strong support for subsequent analysis and prediction.

[0076] S3. Dual-domain jump mixing;

[0077] For the k-th layer, the multi-resolution time series is decomposed into a seasonal part and a trend part The decomposition process uses advanced series decomposition algorithms, such as decomposition methods based on moving average, exponential smoothing, or singular spectrum analysis, to ensure the accuracy and effectiveness of the decomposition;

[0078] For the seasonal part, a multi-layer perceptron is used for feature projection to obtain S kout = MLP(S k ). The structure design of the multi-layer perceptron needs to be carefully adjusted according to the feature dimension, complexity, and required output dimension of the seasonal data. Parameters such as the number of layers, the number of neurons in each layer, and the selection of activation functions need to be strictly optimized and verified;

[0079] For the trend part, it is first divided into multiple patches T k p, the basis for patch division is the time period and internal variation law of the data, and it is further divided into sub-patches The division of sub-patches needs to consider the local stability and correlation of the data;

[0080] After grouping by the same phase, each group is regarded as a token and processed through the self-attention mechanism, and finally the trend feature output T is obtained kout , and the seasonal and trend features are added to obtain X kout , and the addition process needs to ensure the accuracy and consistency of the data.

[0081] For the patch division and sub-patch division of the trend part, a dynamic division strategy needs to be adopted according to the time period and change characteristics of the data. That is, according to the change rate and stability of different time periods or data segments, the size and quantity of patches and sub-patches are adaptively adjusted to ensure that the local features and trend changes of the data can be accurately captured. At the same time, in the calculation process of the self-attention mechanism, in order to improve the calculation efficiency and reduce the consumption of computing resources, sparse attention technology or approximate calculation methods can be adopted, but it is necessary to ensure that the ability to capture key information and prediction accuracy are not significantly affected.

[0082] The Dual-Domain Hybrid (DHM) block is specially designed for the unique properties and complex relationships of seasonal and trend components in time series.

[0083] First, use an efficient series decomposition algorithm to decompose the multi-resolution time series into the seasonal part S k and the trend part T k ;

[0084] The decomposition method based on moving average can effectively smooth the data and separate the trend component by calculating the average value of the data within a certain time window;

[0085] Singular spectrum analysis uses the eigenvalue decomposition of the covariance matrix of the data to decompose the data into subsequences of different frequencies, so as to accurately extract the seasonal and trend components.

[0086] Due to the different characteristics of the seasonal part and the trend part, different processing methods are adopted:

[0087] S301. For the seasonal part S k , due to being affected by various accidental factors and Gaussian noise, it shows high dynamics. In order to remove noise and retain key learning features, a multi-layer perceptron (MLP) is used for feature projection.

[0088] The structural design of the MLP is carefully optimized. Based on the characteristic dimensions, complexity, and expected output dimensions of the seasonal data, the appropriate number of layers, the number of neurons in each layer, and the activation function are determined. For example, for meteorological data with complex seasonal patterns and a lot of noise, an MLP with multiple hidden layers, a gradually decreasing number of neurons in each layer, and a non-linear activation function such as ReLU may need to be designed to enhance the model's expressive ability and learning ability, ensuring that while removing noise, the effective information related to seasons is maximally retained.

[0089] S302. For the trend part T k , considering its complex variability in terms of time periodicity and channel correlation, a unique skip attention mechanism is designed;

[0090] During the patch division process, according to the time period and internal change law of the data, the trend sequence is divided into multiple patches T k p . For example, when analyzing the time series of stock prices, if the stock price fluctuates frequently in the short term but shows a certain periodic trend in the long term, the patch size can be reasonably determined according to factors such as the market cycle (such as bull and bear cycles) and the company's financial report cycle. The patch is further divided into sub-patches When doing so, the local stability and correlation of the data are fully considered to ensure that the sub-patches can reflect the local characteristics of the data.

[0091] After grouping by the same phase, each group is regarded as a token for self-attention mechanism calculation. During the calculation process, the standard attention calculation framework is strictly followed. For the query q i and key k j of each token, the pre-softmax score A i,j =(QK T / √d) i,j is accurately calculated, where d is the feature dimension. After obtaining the attention weights through the softmax function and multiplying them by the value V, the information of the grouped sub-patches is integrated, enabling the flexible associations between different times and variables to be keenly captured. For example, when analyzing global energy market data, through the skip attention mechanism, the potential connections between the energy price trends in different regions and factors such as the global economic situation and geopolitical events can be discovered, providing key information support for accurately predicting the energy price trend.

[0092] Finally, the processed seasonal and trend features are added together to obtain X kout , and a high-precision numerical calculation method is used during the addition process to ensure the accuracy and consistency of the data and avoid prediction deviations caused by calculation errors.

[0093] S4. Multi-layer prediction mixing.

[0094] Feature output X for each granularity kout , and use the corresponding predictor Pred k to make predictions to obtain

[0095] Predictor Pred k Adopts a multi-layer perceptron structure, and its design needs to be customized according to the characteristics of different granularity data and the requirements of the prediction target. Parameters such as the number of layers, the number of neurons, and the connection method need to go through a large number of experiments and optimizations.

[0096] Then accumulate the prediction results of each granularity to obtain the final future time series prediction value The accumulation process needs to adopt a high-precision calculation method to ensure the accuracy of the results.

[0097] To further improve the prediction accuracy and stability, before accumulating the prediction results of each granularity, each prediction result can be weighted. The determination of the weights can be based on the granularity importance of the data, historical prediction errors, or other relevant factors, and can be dynamically adjusted through machine learning algorithms or statistical analysis methods, so that the final prediction result can more reasonably integrate the information of each granularity and improve the overall prediction performance.

[0098] For example, for fine-grained time series data, its predictor may require a more complex structure to capture fast-changing detailed information, such as increasing the number of hidden layers and the density of neurons to improve the sensitivity of the model; while for coarse-grained data, the predictor can be relatively simplified and focus on extracting long-term trend information.

[0099] During the training process, a large amount of historical data is used, and advanced optimization algorithms (such as stochastic gradient descent, Adagrad, Adadelta, etc.) are used to continuously adjust the parameters of the predictor to minimize the prediction error. For example, when dealing with traffic flow prediction problems, traffic flow data from different time periods and different weather conditions in the past few years are used to train the predictor, so that the predictor can learn the changing rules of traffic flow in various situations.

[0100] The determination of the weights can be based on the granularity importance of the data, historical prediction errors, or other relevant factors. For example, if a certain granularity of time series has shown high accuracy in past predictions, its weight in the accumulation process can be appropriately increased; or the weights can be assigned according to the influence degree of different granularity data on the prediction target, so that the final prediction result can more reasonably integrate the information of each granularity and improve the overall prediction performance.

[0101] Those skilled in the art can understand that various aspects of the present invention can be implemented as a system, a method, or a program product. Therefore, various aspects of the present invention can be specifically implemented in the following forms: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuit", "module", or "platform" here.

[0102] Embodiment 2

[0103] The present invention provides a time series data prediction system with a time-varying dual-domain jump attention mechanism, which can be used to implement the time series data prediction method of the above-mentioned time-varying dual-domain jump attention mechanism. Specifically, the time series data prediction system with the time-varying dual-domain jump attention mechanism includes a data module, a preprocessing module, a decomposition module, and a prediction module.

[0104] Among them, the data module receives historical time series observation data containing T time steps and N variables;

[0105] The preprocessing module processes the historical time series observation data with different kernel sizes by iteratively applying 1-D average pooling operations, and obtains a set of time series highlighting different granularity changes while keeping the data length unchanged. Subsequently, normalization and embedding operations are performed on each time series;

[0106] For the k-th layer, the decomposition module decomposes the multi-resolution time series into a seasonal part S k and a trend part T k , projects the features of the seasonal part using a multi-layer perceptron to obtain S kout ; for the trend part, first divide the trend part into multiple patches T k p , then divide it into sub-patches , group each group with the same phase as a token, and process it through the self-attention mechanism to obtain the trend feature output T kout , and add the seasonal and trend features to obtain the feature output X kout ;

[0107] The prediction module, based on the obtained feature output X kout , uses a corresponding predictor to make a prediction to obtain a prediction result, and then accumulates each prediction result to obtain the final future time series prediction value.

[0108] Embodiment 3

[0109] The present invention provides a terminal device, which includes a processor and a memory. The memory is used to store a computer program, and the computer program includes program instructions. The processor is used to execute the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Graphics Processing Units (GPU), Tensor Processing Units (TPU), Digital Signal Processors (DSP), Application Specific Integrated Circuits (ASIC), Field-Programmable Gate Arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is suitable for implementing one or more instructions, specifically suitable for loading and executing one or more instructions to implement the corresponding method flow or corresponding function; the processor described in the embodiments of the present invention can be used for the operation of the time series prediction method of the time-varying dual-domain jump attention mechanism, including:

[0110] Receiving historical time series observation data including T time steps and N variables; processing the historical time series observation data with different kernel sizes by iteratively applying 1-D average pooling operations to obtain a set of time series highlighting different granularity changes while keeping the data length unchanged, and then performing normalization and embedding operations on each time series; for the k-th layer, decomposing the multi-resolution time series into a seasonal part S k and a trend part T k , projecting the features of the seasonal part using a multi-layer perceptron to obtain S kout ; for the trend part, first dividing the trend part into multiple patches T k p , and then dividing them into sub-patches grouping them by the same phase and treating each group as a token, and processing them through the self-attention mechanism to obtain the trend feature output T kout , and adding the seasonal and trend features to obtain the feature output X kout ; based on the obtained feature output X kout , using a corresponding predictor to make a prediction to obtain a prediction result, and then accumulating each prediction result to obtain the final future time series prediction value.

[0111] Please refer to Figure 5, the terminal device is a computer device. The computer device 60 in this embodiment includes: a processor 61, a memory 62, and a computer program 63 stored in the memory 62 and executable on the processor 61. When the computer program 63 is executed by the processor 61, it implements the time-series data prediction method of the time-variable dual-domain jump attention mechanism in the embodiment. To avoid repetition, it will not be elaborated here one by one. Alternatively, when the computer program 63 is executed by the processor 61, it implements the functions of each model / unit in the time-series data prediction system of the time-variable dual-domain jump attention mechanism in the embodiment. To avoid repetition, it will not be elaborated here one by one.

[0112] The computer device 60 can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The computer device 60 may include, but is not limited to, a processor 61 and a memory 62. Those skilled in the art can understand that Figure 5 merely examples of the computer device 60, which do not constitute a limitation on the computer device 60, may include more or fewer components than shown in the figure, or combine some components, or different components. For example, the computer device may also include input / output devices, network access devices, buses, etc.

[0113] The so-called processor 61 may be a central processing unit (CPU), or may also be other general-purpose processors, a graphics processing unit (GPU), a tensor processing unit (TPU), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0114] The memory 62 may be an internal storage unit of the computer device 60, such as the hard disk or memory of the computer device 60. The memory 62 may also be an external storage device of the computer device 60, such as a plug-in hard disk equipped on the computer device 60, a smart media card (SMC), a secure digital (SD) card, a flash card, etc.

[0115] Further, the memory 62 may also include both internal storage units of the computer device 60 and external storage devices. The memory 62 is used to store computer programs and other programs and data required by the computer device. The memory 62 may also be used to temporarily store data that has been output or is to be output.

[0116] Please refer to Figure 6 , the terminal device is the electronic device 600, and the electronic device 600 is presented in the form of a general-purpose computing device. The components of the electronic device may include but are not limited to: at least one processing unit 610, at least one storage unit 620, a bus 630 connecting different platform components (including the storage unit 620 and the processing unit 610), a display unit 640, etc.

[0117] Among them, the storage unit stores program codes, and the program codes can be executed by the processing unit 610, so that the processing unit 610 executes the steps according to various exemplary embodiments of the present invention described in the above method part of this specification. For example, the processing unit 610 can execute steps as shown in Figure 3 .

[0118] The storage unit 620 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 6201 and / or a cache storage unit 6202, and may further include a read-only storage unit (ROM) 6203.

[0119] The storage unit 620 may also include a program / utilities 6204 having a set (at least one) of program modules 6205. Such program modules 6205 include but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.

[0120] The bus 630 may represent one or more of several types of bus structures, including a memory bus or a memory controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the multiple bus structures.

[0121] The electronic device 600 can also communicate with one or more external devices 700 (such as a keyboard, a pointing device, a Bluetooth device, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 600, and / or communicate with any device (such as a router, a modem) that enables the electronic device 600 to communicate with one or more other computing devices. Such communication can be carried out through the input / output interface 650. Moreover, the electronic device 600 can also communicate with one or more networks (such as a local area network, a wide area network, and / or a public network, such as the Internet) through the network adapter 660. The network adapter 660 can communicate with other modules of the electronic device 600 through the bus 630. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 600, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage platforms, etc.

[0122] Embodiment 4

[0123] The present invention also provides a storage medium, specifically a computer-readable storage medium, which is a memory device in a terminal device for storing programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the terminal device and, of course, the extended storage medium supported by the terminal device. It can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. The computer-readable storage medium provides a storage space that stores the operating system of the terminal. And, one or more instructions suitable for being loaded and executed by a processor are stored in this storage space, and these instructions can be one or more computer programs (including program codes). It should be noted that more specific examples of the computer-readable storage medium here include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory, an optical fiber, a portable compact disk read-only memory, an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0124] The computer-readable storage medium also includes a data signal propagated in a baseband or as part of a carrier wave, which carries the readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable storage medium can also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium can be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber cable, radio frequency, etc., or any suitable combination of the above.

[0125] The program code for performing the operations of the present invention can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network or a wide area network, or can be connected to an external computing device (e.g., by using an Internet service provider to connect through the Internet).

[0126] One or more instructions stored in the computer-readable storage medium can be loaded and executed by a processor to implement the corresponding steps of the time series data prediction method related to the time-variable dual-domain jump attention mechanism in the above embodiments; one or more instructions in the computer-readable storage medium are loaded and executed by the processor to perform the following steps:

[0127] Receive historical time series observation data containing T time steps and N variables; process the historical time series observation data with different kernel sizes by iteratively applying 1-D average pooling operations to obtain a set of time series that highlight different granularity changes while keeping the data length unchanged, and then perform normalization and embedding operations on each time series; for the k-th layer, decompose the multi-resolution time series into a seasonal part S k and a trend part T k , project the features of the seasonal part using a multi-layer perceptron to obtain S kout ; for the trend part, first divide the trend part into multiple patches T k p , and then divide it into sub-patches After grouping by the same phase, each group is regarded as a token and processed by the self-attention mechanism to obtain the trend feature output T kout Then, the seasonal and trend features are added together to obtain the feature output X kout Based on the obtained feature output X kout A corresponding predictor is used for prediction to obtain the prediction results, and then the prediction results are accumulated to obtain the final predicted value of the future time series.

[0128] In each of the embodiments provided in the present application, the databases involved may include at least one of a relational database and a non-relational database. The non-relational database may include, but is not limited to, a distributed database based on blockchain, etc. In each of the embodiments provided in the present application, the processors involved may be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., without limitation.

[0129] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the present invention described and shown in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0130] The steps of the time series prediction method based on the time-variable double-domain jump attention mechanism of the present invention are as follows:

[0131] S1. Collect the original time series in multiple fields. Time series data is collected from various data sources, which may include but are not limited to sensor networks, database systems, network logs, etc. For example, in the transportation field, traffic flow, vehicle speed, vehicle type, etc. data are collected in real time through devices such as vehicle sensors and traffic cameras deployed on roads; in the energy field, energy consumption data is obtained from devices such as smart electricity meters and gas meters; in the meteorological field, meteorological elements data such as temperature, air pressure, humidity, and wind speed are collected using meteorological satellites and ground weather stations. Ensure that the collected data has sufficient time span and sample size to support subsequent analysis and prediction work.

[0132] Please refer to Figure 1 , in the curve graph of the time series, the horizontal axis is the time point, and the vertical axis represents the value of the variable.

[0133] S2. Conduct preliminary cleaning and verification on the collected data to remove obvious incorrect data and outliers. Adopt data smoothing techniques (such as moving average, median filtering, etc.) to process noisy data. For missing values, appropriate filling methods can be selected according to the time series characteristics and statistical laws of the data, such as linear interpolation, nearest neighbor filling, or model-based filling methods. For example, when dealing with missing values in meteorological data, if the missing values are in a relatively stable time period, linear interpolation method can be used for filling; if the data has obvious seasonality or periodicity, seasonal decomposition models can be used to fill the missing values to ensure the integrity and reliability of the data.

[0134] S3. Standardize the cleaned data according to the requirements of the MFR block. Select appropriate standardization methods, such as Z-score standardization or Min-Max standardization, according to the distribution characteristics of the data and the needs of subsequent analysis. When performing Z-score standardization, calculate the mean and standard deviation of the data, and then convert each data point to a value under the standard normal distribution; for Min-Max standardization, determine the minimum and maximum values of the data, and map the data to a specified interval (such as [0,1]) to make the data of different variables and different time steps comparable, facilitating subsequent calculations and analysis.

[0135] S4. Scientifically determine the kernel size p of the 1-D average pooling operation according to the time period characteristics of the data, the change frequency of the variables, and the required granularity level. k . In actual operation, preliminary analysis can be carried out on the data first, draw the autocorrelation function graph and power spectral density graph of the data, observe the periodic characteristics and frequency distribution of the data, and select the appropriate kernel size based on this. For example, for power consumption data with obvious daily and weekly cycles, if the change characteristics within the daily cycle are to be extracted, a smaller kernel size (such as 2 - 4 hours) can be selected for the first pooling operation; if the trend information of the weekly cycle is to be further obtained, the kernel size can be gradually increased (such as 12 - 24 hours) for subsequent pooling operations. During the pooling process, calculate strictly according to the formula X k = AvgPool(X k ', p k ), where For the data within each pooling window, precise sampling and average calculation are performed to ensure the accuracy of data processing. Meanwhile, when filling with the average value along the time direction after pooling, according to the time series characteristics and statistical laws of the data, a suitable average value calculation method is selected. If the data shows a linear trend, simple arithmetic mean can be used; if the data has seasonal fluctuations or weight differences, the weighted average method can be adopted, where the weights can be assigned according to the importance of the data or the seasonal intensity to ensure the rationality and effectiveness of the filled data, and finally a set of time series {X0, …, X k}.

[0136] Please refer to Figure 2 (a), Figure 2 (b), Figure 2 (c) and Figure 2 (d) for the results of the Fourier transform of sequences with different granularities. Sequences with different granularities have different numbers of prominent amplitude values, indicating that sequences with different granularities have different prediction capabilities.

[0137] S5. Use a pre-trained professional embedding model or an embedding layer customized according to the data characteristics to perform embedding operations on each time series. When selecting an embedding model, consider the data type, dimension, and semantic information. For time series data with complex semantic structures (such as time series events described in text), a natural language processing embedding model based on deep learning (such as Word2Vec, GloVe, etc.) can be selected for modification and application; for numerical time series data, a dedicated neural network embedding layer can be designed to convert it into a vector representation form suitable for subsequent calculations by learning the internal characteristics and distribution laws of the data When training the embedding layer, a large amount of historical data is used, and a suitable loss function (such as mean squared error loss, cross-entropy loss, etc.) and optimization algorithm (such as Adam, Adagrad, etc.) are adopted for training, continuously adjusting the parameters of the embedding layer to improve the optimality of the embedding effect, so that the embedded vectors can better capture the potential characteristics and semantic information of the data.

[0138] S6. For the multi-resolution time series of the k-th layer Use an advanced series decomposition algorithm to decompose it into a seasonal part S k and a trend part T k. In practical applications, an appropriate decomposition method can be selected according to the characteristics of the data and the prediction requirements. For economic data with obvious seasonality and trend (such as commodity sales data), a decomposition method based on moving average may be a better choice. By setting an appropriate moving average window size (such as determined according to the length of the seasonal cycle), the seasonal component and trend component in the data can be effectively separated; for complex multivariate time series data (such as meteorological data), singular spectrum analysis may have more advantages. Using singular spectrum analysis to perform eigenvalue decomposition on the covariance matrix of the data, the data is decomposed into subsequences of different frequencies, so as to accurately extract the seasonal and trend components S k and T k , ensuring the accuracy and effectiveness of the decomposition.

[0139] Please refer to Figure 3 the overall design diagram of the method. After obtaining the seasonal and trend components, according to their characteristics, different processing methods are carried out respectively.

[0140] S7. For the seasonal part S k , a multi-layer perceptron is used for feature projection to obtain S kout = MLP(S k ). When constructing the multi-layer perceptron, it is carefully adjusted according to the characteristic dimension, complexity and required output dimension of the seasonal data. First, through statistical analysis of the seasonal data, its characteristic dimension is determined. For example, for agricultural production data with multiple seasonal influencing factors, it may be necessary to consider characteristics in multiple dimensions such as temperature, precipitation, and sunlight in different seasons. Then, according to the complexity of the data and the expected output dimension, a suitable multi-layer perceptron structure is designed. If the data is relatively complex, it may be necessary to increase the number of hidden layers of the multi-layer perceptron and the number of neurons in each layer to enhance the expression ability of the model. For example, for meteorological data with complex seasonal patterns and a lot of noise, an MLP with 3 - 5 hidden layers, the number of neurons in each layer between 64 - 256, and using non-linear activation functions such as ReLU may need to be designed. When training the multi-layer perceptron, a large amount of historical seasonal data is used, and a suitable optimization algorithm (such as stochastic gradient descent, Adagrad, etc.) is used to continuously adjust the parameters of the MLP to minimize the prediction error, ensuring that while removing noise, the effective information related to the season is retained to the greatest extent.

[0141] S8. For the trend part T k , it is first divided into multiple patches T k p, the basis for patch division is the time period and internal change law of the data. In actual operation, the appropriate patch size can be determined by analyzing the autocorrelation function of the data and the slope change of the trend line. For example, when analyzing the time series of stock prices, if the stock price fluctuates frequently in the short term but shows a certain periodic trend in the long term, the patch size can be reasonably determined according to factors such as the market cycle (such as bull and bear cycles) and the company's financial report cycle. Further divide the patch into sub-patches When doing so, fully consider the local stability and correlation of the data. By calculating the local variance and covariance of the data, judge the local stability of the data, and divide the regions with similar stability into sub-patches to ensure that the sub-patches can reflect the local characteristics of the data. After grouping by the same phase, each group is regarded as a token and processed through the self-attention mechanism. In the calculation process of the self-attention mechanism, strictly follow the standard attention calculation framework. For the query q of each token i and the key k j , accurately calculate the pre-softmax score A i,j =(QK T / √d) i,j , where d is the feature dimension. After obtaining the attention weights through the softmax function, multiply them by the value V to integrate the information of the grouped sub-patches. To improve the calculation efficiency and reduce the consumption of computing resources, sparse attention techniques or approximate calculation methods can be used, but it is necessary to ensure that the ability to capture key information and the prediction accuracy are not significantly affected. For example, a sparse attention algorithm based on locality-sensitive hashing can be used to reduce unnecessary computational volume while maintaining the attention to important correlation information. Finally, obtain the trend feature output T kout , and add the seasonal and trend features to get X kout . During the addition process, it is necessary to ensure the precision and consistency of the data, and use high-precision numerical calculation methods (such as double-precision floating-point calculation) for the addition operation to avoid prediction deviations caused by calculation errors.

[0142] Please refer to Figure 4 for the schematic diagram of jumping attention. Divide the seasonal component into patches and sub-patches, and apply self-attention to the sub-patches corresponding to the phase.

[0143] S9. For the feature output X of each granularity kout , use the corresponding predictor Pred k for prediction. The predictor Pred kAdopt a multi-layer perceptron structure, and its design needs to be customized according to the characteristics of different granularity data and the requirements of prediction targets. For fine-grained time series data, since its changes are relatively fast and complex, the predictor may require a more complex structure to capture these changes. For example, increase the number of hidden layers (such as 4 - 6 layers) and the density of neurons (such as the number of neurons in each layer is between 128 - 512), and adopt more sensitive activation functions (such as LeakyReLU, ELU, etc.) to improve the model's ability to capture detailed information; while for coarse-grained data, the predictor can be relatively simplified, focusing on extracting long-term trend information, such as using 2 - 3 hidden layers, with the number of neurons in each layer between 32 - 128, and the activation function can choose the relatively simple ReLU function. When training the predictor, use a large amount of historical data and divide the data into training set, validation set, and test set. Adopt appropriate optimization algorithms (such as stochastic gradient descent, Adagrad, Adadelta, etc.) to continuously adjust the parameters of the predictor to minimize the prediction error. For example, when dealing with traffic flow prediction problems, use traffic flow data from different time periods and different weather conditions in the past few years to train the predictor. During the training process, by monitoring the loss function value on the validation set, adjust the learning rate and other hyperparameters to prevent overfitting, so that the predictor can learn the variation rules of traffic flow in various situations.

[0144] S10. Accumulate the prediction results of each granularity to obtain the final future time series prediction value Before accumulation, in order to further improve the prediction accuracy and stability, weights can be assigned to each prediction result The determination of weights can be based on the granularity importance of data, historical prediction errors, or other relevant factors, and can be dynamically adjusted through machine learning algorithms or statistical analysis methods. For example, a simple linear weighting method can be adopted to determine the weights according to the mean absolute error (MAE) or mean square error (MSE) of each granularity in past predictions. If the time series of a certain granularity has shown high accuracy in past predictions (i.e., small MAE or MSE), then its weight in the accumulation process can be appropriately increased; or a more complex machine learning algorithm, such as a decision tree-based model or a neural network model, can be adopted to dynamically allocate weights according to various characteristics of the data and prediction effects. So that the final prediction result can more reasonably integrate the information of each granularity and improve the overall prediction performance. In practical applications, continuously monitor and evaluate the accuracy of the prediction results, and according to new data and feedback information, timely adjust the weight allocation strategy and predictor parameters to adapt to the dynamic changes of the data and improve the reliability of the prediction.

[0145] In summary, the time series data prediction method and system of the present invention with a time-variable dual-domain jumping attention mechanism are different from the previous methods that separately mine time or channel information. This method applies a cross-attention mechanism in the time-channel joint space to achieve two-way interaction between the two domains. It includes three modules: a multi-granularity feature refinement (MFR) module, a dual-domain jump mixing (DHM) module, and a multi-layer prediction mixing (MPM) module. A large number of experiments conducted on public datasets show that this method is significantly superior to existing methods, providing a new solution for time prediction tasks.

[0146] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above-mentioned system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here.

[0147] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0148] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in the present invention can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0149] In the embodiments provided by the present invention, it should be understood that the disclosed device / terminal and method can be implemented in other ways. For example, the device / terminal embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.

[0150] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0151] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0152] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above method embodiments of the present invention, it can also be completed by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device that can carry the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0153] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses, and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a means for implementing the functions specified in one or more flows Figure 1 one or more flows and / or blocks Figure 1 or a means for implementing the functions specified in one or more blocks.

[0154] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction means that implements the functions specified in one or more flows Figure 1 one or more flows and / or blocks Figure 1 or a means for implementing the functions specified in one or more blocks.

[0155] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more flows Figure 1 one or more flows and / or blocks Figure 1 or a means for implementing the functions specified in one or more blocks.

[0156] The above content is only to illustrate the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. Any modification made on the basis of the technical solution according to the technical idea proposed by the present invention falls within the protection scope of the claims of the present invention.

Claims

1. A method for predicting time series data with a time-variable dual-domain jump attention mechanism, characterized in that, It includes the following steps: Receiving historical time series observation data containing T time steps and N variables; Processing the historical time series observation data with different kernel sizes by iteratively applying 1-D average pooling operations to obtain a set of time series highlighting different granularity changes while keeping the data length unchanged, and then performing normalization and embedding operations on each time series; For the k-th layer, the multi-resolution time series is decomposed into a seasonal part and a trend part, and the seasonal part is projected by a multi-layer perceptron to obtain S kout ; For the trend part, first divide the trend part into multiple patches Then divide it into sub-patches After grouping by the same phase, each group is regarded as a token and processed through the self-attention mechanism to obtain the trend feature output T kout , and add the seasonal and trend features to obtain the feature output X kout ; Based on the obtained feature output X kout , corresponding predictors are used for prediction to obtain prediction results, and then the prediction results are accumulated to obtain the final predicted value of the future time series.

2. The time series data prediction method of the time-variable dual-domain jump attention mechanism according to claim 1, wherein Normalization adopts Z-score normalization or Min-Max normalization.

3. The time series data prediction method of the time-variable dual-domain jump attention mechanism according to claim 1, characterized in that By averaging the X k Perform an average pooling operation. During the pooling process, sample and average the data, and fill it with the average value along the time direction after pooling.

4. The time series data prediction method of the time-variable dual-domain jump attention mechanism according to claim 3, wherein When filling, simple arithmetic average or weighted average is selected according to the time series characteristics and statistical laws of the data.

5. The time series data prediction method of the time-variable dual-domain jump attention mechanism according to claim 3, characterized in that Pooled time series X k is: X k ' = AvgPool(X k , p k ) Among them, X k is the input time series, AvgPool is average pooling, which performs a pooling operation on a specified area, and p k is the range of the specified area.

6. The time series data prediction method of the time-variable dual-domain jumping attention mechanism according to claim 1, characterized in that, The multi-layer perceptron has multiple hidden layers, and the number of neurons in each layer gradually decreases and uses a non-linear activation function.

7. The time series data prediction method of the time-variable dual-domain jump attention mechanism according to claim 1, wherein After grouping by the same phase, each group is regarded as a token for self-attention mechanism calculation. During the calculation, following the standard attention calculation framework, for the query qi and key k of each token j , accurately calculate the pre-softmax score A i,j =(QK T / √d) i,j , where d is the feature dimension; after obtaining the attention weights through the softmax function, multiply them by the value V to integrate the information of the grouped sub-patches.

8. The temporal data prediction method of the time-variable dual-domain jump attention mechanism according to claim 1, wherein Before accumulating the prediction results of each granularity, for each prediction result weight distribution is performed.

9. The time series data prediction method of the time-variable dual-domain jumping attention mechanism according to claim 1, wherein Predictor Pred k Adopts a multi-layer perceptron structure.

10. A time series data prediction system with a time-varying dual-domain jump attention mechanism, characterized in that, It includes: A data module that receives historical time series observation data containing T time steps and N variables; A preprocessing module that processes the historical time series observation data with different kernel sizes by iteratively applying 1-D average pooling operations to obtain a set of time series highlighting different granularity changes while keeping the data length unchanged, and then performs normalization and embedding operations on each time series; Decomposition module, for the k-th layer, decompose the multi-resolution time series into the seasonal part S k and the trend part T k , and use a multi-layer perceptron to perform feature projection on the seasonal part to obtain S kout ; For the trend part, first divide the trend part into multiple patches Then divide it into sub-patches After grouping by the same phase, each group is regarded as a token and processed through the self-attention mechanism to obtain the trend feature output T kout , and add the seasonal and trend features to obtain the feature output X kout ; Prediction module, which outputs X based on the obtained features kout , uses the corresponding predictor to make predictions to obtain prediction results, and then accumulates the prediction results to obtain the final predicted value of the future time series.