Long-time sequence prediction method and system for guiding state space attention based on meta-information
By combining meta-information and multi-scale temporal patterns, and utilizing a method that combines selective state-space attention with linear attention, accurate prediction of long-term series is achieved. This solves the problem of insufficient ability of existing methods to capture complex temporal dynamics and improves the accuracy and consistency of prediction.
Patent Information
- Application Number
- CN202511732586.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-24
- Publication Date
- 2026-02-13
AI Technical Summary
Existing long-term series forecasting methods struggle to accurately capture temporal dynamics when dealing with non-stationarity, multi-scale variations, and complex long-term correlations. Furthermore, they rely on insufficient metadata, leading to increased forecast uncertainty and noise interference, which hinders practical applications.
By combining meta-information description with multi-scale temporal patterns, semantically rich representations are generated. By combining selective state-space attention with linear attention, bidirectional temporal dependency modeling is achieved. Diverse temporal patterns are captured using block processing, and the final prediction results are output.
It improves the accuracy and consistency of long-term series forecasts, enhances the understanding and capture of complex temporal dynamics, and reduces forecast uncertainty and noise interference.
Smart Images

Figure QLYQS_8 
Figure QLYQS_17 
Figure QLYQS_18
Abstract
Description
Technical Field
[0001] This invention relates to the field of model development technology, specifically to a long-term series prediction method and system based on meta-information-guided state-space attention. Background Technology
[0002] Long-term series forecasting, as a fundamental task, typically refers to predicting multiple future time steps and has wide applications in climate modeling, energy management, and financial analysis. For example, in climate science, accurately predicting long-term sea surface temperature is crucial for forecasting events such as El Niño and La Niña, which have a significant impact on global weather patterns. However, real-world time series typically exhibit non-stationarity, multi-scale variations, and complex long-term correlations. Furthermore, as the forecast scope and time span expand, capturing the underlying temporal dynamics becomes increasingly challenging, severely impacting its practical application. Therefore, effective forecasting methods are needed to address this issue.
[0003] Deep learning methods possess an excellent ability to learn complex nonlinear correlations and underlying data distributions from large-scale data, making them suitable for time series forecasting. However, existing methods tend to emphasize abrupt changes and cannot well represent the temporal continuity of long series. Furthermore, reliance on simple linear mappings or pure convolutional structures often limits their representational power, resulting in insufficient ability to capture the inherently complex and diverse temporal dynamics of real-world time series data. Moreover, long-term forecasting of time series exacerbates distributional uncertainty and increases the likelihood of encountering distributional shifts or noise interference. In such cases, the model needs to rely on meta-information about the data to maintain temporal consistency and avoid error accumulation, but simply relying on the mean and standard deviation is still insufficient to fully describe the complex dynamics of real-world time series. Summary of the Invention
[0004] This invention is made to solve the above-mentioned problems, and aims to provide a long-term series prediction method and system based on meta-information-guided state-space attention.
[0005] This invention provides a long-term series prediction method based on meta-information-guided state-space attention, characterized by the following steps: S1, input... Rich metadata description Combined with multi-scale temporal patterns, it generates semantically rich representations. S2, through semantically rich representation A time-space flip is performed to achieve bidirectional time-dependent modeling. A block-based bidirectional fusion module is constructed, in which a selective state-space model is combined with linear attention to obtain State-Space Attention (SSA). Block processing is introduced in the feature dimension to output a fused representation. S3, will be a fusion representation The final prediction result is obtained by applying a linear mapping. : ,in, This represents the process of linear mapping.
[0006] The long-term series prediction method based on meta-information-guided state-space attention provided by this invention may also have the following features: S1 includes the following sub-steps: S1.1, Meta-information description construction: By extracting meta-information from the domain level, sample level, and task level, a domain-specific text output Prompt is generated, denoted as the meta-information description; S1.2, Meta-information text encoding: The meta-information description is text-encoded using a frozen large language model, followed by pooling operations to obtain a rich meta-information description. S1.3, Multi-scale temporal pattern construction: Define a set of patch block sizes Each of them Corresponding to specific patch block divisions, different patch block sizes provide different time resolutions for the input. A type of block partitioning with a size of and step length The patch block partitioning operation will divide the input into One patch block, then each patch block from Projected to get ,
[0007] ,
[0008] For each This process is applied independently, and the embeddings obtained at all scales are stitched together to form the final representation of the multi-scale temporal pattern:
[0009] ,
[0010] Describing rich metadata Combined with multi-scale temporal models, it enables rich metadata description. Inserting into the beginning and end of the embedded sequence yields the final semantically rich representation. This enhances the semantic information of the original sequence representation.
[0011] The long-term series prediction method based on meta-information-guided state-space attention provided by this invention may also have the following features: the sample-level meta-information includes basic statistical data, distribution statistical features, and stationarity features; the task-level meta-information includes task-specific objectives and processing instructions; and the domain-level meta-information includes task domain description and dataset features.
[0012] The long-time series prediction method based on meta-information-guided state-space attention provided by this invention may also have the following feature: the construction process of the meta-information description at the sample level is as follows: for each sequence Basic statistical data includes calculating the minimum, maximum, and trend of the sequence; distribution statistics include the skewness of the sequence; stationarity features include analysis using the extended Dick-Fuller test, standard deviation, and structural mutation analysis.
[0013] The long-term sequence prediction method based on meta-information-guided state-space attention provided by this invention may also have the following features: the meta-information description construction process at the task level is as follows: clarifying how to extract the time dependency of the sequence, that is, the input sequence is processed twice, once by flipping the time dimension and once without flipping, thereby realizing the extraction of bidirectional time dependency; the meta-information description construction process at the domain level is as follows: clarifying the source and domain of the dataset, and using information from past time steps to predict information from future time steps.
[0014] The long-time series prediction method based on meta-information-guided state-space attention provided by this invention can also have the following feature: rich meta-information description... The specific calculation method is as follows: The frozen large language model is used as a text encoder to encode and embed the constructed meta-information text description into a hidden representation, and average pooling is used to obtain the final rich meta-information description. :
[0015] ,
[0016] ,
[0017] in This represents the process of text encoding using a pre-trained, frozen large language model. This indicates the average pooling operation applied to the encoded representation; It is the embedding dimension of the text encoder.
[0018] The long-term series prediction method based on meta-information-guided state-space attention provided by this invention may also have the following features: S2 includes the following sub-steps: S2.1, constructing a mechanism for selective state-space attention (SSA): combining a selective state-space model with linear attention through formula derivation to obtain state-space attention (SSA), achieving dynamic information selection and integration across locations; S2.2, constructing a segmented bidirectional fusion module: integrating semantically rich representations... The representations are flipped along the time dimension to obtain forward and reverse representations. After being divided into blocks along the feature dimension, they are input into the state space attention SSA. Finally, the outputs are concatenated into a fused representation. .
[0019] The long-term series prediction method based on meta-information-guided state-space attention provided by this invention may also have the following feature: wherein, in S2.1, the state-space attention (SSA) is calculated as follows:
[0020] ,
[0021] ,
[0022] , , , as well as , In time step The output, Soon As an attention mechanism, information retrieval is performed using queries and keys, where... ; ;and and This represents the m-th row of the corresponding matrix, i.e., a single query, key, and value token. yes Activation function.
[0023] The long-time series prediction method based on meta-information-guided state-space attention provided by this invention may also have the following feature: wherein S2.2 includes the following sub-steps:
[0024] Meta information description Semantically rich representation obtained by fusing with multi-scale temporal patterns To obtain by flipping in the time dimension and For bidirectional time dependency modeling, it introduces block segmentation along the feature dimension to extract different time patterns. Divided into and , Divided into and :
[0025] ,
[0026] ,
[0027] .
[0028] This invention also provides a long-term series prediction system based on meta-information-guided state-space attention, characterized by including: a meta-information-guided embedding module, which takes the input... Rich metadata description Combined with multi-scale temporal patterns, it generates semantically rich representations. The fusion representation output module, through semantically rich representations... A time-space flip is performed to achieve bidirectional time-dependent modeling. A block-based bidirectional fusion module is constructed, in which a selective state-space model is combined with linear attention to obtain State-Space Attention (SSA). Block processing is introduced in the feature dimension to output a fused representation. The prediction module will fuse the representations. The final prediction result is obtained by applying a linear mapping. :
[0029] ,
[0030] in, This represents the process of linear mapping.
[0031] The role and effect of invention
[0032] This invention addresses the limitations of dependency extraction capabilities and insufficient utilization of meta-information over long time series. It provides a long-term series prediction method and system based on meta-information-guided state-space attention. The method integrates expressive statistics and structural descriptions with multi-scale temporal patterns through a meta-information-guided embedding module to generate semantically rich representations. Then, selective state-space attention (SSA) connects the selected state-space equation and the linear attention mechanism, further introducing block-level partitioning along the feature dimension to capture diverse and reliable temporal patterns. Furthermore, sequence flipping enables bidirectional modeling of the time dependency model, ensuring a more comprehensive temporal understanding and ultimately achieving accurate long-term time series prediction. Detailed Implementation
[0033] This invention discloses a long-term series prediction method based on meta-information-guided state-space attention, comprising the following steps:
[0035] S1, Construct the metadata guidance embedding module, and input Rich metadata description Combined with multi-scale temporal patterns, it generates semantically rich representations. .
[0036] S1 includes the following sub-steps:
[0037] S1.1 Meta-information description construction: By extracting meta-information from the domain level, sample level and task level, a domain-specific text output Prompt is generated, which is denoted as meta-information description.
[0038] The metadata at the sample level includes: basic statistical data, distribution statistical characteristics, and stationarity characteristics; the construction process of the metadata description at the sample level is as follows:
[0039] For each sequence Basic statistical data includes calculating the minimum, maximum, and trend of the sequence; distribution statistics include the skewness of the sequence; stationarity features include analysis using the extended Dick-Fuller test, standard deviation, and structural mutation analysis.
[0040] Task-level metadata includes: task-specific objectives and processing instructions.
[0041] The metadata description construction process in the task level is as follows: it clarifies how to extract the time dependency of the sequence, that is, the input sequence is processed twice, once by flipping the time dimension and once without flipping, thereby realizing the extraction of bidirectional time dependency.
[0042] Domain-level meta-information includes: task domain description and dataset features.
[0043] The process of constructing metadata descriptions at the domain level is as follows: clarify the source and domain of the dataset, and use information from past time steps to predict information from future time steps.
[0044] S1.2, Meta-information Text Encoding: The meta-information description is text-encoded using a frozen large language model, followed by pooling to obtain a rich meta-information description. .
[0045] A frozen large language model (such as GPT or Deepseek) is used as a text encoder to encode and embed the constructed meta-information text description into a hidden representation. Average pooling is then used to obtain the final rich meta-information description. :
[0046] .
[0047] .
[0048] in This represents the process of text encoding using a pre-trained, frozen large language model. This indicates the average pooling operation applied to the encoded representation; It is the embedding dimension of the text encoder.
[0049] S1.3, Multi-scale temporal pattern construction: Define a set of patch block sizes Each of them Corresponding to specific patch block divisions, different patch block sizes provide different time resolutions for the input. A type of block partitioning with a size of and step length The patch block partitioning operation will divide the input into One patch block, then each patch block from Projected to get .
[0050] .
[0051] For each This process is applied independently, and the embeddings obtained at all scales are stitched together to form the final representation of the multi-scale temporal pattern:
[0052] .
[0053] Describing rich metadata Combined with multi-scale temporal models, it enables rich metadata description. Inserting into the beginning and end of the embedded sequence yields the final semantically rich representation. This enhances the semantic information of the original sequence representation.
[0054] S2, through semantically rich representation A time-space flip is performed to achieve bidirectional time-dependent modeling. A block-based bidirectional fusion module is constructed, in which a selective state-space model is combined with linear attention to obtain State-Space Attention (SSA). Block processing is introduced in the feature dimension to output a fused representation. .
[0055] S2 includes the following sub-steps:
[0056] S2.1, Mechanism for constructing Selective State-Space Attention (SSA): By deriving formulas, the selective state-space model is combined with linear attention to obtain State-Space Attention (SSA), enabling dynamic information selection and integration across locations.
[0057] Specifically:
[0058] In S2.1, the connection between the selective state-space model and linear attention is established through formula derivation. First, we focus on the core operations of the selective state-space model, considering the input... Specifically as follows:
[0059] In the selective state-space model Matrices are usually diagonal matrices, so they can be equivalently replaced by matrix-vector multiplication: ,in A matrix consisting of diagonal elements and It is an element-wise product; in addition, because and Therefore, the terms can be rearranged to derive the equivalent form. Thus, an equivalent formula can be obtained:
[0060] ,
[0061] ,
[0062] Furthermore, the original selective state-space model only applies to scalars. Perform operations on the above to expand to the sequence. Therefore, its core operations are applied independently to each channel, resulting in the following operations:
[0063] .
[0064] .
[0065] in , , as well as .
[0066] Then focus on the core operation of linear attention, considering the input... Specifically as follows:
[0067] .
[0068] .
[0069] .
[0070] in In time step The output, and It is an auxiliary variable used to calculate the attention output, and , , , It is a kernel function. as well as These represent the query, key, and value matrix, respectively. For and Represents the first element of the corresponding matrix Rows, namely each query, key, and value token.
[0071] By restating the formula for the core operation in linear attention:
[0072] .
[0073] .
[0074] This invention discloses the close structural relationship between the formula for linear attention and the selective state-space model, the selective state-space model through... A forget gate is introduced to filter the hidden state. The key difference between the selective state-space equation and the linear attention equation lies in the output. Generation: The selective state-space equation is obtained by linear projection through Cm to obtain the output. The linear attention equation calculates the output as a weighted sum across all positions. To enhance expressiveness, this invention incorporates the global aggregation capability of attention styles into the selective state-space model, resulting in State-Space Attention (SSA), which enables dynamic information selection and integration across locations.
[0075] .
[0076] .
[0077] Soon As an attention mechanism, information retrieval is performed using queries and keys, where... ; ;and and This represents the m-th row of the corresponding matrix, i.e., a single query, key, and value token. Here, yes Activation function.
[0078] S2.2, Constructing a segmented bidirectional fusion module: integrating semantically rich representations The representations are flipped along the time dimension to obtain forward and reverse representations. After being divided into blocks along the feature dimension, they are input into the state space attention SSA. Finally, the outputs are concatenated into a fused representation. .
[0079] Specifically, the metadata description is first... Semantically rich representation obtained by fusing with multi-scale temporal patterns To obtain by flipping in the time dimension and For bidirectional time dependency modeling, to improve the ability of SSA to capture long-term independence, block segmentation along the feature dimension is introduced to extract different time patterns. Specifically... Divided into and , Divided into and :
[0080] ,
[0081] ,
[0082] .
[0083] S3 will merge representations The final prediction result is obtained by applying a linear mapping. :
[0084]
[0085] in, This represents the process of linear mapping.
[0086] Those skilled in the art should understand that this invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to this invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.
Claims
1. A long-term series prediction method based on meta-information-guided state-space attention, characterized in that, Includes the following steps: S1, will input Rich metadata description Combined with multi-scale temporal patterns, it generates semantically rich representations. ; S2, by using the semantically rich representation A time-space flip is performed to achieve bidirectional time-dependent modeling. A block-based bidirectional fusion module is constructed. In this module, a selective state-space model is combined with linear attention to obtain State-Space Attention (SSA). Block processing is introduced in the feature dimension to output a fused representation. ; S3, the fused representation The final prediction result is obtained by applying a linear mapping. : , in, This represents the process of linear mapping.
2. The long-term series prediction method based on meta-information-guided state-space attention according to claim 1, characterized in that: in, S1 includes the following sub-steps: S1.1 Meta-information description construction: By extracting meta-information from the domain level, sample level and task level, a domain-specific text output Prompt is generated, which is denoted as meta-information description; S1.2, Meta-information Text Encoding: The meta-information description is text-encoded using a frozen large language model, followed by pooling operations to obtain a rich meta-information description. ; S1.3, Multi-scale temporal pattern construction: Define a set of patch block sizes Each of them Corresponding to specific patch block divisions, different patch block sizes provide different time resolutions for the input. A type of block partitioning with a size of and step length The patch block partitioning operation will divide the input into One patch block, then each patch block from Projected to get , , For each This process is applied independently, and the embeddings obtained at all scales are stitched together to form the final representation of the multi-scale temporal pattern: , Describing the rich metadata Combined with the aforementioned multi-scale temporal patterns, this enhances the rich metadata description. Inserting into the beginning and end of the embedded sequence yields the final semantically rich representation. This enhances the semantic information of the original sequence representation.
3. The long-term series prediction method based on meta-information-guided state-space attention as described in claim 2, Its features are: in, The metadata at the sample level includes: basic statistical data, distribution statistical characteristics, and stationarity characteristics; The meta-information at the task level includes: task-specific objectives and processing instructions; The domain-level metadata includes: task domain description and dataset features.
4. The long-term series prediction method based on meta-information-guided state-space attention as described in claim 3, Its features are: The process of constructing the meta-information description in the sample hierarchy is as follows: For each sequence Basic statistical data includes calculating the minimum, maximum, and trend of the sequence; distribution statistical characteristics include the skewness of the sequence. Stationarity characteristics include analysis using the extended Dick-Fuller test, standard deviation, and structural mutation analysis.
5. The long-term series prediction method based on meta-information-guided state-space attention according to claim 3, characterized in that: in, The construction process of the meta-information description in the task hierarchy is as follows: it clarifies how to extract the time dependency of the sequence, that is, the input sequence is processed twice, once by flipping the time dimension and once without flipping, thereby realizing the extraction of bidirectional time dependency; The construction process of the meta-information description in the domain layer is as follows: clarify the source and domain of the dataset, and use information from past time steps to predict information from future time steps.
6. The long-term series prediction method based on meta-information-guided state-space attention according to claim 1, characterized in that: in, The rich metadata description The calculation method is as follows: The frozen large language model is used as a text encoder to encode and embed the constructed meta-information text description into a hidden representation. Average pooling is then used to obtain the final rich meta-information description. : , , in This represents the process of text encoding using a pre-trained, frozen large language model. This indicates the average pooling operation applied to the encoded representation; It is the embedding dimension of the text encoder.
7. The long-term series prediction method based on meta-information-guided state-space attention according to claim 1, characterized in that: in, S2 includes the following sub-steps: S2.1, The mechanism for constructing the selective state space attention SSA: By deriving the formula, the selective state space model is combined with linear attention to obtain the state space attention SSA, thereby realizing dynamic information selection and integration across positions; S2.2, Construct the segmented bidirectional fusion module: This module integrates the semantically rich representation... The representations are flipped along the time dimension to obtain forward and reverse representations, and then divided into blocks along the feature dimension and input into the state space attention SSA. Finally, the outputs are concatenated into a fused representation. .
8. The long-term series prediction method based on meta-information-guided state-space attention according to claim 7, characterized in that: in, In step S2.1, the state-space attention (SSA) is calculated as follows: , , , , , as well as , In time step The output, Soon As an attention mechanism, information retrieval is performed using queries and keys, where... ; ;and and This represents the m-th row of the corresponding matrix, i.e., a single query, key, and value token. yes Activation function.
9. The long-term series prediction method based on meta-information-guided state-space attention according to claim 7, characterized in that: in, S2.2 includes the following sub-steps: Describe the metadata The semantically rich representation obtained by fusing with the multi-scale temporal pattern To obtain by flipping in the time dimension and For bidirectional time dependency modeling, a block-based approach is introduced along the feature dimension to extract different time patterns. Divided into and The Divided into and : , , 。 10. A long-term series prediction system based on meta-information-guided state-space attention, obtained according to any one of claims 1 to 9, characterized in that, include: Meta-information guides the embedded module, which will input Rich metadata description Combined with multi-scale temporal patterns, it generates semantically rich representations. ; The fusion representation output module, through the semantically rich representation A time-space flip is performed to achieve bidirectional time-dependent modeling. A block-based bidirectional fusion module is constructed. In this module, a selective state-space model is combined with linear attention to obtain State-Space Attention (SSA). Block processing is introduced in the feature dimension to output a fused representation. ; The prediction module will use the fused representation The final prediction result is obtained by applying a linear mapping. : , in, This represents the process of linear mapping.