Endogenous-exogenous cooperation-based time sequence prediction method and system
By employing an endogenous-exogenous synergistic time series forecasting method, and utilizing multi-level temporal attention and dynamic interactive attention mechanisms, the dynamic dependency problem in the interaction between endogenous and exogenous variables is solved, thereby improving the accuracy and robustness of multivariate time series forecasting.
Patent Information
- Application Number
- CN202511489547.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-17
- Publication Date
- 2026-01-13
AI Technical Summary
Existing multivariate time series forecasting methods neglect the synergistic effect between endogenous and exogenous variables when dealing with their interaction, making it difficult to fully capture dynamic dependencies and resulting in limited forecasting performance, especially in complex scenarios where accuracy and robustness are insufficient.
We adopt a time series prediction method based on endogenous-exogenous synergy. By integrating endogenous and exogenous inputs into an embedding vector through an endogenous time embedding module and an exogenous structural embedding module, we optimize the representation using a multi-level time attention mechanism and a dynamic interactive attention mechanism to achieve deep coupling modeling of endogenous and exogenous variables and improve prediction accuracy.
By employing a hierarchical system and a multi-level time attention mechanism, the dynamic evolution of endogenous and exogenous variables is accurately characterized, improving the accuracy and robustness of multivariate time series prediction. This solves the problem of matching the heterogeneity of exogenous variables with the dynamic characteristics of endogenous variables, and enhances prediction performance in complex scenarios.
Smart Images

Figure CN121328836A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of multivariate time series prediction, and particularly relates to a time series prediction method and system based on endogenous-external collaborative. BACKGROUND
[0002] In many fields such as finance, transportation, and power, multivariate time series prediction is a key task, and plays a crucial role in decision support, resource allocation, etc. Traditional multivariate time series prediction methods, such as autoregressive models based only on historical endogenous variables, often lack effective consideration of external influencing factors, making it difficult to cope with complex and dynamic real-world scenarios.
[0003] Inspired by the superiority of deep learning technology, various deep learning-based models have been used for multivariate time series prediction tasks, such as some Transformer-based models that capture temporal dependencies in sequences and achieve certain results in prediction accuracy. However, existing deep learning-based methods often use simple concatenation or separate encoding strategies to handle the interaction between endogenous and exogenous variables, ignoring the synergistic effect between endogenous and exogenous variables, and failing to fully capture the dynamic dependencies necessary for modeling complex temporal patterns, resulting in limited prediction performance.
[0004] To solve this problem, some methods attempt to introduce exogenous variables for fusion, but exogenous variables have heterogeneous and irregular temporal characteristics. In practical applications, these variables are often affected by non-standardized observation protocols, fragmented data integrity, and synchronization differences, making it difficult to match the dynamic characteristics of endogenous variables, further challenging multivariate time series prediction.
[0005] Although some Transformer-based models have shown potential in multivariate time series prediction, they can capture long-range dependencies and dynamic interactions, but existing methods usually use direct concatenation of exogenous and endogenous variables as input features, or first process external variables independently and then fuse them at a specific level. Due to non-adaptive processing of exogenous variables, lack of intuitive variable interaction modeling, and insufficient use of exogenous information, performance is often limited. SUMMARY
[0006] To address the above deficiencies in the prior art, the present application provides a time series prediction method and system based on endogenous-external collaborative, which solves the problem of the prior art that cannot accurately depict the synergistic relationship between endogenous and exogenous variables that evolves dynamically over time and has limited prediction accuracy.
[0007] In order to achieve the above-mentioned purposes, the technical scheme adopted by the present application is as follows: a time series prediction method based on endogenous-exogenous cooperation, comprising: obtaining endogenous input and exogenous input of a prediction target; integrating the endogenous input into an endogenous embedding vector; integrating the exogenous input into an exogenous embedding vector; splicing the endogenous embedding vector and the exogenous embedding vector to obtain joint features; optimizing endogenous representation based on the joint features through a multi-level time attention mechanism; fusing the endogenous embedding and the exogenous embedding based on the optimized endogenous representation through a dynamic interactive attention mechanism to obtain cross-modal interactive features; fusing the optimized endogenous representation and the cross-modal interactive features to obtain intermediate dynamic representation; performing time series prediction based on the intermediate dynamic representation; wherein the endogenous input is a historical observation sequence of the prediction target, which is: the trading amount, the trading volume, the highest price, the lowest price and the opening price of a stock market at multiple time steps; or the traffic flow, the speed and the traffic density of a road section at multiple time steps; or the power consumption, the electricity price, the power generation and the power flow data of a region at multiple time steps; the exogenous input includes environmental driving factors and policy control parameters.
[0008] Secondly, the present application also provides a system for implementing the time series prediction method based on endogenous-exogenous cooperation, comprising: an endogenous time embedding module for integrating the endogenous input into an endogenous embedding vector; an exogenous structure embedding module for integrating the exogenous input into an exogenous embedding vector; a multi-level time attention module for splicing the endogenous embedding vector and the exogenous embedding vector, and optimizing the endogenous representation based on the joint features through a multi-level time attention mechanism; a dynamic attention module for fusing the endogenous embedding and the exogenous embedding based on the optimized endogenous representation through a dynamic interactive attention mechanism to obtain cross-modal interactive features; a prediction module for fusing the optimized endogenous representation and the cross-modal interactive features to obtain intermediate dynamic representation, and performing time series prediction based on the intermediate dynamic representation.
[0009] The present application has the following beneficial effects: 1. This invention constructs a hierarchical system to achieve deep coupling modeling of the temporal characteristics of endogenous variables and the structural influence of exogenous variables, solving the problem of difficulty in matching the heterogeneity and irregular temporal characteristics of exogenous variables with the dynamic characteristics of endogenous variables, and improving the accuracy and robustness of multivariate time series prediction in complex scenarios.
[0010] 2. By adopting a collaborative technical approach of "local dependency capture - global trend modeling", the dynamic evolution law is accurately characterized through fine-grained time pattern analysis; for exogenous variables, a structured embedding mechanism is designed to achieve unified representation and dimensional alignment of multi-source heterogeneous features, thereby constructing a joint representation space and dynamic interaction channel for internal and external variables.
[0011] 3. By designing a multi-level temporal attention mechanism, the accuracy of temporal dimension modeling is improved by integrating local short-term associations and global long-term dependencies of endogenous variables through adaptive weight allocation. Furthermore, by designing a dynamic interactive attention mechanism, effective information is filtered based on the strength of cross-variable associations, selectively strengthening key associations between endogenous and exogenous variables while suppressing redundant interference, thereby ensuring efficient fusion and semantic integrity of cross-modal information. Attached Figure Description
[0012] Figure 1 A flowchart of a time series prediction method based on endogenous-exogenous synergy is provided for an embodiment; Figure 2 The following is a structural diagram of a time series prediction system based on endogenous-exogenous synergy, provided as an example. Detailed Implementation
[0013] The specific embodiments of the present invention are described below to enable those skilled in the art to understand the present invention. However, it should be understood that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the present invention as defined and determined by the appended claims. All inventions utilizing the concept of the present invention are protected.
[0014] like Figure 1 As shown, in one embodiment of the present invention, a time series prediction method based on endogenous-exogenous synergy includes the following steps: S1. Obtain the endogenous and exogenous inputs of the prediction target.
[0015] The endogenous input is the historical observation sequence of the prediction target. The prediction target includes, but is not limited to, the closing price, price direction, and volatility of the financial sector over the next N days. In this case, the historical observation sequence can be the trading amount, trading volume, highest price, lowest price, and opening price of a stock market at multiple time steps. The prediction target can also be the hourly traffic flow of a certain road segment over the next 24 hours. In this case, the historical observation sequence can be the traffic flow, speed, and traffic density of that road segment at multiple time steps. The prediction target can also be the total power generation of wind farms in the next few days, the power flow of key transmission lines in the next few hours to days, etc. In this case, the historical observation sequence can include the electricity consumption, electricity price, power generation, and power flow data of a certain region at multiple time steps. The prediction target is determined based on at least one time series data source from power generation, transmission, distribution, electricity consumption, or market transactions. In addition to the above prediction targets, time series data also include market and trading behavior indicators, user-side and integrated energy system indicators, and cybersecurity and abnormal events.
[0016] Exogenous inputs include environmental driving factors and policy control parameters. These factors vary across different sectors. For example, in the power sector, environmental driving factors include, but are not limited to, temperature, humidity, wind speed, wind direction, precipitation, and air quality; policy control parameters include time-of-use pricing, tiered pricing, renewable energy quotas, and energy consumption control policies. In the transportation sector, environmental driving factors include, but are not limited to, temperature, wind speed, visibility, and road conditions; policy control parameters include public transportation scheduling plans, vehicle restrictions, traffic control, and holidays. In the financial sector, environmental driving factors include GDP growth rate, inflation rate, employment data, and commodity prices; policy control parameters include interest rates, tax policies, subsidy policies, and government spending.
[0017] The acquired endogenous and exogenous inputs are used as follows: Figure 2 The time series forecasting system shown is based on endogenous-exogenous synergy.
[0018] S2. The intrinsic input is integrated into an intrinsic embedding vector through the intrinsic temporal embedding module.
[0019] The specific method is as follows: Define the time difference between any two time steps in an endogenous input sequence as:
[0020] in, Indicates the time difference. Indicates the first Each time step Indicates the first One time step; To effectively characterize the impact of different time intervals on feature representation, the time difference is mapped to the relative position encoding space, and its expression is:
[0021] in, Represents a relative position encoding vector. Indicates the time difference The The sinusoidal component is obtained by mapping the frequency components. d The superscript T indicates the embedding dimension, and the superscript T indicates the transpose of the matrix. Each sine component By a pair of trainable coefficients and Parameterization:
[0022] In the formula, and For learnable weights, Here, 'd' represents the preset frequency coefficient and the embedding dimension. (Index) The value of ranges from 1 to d / 2, which allows the model to adaptively capture temporal relationships at different scales. Here, Indicates the time step difference value The Mapping of frequency components. Specifically, first... With preset frequency coefficient Multiply to obtain the phase, then map it to the interval [-1,1] using sine and cosine functions respectively, and finally use a pair of learnable coefficients. and The outputs of these two basis functions are linearly weighted. Therefore, In essence
[0023] Used to capture In the Time relationships at various frequency scales.
[0024] To capture the inherent periodicity of time series and effectively encode recurring seasonal patterns, a periodic code is defined for each timestamp, expressed as:
[0025] in, Represents a periodic encoded vector. The number of periodically coded components. n For the index of periodically coded components, For trainable projection matrices, The default fundamental frequency is defined by sin, which represents the sine function, and cos, which represents the cosine function. This indicates a radius of 180 degrees; By fusing the relative position encoding vector and the periodic encoding vector, a time-aware encoding vector is obtained, the expression of which is:
[0026] in, Represents the time-aware encoding vector; The intrinsic input is integrated into an intrinsic embedding vector based on the time-aware coding vector, and its expression is as follows:
[0027] in, Represents an endogenous embedding vector. Representation layer normalization, This indicates a temporal convolution operation. This indicates endogenous input.
[0028] The above operations enhance the ability to extract local temporal features while ensuring alignment with the dimensions of endogenous variables.
[0029] S3. The exogenous input is integrated into an exogenous embedding vector through the exogenous structure embedding module.
[0030] The specific method is as follows: Define the exogenous input variable at time t as Its expression is:
[0031] M Indicates the number of exogenous variables. Indicates the first M One exogenous variable; The continuous exogenous input variable at time t is smoothed using Gaussian projection to obtain the smoothed features, the expression of which is:
[0032] in, Represents the smoothed features. K Indicates the number of Gaussian functions. k Indicates the Gaussian function index. Indicates the first k The relevant weights corresponding to the Gaussian functions, where exp represents an exponential function with the natural constant e as the base. Indicates the first i One exogenous variable, For the first k The center of the Gaussian function, For the first k The standard deviation corresponding to the center of each Gaussian function Denotes the Euclidean norm; Mapping the discrete exogenous input variable at time t to a low-dimensional space of fixed dimensions is expressed as follows:
[0033] in, Represents a low-dimensional vector. For the first i Learnable embedding matrix of discrete variables; The exogenous variables are integrated into an individual embedding vector, the expression of which is:
[0034] in, Indicates the first i Individual embedding vectors of exogenous variables; and These are all indicator functions, used to determine... Does it belong to a continuous set? or discrete set ,like If the element belongs to the corresponding set, the indicator function takes a value of 1; otherwise, it takes a value of 0. By concatenating the individual embedding vectors of all exogenous variables, we obtain the exogenous embedding vector, which is expressed as:
[0035] in, Represents an exogenous embedding vector. Indicates the first M Individual embedding vectors of exogenous variables.
[0036] S4. The endogenous embedding vector and the exogenous embedding vector are concatenated through a multi-level temporal attention module to obtain joint features.
[0037] To establish the interaction between endogenous and exogenous variables, a learnable global token is introduced. Serving as a cross-modal information hub, it is used to capture global statistical information of the entire sequence and integrate it with endogenous embeddings. and exogenous embedding To concatenate them, the expression is:
[0038] in, Indicates joint features, Indicates splicing; A learnable global token is introduced to adapt to the time dimension through a broadcast mechanism. This is done to establish the interaction between endogenous and exogenous variables. It serves as a cross-modal information hub, used to capture global statistical information for the entire sequence.
[0039] S5. Optimize endogenous representations based on joint features through a multi-level temporal attention mechanism.
[0040] Because time series data exhibits local smoothing characteristics, data within shorter time steps often show stronger correlations. Therefore, this invention designs a multi-level temporal attention mechanism based on a joint representation space to capture short-term local dependencies and long-term global trends in endogenous variables.
[0041] Specifically: Local features are generated based on the intrinsic embedding components in the joint features using a learnable projection matrix, including: Time step Corresponding local query features Its expression is: ; Time step neighborhood Corresponding local bond features Its expression is: ; Time step neighborhood Corresponding local value features Its expression is: ; in, This represents the endogenous embedding component in the joint features. Endogenous embedding vector In the neighborhood The feature matrix within; , , All are learnable projection matrices; Based on the generated local features, a local attention mechanism is used to capture local dependencies. The local attention is calculated, and its expression is as follows:
[0042] in, This represents the local attention output at time step t, and softmax represents the softmax activation function. To hide spatial dimensions; A global anchor vector is introduced and interacts with the endogenous embedding components in the joint features through a cross-channel gating mechanism. Its expression is as follows:
[0043]
[0044]
[0045] in, This represents the global anchor vector, used to capture long-range dependencies across windows; This represents the global query vector, used to extract key information from the global context; The global key vector is obtained by projecting onto the intrinsic input and filtering it using ReLU and TopK, and is used to represent the key features of global matching; The global value vector is obtained by interacting with the projection matrix and then GELU activation, and is used to carry the representation related to global attention; , , All are projection matrices used to project input features onto the query (Q), key (K), and value (V) space; TopK indicates that the first K time steps are retained, ReLU indicates the ReLU activation function, and GELU indicates the GELU activation function.
[0046] The global attention is calculated using the following expression:
[0047] in, This represents the output of global attention. Indicates the first i One global key vector; Indicates the first i A global value vector; I Indicates the total number of global key-value pairs; This represents the softmax function; By integrating global attention, local attention, and endogenous embedding vectors through residual connections and layer normalization, an optimized endogenous representation is obtained, expressed as:
[0048] in, This represents the optimized endogenous characterization. Indicates the gating coefficient. It represents the Hadamardi (or Hadama) stack.
[0049] Unlike the standard global computation in Transformers (where long-range dependencies may be masked by local noise), the method proposed in this invention improves noise robustness while reducing computational complexity.
[0050] S6. Based on the optimized intrinsic representation, the dynamic attention module fuses the intrinsic and extrinsic embeddings through the dynamic interactive attention mechanism to obtain cross-modal interaction features.
[0051] The specific method is as follows: Considering the differences in the distribution of endogenous and exogenous variables in the semantic space, as well as the cross-modal dependency with exogenous variables, this invention uses endogenous features optimized by multi-level temporal attention. As the core query source, a context-aware query representation is obtained through multi-head projection, and its expression is:
[0052] in, This represents the query representation. Indicates multi-head projection operation; By compressing the exogenous embedding vector using a low-rank bottleneck, key-value pairs are generated, the expression of which is:
[0053] in, Represents the key vector. Represents a value vector. , Both represent weight matrices. , Both represent low-rank projection matrices; To mitigate the impact of noise and enhance the relevance of interactions between variables, this paper designs a method to calculate the cross-variable association score matrix, focusing on retaining the most meaningful associations while suppressing irrelevant or noisy connections. The cross-variable score matrix is calculated as follows:
[0054] in, Represents the cross-variable score matrix; Cross-modal interaction features are obtained by adaptively transforming the cross-variable score matrix and value vector through linear projection. The expression for this feature is as follows:
[0055] in, Indicates cross-modal interaction features. Represents the linear projection matrix. This represents a sparse mask.
[0056] S7. The optimized endogenous representation and cross-modal interaction features are fused through the prediction module to obtain the intermediate dynamic representation.
[0057] The specific method is as follows: The optimized intrinsic representation is subjected to average pooling and an activation function is applied to obtain a learnable gating vector, the expression of which is:
[0058] in, Represents a learnable gated vector. This represents the sigmoid activation function. Indicates average pooling. Represents the weight matrix; By scaling cross-modal interaction features using learnable gating vectors and fusing them with optimized intrinsic representations, an intermediate dynamic representation is obtained, the expression of which is:
[0059] in, This represents the intermediate dynamic representation.
[0060] S8. Time series forecasting based on intermediate dynamic representation.
[0061] The specific method is as follows: Layer normalization is performed on the intermediate dynamic representation; The intermediate dynamic representation after layer normalization is input into the feedforward network to extract deep features, further exploring the potential patterns after the fusion of endogenous and exogenous features, and enhancing the feature representation capability. Its expression is as follows:
[0062] in, Representing depth features, , All are learnable weight matrices. , All are bias terms; The deep features are mapped from the current dimension to a space that matches the target strategy dimension through a mapping layer, resulting in a time prediction sequence, the expression of which is:
[0063] in, For time-predicted series, Let be the projection weight matrix. This is a bias term.
[0064] In another feasible embodiment of the present invention, taking the prediction of traffic flow on a certain road segment within the next 24 hours as an example: The process involves acquiring historical traffic flow observation data (endogenous input) for a specific road segment, along with the segment's temperature, wind speed, visibility, road surface conditions, and corresponding public transportation scheduling plans, license plate restrictions, traffic control measures, and holiday parameters (exogenous input). The acquired endogenous inputs corresponding to traffic flow are integrated into an endogenous embedding vector; the acquired exogenous inputs corresponding to traffic flow are also integrated into an exogenous embedding vector; the endogenous and exogenous embedding vectors are concatenated to obtain a joint feature; based on the optimized endogenous representation, the endogenous and exogenous embeddings are fused using a dynamic interactive attention mechanism to obtain cross-modal interactive features; the optimized endogenous representation is then fused with the cross-modal interactive features to obtain an intermediate dynamic representation; time series prediction is performed based on the intermediate dynamic representation to obtain the traffic flow data sequence for the road segment within the next 24 hours. It should be noted that the specific implementation steps in this embodiment are logically identical to the aforementioned steps S1-S8, and will not be repeated here.
[0065] This invention constructs a hierarchical modeling system to achieve deep coupling modeling of the temporal characteristics of endogenous variables and the structural influence of exogenous variables. This solves the problem of difficulty in matching the heterogeneity and irregular temporal characteristics of exogenous variables with the dynamic characteristics of endogenous variables, thereby improving the accuracy and robustness of multivariate time series prediction in complex scenarios.
Claims
1. A time series forecasting method based on endogenous-exogenous synergy, characterized in that, include: Obtain the endogenous and exogenous inputs of the prediction target; Integrate intrinsic inputs into intrinsic embedding vectors; Integrate exogenous inputs into exogenous embedding vectors; The endogenous embedding vector and the exogenous embedding vector are concatenated to obtain the joint feature; Optimize endogenous representations based on joint features using a multi-level temporal attention mechanism; Based on the optimized intrinsic representation, the intrinsic and extrinsic embeddings are fused through a dynamic interactive attention mechanism to obtain cross-modal interaction features; The optimized endogenous representation is fused with cross-modal interaction features to obtain an intermediate dynamic representation; Time series forecasting based on intermediate dynamic representation; Among them, the endogenous input is the historical observation sequence of the prediction target, which is: the transaction amount, transaction volume, highest price, lowest price and opening price of a stock market at multiple time steps; or the traffic flow, speed and traffic density of a road segment at multiple time steps; or the electricity consumption, electricity price, power generation and power flow data of a region at multiple time steps. Exogenous inputs include environmental driving factors and policy control parameters.
2. The method according to claim 1, characterized in that, The specific method for integrating intrinsic inputs into intrinsic embedding vectors is as follows: Define the time difference between any two time steps in an endogenous input sequence as _____. ; Mapping the time difference to the relative position encoding space, the expression is: in, Represents a relative position encoding vector. Indicates the time difference The The sinusoidal component is obtained by mapping the frequency components. d The superscript T indicates the embedding dimension, and the superscript T indicates the transpose of the matrix. Define a periodic code for each timestamp, with the following expression: in, Represents a periodic encoded vector. The number of periodically coded components. n For the index of periodically coded components, For trainable projection matrices, The default fundamental frequency is defined by sin, which represents the sine function, and cos, which represents the cosine function. This indicates a radius of 180 degrees; By fusing the relative position encoding vector and the periodic encoding vector, a time-aware encoding vector is obtained, the expression of which is: in, Represents the time-aware encoding vector; The intrinsic input is integrated into an intrinsic embedding vector based on the time-aware coding vector, and its expression is as follows: in, Represents an endogenous embedding vector. Representation layer normalization, This indicates a temporal convolution operation. This indicates endogenous input.
3. The method according to claim 2, characterized in that, The specific method for integrating exogenous inputs into exogenous embedding vectors is as follows: Define the exogenous input variable at time t as Its expression is: M Indicates the number of exogenous variables. Indicates the first M One exogenous variable; The continuous exogenous input variable at time t is smoothed using Gaussian projection to obtain the smoothed features, the expression of which is: in, Represents the smoothed features. K Indicates the number of Gaussian functions. k Indicates the Gaussian function index. Indicates the first k The relevant weights corresponding to the Gaussian functions, where exp represents an exponential function with the natural constant e as the base. Indicates the first i One exogenous variable, For the first k The center of the Gaussian function, For the first k The standard deviation corresponding to the center of each Gaussian function Denotes the Euclidean norm; Mapping the discrete exogenous input variable at time t to a low-dimensional space of fixed dimensions is expressed as follows: in, Represents a low-dimensional vector. For the first i Learnable embedding matrix of discrete variables; The exogenous variables are integrated into an individual embedding vector, the expression of which is: in, Indicates the first i Individual embedding vectors of exogenous variables; and These are all indicator functions, used to determine... Does it belong to a continuous set? or discrete set ,like If the element belongs to the corresponding set, the indicator function takes a value of 1; otherwise, it takes a value of 0. By concatenating the individual embedding vectors of all exogenous variables, we obtain the exogenous embedding vector, which is expressed as: in, Represents an exogenous embedding vector. Indicates the first M Individual embedding vectors of exogenous variables.
4. The method according to claim 3, characterized in that, The endogenous embedding vector and the exogenous embedding vector are concatenated to obtain the joint feature, which is expressed as follows: in, Indicates joint features, Indicates splicing; It is a learnable global token that adapts to the time dimension through a broadcast mechanism.
5. The method according to claim 4, characterized in that, The specific method for optimizing endogenous representations based on joint features through a multi-level temporal attention mechanism is as follows: Local features are generated based on the intrinsic embedding components in the joint features using a learnable projection matrix, including: Time step Corresponding local query features Its expression is: ; Time step neighborhood Corresponding local bond features Its expression is: ; Time step neighborhood Corresponding local value features Its expression is: ; in, This represents the endogenous embedding component in the joint features. Endogenous embedding vector In the neighborhood The feature matrix within; , , All are learnable projection matrices; Based on the generated local features, a local attention mechanism is used to capture local dependencies. The local attention is calculated, and its expression is as follows: in, This represents the local attention output at time step t, and softmax represents the softmax activation function. To hide spatial dimensions; A global anchor vector is introduced and interacts with the endogenous embedding components in the joint features through a cross-channel gating mechanism. Its expression is as follows: in, Represents the global query vector; Represents the global anchor vector; Represents the global key vector; Represents the global value vector, , , All are projection matrices; TopK means retaining the first K time steps, ReLU means ReLU activation function, and GELU means GELU activation function; The global attention is calculated using the following expression: in, This represents the output of global attention. Indicates the first i One global key vector; Indicates the first i A global value vector; I Indicates the total number of global key-value pairs; This represents the softmax function; By integrating global attention, local attention, and endogenous embedding vectors through residual connections and layer normalization, an optimized endogenous representation is obtained, expressed as: in, This represents the optimized endogenous characterization. Indicates the gating coefficient. It represents the Hadamardi (or Hadama) stack.
6. The method according to claim 5, characterized in that, The specific method for obtaining cross-modal interaction features by fusing endogenous and exogenous embeddings based on optimized endogenous representations through a dynamic interactive attention mechanism is as follows: The query representation with context-aware capabilities is obtained through multi-head projection, and its expression is: in, This represents the query representation. Indicates multi-head projection operation; By compressing the exogenous embedding vector using a low-rank bottleneck, key-value pairs are generated, the expression of which is: in, Represents the key vector. Represents a value vector. , Both represent weight matrices. , Both represent low-rank projection matrices; The cross-variable score matrix is calculated using the following expression: in, Represents the cross-variable score matrix; Cross-modal interaction features are obtained by adaptively transforming the cross-variable score matrix and value vector through linear projection. The expression for this feature is as follows: in, Indicates cross-modal interaction features. Represents the linear projection matrix. This represents a sparse mask.
7. The method according to claim 6, characterized in that, The specific method for fusing the optimized endogenous representation with cross-modal interaction features to obtain the intermediate dynamic representation is as follows: The optimized intrinsic representation is subjected to average pooling and an activation function is applied to obtain a learnable gating vector, the expression of which is: in, Represents a learnable gated vector. This represents the sigmoid activation function. Indicates average pooling. Represents the weight matrix; By scaling cross-modal interaction features using learnable gating vectors and fusing them with optimized intrinsic representations, an intermediate dynamic representation is obtained, the expression of which is: in, This represents the intermediate dynamic representation.
8. The method according to claim 7, characterized in that, The specific method for time series forecasting based on intermediate dynamic representation is as follows: Layer normalization is performed on the intermediate dynamic representation; The intermediate dynamic representation after layer normalization is input into the feedforward network to extract deep features, the expression of which is: in, Representing depth features, , All are learnable weight matrices. , All are bias terms; The deep features are mapped from the current dimension to a space that matches the target strategy dimension through a mapping layer, resulting in a time prediction sequence, the expression of which is: in, For time-predicted series, Let be the projection weight matrix. This is a bias term.
9. A system for implementing the time series forecasting method based on endogenous-exogenous synergy as described in any one of claims 1 to 8, characterized in that, include: The intrinsic temporal embedding module is used to integrate intrinsic inputs into intrinsic embedding vectors; The exogenous structure embedding module is used to integrate exogenous inputs into exogenous embedding vectors; A multi-level temporal attention module is used to concatenate intrinsic and extrinsic embedding vectors and optimize the intrinsic representation based on joint features through a multi-level temporal attention mechanism. The dynamic attention module is used to fuse endogenous and exogenous embeddings based on optimized endogenous representations through a dynamic interactive attention mechanism to obtain cross-modal interactive features. The prediction module is used to fuse the optimized endogenous representation with cross-modal interaction features to obtain intermediate dynamic representation, and to perform time series prediction based on the intermediate dynamic representation.