Time sequence prediction method based on multi-branch structure and dual-prediction head gating mechanism
Through the time series prediction method with multi-branch structure and dual prediction head gating mechanism, the problem of insufficient multi-scale modeling coordination mechanism in the existing technology is solved, accurate prediction of high-dimensional, multi-variable and long-period data is achieved, and the adaptability and robustness of the model are improved.
Patent Information
- Application Number
- CN202510751993.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-09-12
AI Technical Summary
Existing time series modeling methods lack an effective multi-scale modeling coordination mechanism when processing high-dimensional, multivariate, and long-period data. The single prediction head has limited expressive power, lacks adaptive control capabilities, and has insufficient position modeling capabilities, making it difficult to balance the modeling of local details and overall trends.
A multi-branch structure and a dual-prediction head gating mechanism are adopted to extract features through a multi-branch blocking module and a multi-scale fusion module, combined with Reversible Instance Normalization for standardized preprocessing, and adaptively fused through the dual-prediction head output and gating mechanism to achieve collaborative expression and dynamic control of multi-scale information.
It improves the accuracy and robustness of time series forecasting, can effectively capture local changes and long-term trends, enhances the adaptability and generalization ability of the model, and is suitable for forecasting tasks in complex dynamic scenarios.
Smart Images

Figure CN120632780A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer technology, and in particular relates to a time series prediction method based on a multi-branch structure and a dual-prediction head gating mechanism. Background Art
[0002] With the development of the Internet of Things, big data, and automation technologies, time series data continues to grow rapidly in multiple practical application areas, covering a wide range of scenarios, including power load forecasting, weather monitoring, financial market modeling, and industrial equipment monitoring. In these applications, extracting effective features from high-dimensional, multivariate, and long-period time series and making accurate predictions has become a core issue that needs to be addressed in intelligent perception and decision-making systems.
[0003] Early time series modeling methods were mostly based on traditional statistical models, such as AR, ARIMA, and Holt-Winters. These methods are simple in structure, highly interpretable, and capable of handling both linear trends and cyclical fluctuations. However, faced with the vast amount of nonlinear, nonstationary, and complexly interacting time series data in the real world, these classic models face significant bottlenecks in their predictive performance and struggle to adapt to the dynamic changes and multi-scale patterns among multiple variables.
[0004] In recent years, the application of deep learning in time series modeling has rapidly expanded. In particular, models based on recurrent neural networks (RNNs), long short-term memory networks (LSTMs), convolutional neural networks (CNNs), and Transformers have demonstrated superior capabilities in capturing long-term dependencies and extracting complex features. Among them, the Transformer family (such as Informer, FEDformer, and PatchTST) excels in modeling long sequences, while multi-branch structures (such as MTST and Crossformer) introduce different patch granularities to model multi-scale features. Furthermore, methods such as PatchMixer and TimesNet improve modeling efficiency by integrating features through deep convolutional modules.
[0005] Although the above models have made progress in some specific tasks, they still have many common problems:
[0006] 1. Lack of effective multi-scale modeling coordination mechanism: Most current multi-scale structures use multiple independent branches, and information is not shared between branches. The fusion method is often relatively simple (such as averaging and weighted summation), which limits the ability to coordinate the expression of cross-scale information.
[0007] 2. Only a single prediction head output is used, with limited expressive power: Most models rely on a unified prediction head for final result mapping, making it difficult to model features at different semantic levels (such as local details and overall trends).
[0008] 3. The information fusion and modeling path is single and lacks adaptive control capabilities: There is a lack of dynamic modeling mechanism for branch importance and prediction path reliability, and it is impossible to flexibly adjust the contribution of different information sources in practical applications.
[0009] 4. Weak position modeling capabilities: Most models use static position coding or no position coding, which cannot fully express the dynamic changing characteristics of position information in time series. Summary of the Invention
[0010] In view of this, the present invention aims to overcome the shortcomings of the above-mentioned problems in the prior art, and proposes a time series prediction method based on a multi-branch structure and a dual-prediction head gating mechanism, by integrating a multi-branch multi-scale structure, a dual-prediction path design and a gating fusion mechanism, and combining Reversible Instance Normalization (RevIN) for standardized preprocessing. Each branch independently models temporal features of different scales, and outputs prediction results through the main prediction head (head0) and the enhanced prediction head (head1), respectively, and then adaptively fused by the gating mechanism, so as to fully capture information on local changes and long-term trends. Finally, the outputs of all branches are integrated through a global weighted fusion strategy to achieve a dual improvement in accuracy and robustness.
[0011] To achieve the above object, the technical solution of the present invention is achieved as follows:
[0012] A first aspect of the present invention provides a time series prediction method based on a multi-branch structure and a dual prediction head gating mechanism, comprising:
[0013] Constructing a time series prediction model, the model includes a multi-branch time segment feature extraction module, a convolutional hybrid coding feature extraction module, and a dual-path fusion prediction module;
[0014] The multi-branch time segment feature extraction module includes a multi-branch block module and a multi-scale fusion module; the multi-branch block module is used to divide the same time series into blocks of different sizes according to different scales, and the multi-scale fusion module is used to fuse the branches after block processing;
[0015] The convolutional hybrid coding feature extraction module includes a deep convolution feature extraction block and a point-by-point convolution feature extraction block; the deep convolution feature extraction block is used to extract local features within a patch, and the point-by-point convolution feature extraction block is used to perform feature interaction between patches;
[0016] The dual-path fusion prediction module includes a dual-prediction head collaborative learning module and a fusion prediction module based on a gating mechanism; the dual-prediction head collaborative learning module is used to output respective prediction result sequences based on different strategies, and the fusion prediction module based on a gating mechanism is used to perform weighted fusion on the outputs of the dual prediction paths;
[0017] The time series prediction model is trained to obtain the final detection model, and the time series prediction is performed based on the detection model.
[0018] Furthermore, the multi-branch time segment feature extraction module adopts a patch-based representation method to slice the input sequence into blocks, and uses linear mapping to map the original sequence to a high-dimensional feature space to capture local time structure information.
[0019] Furthermore, the multi-branch time segment feature extraction module is implemented as follows:
[0020] make Indicates that in the (n-1)th layer of the model, the d of one of the variables in the multivariate time series n-1 dimensional output representation;
[0021] Let P bn Indicates the size of the patch, S bn Indicates branch b n The length of the non-overlapping region between two consecutive patches in (stride);
[0022] In the marker T bn : In the example, the number of blocks is converted into overlapping blocks generated by sliding windows, and the number of blocks is J. bn With P bn 、S bn d n-1 The relationship is as follows:
[0023]
[0024] Different branches set different patch sizes P bn , achieving multi-scale analysis.
[0025] Furthermore, each branch contains a linear layer (Linear), a convolution block (ConvBlock) and an expansion (Flatten) operation, which is as follows:
[0026] f i =Flatten(ConvBlock(Linear(x i )))
[0027] Among them, Linear() represents the function that maps the original patch to a specified dimension, fi represents the input time series Patch size, and Flatten() represents the integration of the output features of all branches to generate the final time series representation y (n) As the input of the next layer or the final prediction result.
[0028] Furthermore, the convolutional hybrid coding feature extraction module includes a one-dimensional convolution kernel, a 1×1 convolution kernel, an activation function and normalized detail features.
[0029] Furthermore, the prediction model also includes normalization preprocessing of the input original time series through the ReVIN module.
[0030] The second aspect of the present invention provides an electronic device, comprising a processor and a memory connected to the processor for storing instructions executable by the processor, wherein the processor is used to execute the above-mentioned time series prediction method based on a multi-branch structure and a dual prediction head gating mechanism.
[0031] A third aspect of the present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the above-mentioned time series prediction method based on a multi-branch structure and a dual-prediction head gating mechanism.
[0032] Compared with the existing technology, the time series prediction method based on the multi-branch structure and dual prediction head gating mechanism described in the present invention has the following advantages:
[0033] This paper proposes a multi-scale time series modeling mechanism based on Patch representation, which can characterize the local details and global trend information of time series from multiple granularities, realize the effective separation and modeling of multiple potential time series patterns, and enhance the time decoupling capability of the model.
[0034] The present invention adopts a hybrid operation of channel interaction and cross-patch, which significantly improves the model's ability to model multivariate interaction relationships and local temporal dependencies, while reducing computational complexity.
[0035] The present invention constructs a multi-scale feature extraction framework containing multiple depth-adjustable branches. Each branch models a different temporal receptive field, enhancing the model's adaptability to multi-timescale structural changes and improving the model's performance in complex dynamic scenarios.
[0036] The present invention introduces a learnable gating mechanism to dynamically fuse multi-branch features, enabling the model to adaptively select more discriminative temporal features based on the pattern differences of specific input sequences, thereby improving the generalization ability and robustness of the model.
[0037] The model constructed by the present invention has shown significantly better accuracy and efficiency than existing methods in multiple real industrial forecasting tasks and multi-step forecasting tasks, and has good application prospects in practical scenarios such as energy forecasting and industrial fault prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] The accompanying drawings, which constitute part of the present invention, are provided to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are provided to explain the present invention and do not constitute an undue limitation of the present invention. In the accompanying drawings:
[0039] Figure 1 It is the overall framework diagram of the method of the present invention;
[0040] Figure 2 Schematic diagram of the convolutional hybrid coding feature extraction module structure. DETAILED DESCRIPTION
[0041] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments may be combined with each other.
[0042] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "lateral", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside" and the like indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, features defined as "first", "second", etc. may explicitly or implicitly include one or more of the features. In the description of the present invention, unless otherwise specified, "multiple" means two or more.
[0043] In the description of the present invention, it should be noted that, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood in a broad sense. For example, they may refer to fixed connections, detachable connections, or integral connections; mechanical connections or electrical connections; direct connections or indirect connections through an intermediate medium; and internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on specific circumstances.
[0044] The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.
[0045] Example 1:
[0046] refer to Figure 1 Shown is the overall framework diagram of the present invention. The present invention constructs a prediction model, which includes three parts. First, the input original time series will be normalized and pre-processed by the ReVIN module. ReVIN (Reversible Instance Normalization) can normalize each variable channel by channel and restore the original scale (denormalization) after model prediction, effectively alleviating the interference of inconsistent dimensions of different variables on model learning and improving cross-scene generalization capabilities. Afterwards, in the multi-branch time segment feature extraction module, the core idea of this module is to divide the time series into local segments of equal length through a blocking mechanism, thereby converting them into two-dimensional structure input. Through parallel multi-path processing, the model can learn diverse local dynamic features at different scales and different perspectives. Next, the blocked data will be input into the convolutional hybrid coding feature extraction module, which uses a deep convolutional structure to further enhance the receptive field and cross-channel modeling capabilities of the model. The module includes two convolution operations: a depthwise convolution module and a pointwise convolution module. The depthwise convolution module focuses on modeling the internal temporal dependencies of a single variable, while the pointwise convolution module captures cross-channel feature interactions and is supplemented by GELU activation and batch normalization to enhance nonlinear expression capabilities and training stability. This module deeply abstracts and encodes the fused features, providing a highly expressive, high-dimensional semantic representation for subsequent predictions. The processed features are input to the dual-path fusion prediction module. This module is designed with two parallel prediction paths, each capturing different types of information (such as trend terms and cyclical terms). Each path includes a prediction head in the form of a convolution or MLP. Ultimately, the outputs of the two paths are adaptively fused through a gating mechanism. The fused results are first flattened and concatenated, and then sent to the ReVIN module for denormalization, restoring the data to the same scale as the original input, ready for final prediction or downstream analysis tasks.
[0047] Specifically, for the multi-branch time segment feature extraction module, first, let Indicates that in the (n-1)th layer of the model, the d of one of the variables in the multivariate time series n-1 dimensional output representation. Let P bn Indicates the size of the patch, S bn Indicates branch b n The length (stride) of the non-overlapping region between two consecutive patches in . bn : In the example, the number of blocks J is converted into overlapping blocks generated by sliding windows. bn With Pbn 、S bn d n-1 The relationship is shown in the following formula (1):
[0048]
[0049] Different branches set different patch sizes P bn , thus achieving multi-scale analysis:
[0050] Small Patch: Used to capture high-frequency, short-term patterns with higher resolution.
[0051] Large Patch: Used to capture low-frequency, long-term patterns with lower resolution.
[0052] Each branch includes a linear layer (Linear), a convolution block (ConvBlock), and a flattening (Flatten) operation. Its form is shown in formula (2):
[0053] f i =Flatten(ConvBlock(Linear(x i ))) (2)
[0054] Among them, Linear() represents the function that maps the original patch to a specified dimension, fi represents the input time series Patch size, and Flatten() represents the integration of the output features of all branches to generate the final time series representation y (n) As the input of the next layer or the final prediction result.
[0055] After the block division, we advance the linear mapping process into the convolutional hybrid coding feature extraction module, which is divided into two convolution operations: depthwise convolution module and pointwise convolution module. The form is shown in formula (3):
[0056] f′ i =DWConv(PWConv(f i )) (2)
[0057] After this, the dual-path fusion prediction module in the branch performs nonlinear mapping on the data processed by the convolutional hybrid coding feature extraction module. In addition, there is also direct output linear prediction, and the two prediction heads output the final prediction results through the gating mechanism.
[0058] Next, the outputs of each branch are fused, flattened, and spliced, and then denormalized to generate the final time series representation y (n)As the input of the next layer or the final prediction result.
[0059] The present invention uses different convolution kernels in multiple branches to model multiple potential temporal patterns of the original time series. For each branch, depthwise convolution and pointwise convolution are used. This includes a one-dimensional convolution kernel, a 1×1 convolution kernel, an activation function, and normalized detail features. On this basis, a deep feature extraction block containing a one-dimensional depthwise convolution and an overall feature extraction block containing a 1×1 convolution are used to extract the global correlation of the temporal series. Finally, a merging module is used to fuse the information from branches of different scales.
[0060] Specifically, such as Figure 2 The figure shows the main structure of the convolutional mixed coding feature extraction module. This module first performs a deep convolution operation on the input time series. The core of this operation is to use a one-dimensional convolution kernel to perform convolution processing on each channel, so as to model local information only in the time dimension and not mix between channels. This process can effectively extract the temporal detail pattern of each channel itself. For example, for time series data with an input of N×D (where N represents the time length and D represents the feature dimension), Depthwise convolution can perceive the local correlation between consecutive time steps with a low number of parameters. The depthwise convolution processing process is shown in formula (4):
[0061]
[0062] In addition, by adding the GELU nonlinear activation function and Batch Normalization, this module further enhances the feature expression capability and the stability of model training.
[0063] Among them, the function BN() (Batch Normalization) represents a normalization operation commonly used in neural networks. Its main function is to accelerate convergence, improve stability and prevent overfitting.
[0064] Next, pointwise convolution (i.e., 1×1 convolution) is used to fuse the output of depthwise convolution in the channel direction. This operation is equivalent to performing feature transformation at each time step, which plays a role similar to full connection, allowing originally independent channel information to interact, thereby capturing the synergistic relationship between different time series variables. This stage is also equipped with GELU activation function and batch normalization to improve nonlinear modeling capabilities and training stability. The convolution process is shown in formula (5):
[0065]
[0066] Then the fusion operation is performed in the dual-path fusion prediction module:
[0067]
[0068] in is the prediction result from the direct linear mapping, It is the prediction result from the nonlinear mapping of the MLP module. λ∈[0, 1] represents a learnable gating parameter that controls the fusion ratio of the two branch outputs and is generated by sigmoid().
[0069] The present invention solves the problem that the traditional Transformer structure is difficult to simultaneously take into account local detail changes and overall trend modeling through innovative multi-branch multi-scale patch modeling, PatchMixer feature extraction, and dual prediction head gated fusion design. The present invention introduces a multi-scale branching mechanism to simultaneously extract local and long-term features from different granularities, effectively improving the fitting and prediction capabilities of complex time series data. Compared with the problem that a single prediction head structure is difficult to take into account short-term fluctuations and long-term trends, the present invention solves the defect of limited single-path expression of the model through dual prediction head collaborative modeling.
[0070] The present invention has verified stable high performance in a variety of public data sets (such as Weather, ETTh1, ETTh2, etc.) and various forecasting tasks (short-term forecasting, long-term trend modeling), indicating that it has good universality and robustness.
[0071] The effectiveness of the solution of the present invention is verified by experiments below.
[0072] The experiment used seven real-world public datasets for model training and testing, namely: (1) Weather data. This dataset collects 21 meteorological indicators in Germany, such as temperature and humidity, with 21 dimensions and 52,696 time steps; (2) Traffic data. This dataset records the road occupancy rate of different sensors on San Francisco highways, with 862 dimensions and 17,544 time steps; (3) Electricity data. This dataset describes the hourly electricity consumption of 321 customers, with 321 dimensions and 26,304 time steps; (4) Electricity Transformer Temperature (ETT). ETT contains data on two power transformers, each with two resolutions (15 minutes and 1 hour), so it is divided into four subsets:
[0073] ETTh1 and ETTh2: These two subsets correspond to transformer 1 and transformer 2 respectively, updated every hour, each with 7 dimensions and 17420 time steps;
[0074] ETTm1 and ETTm2: These two subsets also correspond to transformer 1 and transformer 2 respectively.
[0075] But it is updated at a frequency of 15 minutes, each with 7 dimensions and 69,680 time steps;
[0076] The overall information is shown in Table 1.
[0077] Table 1
[0078]
[0079] In order to verify the model detection effect, this paper uses mean squared error (MSE) and mean absolute error (MAE) as evaluation indicators of model performance. The smaller the value, the better the model performance. The specific calculation formulas for each are:
[0080]
[0081] The experimental environment used by the model is based on the Intel Xeon E5-2678 v3 CPU and GeForce RTX2080Ti graphics card. The software environment is based on the Ubuntu 18.04 operating system, Python 3.10.13, and Pytorch 1.10.2.
[0082] The experimental results, shown in Table 2, show that the prediction model performs well across seven datasets (e.g., Weather, Etth1, Etth2, Ettm1, Ettm2, Electricity, and Traffic), achieving the lowest mean square error (MSE) and mean average error (MAE) for most prediction steps, validating its robust performance in time series modeling. This outstanding performance is attributed to the multi-branch time-segment feature extraction module within the model structure, which extracts both fine-grained and coarse-grained information from multiple temporal scales, balancing short-term local variations with long-term global trends. Furthermore, the convolutional hybrid coding module efficiently captures local dependencies through depthwise separable convolutions, reducing computational complexity and achieving excellent modeling efficiency. The dual-path fusion prediction module, which fuses linear and nonlinear prediction paths through a gating mechanism, effectively enhances the model's nonlinear expressiveness and robustness.
[0083] Table 2
[0084]
[0085]
[0086] In particular, in datasets with strong periodicity, such as Weather and ETTh1, the forecasting model demonstrates excellent characterization of multi-periodic, multi-frequency sequences, accurately capturing patterns and making stable predictions. Furthermore, in industrial energy consumption datasets, such as ETTh1 and ETTh2, the forecasting model demonstrates strong generalization capabilities for data with both rapid and gradual changes.
[0087] To verify the effectiveness of the modules in this paper, ablation experiments were conducted on the ETTh2 dataset. The experimental settings were as follows:
[0088] Experiment A is to change the multi-branch time segment feature extraction module into a single-branch time segment feature extraction module.
[0089] Experiment b is to remove the design of the convolutional mixed coding feature extraction module
[0090] Experiment C replaces the gating mechanism in the dual-path fusion prediction module with a common residual connection design.
[0091] Experiment d is the original design of this invention
[0092] The results of the ablation experiments are shown in Table 3, demonstrating significant differences in the contributions of the various components of the prediction model to overall performance. Replacing the multi-branch time-segment feature extraction module with a single-branch time-segment feature extraction module (Experiment a) shows a significant drop in model performance. This demonstrates that this module can extract key information from multiple time scales through different patch partitioning strategies, effectively enhancing the model's ability to model periodicity and trends in complex time series. Removing the convolutional hybrid coding feature extraction module (Experiment b) also significantly degrades model accuracy, demonstrating the importance of combining deep and shallow convolutional structures (depthwise and pointwise) in capturing local patterns and cross-temporal dependencies in the sequence. Replacing the gating mechanism in the dual-pathway fusion prediction module with a standard residual connection (Experiment c) results in a slight drop in model performance. This demonstrates that the gating mechanism can adaptively adjust the contributions of different prediction pathways (linear and nonlinear) to the final output, making it more effective in capturing complex variations in the time series. In contrast, the fully designed prediction model (Experiment d) achieves optimal results across all prediction lengths, validating the effectiveness and synergy of the three proposed modules. It can be seen that the design of each module of the prediction model is not only reasonable but also cooperates with each other, which significantly improves the accuracy of time series modeling and prediction.
[0093] Table 3
[0094]
[0095] This paper experimentally evaluated the model's performance and compared it with multiple models. The experimental results show that the proposed model provides better detection results than other models. Across multiple scenarios, the model achieved good results and demonstrated good generalization capabilities.
[0096] Example 2:
[0097] An electronic device includes a processor and a memory connected to the processor for storing instructions executable by the processor, wherein the processor is used to execute the above-mentioned time series prediction method based on a multi-branch structure and a dual-prediction head gating mechanism.
[0098] Example 3:
[0099] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the above-mentioned time series prediction method based on a multi-branch structure and a dual-prediction head gating mechanism.
[0100] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A time series prediction method based on a multi-branch structure and a dual-prediction head gating mechanism, characterized by: include: Constructing a time series prediction model, the model includes a multi-branch time segment feature extraction module, a convolutional hybrid coding feature extraction module, and a dual-path fusion prediction module; The multi-branch time segment feature extraction module includes a multi-branch block module and a multi-scale fusion module; the multi-branch block module is used to divide the same time series into blocks of different sizes according to different scales, and the multi-scale fusion module is used to fuse the branches after block processing; The convolutional hybrid coding feature extraction module includes a deep convolution feature extraction block and a point-by-point convolution feature extraction block; the deep convolution feature extraction block is used to extract local features within a patch, and the point-by-point convolution feature extraction block is used to perform feature interaction between patches; The dual-path fusion prediction module includes a dual-prediction head collaborative learning module and a fusion prediction module based on a gating mechanism; the dual-prediction head collaborative learning module is used to output respective prediction result sequences based on different strategies, and the fusion prediction module based on a gating mechanism is used to perform weighted fusion on the outputs of the dual prediction paths; The time series prediction model is trained to obtain the final detection model, and the time series prediction is performed based on the detection model.
2. The time series prediction method based on a multi-branch structure and a dual-prediction head gating mechanism according to claim 1, characterized in that: The multi-branch time segment feature extraction module adopts a patch-based representation method to slice the input sequence into blocks, and uses linear mapping to map the original sequence to a high-dimensional feature space to capture local time structure information.
3. The time series prediction method based on a multi-branch structure and a dual-prediction head gating mechanism according to claim 1, characterized in that: The multi-branch time segment feature extraction module is implemented as follows: make Indicates that in the (n-1)th layer of the model, the d of one of the variables in the multivariate time series n-1 dimensional output representation; Let P bn Indicates the size of the patch, S bn Indicates branch b n The length of the non-overlapping region between two consecutive patches in (stride); In the marker T bn : In the example, the number of blocks is converted into overlapping blocks generated by sliding windows, and the number of blocks is J. bn With P bn 、S bn d n-1 The relationship is as follows: Different branches set different patch sizes P bn , achieving multi-scale analysis.
4. The time series prediction method based on a multi-branch structure and a dual prediction head gating mechanism according to claim 1, characterized in that: Each branch contains a linear layer (Linear), a convolution block (ConvBlock) and an expansion (Flatten) operation, which is as follows: f i =Flatten(ConvBlock(Linear(x i ))) Among them, Linear() represents the function that maps the original patch to a specified dimension, fi represents the input time series Patch size, and Flatten() represents the integration of the output features of all branches to generate the final time series representation y (n) As the input of the next layer or the final prediction result.
5. The time series prediction method based on a multi-branch structure and a dual prediction head gating mechanism according to claim 1, characterized in that: The convolutional hybrid coding feature extraction module includes a one-dimensional convolution kernel, a 1×1 convolution kernel, an activation function, and normalized detail features.
6. The time series prediction method based on a multi-branch structure and a dual prediction head gating mechanism according to claim 1, characterized in that: The prediction model also includes normalization preprocessing of the input original time series through the ReVIN module.
7. An electronic device comprising a processor and a memory in communication with the processor and configured to store instructions executable by the processor, wherein: The processor is used to execute the time series prediction method based on the multi-branch structure and dual prediction head gating mechanism described in any one of claims 1-6.
8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the time series prediction method based on a multi-branch structure and a dual-prediction head gating mechanism according to any one of claims 1 to 6 is implemented.