Time sequence prediction method and system based on multi-source feature fusion
By fusing historical observation data and short-term prediction data through independent encoding and cross-modal attention mechanisms, an efficient prediction time series is generated, which solves the problem of insufficient multi-source information fusion in existing technologies and achieves more accurate and robust prediction results.
Patent Information
- Application Number
- CN202510847634.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-09-26
AI Technical Summary
Existing time series forecasting methods fail to effectively integrate multi-source information, resulting in insufficient forecast accuracy and robustness.
A method based on multi-source feature fusion is adopted to generate query vectors and key-value vectors through independent encoders, and feature fusion is performed using a cross-modal attention mechanism to generate fused feature representations, and finally the target prediction time series is generated through the decoder.
It improves the accuracy and robustness of predictions, enhances adaptability to complex dynamic environments, reduces the deviation caused by a single data source, and improves the robustness and accuracy of prediction outputs.
Smart Images

Figure CN120705813A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the technical field of wind power prediction, and specifically relates to a time series prediction method and system based on multi-source feature fusion. Background Art
[0002] Time series (abbreviated as "time series") forecasting is an important research direction in machine learning, widely used in financial market forecasting, weather forecasting, equipment failure prediction, energy load forecasting, and other fields. With the development of data acquisition technology, practical applications often require the simultaneous processing of multiple types of time series data, such as historical observations, short-term trend data, and external influencing factors.
[0003] Existing methods typically rely solely on historical observations for forecasting. However, in many real-world scenarios, auxiliary information may be present, such as short-term forecasts generated by other models and relevant external variables. Effectively fusing this multi-source information to improve forecast accuracy and robustness is a key research direction in the field of time series forecasting. Existing simple information fusion methods (such as concatenation and weighted averaging) may fail to fully capture the complex interactions between different information sources, resulting in poor fusion results.
[0004] Therefore, it is urgent to propose a new time series prediction method that can effectively integrate historical observation data and auxiliary prediction information. Summary of the Invention
[0005] This application proposes a time series prediction method and system based on multi-source feature fusion to address the above-mentioned defects of the prior art.
[0006] According to a first aspect of an embodiment of the present application, a time series prediction method based on multi-source feature fusion is provided, comprising:
[0007] receiving first time series data and second time series data;
[0008] Encoding the first time series data using an independent first encoder to generate a first encoded representation, where the first encoding is a query vector;
[0009] Parallel encoding of the second time series data by an independent second encoder to generate a second encoded representation, wherein the second encoding is a key vector and a value vector;
[0010] Inputting the first encoding and the second encoding into the cross-modal attention mechanism for fusion to generate a fused feature representation;
[0011] Based on the fused feature representation, a target prediction time series is generated through a decoder.
[0012] In some implementations, the first time series data includes historical observation data, and the second time series data includes short-term forecast data.
[0013] In some implementations, inputting the first encoding and the second encoding into a cross-modal attention mechanism for fusing to generate a fused feature representation includes:
[0014] Calculating an attention weight based on the similarity between the query vector and the key vector;
[0015] The value vectors are weighted and summed based on the attention weights to obtain the fused feature representation.
[0016] In some embodiments, the attention mechanism is a single-head or multi-head attention mechanism.
[0017] In some implementations, inputting the first encoding and the second encoding into a cross-modal attention mechanism for fusing to generate a fused feature representation includes:
[0018] Perform multiple attention weight calculations of the first encoding and the second encoding in parallel through the multi-head attention mechanism;
[0019] The results of the multiple attention weight calculations are concatenated and linearly transformed into the fused feature representation.
[0020] In some embodiments, the first encoder and / or the second encoder comprises a convolutional neural network layer and a recurrent neural network layer.
[0021] In some embodiments, the decoder operates in an autoregressive manner.
[0022] According to a second aspect of an embodiment of the present application, a time series prediction system based on multi-source feature fusion is provided, comprising:
[0023] A data receiving module, configured to receive first time series data and second time series data;
[0024] a first encoding representation module, configured to encode the first time series data using an independent first encoder to generate a first encoding representation, where the first encoding is a query vector;
[0025] a second encoding representation module, configured to perform parallel encoding on the second time series data using an independent second encoder to generate a second encoding representation, wherein the second encoding is a key vector and a value vector;
[0026] a feature fusion module, configured to input the first encoding and the second encoding into a cross-modal attention mechanism for fusion to generate a fused feature representation;
[0027] The target sequence generation module is used to generate a target prediction time series through a decoder based on the fused feature representation.
[0028] According to a third aspect of an embodiment of the present application, a computer-readable storage medium is provided, storing a computer program, characterized in that when the computer program is executed by a processor, the steps of the time series prediction method based on multi-source feature fusion as described above are implemented.
[0029] According to a fourth aspect of an embodiment of the present application, an electronic device is provided, characterized in that it includes: a memory for storing a computer program; and a processor for implementing the above-mentioned time series prediction method based on multi-source feature fusion when executing the computer program.
[0030] The beneficial effects of the time series prediction method and system based on multi-source feature fusion according to the embodiments of the present application include at least:
[0031] The embodiment of the present application allows for simultaneous acquisition of multiple information sources (such as historical observations and short-term forecast data) by receiving first time series data and second time series data, enriching the input background of the model, thereby reducing the prediction uncertainty caused by the lack or limitation of a single data. In actual scenarios (such as equipment failure warning or energy load forecasting), the complementarity of multi-source data can enhance adaptability to complex dynamic environments and avoid the bias problem inherent in a single data source; through independent encoding design (the first encoder generates a query vector, and the second encoder generates a key-value vector), the coupling of different source information in the feature extraction stage can be avoided, ensuring that their respective features (such as the long-term trend of historical time series and the transient fluctuations of short-term forecasts) are retained more purely; through the cross-modal attention mechanism, the correlation between multi-source features is dynamically captured based on the query-key-value relationship, and the feature fusion process can be optimized (weights are assigned more accurately than traditional methods); finally, the fused features and decoding generation are combined to achieve efficient utilization of multi-source information and improve the robustness and accuracy of the prediction output in real scenarios (such as market volatility response). BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 Schematic diagram of the flow of a time series prediction method based on multi-source feature fusion according to an embodiment of the present application;
[0033] Figure 2 Schematic diagram of the structure of a time series prediction system based on multi-source feature fusion according to an embodiment of the present application. DETAILED DESCRIPTION
[0034] In order to enable those skilled in the art to better understand the technical solution of the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0035] The following detailed description of the embodiments of the present application is provided in conjunction with the accompanying drawings and examples. The following detailed description of the embodiments and the accompanying drawings are used to illustrate the principles of the present application, but are not intended to limit the scope of the present application, that is, the present application is not limited to the described embodiments.
[0036] Refer to the attached Figure 1 As shown, the embodiment of the present application discloses specific implementation steps of a time series prediction method based on multi-source feature fusion. The method is performed based on a time series prediction system based on multi-source feature fusion. The system is configured in a readable storage medium and an electronic device to ensure that those skilled in the art can implement the technical solution of the present application accordingly. The method specifically includes the following steps 110-150.
[0037] Step 110: Receive first time series data and second time series data.
[0038] In some implementations, the first time series data includes historical observation data, and the second time series data includes short-term forecast data.
[0039] Step 120: Encode the first time series data using an independent first encoder to generate a first encoded representation.
[0040] For example, the first code is a query vector.
[0041] In some embodiments, the first encoder comprises a convolutional neural network layer and a recurrent neural network layer.
[0042] Step 130 : Parallel encode the second time series data using an independent second encoder to generate a second encoded representation.
[0043] For example, the second encoding is a key vector and a value vector.
[0044] In some embodiments, the second encoder comprises a convolutional neural network layer and a recurrent neural network layer.
[0045] In some embodiments, the first encoder and the second encoder process their respective input data in parallel. The first encoder and the second encoder are configured to independently and parallelly encode the two types of received time series data (the first time series data and the second time series data). The manner in which the first and second encoders generate the first and second encoded representations can be referred to as separated encoding. The first encoder and / or the second encoder are separated encoders, which are configured to design independent feature extraction networks for input data of different properties to avoid confusion between heterogeneous data features.
[0046] In some embodiments, the first encoder and / or the second encoder comprises a convolutional neural network layer (CNN) and a recurrent neural network layer (RNN). Preferably, the convolutional neural network layer is a one-dimensional convolutional neural network (1D-CNN), which is a convolutional network specifically designed to process sequence data and is used to extract local temporal patterns. In addition, the recurrent neural network layer is configured with a gated recurrent unit (GRU), which is an improved recurrent neural network structure with better long-term dependency modeling capabilities.
[0047] In step 140, the first encoding and the second encoding are input into a cross-modal attention mechanism for fusion to generate a fused feature representation.
[0048] Among them, Cross-Modal Attention is an attention mechanism used to establish correlations between different data modalities and realize the intelligent fusion of heterogeneous information.
[0049] In addition, fused feature representation is also called fused temporal feature representation (Temporal Feature Representation). Temporal feature representation is used to map original time series data into an encoding representation of high-dimensional semantic features. Fused feature representation is also called multi-modal data fusion (Multi-Modal Data Fusion) of temporal features. Multi-modal data fusion is used to integrate information from different data sources or with different characteristics for joint modeling.
[0050] In some embodiments, the first encoding and the second encoding are input into a cross-modal attention mechanism for fusion to generate a fused feature representation, which includes: calculating an attention weight based on the similarity between the query vector and the key vector; and weighted summing the value vector based on the attention weight to obtain the fused feature representation.
[0051] In some embodiments, the attention mechanism is a single-head or multi-head attention mechanism.
[0052] Among them, Multi-Head Attention is the core component of the Transformer architecture, which can focus on multiple position information in the time series in parallel.
[0053] In some embodiments, the first encoding and the second encoding are input into a cross-modal attention mechanism for fusion to generate a fused feature representation, which includes: performing multiple attention weight calculations of the first encoding and the second encoding in parallel through a multi-head attention mechanism; and concatenating the results of the multiple attention weight calculations and linearly transforming them into a fused feature representation.
[0054] Step 150: Generate a target prediction time series through a decoder based on the fused feature representation.
[0055] In some embodiments, the decoder operates in an autoregressive manner. Autoregressive decoding is a recursive prediction method that gradually predicts the value of the next moment based on a time series of generated multivariate feature fusions.
[0056] The embodiment of the present application allows for simultaneous acquisition of multiple information sources (such as historical observations and short-term forecast data) by receiving first time series data and second time series data, enriching the input background of the model, thereby reducing the prediction uncertainty caused by the lack or limitation of a single data. In actual scenarios (such as equipment failure warning or energy load forecasting), the complementarity of multi-source data can enhance adaptability to complex dynamic environments and avoid the bias problem inherent in a single data source; through independent encoding design (the first encoder generates a query vector, and the second encoder generates a key-value vector), the coupling of different source information in the feature extraction stage can be avoided, ensuring that their respective features (such as the long-term trend of historical time series and the transient fluctuations of short-term forecasts) are retained more purely; through the cross-modal attention mechanism, the correlation between multi-source features is dynamically captured based on the query-key-value relationship, and the feature fusion process can be optimized (weights are assigned more accurately than traditional methods); finally, the fused features and decoding generation are combined to achieve efficient utilization of multi-source information and improve the robustness and accuracy of the prediction output in real scenarios (such as market volatility response).
[0057] Refer to the attached Figure 2 As shown, the embodiment of the present application discloses a time series prediction system based on multi-source feature fusion. The system includes a data receiving module 210, a first encoding representation module 220, a second encoding representation module 230, a feature fusion module 240 and a target sequence generation module 250.
[0058] The data receiving module 210 is configured to receive first time series data and second time series data.
[0059] The first encoding representation module 220 is configured to encode the first time series data using an independent first encoder to generate a first encoding representation, where the first encoding is a query vector.
[0060] The second encoding representation module 230 is configured to perform parallel encoding on the second time series data through an independent second encoder to generate a second encoding representation, where the second encoding is a key vector and a value vector.
[0061] The feature fusion module 240 is used to input the first encoding and the second encoding into the cross-modal attention mechanism for fusion to generate a fused feature representation.
[0062] The target sequence generation module 250 is configured to generate a target prediction time series through a decoder based on the fused feature representation.
[0063] In some embodiments, the feature fusion module is a cross-modal fusion module, preferably adopts a multi-head attention mechanism, and uses the output of the first encoder as the query and the output of the second encoder as the key and value.
[0064] The embodiments of the present application effectively avoid mutual interference of multi-source information in the feature extraction stage through independent feature extraction, especially through the separate encoder structure to independently process historical observation data and short-term prediction data; this design can more purely retain the inherent characteristics of each type of data (such as the long-term trend of historical data and the transient characteristics of short-term predictions), provide highly discriminative input representation for subsequent fusion, and improve feature quality from the source; through intelligent fusion mechanisms, such as the use of cross-modal attention mechanisms to dynamically model the correlation between two information sources (such as the temporal dependency between historical observations and short-term predictions), achieve feature interaction capabilities that surpass traditional fusion methods (such as splicing or weighted averaging); through the query-key matching mechanism to adaptively allocate fusion weights, significantly improve the effectiveness of multi-source information fusion, and thereby enhance the robustness and accuracy of the prediction results.
[0065] An embodiment of the present application also provides a computer-readable storage medium having a computer program stored thereon, characterized in that when the program is executed by a processor, the above-mentioned time series prediction method based on multi-source feature fusion is implemented. For example, the time series prediction method based on multi-source feature fusion of the present application can be implemented by computer program instructions, and the relevant code can be stored in a computer-readable storage medium (such as a hard disk, SSD or cloud server). When the program is executed by the processor, the steps of the above-mentioned time series prediction method based on multi-source feature fusion are automatically executed. For example, it can include: driving the processor to receive first time series data and second time series data; generating a first encoded representation and a second encoded representation in parallel through an independent encoder; performing cross-modal attention fusion with the first encoded representation as a query vector and the second encoded representation as a key-value vector; decoding and generating a target prediction time series based on the fused feature representation.
[0066] The readable storage medium of the present application converts the time series prediction method (including separable encoding and cross-modal attention fusion) into an executable computer program, realizing the standardized deployment and batch replication of the technical solution; the storage medium (such as disk, cloud storage) enables the prediction method to run independently of the development environment, adapt to heterogeneous platforms such as edge computing devices or server clusters, and improve deployment flexibility.
[0067] An embodiment of the present application further provides an electronic device comprising: a memory for storing a computer program; and a processor for implementing the aforementioned method for time series prediction based on multi-source feature fusion when executing the computer program. For example, the processor is configured to invoke a separate encoder module to independently encode dual-source data; implement matrix operations of a cross-modal attention mechanism through a hardware acceleration unit; and output a predicted time series to an output interface based on autoregressive decoding.
[0068] The electronic device of the present application calls hardware resources (such as GPU accelerated matrix operations) through the processor, significantly improving the computational efficiency of cross-modal attention fusion, meeting real-time prediction needs, and achieving end-to-end efficient execution; by integrating memory, processor and output interface, a closed-loop prediction system (such as an industrial equipment fault warning terminal) is formed, and the target prediction sequence is directly output to the user terminal or actuator, thereby improving system integration.
[0069] It is understood that the above embodiments are merely exemplary embodiments for illustrating the principles of the present application, and the present application is not limited thereto. Those skilled in the art may make various modifications and improvements without departing from the spirit and substance of the present application, and such modifications and improvements are also considered to be within the scope of protection of the present application.
Claims
1. A time series prediction method based on multi-source feature fusion, characterized in that: include: receiving first time series data and second time series data; Encoding the first time series data using an independent first encoder to generate a first encoded representation, where the first encoding is a query vector; Parallel encoding of the second time series data by an independent second encoder to generate a second encoded representation, wherein the second encoding is a key vector and a value vector; Inputting the first encoding and the second encoding into the cross-modal attention mechanism for fusion to generate a fused feature representation; Based on the fused feature representation, a target prediction time series is generated through a decoder.
2. The method according to claim 1, characterized in that The first time series data includes historical observation data, and the second time series data includes short-term forecast data.
3. The method according to claim 1, characterized in that The step of inputting the first encoding and the second encoding into a cross-modal attention mechanism to fuse the first encoding and the second encoding to generate a fused feature representation includes: Calculating an attention weight based on the similarity between the query vector and the key vector; The value vectors are weighted and summed based on the attention weights to obtain the fused feature representation.
4. The method according to claim 1, wherein The attention mechanism is a single-head or multi-head attention mechanism.
5. The method according to claim 4, characterized in that The step of inputting the first encoding and the second encoding into a cross-modal attention mechanism to fuse the first encoding and the second encoding to generate a fused feature representation includes: Perform multiple attention weight calculations of the first encoding and the second encoding in parallel through the multi-head attention mechanism; The results of the multiple attention weight calculations are concatenated and linearly transformed into the fused feature representation.
6. The method according to claim 1, wherein The first encoder and / or the second encoder comprises a convolutional neural network layer and a recurrent neural network layer.
7. The method according to claim 1, characterized in that The decoder operates in an autoregressive manner.
8. A time series prediction system based on multi-source feature fusion, characterized in that: include: A data receiving module, configured to receive first time series data and second time series data; a first encoding representation module, configured to encode the first time series data using an independent first encoder to generate a first encoding representation, where the first encoding is a query vector; a second encoding representation module, configured to perform parallel encoding on the second time series data using an independent second encoder to generate a second encoding representation, wherein the second encoding is a key vector and a value vector; a feature fusion module, configured to input the first encoding and the second encoding into a cross-modal attention mechanism for fusion to generate a fused feature representation; The target sequence generation module is used to generate a target prediction time series through a decoder based on the fused feature representation.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the time series prediction method based on multi-source feature fusion according to any one of claims 1 to 7 are implemented.
10. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the time series prediction method based on multi-source feature fusion according to any one of claims 1 to 7 when executing the computer program.
Citation Information
Patent Citations
Interaction behavior prediction method and device, electronic equipment and readable storage medium
CN118013115A
Multi-element load prediction method and device, storage medium and program product
CN120067995A
Cited By
Image feature fusion and target detection method, electronic equipment and readable storage medium
CN121438052A