A dynamic multi-modal time series modeling and fusion method and system

CN122839281APending Publication Date: 2026-09-29STATE GRID JIANGSU ELECTRIC POWER CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202611072452.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-20
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

部分方法仅依赖插值、重采样等传统手段完成对齐,不仅过程繁琐,而且可能引入额外误差,导致对多模态数据的真实时序关系刻画不够准确

Benefits of technology

1.本发明通过模态感知的自适应周期性编码机制,在绝对时间编码的基础上引入面向不同模态的周期相位校准网络,自适应生成日周期、周周期、季节周期等不同周期分量对应的相位偏移量。该机制使不同模态的周期编码能够结合各模态自身的历史响应特征表达其相对于全局时间基准的提前、同步或滞后状态,从而缓解气象变化、电力负荷波动、设备状态响应等多源数据在周期语义空间中的错位问题,提高异步多模态时序数据的表征精度和后续跨模态融合的准确性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122839281A_ABST
    Figure CN122839281A_ABST
Patent Text Reader

Abstract

The application discloses a dynamic multi-modal time series modeling and fusion method and system, and belongs to the cross technical field of artificial intelligence and power systems. Firstly, the power load, meteorological image and static equipment data are mapped and fused to generate adaptive periodic encoding for time enhancement; secondly, a time attention bias matrix based on the relative sampling time relationship is constructed, which is injected into the cross attention mechanism to adjust the correlation strength, and then a multi-modal unified representation is obtained through a gating mechanism; finally, under the federated learning framework, a dynamic alignment mechanism based on the maximum mean difference is introduced to calculate the cross-region distribution alignment loss, and the local parameters are updated independently in combination with the private module. The application effectively overcomes the multi-modal time series misalignment noise, and realizes efficient cooperation and high-precision personalized prediction of the cross-region model while protecting data privacy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of interdisciplinary technology of artificial intelligence and power systems, and more specifically, to a dynamic multimodal time series modeling and fusion method and system. Background Technology

[0002] Power data, as a crucial foundation for power system operation and dispatch, typically comprises various forms, including structured, semi-structured, and unstructured data. These data originate from diverse devices and channels, exhibiting strong heterogeneity. Traditional methods often rely on single-modal data for analysis and modeling, failing to fully leverage the correlations between different data types. This is particularly true when dealing with multimodal data exhibiting time-series characteristics, where the methods cannot effectively capture the temporal dependencies and dynamic evolution patterns across modalities. This, to some extent, limits the model's predictive accuracy and its ability to understand complex scenarios.

[0003] To address this issue, current research focuses on how to effectively integrate information from different time scales and data modalities (such as numerical, text, and image data) to capture cross-modal dependencies and dynamic evolution patterns, thereby improving the model's performance in downstream tasks such as classification, prediction, and diagnosis.

[0004] With the deep integration of artificial intelligence and Internet of Things technologies, in order to achieve the above goals in real-world scenarios, researchers have proposed a series of models and methods with independent intellectual property rights to address the heterogeneity, temporal asynchrony, and complexity of cross-modal interactions in multimodal data.

[0005] For example, patent application CN202411088310.3 discloses a multimodal fusion modeling method for power load forecasting. This method uses multi-source data, including historical electricity consumption records, seasonal fluctuation information, meteorological conditions, and special event factors, as input. In the data preprocessing stage, the data is first standardized to ensure consistency and comparability of the input data. Subsequently, different types of data are transformed into feature vectors or feature mapping is performed via a fully connected network to adapt to the input requirements of deep learning models. Regarding model structure, this patent uses a long short-term memory network to process electricity consumption time series, thereby capturing the dynamic characteristics of load evolution over time. Simultaneously, a cross-attention mechanism is introduced, combined with a multi-head self-attention structure, to achieve correlation modeling and information fusion of multimodal data, enabling the prediction model to better uncover complementary relationships between different modalities and improve prediction accuracy.

[0006] While existing technologies can utilize multimodal data for downstream applications to some extent, they still have the following shortcomings: First, in terms of temporal dependency modeling, existing technologies often focus more on the dependency relationship between features between modalities when fusing multimodal data, but fail to fully consider the dynamic temporal features within a modality, resulting in the loss of the temporal evolution pattern within the modality and thus unsatisfactory fusion results.

[0007] Secondly, regarding modal alignment, most multimodal fusion frameworks assume that the input data is time-aligned offline batch data, neglecting the dynamic arrival of data. Data streams from different modalities do not arrive at the processing center at the same time, with identical timestamps and frequencies; they are often asynchronous, at different speeds, and may even be lost. Some methods rely solely on traditional methods such as interpolation and resampling for alignment, which is not only cumbersome but may also introduce additional errors, resulting in an inaccurate characterization of the true temporal relationships of the multimodal data.

[0008] Finally, regarding model generalization and transfer capabilities, existing models often rely on specific datasets for training, lacking adaptability across regions and scenarios. When faced with new scenarios where power network topology and user load characteristics vary significantly, models typically require complex retraining or tuning, making it difficult to directly generalize and apply them to diverse power application scenarios.

[0009] In summary, existing power load forecasting technologies still have shortcomings in areas such as deep time-series modeling with multimodal fusion, efficient alignment mechanisms for asynchronous data, and cross-scenario model generalization capabilities. Therefore, there is an urgent need for a multimodal collaborative modeling method for power load that can deeply mine intra-modal time-series features, support adaptive alignment of asynchronous modes, and possess good cross-regional migration performance. This method aims to overcome the limitations of existing technologies and meet the refined scheduling needs of large-scale power systems. Summary of the Invention

[0010] This invention addresses the shortcomings of existing technologies in multimodal temporal modeling, asynchronous data alignment, and model generalization capabilities across scenarios by providing a dynamic multimodal temporal modeling and fusion method and system. First, this invention maps power load, meteorological images, and static equipment data, and fuses adaptive periodic codes generated by phase offset prediction to achieve temporal enhancement. Second, it constructs a temporal attention bias matrix based on relative sampling time relationships, injects it into a cross-attention mechanism to adjust the correlation strength, and then dynamically aggregates it through a gating mechanism to obtain a unified multimodal representation. Finally, within a federated learning framework, it introduces a dynamic alignment mechanism based on the maximum mean difference to calculate cross-regional distribution alignment loss, and combines this with a private module to independently update local parameters. This invention effectively overcomes multimodal temporal misalignment noise, achieving efficient collaboration and high-precision personalized prediction across regions while protecting data privacy.

[0011] The present invention adopts the following technical solution.

[0012] In a first aspect, the present invention provides a dynamic multimodal temporal modeling and fusion method, comprising: The time series sequence of the power load target mode, the image sequence of the meteorological image auxiliary mode, and the static equipment category data are acquired. Feature mapping is performed on each of them to obtain a power load target feature sequence, a meteorological image auxiliary feature sequence, and a static equipment feature vector with a unified feature dimension. Based on the power load target feature sequence and the meteorological image auxiliary feature sequence, the corresponding phase offset is predicted to generate an adaptive periodic code, which is then fused with the corresponding feature sequence to obtain a time-enhanced feature sequence. A time-related constraint matrix is ​​constructed based on the relative sampling time relationship between the power load target mode and the meteorological image auxiliary mode, and mapped to a time attention bias matrix. A query matrix, a key matrix, and a value matrix are generated based on the time-enhanced feature sequence. The time attention bias matrix is ​​directly injected as an additional term into the inner product calculation result of the query matrix and the key matrix to adjust the attention correlation strength and obtain the fusion feature branch corresponding to each meteorological image auxiliary mode. Centered on the power load target feature sequence, each fused feature branch is dynamically weighted and aggregated to obtain a unified multimodal time series representation; the unified multimodal time series representation is fused with the static equipment feature vector and input into the prediction network to obtain the load prediction result.

[0013] Preferably, feature mapping is performed separately to obtain a power load target feature sequence, a meteorological image auxiliary feature sequence, and a static equipment feature vector with a unified feature dimension, including: For the time series sequence of the power load target mode, a fully connected layer is used for linear projection to obtain a continuous vector sequence, which is used as the power load target feature sequence; For the image sequence of the meteorological image auxiliary mode, a two-dimensional image block embedding mechanism is introduced. A two-dimensional convolution kernel is used to extract local features from the single frame image in the sequence. The output local feature map is flattened in the spatial dimension and obtained by linear projection mapping. For the static device category data, it is parsed into multiple discrete attribute indices, and a corresponding learnable embedding matrix is ​​constructed for each discrete attribute index. The discrete attribute indexes are transformed into corresponding dense continuous vectors through table lookup operations. The dense continuous vectors corresponding to each attribute index are concatenated and fused into the static device feature vector through linear projection operations.

[0014] Preferably, the corresponding phase offset is predicted based on the power load target feature sequence and the meteorological image auxiliary feature sequence, respectively, and an adaptive periodic code is generated and fused with the corresponding feature sequence to obtain a time-enhanced feature sequence, including: For any feature sequence in the power load target feature sequence and the meteorological image auxiliary feature sequence, the input is given to the phase offset prediction module composed of a feedforward neural network to obtain the phase offset scalar corresponding to the feature sequence; The phase offset scalar is used as a phase parameter and substituted into a preset periodic sine and cosine function to calculate the adaptive periodic code corresponding to the feature sequence. The adaptive periodic encoding is added element-wise to the feature sequence to obtain the time-enhanced feature sequence.

[0015] Preferably, the relative sampling time relationship between the target power load mode and the meteorological image auxiliary mode is extracted to construct a time correlation constraint matrix and mapped to a time attention bias matrix, including: The first sampling timestamp of the power load target mode at each time step of the target sequence and the second sampling timestamp of the meteorological image auxiliary mode at each time step of the auxiliary sequence are extracted respectively. Calculate the relative timestamp difference between each of the first sampling timestamp and each of the second sampling timestamps, and arrange them according to the time step index of the target sequence and the auxiliary sequence to obtain the time correlation constraint matrix; The relative timestamp difference elements in the time-related constraint matrix are transformed into bias scalars through a nonlinear mapping function that includes sine and cosine coding and multilayer perceptron, thus forming the time attention bias matrix that matches the inner product dimension of the cross-attention calculation.

[0016] Preferably, taking the power load target feature sequence as the center, each fused feature branch is dynamically weighted and aggregated to obtain a unified multimodal time series representation, including: The target feature sequence of the power load is spliced ​​with each of the fused feature branches along the channel dimension to obtain the splicing tensor corresponding to each meteorological image auxiliary mode; Each of the splicing tensors is input into a gated network containing a fully connected layer and a nonlinear activation function, and the nonlinear activation function outputs the dynamic importance weights corresponding to each meteorological image auxiliary mode. After expanding the dynamic importance weights corresponding to each meteorological image auxiliary mode along the feature dimension using tensor broadcasting, they are multiplied element-wise with the corresponding fusion feature branches. All multiplied fusion feature branches are then added to the power load target feature sequence to obtain the unified multimodal time series representation.

[0017] Preferably, the prediction network is deployed in a federated learning framework that includes a central server and multiple clients. During joint training of the prediction network, a dynamic alignment mechanism based on the maximum mean difference is introduced to calculate the cross-regional feature distribution alignment loss, specifically including: Each client uses a clustering algorithm to cluster the meteorological image auxiliary feature sequence within the current time window, and uses the main cluster center as the meteorological prototype vector of the current time window and uploads it to the central server; The central server calculates the similarity between the received meteorological prototype vectors, determines the client node pairs with similarity greater than a preset threshold as collaborating pairs, and implements alignment blocking for client node pairs with similarity not greater than the preset threshold. For the cooperating pair, a monotonically decreasing confidence function is constructed based on the relative timestamp difference in the time association constraint matrix, and the asynchronous confidence factor is calculated. By combining the batch feature mean vector of each client in the collaborative pair with the asynchronous confidence factor, the weighted cross-regional feature distribution alignment loss is calculated on the central server, and the parameter gradient calculated based on the cross-regional feature distribution alignment loss is sent to the corresponding client for network parameter updates.

[0018] Preferably, the prediction network includes a shared module and a private module, wherein the private module is equipped with a bottleneck adapter, and the bottleneck adapter includes a dimension-reducing linear layer, a nonlinear activation layer and a dimension-increasing linear layer connected in sequence. During the local optimization phase of the joint training of the prediction network, each client, under the premise of freezing or fine-tuning the parameters of the shared module, performs local feature calibration by sequentially passing the feature vector output by the shared module through the dimension reduction linear layer, the nonlinear activation layer and the dimension increase linear layer. Each client constructs a joint loss function based on the local basic prediction task loss, the cross-regional feature distribution alignment loss, and the regularization constraint term for the private module parameters, and independently updates the network parameters of the local private module based on the joint loss function.

[0019] Secondly, the present invention provides a dynamic multimodal temporal modeling and fusion system, comprising: The multimodal feature mapping and temporal enhancement module is used to acquire the time series sequence of the power load target mode, the image sequence of the meteorological image auxiliary mode, and the static equipment category data, and perform feature mapping on them respectively to obtain the power load target feature sequence, the meteorological image auxiliary feature sequence, and the static equipment feature vector with a unified feature dimension; predict the corresponding phase offset based on the power load target feature sequence and the meteorological image auxiliary feature sequence respectively, generate adaptive periodic codes, and fuse them with the corresponding feature sequences to obtain the temporal enhancement feature sequence; The bias-aware cross-attention module is used to construct a time-related constraint matrix based on the relative sampling time relationship between the power load target mode and the meteorological image auxiliary mode, and map it into a time-attention bias matrix; it generates a query matrix, a key matrix, and a value matrix based on the time-enhanced feature sequence, and directly injects the time-attention bias matrix as an additional term into the inner product calculation result of the query matrix and the key matrix to adjust the attention correlation strength and obtain the fusion feature branch corresponding to each meteorological image auxiliary mode; The dynamic gating aggregation and prediction module is used to dynamically weight and aggregate each fused feature branch with the power load target feature sequence as the center to obtain a unified multimodal time series representation; the unified multimodal time series representation is fused with the static equipment feature vector and input into the prediction network to obtain the load prediction result.

[0020] Thirdly, the present invention provides a terminal, including a processor and a storage medium; the storage medium is used to store instructions. The processor is configured to operate according to the instructions to execute the steps of the method.

[0021] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the steps of the method.

[0022] The beneficial effects of this invention are compared with those of the prior art: 1. This invention employs a modality-aware adaptive periodic encoding mechanism, introducing a periodic phase calibration network oriented towards different modalities on top of absolute time encoding. This adaptively generates phase offsets corresponding to different periodic components, such as daily, weekly, and seasonal periods. This mechanism enables the periodic encoding of different modalities to express their advanced, synchronized, or lagging states relative to the global time reference by incorporating their own historical response characteristics. This alleviates the misalignment problem of multi-source data such as weather changes, power load fluctuations, and equipment status responses in the periodic semantic space, improving the representation accuracy of asynchronous multimodal time series data and the accuracy of subsequent cross-modal fusion.

[0023] 2. This invention utilizes a cross-attention mechanism with a relative time-aware bias to incorporate the timestamp difference between modalities as an attention bias in weight calculation during cross-modal interactions. This enables asynchronous temporal dependency modeling without the need for forced interpolation or hard alignment, thereby reducing information distortion and error accumulation caused by traditional time alignment.

[0024] 3. This invention uses cross-attention and adaptive gating aggregation mechanisms to dynamically weight and fuse the interaction features of modalities from different sources. This enables the model to adaptively adjust the fusion ratio based on the information contribution of each modality within the current time window, thereby suppressing redundant, noisy, or missing modalities and enhancing the stability and robustness of the multimodal fusion representation.

[0025] 4. This invention reconstructs a cross-regional collaborative model, combining meteorological prototype similarity gating and regional collaborative alignment strategies. While preserving the individual characteristics of each region, it promotes the sharing of effective knowledge among similar regions, achieving a balance between global rule transfer and local difference adaptation. This enhances the model's generalization ability, prediction stability, and collaborative training efficiency in multi-regional deployment scenarios. Attached Figure Description

[0026] Figure 1 This is an overall flowchart of a dynamic multimodal temporal modeling and fusion method provided by the present invention. Detailed Implementation

[0027] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of this invention. The embodiments described in this application are merely some embodiments of this invention, and not all embodiments. Based on the spirit of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of this invention.

[0028] Example 1: like Figure 1 As shown, this invention provides a dynamic multimodal temporal series modeling and fusion method, including: A: Obtain the time series sequence of the power load target mode, the image sequence of the meteorological image auxiliary mode, and the static equipment category data, and perform feature mapping on them respectively to obtain the power load target feature sequence, the meteorological image auxiliary feature sequence, and the static equipment feature vector with unified feature dimensions; predict the corresponding phase offset based on the power load target feature sequence and the meteorological image auxiliary feature sequence respectively, generate adaptive periodic coding, and fuse it with the corresponding feature sequence to obtain the time-enhanced feature sequence.

[0029] A1. First, map different forms of modal input to a high-dimensional vector space. Then, perform feature mapping for data of different modalities to obtain a power load target feature sequence, a meteorological image auxiliary feature sequence, and a static equipment feature vector with a unified feature dimension. For time-series data of power load target modes, i.e., numerical monitoring data continuously sampled over time, such as real-time voltage and current of substations, a fully connected layer is used for linear projection. Let the dimension of the input structured time-series data be... (in (original feature dimension), through dimension The weight matrix is ​​linearly transformed along the feature dimension, mapping it to a dimension of... A continuous vector sequence is used as the target feature sequence of the power load; For static device category data, i.e., discrete label data that lacks temporal evolution characteristics and characterizes the inherent physical attributes of the system, such as device model codes, voltage levels, and geographical topology nodes, this invention does not employ temporal network processing due to its lack of continuous time dependence. Instead, it constructs independent static representation branches. Let the number of input static category features be... (i.e., the input tensor dimension is) Since neural networks cannot directly process discrete symbols, this embodiment parses them into multiple discrete attribute indices and constructs a corresponding learnable embedding matrix for each discrete attribute index. The discrete attribute indices are transformed into dense continuous vectors through a lookup table operation, and then merged into a single vector with a fixed dimension through concatenation and linear projection operations. The static global feature vector is used as the static device feature vector.

[0030] For image sequences in meteorological image auxiliary modalities, which are continuously acquired visual data over time to capture the evolution of equipment appearance or meteorological changes over large areas, such as infrared thermal imaging temperature measurement videos and cloud image sequences, a two-dimensional image patch embedding mechanism is introduced to independently extract local features from each frame of the sequence. Let the input image sequence have a dimension of... ,in These represent the sequence length, height, width, and number of channels of the image, respectively, using a size of... Step size is The number of output channels is A two-dimensional convolutional kernel is used to extract features from a single frame of image. The output local feature map is then flattened in space to obtain a dimension of... The spatial feature tensor, in which The number of spatial feature blocks in a single frame; if or Unable to be If divisible, it is padded with zeros. After flattening, the meteorological image auxiliary feature sequence is obtained through linear projection mapping.

[0031] A2. After obtaining the embedding representations of each modality, for features with temporal continuity, namely the target feature sequence of power load and the auxiliary feature sequence of meteorological images, multiple time enhancement mechanisms are introduced to fully utilize temporal information and improve the time-series modeling capability of prediction. Specifically, these include: A2.1. Absolute Time Encoding: For the discrete calendar time information corresponding to each time step in the sequence, such as the hour range 0-23, weekdays 1-7, and holiday markers, similar to the processing method for the static categorical data mentioned above, this step constructs a specialized time feature learnable embedding matrix. Using this matrix, the discrete time attribute labels are mapped to low-dimensional dense vectors. After concatenation along the feature dimension, these vectors are linearly projected through a fully connected layer to generate an absolute time feature tensor with the same length as the input sequence.

[0032] Subsequently, differentiated time fusion strategies were implemented for different modalities: For the power load target feature sequence, the absolute time feature tensor and the structured feature tensor are directly concatenated along the feature channel dimension and then fused through a fully connected layer for dimensionality reduction. For the meteorological image auxiliary feature sequence containing multiple spatial feature blocks, a spatial dimension broadcasting mechanism is employed to... Copying the absolute time vector along the spatial dimension Next, it is expanded to match the spatial features of the image. Tensors are used for dimensionality reduction and fusion through a fully connected layer with a weight dimension of 2D×D, restoring the channel dimension to D. This mechanism not only endows each local visual patch with global temporal semantics, but also effectively avoids the microscopic physical features (such as local transient hotspots) caused by direct addition being masked by the temporal signal, enabling the model to perceive long-term trends and holiday effects without loss.

[0033] A2.2. Modality-aware adaptive periodic coding: For certain time-series data exhibiting distinct daily, weekly, or seasonal cycles, explicit injection of periodic patterns is achieved using sine / cosine functions or learnable embeddings. Considering the inherent hysteresis in the physical response of multimodal data, such as the lag in numerical temperature response due to equipment thermal inertia compared to meteorological image changes, this embodiment proposes a modality-aware periodic phase calibration mechanism to adaptively adjust the phase of the periodic encoding based on the historical response characteristics of different modalities.

[0034] Specifically, for any feature sequence fused in step A2.1, denoted as the k-th mode, the temporal feature vector of this mode that has not yet been fused after mapping in step A1 is first used as input. Then, the phase offset corresponding to this mode is generated through the phase calibration network, i.e., the phase offset prediction module composed of a feedforward neural network. The phase calibration network comprises two fully connected linear layers with a ReLU activation function in between. The output uses a tanh function or other bounded activation function to limit the phase shift within a preset range, thus avoiding excessive shifts that do not conform to actual periodic patterns. For different periodic components such as daily, weekly, and seasonal periods, corresponding phase shifts can be generated, enabling the model to express modal response hysteresis relationships at different periodic scales.

[0035] When generating periodic codes, the phase offset is added to the phase term of the corresponding periodic function, i.e. ,in For the current time step, angular frequency and initial phase The periodic encoding hyperparameters are pre-defined. This allows different modalities to express their respective response offset states in the periodic semantic space, thus mitigating the periodic misalignment problem caused by direct fusion of asynchronous multimodal data and improving the accuracy of subsequent cross-modal attention computation and multimodal fusion representation. The generated adaptive periodic encoding is then fused with the corresponding feature sequence (e.g., element-wise addition) to obtain the temporally enhanced feature sequence of the corresponding modality.

[0036] The time-enhanced feature sequences are used as input to subsequent temporal coding networks, enabling these networks to further learn the dynamic dependencies within a modality in a unified temporal semantic space.

[0037] A3. Deep Feature Temporal Modeling Structure: After completing the modality embedding and temporal augmentation described above, a specialized temporal modeling structure is used to further process the temporal augmentation feature sequences of the corresponding modalities, taking into account the characteristics of different modalities: A3.1. For the time-enhanced feature sequence corresponding to the power load target mode, a Transformer encoder is used for modeling. To ensure that the model strictly distinguishes the absolute order of features at different time steps within the current local window during self-attention calculation, a dimension is constructed. (in For sequence length, To unify the feature dimensions, a learnable location parameter matrix (or an absolute location encoding based on sine and cosine harmonic functions) is used. This sequence index encoding is then added element-wise to the time-enhanced feature sequence before being input into the self-attention module. Subsequently, the multi-head self-attention mechanism within the Transformer is used to compute the correlation weights between features at different time steps in parallel, thereby capturing long-term dependencies. Finally, the features are processed through residual connections and layer normalization, then input into a feedforward neural network, outputting a refined time-enhanced feature representation of the power load, which is used to generate the query matrix in subsequent steps.

[0038] A3.2. For the time-enhanced feature sequence corresponding to the meteorological image auxiliary mode, the dimension of the feature vector is changed from... Rearranged as That is, spatial and feature dimensions are used as batch and channel, and temporal dimension is used as sequence length. Then, a one-dimensional temporal convolutional network is introduced to address each independent spatial location (i.e., (any one of the local images), using a one-dimensional convolution kernel along the time dimension ( The feature map is slid along an axis to extract its temporal evolution across multiple consecutive frames. After temporal convolution, the spatial dimension of the feature map is... By keeping the original local spatial topology unchanged, the microscopic dynamic evolution law is successfully integrated into the high-dimensional features, and a deeply refined meteorological image temporal enhancement feature representation is output, which is used to generate the key matrix and value matrix in subsequent steps.

[0039] A3.3. Bypass processing: For the static device feature vector, since it does not have continuous temporal dependency characteristics, it is directly skipped from the temporal modeling network in this step and is used as an independent static feature branch to be input into the subsequent modal fusion stage.

[0040] B: Construct a time-related constraint matrix based on the relative sampling time relationship between the target power load mode and the meteorological image auxiliary mode, and map it to a time-attention bias matrix; generate a query matrix, a key matrix, and a value matrix based on the time-enhanced feature sequence, and directly inject the time-attention bias matrix as an additional term into the inner product calculation result of the query matrix and the key matrix to adjust the attention correlation strength, thereby obtaining the fusion feature branches corresponding to each meteorological image auxiliary mode; dynamically weight and aggregate each fusion feature branch with the target power load feature sequence as the center to obtain a unified multimodal time series representation; fuse the unified multimodal time series representation with the static equipment feature vector, input it into the prediction network, and obtain the load prediction result.

[0041] B1. Construction of the Relative Time Bias Matrix: Considering the asynchronous sampling characteristics of multimodal data, for example, the target primary mode's load sequence is sampled at high frequency, while the auxiliary mode's meteorological images are sampled at low frequency, the first sampling timestamp of the target power load mode (primary mode) at each time step of the time series is extracted, and the second sampling timestamp of the meteorological image auxiliary mode with time-series attributes at each time step of the image sequence is also extracted. For each auxiliary mode, the relative timestamp difference between each first sampling timestamp and each second sampling timestamp is calculated, and these differences are arranged according to the time step indices of the target sequence and the auxiliary sequences, combining them to obtain the time correlation constraint matrix. ,in The length of the main mode sequence. To determine the auxiliary modal sequence length, each relative time difference element in the matrix is ​​first expanded into a time feature vector using sine and cosine encoding, and then passed through a nonlinear mapping function incorporating a multilayer perceptron (MLP). This is transformed into a bias scalar, thus forming a time attention bias matrix with the same dimension as the attention inner product matrix. The elements in this matrix Characterizing the first dominant mode The time step and the i-th auxiliary mode Pure time distance decay or periodic correlation weight between time steps.

[0042] B2. Cross-attention mechanism with relative time-aware bias: Let the time-enhanced feature sequence corresponding to the target mode of the power load, output after deep refinement in step A3.1, be represented as follows: The time-enhanced feature sequence corresponding to the i-th meteorological image auxiliary mode output after refinement in step A3.2 is represented as follows: First, a Query, Key, and Value matrix is ​​generated using linear projection:

[0043]

[0044]

[0045] in The weight matrix is ​​a learnable matrix. For the dimension of attention head.

[0046] The time attention bias matrix constructed in B1 As an additional feature, it is directly injected into the inner product calculation result (attention Logits) of Query and Key, and the calculation formula is as follows:

[0047] This mechanism will use "feature semantic similarity" ")" and "physical time proximity" "Dynamic game" is played within the same attention probability space. Even if the features of a historical weather image in the auxiliary modality are semantically very similar to the current load, if it occurred too long ago ( (If the bias term is too large, it will be negative infinity). The Softmax mechanism will also adaptively weaken its attention weight, thereby solving the "temporal mismatch" problem caused by multimodal low-frequency asynchronous sampling.

[0048] Special Note: When processing image-assisted modalities, there is a dimensionality mismatch between the 3D spatiotemporal tensor and the 2D temporal tensor in the cross-attention mechanism. Therefore, a cross-modal spatiotemporal flattening mechanism is introduced. First, the image-assisted modality is flattened along the time and space dimensions into a format with dimensions of... The spatiotemporal long sequence is used as the key / value pair. Simultaneously, since each time step corresponds to S spatial feature blocks, the calculated... Time bias matrix along spatial dimension This involves replication and broadcasting, whereby the offset scalar at each time step is shared with all spatial locations at that time, ensuring that it aligns with... Inner product matrix Strict dimensional alignment. This mechanism enables fine-grained (pixel-level) cross-timestep addressing of target sequence features to auxiliary spatiotemporal features.

[0049] B3. Star-shaped cross-attention and adaptive gating aggregation of multi-source auxiliary modes: using the deep temporal augmentation feature representation corresponding to the power load target feature sequence. As the absolute interaction center node, it is respectively connected to the first... Each auxiliary modality performs the time-aware cross-attention calculation in step B2 above. If there exists If there is an auxiliary mode, then the output will be parallel. A separate fusion feature branch (All branch dimensions are) This design effectively avoids attention weight collapse caused by excessive differences in variance among different modal features.

[0050] After obtaining all independent fusion branches, considering that the contribution of each auxiliary mode to the prediction task changes dynamically at different times (for example, the image cloud map mode is more important during thunderstorms, while the historical load mode is more important during stable weather), this embodiment introduces a cross-modal gating network to represent the deep temporal augmentation features corresponding to the power load target feature sequence. With all fusion branches By splicing along the channels and using a fully connected layer and a sigmoid activation function, the dynamic importance weight of each auxiliary mode at the current time step is dynamically generated. The mathematical formula for its calculation is as follows:

[0051] In the formula, For the first A column vector of dynamic importance weights for each auxiliary mode; This represents the concatenation operation of the target mode and auxiliary mode of power load along the channel dimension, and the tensor dimension of the concatenated tensor is... ; and These represent the learnable weight matrix and bias term in the gated network, respectively. The dimension is (Used to compress channel dimensions into a single weight); The Sigmoid non-linear activation function is used to normalize the output value to... This allows for the soft shielding of noise modes and adaptive focusing on key modes within a specific range.

[0052] Finally, taking the power load target feature sequence as the center, the fused feature branches are dynamically weighted and aggregated, and the final multimodal time series unified representation is output by weighted summation. :

[0053] Among them, here This is a broadcast multiplication operation, that is, multiplying by a dimension of... Weight column vector Along feature dimension After copying and expanding, then with Feature tensor Perform element-wise multiplication.

[0054] B4. Fuse the multimodal time-series unified representation with the static equipment feature vector, input it into the prediction network, and obtain the load prediction result: Obtain the multimodal time-series unified representation of the power load target mode under time awareness (dimension: ...). After that, it is first reduced in dimensionality along the time axis by a mechanism of global average pooling or extracting the last feature block, and compressed into a one-dimensional dynamic feature vector (with dimension 1) representing the comprehensive dynamic evolution state within the current historical time window. ).

[0055] Subsequently, it is compared with the structured static categorical data that did not participate in temporal modeling (i.e., the independent static feature branches that did not participate in temporal modeling in step A3.3, with a dimension of...). Perform feature concatenation to form a dimension of The global context feature vector.

[0056] Finally, the global context feature vector is input into the prediction network (a deep feedforward neural network is used in this embodiment) for task decoding. This network contains multiple fully connected hidden layers, utilizing the hidden layer weight matrix for high-order nonlinear feature crossing. Between each hidden layer, a GELU nonlinear activation function is introduced to improve the network's fitting and expressive capabilities, and a Dropout regularization layer is forcibly interspersed (with a random inactivation probability set to a preset value). This effectively prevents the model from overfitting under the concatenation of heterogeneous features (dynamic and static features). After deep feature cross-validation, the last linear projection head of the prediction network maps the hidden layer features to the target prediction space, outputting the future power load prediction result.

[0057] C: Based on the load forecast results and the actual power load data, calculate the basic forecast task loss, combine the cross-regional feature distribution alignment loss and the regularization constraint term for the private module parameters, and construct a joint loss function; based on the joint loss function, update the network parameters through the backpropagation algorithm.

[0058] C1. Predictive network architecture under the federated learning framework: To achieve cross-regional collaborative modeling and localized optimization, the prediction network is deployed in a federated learning framework comprising a central server and multiple clients, and the network is divided into shared and private modules. The shared module includes an intra-modal temporal modeling layer and a cross-attention module, used to learn temporal dependencies and modal interactions common across different regions. The private module is equipped with a bottleneck adapter to capture region-specific features. The bottleneck adapter consists of a sequentially connected dimensionality-reducing linear layer, a non-linear activation layer, and an dimensionality-increasing linear layer. During the local optimization phase, each client sequentially passes the feature vector output from the shared module through: 1. Dimensionality Compression: A linear dimensionality reduction layer is used to reduce the dimension of the shared module's output to [value missing]. Projecting feature vectors into a low-dimensional bottleneck space (in This allows for the filtering out of redundancy in global common information and focuses on region-specific characteristics. 2. Nonlinear mapping: Feature enhancement is performed through activation layers (such as ReLU or GELU); 3. Dimension Restoration: Finally, a linear layer is used to restore the features to their original dimensions. .

[0059] This "compression-activation-recovery" structure achieves localized calibration of global features with a very small number of learnable parameters (far smaller than the number of parameters in the shared layer), ensuring cross-regional knowledge collaboration while avoiding overfitting caused by insufficient local data.

[0060] C2. Dynamic alignment mechanism based on maximum mean difference: To avoid forced aggregation model bias caused by differences in data distribution among different clients, a dynamic alignment mechanism based on the maximum mean difference is introduced in the shared feature space to calculate the cross-regional feature distribution alignment loss. The specific steps are as follows: C2.1. Conditional Trigger Alignment Strategy Based on Meteorological Prototypes: To avoid negative model transfer caused by forced alignment under different climatic backgrounds (e.g., region A experiencing heavy rain while region B experiences extreme heat), a meteorological prototype similarity gating is introduced before calculating the maximum mean difference. Specifically, each client uses the image features (e.g., meteorological cloud images) or structured category features (e.g., weather tags) mapped in step A to cluster the feature vectors within the current time window. In this embodiment, the K-Means algorithm is preferably used to obtain K cluster centers, and the main cluster center is used as the meteorological prototype vector for the current time window and uploaded to the central server. The central server calculates the similarity between the received meteorological prototype vectors, and selects those with similarity greater than a preset threshold. Client node pairs are determined to be "collaborative pairs" if their similarity is no greater than 1. The client nodes implement alignment blocking. This strategy achieves precise collaboration by "aligning similar nodes and decoupling dissimilar nodes".

[0061] C2.2 Asynchronous Perception-Driven Dynamic MMD Distribution Alignment: To mitigate feature alignment noise caused by the highly irregular sampling of low-frequency auxiliary modes, an asynchronous confidence factor was introduced. For the aforementioned collaborative pair, based on the relative timestamp difference in the time association constraint matrix... Define a monotonically decreasing confidence function. In this embodiment, a negative exponential decay function is preferred to calculate the asynchronous confidence factor.

[0062] Each client inputs multimodal data into its local shared layer to calculate the statistical expectation (i.e., mean vector) of the current batch feature distribution. (Combining the batch feature mean vectors of each client in the collaborative pair) and Combined with asynchronous confidence factor The weighted cross-regional feature distribution alignment loss is centrally calculated on the central server:

[0063] In the formula, It is a kernel function (in this embodiment, the Gaussian RBF kernel function is preferred). This mechanism ensures that the model applies a strong distributional alignment penalty only to valid features with high time alignment quality and high confidence, thereby avoiding noisy alignment.

[0064] C3. Personalized local optimization of the private layer: The personalized optimization mechanism primarily targets the training process of the private layer. Each client, with the shared layer parameters frozen or fine-tuned at a low learning rate, independently iterates the private prediction head using local, real-world load-label data. The feature vector output from the shared module is then sequentially passed through the dimensionality-reducing linear layer, the non-linear activation layer, and the dimensionality-increasing linear layer for local feature calibration. Clients can flexibly adjust the lightweight structure or training strategy of the private layer (such as local learning rate or early stopping mechanism) based on the scale of their local data.

[0065] C4. Joint Loss Function and Parameter Update: During the overall training process, the joint loss function for each client can be defined as:

[0066] in, The basic forecast task loss (such as mean square error MSE or mean absolute error MAE) is used to measure the deviation between the final load forecast and the actual value. The above is a cross-regional feature distribution alignment loss that incorporates asynchronous confidence. Regularization constraints (such as L2 weight decay) are applied to the parameters of the private module to prevent the local private layer from overfitting in small samples. and This is a weighted hyperparameter used to dynamically balance "local task prediction accuracy", "cross-regional knowledge consistency", and "local personalized anti-overfitting". The preferred value range is: , ; During reverse propagation, the central server will be based on The calculated gradients of the shared layer parameters are sent to each client; each client compares the received global gradients with its local gradients. and The calculated local gradients are weighted and fused, and the network parameters of the local private module are independently updated based on the joint loss function.

[0067] Example 2: In the fields of power marketing and grid dispatching, load forecasting is a crucial foundation for achieving power supply-demand balance and designing differentiated electricity pricing strategies. With the diversification of the power customer base and the uncertainty of load fluctuations, traditional forecasting methods relying on single historical load sequences are insufficient to meet the high-precision forecasting needs in complex scenarios. Therefore, this invention introduces multimodal time-series data mining and dynamic collaborative modeling techniques that can comprehensively utilize multi-source heterogeneous data such as customer electricity consumption history, meteorological environment, and holiday activities to capture the dynamic dependencies between multiple modalities. The following lists specific applications of this invention in different typical scenarios: Application 1: High-frequency composite load forecasting in residential areas In residential scenarios, the multimodal data required for prediction mainly includes high-frequency collected smart meter data (such as current, voltage, and active power), local meteorological data (such as temperature, humidity, and satellite cloud images), and holiday tag information.

[0068] To address the characteristics of high-frequency and naturally easy-to-align sensor data acquisition in this scenario, the dynamic multimodal time-series modeling method of this invention exhibits strong adaptive robustness. Specifically, when the timestamps of the collected power operation data and meteorological numerical data are highly consistent (i.e., the relative timestamp difference is approximately 0), the cross-attention mechanism with relative time-aware bias of this invention can adaptively output an attention bias term approaching neutrality, seamlessly processing high-frequency synchronous data without manual judgment or modification of the network structure. Simultaneously, the adaptive gating aggregation mechanism can automatically assign greater fusion weights to high-frequency and high-confidence modalities based on their feature performance within the current time window. Compared to the computational redundancy caused by forced interpolation in traditional preprocessing, the method of this invention, when processing multi-source data from residential areas with mixed synchronous and asynchronous data, not only avoids manual intervention but also eliminates local feature distortions that may be introduced by interpolation, ensuring the microscopic accuracy of load forecasting.

[0069] Application 2: Cross-regional load pattern migration based on federated learning In power dispatching scenarios involving multiple cities or large regions, there is often a lack of historical load data for newly connected areas (such as city B) and difficulties in cold-starting models. This invention achieves efficient model transfer and personalized adaptation across regions through a federated learning framework.

[0070] During the migration process, mature regions (such as City A) and newly accessed regions (such as City B) are connected to the central server as different client nodes. A shared module captures common cross-regional temporal dependencies and modal interaction patterns, and a dynamic MMD alignment mechanism based on meteorological prototypes is used in the cloud to mitigate the negative migration risks caused by differences in climate and electricity consumption habits between the two locations. For the localization adaptation of City B, a bottleneck adapter in its private module is used. Only a very small amount of actual electricity consumption data collected locally in City B is needed to quickly fine-tune lightweight parameters to complete the local calibration of global features. Through the above-mentioned federated collaboration and private fine-tuning mechanism, City B does not need to train a large multimodal network from scratch, nor does it need to upload sensitive local load data. This significantly reduces the modeling cost of the new region while ensuring data privacy and improving the generalization capability of cross-regional scheduling.

[0071] Example 3: This invention provides a terminal, including a processor and a storage medium; the storage medium is used to store instructions; The processor is configured to operate according to the instructions to execute the steps of the method.

[0072] Example 4: The present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method.

[0073] This disclosure can be a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for causing a processor to implement various aspects of this disclosure.

[0074] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example—but not limited to—electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination of the foregoing. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.

[0075] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.

[0076] Computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, etc., and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing the status information of the computer-readable program instructions to implement various aspects of this disclosure.

[0077] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the protection scope of the claims of the present invention.

Claims

1. A dynamic multimodal temporal series modeling and fusion method, characterized in that, include: The time series sequence of the power load target mode, the image sequence of the meteorological image auxiliary mode, and the static equipment category data are acquired. Feature mapping is performed on each of them to obtain a power load target feature sequence, a meteorological image auxiliary feature sequence, and a static equipment feature vector with a unified feature dimension. Based on the power load target feature sequence and the meteorological image auxiliary feature sequence, the corresponding phase offset is predicted to generate an adaptive periodic code, which is then fused with the corresponding feature sequence to obtain a time-enhanced feature sequence. A time-related constraint matrix is ​​constructed based on the relative sampling time relationship between the target power load mode and the meteorological image auxiliary mode, and then mapped to a time attention bias matrix. Based on the time-enhanced feature sequence, a query matrix, a key matrix, and a value matrix are generated. The time attention bias matrix is ​​directly injected as an additional term into the inner product calculation result of the query matrix and the key matrix to adjust the attention association strength and obtain the fusion feature branch corresponding to each meteorological image auxiliary mode. Centered on the power load target feature sequence, each fused feature branch is dynamically weighted and aggregated to obtain a unified multimodal time series representation; the unified multimodal time series representation is fused with the static equipment feature vector and input into the prediction network to obtain the load prediction result.

2. The dynamic multimodal temporal series modeling and fusion method according to claim 1, characterized in that, Feature mapping is performed separately to obtain a power load target feature sequence, a meteorological image auxiliary feature sequence, and a static equipment feature vector with unified feature dimensions, including: For the time series sequence of the power load target mode, a fully connected layer is used for linear projection to obtain a continuous vector sequence, which is used as the power load target feature sequence; For the image sequence of the meteorological image auxiliary mode, a two-dimensional image block embedding mechanism is introduced. A two-dimensional convolution kernel is used to extract local features from the single frame image in the sequence. The output local feature map is flattened in the spatial dimension and obtained by linear projection mapping. For the static device category data, it is parsed into multiple discrete attribute indices, and a corresponding learnable embedding matrix is ​​constructed for each discrete attribute index. The discrete attribute indexes are transformed into corresponding dense continuous vectors through table lookup operations. The dense continuous vectors corresponding to each attribute index are concatenated and fused into the static device feature vector through linear projection operations.

3. The dynamic multimodal temporal modeling and fusion method according to claim 2, characterized in that, Based on the target feature sequence of power load and the auxiliary feature sequence of meteorological images, the corresponding phase offset is predicted, an adaptive periodic code is generated and fused with the corresponding feature sequence to obtain a time-enhanced feature sequence, including: For any feature sequence in the power load target feature sequence and the meteorological image auxiliary feature sequence, the input is given to the phase offset prediction module composed of a feedforward neural network to obtain the phase offset scalar corresponding to the feature sequence; The phase offset scalar is used as a phase parameter and substituted into a preset periodic sine and cosine function to calculate the adaptive periodic code corresponding to the feature sequence. The adaptive periodic encoding is added element-wise to the feature sequence to obtain the time-enhanced feature sequence.

4. The dynamic multimodal temporal series modeling and fusion method according to claim 1, characterized in that, Extracting the relative sampling time relationship between the target power load mode and the meteorological image auxiliary mode to construct a time correlation constraint matrix and mapping it to a time attention bias matrix includes: The first sampling timestamp of the power load target mode at each time step of the target sequence and the second sampling timestamp of the meteorological image auxiliary mode at each time step of the auxiliary sequence are extracted respectively. Calculate the relative timestamp difference between each of the first sampling timestamp and each of the second sampling timestamps, and arrange them according to the time step index of the target sequence and the auxiliary sequence to obtain the time correlation constraint matrix; The relative timestamp difference elements in the time-related constraint matrix are transformed into bias scalars through a nonlinear mapping function that includes sine and cosine coding and multilayer perceptron, thus forming the time attention bias matrix that matches the inner product dimension of the cross-attention calculation.

5. The dynamic multimodal temporal modeling and fusion method according to claim 4, characterized in that, Centered on the power load target feature sequence, dynamic weighted aggregation is performed on each fused feature branch to obtain a unified multimodal time series representation, including: The target feature sequence of the power load is spliced ​​with each of the fused feature branches along the channel dimension to obtain the splicing tensor corresponding to each meteorological image auxiliary mode; Each of the splicing tensors is input into a gated network containing a fully connected layer and a nonlinear activation function, and the nonlinear activation function outputs the dynamic importance weights corresponding to each meteorological image auxiliary mode. After expanding the dynamic importance weights corresponding to each meteorological image auxiliary mode along the feature dimension using tensor broadcasting, they are multiplied element-wise with the corresponding fusion feature branches. All multiplied fusion feature branches are then added to the power load target feature sequence to obtain the unified multimodal time series representation.

6. The dynamic multimodal temporal series modeling and fusion method according to claim 5, characterized in that, The prediction network is deployed in a federated learning framework that includes a central server and multiple clients. During joint training of the prediction network, a dynamic alignment mechanism based on the maximum mean difference is introduced to calculate the cross-regional feature distribution alignment loss, specifically including: Each client uses a clustering algorithm to cluster the meteorological image auxiliary feature sequence within the current time window, and uses the main cluster center as the meteorological prototype vector of the current time window and uploads it to the central server; The central server calculates the similarity between the received meteorological prototype vectors, determines the client node pairs with similarity greater than a preset threshold as collaborating pairs, and implements alignment blocking for client node pairs with similarity not greater than the preset threshold. For the cooperating pair, a monotonically decreasing confidence function is constructed based on the relative timestamp difference in the time association constraint matrix, and the asynchronous confidence factor is calculated. By combining the batch feature mean vector of each client in the collaborative pair with the asynchronous confidence factor, the weighted cross-regional feature distribution alignment loss is calculated on the central server, and the parameter gradient calculated based on the cross-regional feature distribution alignment loss is sent to the corresponding client for network parameter updates.

7. The dynamic multimodal temporal series modeling and fusion method according to claim 6, characterized in that, The prediction network includes a shared module and a private module. The private module is equipped with a bottleneck adapter, which includes a dimension-reducing linear layer, a nonlinear activation layer and a dimension-increasing linear layer connected in sequence. During the local optimization phase of the joint training of the prediction network, the feature vector output by the shared module is sequentially passed through the dimension reduction linear layer, the nonlinear activation layer, and the dimension increase linear layer for local feature calibration. Each client constructs a joint loss function based on the local basic prediction task loss, the cross-regional feature distribution alignment loss, and the regularization constraint term for the private module parameters, and independently updates the network parameters of the local private module based on the joint loss function.

8. A dynamic multimodal temporal series modeling and fusion system, characterized in that, include: The multimodal feature mapping and temporal enhancement module is used to acquire the time series sequence of the power load target mode, the image sequence of the meteorological image auxiliary mode, and the static equipment category data, and perform feature mapping on them respectively to obtain the power load target feature sequence, the meteorological image auxiliary feature sequence, and the static equipment feature vector with a unified feature dimension; predict the corresponding phase offset based on the power load target feature sequence and the meteorological image auxiliary feature sequence respectively, generate adaptive periodic codes, and fuse them with the corresponding feature sequences to obtain the temporal enhancement feature sequence; The bias-aware cross-attention module is used to construct a time-related constraint matrix based on the relative sampling time relationship between the power load target mode and the meteorological image auxiliary mode, and map it into a time-attention bias matrix; Based on the time-enhanced feature sequence, a query matrix, a key matrix, and a value matrix are generated. The time attention bias matrix is ​​directly injected as an additional term into the inner product calculation result of the query matrix and the key matrix to adjust the attention association strength and obtain the fusion feature branch corresponding to each meteorological image auxiliary mode. The dynamic gating aggregation and prediction module is used to dynamically weight and aggregate each fused feature branch with the power load target feature sequence as the center to obtain a unified multimodal time series representation; the unified multimodal time series representation is fused with the static equipment feature vector and input into the prediction network to obtain the load prediction result.

9. A terminal, comprising a processor and a storage medium; characterized in that: The storage medium is used to store instructions; The processor is configured to operate according to the instructions to perform the steps of the method according to any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the steps of the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Power load prediction model and method based on multi-modal fusion

    CN119227859A