A new energy missing data completion method and system based on multi-model fusion
Patent Information
- Application Number
- CN202511008371.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-22
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2045-07-22
AI Technical Summary
[0005]发明目的:针对新能源发电场景中数据缺失率高、时序依赖复杂、传统补全方法准确性不足等问题,本发明的目的在于提供一种基于多模型融合的新能源缺失数据补全方法与系统,通过对输入数据进行线性、非线性特征筛选,通过自适应频域特征提取、多粒度时序建模等策略来学习新能源异构数据特征,提高模型对新能源缺失数据补全的准确率和系统鲁棒性
[0030]有益效果:本发明提供的一种基于多模型融合的新能源缺失数据补全方法,设计的特征筛选方案能有效减少输入数据的维度、降低输入的复杂度,同时增强了模型对特征之间非线性关系的捕捉能力,增强模型的学习能力;本发明设计的自适应频域特征提取模块利用数据驱动的主频提取与频谱抑噪机制,增强了模型对长序列主模态结构的建模能力,提高了对高波动性数据的鲁棒性;所构建的多粒度时序编码模块通过多尺度建模时间序列的短期特征与长期趋势,并通过可学习权重与自适应对齐机制实现语义一致的特征融合,进一步提升模型在复杂时序结构下的泛化能力,能够在单次推理内有效实现对8小时之间缺失数据的补全。
Smart Images

Figure CN120910415B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of new energy missing data completion, and in particular to a method and system for new energy missing data completion based on multi-model fusion. Background Technology
[0002] In new energy power generation systems, data acquisition and monitoring are core components for ensuring safe system operation and energy efficiency optimization. Key parameters such as the bearing temperature of wind turbine generators and the component temperature of photovoltaic arrays are closely related to the power generation efficiency, equipment lifespan, and failure risk of new energy equipment. Because new energy power plants are often located in complex natural environments (such as sand erosion in desert photovoltaic areas and salt spray corrosion in offshore wind farms), sensor failures and communication interruptions frequently lead to random and continuous data gaps in monitoring data. Completing the operational data of new energy equipment can provide important references for operational status assessment and fault early warning.
[0003] New energy missing data completion is a time series interpolation task that involves analyzing past operational data of new energy equipment over time to fill in missing values for certain points in time or time periods. Traditional data completion methods struggle to capture the nonlinear characteristics of new energy data. KNN interpolation is prone to introducing spurious fluctuations when dealing with the spatiotemporal correlations of wind power data. While LSTM-based completion models can capture time-series dependencies, they lack the ability to model cross-modal interactions in multi-source heterogeneous new energy data.
[0004] Current technology faces three core challenges: First, new energy data is highly complex, with intricate interrelationships between different sensors. Traditional Pearson correlation coefficients ignore numerous nonlinear features, leading to the failure of feature association modeling. Second, new energy data is typically long-term series data, where current data often depends on data from a long historical period. Traditional completion models struggle to effectively capture these long-term dependencies, resulting in inaccurate completion results. Finally, due to the instability of energy sources (solar, wind, etc.), new energy data exhibits high volatility, significantly interfering with model feature learning. A major research challenge is how to consider nonlinear features in new energy data completion tasks, fully extract multi-source heterogeneous data features from new energy units (time, multiple sensors, etc.), reduce data volatility, and improve model accuracy and generalization ability. Summary of the Invention
[0005] Purpose of the invention: To address the problems of high data missing rate, complex time-series dependencies, and insufficient accuracy of traditional data completion methods in new energy power generation scenarios, the present invention aims to provide a new energy missing data completion method and system based on multi-model fusion. By performing linear and nonlinear feature filtering on the input data, and by using strategies such as adaptive frequency domain feature extraction and multi-granularity time-series modeling, the invention learns the characteristics of heterogeneous new energy data, thereby improving the accuracy of the model in completing new energy missing data and the robustness of the system.
[0006] Technical solution: To achieve the above-mentioned objectives, the present invention adopts the following technical solution:
[0007] In a first aspect, the present invention provides a method for completing missing new energy data based on multi-model fusion, comprising the following steps:
[0008] The collected new energy datasets were subjected to two-stage feature screening based on linear and nonlinear correlations.
[0009] The dataset is normalized, a missing data mask is constructed based on the missing proportion parameter, and the dataset is divided.
[0010] A fusion model, AFMFormer (Adaptive Frequency-aware Multi-scale Transformer), is constructed, comprising an adaptive frequency-domain feature extraction module, a multi-granularity temporal coding module, and an adaptive fusion completion module. The adaptive frequency-domain feature extraction module, based on an improved Fourier transform design, possesses frequency selectivity and multi-resolution sensing capabilities, achieving adaptive spectral decomposition to extract dominant frequency components from the sequence and suppress high-frequency noise. The multi-granularity temporal coding module includes two parallel coding paths: one is a patch-based Transformer path, which divides the time series into several time windows and performs local attention modeling; the other is a Standard Transformer path, used to extract trend information and overall dynamic changes in the sequence. The adaptive fusion completion module aligns the output features of different paths in the time dimension and dynamically weights and fuses the output features of the two coding paths.
[0011] Train the AFMFormer model, use the trained model to impute missing data, and output the results of the missing data completion.
[0012] Furthermore, the feature selection includes: calculating the Pearson correlation coefficient and the maximum mutual information coefficient among the features in the collected new energy data; if the correlation between any two features does not reach a preset threshold, the data is considered redundant, and features with low correlation to the target variable are ignored.
[0013] Furthermore, the adaptive frequency domain feature extraction module includes an adaptive frequency domain enhancement block and a cross-scale feature intersection block;
[0014] The adaptive frequency domain enhancement block introduces a learnable threshold parameter after performing a fast Fourier transform on the input value to filter high-energy regions in the frequency domain; after filtering out the main frequency domain information, adaptive spectrum enhancement is performed through local and global dual-channel learning filters.
[0015] The cross-scale feature intersection block inputs the adaptively frequency-domain enhanced data into two parallel convolutional layers with different kernel sizes. The outputs of the two convolutional layers interact through element-wise multiplication to enhance the feature representation. The two interacting features are then fused and processed by a fusion convolutional layer to obtain the final output feature.
[0016] Furthermore, the calculation process of the adaptive frequency domain enhancement block is as follows:
[0017]
[0018] P = ||F|| 2
[0019] F filtered =F⊙(P>θ)
[0020] F enhanced =W G ⊙F+W L ⊙F filtered
[0021]
[0022] Where, χ t For the input observations, Let F represent the Fast Fourier Transform, P be the power spectrum, and θ be a trainable threshold. filtered W is the filtered frequency domain representation. G For a global filter, W L For a local filter, χ t ' is the time series representation obtained by the inverse fast Fourier transform, N is the number of variables, T is the length of the input sequence, and ||·|| 2 The frequency domain representation represents the squared amplitude of each frequency component of F, and ⊙ represents element-wise multiplication.
[0023] Furthermore, in the multi-granularity time-series coding module, the Patch-based Transformer path first uses a Patch-based Embedding layer to process the input time-series data. Divide into multiple local patches. The data is mapped to a high-dimensional space through a linear layer, where B is the batch size, T is the sequence length, N is the number of variables, and L is the length of each block; then, learnable positional encodings are added to model the sequential relationships between patches, resulting in the representation. Where d is the embedding dimension; the high-dimensional representation is then encoded by a Patch-based Encoder; the Patch-based Encoder contains multiple stacked Patch-based Transformer encoder modules, each of which consists of four parts: self-attention mechanism, feedforward network, residual connection and one-dimensional batch normalization. One-dimensional batch normalization transposes and expands the input in the time dimension and applies one-dimensional batch normalization to each feature channel to achieve standardization processing across time steps.
[0024] The Standard Transformer path first uses a Standard Embedding layer to map the original input from the time dimension to a high-dimensional space through a linear layer; then, it encodes the high-dimensional space representation through a Standard Encoder. The Standard Encoder contains multiple stacked Standard Transformer encoder modules, each of which consists of four parts: a self-attention mechanism, a feedforward network, residual connections, and standard layer normalization. The multi-granularity temporal coding module encodes the input in parallel through the above two paths.
[0025] Furthermore, the adaptive fusion completion module uses an adaptive average pooling operation to uniformly divide the sequence obtained from the StandardTransformer path and calculates the mean in each sub-interval, uniformly adjusting the output of the StandardTransformer path to the sequence length of the Patch-based Transformer path output. Based on this, a learnable fusion weight parameter is introduced to adaptively adjust the contribution of the two aligned encoding paths to the final completion result, and then the feature representation is fused through this parameter. Finally, a linear mapping head is used to flatten and linearly map the fused features, restoring the high-dimensional features to predicted values in sequence form.
[0026] Secondly, this invention provides a new energy missing data completion system based on multi-model fusion, comprising: a feature selection module for performing two-stage feature selection on the collected new energy dataset using linear and nonlinear correlation; a data preprocessing module for normalizing the dataset, constructing a missing data mask based on the missing proportion parameter, and dividing the dataset; a model building module for constructing a fusion model AFMFormer including an adaptive frequency domain feature extraction module, a multi-granularity time-series coding module, and an adaptive fusion completion module; wherein the adaptive frequency domain feature extraction module, based on an improved Fourier transform design, has frequency selectivity and multi-resolution sensing capabilities, achieving adaptive spectral decomposition, extracting the dominant frequency components in the sequence, and suppressing high-frequency noise; the multi-granularity time-series coding module includes two parallel coding paths: one is a Patch-based Transformer path, which divides the time series into several time windows and performs local attention modeling; the other is a Standard... The Transformer path is used to extract trend information and overall dynamic changes in the sequence; the adaptive fusion completion module is used to align the output features of different paths in the time dimension and dynamically weight and fuse the output features of the two encoded paths; and the model training module is used to train the AFMFormer model, use the trained model to impute missing data, and output the missing data completion results.
[0027] Thirdly, the present invention provides a computer system, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the aforementioned method for completing missing new energy data based on multi-model fusion.
[0028] Fourthly, the present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the method for completing missing new energy data based on multi-model fusion.
[0029] Fifthly, the present invention provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the aforementioned method for completing missing new energy data based on multi-model fusion.
[0030] Beneficial Effects: This invention provides a new energy missing data completion method based on multi-model fusion. The designed feature selection scheme can effectively reduce the dimensionality of input data and reduce input complexity, while enhancing the model's ability to capture nonlinear relationships between features and improving the model's learning ability. The adaptive frequency domain feature extraction module designed in this invention utilizes data-driven main frequency extraction and spectrum noise reduction mechanisms to enhance the model's ability to model the main modal structure of long sequences and improve robustness to highly volatile data. The constructed multi-granularity time series coding module models the short-term features and long-term trends of time series at multiple scales and achieves semantically consistent feature fusion through learnable weights and adaptive alignment mechanisms, further improving the model's generalization ability under complex time series structures and enabling effective completion of missing data between 8 hours in a single inference. Attached Figure Description
[0031] Figure 1 This is a schematic diagram of the overall process of an embodiment of the present invention.
[0032] Figure 2 This is a schematic diagram of the feature selection process in an embodiment of the present invention.
[0033] Figure 3 This is a heatmap of the Pearson correlation coefficient of the dataset in this embodiment of the invention.
[0034] Figure 4 This is a heatmap of the maximum mutual information coefficient of the dataset in this embodiment of the invention.
[0035] Figure 5 This is a schematic diagram of the overall structure of the model in an embodiment of the present invention.
[0036] Figure 6 This is a schematic diagram of the structure of the adaptive frequency domain enhancement block in an embodiment of the present invention.
[0037] Figure 7 This is a schematic diagram of the structure of the cross-scale feature interaction block in an embodiment of the present invention. Detailed Implementation
[0038] The technical solution of the present invention will be further described below with reference to the accompanying drawings and specific embodiments.
[0039] like Figure 1 As shown in the figure, an embodiment of the present invention discloses a method for completing missing new energy data based on multi-model fusion, which includes the following steps:
[0040] (1) Perform linear and nonlinear two-stage feature screening on the collected new energy datasets (such as wind turbine datasets) to select feature data that are strongly correlated with the target variable (such as wind power).
[0041] Specifically, this embodiment uses the Pearson correlation coefficient to screen features with high linear correlation to the target variable, and the maximum mutual information coefficient to screen features with high non-linear correlation to the target variable. The Pearson correlation coefficient and maximum mutual information coefficient can be calculated among all sensor features. If the correlation between any two features does not reach a preset threshold, the data is considered redundant, and features with low correlation to the target variable are ignored.
[0042] The following uses a publicly available dataset as an example to explain in detail the feature selection process of this embodiment. This dataset corresponds to the Beberibe Wind Farm (UEBB) located in the northeastern coastal region of Brazil. The SCADA system sampling interval is 10 minutes. All 32 Enercon E-48 wind turbines in the wind farm are equipped with 100-meter metering masts conforming to IEC standards, featuring five layers of top-quality calibrated cup anemometers and one layer (100 meters) of three-dimensional acoustic anemometers. The dataset contains 3 types of time and location information, 16 types of wind speed and direction information, 6 types of meteorological condition information, 11 types of equipment operating status information, and 6 other auxiliary information. For example... Figure 2 As shown, feature selection specifically includes:
[0043] (1.1) Linear Correlation Feature Screening: Calculate the Pearson Correlation Coefficient (PCC) between features, setting a threshold of 0.8. If the linear correlation between any two features exceeds the threshold, the data is considered redundant, and features with low correlation to the target variable are ignored. Assuming there are variables X and Y, the Pearson correlation coefficient between X and Y is calculated as follows:
[0044]
[0045] Where, x i and y i Let X and Y be the values of the i-th samples, respectively, and n be the number of samples. and These are the sample means for X and Y, respectively.
[0046] The Pearson correlation coefficient heatmap of the wind turbine dataset is as follows: Figure 3 As shown in Table 1, the features obtained after linear feature filtering are as follows:
[0047] Table 1 Results of Linear Feature Selection
[0048]
[0049]
[0050] (1.2) Nonlinear Correlation Feature Screening: For the features in Table 1 that have undergone linear feature screening, calculate the maximum mutual information coefficient between the features. Set the threshold to 0.7. If the nonlinear correlation between any two features exceeds the threshold, the data is considered redundant, and features with low correlation to the target variable are ignored. Assuming there are variables X and Y, the mutual information between X and Y is calculated as follows:
[0051]
[0052] in, and Let X and Y be the value spaces of X and Y respectively, P(x) and P(y) be the marginal probability distributions of X and Y respectively, and P(x,y) be the joint probability distribution of X and Y.
[0053] Given variables X and Y, and a sample set For any mesh partition G (dividing the xy plane into x-directions x... bin An interval, y in the y direction bin The mutual information of the samples within this grid division is I(G). The maximum mutual information coefficient of the samples is calculated as follows:
[0054]
[0055] Here, B is an upper bound related to the sample size n, usually taken as B = n. α (α∈(0.6,0.8)).
[0056] The heatmap of the maximum mutual information coefficient of the wind turbine generator dataset is as follows: Figure 4 As shown in Table 2, the features obtained after nonlinear feature filtering are as follows:
[0057] Table 2 Results of Nonlinear Feature Selection
[0058] wind_speed_max 0.7492 rotor_rpm_max 0.7994 wind_speed_cube 1.0000 rotor_rpm_min 0.7119 sonic_wind_speed 0.9712 active_power_total 0.8595 wind_speed_nacelle 0.8418 active_power_total_max 0.8050 wind_speed_nacelle_max 0.8243 active_power_total_min 0.7116 rotor_rpm 0.8636 wind_speed 1.0000
[0059] (2) Normalize the dataset, construct a missing data mask based on the missing proportion parameter, and divide the dataset. In the dataset, the input is the values of multiple missing features of the wind turbine over a period of time, and the expected output is the complete value of the missing data (such as wind speed) during that period.
[0060] In this embodiment, step (2) specifically includes:
[0061] (2.1) Dataset normalization: Standard deviation normalization is used to process the data, transforming it into a standard normal distribution with a mean of 0 and a standard deviation of 1, thus eliminating the influence of the data's dimensions. The formula for calculating standard deviation normalization is as follows:
[0062]
[0063] Where μ is the mean of the dataset and σ is the standard deviation of the dataset;
[0064] (2.2) Construct a missing data mask based on the missing proportion parameter. Create an array of all 1s according to the shape of the data to represent that all data are present. Then, according to the missing proportion parameter, randomly select the corresponding proportion of element positions in the array and set the values at these positions to 0 to represent missing data, thus obtaining the missing data mask. In this experiment, the missing proportion ε = 0.2 is used.
[0065] (2.3) Dataset partitioning: The dataset and its corresponding mask are partitioned into windows with a time step of 50. The input consists of the values of multiple features of the wind turbine generator within the 50-step window, including those with missing values. The supervision labels are the complete imputation values for the missing data within the 50-step window. 80% of the data is allocated to the training set, and 20% to the test set.
[0066] (3) Construct a fusion model AFMFormer that includes an adaptive frequency domain feature extraction module, a multi-granularity temporal coding module, and an adaptive fusion completion module.
[0067] In this embodiment, the adaptive frequency domain feature extraction module performs frequency domain transformation on the input multidimensional time series data. Combined with a data-driven dominant frequency selection mechanism, it achieves adaptive spectral decomposition. Through a spectral attention mechanism, it extracts the dominant frequency components in the sequence and suppresses high-frequency noise, thereby improving the modeling ability of major patterns in long-period, highly volatile sequences. This module, based on an improved Fourier transform design, possesses frequency selectivity and multi-resolution sensing capabilities. The multi-granularity temporal coding module architecture includes two parallel coding paths: one is a patch-based Transformer path, which divides the time series into several time windows and performs local attention modeling, focusing on capturing short-term features and local structures; the other is a Standard Transformer path, used for global modeling and learning long-term dependencies to extract trend information and overall dynamic changes in the sequence. The adaptive fusion and completion module aligns the output features of different paths in the time dimension through adaptive pooling, ensuring semantic consistency and temporal synchronization of the fusion results, and then dynamically weights and fuses the output features of the two coding paths. Figures 5 to 7 The construction process of each module in step (3) specifically includes:
[0068] (3.1) Construct an adaptive frequency domain feature extraction module, which consists of two steps:
[0069] The first step is to construct an adaptive frequency domain enhancement block. This involves applying the input observation value χ... tPerform a Fast Fourier Transform (FFT) to convert it from its time-domain representation χ t The frequency domain representation F is then transformed. Further, the power spectrum P is calculated, and an adaptive mask is used to preserve the frequency domain instances in the frequency domain representation F where the power spectrum P is greater than a trainable threshold θ, resulting in the filtered frequency domain representation F. filtered Then, two sets of learnable filters are applied simultaneously, with the global filter W... G The local filter W acts on the original frequency domain representation F. L The frequency domain representation F after filtering enhanced The final frequency domain feature F is obtained by adding the two sets of filtering results. filtered Finally, the enhanced frequency domain features F are transformed using the Inverse Fast Fourier Transform (IFFT). enhanced Transforming back to the time domain yields the final time series representation χ. t The calculation process for the adaptive frequency domain enhancement block is shown in the formula:
[0070]
[0071] P = ||F|| 2
[0072] F filtered =F⊙(P>θ)
[0073] F enhanced =W G ⊙F+W L ⊙F filtered
[0074]
[0075] Where N is the number of variables, T is the length of the input sequence, and ||·|| 2 The frequency domain representation represents the squared amplitude of each frequency component of F, and ⊙ represents element-wise multiplication.
[0076] The second step is to construct a cross-scale feature interaction block. This block consists of two parallel convolutional layers, each using a different kernel size. small Using smaller convolutional kernels (e.g., kernel size = 1), the focus is on extracting fine-grained local features from the time series. Convolutional layers (Conv...) big Use a larger convolutional kernel (e.g., kernel size = 3) to focus on capturing long-term dependencies in the sequence. Input data χ tThe features are fed into parallel convolutional layers, and the outputs of the two convolutional layers interact through element-wise multiplication (Hadamard product) to enhance the feature representation. The two interacting features are then fused and passed through a Convolutional layer. fusion The process is then refined to obtain the final output features. The calculation process for cross-scale feature interaction blocks is shown in the formula:
[0077] A ˋ1 =GELU(Conv small (χ t ′))⊙Conv big (χ t ′)
[0078] A2 = GELU(Conv) big (χ t ′))⊙Conv small (χ t ′)
[0079] χ t "=Conv fusion (A ˋ1 +A ˋ2 )
[0080] GELU stands for Gaussian Error Linear Unit.
[0081] (3.2) Constructing a multi-granularity timing coding module involves two steps:
[0082] The first step is to construct a short-term temporal embedding code. First, short-term temporal embedding is performed using a patch-based embedding layer. Specifically, the original input with a batch size of B, a length of T, and N variables is... The time step is divided into multiple local blocks (Patches), each of length L, and there are a total of Each block, after being divided Then, a linear layer is used to map it to a high-dimensional space H, and finally, learnable positional encoding is added to model the order relationship between patches, resulting in a representation. Where d is the embedding dimension. Next, a patch-based encoder is used for short-term temporal encoding. Specifically, the patch-based encoder consists of multiple stacked Transformer encoder modules. Each encoder module comprises four parts: a self-attention mechanism, a feed-forward network (FFN) residual connection, and one-dimensional batch normalization (BatchNorm1d). The cross-time step normalization standardizes each feature channel in the temporal dimension. For the input tensor... First, transpose it using dimensionality transformation. The values of each channel are expanded over time. Then, one-dimensional batch normalization is used to calculate the mean and variance of each feature channel across all samples and time steps, and the result is normalized. Finally, the value is transposed back to its original shape. The normalization across time steps is calculated using the following formula:
[0083] BN T (x)=Transpose -1 °BatchNorm1d°Transpose(x)
[0084] Among them, Transpose -1 For inverse transpose transformation, BatchNorm1d is one-dimensional batch normalization, Transpose is transpose transformation, and ° is a function composition operator.
[0085] The formula for calculating short-term temporal embedding codes is as follows:
[0086]
[0087] FFN(H)=ReLU(Linear1(H))·Linear2(H)
[0088] H′=BN T (H+Self-Attention(H))
[0089] H short =BN T (H′+FFN(H′))
[0090] in, For high-dimensional space embedding features, Q, K, and V are the query, key, and value vectors, respectively, and d k The dimension of the key vector is represented by Linear1 and Linear2, which represent linear layers, and ReLU represents the Rectified Linear Unit (ReLU).
[0091] The second step is to construct long-term temporal embedding encoding. First, long-term temporal embedding is performed using a Standard Embedding layer. Specifically, each time series is independently embedded as a variable label. Specifically, the original input X is mapped from the time dimension to a high-dimensional space H through a linear layer. Second, long-term temporal encoding is performed using a Standard Encoder. Specifically, the Standard Encoder contains multiple stacked Transformer encoder modules. Each encoder module consists of four parts: a self-attention mechanism, feedforward network residual connections, and layer normalization. Layer normalization calculates the mean μ of the feature vector at each time step for each sample along dimension d. b,t and variance They are then used to normalize the vectors, thereby helping the model training to be more stable. The calculation formula is as follows:
[0092]
[0093] Where, x b,t,i Let μ be the embedding feature of the b-th sample in the batch at the t-th time step in the i-th dimension. b,t and respectively Let be the feature mean and variance of the b-th sample at time step t, d be the embedding dimension, and ∈ be a stabilizing term to prevent division by zero.
[0094]
[0095] FFN(H)=ReLU(Linear1(H))·Linear2(H)
[0096] H′=LayerNorm(H+Attention(H))
[0097] H long =LayerNorm(H′+FFN(H′))
[0098] in, For high-dimensional space embedding features, Q, K, and V are the query, key, and value vectors, respectively, and d k is the dimension of the key vector.
[0099] (3.3) Constructing an adaptive fusion completion module: Since encoders for different paths use different modeling granularities, the encoded features differ in the time dimension. Adaptive average pooling is used to uniformly adjust the output of longer paths to the target time length. Specifically, for long sequences Adaptive average pooling divides it evenly into The algorithm iterates through sub-intervals, calculates the mean across each sub-interval, and finally outputs the result. A learnable fusion weight parameter α∈(0,1) is introduced to adaptively adjust the contribution of the two paths to the final completion result. The fused result is H. fused =αH short +(1-α)H′ long Finally, the high-dimensional features are restored to predicted values in sequence form using a flatten head. This module includes two steps: feature flattening and linear projection.
[0100]
[0101] (4) Use the model for training and inference to obtain the missing data completion values of new energy, and test the model results.
[0102] In this embodiment, the model performance quantification evaluation is selected in R. 2 The tests are conducted on five metrics: RSE, RMSE, MAE, and MAPE. 2 The coefficient of determination (R) is a metric used to measure the goodness of fit of a data prediction model. It measures how well a regression model fits the data and explains the proportion of variation in the target variable. 2 The closer the value is to 1, the better the model fit. MSE (Mean Squared Error) is a metric used to measure the goodness of fit of a data prediction model. It is calculated by averaging the squared differences between the predicted and actual values. The smaller the MSE, the smaller the gap between the predicted and actual values, and the better the model fit. RMSE (Root Mean Squared Error) is the square root of MSE, sensitive to outliers, and reflects the distribution of prediction errors. MAE (Mean Absolute Error) is a statistical metric used to measure the accuracy of a prediction model or estimation method. It reflects the average level of the absolute error between the predicted and actual values; a smaller MAE is better. MAPE (Mean Absolute Percentage Error) is a metric for measuring prediction accuracy, calculated by the absolute value of the percentage error between the predicted and actual values.
[0103] To evaluate the predictive performance of the model (AFMFormer) in this embodiment, the following six time series completion prediction models were selected: 1) Transformer: a neural network that relies on a self-attention mechanism to capture long-range dependencies in a sequence; 2) iTransformer: a Transformer structure that uses variable separation modeling, effectively enhancing the ability to perceive structure between variables and improving the modeling performance of multivariate time series; 3) PatchTST: a Transformer structure based on a sliding window that divides the time series into multiple local patches, utilizing a self-attention mechanism within the patch for short-term information modeling. It maintains local temporal consistency between patches and is suitable for modeling short-term local fluctuation features; 4) Informer: A Transformer variant optimized for long time series, which uses the ProbSparse attention mechanism to reduce the complexity of the traditional self-attention mechanism; 5) FreTS: A neural network that transforms time series from the time domain to the frequency domain and uses a frequency domain multilayer perceptron to capture global dependencies; 6) DLINer: A minimal linear model with a core of a single linear layer, which directly regresses the historical sequence, avoiding the information loss and computational overhead caused by the Transformer's self-attention mechanism, and has obvious advantages in long sequence tasks with significant error accumulation.
[0104] Comparative Experiment: AFMFormer and six other mainstream temporal completion models were tested using the UEBB dataset. The five evaluation metrics are shown in Table 3. AFMFormer achieved the best results in all five metrics. This is because the model achieves dynamic screening of dominant frequency components and deep interaction of cross-scale features through the adaptive frequency domain feature extraction module (AFEB+CSFIB). Combined with the patch-based and Standard Transformer dual-path structure, it realizes the collaborative modeling of short-term local changes and long-term trends, significantly improving the structural decoupling ability and expression accuracy of complex temporal structures.
[0105] Table 3. Completion accuracy under different missing rates
[0106]
[0107] Ablation experiments: AFMFormer consists of an adaptive frequency domain feature extraction module and a multi-granularity time series completion backbone network. The backbone network has two branches. To verify the effectiveness of each part, the following models were designed: 1) Base; 2) Patch-only Base; 3) Standard-only Base; 4) AFMFormer. The Base model is a time series completion backbone network (with the adaptive frequency domain feature extraction module removed). Patch-only Base is the Base model with long-term temporal embedding encoding removed. Standard-only Base is the Base model with short-term temporal embedding encoding removed.
[0108] Table 4 Ablation Experiment
[0109]
[0110] Ablation experiments show that introducing the Adaptive Frequency Domain Feature Extraction (AFDFEM) module to the base model significantly improves AFMFormer across all metrics. This demonstrates that the AFDFEM module, through data-driven frequency domain decomposition and major frequency extraction of the input sequence, significantly enhances the model's ability to model multi-scale temporal structures in new energy data, particularly demonstrating a clear advantage in decoupling long-term trends from short-term fluctuations. Removing any encoding path results in a significant performance degradation. While the patch-only Base model possesses some capability in modeling short-term fluctuations, its R-value is significantly lower than that of the full fusion model. 2 The performance decreased by 5.07%, while RMSE and MAE increased to 0.7063 and 0.4869, respectively, indicating a significant performance degradation. Similarly, the Standard-only Base model performed poorly in modeling long-term dependencies. This result demonstrates that the Patch-based and Standard paths are significantly complementary in terms of modeling capabilities, and neither single path can independently handle the task of modeling complex multi-scale temporal features.
[0111] Based on the same inventive concept, this invention discloses a new energy missing data completion system based on multi-model fusion, comprising: a feature screening module for performing two-stage feature screening of the collected new energy dataset using linear and nonlinear correlation; a data preprocessing module for normalizing the dataset, constructing a missing data mask based on the missing proportion parameter, and dividing the dataset; a model building module for constructing a fusion model AFMFormer including an adaptive frequency domain feature extraction module, a multi-granularity time-series coding module, and an adaptive fusion completion module; wherein the adaptive frequency domain feature extraction module, based on an improved Fourier transform design, has frequency selectivity and multi-resolution sensing capabilities, achieving adaptive spectral decomposition, extracting the dominant frequency components in the sequence, and suppressing high-frequency noise; the multi-granularity time-series coding module includes two parallel coding paths: one is a patch-based Transformer path, which divides the time series into several time windows and performs local attention modeling; the other is a Standard... The Transformer path is used to extract trend information and overall dynamic changes in the sequence; the adaptive fusion completion module is used to align the output features of different paths in the time dimension and dynamically weight and fuse the output features of the two encoded paths; the model training module is used to train the AFMFormer model, use the trained model to impute missing data, and output the missing data completion results.
[0112] The specific working processes of each module described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. The division of modules is only a logical functional division, and there may be other division methods in actual implementation. For example, multiple modules may be combined or integrated into another system.
[0113] This invention also discloses a computer system, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps of the method for completing missing new energy data based on multi-model fusion.
[0114] This invention also discloses a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the method for completing missing new energy data based on multi-model fusion.
[0115] This invention also discloses a computer program product, including a computer program that, when executed by a processor, implements the steps of the aforementioned method for completing missing new energy data based on multi-model fusion.
[0116] The program code used to implement the method of the present invention can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the program code causes the steps of the method of the present invention to be performed. The program code can be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a standalone software package, or entirely on a remote machine or server. All aspects not detailed in this invention are well-known to those skilled in the art.
[0117] Although the present invention has been illustrated and described with reference to preferred embodiments, those skilled in the art should understand that various changes and modifications can be made to the present invention without departing from the scope defined by the claims.
Claims
1. A method for completing missing new energy data based on multi-model fusion, characterized in that, Includes the following steps: The collected new energy datasets were subjected to two-stage feature screening based on linear and nonlinear correlations. The dataset is normalized, a missing data mask is constructed based on the missing proportion parameter, and the dataset is divided. A fusion model AFMFormer is constructed, which includes an adaptive frequency domain feature extraction module, a multi-granularity temporal coding module, and an adaptive fusion completion module. in The adaptive frequency domain feature extraction module, based on an improved Fourier transform design, has frequency selectivity and multi-resolution sensing capabilities, enabling adaptive spectral decomposition to extract the dominant frequency components in the sequence and suppress high-frequency noise. The multi-granularity temporal coding module includes two parallel coding paths: one is a Patch-based Transformer path, which divides the time series into several time windows and performs local attention modeling; the other is a Standard Transformer path, which is used to extract trend information and overall dynamic changes in the sequence. The adaptive fusion completion module is used to align the output features of different paths in the time dimension and dynamically weight and fuse the output features of two encoded paths. Train the AFMFormer model, use the trained model to impute missing data, and output the imputed data results. The adaptive frequency domain feature extraction module includes an adaptive frequency domain enhancement block and a cross-scale feature interaction block; The adaptive frequency domain enhancement block introduces a learnable threshold parameter after performing a fast Fourier transform on the input value to filter high-energy regions in the frequency domain; after filtering out the main frequency domain information, adaptive spectrum enhancement is performed through local and global dual-channel learning filters. The cross-scale feature interaction block inputs the adaptively frequency-domain enhanced data into two parallel convolutional layers with different kernel sizes. The outputs of the two convolutional layers interact through element-wise multiplication to enhance the feature representation. The two interacting features are then fused and processed by a fusion convolutional layer to obtain the final output features. The calculation process of the adaptive frequency domain enhancement block is as follows: ; in, For the input observations, Represents the Fast Fourier Transform. For frequency domain representation, Power spectrum, For a trainable threshold, This is the filtered frequency domain representation. For global filters, For local filters, This is the time series representation obtained from the inverse Fast Fourier Transform. For the number of variables, Given the length of the input sequence, Frequency domain representation The amplitudes of each frequency component are squared. This indicates element-wise multiplication.
2. The method for completing missing new energy data based on multi-model fusion according to claim 1, characterized in that, The feature selection includes: calculating the Pearson correlation coefficient and the maximum mutual information coefficient among the features in the collected new energy data. If the correlation between any two features does not reach a preset threshold, the data is considered redundant, and features with low correlation to the target variable are ignored.
3. The method for completing missing new energy data based on multi-model fusion according to claim 1, characterized in that, In the multi-granularity temporal coding module, the Patch-based Transformer path first uses a Patch-based Embedding layer to process the input time series data. Divided into multiple local blocks (Patches): And it is mapped to a high-dimensional space through a linear layer, where For batch size, The length of each block is determined; then learnable positional encodings are added to model the order relationships between patches, resulting in the corresponding embedding representations. ,in To embed the dimension, the high-dimensional representation is then encoded using a Patch-based Encoder. The Patch-based Encoder contains multiple stacked Patch-based Transformer encoder modules. Each encoder module consists of four parts: a self-attention mechanism, a feedforward network, a residual connection, and one-dimensional batch normalization. One-dimensional batch normalization involves transposing and expanding the input in the time dimension and applying it to each feature channel to achieve standardization across time steps. The Standard Transformer path first uses a Standard Embedding layer to map the original input from the time dimension to a high-dimensional space through a linear layer; then, it encodes the high-dimensional space representation through a Standard Encoder. The Standard Encoder contains multiple stacked Standard Transformer encoder modules, each of which consists of four parts: a self-attention mechanism, a feedforward network, residual connections, and standard layer normalization. The multi-granularity temporal coding module encodes the input in parallel through the above two paths.
4. The method for completing missing new energy data based on multi-model fusion according to claim 1, characterized in that, The adaptive fusion completion module uses an adaptive average pooling operation to evenly divide the sequence obtained by the Standard Transformer path and calculates the mean in each sub-interval. The output of the Standard Transformer path is uniformly adjusted to the sequence length of the Patch-based Transformer path. Based on this, a learnable fusion weight parameter is introduced to adaptively adjust the contribution of the two aligned encoding paths to the final completion result, and then the feature representation is fused through this parameter. Finally, a linear mapping head is used to flatten and linearly map the fused features, restoring the high-dimensional features to predicted values in sequence form.
5. A new energy missing data completion system based on multi-model fusion, used to implement the method according to any one of claims 1-4, characterized in that, include: The feature selection module is used to perform two-stage feature selection on the collected new energy dataset, including linear and nonlinear correlations. The data preprocessing module is used to normalize the dataset, construct a missing data mask based on the missing proportion parameter, and divide the dataset. The model building module is used to build the fusion model AFMFormer, which includes an adaptive frequency domain feature extraction module, a multi-granularity temporal coding module, and an adaptive fusion completion module. in The adaptive frequency domain feature extraction module, based on an improved Fourier transform design, has frequency selectivity and multi-resolution sensing capabilities, enabling adaptive spectral decomposition to extract the dominant frequency components in the sequence and suppress high-frequency noise. The multi-granularity temporal coding module includes two parallel coding paths: one is a Patch-based Transformer path, which divides the time series into several time windows and performs local attention modeling; the other is a Standard Transformer path, which is used to extract trend information and overall dynamic changes in the sequence. The adaptive fusion completion module is used to align the output features of different paths in the time dimension and dynamically weight and fuse the output features of two encoded paths. The model training module is used to train the AFMFormer model, perform missing data imputation using the trained model, and output the missing data completion results.
6. A computer system comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the computer program is executed by the processor, it implements the steps of the new energy missing data completion method based on multi-model fusion as described in any one of claims 1-4.
7. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the new energy missing data completion method based on multi-model fusion as described in any one of claims 1-4.
8. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the new energy missing data completion method based on multi-model fusion as described in any one of claims 1-4.
Citation Information
Patent Citations
Sea surface temperature complementation method and system based on missing correlation and physical constraint
CN119107261A
Adaptive timestamp coding enhanced complex equipment incomplete state monitoring data time sequence interpolation method
CN119646555A