Virtual power plant operation parameter prediction method and device, storage medium and computer program product

By using time series decomposition and dynamic clustering representation enhancement of large language models, the accuracy and real-time issues in multidimensional data prediction of virtual power plants are solved, thereby improving the accuracy and efficiency of load and electricity price prediction for virtual power plants.

CN120611137BActive Publication Date: 2025-11-21ELU TECHNOLOGY HOLDINGS (ZHEJIANG)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511114173.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2025-11-21
Estimated Expiration
2045-08-08

AI Technical Summary

Technical Problem

When dealing with multidimensional data, time-series forecasting methods for virtual power plants struggle to effectively capture complex features. Existing methods are insufficient in terms of accuracy and robustness, failing to meet real-time requirements.

Method used

The time series decomposition algorithm is used to decompose the data into trend, seasonal and residual components. Combined with dynamic clustering representation enhancement and model fine-tuning of the large language model, the predictive ability is improved through semantic anchor matching and contrastive learning mechanism.

Benefits of technology

It significantly improves the accuracy and real-time performance of virtual power plant load and electricity price forecasts, optimizes the inference efficiency of the model on edge devices, and supports the intelligent management and optimized operation of virtual power plants.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120611137B_ABST
    Figure CN120611137B_ABST
Patent Text Reader

Abstract

The application provides a virtual power plant operation parameter prediction method and device, a storage medium and a computer program product. The method comprises the following steps: performing data preprocessing on obtained influence factor data; using a time series decomposition algorithm to decompose the influence factor data into three-dimensional sequences containing trend components, seasonal components and residual components, cutting each sequence into sequence blocks and mapping the sequence blocks into time sequence feature vectors; performing K-means clustering on word embeddings used for pre-training of a large language model, selecting K cluster centers as semantic anchor points and splicing the semantic anchor points with the time sequence feature vectors; fine-tuning position embedding parameters of the large language model, weights of a feedforward neural network connected in residual and parameters of a normalization layer; predicting trend components, seasonal components and residual components of each dimension according to the semantic enhanced time sequence feature vectors, and obtaining a prediction result after splicing and inverse normalization processing of the components. The prediction method improves the accuracy and real-time performance of virtual power plant load and electricity price prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the interdisciplinary field of smart grid and machine learning, specifically involving a method and device for predicting the operating parameters of a virtual power plant (VPP), a storage medium, and a computer program product, aiming to improve the accuracy and real-time performance of predicting parameters such as load and electricity price of virtual power plants. Background Technology

[0002] Virtual power plants, by aggregating distributed energy resources (such as solar, wind, and energy storage devices) and controllable loads, enable flexible scheduling and optimized operation of the power system, representing an important development direction for smart grids. Their operation relies on accurate forecasting of key parameters such as load and electricity prices to support optimal resource allocation and electricity market transactions.

[0003] However, the forecasting task of virtual power plants faces multiple challenges: First, the data is multidimensional, noisy, and non-stationary. For example, solar power generation fluctuates intermittently due to weather conditions, and load data is uncertain due to user behavior. Second, traditional time series forecasting methods, such as the Autoregressive Integrated Moving Average (ARIMA) model and Long Short-Term Memory (LSTM) networks, struggle to capture long-term dependencies and complex relationships between cross-dimensional features when dealing with multidimensional heterogeneous data. Furthermore, while Large Language Models (LLMs) based on the Transformer architecture have performed well in natural language processing in recent years, their application in time series forecasting is also challenging due to the structural differences between numerical data and linguistic semantics, and their high computational complexity cannot meet real-time requirements. Finally, existing improvement methods, such as simple block-based input to Transformers or fixed clustering strategies, do not adequately decompose data features or adapt to dynamic changes, resulting in insufficient forecast accuracy and robustness.

[0004] Therefore, there is an urgent need for a prediction method that can comprehensively consider the complexity, semantic relevance, and real-time response requirements of virtual power plant time series data in order to improve the operating efficiency of virtual power plants. Summary of the Invention

[0005] This invention addresses the problem that traditional time-series forecasting methods lack accuracy in multi-dimensional data processing of virtual power plants and cannot effectively capture complex features. This invention proposes a method and device for predicting virtual power plant operating parameters, a storage medium, and a computer program product.

[0006] To achieve the above objectives, according to a first aspect of the present invention, a method for predicting operating parameters of a virtual power plant is provided, comprising the following steps: a data acquisition step, acquiring multi-dimensional influencing factor data for predicting the operating parameters of the virtual power plant; a data preprocessing step, performing preprocessing on the multi-dimensional influencing factor data including normalization and data cleaning; a time-series feature decomposition step, using a time-series decomposition algorithm to decompose each dimension of influencing factor data into a three-dimensional sequence containing a trend component, a seasonal component, and a residual component, dividing each sequence into sequence blocks, and mapping the sequence blocks into time-series feature vectors that can be used as input to a large language model; and a dynamic clustering representation enhancement step, performing K-... Means clustering: Cluster centers are selected and their similarity is matched with the temporal feature vector. The K highest-scoring cluster centers are selected as semantic anchors and concatenated with the temporal feature vector to obtain a semantically enhanced temporal feature vector. Model fine-tuning step: The semantically enhanced temporal feature vector is input into the pre-trained large language model, and the position embedding parameters, weights of the feedforward neural network with residual connections, and parameters of the normalization layer are fine-tuned. Model output step: Based on the semantically enhanced temporal feature vector, trend components, seasonal components, and residual components in each dimension are predicted. The trend components, seasonal components, and residual components in the same dimension are concatenated and denormalized to obtain the prediction result, which is then used as the output of the large language model.

[0007] As a preferred embodiment, in the data acquisition step, the influencing factor data of the multiple dimensions includes at least two of the following data items: weather information, calendar information, historical controllable load sequence, electricity price sequence, distributed energy output sequence, electric vehicle charging and discharging load and uncontrolled load.

[0008] As a preferred embodiment, the preprocessing steps in the data preprocessing step include the following steps: a normalization step, which normalizes the influencing factor data of the multiple dimensions; an outlier detection step, which detects outliers in the dataset of the influencing factor data of each dimension; a missing value imputation step, which imputes missing values ​​in the dataset of the influencing factor data of each dimension; and a granularity unification step, which aligns the time granularity of the influencing factor data of each dimension using either nearest neighbor interpolation or cubic spline interpolation, based on the differences in sampling frequency.

[0009] As a preferred embodiment, in the normalization step, the min-max normalization method is used to normalize the data of the influencing factors in the multiple dimensions:

[0010] ,

[0011] in, The minimum value of the influencing factor data for each dimension. The maximum value of the influencing factor data for each dimension. The data points are normalized from the influencing factors data for each dimension, where x is the original data point.

[0012] As a preferred embodiment, the outlier detection step includes:

[0013] The first outlier detection step uses the Z-score method to calculate the mean of the influencing factor data for each dimension. and standard deviation Data points that deviate from the mean by more than 3 standard deviations are marked as outliers.

[0014] ,

[0015] The second outlier detection step uses the interquartile range method to detect outliers, marking data points that exceed the upper and lower bounds, where the lower bound is... The upper boundary is In the formula It is the first quartile. The third quartile, the interquartile range (IQR) is calculated using the following formula: .

[0016] The third outlier detection step uses the local outlier factor method to detect outliers. It calculates the local density of each data point and compares it to the density of its neighboring points. If the local density of a data point is lower than the density of its neighboring points by a preset threshold, then the data point is considered outlier. The calculation formula is defined as follows:

[0017] ,

[0018] In the formula, p is the target data point. It is the k-neighborhood of p. It is the locally reachable density of p.

[0019] The formula for calculating the locally reachable density is as follows:

[0020] ,

[0021] In the formula, , d(p,o) is the k-distance between point o and point p and point o.

[0022] In the statistical steps, if at least two of the first, second, and third outlier detection steps are marked as outliers for a certain data point, then the data point is confirmed as an outlier.

[0023] As a preferred embodiment, the missing value filling step includes:

[0024] This multivariate time series interpolation method based on the attention mechanism normalizes the data, performs nulling after outlier detection, selects a time window for each missing value, and calculates the correlation between the d-th dimension and other dimensions within the time window, as defined below:

[0025] ,

[0026] in, For missing values, Let be the mean of the d-th dimension within the time window. Let be the mean of the k-th dimension within the time window. Let d be the correlation score between the d-dimensional and k-dimensional dimensions, and t be the current time step. w represents the historical time step within the time window.

[0027] The relevance scores are converted into attention weights using a softmax function. ,

[0028] ,

[0029] The missing values ​​are estimated by using attention weights to perform a weighted average of the non-missing values ​​in other dimensions. The formula for missing value estimation is defined as follows:

[0030] ,

[0031] If the k-th dimension is also missing at time t, then the mean of that dimension within the window is used. Alternative The filled values ​​are then denormalized to the original scale. The denormalization formula is defined as follows:

[0032] ,

[0033] in, To represent the missing values ​​estimated by the attention mechanism interpolation method at time t and dimension d, This refers to the original scaled data values ​​after denormalization, i.e., the supplemented true values. is the set of all original data of dimension d, used to calculate the minimum and maximum values ​​during the denormalization process.

[0034] As a preferred embodiment, the granularity unification step includes: for data with a sampling frequency higher than the target granularity, directly retaining the original data; if the time points are not perfectly aligned, using nearest neighbor interpolation to select the closest data point in time to fill the gap; for data with a sampling frequency lower than the target granularity, using cubic spline interpolation to generate high-frequency data points, i.e., for low-frequency data points... Fitting a piecewise cubic polynomial ,Require It is continuous at each point and its first and second derivatives.

[0035] As a preferred embodiment, the time-series feature decomposition step includes: receiving data on influencing factors from multiple dimensions. ,in Let d represent the data of dimension d, where T is the time step and d is the number of dimensions. Then, perform instance normalization again:

[0036] ,

[0037] For each dimension of the normalized data The time series decomposition algorithm is used to decompose it into seasonal cycles, trend windows, and low-pass filter windows, as shown below:

[0038] ,

[0039] in, This is a trend component, reflecting long-term changes.

[0040] Seasonal components, capturing cyclical fluctuations.

[0041] The residual component contains noise and aperiodic fluctuations.

[0042] As a preferred embodiment, the time-series feature decomposition step further includes:

[0043] Set the block length L and stride S, set the preset overlap rate, and independently segment the sequence into blocks for each trend component, seasonal component, and residual component obtained by the time series decomposition algorithm to generate a block set:

[0044] ,

[0045] When the sequence ends with fewer than L points, zeros are added to ensure that all sequence blocks have the same length.

[0046] The trend component sequence block, seasonal component sequence block, and residual component sequence block within the same time window are concatenated along the feature dimension to form a fused sequence block as the time-series feature vector. The concatenation formula is defined as follows:

[0047] ,

[0048] in, For fused sequence blocks, For trend component sequence blocks, For seasonal component sequence blocks, This is a block of residual component sequences.

[0049] As a preferred embodiment, the dynamic clustering representation enhancement step includes: a cluster center partitioning step, which involves dividing the word embedding matrix of the pre-trained large language module. The K-means algorithm is used to divide the word embedding matrix into K clusters, and the intra-cluster variance is minimized using the following formula:

[0050] ,

[0051] in, As the cluster center, K is determined using the elbow method, and the sum of squares within the cluster is calculated:

[0052] ,

[0053] Where WCSS is the sum of squares within the cluster, plotted... Follow For a changing curve, select the point with the largest change in curvature as the optimal number of clusters. The following conditions must be met:

[0054] ,

[0055] The cluster center selection step involves calculating the silhouette coefficient of each cluster and retaining the top clusters. The cluster center with the highest discriminative power ;

[0056] The feature concatenation step involves calculating the temporal feature vector and the cluster centers. The cosine similarity is selected. The most relevant cluster centers are concatenated to form the semantically enhanced temporal feature vector. The feature concatenation formula is defined as follows:

[0057] ,

[0058] in, For semantically enhanced temporal feature vectors, This is a time-series feature vector.

[0059] As an optional approach, a modeling step is included before the model fine-tuning step. This modeling step includes: a prediction target definition step, which designs global prompts to guide the large language model to focus on the prediction target and establishes the association between local features and the global task; a text description step, which inputs task description text; and an encoding step, which uses the large language model's word segmenter to encode prompt words. Then the prompt words Mapped to cue embedding vector The feature concatenation step involves embedding the prompt into the vector. With the semantically enhanced temporal feature vector The input sequence is obtained by concatenating along the sequence dimension. The splicing formula is defined as follows:

[0060] ,

[0061] The positional encoding step involves injecting absolute positional codes into the input sequence to preserve temporal order information.

[0062] ,

[0063] in, For the position encoding matrix, The input sequence is .

[0064] As an optional approach, the model fine-tuning steps include:

[0065] With the backbone parameters of the pre-trained large language model frozen, we optimize the position embeddings, the weights of the feedforward network in the residual connections, and the scaling factor and bias in the normalization layer. The parameter optimization formula for the normalization layer is defined as follows:

[0066] ,

[0067] in, The mean of the input data. The variance of the mean of the input data. The scaling factor in the normalization layer. The bias in the normalization layer. To hide the input state, To output the hidden state.

[0068] As an optional approach, the modeling step further includes:

[0069] The sample setup steps involve setting positive and negative samples. Positive samples are the K temporal feature vectors that match the Top-K cluster centers, while negative samples are other temporal feature vectors in the same batch that do not belong to the Top-K cluster centers.

[0070] The steps for setting the contrastive learning loss function are as follows: Based on the optimization objectives of narrowing the distance between positive samples and widening the distance between negative samples, a contrastive learning loss function is set to semantically align the temporal feature vector with the large language model.

[0071] ,

[0072] in, As a positive sample, The cosine similarity function is used. The temperature coefficient controls the steepness of the distribution. M represents the number of negative samples, and M represents the total number of samples.

[0073] The steps for setting the training loss function are as follows: Based on the contrastive learning loss function, the training loss function is defined as follows:

[0074] ,

[0075] in, To predict losses, To compare learning loss, Weights for predicting loss, The weights are used to compare the learning loss.

[0076] As an optional approach, the model output step includes:

[0077] The semantically enhanced temporal feature vector is positionally encoded to obtain a positional encoding matrix. The input is fed into a pre-trained large language model, which generates hidden states H through multiple Transformer layers. The hidden states of the corresponding temporal segments are then extracted based on the cue positions. Hidden state is hidden through a fully connected layer. Mapped to predicted values The result is then denormalized to its original dimensions and used as the prediction result. The denormalization formula is defined as follows:

[0078] ,

[0079] in, For predicted values, Let d be the mean of the original data in the d-th dimension. Let d be the standard deviation of the original data in the d-th dimension. This is the predicted value of the d-th dimension after inverse normalization.

[0080] According to a second aspect of the present invention, a prediction apparatus for predicting operating parameters of a virtual power plant is provided, comprising: a data acquisition unit for acquiring multi-dimensional influencing factor data for predicting operating parameters of the virtual power plant; a data preprocessing unit for performing preprocessing on the multi-dimensional influencing factor data, including normalization and data cleaning; a time-series feature decomposition unit for using a time-series decomposition algorithm to decompose each dimension of influencing factor data into a three-dimensional sequence containing a trend component, a seasonal component, and a residual component, dividing each sequence into sequence blocks, and mapping the sequence blocks into time-series feature vectors that can be used as input to a large language model; and a dynamic clustering representation enhancement unit for performing K-me on word embeddings used for pre-training the large language model. Ans clustering is used to select cluster centers and perform similarity matching with the temporal feature vector. The K cluster centers with the highest scores are selected as semantic anchors and concatenated with the temporal feature vector to obtain a semantically enhanced temporal feature vector. The model fine-tuning unit is used to input the semantically enhanced temporal feature vector into the pre-trained large language model and fine-tune the position embedding parameters, the weights of the feedforward neural network of the residual connection, and the parameters of the normalization layer of the large language model. The model output unit is used to predict the trend component, seasonal component, and residual component of each dimension based on the semantically enhanced temporal feature vector, and concatenate and denormalize the trend component, seasonal component, and residual component under the same dimension to obtain the prediction result, which is used as the output of the large language model.

[0081] According to a third aspect of the invention, a non-transitory storage medium is provided, which stores a computer program that, when executed by a processor, enables the prediction method for virtual power plant parameters according to the first aspect of the invention.

[0082] According to a fourth aspect of the present invention, a computer program product is provided, comprising computer instructions that, when executed by a processor, enable the prediction method for virtual power plant parameters according to a first aspect of the present invention.

[0083] Compared with the prior art, the present invention has at least the following beneficial effects:

[0084] First, this invention proposes a multi-component block method based on time series decomposition algorithms, overcoming the shortcomings of traditional time series forecasting methods in effectively capturing trend, seasonal, and residual features when processing multidimensional data from virtual power plants. Second, word embedding clustering from a pre-trained large language model is applied to enhance the representation of time series data. Through a large language model prediction and contrastive learning mechanism based on a semantic anchor matching strategy, the model's ability to understand complex time series patterns is significantly improved, overcoming the deficiencies of existing methods in semantic association and dynamic adaptability. Finally, with prediction accuracy and computational efficiency as optimization objectives, a residual fine-tuning and quantization deployment strategy for the virtual power plant prediction model is established, effectively solving the challenge of real-time inference of large models on edge devices. This method can more comprehensively improve the accuracy and real-time performance of virtual power plant load and electricity price forecasts, and has significant practical implications for the intelligent management and optimized operation of virtual power plants.

[0085] Furthermore, based on the STL decomposition results, the time series data is divided into multi-component blocks, and the trend, seasonal and residual features are processed and spliced ​​independently, which makes it easier for the model to accurately capture the change patterns at different time scales.

[0086] Furthermore, based on cluster-enhanced temporal token generation, the representation enhancement strategy is dynamically adjusted by matching the similarity with the cluster centers of LLM word embeddings, thereby improving the model's semantic understanding of temporal data.

[0087] Furthermore, by combining residual connections, normalization layers, and fine-tuning of location embeddings, along with INT8 quantization deployment, the inference efficiency of the model on edge devices is optimized to ensure real-time prediction.

[0088] In summary, this invention significantly improves the accuracy and efficiency of multi-dimensional time-series data prediction for virtual power plants through multi-component decomposition, block partitioning, characterization enhancement, and efficient model optimization. It provides reliable technical support for resource scheduling and market transactions of virtual power plants and supports migration prediction with limited data, demonstrating promising application prospects. Attached Figure Description

[0089] Figure 1 An overall flowchart illustrating the method for predicting virtual power plant parameters according to the present invention is provided.

[0090] Figure 2 A flowchart illustrating the data preprocessing steps according to the present invention is shown.

[0091] Figure 3 A flowchart illustrating the modeling steps according to the present invention is provided.

[0092] Figures 4A to 4F An example is shown comparing the MAE predicted by the LLM according to the present invention with those predicted by various conventional models.

[0093] Figure 5A A comparison chart illustrating the load forecasting performance of the forecasting method according to the present invention and a conventional forecasting method without STL decomposition is provided.

[0094] Figure 5B A comparison chart illustrating the electricity price prediction performance of the prediction method according to the present invention and a conventional prediction method without STL decomposition is provided.

[0095] Figure 6 The diagram illustrates the effect of changing the parameters of the comparative learning part of the loss function according to the present invention.

[0096] Figure 7A The diagram illustrates a comparison of load forecasting performance between models with and without fine-tuning based on the forecasting method according to the present invention.

[0097] Figure 7B A comparison chart of the electricity price prediction performance of the model with and without fine-tuning according to the prediction method of the present invention is shown.

[0098] Figure 8 A software structure block diagram illustrating the method for predicting virtual power plant parameters according to the present invention is shown. Detailed Implementation

[0099] Exemplary embodiments of the present invention will now be described in detail with reference to the accompanying drawings. It should be noted that, unless otherwise specifically stated, the relative configuration of components, numerical representations, and values ​​described in these embodiments does not limit the scope of the invention.

[0100] In this invention, the term "unit" can refer to a software environment, a hardware environment, or a combination of both. In a software environment, the term "unit" refers to a function, application, software module, feature, routine, set of instructions, or program that can be executed by a programmable processor (such as a microprocessor, central processing unit (CPU), or specially designed programmable device) or controller. Memory contains instructions or programs that, when executed by the CPU, cause the CPU to perform operations corresponding to the unit or function. In a hardware environment, the term "unit" refers to a hardware element, circuit, component, physical structure, system, module, or subsystem. According to a particular embodiment, the term "unit" can include mechanical, optical, or electrical components, or any combination thereof. The term "unit" can include active (e.g., transistors) or passive (e.g., capacitors) components. The term "unit" can include a semiconductor device having a substrate and other material layers having various conductivity concentrations. It can include a CPU or programmable processor that can execute programs stored in memory to perform a specified function. The term "unit" can include logic elements (e.g., AND, OR) implemented by transistor circuitry or any other switching circuitry. In the context of a combination of software and hardware environments, the term "unit" or "circuit" refers to any combination of software and hardware environments as described above. Additionally, the terms "element," "component," "part," or "device" may also refer to a "circuit" integrated with or not integrated with packaging material.

[0101] The virtual power plant described in this invention will be explained below.

[0102] [Virtual Power Plant]

[0103] A virtual power plant (VPP) is a power coordination and management system that uses advanced information and communication technologies and software systems to aggregate and coordinate distributed energy sources such as distributed power sources, energy storage systems, controllable loads, electric vehicles, and charging piles. It functions as a special type of power plant participating in the electricity market and grid operation. A virtual power plant is not a real power plant, but rather a smart grid technology that applies a distributed power management system to participate in grid operation and dispatch, achieving "source-load-grid" aggregation and optimization. It balances the grid by coordinating power generation resources and adjusting some electricity demand, reducing peak-hour electricity consumption and increasing the flexibility of power dispatch.

[0104] The following section will describe in detail the method for predicting the operating parameters of a virtual power plant.

[0105] [Prediction Methods for Virtual Power Plant Operating Parameters]

[0106] The method for predicting virtual power plant operating parameters according to the present invention can be implemented by a processor in the virtual power plant operating parameter prediction system executing a computer program stored in the memory of the virtual power plant operating parameter prediction system. Alternatively, the virtual power plant operating parameter prediction system can communicate with a server, and a processor in the server can execute a computer program stored on the server or in the cloud, feeding back the program execution results to the virtual power plant operating parameter prediction system in real time.

[0107] In this invention, the method for predicting the operating parameters of a virtual power plant mainly includes a data acquisition step S100, a data preprocessing step S200, a time-series feature decomposition step S300, a dynamic clustering representation enhancement step S400, a model fine-tuning step S500, and a model output step S600. (See below for further details.) Figure 1 The method for predicting virtual power plant parameters according to the present invention will be described.

[0108] like Figure 1 As shown, the data acquisition step S100 is executed first to acquire data on multiple dimensions of influencing factors used to predict the operating parameters of the virtual power plant.

[0109] The multi-dimensional influencing factor data in this invention can include weather data (such as temperature, humidity, air pressure, wind speed, and solar radiation), calendar information (such as dates and holidays), historical controllable load sequences (such as regional load demand sequences), electricity price sequences (such as regional electricity market price sequences), distributed energy output sequences (such as regional wind and solar renewable energy output and V2G load), and uncontrolled loads (such as inter-regional transmission loads). These influencing factor data allow for comprehensive prediction of key parameters such as load and electricity price of the virtual power plant from multiple perspectives.

[0110] Next, perform data preprocessing step S200 to preprocess the data on influencing factors across multiple dimensions, including normalization and data cleaning.

[0111] This invention normalizes the influencing factor data for each dimension (such as temperature, humidity, wind speed, etc.) to ensure that data of different dimensions are on the same scale, facilitating comparison. Data cleaning identifies erroneous data in the influencing factor data and is a process of re-examining and verifying the data. The goal of data cleaning is to remove duplicate records, eliminate abnormal data, correct erroneous data, ensure data consistency, and improve data quality.

[0112] Further execute the time series feature decomposition step S300, using a time series decomposition algorithm to decompose each dimension of influencing factor data into a three-dimensional sequence containing trend component, seasonal component and residual component, divide each sequence into sequence blocks, and map the sequence blocks into time series feature vectors that can be used as input to a large language model.

[0113] Specifically, this invention employs the STL (Seasonal-Trend Decomposition using Loess) algorithm to decompose time series data. First, multi-dimensional data on influencing factors are acquired, and then the STL algorithm decomposes this data into three components: seasonal cycle, trend window, and low-pass filter window. The trend component reflects long-term changes, the seasonal component captures periodic fluctuations, and the residual component contains noise and aperiodic fluctuations. The STL decomposition algorithm is highly flexible, applicable to various seasonal patterns, and can identify long-term climate change trends, seasonal fluctuations, and short-term anomalies or noise.

[0114] Furthermore, after STL decomposition to obtain the trend component (T), seasonal component (S), and residual component (R), the time series of each dimension is divided into blocks. For each component, an independent block set is generated, and then the patch is mapped to a temporal feature vector (i.e., a token) that can be recognized by a Large Language Model (LLM). LLM is essentially a model for processing sequential data and cannot directly understand raw data such as pixel values ​​or video frames. Mapping patches of non-textual data such as images or videos to tokens, which are then input into the model for processing, primarily aims to transform unstructured data into a structured semantic representation that the model can understand, thereby achieving the fusion and processing of multimodal information.

[0115] In the dynamic clustering representation enhancement step S400, K-means clustering is performed on the word embeddings used for pre-training the large language model. Cluster center representations are selected and similarity matching is performed with temporal feature vectors. The K cluster centers with high scores are selected as semantic anchors and concatenated with temporal feature vectors to obtain semantically enhanced temporal feature vectors.

[0116] Specifically, for the word embedding matrix of the pre-trained LLM, the K-means algorithm is used to divide the word embedding matrix into K clusters. The Top-K cluster centers most relevant to the temporal token (i.e., temporal feature vector) are selected and concatenated to form the enhanced token (i.e., semantically enhanced temporal feature vector). By dynamically associating the LLM semantic space with local temporal features, semantic features are integrated into the temporal features, improving the interpretability of large language models for complex patterns. By decomposing temporal features using STL and combining them with LLM semantic anchors, cross-modal enhancement from numerical features to semantic features is achieved, overcoming the LLM's obstacle to understanding purely numerical data.

[0117] In the model fine-tuning step S500, the semantically enhanced temporal feature vector is input into the pre-trained large language model, and the position embedding parameters, the weights of the feedforward neural network of the residual connection, and the parameters of the normalization layer of the large language model are fine-tuned.

[0118] This invention, while freezing the pre-trained LLM backbone parameters, optimizes residual connections, normalization layers, and position embedding parameters in a targeted manner. This enables the large language model to adapt to the prediction task of virtual power plant parameters, such as predicting key indicators like electricity prices and loads. In this embodiment, only parameters strongly correlated with time-series characteristics are optimized, preserving the semantic understanding capabilities of the LLM pre-training and avoiding redundant training. This significantly reduces computational costs while retaining pre-training capabilities, thus meeting the real-time requirements of virtual power plants.

[0119] In the model output step S600, the trend component, seasonal component and residual component of each dimension are predicted based on the semantically enhanced temporal feature vector. The trend component, seasonal component and residual component under the same dimension are then concatenated and inversely normalized to obtain the prediction result, which is used as the output of the large language model.

[0120] As mentioned earlier, STL decomposition divides the influencing factor data for each dimension into three components: trend, seasonality, and residuals. Each component is processed independently after being segmented. These three components are then input into the LLM after being segmented and concatenated with semantic anchors. When the LLM outputs, these components need to be recombine because the prediction target is the data of the original dimension, such as load and electricity price. Each dimension has its own three components, and the three components corresponding to each dimension must be concatenated according to the dimension during output. For example, the trend, seasonality, and residual components of load are concatenated to form the load prediction sequence, and the same applies to the prediction sequences of electricity price or other dimensions. The purpose of this is to maintain the temporal alignment of each component, ensuring that the trend, seasonality, and residuals correspond in time and that the combined result correctly reflects the pattern of the original data. Next is denormalization. Since normalization was used in the previous data preprocessing steps, the normalized values ​​predicted by the model need to be converted back to their original units during output, such as converting load to "MW" or electricity price to "yuan / kWh". This step is crucial because users need meaningful numerical values, not just normalized values.

[0121] This invention proposes a multi-component block method based on STL decomposition, which overcomes the shortcomings of traditional time series prediction methods in effectively capturing trend, seasonal, and residual features when processing multidimensional data from virtual power plants. Secondly, it is the first to apply word embedding clustering of a pre-trained large language model to enhance the representation of time series data. Through large language model prediction and contrastive learning mechanisms based on semantic anchor matching strategies, it significantly improves the model's ability to understand complex time series patterns, overcoming the deficiencies of existing methods in semantic association and dynamic adaptability. Finally, with prediction accuracy and computational efficiency as optimization goals, the virtual power plant prediction model combines residual connections, normalization layers, fine-tuning of position embeddings, and INT8 quantization deployment strategies to optimize the inference efficiency of the model on edge devices, ensuring real-time prediction and effectively solving the problem of real-time inference of large models on edge devices.

[0122] The prediction method of this invention can comprehensively improve the accuracy and real-time performance of virtual power plant load and electricity price prediction, significantly enhancing the accuracy and efficiency of multi-dimensional time-series data prediction for virtual power plants. This has significant practical implications for the intelligent management and optimized operation of virtual power plants. The prediction method provides reliable technical support for resource scheduling and market transactions in virtual power plants and supports migration prediction with limited data, demonstrating promising application prospects.

[0123] In this invention, the data preprocessing step S200 further includes a normalization step S210, an outlier detection step S220, a missing value imputation step S230, and a granularity unification step S240. See below for further details. Figure 2 The data preprocessing step S200 of the present invention will be described.

[0124] In the normalization step S210, the data of influencing factors in multiple dimensions are normalized.

[0125] Specifically, in the normalization step, the min-max normalization method is used to normalize the data of influencing factors across multiple dimensions:

[0126] (1)

[0127] In formula (1), The minimum value of the influencing factor data for each dimension. The maximum value of the influencing factor data for each dimension. The data points are normalized from the influencing factors data for each dimension, where x is the original data point.

[0128] In the outlier detection step S220, outliers in the dataset of influencing factor data for each dimension are detected.

[0129] In this embodiment, the outlier detection step can simultaneously employ three detection methods: the Z-score method, the interquartile range method, and the local outlier factor method.

[0130] In the first outlier detection step, the Z-score method is used to calculate the mean of the influencing factor data for each dimension. and standard deviation Data points that deviate from the mean by more than 3 standard deviations are marked as outliers.

[0131] (2)

[0132] In the second outlier detection step, the interquartile range (IQR) method is used to detect outliers, marking data points that exceed the upper and lower bounds, where the lower bound is... The upper boundary is In the formula It is the first quartile. The third quartile, the interquartile range (IQR) is calculated using the following formula: .

[0133] In the third outlier detection step, the local outlier factor method is used to detect outliers. The local density of each data point is calculated and compared with the density of its neighboring points. If the local density of a data point is lower than the density of its neighboring points by a preset threshold, then the data point is considered outlier. The calculation formula is defined as follows:

[0134] (3)

[0135] In the formula, p is the target data point. It is the k-neighborhood of p. It is the local reachability density of p. The larger the LOF value, the more likely point p is to be an outlier. This method sets the threshold to 0.7.

[0136] The formula for calculating the locally reachable density is as follows:

[0137] (4)

[0138] In formula (4), , d(p,o) is the k-distance between point o and point p and point o.

[0139] In the statistical steps, if a data point is marked as an outlier in at least two of the first, second, and third outlier detection steps, then the data point is confirmed as an outlier. That is, if a data point is marked as an outlier by at least two of the three detection methods (Z-score, box plot, and LOF), it is confirmed as an outlier. Combining multiple detection methods can improve the accuracy of outlier detection.

[0140] In the missing value imputation step S230, missing values ​​in the dataset of influencing factor data for each dimension are imputed.

[0141] The missing value filling step includes:

[0142] This multivariate time series interpolation method based on the attention mechanism normalizes the data, performs nulling after outlier detection, selects a time window for each missing value, and calculates the correlation between the d-th dimension and other dimensions within the time window, as defined below:

[0143] (5)

[0144] In formula (5), For missing values, Let be the mean of the d-th dimension within the time window. Let be the mean of the k-th dimension within the time window. Let d be the correlation score between the d-dimensional and k-dimensional dimensions, and t be the current time step. w represents the historical time step within the time window.

[0145] The relevance scores are converted into attention weights using a softmax function. ,

[0146] (6)

[0147] The missing values ​​are estimated by using attention weights to perform a weighted average of the non-missing values ​​in other dimensions. The formula for missing value estimation is defined as follows:

[0148] (7)

[0149] If the k-th dimension is also missing at time t, then the mean of that dimension within the window is used. Alternative The filled values ​​are then denormalized to the original scale. The denormalization formula is defined as follows:

[0150] (8)

[0151] in, To represent the missing values ​​estimated by the attention mechanism interpolation method at time t and dimension d, This refers to the original scaled data values ​​after denormalization, i.e., the supplemented true values. is the set of all original data of dimension d, used to calculate the minimum and maximum values ​​during the denormalization process.

[0152] In the granularity unification step S240, the time granularity of the influencing factor data of each dimension is aligned according to the sampling frequency difference using either the nearest neighbor interpolation method or the cubic spline interpolation method.

[0153] Specifically, in this embodiment, for data with a sampling frequency higher than the target granularity (such as load, with a granularity of 30 minutes), the original data is directly retained. If the time points are not perfectly aligned, nearest neighbor interpolation is used to select the data point with the closest time interval to fill the gap. For data with a sampling frequency lower than the target granularity (such as weather, with a granularity of 1 hour), cubic spline interpolation is used to generate high-frequency data points, i.e., for low-frequency data points... Fitting a piecewise cubic polynomial ,Require It is continuous at each point and its first and second derivatives.

[0154] As an optional implementation, the time-series feature decomposition step S300 includes:

[0155] First, we receive data on influencing factors across multiple dimensions; this dataset is represented as follows: ,in, This represents the d-th dimension data (such as temperature or load), where T is the time step and d is the number of dimensions. Instance normalization is then performed again.

[0156] (9)

[0157] For each dimension of the normalized data The time series decomposition algorithm is used to decompose it into seasonal cycles, trend windows, and low-pass filter windows, as shown below:

[0158] (10)

[0159] in, This is a trend component, reflecting long-term changes.

[0160] Seasonal components, capturing cyclical fluctuations.

[0161] The residual component contains noise and aperiodic fluctuations.

[0162] This invention uses the STL decomposition algorithm to identify long-term climate change trends, seasonal fluctuations, and short-term anomalies or noise, thus overcoming the shortcomings of traditional time series forecasting methods in effectively capturing trend, seasonal, and residual characteristics when processing multidimensional data from virtual power plants.

[0163] As an optional implementation, the time-series feature decomposition step S300 further includes:

[0164] Set the block length L and stride S, set the preset overlap rate, and independently segment the sequence into blocks for each trend component, seasonal component, and residual component obtained by the time series decomposition algorithm to generate a block set:

[0165] (11)

[0166] When the sequence ends with fewer than L points, zeros are added to ensure that all sequence blocks have the same length.

[0167] The trend component sequence block, seasonal component sequence block, and residual component sequence block within the same time window are concatenated along the feature dimension to form a fused sequence block as a time-series feature vector. The concatenation formula is defined as follows:

[0168] (12)

[0169] In formula (12), For fused sequence blocks, For trend component sequence blocks, For seasonal component sequence blocks, This is a block of residual component sequences.

[0170] In this embodiment, the trend component, seasonal component, and residual component are divided into blocks (with a 50% overlap rate to preserve continuity). The T / S / R blocks of the same time window are spliced ​​along the feature dimension to form a basic token that integrates time-series features.

[0171] As an optional implementation, the dynamic clustering representation enhancement step S400 includes:

[0172] First, the cluster center partitioning step S410 is performed, which involves dividing the word embedding matrix of the pre-trained large language module. The K-means algorithm is used to divide the word embedding matrix into K clusters, and the intra-cluster variance is minimized using the following formula:

[0173] (13)

[0174] in, As the cluster center, K is determined using the elbow method, and the sum of squares within the cluster is calculated:

[0175] (14)

[0176] Where WCSS is the sum of squares within the cluster, plotted... Follow For a changing curve, select the point with the largest change in curvature as the optimal number of clusters. The following conditions must be met:

[0177] (15)

[0178] Further, the cluster center selection step S420 is performed, the silhouette coefficient of each cluster is calculated, and the top clusters are retained. The cluster center with the highest discriminative power ;

[0179] Finally, the feature concatenation step S430 is performed to calculate the temporal feature vector and cluster centers. The cosine similarity is selected. The most relevant cluster centers are concatenated to form a semantically enhanced temporal feature vector. The feature concatenation formula is defined as follows:

[0180] (16)

[0181] in, For semantically enhanced temporal feature vectors, This is a time-series feature vector.

[0182] In this embodiment of the invention, the interpretability of large language models for complex patterns is improved by dynamically associating the LLM semantic space with local temporal features.

[0183] [Methods for constructing and training large language models]

[0184] In this invention, before performing the model fine-tuning step S500, a modeling step S700 is also included. The modeling step S700 further includes a prediction target definition step S710, a text description step S720, an encoding step S730, a feature concatenation step S740, and a position encoding step S750. (Refer to the following...) Figure 3 The modeling step S700 of the present invention will be described.

[0185] In the prediction target definition step S710, global prompts are designed to guide the large language model to focus on the prediction target and establish the association between local features and the global task.

[0186] Local features refer to the trend, seasonal, and residual components after STL decomposition, as well as the time-series tokens enhanced by dynamic clustering, which only describe local patterns in the data (e.g., "seasonal fluctuations in load at the current moment"). The global task requires integrating all local features to output a comprehensive forecast of "load and electricity price for the next 24 hours," informing the model what to predict. Designing global hints and encoding them as hint tokens is a crucial step in enabling large language models to focus on specific forecasting tasks. Its core purpose is to guide the model to focus on the target (e.g., load and electricity price forecasts) through natural language descriptions and to establish semantic connections between local time-series features (e.g., trend and seasonal components) and the global task.

[0187] In the text description step S720, the task description text is entered.

[0188] By inputting task description text (e.g., "predict electricity prices"), relevant knowledge stored in the model (e.g., "load is affected by weather" and "electricity prices are related to supply and demand") can be activated. However, native LLM lacks direct understanding of time-series data (numerical data). It is necessary to use text prompts to transform abstract prediction tasks (e.g., "the next 24 hours") into semantic instructions that the model can parse, and establish a mapping from numerical features to task objectives.

[0189] In encoding step S730, the word segmenter of the large language model is used to encode the prompt word. Then the prompt words Mapped to cue embedding vector .

[0190] In the feature concatenation step S740, the cue embedding vector is... With semantically enhanced temporal feature vectors The input sequence is obtained by concatenating along the sequence dimension. The splicing formula is defined as follows:

[0191] (17)

[0192] In the position encoding step S750, absolute position encoding is injected into the input sequence to preserve the temporal order information:

[0193] (18)

[0194] in, For the position encoding matrix, The input sequence is .

[0195] Specifically, the cue token is usually placed at the very beginning of the input sequence (e.g., at the start of the sequence). The attention mechanism of the Transformer allows the large language model to prioritize the processing of task semantics, which is then combined with the subsequent temporal token (semantic-enhanced temporal feature vector). The dimension of the cue embedding must be consistent with the dimensions of the temporal token and the positional encoding to ensure compatibility of the concatenation operation (e.g., concatenation along the sequence dimension).

[0196] As an optional implementation, the model fine-tuning step S500 includes:

[0197] With the backbone parameters of the pre-trained large language model frozen, we optimize the positional embedding, the weights of the feedforward neural network (FFN) in the residual connection, and the scaling factor and bias in the layer normalization layer. The parameter optimization formula for the normalization layer is defined as follows:

[0198] (19)

[0199] in, The mean of the input data. The variance of the mean of the input data. The scaling factor in the normalization layer. The bias in the normalization layer. To hide the input state, To output the hidden state.

[0200] This invention optimizes residual connections, normalization layers, and position embedding parameters in a targeted manner while freezing the backbone parameters of the pre-trained LLM, enabling the large language model to adapt to the prediction task of a virtual power plant without requiring overall training of the LLM, thus shortening the training time.

[0201] As an optional implementation, modeling step S700 further includes:

[0202] In sample setting step S760, positive samples and negative samples are set. Positive samples are the K time-series feature vectors that match the Top-K cluster centers, and negative samples are other time-series feature vectors in the same batch that do not belong to the Top-K cluster centers.

[0203] Step S770, setting the contrastive learning loss function, is based on the optimization objective of narrowing the distance between positive samples and widening the distance between negative samples. The contrastive learning loss function, used to semantically align the temporal feature vectors with the large language model, is defined as follows:

[0204] (20)

[0205] in, As a positive sample, The cosine similarity function is used. This is a temperature coefficient (default 0.1) that controls the steepness of the distribution. is the number of negative samples, and M is the total number of samples;

[0206] The model training in this invention introduces a contrastive learning strategy, which uses the InfoNCE loss function (contrastive learning loss function) to bring positive samples closer and push negative samples further apart.

[0207] Step S780: Set the training loss function based on the contrastive learning loss function, as defined below:

[0208] ,(twenty one)

[0209] in, To predict losses, To compare learning loss, Weights for predicting loss (mean squared error loss, MSE), The weights are used to compare the learning loss (Lcontrast).

[0210] In the above formula (21), the first part of the training loss function is the prediction loss, which is used to improve the prediction accuracy, and the second part is used to align the temporal sequence with the semantics of the large model and enhance the consistency of representation.

[0211] As an optional implementation, the model output step includes:

[0212] Position encoding matrix is ​​obtained by performing position encoding on semantically enhanced temporal feature vectors. The input is fed into a pre-trained large language model, which generates hidden states H through multiple Transformer layers. The hidden states of the corresponding temporal segments are then extracted based on the cue positions. Hidden state is hidden through a fully connected layer. Mapped to predicted values The result is then denormalized to its original dimensions and used as the prediction result. The denormalization formula is defined as follows:

[0213] ,(twenty two)

[0214] in, For predicted values, Let d be the mean of the original data in the d-th dimension. Let d be the standard deviation of the original data in the d-th dimension. This is the predicted value of the d-th dimension after inverse normalization.

[0215] Finally, through reasoning and output parsing of the large language model, the system completes the multivariate input and output of load and electricity price in the context of a virtual power plant.

[0216] The prediction accuracy of the virtual power plant parameter prediction method is verified by designing an ablation experiment and comparing parameters.

[0217] Experimental Verification of the Prediction Method for Virtual Power Plant Parameters

[0218] The ablation experiments and parameter comparisons were designed, including: verification of the effectiveness of multi-component block stitching, i.e., a comparison experiment between directly splitting the original time series data into blocks (without STL decomposition) and using trend, seasonal, and residual block stitching; comparison of the learning part by changing the loss function, and comparing the impact of forcibly aligning the time series tokens with the semantic anchors of the large model on the prediction accuracy; comparison of canceling fine-tuning, freezing all parameters and fine-tuning the parameters; and verification of the feasibility of the model design.

[0219] The prediction method of this invention was validated using publicly available datasets released by the Australian Energy Market Operator (AEMO).

[0220] Experimental setup: Forecasting tasks (pred_len): 24-hour (short-term), 48-hour (medium-term), and 168-hour (long-term) load and electricity price forecasts. Input sequence length (seq_len): 24 (short-term), 48 (medium-term), and 168 (long-term).

[0221] For electricity price parameters, the various prediction accuracy indicators are shown in Table 1:

[0222]

[0223] Table 1. Comparison of Electricity Price Forecasting Performance (MAE / RMSE)

[0224] For the load parameters, the various prediction accuracy indicators are shown in Table 2:

[0225]

[0226] Table 2. Comparison of Load Forecasting Performance (MAE / RMSE)

[0227] Reference Figures 4A to 4F The image shows a comparison of the Mean Absolute Error (MAE) of the LLM model of this invention with that of various traditional models. MAE is a commonly used metric for measuring the performance of a prediction model. It is evaluated by calculating the average of the absolute differences between the predicted and actual values. Figures 4A to 4FIn this study, the MAE of the LLM model of this invention decreased by 26.3%, 26.3%, 25%, 28.9%, 21%, and 25.5% respectively compared with the MAE predicted by the traditional model. Figure 5A and 5B These are two sets of bar charts comparing load forecasting performance and electricity price forecasting performance. The first chart compares load forecasting and electricity price forecasting (input sequence length seq_len=24, predicted sequence length pred_len=24). The five bars in the charts represent the MAE and RMSE error values ​​for the traditional method without STL decomposition (leftmost bar), trend component, seasonal component, residual component, and the multi-component concatenation method of this invention (rightmost bar). The charts also show the performance improvement of the multi-component concatenation method, demonstrating its superiority over single-component and no-decomposition methods on the AEMO dataset. Figure 5A In load forecasting, the multi-component splicing method of the present invention ( Figure 5A The rightmost method (MAE) has a value of 0.26 and an RMSE of 0.43, with significantly lower errors than other methods, resulting in a performance improvement of 31.6%. Figure 5B In electricity price forecasting, the multi-component splicing method of the present invention ( Figure 5B The MAE (on the far right) is 0.3, RMSE is 0.48, and the error is significantly lower than other methods, resulting in a performance improvement of 33.3%. In summary, the STL decomposition and block splicing method of this invention is significantly superior to traditional methods without STL decomposition and other single-component methods.

[0228] like Figure 6 As shown in the figure, the effect of changing the parameters of the contrastive learning loss function is compared. In the experiment comparing the input parameters of the contrastive learning loss function, When =0, it is equivalent to learning without comparison. An excessively large sample size can lead to an overemphasis on sample comparisons and neglect of the prediction task; optimal... The value is around 0.3. The results of GPT-2 experiments with fine-tuning and full freezing are as follows: Figure 7A and 7B As shown, the left side represents the error value without fine-tuning (i.e., freezing all model parameters), while the right side represents the error value after fine-tuning the model according to the invented fine-tuning method. In the fine-tuning strategy comparison experiment, fine-tuning the model parameters can significantly improve the model performance, with the load MAE decreasing by 31.6%, the electricity price MAE decreasing by 27.3%, the load RMSE decreasing by 26.2%, and the electricity price RMSE decreasing by 23.5%.

[0229] In summary, this invention proposes a prediction method for virtual power plant parameters (such as load and electricity price) based on a large language model by employing techniques such as multi-component block partitioning, representation enhancement, and large language model fine-tuning. This method significantly improves the prediction accuracy and real-time performance of multi-dimensional time-series data of virtual power plants, and can provide high-precision, low-latency prediction support for virtual power plant resource scheduling and market transactions, thus possessing significant engineering application value.

[0230] In addition, the present invention also provides a device for predicting virtual power plant parameters.

[0231] [Software Structure of the Prediction Device for Predicting Operating Parameters of a Virtual Power Plant According to the Invention]

[0232] The following is for reference Figure 8 The software structure of the prediction device for predicting operating parameters of a virtual power plant according to the present invention will be described. For example... Figure 8 As shown, the prediction device 800 for predicting the operating parameters of a virtual power plant according to the present invention includes a data acquisition unit 801, a data preprocessing unit 802, a time-series feature decomposition unit 803, a dynamic clustering representation enhancement unit 804, a model fine-tuning unit 805, and a model output unit 806. The processing performed by each unit is described in detail below.

[0233] like Figure 8 As shown, firstly, the data acquisition unit 801 acquires data on multiple dimensions of influencing factors used to predict the operating parameters (such as load and electricity price) of the virtual power plant.

[0234] The multi-dimensional influencing factor data in this invention can include weather data (such as temperature, humidity, air pressure, wind speed, and solar radiation), calendar information (such as dates and holidays), historical controllable load sequences (such as regional load demand sequences), electricity price sequences (such as regional electricity market price sequences), distributed energy output sequences (such as regional wind and solar renewable energy output and V2G load), and uncontrolled loads (such as inter-regional transmission loads). Through this influencing factor data, key operating parameters such as load and electricity price of the virtual power plant can be comprehensively predicted from multiple perspectives.

[0235] Then, the data preprocessing unit 802 performs preprocessing on the multi-dimensional influencing factor data, including normalization and data cleaning.

[0236] The data preprocessing unit 802 normalizes the influencing factor data (such as temperature, humidity, wind speed, etc.) for each dimension, ensuring that data of different dimensions are on the same scale for easy comparison. Data cleaning identifies erroneous data in the influencing factor data and is a process of re-checking and verifying the data. The goal of data cleaning is to remove duplicate records, eliminate abnormal data, correct erroneous data, ensure data consistency, and improve data quality.

[0237] Furthermore, the time series feature decomposition unit 803 uses a time series decomposition algorithm to decompose each dimension of influencing factor data into a three-dimensional sequence containing trend component, seasonal component and residual component, divides each sequence into sequence blocks, and maps the sequence blocks into time series feature vectors that can be used as input to a large language model.

[0238] Specifically, the time series feature decomposition unit 803 uses the STL algorithm to decompose the time series data. First, it acquires data on influencing factors across multiple dimensions and decomposes them into three components using the STL algorithm: seasonal cycle, trend window, and low-pass filter window. The trend component reflects long-term changes, the seasonal component captures periodic fluctuations, and the residual component contains noise and aperiodic fluctuations. The STL decomposition algorithm is highly flexible, applicable to various seasonal patterns, and can identify long-term climate change trends, seasonal fluctuations, and short-term anomalies or noise.

[0239] Furthermore, after STL decomposition to obtain the trend component (T), seasonal component (S), and residual component (R), the time series of each dimension is divided into blocks. For each component, an independent block set is generated, and then the patch is mapped to a temporal feature vector (i.e., a token) that can be recognized by a Large Language Model (LLM). LLM is essentially a model for processing sequential data and cannot directly understand raw data such as pixel values ​​or video frames. Mapping patches of non-textual data such as images or videos to tokens, which are then input into the model for processing, primarily aims to transform unstructured data into a structured semantic representation that the model can understand, thereby achieving the fusion and processing of multimodal information.

[0240] Next, the dynamic clustering representation enhancement unit 804 is used to perform K-means clustering on the word embeddings used for pre-training the large language model. It selects the cluster center representation and performs similarity matching with the temporal feature vector, selects the K cluster centers with high scores as semantic anchors, and concatenates them with the temporal feature vector to obtain the semantically enhanced temporal feature vector.

[0241] Specifically, for the word embedding matrix of the pre-trained LLM, the K-means algorithm is used to divide the word embedding matrix into K clusters. The Top-K cluster centers most relevant to the temporal token (i.e., temporal feature vector) are selected and concatenated to form the enhanced token (i.e., semantically enhanced temporal feature vector). By dynamically associating the LLM semantic space with local temporal features, semantic features are integrated into the temporal features, improving the interpretability of large language models for complex patterns. By decomposing temporal features using STL and combining them with LLM semantic anchors, cross-modal enhancement from numerical features to semantic features is achieved, overcoming the LLM's obstacle to understanding purely numerical data.

[0242] Furthermore, the model fine-tuning unit 805 is used to input the semantically enhanced temporal feature vector into the pre-trained large language model and fine-tune the position embedding parameters, the weights of the feedforward neural network of the residual connection, and the parameters of the normalization layer of the large language model.

[0243] During training, the model fine-tuning unit 805 optimizes residual connections, normalization layers, and position embedding parameters while freezing the pre-trained LLM backbone parameters. This allows the large language model to adapt to the prediction task of virtual power plant parameters, such as predicting key indicators like electricity prices and loads. In this embodiment, only parameters strongly correlated with time-series characteristics are optimized, preserving the semantic understanding capabilities of the pre-trained LLM and avoiding redundant training. This significantly reduces computational costs while retaining pre-training capabilities, thus meeting the real-time requirements of virtual power plants.

[0244] Finally, the model output unit 806 is used to predict the trend component, seasonal component and residual component of each dimension based on the semantically enhanced temporal feature vector, and to concatenate and inversely normalize the trend component, seasonal component and residual component under the same dimension to obtain the prediction result, which is used as the output of the large language model.

[0245] As mentioned earlier, the temporal feature decomposition unit 803 uses the STL decomposition algorithm to divide the influencing factor data of each dimension into three components: trend, seasonality, and residual. Each component is processed independently after being segmented. These three components are segmented and concatenated with semantic anchors before being input into the LLM. When the LLM outputs, these components need to be recombine because the prediction target is the data of the original dimension, such as load and electricity price. Each dimension has its own three components, and the three components corresponding to each dimension also need to be concatenated according to the dimension during output. For example, the trend, seasonality, and residual components of load are concatenated to form the load prediction sequence, and the prediction sequences of electricity price or other dimensions are processed in the same way. The purpose of doing this is to maintain the time alignment of each component, ensuring that the trend, seasonality, and residual correspond in time and that the combined result can correctly reflect the pattern of the original data. Then comes denormalization. Normalization was used in the previous data preprocessing steps, so the normalized values ​​predicted by the model need to be converted back to the original units during output, such as converting the load to "MW" or the electricity price to "yuan / kWh". This step is very important because users need actual meaningful values, not normalized values.

[0246] The prediction device in this invention overcomes the shortcomings of traditional time-series prediction devices in effectively capturing trend, seasonal, and residual features when processing multidimensional data from virtual power plants. Secondly, it is the first to apply word embedding clustering of a pre-trained large language model to enhance the representation of time-series data. Through prediction and contrastive learning mechanisms based on a semantic anchor matching strategy, it significantly improves the model's ability to understand complex time-series patterns, overcoming the deficiencies of existing methods in semantic association and dynamic adaptability. Finally, with prediction accuracy and computational efficiency as optimization goals, the virtual power plant prediction model combines residual connections, normalization layers, fine-tuning of position embeddings, and an INT8 quantization deployment strategy to optimize the model's inference efficiency on edge devices, ensuring real-time prediction and effectively solving the problem of real-time inference of large models on edge devices.

[0247] The forecasting device of this invention can comprehensively improve the accuracy and real-time performance of virtual power plant load and electricity price forecasts, significantly enhancing the accuracy and efficiency of multi-dimensional time-series data forecasting for virtual power plants. This has significant practical implications for the intelligent management and optimized operation of virtual power plants. The forecasting method provides reliable technical support for resource scheduling and market transactions in virtual power plants and supports migration forecasting with limited data, demonstrating promising application prospects.

[0248] [Other Implementation Methods]

[0249] Embodiments of the invention can also be implemented by a computer that reads and executes computer-executable instructions (e.g., one or more programs) recorded on a storage medium (also more fully referred to as a "non-transitory computer-readable storage medium") to perform one or more functions in the above embodiments, and / or includes one or more circuits (e.g., application-specific integrated circuits (ASICs)) for performing one or more functions in the above embodiments. Furthermore, embodiments of the invention can be implemented using a method by which the computer of the system or device, for example, reads and executes the computer-executable instructions from the storage medium to perform one or more functions in the above embodiments, and / or controls the one or more circuits to perform one or more functions in the above embodiments. The computer may include one or more processors (e.g., a central processing unit (CPU), a microprocessor unit (MPU)) and may include separate computers or a network of separate processors to read and execute the computer-executable instructions. The computer-executable instructions may be provided to the computer, for example, from a network or the storage medium. The storage medium may include one or more of the following: hard disk, random access memory (RAM), read-only memory (ROM), memory of a distributed computing system, optical disc (such as compressed optical disc (CD), digital versatile optical disc (DVD) or Blu-ray disc (BD)™), flash memory device, and memory card.

[0250] While the present invention has been described above with reference to exemplary embodiments, these embodiments are only for illustrating the technical concept and features of the present invention and should not be construed as limiting the scope of protection of the present invention. Any equivalent variations or modifications made in accordance with the spirit and essence of the present invention should be covered within the scope of protection of the present invention.

Claims

1. A method for predicting operating parameters of a virtual power plant, characterized in that, Includes the following steps: The data acquisition step involves acquiring multi-dimensional influencing factor data for predicting the operating parameters of the virtual power plant, wherein the multi-dimensional influencing factor data includes at least weather information and distributed energy output sequences; The data preprocessing step involves performing preprocessing on the data of influencing factors across the multiple dimensions, including normalization and data cleaning. The time series feature decomposition step uses a time series decomposition algorithm to decompose each dimension of influencing factor data into a three-dimensional sequence containing trend component, seasonal component and residual component. Each sequence is divided into sequence blocks, and the sequence blocks are spliced ​​along the feature dimension to form a fused sequence block as the time series feature vector input to the large language model. The dynamic clustering representation enhancement step involves performing K-means clustering on the word embeddings used to pre-train the large language model, selecting cluster center representations and matching them with the temporal feature vectors for similarity, selecting the K cluster centers with high scores as semantic anchors, and concatenating them with the temporal feature vectors to obtain a semantically enhanced temporal feature vector. The model fine-tuning step involves inputting the semantically enhanced temporal feature vector into the pre-trained large language model and fine-tuning the position embedding parameters, the weights of the feedforward neural network with residual connections, and the parameters of the normalization layer of the large language model. The model output step involves predicting the trend component, seasonal component, and residual component of each dimension based on the semantically enhanced temporal feature vector. The trend component, seasonal component, and residual component under the same dimension are then concatenated and inversely normalized to obtain the prediction result, which is used as the output of the large language model.

2. The prediction method according to claim 1, characterized in that, In the data acquisition step, the data on influencing factors across multiple dimensions further includes at least one of the following data items: Calendar information, historical controllable load sequences, electricity price sequences, electric vehicle charging and discharging loads, and uncontrolled loads.

3. The prediction method according to claim 1, characterized in that, The preprocessing performed in the data preprocessing step includes the following steps: The normalization step involves normalizing the data of influencing factors across the multiple dimensions. The outlier detection step detects outliers in the dataset of influencing factor data for each dimension. The missing value imputation step is used to impute missing values ​​in the dataset of the influencing factors for each dimension. In the granularity unification step, based on the differences in sampling frequency, the time granularity of the influencing factor data in each dimension is aligned using either the nearest neighbor interpolation method or the cubic spline interpolation method.

4. The prediction method according to claim 3, characterized in that, In the normalization step, the min-max normalization method is used to normalize the data of the influencing factors in the multiple dimensions: , in, The minimum value of the influencing factor data for each dimension. For each dimension The maximum value of the data on the influencing factors of degree. The data points are normalized from the influencing factors data for each dimension, where x is the original data point.

5. The prediction method according to claim 3, characterized in that, The outlier detection steps include: The first outlier detection step uses the Z-score method to calculate the mean of the influencing factor data for each dimension. and standard deviation Data points that deviate from the mean by more than 3 standard deviations are marked as outliers. , The second outlier detection step uses the interquartile range (ICM) method to detect outliers, marking data points that exceed the upper and lower bounds. The lower bound is... The upper boundary is In the formula It is the first quartile. The third quartile, the interquartile range (IQR) is calculated using the following formula: ; The third outlier detection step uses the local outlier factor method to detect outliers. It calculates the local density of each data point and compares it to the density of its neighboring points. If the local density of a data point is lower than the density of its neighboring points by a preset threshold, then the data point is considered outlier. The calculation formula is defined as follows: , In the formula, p For the target data point, It is the k-neighborhood of p. yes p Locally achievable density, The formula for calculating the locally reachable density is as follows: , In the formula, , It is a point o of k -distance, d ( p , o ) is a point p and points o The Euclidean distance between them; In the statistical steps, if at least two of the first, second, and third outlier detection steps are marked as outliers for a certain data point, then the data point is confirmed as an outlier.

6. The prediction method according to claim 3, characterized in that, The missing value filling step includes: This multivariate time series interpolation method based on the attention mechanism normalizes the data, performs nulling after outlier detection, selects a time window for each missing value, and calculates the correlation between the d-th dimension and other dimensions within the time window, as defined below: , in, For missing values, Let be the mean of the d-th dimension within the time window. Let be the mean of the k-th dimension within the time window. Let d be the correlation score between the d-dimensional and k-dimensional dimensions, and t be the current time step. w represents the historical time step within the time window. The relevance scores are converted into attention weights using a softmax function. , , The missing values ​​are estimated by using attention weights to perform a weighted average of the non-missing values ​​in other dimensions. The formula for missing value estimation is defined as follows: , If the first k Dimension in time t If it is also missing, then use the mean of that dimension within the window. Alternative The filled values ​​are then denormalized to the original scale. The denormalization formula is defined as follows: , in, To represent the missing values ​​estimated by the attention mechanism interpolation method at time t and dimension d, This refers to the original scaled data values ​​after denormalization, i.e., the supplemented true values. is the set of all original data of dimension d, used to calculate the minimum and maximum values ​​during the denormalization process.

7. The prediction method according to claim 3, characterized in that, The granularity unification step includes: For data with a sampling frequency higher than the target granularity, the original data is directly retained. If the time points are not perfectly aligned, nearest neighbor interpolation is used to select the data point that is closest in time to fill the gap. For data with a sampling frequency lower than the target granularity, cubic spline interpolation is used to generate high-frequency data points, i.e., for low-frequency data points... Fitting a piecewise cubic polynomial ,Require It is continuous at each point and its first and second derivatives.

8. The prediction method according to claim 1, characterized in that, The time-series feature decomposition step includes: Receive data on influencing factors from multiple dimensions ,in Let d represent the data of dimension d, where T is the time step and d is the number of dimensions. Then, perform instance normalization again: , For each dimension of the normalized data The time series decomposition algorithm is used to decompose it into seasonal cycles, trend windows, and low-pass filter windows, as shown below: , in, This is a trend component, reflecting long-term changes. To capture seasonal fluctuations, The residual component contains noise and aperiodic fluctuations.

9. The prediction method according to claim 1 or 8, characterized in that, The time-series feature decomposition step further includes: Set block length L stride S Set a preset overlap rate, and independently segment the sequence into blocks for each trend component, seasonal component, and residual component obtained by the time series decomposition algorithm, generating the following block set: , Among them, the end is insufficient L Zero-padding is used to ensure that all sequence blocks have the same length. The trend component sequence block, seasonal component sequence block, and residual component sequence block within the same time window are concatenated along the feature dimension to form a fused sequence block as the time-series feature vector. The concatenation formula is defined as follows: , in, For fused sequence blocks, For trend component sequence blocks, For seasonal component sequence blocks, This is a block of residual component sequences.

10. The prediction method according to claim 1, characterized in that, The dynamic clustering representation enhancement steps include: The clustering center partitioning steps are for the word embedding matrix of the pre-trained large language module. The K-means algorithm is used to divide the word embedding matrix into K clusters, and the intra-cluster variance is minimized using the following formula: , in, As the cluster center, K is determined using the elbow method, and the sum of squares within the cluster is calculated: , Where WCSS is the sum of squares within the cluster, plotted... Follow For a changing curve, select the point with the largest change in curvature as the optimal number of clusters. The following conditions must be met: , The cluster center selection step involves calculating the silhouette coefficient of each cluster and retaining the top clusters. The cluster center with the highest discriminative power ; The feature concatenation step involves calculating the temporal feature vector and the cluster centers. The cosine similarity is selected. The most relevant cluster centers are concatenated to form the semantically enhanced temporal feature vector. The feature concatenation formula is defined as follows: , in, For semantically enhanced temporal feature vectors, This is a time-series feature vector.

11. The prediction method according to claim 1, characterized in that, Prior to the model fine-tuning step, a modeling step is also included, which includes: The steps for defining the prediction target include designing global prompts, guiding the large language model to focus on the prediction target, and establishing the relationship between local features and the global task. Text description steps: Enter task description text; The encoding step involves using a large language model's word segmenter to encode prompt words. Then the prompt words Mapped to cue embedding vector ; The feature concatenation step embeds the cue into the vector. With the semantically enhanced temporal feature vector The input sequence is obtained by concatenating along the sequence dimension. The splicing formula is defined as follows: , The positional encoding step involves injecting absolute positional codes into the input sequence to preserve temporal order information. , in, For the position encoding matrix, The input sequence is .

12. The prediction method according to claim 11, characterized in that, The model fine-tuning steps include: With the backbone parameters of the pre-trained large language model frozen, we optimize the position embeddings, the weights of the feedforward network in the residual connections, and the scaling factor and bias in the normalization layer. The parameter optimization formula for the normalization layer is defined as follows: , in, The mean of the input data. The variance of the mean of the input data. The scaling factor in the normalization layer. The bias in the normalization layer. To hide the input state, To output the hidden state.

13. The prediction method according to claim 11, characterized in that, The modeling steps also include: The sample setup steps involve setting positive and negative samples. Positive samples are the K temporal feature vectors that match the Top-K cluster centers, while negative samples are other temporal feature vectors in the same batch that do not belong to the Top-K cluster centers. The steps for setting the contrastive learning loss function are as follows: Based on the optimization objectives of narrowing the distance between positive samples and widening the distance between negative samples, a contrastive learning loss function is set to semantically align the temporal feature vector with the large language model, defined as follows: , in, As a positive sample, The cosine similarity function is used. The temperature coefficient controls the steepness of the distribution. M represents the number of negative samples, and M represents the total number of samples. The steps for setting the training loss function are as follows: Based on the contrastive learning loss function, the training loss function is defined as follows: , in, To predict losses, To compare learning loss, Weights for predicting loss, The weights are used to compare the learning loss.

14. The prediction method according to claim 12, characterized in that, The model output steps include: The semantically enhanced temporal feature vector is positionally encoded to obtain a positional encoding matrix. The input is fed into a pre-trained large language model, which generates hidden states H through multiple Transformer layers. The hidden states of the corresponding temporal segments are then extracted based on the cue positions. Hidden state is hidden through a fully connected layer. Mapped to predicted values The result is then denormalized to its original dimensions and used as the prediction result. The denormalization formula is defined as follows: , in, For predicted values, Let d be the mean of the original data in the d-th dimension. Let d be the standard deviation of the original data in the d-th dimension. This is the predicted value of the d-th dimension after inverse normalization.

15. A predictive device for predicting operating parameters of a virtual power plant, characterized in that, include: The data acquisition unit is used to acquire multi-dimensional influencing factor data for predicting the operating parameters of the virtual power plant, wherein the multi-dimensional influencing factor data includes at least weather information and distributed energy output sequence; The data preprocessing unit is used to perform preprocessing on the data of influencing factors in the multiple dimensions, including normalization and data cleaning. The temporal feature decomposition unit is used to decompose each dimension of influencing factor data into a three-dimensional sequence containing trend component, seasonal component and residual component using a time series decomposition algorithm. Each sequence is divided into sequence blocks, and the sequence blocks are spliced ​​along the feature dimension to form a fused sequence block as the temporal feature vector input to the large language model. The dynamic clustering representation enhancement unit is used to perform K-means clustering on the word embeddings used to pretrain the large language model, select the cluster center representations and perform similarity matching with the temporal feature vector, select the K cluster centers with high scores as semantic anchors, and concatenate them with the temporal feature vector to obtain the semantically enhanced temporal feature vector. The model fine-tuning unit is used to input the semantically enhanced temporal feature vector into the pre-trained large language model and fine-tune the position embedding parameters, the weights of the feedforward neural network of the residual connection, and the parameters of the normalization layer of the large language model. The model output unit is used to predict the trend component, seasonal component and residual component of each dimension based on the semantically enhanced temporal feature vector, and to concatenate and inversely normalize the trend component, seasonal component and residual component under the same dimension to obtain the prediction result, which is used as the output of the large language model.

16. A non-transitory storage medium storing a computer program that, when executed by a processor, enables the prediction method for virtual power plant parameters according to any one of claims 1-14.

17. A computer program product comprising computer instructions that, when executed by a processor, enable the prediction method for virtual power plant parameters according to any one of claims 1-14.

Citation Information

Patent Citations

  • Power system anomaly prediction method based on machine learning and big data analysis

    CN112084237A

  • Personalized soft skill training method and system based on artificial intelligence

    CN120336498A