Transformer area distributed photovoltaic power ultra-short-term prediction method and system
By combining agglomerative hierarchical clustering with GraphSAGE graph neural network and Transformer encoder, the problems of insufficient prediction accuracy and limited cross-regional adaptability in distributed photovoltaic systems are solved, and efficient and low-cost photovoltaic power prediction is achieved to support grid stability analysis.
Patent Information
- Application Number
- CN202510971656.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-15
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-07-15
AI Technical Summary
When dealing with distributed photovoltaic systems, existing photovoltaic power prediction methods have problems such as insufficient prediction accuracy, limited cross-regional adaptability, insufficient utilization of meteorological collaborative information, and insufficient characterization of sub-regional heterogeneity. These problems lead to large errors in the prediction results and make it difficult to meet the needs of grid scheduling.
An agglomerative hierarchical clustering algorithm is used to divide the photovoltaic power generation area into sub-areas with consistent output. The GraphSAGE graph neural network and the Transformer encoder model are combined to achieve decoupled representation of spatiotemporal features. A migration modeling strategy is constructed by sharing features to quickly and lightweight model the photovoltaic power prediction of each sub-area.
It improves the accuracy and consistency of photovoltaic power forecasts, reduces modeling costs, adapts to meteorological characteristics and grid topology differences in different geographical regions, realizes rapid collaborative forecasting of massive substations, and supports regional power dispatching and grid stability analysis.
Smart Images

Figure CN120806264A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to a transformer area distributed photovoltaic power ultra-short-term prediction method and system, and belongs to the technical field of distributed photovoltaic power prediction. BACKGROUND
[0002] As a key path of clean energy transformation, distributed photovoltaic power generation is accelerating penetration into the end of the distribution network in a large-scale trend. This trend, while bringing many advantages to the power grid, also poses new challenges to the safe operation and regulation and management of the power grid. Photovoltaic output is easily affected by weather factors and has strong randomness and volatility, and the confidence of the prediction results is low. This uncertainty exacerbates the short-term load fluctuation of the distribution network, significantly increasing the difficulty of load prediction. Because the massive low-voltage distributed photovoltaic power is still in the "blind adjustment state" that is unobservable and uncontrollable, it further exacerbates the difficulty of power grid consumption and increases the adjustment pressure of power balance. In order to accurately depict the distribution of photovoltaic power generation in the distribution network, high-precision power prediction technology for low-voltage distributed photovoltaic power is urgently needed.
[0003] Existing photovoltaic power prediction methods can be divided into physical modeling, statistical methods, and data-driven methods. Physical methods mainly focus on the internal physical modules of photovoltaic power generation systems, establish mathematical models, and directly calculate photovoltaic power generation output. Statistical methods establish photovoltaic power output prediction mathematical models based on statistical rules in acquired meteorological data and historical output data. The prediction models established by this method are relatively simple, but the prediction accuracy and stability are relatively poor. Data-driven methods achieve power prediction by mining the nonlinear mapping relationship between meteorological data and power data. This method can effectively capture the complex nonlinear relationship between meteorological and power factors, and the prediction performance is often good, so it has become the most mainstream method at present. Although the above methods have achieved good prediction results, they mainly rely on the data of a single station for modeling, resulting in prediction models that are only applicable to specific stations and are prone to overfitting. Therefore, these methods are not completely suitable for the distributed photovoltaic power prediction application scenario of "many points and wide surfaces".
[0004] It is worth noting that due to the similar fluctuation trend of the distributed photovoltaic output with close spatial position, the output is also inherently autocorrelated in the time dimension, so the existing researches are mostly based on the analysis of the spatio-temporal correlation of the distributed photovoltaic output and the centralized photovoltaic output and the distributed photovoltaic output. The method of predicting the target station output by using the output of the adjacent centralized photovoltaic station is relatively simple, but it does not consider the strength of the correlation between the stations, and in practice, it is impossible to ensure that there are centralized photovoltaic stations around all distributed photovoltaic stations, so the applicability is limited. In contrast, making full use of the output correlation characteristics between the distributed photovoltaic stations for modeling can not only effectively improve the prediction accuracy, but also has a relatively wide range of application. However, there are still some limitations in such methods at present: (1) Insufficient description of the heterogeneity of sub-regions: ignoring the differences in the output distribution characteristics of different small areas, the same set of prediction systems is used for all the stations in the region, which leads to the fact that the model focuses on global characteristics and ignores local characteristics of sub-regions, and a single global model cannot effectively capture these differences, thereby increasing the error of the prediction results. (2) Inadequate use of meteorological coordination information: only relying on historical output data to mine the spatio-temporal correlation characteristics, without effectively using the prediction meteorological coordination information in the region, ignoring the important influence of the prediction meteorological on the ultra-short-term power prediction, and the prediction model based on historical data alone cannot fully capture the real-time feedback of meteorological changes on power fluctuations, which may lead to large deviations in the prediction results. (3) Limited adaptability across regions: most power prediction models focus on specific stations in a particular region, and it is difficult to adapt to the differences in meteorological characteristics and power grid topology in different geographical regions, which limits the range of engineering applications. In summary, there is an urgent need for a low-cost and efficient technology that can quickly build a targeted prediction model for a large number of photovoltaic stations, while fully utilizing the power generation synergy between the stations, to accurately aggregate the power of each station, in order to better serve the regional power dispatching and power grid stability analysis. SUMMARY
[0005] In view of the deficiencies of the prior art, the present application provides a kind of substation distributed photovoltaic power ultra-short-term prediction method and system, based on condensation hierarchical clustering algorithm, the photovoltaic of station in region is clustered into the homogeneous sub-region set of output characteristics, combining GraphSAGE graph neural network and Transformer encoder model, realize the decoupling representation of the space-time characteristics in sub-region, based on composite space-time characteristics finally synchronously output the photovoltaic power prediction result of each station in sub-region, further based on the shared characteristics between sub-regions, build suitable migration modeling strategy, realize the quick lightweight modeling of each sub-region prediction model, finally obtain the regional total power prediction result through the space aggregation of sub-region prediction power.
[0006] The technical scheme of the present application is as follows: A substation distributed photovoltaic power ultra-short-term prediction method, the steps are as follows: (1) Collect the historical measured power data of all sub-areas in the region to be predicted and the meteorological data of the corresponding position of each sub-area extracted through the measured reanalysis data, and perform data processing; (2) Calculate the dynamic time warping distance using the historical measured power data of the sub-area as the similarity measurement index between the output of the sub-area, and divide the photovoltaic sub-areas in the region into sub-regions with consistent output by using the condensed hierarchical clustering algorithm; (3) Construct a collaborative architecture of Transformer and GraphSAGE, realize decoupling of time and space features and tensor fusion, capture time-varying features by using Transformer self-attention, model spatial super-short-term meteorological correlation by using GraphSAGE neighborhood aggregation, fit the composite time and space features by using BiLSTM model, and simultaneously output the photovoltaic power prediction results of each sub-area in the sub-region, avoiding repeated modeling for each sub-area; (4) Based on the shared features between the sub-regions, a suitable transfer modeling strategy is constructed, the time and space feature extraction backbone network parameters are frozen, only the region-specific connection layer is subjected to lightweight adaptation operation, rapid and low-cost modeling of each sub-region is realized, and finally the total power of the region is obtained by accumulating the power of each sub-region.
[0007] According to the application, in step (1), the data processing specifically includes missing value processing, abnormal value processing and normalization processing; The KNN algorithm is selected as the missing value repair method, and the mathematical expression is as follows: (1) (2) In the formula: and are two sample points, each having n characteristic values, and are the values of points and on the first characteristic, is the value of the K closest sample to the missing value, is the Euclidean distance between the two samples, used to select the K nearest neighbors of the missing data point, and the missing value is filled with the average value of the K nearest neighbors.
[0008] The abnormal value processing adopts the box plot method (IQR) to detect abnormal values; The normalization processing adopts the maximum and minimum value standard method to standardize the sample data, and the sample data is linearly mapped to [0, 1]; (3) In the formula: represents the normalized sample data; represents the sample data to be normalized; and respectively represent the maximum value and the minimum value in the sample data.
[0009] According to the application, in step (2), specifically, the photovoltaic output of the transformer area under clear sky, cloudy sky and rainy sky is collectively taken as the output feature, the output data time axis under the three kinds of weather is strictly aligned, the DTW (Dynamic Time Warping) distance is taken as the measurement index of the similarity of the transformer area distributed photovoltaic output, the AHC clustering algorithm is adopted, the transformer area photovoltaic is classified based on the power grid topology, and the logical consistency of the clustering result and the physical connection is ensured. Suppose the photovoltaic power sequences of two transformer area photovoltaics are respectively: (4) (5) In the formula: represents the power sequence ; represents the power value at time point in the power sequence ; The normalized path can reveal the similar time points in the power sequence, so as to realize the flexible alignment of the power sequence, and the normalized path of and is defined as : (6) In the formula: is in the form of , the path starts from and ends at , and finally the normalized path to be obtained is the normalized path with the shortest distance and monotonically increasing; The similarity between the photovoltaic power sequences and is: (7) (8) Wherein, represents the initial similarity between time points and in the given two photovoltaic power sequences and ; represents the updated two photovoltaic power sequences and At the moment and The similarity between them is recursively updated by minimizing the previous similarity metric to calculate the sequence similarity of the best match, and finally the cumulative distance of the dynamic curved path that meets the constraints is calculated through a recursive algorithm. The specific steps based on AHC and DTW algorithms are as follows: 1) Treat each photovoltaic area as a separate cluster; 2) Calculate the similarity between any two stations according to the DTW algorithm, represented by the matrix M, where , 、 Representing different regions respectively; 3) Find the two clusters with the greatest similarity and merge them into a new cluster ; 4) According to the newly formed cluster, move the number of the subsequent cluster forward one and delete the matrix No. Line and Column, update matrix , , update the current number of clusters ; 5) Repeat steps 3) and 4) until the distance between different clusters reaches the preset threshold; 6) Based on the physical connection relationship between substations in the power grid topology, ensure that the clustering results of each substation are logically consistent with its physical connection in the power grid, and finally output the clustering results; After clustering, the Silhouette coefficient (SC) is introduced as the evaluation function of the clustering results until the final optimal clustering result is obtained. The calculation formula of the Silhouette coefficient is as follows: (9) in, Taiwan area Silhouette coefficient; For Taiwan sample The average distance to other samples in the same cluster is called cohesion; for The average distance from all samples in other clusters is called separation; the average silhouette coefficient is the average of the silhouette coefficients of all samples, and its value range is [-1, 1]. The larger its value, the smaller the intra-cluster distance, the larger the inter-cluster distance, and the better the clustering effect.
[0010] According to the present invention, preferably, in step (3), The transformer encoder is stacked by multiple encoder layers, each of which includes an attention sublayer and a feed-forward neural network sublayer, the attention sublayer is composed of multi-head self-attention mechanism, residual connection and layer normalization; the feed-forward neural network sublayer contains feed-forward neural network, residual connection and layer normalization, wherein the multi-head self-attention mechanism is the core part of the transformer architecture, which captures the global information in the sequence by modeling the dependency between each position and other positions in the input sequence, the attention score calculation adopts the key-value-query mode, which makes the key only focus on the first n important queries, that is: (10) In the formula: Attention(·) is an attention score calculation function, Q 、 K 、 V Query, key and value matrices respectively; is the dimension of K ; T is the matrix transpose operation, the attention score is calculated by assigning weights to each position in the input sequence, and Softmax is a normalization function; The GraphSage network defines two key functions, AGGREGATE(·) and CONCAT(·), AGGREGATE(·) is used to aggregate information from node neighbors, and CONCAT(·) is used to combine the current node features with the aggregated node features, abstract the district photovoltaic as a node in the graph, construct edge connection according to meteorological similarity, and form a spatial correlation graph, let the graph , v represents one of the nodes, for any node neighbor , the node embedding of the Kth layer is , then the update process is: (11) (12) AGGREGATE(·) represents the aggregation operation, and the average aggregation method is selected here, and CONCAT(·) represents the connection operation; σ represents an activation function; and respectively represent the aggregated node features of the adjacent neighborhood of node v and the node features of the adjacent nodes of node v, wherein ; A time-space correlation feature synchronous extraction model framework for district distributed photovoltaic is constructed, and synchronous prediction of the output of each district in the sub-region is realized, the specific steps are as follows: (31) Based on the historical The time series characteristics of the photovoltaic output are extracted by using a Transformer encoder module at the power data of the time step. The Transformer encoder effectively captures complex time series characteristics in the photovoltaic output sequence, including diurnal periodic fluctuations, seasonal trend changes and other modes, through a multi-layer stacking structure combined with position encoding and self-attention mechanisms. (32) Based on the prediction of each area The spatial correlation characteristics of each area under the influence of the weather are dynamically learned by using a GraphSage network to extract the spatial correlation characteristics of the weather of each area in the sub-region, abstracting each area in the sub-region as a node in the graph structure, taking the similarity of the meteorological elements as the node connection weight, and performing neighborhood sampling and information aggregation operations. Through multiple iterations and optimization, the GraphSage network outputs the spatial feature vector of each area. (33) The time series characteristics extracted by the Transformer encoder and the spatial characteristics generated by the GraphSage network are spliced into a joint feature vector containing spatio-temporal coupling information. (34) Based on the extracted spatio-temporal aggregation features, a Bilstm module is used to fit the mapping relationship between the aggregation features and the power, and simultaneously output the photovoltaic power of all areas in the sub-region. Considering the complex causal relationship contained in the spatio-temporal aggregation features and the bidirectional dependence of information in the time series transmission, the BiLSTM can simultaneously capture the forward and backward dependent information of the data through the parallel architecture of the forward and backward LSTM units, effectively solving the limitations of the one-way recurrent neural network in information utilization. The model is trained in an end-to-end manner, with the actual power values of multiple areas as the supervision signal, and the network parameters are iteratively optimized until the prediction error converges to a preset threshold, finally realizing high-precision collaborative prediction of the photovoltaic output of multiple areas.
[0011] According to the application, in step (4), specifically, The spatio-temporal feature extraction module is jointly trained on the source area data S, and the optimization target is to minimize the prediction loss function : (13) (14) Wherein, is the prediction function of the complete model, is the input composed of historical power and weather data, is the real power output; after the BiLSTM module fits the spatio-temporal aggregation features, it is mapped to the prediction space through a fully connected layer , represents the predicted value of all samples on the source area training data S The mean squared error (MSE) is the average of the squared differences between the predicted and actual values. The core parameters of the frozen Transformer encoder and the GraphSAGE model are θ T ∗ , θ G ∗ , which retains the general feature extraction capability: (15) (16) A lightweight backbone network is constructed , only the BiLSTM and fully connected layer parameters are retained: (17) On the target sub-region data, only the BiLSTM parameters and the fully connected layer parameters are locally fine-tuned, and the optimization target is to minimize the sub-region specific loss function: (18) Through iterative updates and , the model adapts to the differences in power grid topology and weather in different sub-regions, while maintaining a lightweight architecture, achieving fast modeling across regions, which corresponds to the average loss on the th sub-region data , ensuring that the model focuses on the sample characteristics of the region when adapting to new regions.
[0012] Further, during the fine-tuning process, if a large batch is used for training, the model may quickly adjust its weights and overfit the data of the target sub-region, causing the model to forget the knowledge learned from the source sub-region. Similarly, the learning rate controls the step size of each update. If a too high learning rate is used, it may lead to over-adjustment and loss of knowledge from the source sub-region. Therefore, this embodiment uses small batch data and stepwise learning rate to fine-tune the model parameters. The learning rate is initially set appropriately and gradually reduced at the end of each cycle. The learning rate decay step and learning rate decay ratio are gradually fine-tuned according to the prediction performance of the model in different sub-regions, helping the model to gradually converge and avoiding oscillation in the training process.
[0013] Further, in order to evaluate the accuracy of the model prediction, the Root Mean Square Error (RMSE) and the R-squared (R 2 ) are used to evaluate the model's photovoltaic power prediction performance. The specific calculation formula is as follows: (19) (20) In the formula: , Respectively represent the RMSE, R 2 Value between predicted power and real power, Indicates the number of samples, Indicates the sample sequence number, , Respectively represent the power real value and predicted value of the Sample, Indicates the average value of the real value.
[0014] A transformer area distributed photovoltaic power ultra-short-term prediction system, comprising: A data processing module is used for collecting historical measured power data of all transformer areas in a to-be-predicted area and meteorological data corresponding to the position of each transformer area, and performing data processing; A sub-area division module is used for calculating a dynamic time warping distance by using the historical measured power data of the transformer area, as a similarity measurement index for evaluating the output of the transformer area, and adopting a condensed hierarchical clustering algorithm to divide the transformer area photovoltaic in the area into a sub-area with consistent output; A prediction module is used for constructing a Transformer and GraphSAGE collaborative architecture, realizing time-space feature decoupling and tensor fusion, capturing time-varying features by using Transformer self-attention, modeling spatial ultra-short-term meteorological correlation by using GraphSAGE neighborhood aggregation, fitting the composite time-space features by using a BiLSTM model, and synchronously outputting the photovoltaic power prediction results of each transformer area in the sub-area, thereby avoiding repeated modeling of each transformer area; A total power calculation module is used for constructing a suitable transfer modeling strategy based on the shared features between sub-areas, freezing the time-space feature extraction backbone network parameters, and only performing lightweight adaptive operation on the region-specific connection layer, thereby realizing fast and low-cost modeling of each sub-area, and finally obtaining the total power of the region by accumulating the power of each sub-area.
[0015] The beneficial effects of the present application are: 1、The present application uses a dynamic time warping distance method to measure the similarity of the output of the transformer area, which breaks through the limitation of nonlinear representation of Euclidean distance, solves the time sequence shift problem, and accurately measures the output characteristics of the transformer area photovoltaic.
[0016] 2. The present invention constructs a feature extraction model that collaborates with the Transformer encoder and the GraphSAGE graph neural network. Through the decoupling of spatiotemporal features and the tensor fusion mechanism, combined with the BiLSTM network, it realizes the joint fitting of the power of the entire sub-region, breaking through the bottleneck of the traditional single model in characterizing spatiotemporal coupling features. At the same time, it ensures the accuracy and consistency of the prediction results, and effectively improves the overall performance of regional photovoltaic power prediction.
[0017] 3. This invention employs a lightweight transfer learning strategy based on sub-region feature sharing. By freezing backbone network parameters and adaptively fine-tuning region-specific connection layers, it enables rapid cross-sub-region model building. Ultimately, the regional total power is obtained by summing the power of each sub-region. This overcomes the regional limitations of a single model, reduces training overhead, and enables rapid collaborative prediction of PV power across a large number of sub-regions. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 This is a flow chart of the ultra-short-term prediction method for distributed photovoltaic power in a substation area proposed in an embodiment of the present invention.
[0019] Figure 2 This is a comparison chart of the output of adjacent substations on sunny and cloudy days analyzed in the embodiment of the present invention, where: Figure 2 The middle (a) is the sunny day output diagram of adjacent substations. Figure 2 Middle (b) is the output diagram of adjacent substations on a sunny day.
[0020] Figure 3 This is a flowchart of photovoltaic area division based on the agglomerative hierarchical clustering algorithm according to an embodiment of the present invention.
[0021] Figure 4 This is a diagram showing the Transformer embedding method of an embodiment of the present invention.
[0022] Figure 5 This is a structural diagram of the spatiotemporal feature decoupling module proposed in an embodiment of the present invention. DETAILED DESCRIPTION
[0023] The present invention will be further described below with reference to embodiments and accompanying drawings, but is not limited thereto.
[0024] Example 1: like Figure 1 As shown, this embodiment provides a method for ultra-short-term prediction of distributed photovoltaic power in a substation, and the steps are as follows: (1) Collect historical measured power data of all substations in the forecast area and extract meteorological data of the corresponding location of each substation through measured reanalysis data, and perform data processing; In the operation process of the photovoltaic power generation system, due to the influence of factors such as fault of the collecting device and human operation error, there may be inconsistent deviation between the actual power output and the collected data, such deviation will interfere with the learning and training process of the prediction model, from the perspective of data characteristics, the common interference data mainly shows missing values and abnormal values in the time series. If the unprocessed interference data is directly applied to the prediction model construction, it may cause the problem of non-convergence of the model iteration process, or cause the significant decline of the prediction accuracy, therefore, after the data collection is completed, the data must be cleaned.
[0025] Specifically, the data processing includes missing value processing, abnormal value processing and normalization processing. In the photovoltaic power time series data, the numerical evolution not only follows the dynamic change rule of the time dimension, but also presents complex space-time coupling characteristics due to the nonlinear fluctuation of meteorological conditions. Influenced by the instantaneousness of meteorological elements such as solar irradiance and environmental temperature, the photovoltaic power data shows significant autocorrelation in the time series, that is, the power value at a certain sampling time has a close space-time dependence relationship with the power and irradiance at adjacent time. The KNN algorithm is based on the similarity measurement principle between data samples, and can effectively capture the local similarity of photovoltaic data in the time series by constructing a distance measurement model in a multi-dimensional feature space. The method takes the K nearest neighbor data of the target sample as a reference, and fills the missing values through weighted or non-weighted aggregation strategy, so as to retain the space-time evolution characteristics and original distribution characteristics of the data to the greatest extent. Based on this, the present application selects KNN algorithm as the missing value repair method to ensure the integrity of the data set and the reliability of the analysis result, which is mathematically expressed as follows:
[0026] (1) (2) In the formula: and are two sample points, each has n characteristic values, and are the values of points and on the first characteristic, is the value of the K nearest neighbor of the missing value, is the Euclidean distance between the two samples, used to select the K nearest neighbor of the missing data point, and the missing value is filled with the average value of the K nearest neighbors, and the too small K value is easy to be interfered by the instantaneous abnormal value, and the too large KThe value can blur the short-term change characteristics of the data. Through multiple cross-validation experiments, the embodiment selects K = 4.
[0027] The outlier processing adopts the box plot method (IQR) to detect outliers. The core idea is to identify extreme values deviating from the main distribution range by describing the distribution form and dispersion degree of the data. The key steps of the outlier detection method based on IQR are to set an outlier judgment standard. Generally, 1.5 times IQR is used to judge whether the data is an outlier. Any data point less than the lower limit or greater than the upper limit (Upper Bound) is regarded as an outlier. In addition, the embodiment also judges the case that the photovoltaic power is negative as an outlier. For the detected outliers, the embodiment uses the missing value interpolation method to replace these outliers by calculating new data. In this way, the integrity of the data can be maintained, and the accuracy of the subsequent analysis and modeling process can be ensured.
[0028] The normalization processing adopts the maximum and minimum value standard method to standardize the sample data. The sample data is linearly mapped to [0, 1]; (3) In the formula: X represents the standardized sample data; X represents the sample data to be standardized; and respectively represent the maximum value and the minimum value in the sample data of this type.
[0029] (2) Calculate the dynamic time warping distance using the historical measured power data of the transformer area as an evaluation index of the similarity between the transformer area outputs. The agglomerative hierarchical clustering algorithm is used to divide the transformer area photovoltaics in the region into sub-regions with consistent outputs; The strong correlation between meteorological factors and photovoltaic power generation has been widely studied and confirmed. Among them, the short-time scale fluctuations of key environmental parameters such as irradiance and temperature have a significant synchronous influence effect on the photovoltaic output fluctuation characteristics of adjacent transformer areas. As shown in Figure 2 , by analyzing the photovoltaic power generation data of two adjacent transformer areas under typical sunny (left) and cloudy (right) conditions, it can be found that within the same natural day, the power curves of the two adjacent transformer areas show a high degree of morphological consistency. Although affected by factors such as cloud movement, there is a time series shift phenomenon in local time periods, but overall, it still shows significant spatio-temporal correlation. Based on the above findings, the transformer area photovoltaics with coordinated output trend changes can be spatially aggregated to further explore the spatio-temporal coupling rules between adjacent transformer areas from the regional scale. This regional aggregation strategy can not only weaken the random fluctuation influence of a single transformer area with the help of group effect, but also enhance the prediction reliability relying on spatial correlation.
[0030] In addition, in the operation environment of the power distribution network station area, the power generation characteristics of the distributed photovoltaic users are heterogeneous, causing the photovoltaic output curve of each user to present complex fluctuation patterns and scale expansion characteristics in the time dimension. Such non-uniform output characteristics pose a severe challenge to traditional power prediction methods based on a unified model. To effectively solve the above problems, the embodiment adopts a hybrid method combining the agglomerative hierarchical clustering algorithm (AHC) and dynamic time warping to construct a refined station area photovoltaic output feature classification system, clusters the station area photovoltaics with similar output characteristics into several feature groups, so that the photovoltaic output data of the users in the same cluster exhibit high homogeneity in fluctuation trends, peak distribution, and other dimensions.
[0031] Specifically, the station area photovoltaic output under clear sky, overcast sky, and rainy day typical weather is collectively taken as the output feature, the output data time axis under the three kinds of weather is strictly aligned, the DTW (Dynamic Time Warping) distance is taken as the measurement index of the similarity of the station area distributed photovoltaic output, the AHC clustering algorithm is adopted, and the station area photovoltaics are classified based on the grid topology, ensuring that the clustering results have logical consistency with the physical connection; DTW is a widely used nonlinear distance measurement method in time series analysis, which can effectively deal with the local misalignment of the output curve on the time axis caused by weather disturbance, user behavior, or other external conditions. Unlike the traditional Euclidean distance, DTW performs dynamic matching on the time axis, so that two similar but asynchronous sequences can also be identified as similar, thus more truly reflecting the output characteristics of the household photovoltaic system. Based on this distance definition, a similarity matrix between users can be constructed. Then the agglomerative hierarchical clustering algorithm is introduced, which gradually merges the most similar station area photovoltaic sets from bottom to top, and constructs a hierarchical user clustering tree diagram. This method does not need to pre-set the number of categories, and can naturally form clustering levels according to the DTW distance, adapting to the complexity and diversity of the output mode in the spatial and temporal dimensions. This clustering algorithm regards each sample as a cluster, then starts to merge clusters with high similarity according to certain rules, and finally all samples form a cluster or reach a certain condition, and the algorithm ends. In practical applications, by setting a distance threshold or specifying a clustering level, the required user classification results can be obtained flexibly.
[0032] Suppose the photovoltaic power sequences of two station area photovoltaics are: (4) (5) In the formula: represents the power sequence the power value at time t; representing the power sequence in the power value at time t, the warping path can reveal the similar time points in the power sequence, so as to realize the flexible alignment of the power sequence, define and the warping path of : (6) In the formula: the form of , the path starts from to end, and finally the obtained warping path is the shortest one and monotonically increasing; photovoltaic power sequence and between the similarity is: (7) (8) wherein, represent the initial similarity between time and in the given two photovoltaic power sequences and ; represent the similarity between time and of the updated two photovoltaic power sequences and , recursively updated by minimizing the previous similarity measure value, to calculate the best matching sequence similarity, and finally calculate the cumulative distance of the dynamic bending path meeting the constraint condition through the recursive algorithm, the specific steps based on AHC and DTW algorithm, as shown in Figure 3 , the specific steps are as follows: 1) take each district photovoltaic as a separate cluster; 2) calculate the similarity between any two districts according to the DTW algorithm, represented by matrix M, wherein , , represent different districts respectively; 3) find the two clusters with the largest similarity, and merge them into a new cluster ; 4) according to the newly formed cluster, move the number of the following cluster one position forward, delete the first row and column of matrix , and update the matrix , , update the current number of clusters ; 5) Repeat step 3) and step 4) until the preset distance threshold is reached between different clusters; 6) Based on the physical connection relationship between the transformer areas in the power grid topology, ensure that the clustering result of each transformer area is logically consistent with its physical connection in the power grid, and finally output the clustering result; After clustering, the silhouette coefficient (SC) is introduced as the evaluation function of the clustering result, until the final best clustering result is obtained, and the silhouette coefficient calculation formula is as follows: (9) Wherein, the silhouette coefficient of the transformer area ; is the average distance of the transformer area sample and other samples in the same cluster, called cohesion degree; is the average distance of all samples in other clusters, called separation degree; the average silhouette coefficient is the average value of the silhouette coefficients of all samples, the value range is [-1, 1], the larger the value is, the smaller the distance within the cluster is, the larger the distance between clusters is, and the better the clustering effect is.
[0033] (3) Construct a collaborative architecture of Transformer and GraphSAGE to realize spatiotemporal feature decoupling and tensor fusion, use Transformer self-attention to capture time-varying features, model spatial ultra-short-term meteorological correlation through GraphSAGE neighborhood aggregation, use BiLSTM model to fit the composite spatiotemporal features, and simultaneously output the photovoltaic power prediction results of each transformer area in the sub-region, avoiding repeated modeling for each transformer area; The photovoltaic output fluctuation characteristics mainly have two parts: one part is the internal time sequence output characteristics affected by the earth's rotation, showing obvious daily periodicity, and the other part is the external fluctuation characteristics affected by weather changes such as cloud clusters, showing great uncertainty, both of which have strong direct influence on photovoltaic output. Due to the geographical distribution characteristics of transformer areas, there is rich spatial dependence information between them, and through in-depth mining of the two characteristics, it is helpful to help the model better understand the spatiotemporal interaction relationship between meteorology and meteorology, meteorology and power in each transformer area. Therefore, this embodiment constructs a framework for synchronous extraction of spatiotemporal correlation features. This framework can not only capture the complex interaction relationship between spatial and temporal features, but also efficiently process large-scale data of multiple transformer areas, effectively reducing the calculation time.
[0034] Although the traditional LSTM, GRU and other recurrent neural network models have time series modeling capabilities, their serial computing structure limits parallel efficiency and makes it difficult to effectively capture the cross-correlation characteristics of the time dimension between the transformer models. The Transformer model takes the self-attention mechanism as the core, which can model the global dependence between any time steps in the time series. Its full-parallel computing architecture supports batch processing of multiple transformer time series data as a unified input matrix, modeling the time series characteristics of different transformer under the same architecture, and explicitly constructing the time interaction path across the transformer. The feature embedding strategy adopted by the Transformer model integrates the information of multiple variables at the same timestamp into a single time marker, as shown in Figure 4 The encoder of the Transformer model is responsible for capturing the global dependence between input variables, while the decoder is used to map the extracted deep time series features to the prediction output. Considering the redundancy of the traditional Transformer encoding and decoding architecture in the single time series feature extraction task, a lightweight Transformer encoder is used as the core modeling framework.
[0035] The Transformer encoder is stacked by multiple encoder layers, each of which includes an attention sublayer and a feedforward neural network sublayer. The attention sublayer is composed of multi-head self-attention mechanism, residual connection and layer normalization; the feedforward neural network sublayer contains feedforward neural network, residual connection and layer normalization. The multi-head self-attention mechanism is the core part of the Transformer architecture, which captures the global information in the sequence by modeling the dependence between each position and other positions of the input sequence. The attention score calculation adopts the key-value-query mode, which allows the key to focus only on the first n important queries, i.e.: (10) where Attention(·) is the attention score calculation function, Q , K , V are the query, key and value matrices, respectively; is the dimension of K ; T is the matrix transpose operation, which calculates the attention score by assigning weights to each position in the input sequence, and Softmax is the normalization function. In recent years, with the gradual rise of Graph Neural Network (GNN) and its derivative models, researchers have obtained powerful tools in exploring the spatial correlation of time series problems. GraphSAGE is a graph neural network model based on inductive learning. By sampling and aggregating the features of adjacent nodes, it can efficiently learn the spatial embedding representation of nodes without relying on the full graph structure, capturing the disturbance propagation effect of spatial neighbors. When applied to the modeling of district photovoltaic power prediction, it can enhance the model's ability to perceive regional collaborative change patterns. In addition, in multi-region modeling tasks, GraphSAGE can adaptively capture the topological structure differences of PSDTAs in different sub-regions, without the need to reconstruct the graph structure and retrain in each sub-region, achieving shared modeling with good transferability.
[0036] GraphSage network defines two key functions, AGGREGATE(·) and CONCAT(·), AGGREGATE(·) is used to aggregate information from node neighbors, and CONCAT(·) is used to combine the current node features with the aggregated node features. Abstract district photovoltaic power as nodes in the graph, construct edge connection according to meteorological similarity, and form a spatial correlation graph. Let G = (V, E) be a spatial correlation graph, v represents one of the nodes, and for any node neighbor , the K-th layer node embedding representation is , then the update process is: (11) (12) where AGGREGATE(·) represents the aggregation operation, and the average aggregation method is selected here, and CONCAT(·) represents the connection operation; σ represents the activation function; and respectively represent the aggregated node features of the adjacent neighborhood of node v and the node features of the adjacent nodes of node v, where ; There is rich spatio-temporal correlation information between the photovoltaic output of the transformer area, and in-depth mining and understanding of it helps to improve the accuracy of photovoltaic power prediction. The temporal dependence of photovoltaic system power output is usually reflected in short-term historical data, which contains the working state of the photovoltaic system in the past, can reflect the influence of device performance and climate change on power generation, and help to capture the recent power generation trend and fluctuation pattern, thereby providing good time sequence features for future prediction. However, historical photovoltaic output data can indeed reflect the photovoltaic output level between transformer areas to some extent, and the output of the photovoltaic system is not only affected by historical factors, but also driven by many spatial factors such as meteorological conditions, geographical location and environmental changes. The differences in these factors make the spatial correlation between different transformer areas show great differences. In addition, historical data cannot fully consider the spatial influence of future meteorological changes on photovoltaic power output, which is particularly important for ultra-short-term prediction. Therefore, relying on historical output data to mine the spatial correlation between photovoltaic systems has certain disadvantages. In contrast, using future predicted meteorological data to mine the spatial correlation between transformer area photovoltaics has more obvious advantages. Through future meteorological data, the model can capture these spatial differences in advance, and then more accurately predict the photovoltaic output of different transformer areas. The spatial characteristics revealed by historical output data are often static, while the dynamic characteristics of predicted meteorology can adapt to changes in spatial characteristics between transformer areas, and can enhance the sensitivity of the model to spatial differences between transformer areas.
[0037] Therefore, by constructing a model that combines the output data of each transformer area at historical time and the meteorological data at future time, the model can consider the influence of temporal variation and future environmental changes, and capture more rich spatio-temporal features. Therefore, the embodiment constructs a spatio-temporal correlation feature synchronous extraction model framework for distributed photovoltaics in transformer areas, and realizes synchronous prediction of the output of each transformer area in the sub-area, the specific steps are as follows:
[0038] (31) Based on the historical power data of each transformer area at each time step, the Transformer encoder module is used to extract the time sequence features of photovoltaic output. The Transformer encoder effectively captures complex temporal features in the photovoltaic output sequence through a multi-layer stacking structure combined with position encoding and self-attention mechanism, including diurnal periodic fluctuations, seasonal trend changes and other patterns; (32) Based on the predicted The meteorological data of each time step is used to extract the spatial correlation characteristics of the meteorology of each substation by using the GraphSage network. Each substation in the subregion is abstracted as a node in the graph structure, and the similarity of meteorological elements is used as the node connection weight. Through neighborhood sampling and information aggregation operations, the spatial correlation characteristics of each substation under the influence of meteorology are dynamically learned. Through multiple iterations of optimization, the GraphSage network outputs the spatial feature vector of each substation. (33) The time sequence features extracted by the Transformer encoder are spliced with the spatial features generated by the GraphSage network to form a joint feature vector containing spatio-temporal coupling information, as shown in the structure Figure 5 The spatio-temporal aggregation features not only integrate the time evolution law of photovoltaic output, but also incorporate the spatial correlation characteristics under the influence of meteorological factors, providing rich feature expression for subsequent prediction. (34) Based on the extracted spatio-temporal aggregation features, a Bilstm module is used to fit the mapping relationship between the aggregation features and the power, and the photovoltaic power of all substations in the subregion is output simultaneously. Considering the complex causal relationship contained in the spatio-temporal aggregation features and the bidirectional dependency of information in time series transmission, BiLSTM can capture forward and backward dependency information simultaneously through the parallel architecture of forward and backward LSTM units, effectively solving the limitations of one-way recurrent neural networks in information utilization. The model uses the actual power values of multiple substations as the supervision signal to iteratively optimize the network parameters until the prediction error converges to the preset threshold, finally realizing high-precision collaborative prediction of multi-substation photovoltaic output.
[0039] (4) Based on the shared features between subregions, a suitable transfer modeling strategy is constructed. The parameters of the spatio-temporal feature extraction backbone network are frozen, and only lightweight adaptation operations are performed on the region-specific connection layer to realize fast and low-cost modeling of each subregion. Finally, the total power of the region is obtained by accumulating the power of each subregion.
[0040] Specifically, Due to the joint action of time periodicity, climate factors, solar radiation and other factors, the power generation time sequence characteristics between the stations often have consistent regularity and periodicity. For example: photovoltaic power generation is mainly concentrated in the daytime, and the output increases first and then decreases with time, generally reaching the maximum at noon, the strongest at noon, and gradually weakening in the afternoon. Or the power generation at night or on cloudy days is close to zero. The difference in spatial characteristics is a key factor affecting photovoltaic output. In addition, the change of the solar angle will directly affect the receiving light of the photovoltaic panel, but its change also has a certain regularity. Although the specific values of the solar angle are different in different stations, the trend and mode of its change are consistent in different regions. This consistency provides a theoretical basis for extracting the time sequence characteristics by fixing the time sequence model parameters between different stations. Although the stations are relatively close, due to local meteorological differences and local cloud distribution differences, the spatial dependence relationship between the station outputs changes dynamically, which will directly affect the time sequence characteristics of photovoltaic power generation. Therefore, it is necessary to model the spatial characteristics of meteorological fluctuations between stations in a small area in the photovoltaic power prediction model.
[0041] From step S1, the clustered photovoltaic stations are divided into different sub-regions, and the time sequence characteristics of each sub-region have strong consistency, but there is weak spatial correlation due to the geographical distance. If an independent model is trained for each sub-region, it will lead to too long training time and waste of computing resources. Therefore, it is possible to consider sharing the consistent characteristics between stations and only locally adjusting the individual differences between sub-regions, thereby greatly reducing the training time and the demand for computing resources, and enhancing the generalization ability of the model. In view of the above analysis, the embodiment constructs a cross-regional lightweight migration modeling strategy, which provides theoretical support and technical path for efficient collaborative modeling of regional-level distributed photovoltaic systems. Specifically, in the time dimension, based on the time sequence feature extraction ability of the Transformer encoder learned in the typical sub-region, the parameter reuse strategy is directly applied to other stations to efficiently capture the common time-varying regularity, and the GraphSAGE algorithm can dynamically adapt to node representation learning under different topological structures, by weakening the influence of power grid topology, the unified model architecture is used to efficiently extract the spatial characteristics of different regions.
[0042] The core idea of transfer learning is to transfer existing knowledge to new tasks, thereby reducing the sample requirements and training time for the new tasks while improving model performance. This provides a way to migrate the parameters of time series feature extraction models between different sub-regions. To meet the requirements of efficient modeling of multi-sub-region photovoltaic systems, a lightweight parameter transfer optimization strategy is proposed. By freezing the core parameter group responsible for decoupling spatiotemporal features in the Transformer encoder and GraphSAGE model, a lightweight universal feature extraction backbone network is constructed to maximize the reuse of the spatiotemporal feature expression capabilities learned during training in the source region. Furthermore, only the BILSTM modules in the model that are specific to the sub-region are locally adaptively fine-tuned to adapt to the differences in grid topology and changing meteorological conditions across different sub-regions.
[0043] Jointly train the spatiotemporal feature extraction module on the source region data S, and the optimization goal is to minimize the prediction loss function : (13) (14) in, is the complete model prediction function, The input consists of historical power and meteorological data, is the real power output; after the BiLSTM module fits the spatiotemporal aggregation features, it passes through the fully connected layer Mapped to the prediction space, Represents the predicted value of all samples on the source region training data S The square of the difference from the true value y is averaged, which is the expected form of the mean square error (MSE); Freeze the core parameters θ of the Transformer encoder and GraphSAGE models T ∗ ,θ G ∗ , retaining general feature extraction capabilities: (15) (16) Building a lightweight backbone network , only retaining the fine-tunable parameters of BiLSTM and fully connected layers: (17) On the target sub-region data, only the BiLSTM parameters and fully connected layer parameters Perform local fine-tuning, and the optimization goal is to minimize the sub-region specific loss function: (18) by iterative updating and , the model adapts to the differences in power grid topology and weather in different sub-regions, while maintaining a lightweight architecture, achieving fast modeling across regions, corresponding to the average loss on the data of the first sub-region, ensuring that the model focuses on the sample characteristics of the new region when adapting.
[0044] During fine-tuning, if a large batch is used for training, the model may quickly adjust its weights, overfitting the data of the target sub-region, causing the model to forget the knowledge learned from the source sub-region. Similarly, the learning rate controls the step size of each update, and if a too high learning rate is used, it may lead to over-adjustment and loss of knowledge from the source sub-region. Therefore, this embodiment uses small batch data and stepwise learning rate to fine-tune the model parameters, initially setting a suitable learning rate, which is gradually reduced at the end of each cycle. The learning rate decay step and learning rate decay ratio are gradually fine-tuned according to the prediction performance of the model in different sub-regions, helping the model to gradually converge and avoiding oscillation in the training process.
[0045] The migration object of this embodiment is the prediction model between sub-regions within the region, and the effectiveness of migration depends mainly on the similarity between the source domain and the target domain. Since the solar elevation angle of the sub-regions at the same latitude changes similarly, leading to consistent overall trends in their time series characteristics. Therefore, this embodiment uses the idea of transfer learning to improve the generalization ability of the model, without making a detailed division between the source domain and the target domain. Specifically, this embodiment only selects the sub-region with the best clustering effect as the source domain for modeling, maximizing the similarity between the source domain and the target domain, and ensuring that transfer learning can be effectively applied. After obtaining the prediction values of each sub-region substation, the power prediction values of each substation at future time points are accumulated and summed, and then the total regional power prediction value is obtained.
[0046] To evaluate the accuracy of the model's prediction, the Root Mean Square Error (RMSE) and the R-squared (R 2 ) are used to evaluate the model's photovoltaic power prediction performance. The specific calculation formulas are as follows: (19) (20) In the formula: , respectively represent the RMSE and R 2 value between the predicted power and the true power, represents the number of samples, denotes the sample sequence number, , denotes the power true value and the predicted value of the i-th sample, respectively, denotes the power true value and the predicted value of the i-th sample, respectively, denotes the average value of the true value.
[0047] Embodiment 2: The embodiment provides a distributed photovoltaic power ultra-short-term prediction system for a transformer area, comprising: A data processing module is configured to collect historical measured power data of all transformer areas in a to-be-predicted region and meteorological data corresponding to the positions of each transformer area, and perform data processing. A sub-region division module is configured to calculate a dynamic time warping distance by using the historical measured power data of the transformer areas, as a similarity measurement index for evaluating the output of the transformer areas, and divide the transformer areas in the region into sub-regions with consistent output by using a condensed hierarchical clustering algorithm. A prediction module is configured to construct a Transformer and GraphSAGE collaborative architecture, realize time-space feature decoupling and tensor fusion, capture time-varying features by using a Transformer self-attention, model spatial ultra-short-term meteorological correlation by using a GraphSAGE neighborhood aggregation, fit a composite time-space feature by using a BiLSTM model, and simultaneously output photovoltaic power prediction results of each transformer area in the sub-region, thereby avoiding repeated modeling for each transformer area. A total power calculation module is configured to construct a suitable transfer modeling strategy based on shared features between sub-regions, freeze time-space feature extraction backbone network parameters, and only perform lightweight adaptive operations on region-specific connection layers, so as to realize rapid and low-cost modeling of each sub-region, and finally obtain a total power of the region by accumulating the power of each sub-region.
Claims
1. A method for ultra-short-term prediction of distributed photovoltaic power in a substation, characterized by: Here are the steps: (1) Collect historical measured power data of all substations in the forecast area and meteorological data of the corresponding location of each substation, and process the data; (2) The dynamic time warping distance is calculated using the historical measured power data of the substation, which is used as a similarity metric to evaluate the output of the substations. The agglomerative hierarchical clustering algorithm is used to divide the photovoltaic substations in the region into sub-regions with consistent output. (3) Construct a Transformer and GraphSAGE collaborative architecture to achieve spatiotemporal feature decoupling and tensor fusion. Use Transformer self-attention to capture time-varying features, model spatial ultra-short-term meteorological correlations through GraphSAGE neighborhood aggregation, use the BiLSTM model to fit composite spatiotemporal features, and synchronously output the photovoltaic power prediction results of each substation in the subregion to avoid repeated modeling for each substation. (4) Based on the shared features among sub-regions, a suitable migration modeling strategy is constructed, the parameters of the spatiotemporal feature extraction backbone network are frozen, and only lightweight adaptation operations are performed on the region-specific connection layer to achieve fast and low-cost modeling of each sub-region. Finally, the total regional power is obtained by accumulating the power of each sub-region.
2. The method for ultra-short-term prediction of distributed photovoltaic power in a metropolitan area according to claim 1, characterized in that: In step (1), specifically, data processing includes missing value processing, outlier processing and normalization processing.
3. The method for ultra-short-term prediction of distributed photovoltaic power in a substation area according to claim 2, wherein: In step (2), specifically, the photovoltaic output of the substation under clear sky, cloudy day and rainy day is taken as the output feature, the output data under the three weather conditions are time-aligned, the DTW distance is used as the measurement index of the similarity of the distributed photovoltaic output of the substation, and the AHC clustering algorithm is used to classify the photovoltaic in the substation based on the grid topology to ensure that the clustering results are logically consistent with the physical connection.
4. The method for ultra-short-term prediction of distributed photovoltaic power in a substation area according to claim 1, wherein: In step (3), The Transformer encoder is composed of multiple stacked encoder layers. Each encoder layer consists of an attention sublayer and a feedforward neural network sublayer. The attention sublayer consists of a multi-head self-attention mechanism, residual connections, and layer normalization. The feedforward neural network sublayer contains a feedforward neural network, residual connections, and layer normalization. The multi-head self-attention mechanism captures global information in the sequence by modeling the dependency of each position in the input sequence with other positions. The attention score calculation adopts a key-value-query mode, so that the key only focuses on the top n important queries, that is: (10) Where: Attention(·) is the attention score calculation function, Q 、 K 、 V are query, key, and value matrices respectively; yes K Dimensions; T is the matrix transpose operation, the attention score is calculated to assign weights to each position in the input sequence, and Softmax is the normalization function.
5. The method for ultra-short-term prediction of distributed photovoltaic power in a substation area according to claim 4, characterized in that: GraphSage network defines two key functions, namely AGGREGATE(·) and CONCAT(·). AGGREGATE(·) is used to aggregate information from node neighbors, and CONCAT(·) is used to combine the current node features with the aggregated node features. The photovoltaic area is abstracted as a node in the graph, and edge connections are constructed based on meteorological similarity to form a spatial association graph. Suppose , v represents one of the nodes, for any node neighbor , the node embedding of the K-th layer is expressed as , the update process is: (11) (12) Among them, AGGREGATE(·) represents the aggregation operation, and the average aggregation method is selected here. CONCAT(·) represents the connection operation; σ represents the activation function; and They represent the aggregate node features of the adjacent neighborhood of node v and the node features of the adjacent nodes of node v, respectively, where .
6. The method for ultra-short-term prediction of distributed photovoltaic power in a substation area according to claim 5, characterized in that: In step (3), the specific steps are as follows: (31) Based on the history of each district The Transformer encoder module is used to extract the temporal features of photovoltaic output from the power data of each time step. The Transformer encoder uses a multi-layer stacked structure, combined with position encoding and self-attention mechanism, to effectively capture the complex temporal features of photovoltaic output sequences, including diurnal periodic fluctuations and seasonal trend changes. (32) Based on the forecast of each substation Based on the meteorological data of each time step, the GraphSage network is used to extract the spatial correlation characteristics of the meteorological conditions in each substation. Each substation in the sub-region is abstracted as a node in the graph structure, and the similarity of meteorological elements is used as the node connection weight. Through neighborhood sampling and information aggregation operations, the spatial correlation characteristics of each substation under the influence of meteorological factors are dynamically learned. After multiple iterative optimizations, the GraphSage network outputs the spatial feature vector of each substation. (33) Tensor splicing of the temporal features extracted by the Transformer encoder and the spatial features generated by the GraphSage network to form a joint feature vector containing spatiotemporal coupling information; (34) Based on the extracted spatiotemporal aggregation features, the Bilstm module is used to fit the mapping relationship between the aggregation features and power, and the photovoltaic power of all substations in the sub-region is output synchronously.
7. The method for ultra-short-term prediction of distributed photovoltaic power in a substation area according to claim 6, characterized in that: In step (4), specifically, Jointly train the spatiotemporal feature extraction module on the source region data S, and the optimization goal is to minimize the prediction loss function : (13) (14) in, is the complete model prediction function, The input consists of historical power and meteorological data, is the real power output; after the BiLSTM module fits the spatiotemporal aggregation features, it passes through the fully connected layer Mapped to the prediction space, Represents the predicted value of all samples on the source region training data S The square of the difference from the true value y is averaged, which is the expected form of the mean square error; Freeze the core parameters θ of the Transformer encoder and GraphSAGE models T ∗ ,θ G ∗ , retaining general feature extraction capabilities: (15) (16) Building a lightweight backbone network , only retaining the fine-tunable parameters of BiLSTM and fully connected layers: (17) On the target sub-region data, only the BiLSTM parameters and fully connected layer parameters Perform local fine-tuning, and the optimization goal is to minimize the sub-region specific loss function: (18) Update through iteration and The model adapts to the grid topology and meteorological differences of different sub-regions, and achieves rapid cross-regional modeling while maintaining a lightweight architecture. Then the corresponding Sub-region data The average loss on , ensures that the model focuses on the sample characteristics of the region when adapting to the new region.
8. The method for ultra-short-term prediction of distributed photovoltaic power in a substation area according to claim 7, characterized in that: In step (4), during the fine-tuning process, small batch data and stepped learning rate are used to fine-tune the model parameters. An appropriate learning rate is initially set and gradually reduced at the end of each cycle. The learning rate decay step size and learning rate decay ratio are gradually fine-tuned according to the prediction performance of the model in different sub-regions to help the model gradually converge and avoid oscillations during training.
9. The method for ultra-short-term prediction of distributed photovoltaic power in a substation area according to claim 8, characterized in that: In step (4), in order to evaluate the accuracy of the model prediction, the root mean square error and determination coefficient are used to evaluate the model and test the photovoltaic power prediction performance. The specific calculation formula is shown as follows: (19) (20) Where: 、 Represents the RMSE and R between the predicted power and the true power respectively. 2 value, represents the number of samples, Indicates the sample serial number, 、 Respectively represent The true and predicted power values of samples, Represents the mean of the true values.
10. A distributed photovoltaic power ultra-short-term prediction system for a substation, characterized by: include: The data processing module is used to collect the historical measured power data of all substations in the forecast area and the meteorological data of the corresponding location of each substation, and perform data processing; The sub-region division module is used to calculate the dynamic time warping distance using the historical measured power data of the substation area. This is used as a similarity metric to evaluate the output of the substations. The agglomerative hierarchical clustering algorithm is used to divide the photovoltaic substations in the region into sub-regions with consistent output. The prediction module is used to build a Transformer and GraphSAGE collaborative architecture to achieve spatiotemporal feature decoupling and tensor fusion. It uses Transformer self-attention to capture time-varying features, models spatial ultra-short-term meteorological correlations through GraphSAGE neighborhood aggregation, and uses a BiLSTM model to fit composite spatiotemporal features. It simultaneously outputs the photovoltaic power prediction results for each sub-region, avoiding repeated modeling for each sub-region. The total power calculation module builds a suitable migration modeling strategy based on the shared features between sub-regions, freezes the parameters of the spatiotemporal feature extraction backbone network, and only performs lightweight adaptation operations on the region-specific connection layer to achieve fast and low-cost modeling of each sub-region. Finally, the total regional power is obtained by accumulating the power of each sub-region.
Citation Information
Patent Citations
Transformer reverse weight / overload early warning method and system based on neural network
CN118396193A
Distribution area distributed photovoltaic power prediction method and system
CN120109786A
Photovoltaic power generation capability prediction method and system based on space-time diagram neural network, and medium
CN120146291A
Cited By
Secondary frequency control method and device based on graph space-time Transform and storage medium
CN122292412A
A method, apparatus, and storage medium for secondary frequency control based on graph-space-time Transformer
CN122292412B