A transformer area distributed photovoltaic power ultra-short-term prediction method and system

By employing agglomerative hierarchical clustering and graph neural network models, the problems of insufficient prediction accuracy and limited cross-regional adaptability in distributed photovoltaic systems are solved, achieving efficient and accurate photovoltaic power prediction. This method is applicable to grid topology differences in different geographical regions and improves the accuracy of grid dispatching.

CN120806264BActive Publication Date: 2026-03-24SHANDONG UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-15
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Existing photovoltaic power prediction methods suffer from insufficient prediction accuracy, limited cross-regional adaptability, inadequate utilization of meteorological collaborative information, and insufficient characterization of sub-regional heterogeneity when dealing with distributed photovoltaic systems. This results in large prediction errors and makes it difficult to adapt to the differences in grid topology in different geographical regions.

Method used

A condensed hierarchical clustering algorithm is used to divide the photovoltaic power distribution area into sub-regions with homogeneous power output characteristics. By combining the GraphSAGE graph neural network and the Transformer encoder model, spatiotemporal feature decoupling representation is achieved. Furthermore, a transfer modeling strategy is constructed by sharing features to quickly and lightweightly model the photovoltaic power prediction of each sub-region.

Benefits of technology

It improves the accuracy and consistency of photovoltaic power prediction, reduces training overhead, enables rapid collaborative prediction of photovoltaic power in massive distribution areas, adapts to the differences in grid topology in different geographical regions, and improves the accuracy of grid dispatch.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120806264B_ABST
    Figure CN120806264B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of table area distributed photovoltaic power ultra-short term prediction method and system, belong to distributed photovoltaic power prediction technical field.Based on the condensed hierarchical clustering algorithm, the table area photovoltaic in region is clustered into the homogeneous sub-region set of output characteristics, combined with GraphSAGE graph neural network and Transformer encoder model, the decoupling representation of space-time characteristics in sub-region is realized, based on composite space-time characteristics, finally synchronously output the photovoltaic power prediction result of each table area in sub-region, further based on the shared feature between sub-region, construct suitable migration modeling strategy, realize the fast lightweight modeling of each sub-region prediction model, finally, the spatial aggregation of sub-region prediction power obtains regional total power prediction result.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to a transformer area distributed photovoltaic power ultra-short-term prediction method and system, and belongs to the technical field of distributed photovoltaic power prediction. BACKGROUND

[0002] As a key path of clean energy transformation, distributed photovoltaic power generation is accelerating penetration into the end of the distribution network in a large-scale trend. This trend, while bringing many advantages to the power grid, also poses new challenges to the safe operation and regulation and management of the power grid. Photovoltaic output is easily affected by weather factors and has strong randomness and volatility, and the confidence of the prediction results is low. This uncertainty exacerbates the short-term load fluctuation of the distribution network, significantly increasing the difficulty of load prediction. Because the massive low-voltage distributed photovoltaic power is still in the "blind adjustment state" that is unobservable and uncontrollable, it further exacerbates the difficulty of power grid consumption and increases the adjustment pressure of power balance. In order to accurately depict the distribution of photovoltaic power generation in the distribution network, high-precision power prediction technology for low-voltage distributed photovoltaic power is urgently needed.

[0003] Existing photovoltaic power prediction methods can be divided into physical modeling, statistical methods, and data-driven methods. Physical methods mainly focus on the internal physical modules of photovoltaic power generation systems, establish mathematical models, and directly calculate photovoltaic power generation output. Statistical methods establish photovoltaic power output prediction mathematical models based on statistical rules in acquired meteorological data and historical output data. The prediction models established by this method are relatively simple, but the prediction accuracy and stability are relatively poor. Data-driven methods achieve power prediction by mining the nonlinear mapping relationship between meteorological data and power data. This method can effectively capture the complex nonlinear relationship between meteorological and power factors, and the prediction performance is often good, so it has become the most mainstream method at present. Although the above methods have achieved good prediction results, they mainly rely on the data of a single station for modeling, resulting in prediction models that are only applicable to specific stations and are prone to overfitting. Therefore, these methods are not completely suitable for the distributed photovoltaic power prediction application scenario of "many points and wide surfaces".

[0004] It is worth noting that due to the similar fluctuation trend of the distributed photovoltaic output with close spatial position, the output is also inherently autocorrelated in the time dimension, so the existing researches are mostly based on the analysis of the spatio-temporal correlation of the distributed photovoltaic output and the centralized photovoltaic output and the distributed photovoltaic output. The method of predicting the target station output by using the output of the adjacent centralized photovoltaic station is relatively simple, but it does not consider the strength of the correlation between the stations, and in practice, it is impossible to ensure that there are centralized photovoltaic stations around all distributed photovoltaic stations, so the applicability is limited. In contrast, making full use of the output correlation characteristics between the distributed photovoltaic stations for modeling can not only effectively improve the prediction accuracy, but also has a relatively wide range of application. However, there are still some limitations in such methods at present: (1) Insufficient description of the heterogeneity of sub-regions: ignoring the differences in the output distribution characteristics of different small areas, the same set of prediction systems is used for all the stations in the region, which leads to the fact that the model focuses on global characteristics and ignores local characteristics of sub-regions, and a single global model cannot effectively capture these differences, thereby increasing the error of the prediction results. (2) Inadequate use of meteorological coordination information: only relying on historical output data to mine the spatio-temporal correlation characteristics, without effectively using the prediction meteorological coordination information in the region, ignoring the important influence of the prediction meteorological on the ultra-short-term power prediction, and the prediction model based on historical data alone cannot fully capture the real-time feedback of meteorological changes on power fluctuations, which may lead to large deviations in the prediction results. (3) Limited adaptability across regions: most power prediction models focus on specific stations in a particular region, and it is difficult to adapt to the differences in meteorological characteristics and power grid topology in different geographical regions, which limits the range of engineering applications. In summary, there is an urgent need for a low-cost and efficient technology that can quickly build a targeted prediction model for a large number of photovoltaic stations, while fully utilizing the power generation synergy between the stations, to accurately aggregate the power of each station, in order to better serve regional power dispatching and power grid stability analysis. SUMMARY

[0005] In view of the deficiencies of the prior art, the present application provides a kind of substation distributed photovoltaic power ultra-short-term prediction method and system, based on condensation hierarchical clustering algorithm, the photovoltaic of station in region is clustered into the output characteristic homogenization sub-region set, combining GraphSAGE chart neural network and Transformer encoder model, realize the decoupling representation of the space-time characteristics in sub-region, based on composite space-time characteristics finally synchronously output the photovoltaic power prediction result of each station in sub-region, further based on the shared characteristics between sub-regions, build suitable migration modeling strategy, realize the quick lightweight modeling of each sub-region prediction model, finally obtain the regional total power prediction result through the space aggregation of sub-region prediction power.

[0006] The technical scheme of the present application is as follows:

[0007] A kind of substation distributed photovoltaic power ultra-short-term prediction method, steps are as follows:

[0008] (1) Collect the historical measured power data of all sub-stations in the region to be predicted and the meteorological data of each sub-station corresponding position extracted through measured reanalysis data, and perform data processing;

[0009] (2) Calculate the dynamic time warping distance using the historical measured power data of the sub-station as the similarity measurement index between the output of the sub-station, and divide the photovoltaic sub-stations in the region into sub-regions with consistent output by using the condensed hierarchical clustering algorithm;

[0010] (3) Construct a collaborative architecture of Transformer and GraphSAGE to realize the decoupling and tensor fusion of time and space features, capture time-varying features by using Transformer self-attention, model spatial super-short-term meteorological correlation through GraphSAGE neighborhood aggregation, fit the composite time and space features by using BiLSTM model, and simultaneously output the photovoltaic power prediction results of each sub-station in the sub-region, avoiding repeated modeling for each sub-station;

[0011] (4) Based on the shared features between the sub-regions, a suitable transfer modeling strategy is constructed, the time and space feature extraction backbone network parameters are frozen, only the region-specific connection layer is subjected to lightweight adaptation operation, the rapid and low-cost modeling of each sub-region is realized, and finally the total power of the region is obtained by accumulating the power of each sub-region.

[0012] According to the present application, in step (1), the data processing specifically includes missing value processing, outlier processing and normalization processing;

[0013] The KNN algorithm is selected as the missing value repair method, and the mathematical expression is as follows:

[0014] (1)

[0015] (2)

[0016] In the formula: and are two sample points, each having n characteristic values, and are the values of points and on the first characteristic, is the value of the K closest sample to the missing value, is the Euclidean distance between the two samples, used to select the K nearest neighbors of the missing data point, and the missing value is filled with the average value of the K nearest neighbors.

[0017] Outlier handling employs the box plot method (IQR) to detect outliers;

[0018] Normalization is performed by standardizing the sample data using the maximum-minimum standard method, which maps the sample data linearly to [0,1].

[0019] (3)

[0020] In the formula: This represents the standardized sample data; This represents the sample data to be standardized. and These represent the maximum and minimum values ​​in this type of sample data, respectively.

[0021] According to a preferred embodiment of the present invention, in step (2), specifically, the photovoltaic output of the distribution area under typical weather conditions of clear sky, cloudy day and rainy day is taken as the output feature. The time axis of the output data under the three weather conditions is strictly aligned. The DTW (Dynamic Time Warping) distance is used as the metric for the similarity of the distributed photovoltaic output of the distribution area. The AHC clustering algorithm is used to classify the photovoltaic of the distribution area based on the power grid topology to ensure that the clustering results have logical consistency with the physical connection.

[0022] Assume the photovoltaic power sequences of the two photovoltaic substations are as follows:

[0023] (4)

[0024] (5)

[0025] In the formula: Represents power sequence middle The power value at any given time; Represents power sequence middle The power value at a given time, and the normalized path can reveal similar time points in the power sequence, thus achieving flexible alignment of the power sequence. (Definition) and The regularized path is :

[0026] (6)

[0027] In the formula: The form is The path starts from Start to The final regularized path is the one with the shortest distance and is monotonically increasing.

[0028] Photovoltaic power sequence The similarity between is:

[0029] (7)

[0030] (8)

[0031] wherein, represents the initial similarity between the time and in the given two photovoltaic power sequences and ; represents the similarity between the time and of the updated two photovoltaic power sequences and , which is recursively updated by minimizing the previous similarity measure value, to calculate the best matching sequence similarity, and finally the cumulative distance of the dynamic bending path satisfying the constraint condition is calculated by a recursive algorithm,

[0032] The specific steps based on AHC and DTW algorithm are as follows:

[0033] 1) Take each photovoltaic in the area as a separate cluster;

[0034] 2) Calculate the similarity between any two areas according to the DTW algorithm, represented by matrix M, wherein , , represent different areas respectively;

[0035] 3) Find the two clusters with the largest similarity and merge them into a new cluster ;

[0036] 4) According to the newly formed cluster, move the number of the following cluster one position forward, delete the first row and column of matrix , update matrix , , and update the current number of clusters ;

[0037] 5) Repeat steps 3) and 4) until the distance threshold between different clusters is reached;

[0038] 6) Based on the physical connection relationship between the areas in the power grid topology, ensure that the clustering result of each area is logically consistent with its physical connection in the power grid, and finally output the clustering result;

[0039] The silhouette coefficient (SC) is introduced as an evaluation function of the clustering result after clustering until the final best clustering result is obtained, and the calculation formula of the silhouette coefficient is as follows:

[0040] (9)

[0041] Wherein, The silhouette coefficient of the transformer area; The average distance between the transformer area and other samples in the same cluster, called cohesion degree; The average distance between the transformer area and all samples in other clusters, called separation degree; the average silhouette coefficient is the average value of the silhouette coefficients of all samples, and the value range is [-1, 1], the greater the value, the smaller the distance in the cluster, the greater the distance between clusters, and the better the clustering effect. According to the application, in step (3),

[0042] The transformer encoder is stacked by multiple encoder layers, each encoder layer includes an attention sublayer and a feedforward neural network sublayer, the attention sublayer is composed of a multi-head self-attention mechanism, a residual connection and layer normalization; the feedforward neural network sublayer includes a feedforward neural network, a residual connection and layer normalization, wherein the multi-head self-attention mechanism is the core part in the transformer architecture, which models the dependency relationship between each position and other positions of the input sequence to capture the global information in the sequence, and the attention score calculation adopts the key-value-query mode, which makes the key only focus on the first n important queries, that is:

[0043]

[0044] (10)

[0045] In the formula: Attention(·) is an attention score calculation function, Q , K , V Query, key and value matrices respectively; is the dimension of K ; T is the matrix transposition operation, the attention score is calculated by assigning weights to each position in the input sequence, and Softmax is a normalization function;

[0046] ​​The GraphSage network defines two key functions: AGGREGATE(·) and CONCAT(·). AGGREGATE(·) aggregates information from a node's neighbors, while CONCAT(·) combines the features of the current node with the aggregated features. The photovoltaic power station area is abstracted as nodes in the graph, and edge connections are constructed based on meteorological similarity to form a spatial relational graph. v represents one of the nodes, and for any node's neighbors... The node embedding of the Kth layer is represented as The update process is as follows:

[0047] (11)

[0048] (12)

[0049] Where AGGREGATE(·) represents the aggregation operation, and the average aggregation method is selected here; CONCAT(·) represents the join operation; σ represents the activation function; and Let represent the aggregated node features of the neighboring neighborhood of node v and the node features of the neighboring nodes of node v, respectively. ;

[0050] A model framework for synchronous extraction of spatiotemporal correlation features for distributed photovoltaic power generation in sub-regions is constructed, and synchronous prediction of power output for each sub-region is achieved. The specific steps are as follows:

[0051] (31) Based on the history of each station area The power data at each time step is used to extract the temporal features of photovoltaic power output using the Transformer encoder module. The Transformer encoder effectively captures complex temporal features in the photovoltaic power output sequence, including diurnal periodic fluctuations and seasonal trend changes, through a multi-layer stacked structure combined with position encoding and self-attention mechanism.

[0052] (32) Forecast based on each substation area Meteorological data at each time step were used to extract the spatial correlation features of meteorological data in each station area using the GraphSage network. Each station area in the sub-region was abstracted as a node in the graph structure. The similarity of meteorological elements was used as the node connection weight. Through neighborhood sampling and information aggregation operations, the spatial correlation features of each station area under the influence of meteorology were dynamically learned. Through multiple iterations and optimizations, the GraphSage network outputs the spatial feature vector of each station area.

[0053] (33) The time sequence features extracted by the Transformer encoder are tensor spliced with the spatial features generated by the GraphSage network to form a joint feature vector containing space-time coupling information;

[0054] (34) Based on the extracted space-time aggregation features, a Bilstm module is used to fit the mapping relationship between the aggregation features and the power, and the photovoltaic power of all sub-regions is output simultaneously. Considering the complex causal relationship contained in the space-time aggregation features and the bidirectional dependence of information in time series transmission, the BiLSTM can simultaneously capture the forward and backward dependent information of the data through the parallel architecture of the forward and backward LSTM units, effectively solving the limitations of the one-way recurrent neural network in information utilization. Through end-to-end training, the model takes the actual power value of multiple sub-regions as the supervision signal, iteratively optimizes the network parameters until the prediction error converges to the preset threshold, and finally realizes high-precision collaborative prediction of the photovoltaic output of multiple sub-regions.

[0055] According to the application, in step (4), specifically,

[0056] The space-time feature extraction module is jointly trained on the source area data S, and the optimization target is to minimize the prediction loss function :

[0057] (13)

[0058] (14)

[0059] wherein, is the complete model prediction function, is the input composed of historical power and meteorological data, is the real power output; after the BiLSTM module fits the space-time aggregation features, it is mapped to the prediction space through a fully connected layer , represents the difference between the predicted value and the true value y of all samples on the source area training data S, and the average of the square of the difference is the mean square error (MSE) expectation form;

[0060] The core parameters θ of the Transformer encoder and the GraphSAGE model are frozen T ∗ , θ G ∗ , the general feature extraction capability is retained:

[0061] (15)

[0062] (16)

[0063] Constructing lightweight backbone network , only the BiLSTM and fully connected layer parameters are kept:

[0064] (17)

[0065] On the target sub-region data, only the BiLSTM parameters and the fully connected layer parameters are locally fine-tuned, and the optimization goal is to minimize the sub-region-specific loss function:

[0066] (18)

[0067] Through iterative updates and , the model adapts to the differences in power grid topology and weather in different sub-regions, achieving fast modeling across regions while maintaining a lightweight architecture, , the average loss on the th sub-region data , ensuring that the model focuses on the sample characteristics of the region when adapting to new regions.

[0068] Further, during the fine-tuning process, if a large batch is used for training, the model may quickly adjust its weights, overfitting the data of the target sub-region, causing the model to forget the knowledge learned from the source sub-region. Similarly, the learning rate controls the step size of each update, and if a too high learning rate is used, it may lead to over-adjustment and loss of knowledge from the source sub-region. Therefore, this embodiment uses small batch data and stepwise learning rate to fine-tune the model parameters, initially setting a suitable learning rate, which is gradually reduced at the end of each cycle. The learning rate decay step and learning rate decay ratio are fine-tuned according to the prediction performance of the model in different sub-regions, helping the model to gradually converge and avoiding oscillation in the training process.

[0069] Further, in order to evaluate the accuracy of the model prediction, the Root Mean Square Error (RMSE) and the R-squared (R 2 ) are used to evaluate the performance of the model in predicting photovoltaic power, and the specific calculation formula is as follows:

[0070] (19)

[0071] (20)

[0072] In the formula: , respectively represent the RMSE and R 2 value between the predicted power and the true power, Indicates the number of samples. Indicates the sample sequence number. , They represent the first The true and predicted power values ​​for each sample. This represents the average of the true values.

[0073] A distributed photovoltaic power ultra-short-term forecasting system for a transformer substation includes:

[0074] The data processing module is used to collect historical power measurement data of all stations in the area to be predicted and meteorological data of the corresponding location of each station, and to process the data.

[0075] The sub-region division module is used to calculate the dynamic time warping distance using the historical power measurement data of the transformer area, which serves as a similarity metric for evaluating the power output between transformer areas. It uses agglomerative hierarchical clustering algorithm to divide the photovoltaic transformer areas within the region into sub-regions with consistent power output.

[0076] The prediction module is used to build a collaborative architecture of Transformer and GraphSAGE, realize the decoupling of spatiotemporal features and tensor fusion, use Transformer self-attention to capture time-varying features, model spatial ultra-short-term meteorological correlations through GraphSAGE neighborhood aggregation, use BiLSTM model to fit composite spatiotemporal features, and simultaneously output the photovoltaic power prediction results of each sub-region, avoiding repeated modeling of each sub-region.

[0077] The total power calculation module constructs a suitable transfer modeling strategy based on the shared features between sub-regions, freezes spatiotemporal features to extract backbone network parameters, performs lightweight adaptation operations only on region-specific connection layers, realizes fast and low-cost modeling of each sub-region, and finally obtains the total regional power by accumulating the power of each sub-region.

[0078] The beneficial effects of this invention are as follows:

[0079] 1. This invention employs a dynamic time warping distance method to measure the similarity of power output in photovoltaic power distribution areas. This method overcomes the limitations of nonlinear representation by Euclidean distance, solves the time-series offset problem, and accurately measures the photovoltaic power output characteristics of power distribution areas. Furthermore, it combines agglomerative hierarchical clustering algorithms to divide sub-regions with consistent power output, ultimately determining the optimal sub-region division. This adapts to complex weather conditions and the heterogeneity of power distribution areas, providing a homogenized basis for decoupling spatiotemporal features.

[0080] 2、The application constructs a feature extraction model of the collaborative Transformer encoder and the GraphSAGE graph neural network, realizes the joint fitting of the sub-regional whole-area power by combining the BiLSTM network through the space-time feature decoupling and tensor fusion mechanism, breaks through the representation bottleneck of the traditional single model for the space-time coupled features, guarantees the accuracy and consistency of the prediction results, and effectively improves the overall performance of the regional photovoltaic power prediction.

[0081] 3、The application designs a lightweight transfer learning strategy based on sub-regional feature sharing, realizes the rapid modeling of the cross-sub-regional model by freezing the backbone network parameters and adaptively fine-tuning the region-specific connection layer, and finally obtains the regional total power by accumulating the power of each sub-region. Breakthrough the regional limitations of single model, reduce the training cost, realize the rapid collaborative prediction of massive transformer area photovoltaic. BRIEF DESCRIPTION OF DRAWINGS

[0082] Figure 1 The flow chart of the transformer area distributed photovoltaic power ultra-short term prediction method provided by the embodiment of the application.

[0083] Figure 2 The output comparison chart of adjacent transformer areas on sunny and cloudy days is analyzed for the embodiment of the application, wherein, Figure 2 Fig. (a) is a sunny output chart of adjacent transformer areas, Figure 2 Fig. (b) is a sunny output chart of adjacent transformer areas.

[0084] Figure 3 The flow chart of the transformer area distributed photovoltaic power ultra-short term prediction method provided by the embodiment of the application.

[0085] Figure 4 The display chart of the Transformer embedding method of the embodiment of the application.

[0086] Figure 5 The structure chart of the space-time feature decoupling module of the embodiment of the application. DETAILED DESCRIPTION

[0087] The application will be further described below by embodiments and in conjunction with the drawings, but is not limited thereto.

[0088] Embodiment 1:

[0089] As shown in the figure, the embodiment provides a transformer area distributed photovoltaic power ultra-short term prediction method, and the steps are as follows: Figure 1 (1) Collect the historical measured power data of all transformer areas in the region to be predicted and the meteorological data of the corresponding position of each transformer area extracted through the measured reanalysis data, and perform data processing;

[0090]

[0091] ​In the operation process of the photovoltaic power generation system, due to the influence of factors such as fault of the collection device and human operation error, there may be inconsistent deviation between the actual power output and the collected data, such deviation will interfere with the learning and training process of the prediction model, from the perspective of data characteristics, the common interference data mainly shows missing values and abnormal values in the time series. If the unprocessed interference data is directly applied to the prediction model construction, it may cause the problem of non-convergence of the model iteration process, or cause the significant decline of the prediction accuracy, therefore, after the data collection is completed, the data must be cleaned.

[0092] Specifically, the data processing includes missing value processing, abnormal value processing and normalization processing.

[0093] In the photovoltaic power time series data, the numerical evolution not only follows the dynamic change rule of the time dimension, but also presents complex space-time coupling characteristics due to the nonlinear fluctuation of meteorological conditions. Influenced by the instantaneousness of meteorological elements such as solar irradiance and environmental temperature, the photovoltaic power data shows significant autocorrelation in the time series, that is, the power value at a certain sampling time has a close space-time dependence relationship with the power and irradiance at adjacent time. The KNN algorithm is based on the similarity measurement principle between data samples, and can effectively capture the local similarity of photovoltaic data in the time series by constructing a distance measurement model in a multi-dimensional feature space. The method takes the K nearest neighbor data of the target sample as a reference, and fills the missing values through weighted or non-weighted aggregation strategy, so as to retain the space-time evolution characteristics and original distribution characteristics of the data to the greatest extent. Based on this, the present application selects KNN algorithm as the missing value repair method to ensure the integrity of the data set and the reliability of the analysis result, which is mathematically expressed as follows:

[0094] (1)

[0095] (2)

[0096] In the formula: and are two sample points, each has n characteristic values, and are the values of points and on the first characteristic, is the value of the K nearest neighbor of the missing value, is the Euclidean distance between the two samples, which is used to select the K nearest neighbor of the missing data point, and the missing value is filled with the average value of the K nearest neighbors, and the KThe value is susceptible to instantaneous outliers, and too large K value may obscure the short-term variation characteristics of the data. After multiple cross-validation experiments, the embodiment selects K = 4.

[0097] outliers, the core idea is to identify extreme values that deviate from the main distribution range by describing the distribution form and dispersion degree of the data. The key step of the outlier detection method based on IQR is to set an outlier judgment standard. Generally, 1.5 times IQR is used to judge whether the data is an outlier. Any data point less than the lower limit or greater than the upper limit (Upper Bound) is considered an outlier. In addition, this embodiment also judges the case where the photovoltaic power is negative as an outlier. For the detected outliers, this embodiment uses the method of missing value interpolation to replace these outliers by calculating new data. In this way, the integrity of the data can be maintained, and the accuracy of the subsequent analysis and modeling process can be ensured.

[0098] Normalization, the sample data is standardized by using the maximum and minimum value standard method, and the sample data is linearly mapped to [0, 1];

[0099] (3)

[0100] In the formula: denotes the standardized sample data; denotes the sample data to be standardized; and respectively represent the maximum and minimum values of the sample data.

[0101] (2) Calculate the dynamic time warping distance using the historical measured power data of the transformer area as an evaluation index of the similarity between the transformer area output. The condensation hierarchical clustering algorithm is used to divide the transformer area photovoltaic in the region into a sub-region with consistent output;

[0102] The strong correlation between meteorological factors and photovoltaic power generation has been widely studied and confirmed. Among them, the short-time scale fluctuations of key environmental parameters such as irradiance and temperature have a significant synchronous influence effect on the photovoltaic output fluctuation characteristics of adjacent transformer areas. For example Figure 2As shown, by analyzing the photovoltaic power generation output data of two adjacent stations under typical sunny (left) and cloudy (right) conditions, it can be found that in the same natural day, the power generation curves of the two adjacent stations show high morphological consistency. Although affected by factors such as cloud movement, there is a power time sequence offset phenomenon in local period, but overall it still shows significant spatio-temporal correlation. Based on the above findings, the photovoltaic power stations with coordinated output trends can be spatially aggregated to further explore the spatio-temporal coupling rules between adjacent stations from the regional scale. This regional aggregation strategy can not only weaken the random fluctuation influence of a single station by relying on group effect, but also enhance the prediction reliability by relying on spatial correlation.

[0103] In addition, in the operation environment of the distribution network station area, the distributed photovoltaic user's power generation characteristics show heterogeneity, which causes the photovoltaic output curve of each user to show complex fluctuation patterns and scale expansion characteristics in the time dimension. This non-uniform output characteristic poses a serious challenge to the traditional power prediction method based on a unified model. To effectively solve the above problems, the embodiment adopts a hybrid method combining agglomerative hierarchical clustering (AHC) and dynamic time warping to construct a refined station photovoltaic output feature classification system. The photovoltaic power stations with similar output characteristics are clustered into several feature groups, so that the photovoltaic output data of users in the same cluster show high homogeneity in fluctuation trend, peak distribution and other dimensions.

[0104] Specifically, the station photovoltaic output under typical sunny, cloudy and rainy weather is taken as the output feature, the output data time axis under the three kinds of weather is strictly aligned, the DTW (Dynamic Time Warping) distance is taken as the measurement index of the similarity of the distributed photovoltaic output of the station, the AHC clustering algorithm is used, and the station photovoltaic is classified based on the grid topology, to ensure that the clustering result is logically consistent with the physical connection;

[0105] DTW is a nonlinear distance measure method widely used in time series analysis, which can effectively deal with the local misalignment of output curves on the time axis caused by weather disturbance, user behavior or other external conditions. Unlike the traditional Euclidean distance, DTW can identify two similar but asynchronous sequences as similar by dynamic matching on the time axis, thus more truly reflecting the output characteristics of residential photovoltaic systems. Based on this distance definition, a similarity matrix between users can be constructed. Then the agglomerative hierarchical clustering algorithm is introduced, which gradually merges the most similar photovoltaic sets in the district from bottom to top, and constructs a hierarchical user clustering tree diagram. This method does not need to pre-set the number of categories, and can naturally form clustering levels according to the DTW distance, adapting to the complexity and diversity of output patterns in space and time dimensions. This clustering algorithm regards each sample as a cluster, then starts to merge clusters with high similarity according to certain rules, and finally all samples form a cluster or reach a certain condition, and the algorithm ends. In practical application, by setting the distance threshold or specifying the clustering level, the required user classification results can be obtained flexibly.

[0106] Suppose the photovoltaic power sequences of two district photovoltaics are:

[0107] (4)

[0108] (5)

[0109] In the formula: represents the power sequence ; represents the power value at time in the power sequence ; represents the power value at time in the power sequence , the warping path can reveal the similar time points in the power sequence, so as to realize the flexible alignment of the power sequence, and the definition of is

[0110] (6)

[0111] In the formula: is in the form of , the path starts from to , and the final warping path to be obtained is the shortest one and monotonically increasing;

[0112] The similarity between the photovoltaic power sequences and is:

[0113] (7)

[0114] (8)

[0115] wherein, denotes the initial similarity between the time instants and in the given two photovoltaic power sequences and ; denotes the similarity between the time instants and in the updated two photovoltaic power sequences and , which is recursively updated by minimizing the previous similarity measure, to calculate the sequence similarity of the best match, and finally the cumulative distance of the dynamic bending path satisfying the constraint condition is calculated by a recursive algorithm,

[0116] The specific steps based on the AHC and DTW algorithms are as follows, as shown in Figure 3 :

[0117] 1) Take each photovoltaic in a district as a separate cluster;

[0118] 2) Calculate the similarity between any two districts according to the DTW algorithm, represented by matrix M, wherein , , represent different districts, respectively;

[0119] 3) Find the two clusters with the greatest similarity and merge them into a new cluster ;

[0120] 4) According to the newly formed cluster, move the number of the following cluster one position forward, delete the first row and column of matrix , update matrix , , and update the current number of clusters ; 5) Repeat steps 3) and 4) until the distance threshold between different clusters is reached;

[0121] 6) Based on the physical connection relationship between the districts in the power grid topology, ensure that the clustering result of each district is logically consistent with its physical connection in the power grid, and finally output the clustering result;

[0122]

[0123] ​​After clustering, the silhouette coefficient (SC) is introduced as an evaluation function for the clustering results until the final optimal clustering result is obtained. The formula for calculating the silhouette coefficient is as follows:

[0124] (9)

[0125] in, Taiwan The profile coefficient; Taiwan District sample The average distance from other samples in the same cluster is called cohesion. for The average distance between all samples in other clusters is called the separation degree; the average silhouette coefficient is the average of the silhouette coefficients of all samples, and its value ranges from [-1, 1]. The larger the value, the smaller the intra-cluster distance and the larger the inter-cluster distance, and the better the clustering effect.

[0126] (3) Construct a collaborative architecture of Transformer and GraphSAGE to achieve decoupling of spatiotemporal features and tensor fusion. Use Transformer self-attention to capture time-varying features, model spatial ultra-short-term meteorological correlations through GraphSAGE neighborhood aggregation, and use BiLSTM model to fit composite spatiotemporal features. Simultaneously output the photovoltaic power prediction results of each sub-region, avoiding repeated modeling of each sub-region.

[0127] The fluctuation characteristics of photovoltaic (PV) power output mainly consist of two parts: one is the internal temporal output characteristic influenced by the Earth's rotation, exhibiting a clear diurnal periodicity; the other is the external fluctuation characteristic influenced by weather changes such as cloud cover and sound effects, exhibiting significant uncertainty. Both have a strong direct correlation with PV power output. Due to the geographical distribution of PV power distribution areas, there is rich spatial dependency information among them. In-depth mining of these two characteristics helps the model better understand the spatiotemporal interactions between different weather conditions and between weather and power output in each distribution area. Therefore, this embodiment constructs a framework for the simultaneous extraction of spatiotemporal correlation features. This framework can capture the complex interactions between spatial and temporal features and efficiently process large-scale data from multiple distribution areas, effectively reducing computation time.

[0128] Although traditional recurrent neural network models such as LSTM and GRU have time series modeling capabilities, their serial computing structure limits parallel efficiency and makes it difficult to effectively capture the cross-correlation characteristics of the time dimension between the transformer models. The Transformer model takes the self-attention mechanism as the core, which can model the global dependence between any time steps in the time series. Its fully parallel computing architecture supports batch processing of multiple transformer time series data as a unified input matrix, modeling the time series characteristics of different transformer areas under the same architecture, and explicitly constructing the time interaction path across transformer areas. The feature embedding strategy adopted by the Transformer model integrates the information of multiple variables at the same timestamp into a single time marker, as shown in Figure 4 The encoder of the Transformer model is responsible for capturing the global dependence between input variables, while the decoder is used to map the extracted deep time series features to the prediction output. Considering the redundancy of the traditional Transformer encoding and decoding architecture in the single time series feature extraction task, a lightweight Transformer encoder is used as the core modeling framework.

[0129] The Transformer encoder is stacked by multiple encoder layers, each of which includes an attention sublayer and a feedforward neural network sublayer. The attention sublayer is composed of multi-head self-attention mechanism, residual connection and layer normalization. The feedforward neural network sublayer contains feedforward neural network, residual connection and layer normalization. The multi-head self-attention mechanism is the core part of the Transformer architecture, which captures the global information in the sequence by modeling the dependence between each position and other positions of the input sequence. The attention score calculation adopts the key-value-query mode, which allows the key to focus only on the first n important queries, that is:

[0130] (10)

[0131] where Attention(·) is the attention score calculation function, Q , K , V are the query, key, and value matrices, respectively; is the dimension of K ; T is the matrix transpose operation, which calculates the attention score by assigning weights to each position in the input sequence, and Softmax is the normalization function.

[0132] In recent years, with the gradual rise of Graph Neural Network (GNN) and its derivative models, researchers have obtained powerful tools in exploring the spatial correlation of time series problems. GraphSAGE is a graph neural network model based on inductive learning. By sampling and aggregating the features of adjacent nodes, it can efficiently learn the spatial embedding representation of nodes without relying on the full graph structure, capturing the disturbance propagation effect of spatial neighbors. When applied to the modeling of district photovoltaic power prediction, it can enhance the model's ability to perceive regional collaborative change patterns. In addition, in multi-region modeling tasks, GraphSAGE can adaptively capture the topological structure differences of PSDTAs in different sub-regions, without the need to reconstruct the graph structure and retrain in each sub-region, achieving shared modeling and good transferability.

[0133] GraphSage network defines two key functions, AGGREGATE(·) and CONCAT(·), AGGREGATE(·) is used to aggregate information from node neighbors, and CONCAT(·) is used to combine the current node features with the aggregated node features. Abstract district photovoltaic power as nodes in the graph, construct edge connection according to meteorological similarity, and form a spatial correlation graph. Let G = (V, E) be a spatial correlation graph, v represents one of the nodes, and for any node neighbor , the K-th layer node embedding representation is , then the update process is:

[0134] (11)

[0135] (12)

[0136] where AGGREGATE(·) represents the aggregation operation, and the average aggregation method is selected here, and CONCAT(·) represents the connection operation; σ represents the activation function; and respectively represent the aggregated node features of the adjacent neighborhood of node v and the node features of the adjacent nodes of node v, where ;

[0137] ​There is rich spatiotemporal correlation information among photovoltaic (PV) outputs in different distribution substations, and in-depth mining and understanding of this information can help improve the accuracy of PV power prediction. The temporal dependence of PV system power output is usually reflected in short-term historical data. This data includes the past operating status of PV systems, reflecting the impact of equipment performance and climate change on power generation, thus helping to capture recent power generation trends and fluctuation patterns, providing good time-series characteristics for future predictions. However, while historical PV output data can reflect the PV output levels between distribution substations to some extent, the output of PV systems is not only affected by historical factors but also driven by spatial factors such as meteorological conditions, geographical location, and environmental changes. The differences in these factors lead to significant variations in spatial correlations between different distribution substations. Furthermore, historical data cannot fully account for the spatial impact of future weather changes on PV power output, which is particularly important for ultra-short-term predictions. Therefore, relying on historical output data to mine spatial correlations between PV systems has certain drawbacks. In contrast, using future forecast meteorological data to mine spatial correlations between PV systems has a more significant advantage. Through future meteorological data, models can pre-capture these spatial differences, thereby more accurately predicting the PV output of different distribution substations. Historical power output data often reveals static spatial characteristics, while the dynamic characteristics of weather forecasting can adapt to changes in spatial characteristics between power stations, thus enhancing the model's sensitivity to spatial differences between power stations.

[0138] Therefore, by constructing a model that combines historical power output data and future meteorological data for each distribution area, the model can simultaneously consider temporal variation patterns and the impact of future environmental changes, capturing richer spatiotemporal features. To this end, this embodiment constructs a spatiotemporal correlation feature synchronous extraction model framework for distributed photovoltaic power generation in distribution areas, and achieves synchronous prediction of power output for each distribution area within a sub-region. The specific steps are as follows:

[0139] (31) Based on the history of each station area The power data at each time step is used to extract the temporal features of photovoltaic power output using the Transformer encoder module. The Transformer encoder effectively captures complex temporal features in the photovoltaic power output sequence, including diurnal periodic fluctuations and seasonal trend changes, through a multi-layer stacked structure combined with position encoding and self-attention mechanism.

[0140] (32) Forecast based on each substation area The meteorological data of each time step is used to extract the spatial correlation characteristics of the meteorology of each substation by using the GraphSage network. Each substation in the subregion is abstracted as a node in the graph structure, and the similarity of meteorological elements is used as the connection weight of the node. Through neighborhood sampling and information aggregation operations, the spatial correlation characteristics of each substation under the influence of meteorology are dynamically learned. Through multiple iterations of optimization, the GraphSage network outputs the spatial feature vector of each substation.

[0141] (33) The time sequence features extracted by the Transformer encoder are spliced with the spatial features generated by the GraphSage network to form a joint feature vector containing spatio-temporal coupling information, as shown in Figure 5 The spatio-temporal aggregation features not only integrate the time evolution law of photovoltaic power output, but also incorporate the spatial correlation characteristics under the influence of meteorological factors, providing rich feature expression for subsequent prediction.

[0142] (34) Based on the extracted spatio-temporal aggregation features, a Bilstm module is used to fit the mapping relationship between the aggregation features and the power, and simultaneously output the photovoltaic power of all substations in the subregion. Considering the complex causal relationship contained in the spatio-temporal aggregation features and the bidirectional dependency of information in time series transmission, BiLSTM can capture forward and backward dependency information through the parallel architecture of forward and backward LSTM units, effectively solving the limitations of one-way recurrent neural networks in information utilization. The model uses the actual power values of multiple substations as the supervision signal to iteratively optimize the network parameters until the prediction error converges to the preset threshold, finally realizing high-precision collaborative prediction of multi-substation photovoltaic power output.

[0143] (4) Based on the shared features between subregions, a suitable transfer modeling strategy is constructed. The parameters of the spatio-temporal feature extraction backbone network are frozen, and only lightweight adaptation operations are performed on the region-specific connection layer to realize fast and low-cost modeling of each subregion. Finally, the total power of the region is obtained by accumulating the power of each subregion.

[0144] Specifically,

[0145] Due to the joint action of time periodicity, climate factors, solar radiation and other factors, the power generation time sequence characteristics between the stations often have consistent regularity and periodicity. For example: photovoltaic power generation is mainly concentrated in the daytime, and the output increases first and then decreases with time, generally reaching the maximum at noon, the strongest at noon, and gradually weakening in the afternoon. Or the power generation at night or on cloudy days is close to zero. The difference in spatial characteristics is a key factor affecting photovoltaic output. In addition, the change of the solar angle will directly affect the receiving light of the photovoltaic panel, but its change also has a certain regularity. Although the specific values of the solar angle are different in different stations, the trend and mode of its change are consistent in different regions. This consistency provides a theoretical basis for extracting the time sequence characteristics by fixing the time sequence model parameters between different stations. Although the stations are relatively close, due to local meteorological differences and local cloud distribution differences, the spatial dependence relationship between the station outputs changes dynamically, which will directly affect the time sequence characteristics of photovoltaic power generation. Therefore, it is necessary to model the spatial characteristics of meteorological fluctuations between stations in a small area in the photovoltaic power prediction model.

[0146] From step S1, the clustered photovoltaic stations are divided into different sub-regions, and the time sequence characteristics of each sub-region have strong consistency, but there is weak spatial correlation due to the geographical distance. If an independent model is trained for each sub-region, it will lead to too long training time and waste of computing resources. Therefore, it is possible to consider sharing the consistent characteristics between stations and only locally adjusting the individual differences between sub-regions, thereby greatly reducing the training time and the demand for computing resources, and enhancing the generalization ability of the model. In view of the above analysis, the embodiment constructs a cross-regional lightweight migration modeling strategy, which provides theoretical support and technical path for efficient collaborative modeling of regional-level distributed photovoltaic systems. Specifically, in the time dimension, based on the time sequence feature extraction ability of the Transformer encoder learned in the typical sub-region, the parameter reuse strategy is directly applied to other stations to efficiently capture the common time-varying regularity, and the GraphSAGE algorithm can dynamically adapt to node representation learning under different topological structures, by weakening the influence of power grid topology, based on a unified model architecture, efficient extraction of spatial characteristics in different regions is realized.

[0147] The core idea of ​​transfer learning is to transfer existing knowledge to new tasks, thereby reducing the sample requirements and training time of the new tasks while improving model performance. This provides a framework for transferring temporal feature extraction model parameters between different sub-regions. To address the need for efficient modeling of multi-sub-region photovoltaic systems, a lightweight parameter transfer optimization strategy is proposed: by freezing the core parameter groups responsible for decoupling spatiotemporal features in the Transformer encoder and GraphSAGE model, a lightweight general feature extraction backbone network is constructed, maximizing the reuse of spatiotemporal feature representation capabilities learned during training in the source region; simultaneously, only the Bilstm module in the model, which is specific to sub-regions, is locally adaptively fine-tuned to adapt to differences in grid topology and meteorological conditions in different sub-regions.

[0148] A spatiotemporal feature extraction module is jointly trained on the source region data S, with the optimization objective being to minimize the prediction loss function. :

[0149] (13)

[0150] (14)

[0151] in, For the complete model prediction function, The input consists of historical power and meteorological data. To obtain the true power output; after fitting the spatiotemporal aggregation features using the BiLSTM module, the output is then processed through a fully connected layer. Mapped to the prediction space, This represents the predicted value for all samples on the source region training data S. The average of the squared differences between the true value y and the actual value y is the expected form of the mean squared error (MSE).

[0152] Freeze the core parameter θ of the Transformer encoder and GraphSAGE model T ∗ θ G ∗ Retains general feature extraction capabilities:

[0153] (15)

[0154] (16)

[0155] Building a lightweight backbone network Only the fine-tunable parameters of BiLSTM and fully connected layers are retained:

[0156] (17)

[0157] On target sub-region data, only BiLSTM parameters are fine-tuned locally with full connection layer parameters being fine-tuned, and the optimization objective is to minimize the sub-region specific loss function:

[0158] (18)

[0159] By iteratively updating and , the model adapts to the differences in power grid topology and weather between different sub-regions, achieving fast modeling across regions while maintaining a lightweight architecture, which corresponds to the average loss on the data of the th sub-region , ensuring that the model focuses on the sample characteristics of the target sub-region when adapting to new regions.

[0160] During fine-tuning, if large batches are used for training, the model may quickly adjust its weights, overfitting the data of the target sub-region, causing the model to forget the knowledge learned from the source sub-region. Similarly, the learning rate controls the step size of each update, and if a too high learning rate is used, it may lead to over-adjustment and loss of knowledge from the source sub-region. Therefore, this embodiment uses small batch data and stepwise learning rate to fine-tune the model parameters, initially setting a suitable learning rate, which is gradually reduced at the end of each cycle. The learning rate decay step and learning rate decay ratio are gradually fine-tuned according to the prediction performance of the model in different sub-regions, helping the model to gradually converge and avoiding oscillation in the training process.

[0161] The migration object of this embodiment is the prediction model between sub-regions within a region, and the effectiveness of migration depends mainly on the similarity between the source domain and the target domain. Since the solar altitude angle of the sub-regions at the same latitude changes similarly, leading to consistent overall trends in their time series characteristics. Therefore, this embodiment uses the idea of transfer learning to improve the generalization ability of the model, without making a detailed division between the source domain and the target domain. Specifically, this embodiment only selects the sub-region with the best clustering effect as the source domain for modeling, maximizing the similarity between the source domain and the target domain, and ensuring that transfer learning can be effectively applied. After obtaining the prediction values of each sub-region substation, the future power prediction values of each substation are accumulated and summed to obtain the total regional power prediction value.

[0162] To evaluate the accuracy of the model's predictions, the Root Mean Square Error (RMSE) and the R-squared (R 2 ) are used to evaluate the model's photovoltaic power prediction performance. The specific calculation formulas are as follows:

[0163] (19)

[0164] (20)

[0165] In the formula: , respectively represent the RMSE, R 2 value between the predicted power and the real power, represents the number of samples, represents the sample sequence number, , respectively represent the real value and the predicted value of the power of the sample, represents the average value of the real value.

[0166] Embodiment 2:

[0167] The embodiment provides a distributed photovoltaic power ultra-short-term prediction system for a transformer area, comprising:

[0168] a data processing module, configured to collect historical measured power data of all transformer areas in a to-be-predicted area and meteorological data corresponding to the positions of each transformer area, and perform data processing;

[0169] a sub-area division module, configured to calculate a dynamic time warping distance by using the historical measured power data of the transformer areas, as a similarity measurement index for evaluating the output of the transformer areas, and divide photovoltaic transformer areas in the area into sub-areas with consistent output by using a condensed hierarchical clustering algorithm;

[0170] a prediction module, configured to construct a Transformer and GraphSAGE collaborative architecture, realize time-space feature decoupling and tensor fusion, capture time-varying features by using Transformer self-attention, model spatial ultra-short-term meteorological correlation by using GraphSAGE neighborhood aggregation, fit composite time-space features by using a BiLSTM model, and simultaneously output photovoltaic power prediction results of each transformer area in the sub-area, so as to avoid repeated modeling for each transformer area;

[0171] a total power calculation module, configured to construct a suitable transfer modeling strategy based on shared features between sub-areas, freeze time-space feature extraction backbone network parameters, and only perform lightweight adaptive operation on a region-specific connection layer, so as to realize rapid and low-cost modeling of each sub-area, and finally obtain a total power of the area by accumulating the power of each sub-area.

Claims

1. A method for ultra-short-term prediction of distributed photovoltaic power in a transformer substation, characterized in that, The steps are as follows: (1) Collect historical power measurement data of all stations in the area to be predicted and meteorological data of the corresponding location of each station, and perform data processing; (2) Calculate the dynamic time warping distance using the historical power measurement data of the transformer area as a similarity index for evaluating the power output of the transformer area, and use the agglomerative hierarchical clustering algorithm to divide the photovoltaic transformer areas in the region into sub-regions with consistent power output. (3) Construct a collaborative architecture of Transformer and GraphSAGE to achieve decoupling of spatiotemporal features and tensor fusion. Use Transformer self-attention to capture time-varying features, model spatial ultra-short-term meteorological correlations through GraphSAGE neighborhood aggregation, and use BiLSTM model to fit composite spatiotemporal features. Simultaneously output the photovoltaic power prediction results of each sub-region, avoiding repeated modeling of each sub-region. The specific steps are as follows: (31) Based on the history of each station area The power data at each time step is used to extract the temporal features of photovoltaic power output using the Transformer encoder module. The Transformer encoder effectively captures the complex temporal features in the photovoltaic power output sequence, including diurnal periodic fluctuations and seasonal trend changes, through a multi-layer stacked structure combined with position encoding and self-attention mechanism. (32) Forecast based on each substation area Meteorological data at each time step were used to extract the spatial correlation features of meteorological data in each station area using the GraphSage network. Each station area in the sub-region was abstracted as a node in the graph structure. The similarity of meteorological elements was used as the node connection weight. Through neighborhood sampling and information aggregation operations, the spatial correlation features of each station area under the influence of meteorology were dynamically learned. Through multiple iterations and optimizations, the GraphSage network outputs the spatial feature vector of each station area. (33) The temporal features extracted by the Transformer encoder and the spatial features generated by the GraphSage network are concatenated by tensors to form a joint feature vector containing spatiotemporal coupling information; (34) Based on the extracted spatiotemporal aggregation features, the BiLSTM model is used to fit the mapping relationship between the aggregation features and the power, and the photovoltaic power of all substations in the sub-region is output synchronously. (4) Based on the shared features between sub-regions, construct a suitable transfer modeling strategy, freeze the spatiotemporal features to extract backbone network parameters, perform lightweight adaptation operation only on the region-specific connection layer, realize fast and low-cost modeling of each sub-region, and finally obtain the total power of the region by accumulating the power of each sub-region. A spatiotemporal feature extraction module is jointly trained on the source region data S, with the optimization objective being to minimize the prediction loss function. : (13) (14) in, For the complete model prediction function, The input consists of historical power and meteorological data. To obtain the true power output; after fitting the spatiotemporal aggregation features, the BiLSTM model is passed through a fully connected layer. Mapped to the prediction space, This represents the predicted value for all samples on the source region training data S. The average of the squared differences between the mean squared error and the true value y is the expected form of the mean squared error. Freezing the core parameter θ of the Transformer encoder and GraphSAGE model T ∗ θ G ∗ Retains general feature extraction capabilities: (15) (16) Building a lightweight backbone network Only the fine-tunable parameters of BiLSTM and fully connected layers are retained: (17) On the target sub-region data, only the BiLSTM parameters... With fully connected layer parameters Perform local fine-tuning, with the optimization objective being to minimize the loss function specific to a sub-region: (18) Through iterative updates and The model adapts to the differences in power grid topology and meteorology in different sub-regions, achieving rapid cross-regional modeling while maintaining a lightweight architecture. Then the corresponding number Sub-region data The average loss is used to ensure that the model focuses on the sample characteristics of the new region when adapting to it.

2. The method for ultra-short-term forecasting of distributed photovoltaic power in a transformer substation as described in claim 1, characterized in that, In step (1), specifically, data processing includes missing value processing, outlier processing, and normalization processing.

3. The method for ultra-short-term forecasting of distributed photovoltaic power in a transformer substation as described in claim 2, characterized in that, In step (2), specifically, the photovoltaic output of the distribution area under clear sky, cloudy day and rainy day is used as the output data. The time axis of the output data under the three weather conditions is aligned. DTW distance is used as the measure of the similarity of distributed photovoltaic output of the distribution area. The AHC clustering algorithm is used to classify the photovoltaic of the distribution area based on the power grid topology to ensure that the clustering results are logically consistent with the physical connection.

4. The method for ultra-short-term forecasting of distributed photovoltaic power in a transformer substation as described in claim 1, characterized in that, In step (3), The Transformer encoder consists of multiple stacked encoder layers. Each encoder layer comprises an attention sublayer and a feedforward neural network sublayer. The attention sublayer consists of a multi-head self-attention mechanism, residual connections, and layer normalization. The feedforward neural network sublayer contains a feedforward neural network, residual connections, and layer normalization. The multi-head self-attention mechanism captures global information in the sequence by modeling the dependencies between each position in the input sequence and other positions. The attention score is calculated using a key-value-query model, ensuring that the key only focuses on the first n important queries. (10) In the formula: Attention(·) is the attention score calculation function. Q , K , V These are the query, key, and value matrices, respectively; d s yes K The dimension; T is the matrix transpose operation, the attention score is calculated by assigning weights to each position in the input sequence, and softmax is the normalization function.

5. The method for ultra-short-term forecasting of distributed photovoltaic power in a transformer substation as described in claim 4, characterized in that, The GraphSage network defines two key functions: AGGREGATE(·) and CONCAT(·). AGGREGATE(·) is used to aggregate information from the neighboring nodes, and CONCAT(·) is used to combine the features of the current node with the aggregated node features. The photovoltaic power station area is abstracted as a node in the graph, and edge connections are constructed based on meteorological similarity to form a spatial association graph. Let graph G = (N, E). n This represents one of the nodes, and for any node's neighbors... , No. k The node features of a layer are represented as The update process is as follows: (11) (12) Where AGGREGATE(·) represents the aggregation operation, and the average aggregation method is selected here; CONCAT(·) represents the join operation; σ represents the activation function; and Let represent the aggregated node features of the neighboring neighborhood of node v and the node features of the neighboring nodes of node v, respectively. .

6. The method for ultra-short-term forecasting of distributed photovoltaic power in a transformer substation as described in claim 5, characterized in that, In step (4), during the fine-tuning process, small batch data and a step-wise learning rate are used to fine-tune the model parameters. An appropriate learning rate is initially set and gradually reduced at the end of each cycle. The learning rate decay step size and learning rate decay ratio are gradually fine-tuned according to the prediction performance of the model in different sub-regions to help the model gradually converge and avoid oscillations during the training process.

7. The method for ultra-short-term forecasting of distributed photovoltaic power in a transformer substation as described in claim 6, characterized in that, In step (4), to evaluate the accuracy of the model prediction, the root mean square error and the coefficient of determination are used to evaluate the model's photovoltaic power prediction performance. The specific calculation formula is shown in the following formula: (19) (20) In the formula: , Representing the RMSE and R between predicted power and actual power, respectively. 2 The values ​​are: m represents the number of samples, j represents the sample sequence number, and y represents the sample number. j , Let d represent the true power value and the predicted power value of the j-th sample, respectively. j This represents the average of the true values.

8. A distributed photovoltaic power ultra-short-term forecasting system for a transformer substation, characterized in that, The method for ultra-short-term forecasting of distributed photovoltaic power in a transformer substation, as described in claim 1, includes: The data processing module is used to collect historical power measurement data of all stations in the area to be predicted and meteorological data of the corresponding location of each station, and to process the data. The sub-region division module is used to calculate the dynamic time warping distance using the historical power measurement data of the transformer area, which serves as a similarity metric for evaluating the power output between transformer areas. It uses agglomerative hierarchical clustering algorithm to divide the photovoltaic transformer areas within the region into sub-regions with consistent power output. The prediction module is used to build a collaborative architecture of Transformer and GraphSAGE, realize the decoupling of spatiotemporal features and tensor fusion, use Transformer self-attention to capture time-varying features, model spatial ultra-short-term meteorological correlations through GraphSAGE neighborhood aggregation, use BiLSTM model to fit composite spatiotemporal features, and simultaneously output the photovoltaic power prediction results of each sub-region, avoiding repeated modeling of each sub-region. The total power calculation module constructs a suitable transfer modeling strategy based on the shared features between sub-regions, freezes spatiotemporal features to extract backbone network parameters, performs lightweight adaptation operations only on region-specific connection layers, realizes fast and low-cost modeling of each sub-region, and finally obtains the total regional power by accumulating the power of each sub-region.

Citation Information

Patent Citations

  • Transformer reverse weight / overload early warning method and system based on neural network

    CN118396193A

  • Photovoltaic power generation capability prediction method and system based on space-time diagram neural network, and medium

    CN120146291A