Harmonic feature-based unreported photovoltaic identification and capacity estimation method and system
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO
- Filing Date
- 2026-07-09
- Publication Date
- 2026-08-07
AI Technical Summary
[0005]本发明提供了基于谐波特征的未报装光伏识别与容量估算方法及系统,用于解决现有技术难以从用户净负荷曲线中准确识别未报装光伏用户,且无法有效估算其装机容量的问题
本发明的技术方案首先基于目标区域内已报装光伏用户的谐波监测数据和备案装机容量,构建反映谐波幅值与装机容量映射关系的谐波-容量关系模型,以及按辐照度等级和曲线类别索引的归一化谐波曲线库。为后续识别与估算提供区域化、本地化的比对基准和定量换算依据。
Smart Images

Figure CN122532926A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power quality monitoring technology, specifically to a method and system for identifying and estimating the capacity of unregistered photovoltaic systems based on harmonic characteristics. Background Technology
[0002] With the rapid development of distributed photovoltaic (PV) power generation, a large number of residential and commercial users are installing PV systems on their rooftops or ground floors. According to power management regulations, users should register their PV systems with the power supply department after installation to facilitate grid dispatching, safety management, and electricity billing. However, in practice, some users install PV systems without completing the registration process, i.e., they are not registered PV users. This not only affects the accuracy of grid load forecasting and power flow calculations but may also lead to power quality problems and safety hazards due to the inconsistent quality of PV equipment, resulting in economic losses and management difficulties for power supply companies.
[0003] Currently, methods for identifying users who haven't registered for photovoltaic (PV) installation mainly include on-site inspections and electricity consumption characteristic analysis. On-site inspections rely on manual verification of each household, which is not only costly in terms of manpower and resources but also difficult to achieve real-time detection and comprehensive coverage, especially for concealed PV panels where identification efficiency is low. Electricity consumption characteristic analysis typically involves analyzing the net load curve changes of the user's meter, such as observing whether the user's active power shows reverse or abnormal decline characteristics during the day, to determine if there are unregistered PV installations. However, these methods are easily affected by fluctuations in user electricity consumption habits and the diversity of load types. When the user's electricity consumption and PV output cancel each other out, the net load curve characteristics may not be obvious, leading to missed or incorrect detections.
[0004] Therefore, how to provide a method that can effectively identify unregistered photovoltaic users and accurately estimate their installed capacity has become a technical problem that urgently needs to be solved in this field. Summary of the Invention
[0005] This invention provides a method and system for identifying and estimating the capacity of unregistered photovoltaic users based on harmonic characteristics, which solves the problem that existing technologies are unable to accurately identify unregistered photovoltaic users from user net load curves and cannot effectively estimate their installed capacity.
[0006] In view of the above problems, the present invention provides a method and system for identifying and estimating the capacity of unregistered photovoltaic systems based on harmonic characteristics.
[0007] In a first aspect, the present invention provides a method for identifying and estimating the capacity of unregistered photovoltaic systems based on harmonic characteristics, including: Based on harmonic monitoring data and registered installed capacity of multiple photovoltaic users in the target area, a harmonic-capacity relationship model and a normalized harmonic curve library for the target area are constructed. The load-side harmonic stripping and normalization processing of the harmonic monitoring data of the user to be identified is performed to obtain normalized harmonic features. The weighted correlation coefficient between the normalized harmonic features and the curves in the normalized harmonic curve library is calculated, and the cumulative confidence level is evaluated and determined. If the cumulative confidence level is greater than or equal to a preset threshold, the user to be identified is determined to be a suspected user who has not registered for photovoltaic installation. The measured harmonic content amplitude of the user to be identified under the set illumination conditions is extracted and input into the harmonic-capacity relationship model, and the estimated installed capacity is output.
[0008] Secondly, the present invention provides a system for identifying and estimating the capacity of unregistered photovoltaic systems based on harmonic characteristics, comprising: The model building module is used to build a harmonic-capacity relationship model and a normalized harmonic curve library for the target area based on harmonic monitoring data and registered installed capacity of multiple photovoltaic users who have registered within the target area. The identification and determination module is used to perform load-side harmonic stripping and normalization processing on the harmonic monitoring data of the user to be identified to obtain normalized harmonic features, calculate the weighted correlation coefficient between the normalized harmonic features and the curves in the normalized harmonic curve library, and evaluate and determine the cumulative confidence level. The capacity estimation module is used to determine that the user to be identified is a suspected unregistered photovoltaic user if the cumulative confidence level is greater than or equal to a preset threshold, extract the measured harmonic content amplitude of the user to be identified under set illumination conditions and input it into the harmonic-capacity relationship model, and output the estimated installed capacity.
[0009] One or more technical solutions provided in this invention have at least the following technical effects or advantages: The technical solution of this invention first constructs a harmonic-capacity relationship model reflecting the mapping relationship between harmonic amplitude and installed capacity based on harmonic monitoring data and registered installed capacity of photovoltaic users already registered in the target area, as well as a normalized harmonic curve library indexed by irradiance level and curve category. This provides regionalized and localized comparison benchmarks and quantitative conversion basis for subsequent identification and estimation.
[0010] Furthermore, load-side harmonic stripping and time-frequency domain normalization are performed on the users to be identified to eliminate interference from their own power load and extract normalized harmonic features that reflect the characteristics of the photovoltaic inverter. Then, a weighted correlation comparison is performed with a curve library, and the cumulative confidence level is evaluated by comprehensively considering results from multiple consecutive days. This improves the purity of photovoltaic feature extraction and the robustness of identification, reducing the risk of misjudgment caused by occasional factors on a single day.
[0011] Finally, when the cumulative confidence level reaches the target, the user is identified as a suspected non-registered photovoltaic user. Under a set illumination period consistent with the model training conditions, the measured harmonic content is extracted, input into the harmonic-capacity relationship model, and the estimated installed capacity is output. Furthermore, it can generate on-site verification work orders by combining the deviation rate of the registered capacity. This achieves a complete closed loop from qualitative identification to quantitative estimation, providing power supply departments with accurate verification basis.
[0012] In summary, the technical solution of the present invention effectively overcomes the shortcomings of traditional net load curve analysis methods, such as susceptibility to fluctuations in user electricity consumption and difficulty in quantitatively estimating capacity. It improves the accuracy of identifying unregistered photovoltaic users while providing reliable capacity data support. Attached Figure Description
[0013] Figure 1 This is a flowchart illustrating the method for identifying and estimating the capacity of unregistered photovoltaic systems based on harmonic characteristics provided by this invention.
[0014] Figure 2 This is a schematic diagram illustrating the calculation logic of cumulative confidence in the method for identifying and estimating the capacity of unregistered photovoltaic systems based on harmonic characteristics provided by this invention.
[0015] Figure 3 This is a schematic diagram of the structure of the unregistered photovoltaic identification and capacity estimation system based on harmonic characteristics provided by the present invention.
[0016] In the attached diagram, the labels representing each component are as follows: Model building module 11, identification and judgment module 12, capacity estimation module 13. Detailed Implementation
[0017] This invention provides a method and system for identifying and estimating the capacity of unregistered photovoltaic users based on harmonic characteristics, which solves the problem that existing technologies are unable to accurately identify unregistered photovoltaic users from user net load curves and cannot effectively estimate their installed capacity.
[0018] It should be noted that the terms "comprising" and "having" are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or modules that are not explicitly listed or that are inherent to these processes, methods, products, or devices.
[0019] Example 1, as Figure 1 As shown, this invention provides a method for identifying and estimating the capacity of unregistered photovoltaic systems based on harmonic characteristics. The method includes: S100: Based on the harmonic monitoring data and registered installed capacity of multiple photovoltaic users in the target area, construct a harmonic-capacity relationship model and a normalized harmonic curve library for the target area.
[0020] This step constructs a harmonic-capacity relationship model that reflects the mapping relationship between the amplitude of photovoltaic harmonic injection and the installed capacity in the region; and a curve library composed of normalized standard harmonic curves generated by each registered user under different operating conditions. This provides a regionalized judgment basis that conforms to the characteristics of the local power grid for the subsequent identification of unregistered photovoltaic systems, and provides a quantitative conversion basis for the estimation of installed capacity.
[0021] In this step, a harmonic-capacity relationship model for the target region is constructed, including: The harmonic content amplitude of each photovoltaic user who has applied for installation is extracted at noon as the characteristic amplitude when the irradiation conditions are similar; Using the registered installed capacity of each photovoltaic user as the independent variable and the characteristic amplitude as the dependent variable, a training sample set is formed by combining the registered installed capacity of all photovoltaic users with the corresponding characteristic amplitude. A machine learning regression model is trained using the training sample set, and the trained model is validated using cross-validation. The model parameters with the smallest root mean square error are selected as the final model parameters. The mapping function between installed capacity and harmonic content amplitude is determined using the final model parameters, thus forming the harmonic-capacity relationship model.
[0022] Specifically, the harmonic content amplitude data during the noon period is first selected from the historical harmonic monitoring data of each photovoltaic user who has applied for installation. The solar irradiance during the noon period is relatively stable and close to the peak. The irradiance conditions at the same time on different dates have strong consistency. Therefore, the average value of the harmonic content amplitude of each day during this period or the corresponding value of a typical sunny day is taken as the characteristic amplitude of the user under similar irradiance conditions.
[0023] For example, there are three photovoltaic users, A, B, and C, in the target area. User A has a registered installed capacity of 3kW. The harmonic current amplitude of each sampling point within the time window of 11:45 to 12:15 every day over the past 30 days is retrieved. After removing cloudy and rainy days, the average value is taken, and its characteristic amplitude is 0.45A. Similarly, User B has a registered capacity of 5kW and a characteristic amplitude of 0.72A; User C has a registered capacity of 8kW and a characteristic amplitude of 1.15A.
[0024] Next, using the registered installed capacity of each registered photovoltaic user as the independent variable and the extracted feature amplitude as the dependent variable, each registered photovoltaic user in the target area is treated as an independent sample record. Each sample consists of a data pair composed of the user's registered installed capacity and its feature amplitude. The data pairs of all registered photovoltaic users are summarized to form a training sample set. For example, the registered capacity and feature amplitude of the above three users are treated as three independent sample records to form a training sample set: [(3kW,0.45A),(5kW,0.72A),(8kW,1.15A)].
[0025] Finally, the selected machine learning regression model is trained using the training sample set. Cross-validation is employed during training to determine the mapping function between installed capacity and harmonic emission amplitude, i.e., the harmonic-capacity relationship model.
[0026] Preferably, the above-mentioned machine learning regression model uses the support vector regression model because the number of photovoltaic users who have registered in this area is usually limited and the sample set is relatively small. Support vector regression has good generalization ability under small sample conditions, can effectively avoid overfitting, and can flexibly fit the nonlinear mapping relationship that may exist between the harmonic content amplitude and the installed capacity through the kernel function.
[0027] Preferably, the cross-validation method uses K-fold cross-validation, which divides the sample set into K subsets, and alternately uses K-1 subsets as the training set and the remaining 1 subset as the validation set, repeating this process K times and calculating the root mean square error each time. Finally, the model parameters with the smallest root mean square error among all validation rounds are selected as the final model parameters.
[0028] The value of K needs to strike a balance between bias and variance. If K is too small, such as K=2, the training set size will be too small each time, resulting in insufficient model training and increased estimation bias. If K is too large, such as K=10, it will increase the number of training epochs and significantly increase computational cost. Since the number of registered photovoltaic users is usually in the tens to hundreds, the sample set size is moderate. Therefore, K=5 is a widely adopted compromise value in the industry for this data volume. It can ensure that the validation set of each epoch has a certain size to stably evaluate the model performance, while also ensuring that the training set accounts for 80% of the total data, ensuring that the model fully learns the harmonic-capacity mapping law in the sample, and keeping the computational cost under control.
[0029] For example, a sample set containing 80 data points is input into a support vector regression model, and a radial basis function is selected as the kernel function. Five-fold cross-validation is used, with 64 training sets and 16 validation sets per round. After five rounds of training and validation, the five root mean square error values are 0.08, 0.06, 0.07, 0.05, and 0.09, respectively. The minimum value is 0.05, so the model parameters corresponding to that round are selected as the final parameters to determine the mapping function. The resulting harmonic-capacity relationship model represents the variation of photovoltaic harmonic amplitude with installed capacity in the region. Subsequently, by inputting the harmonic amplitude of the user to be identified, the estimated capacity can be output.
[0030] In this step, a normalized harmonic curve library for the target region is constructed, including: Calculate the daily average value of harmonic content amplitude for each photovoltaic user registered in the target area on each monitoring day, and divide all monitoring days into several irradiance levels according to the preset quantile interval based on the level of the daily average value. Within each irradiance level, the harmonic content data of all photovoltaic users who have registered for installation under that irradiance level are extracted on the corresponding monitoring day. The harmonic content data is then normalized to obtain a set of normalized harmonic curves under that irradiance level. Unsupervised clustering is performed on the normalized harmonic curve set, and the center curve of each cluster is used as a typical normalized harmonic curve under the irradiance level. All irradiance levels and their corresponding typical normalized harmonic curves are organized into a two-dimensional typical sub-library matrix indexed by irradiance level and curve category, thus forming the normalized harmonic curve library.
[0031] Specifically, the process begins by using the monitoring day as the smallest unit. For each monitoring day, the regional average of the daily harmonic content for all registered photovoltaic users within the target area is calculated, and this regional average represents the overall irradiance level for that monitoring day. The regional averages for all monitoring days are then aggregated, sorted from lowest to highest value, and divided into several irradiance levels according to a preset quantile interval. Each monitoring day corresponds to a unique irradiance level, and the harmonic data for all registered users within that day are assigned to that level.
[0032] Preferably, the aforementioned preset quantile intervals are quartile intervals. Quartiles can adapt to regional data distribution, ensuring a balanced sample size across all levels and achieving an engineering-usable balance between level granularity and model complexity. Using quartile intervals means dividing the sorted data into four levels based on four quantile intervals: 0-25%, 25%-50%, 50%-75%, and 75%-100%, corresponding to low irradiance, low-to-medium irradiance, medium-to-high irradiance, and high irradiance levels, respectively.
[0033] For example, users A, B, and C have registered within the target area, with a total monitoring period of 60 days. On May 1st, the daily average harmonic radiation for user A is 0.38A, for user B it is 0.72A, and for user C it is 0.55A, with a regional average of 0.55A. Similarly, the regional averages for the remaining 59 days are calculated. Ranking the 60 regional averages, the 25th percentile is 0.40A, the 50th percentile is 0.62A, and the 75th percentile is 0.88A. The regional average of 0.55A on May 1st falls within the 25%-50% range; therefore, May 1st is classified as a low to medium irradiance level, and all harmonic data for users A, B, and C on that day are also classified as low to medium irradiance levels. If the regional average on May 15th is 0.91A, falling within the 75%-100% range, then May 15th is classified as a high irradiance level.
[0034] Secondly, within each irradiance level, complete harmonic content time-series data for all registered photovoltaic users were collected for the corresponding monitoring day, i.e., the harmonic content amplitude sequence of each sampling point. Each harmonic content time-series data was normalized to eliminate the differences in absolute amplitude values caused by different installed capacities, ensuring that the normalized curve only reflects the morphological characteristics of the harmonic distribution.
[0035] The normalization method can be achieved by dividing the amplitude of each sampling point in each time series data by the maximum value of that time series data. After processing, a set of normalized harmonic curves for the current irradiance level is obtained.
[0036] For example, in the low to medium irradiance level, harmonic time-series data were collected for 45 monitoring days, including user A on day 10, user B on days 22 and 45, and user C on day 3. Taking user A's data on day 10 as an example, data was collected every 5 minutes, totaling 288 sampling points. The maximum harmonic amplitude was 0.52A. Dividing the amplitude of each of the 288 points by 0.52 yielded a normalized harmonic curve with a value range of [0,1]. The same operation was performed on all 45 data points to form a set of normalized harmonic curves for the low to medium irradiance level.
[0037] Next, an unsupervised clustering algorithm is performed on the set of normalized harmonic curves for each irradiance level, grouping curves with similar shapes into the same cluster. After clustering, the cluster center curve of each cluster is taken as the representative of that cluster, and the center curve is a typical normalized harmonic curve for the corresponding irradiance level.
[0038] The unsupervised clustering of the normalized harmonic curve set includes: Within a preset range of cluster numbers, the cluster numbers are set to different values in turn, and cluster calculations are performed on the normalized harmonic curve set respectively. For each clustering number value, the silhouette coefficient of the clustering result is calculated, and the clustering number corresponding to the maximum silhouette coefficient is determined as the optimal clustering number. The normalized harmonic curve set is then divided into final clusters using the optimal clustering number. The mean curve of all normalized harmonic curves within each cluster is calculated as the center curve of the cluster, and is determined as a typical normalized harmonic curve under the irradiance level.
[0039] Specifically, a preset traversal range is set for the number of clusters K, typically with a lower limit of 2 and an upper limit that can be set based on the sample size, such as taking the square root of the total number of curves N in the set. Within the preset range, different integer values are taken for K, and clustering calculations are performed on the set of normalized harmonic curves under the current irradiance level, recording the results of each clustering.
[0040] Preferably, the unsupervised clustering algorithm described above employs the K-medoids clustering algorithm. Compared to K-means, which uses a virtual mean vector as the cluster centers, K-medoids selects curves that actually exist in the dataset as the cluster centers, making it more robust to outliers and noise.
[0041] For example, there are 45 normalized harmonic curves under low to medium irradiance levels. Taking the square root of the harmonics, which is approximately 6.7, the upper limit of the number of clusters is set to 7, and the number of clusters ranges from K=2 to 7. K-medoids clustering is then performed on the 45 curves sequentially with K=2, 3, 4, 5, 6, and 7, resulting in a total of 6 clustering results.
[0042] Next, for each K value, the average silhouette coefficient of that cluster partition is calculated. The silhouette coefficient comprehensively measures the intra-cluster compactness and inter-cluster separation, and its value ranges from [-1, 1]. The closer the value is to 1, the better the clustering effect. After traversing all K values, the average silhouette coefficients corresponding to each K value are compared, and the K value with the largest silhouette coefficient is selected as the optimal number of clusters. The cluster partition corresponding to this K value is then used as the final clustering result.
[0043] The silhouette coefficient of a cluster is calculated as follows: First, calculate the average distance from a curve in a cluster to all other curves within its cluster. Then, calculate the average distance from the curve to all curves in the nearest other cluster. The silhouette coefficient of a single sample is then calculated as (average distance from the curve to all curves in the nearest other cluster - average distance from the curve to all other curves in its cluster) / the larger of the two. The arithmetic mean of the silhouette coefficients of all curves under the current clustering result is the average silhouette coefficient corresponding to the number of clusters K.
[0044] For example, for 45 curves at the low to medium irradiance level, the clustering results at K=3 are: Cluster 1: 16 curves, Cluster 2: 18 curves, and Cluster 3: 11 curves. Taking curve m in Cluster 1 as an example, the average Euclidean distance from m to the other 15 curves in the same cluster is calculated, resulting in an average distance of 0.12 from the curve to all other curves in its cluster. The average Euclidean distance from m to the 18 curves in its nearest neighbor cluster 2 is calculated, resulting in an average distance of 0.35 from the curve to all curves in the nearest other cluster. Therefore, the single-sample silhouette coefficient is approximately (0.35 - 0.12) / 0.35 ≈ 0.66. After calculating the single-sample silhouette coefficient for each of the 45 curves and taking the arithmetic mean, the average silhouette coefficient at K=3 is 0.61. The average silhouette coefficients for the six clustering results from K=2 to 7 are calculated as follows: 0.42, 0.61, 0.53, 0.48, 0.39, and 0.35, respectively. The silhouette coefficient of 0.61 is the maximum value when K=3, so the optimal number of clusters is determined to be 3. The clusters obtained when K=3 are cluster 1: 16, cluster 2: 18, and cluster 3: 11 as the final clustering division.
[0045] Finally, for each cluster obtained under the optimal number of clusters, all normalized harmonic curves in the cluster are collected, and the arithmetic mean of the amplitude of each curve at the same time is calculated for each sampling point. The mean values of each sampling point are connected to form a mean curve, which is the center curve of the cluster and is determined as a typical normalized harmonic curve under the irradiance level.
[0046] For example, for cluster 1, with 16 curves, at time 00:00, the normalized amplitude of each of the 16 curves is calculated, and the arithmetic mean is 0.13; at time 00:05, the arithmetic mean of the normalized amplitude of the 16 curves is also taken; and so on, traversing all 288 sampling points, connecting the mean values of each point to form the mean curve of cluster 1, which serves as the first typical normalized harmonic curve under the medium and low irradiance level.
[0047] Finally, using all irradiance levels as one dimension of the matrix, and the typical normalized harmonic curves and their categories obtained from clustering under each level as another dimension, a two-dimensional typical sub-library matrix is constructed, indexed by irradiance level and curve category. Each matrix element corresponds to a typical normalized harmonic curve of a specific category under a specific irradiance level, and all matrix elements together constitute the normalized harmonic curve library of the target region.
[0048] For example, if each of the four irradiance levels yields three typical curves, a two-dimensional matrix of 4 rows × 3 columns is formed. The (2,1)th element in the matrix represents the first type of typical curve for the low-to-medium irradiance level, and the (4,3)th element represents the third type of typical curve for the high irradiance level. When comparing the user to be identified, first determine their irradiance level, then extract all typical curves at that level and perform correlation calculations with the normalized harmonic features of the user to be identified, eliminating the need to traverse the entire database and improving comparison efficiency.
[0049] It should be noted that the above values are for illustrative purposes only and do not constitute a limitation on the present invention.
[0050] In summary, this step utilizes harmonic monitoring data and registered installed capacity of photovoltaic users already registered in the target area to construct a harmonic-capacity relationship model reflecting the mapping relationship between harmonic amplitude and installed capacity in the area, as well as a normalized harmonic curve library indexed by irradiance level and curve category. This provides a regionalized comparison benchmark and quantitative estimation basis for subsequent identification of unregistered photovoltaic systems, making identification and judgment based on evidence and capacity estimation based on a model.
[0051] S200: Perform load-side harmonic stripping and normalization processing on the harmonic monitoring data of the user to be identified to obtain normalized harmonic features, calculate the weighted correlation coefficient between the normalized harmonic features and the curves in the normalized harmonic curve library, and evaluate and determine the cumulative confidence level.
[0052] like Figure 2 As shown, this step extracts harmonic components generated by the user's internal load from the harmonic monitoring data collected at the user's meter, retains residual harmonics reflecting the characteristics of the photovoltaic inverter, and normalizes them to obtain normalized harmonic features. The normalized harmonic features are then compared with typical curves at corresponding irradiance levels in the constructed normalized harmonic curve library using weighted correlation calculations, emphasizing harmonic components with significant contributions to identification at specific frequencies or time periods. The correlation results from multiple comparisons are then combined to evaluate a quantitative cumulative confidence index using an accumulative approach.
[0053] In this step, the harmonic monitoring data of the user to be identified is subjected to load-side harmonic stripping and normalization to obtain normalized harmonic characteristics, including: Harmonic content data of the user to be identified during the nighttime hours when there is no photovoltaic power generation is extracted each day, and the average value of the harmonic content data during the nighttime hours of multiple consecutive days is used as the load baseline value of the user to be identified. The net photovoltaic harmonic component sequence is obtained by subtracting the load baseline value from the measured harmonic content at each sampling time of the user to be identified during the day. Read the continuous period in the net photovoltaic harmonic component sequence where the amplitude is continuously higher than the load baseline value, and determine the continuous period as the effective sunshine window for the day; Within the effective illumination window of the day, the maximum value of the net photovoltaic harmonic component sequence is extracted as the normalization reference value. Each data point in the net photovoltaic harmonic component sequence within the effective illumination window of the day is divided by the normalization reference value to complete the normalization process and obtain the time-domain normalized harmonic curve, which serves as the normalized harmonic feature of the user to be identified.
[0054] Specifically, firstly, harmonic content data for the nighttime periods without photovoltaic power generation for the user to be identified is extracted. Nighttime periods typically range from sunset to sunrise the following day, such as 9:00 PM to 5:00 AM. Harmonic content data for this nighttime period is then collected over several consecutive days, such as the last 7 to 14 days. A fixed window length is used to slide along the time axis, and the arithmetic mean of the harmonic content of all sampling points within each window is calculated sequentially. This sliding window mean is used as the current load baseline value for the user, reflecting the inherent harmonic levels of various electrical loads within the user when there is no photovoltaic output. The fixed window length equals the number of sampling points for a complete nighttime period; for example, with a sampling interval of 5 minutes, a total of 96 points are collected over 8 hours.
[0055] For example, for user X to be identified, with a sampling interval of 5 minutes, a total of 96 sampling points are collected from 21:00 on a single day to 05:00 the next day, with a fixed window length of 96. The average value of the 96 points is calculated daily for the past 7 days, yielding values of 0.26A, 0.29A, 0.27A, 0.30A, 0.28A, 0.31A, and 0.25A respectively. The 7-day average value of 0.28A is taken as the baseline value of the load.
[0056] Secondly, the measured harmonic content of the user to be identified at each sampling time during the daytime, such as from 05:00 to 21:00, is subtracted from the load baseline value obtained in the first step at each sampling point to obtain the difference sequence, i.e., the net photovoltaic harmonic component sequence. If the difference at a certain sampling point is less than zero, it is set to zero. Theoretically, the net photovoltaic harmonic component sequence mainly reflects the harmonic components injected by the photovoltaic inverter, and the influence of harmonics from the user's conventional load has been removed.
[0057] For example, if user X's measured harmonic content at 12:00 on a certain day is 0.65A, subtracting the load baseline value of 0.28A, the net photovoltaic harmonic component at that moment is 0.37A. By iterating through all 192 sampling points during the day and calculating the difference at each point, the net photovoltaic harmonic component sequence for that day is obtained.
[0058] Secondly, in the net photovoltaic harmonic component sequence, scanning begins at a certain time after sunrise, detecting continuous periods where the amplitude is consistently higher than the load baseline value. Since the net photovoltaic harmonic component is the result after subtracting the baseline value, continuous sampling intervals with amplitudes consistently greater than zero can be directly detected. These continuous periods meeting the criteria are defined as the effective solar illumination window for the day, corresponding to the actual photovoltaic output time period, thus eliminating interference from intermittent cloud cover or periods of weak morning and evening irradiance.
[0059] For example, if the amplitude of the net photovoltaic harmonic component sequence for user X is continuously greater than zero from 08:15 to 16:45 without interruption, then the effective sunshine window for that day is determined to be from 08:15 to 16:45, lasting 8.5 hours, with a total of 102 sampling points. If the amplitude drops below zero due to cloud cover from 13:00 to 13:30, then the effective sunshine window may be divided into two segments: 08:15 to 13:00 and 13:30 to 16:45, and processed separately.
[0060] Finally, within the effective illumination window of the day, the maximum value of the net photovoltaic harmonic component sequence is found and used as the normalization reference value. Each data point in the net photovoltaic harmonic component sequence within the effective illumination window is divided by this normalization reference value to map the value range of all data points to [0,1], thus completing the normalization process and obtaining the time-domain normalized harmonic curve, which is the normalized harmonic characteristic of the user to be identified. This eliminates the influence of the difference in installed capacity on the absolute value of the amplitude and only retains the morphological characteristics of the harmonic distribution.
[0061] For example, user X has 102 sampling points within the effective lighting window of 08:15 to 16:45 on a given day. The net photovoltaic harmonic component reaches its maximum value of 0.37A at 12:00, which is the normalized baseline value. Dividing each of the 102 sampling points by 0.37 yields a time-domain curve with a normalized value range of [0,1]. For instance, the original value at 08:15 is 0.15A, and the normalized value is 0.15 / 0.37≈0.41; the original value at 12:00 is 0.37A, and the normalized value is 1.00. These time-domain curves serve as the normalized harmonic characteristics of user X on that day.
[0062] In this step, the normalized harmonic characteristics also include frequency domain characteristics, and the steps for obtaining the frequency domain characteristics include: At each sampling time, the amplitudes of the 5th and 7th harmonics are extracted, and the ratio between the amplitudes of the 5th and 7th harmonics is calculated, as well as the ratios between the amplitudes of the 5th and 7th harmonics and the total harmonic distortion rate, to obtain the proportional characteristics between the harmonic orders at the sampling time. Normalization is performed on the time series based on the proportional characteristics between harmonic orders at all sampling times to obtain a frequency domain normalized sequence; Determine whether the amplitude of the 5th harmonic and the amplitude of the 7th harmonic at each sampling time exceed the product of the load baseline value and the preset gate multiple. If they do, the frequency domain normalized feature at the sampling time is valid. The frequency domain normalized feature is then concatenated with the corresponding values of the time domain normalized harmonic curve to form a two-dimensional feature vector at the sampling time. If the sampling time does not exceed the limit, the frequency domain dimension of the sampling time is set to zero, and the time domain normalized feature of the sampling time is used as the feature vector of the sampling time. The feature vectors of all sampling times are combined in chronological order to form the normalized harmonic features of the user to be identified.
[0063] Specifically, at each sampling moment within the effective illumination window for the user to be identified, the amplitudes of the 5th and 7th harmonics are extracted from the harmonic monitoring data. Three ratios are then calculated: the ratio of the 5th to the 7th harmonic amplitude, and the ratios of each of the 5th and 7th harmonic amplitudes to the total harmonic distortion rate at the current sampling moment. These ratios together constitute the proportional characteristics of the harmonic orders at that sampling moment, reflecting the relative relationships of the energy distribution of different harmonic frequencies.
[0064] The amplitude of the 5th harmonic refers to the magnitude of the harmonic component in the current or voltage with a frequency five times the fundamental frequency of 50Hz, i.e., 250Hz; the amplitude of the 7th harmonic refers to the magnitude of the harmonic component in the current or voltage with a frequency seven times the fundamental frequency, i.e., 350Hz. Both are expressed in amperes or volts. The 5th and 7th harmonics are the two most significant low-order characteristic harmonics, with relatively stable amplitudes that are closely related to the inverter type and operating state. Compared to higher-order harmonics, such as the 11th and 13th harmonics, the 5th and 7th harmonics attenuate more slowly and are easier to detect and distinguish when there is less interference on the load side. Compared to the 3rd harmonic, the 5th and 7th harmonics are not affected by zero-sequence component cancellation in three-phase three-wire systems, making them more widely applicable. Therefore, selecting the amplitudes of the 5th and 7th harmonics as the core parameters for frequency domain feature extraction can effectively characterize the harmonic injection characteristics of photovoltaic inverters.
[0065] Total Harmonic Distortion (THD) is a comprehensive indicator that measures the proportion of harmonic content in a current or voltage waveform relative to the fundamental component. It is usually expressed as a percentage and is defined as the ratio of the square root of the sum of the squares of the effective values of all harmonic components to the effective value of the fundamental component. In power monitoring, THD is used to assess the degree to which a waveform deviates from a standard sine wave; a higher value indicates more severe harmonic pollution.
[0066] For example, user X has 102 sampling points during the effective illumination window from 08:15 to 16:45 on a given day. At the sampling time of 12:00, the amplitude of the 5th harmonic is 0.22A, the amplitude of the 7th harmonic is 0.10A, and the total harmonic distortion rate is 4.8%. The following three ratios are calculated: the ratio between the amplitudes of the 5th and 7th harmonics = 0.22 / 0.10 = 2.20; the ratio between the amplitude of the 5th harmonic and the total harmonic distortion rate = 0.22 / 4.8% ≈ 4.58; and the ratio between the amplitude of the 7th harmonic and the total harmonic distortion rate = 0.10 / 4.8% ≈ 2.08. Therefore, the harmonic order ratio characteristics at this time are (2.20, 4.58, 2.08). Extracting these values from all 102 sampling points creates a time series of harmonic order ratio characteristics.
[0067] Secondly, the time series composed of the proportional characteristics of harmonic orders at all sampling times is normalized. The normalization method is as follows: for each of the three proportional dimensions, the maximum value of each dimension within the entire effective illumination window is taken as the normalization reference value for that dimension, and the value of that dimension at each sampling time is divided by the corresponding reference value. After normalization, the value range of each dimension is mapped to [0,1], resulting in a frequency domain normalized sequence.
[0068] For example, within the effective illumination window of user X, 102 sampling points show that the maximum dimension of the ratio between the 5th harmonic amplitude and the 7th harmonic amplitude is 2.80, the maximum dimension of the ratio between the 5th harmonic amplitude and the total harmonic distortion (THD) is 5.20, and the maximum dimension of the ratio between the 7th harmonic amplitude and the THD is 3.10. The proportional characteristics (2.20, 4.58, 2.08) at time 12:00 are normalized to (2.20 / 2.80≈0.79, 4.58 / 5.20≈0.88, 2.08 / 3.10≈0.67). Processing each of the 102 sampling points individually yields the frequency domain normalized sequence.
[0069] Next, for each sampling time, it is determined whether the amplitudes of the 5th and 7th harmonics both exceed the product of the load baseline value and the preset threshold. If the amplitudes of both the 5th and 7th harmonics exceed the threshold, the frequency domain feature at that time is considered valid. The frequency domain normalized feature is concatenated with the time domain normalized value at the same time to form a two-dimensional feature vector, i.e., (time domain value, frequency domain composite value). If they do not exceed the threshold, the frequency domain dimension is set to zero, and only the time domain normalized feature at that time is used as the feature vector, i.e., (time domain value, 0).
[0070] The preset gate factor is set empirically to a value between 1.5 and 3.0, with a preferred preset gate factor of 2.0. The frequency domain composite value can be obtained by taking the arithmetic mean of the three frequency domain normalization ratios.
[0071] For example, if the baseline load is 0.28A and the preset gating factor is 2.0, then the gating threshold is 0.28 × 2.0 = 0.56A. At 12:00, the amplitudes of the 5th harmonic (0.22A) and 7th harmonic (0.10A) are both below 0.56A, so the frequency domain dimension is set to zero, and the feature vector at this time is (1.00, 0). At 13:30, the amplitudes of the 5th harmonic (0.68A) and 7th harmonic (0.35A) are both above 0.56A, so the frequency domain feature is valid. The frequency domain composite value is the average of the three normalization ratios (0.82 + 0.75 + 0.71) / 3 ≈ 0.76, which is concatenated with the time domain normalized value of 0.92 at this time, resulting in a feature vector of (0.92, 0.76).
[0072] Finally, the feature vectors from all sampling times are arranged sequentially in chronological order to form the normalized harmonic feature matrix of the user to be identified. Each row of the matrix corresponds to a sampling time, and the columns correspond to the time domain and frequency domain dimensions, respectively. The normalized harmonic feature matrix is the normalized harmonic feature of the user to be identified, which is used for subsequent comparison with the normalized harmonic curve library.
[0073] For example, within the effective illumination window of user X, there are 102 sampling points, of which 35 sampling points have valid frequency domain features and 67 sampling points have invalid frequency domain features. The 102 feature vectors are arranged in time order from 08:15 to 16:45 to form a 102×2 normalized harmonic feature matrix, which serves as the normalized harmonic feature of user X for that day.
[0074] In this step, the weighted correlation coefficient between the normalized harmonic characteristics and the curves in the normalized harmonic curve library is calculated, and the cumulative confidence level is evaluated and determined, including: During the period when the net photovoltaic harmonic amplitude of the user to be identified exceeds the load baseline value, the net photovoltaic harmonic amplitude corresponding to each sampling time is used as the calculation weight of the sampling time. The correlation between the normalized harmonic characteristics of the user to be identified and each typical normalized harmonic curve in the normalized harmonic curve library is weighted and calculated. The maximum value of the weighted correlation coefficient is taken as the weighted correlation coefficient of the user to be identified on that day. The weighted correlation coefficient is calculated for the user to be identified for K consecutive days. The number of days in which the daily weighted correlation coefficient exceeds the first preset threshold is counted, and the ratio of the number of days to K is calculated as the cumulative confidence of the user to be identified, where K is a preset constant.
[0075] Specifically, within the effective illumination window of the user to be identified for the day, the net photovoltaic harmonic amplitude corresponding to each sampling moment is used as the calculation weight for that sampling moment. The larger the net photovoltaic harmonic amplitude, the stronger the photovoltaic output and the higher the signal-to-noise ratio at that moment, and it should be given a higher weight in the correlation calculation. The comparison range is all typical normalized harmonic curves corresponding to the irradiance level of the user in the normalized harmonic curve library.
[0076] For example, if user X is classified as having a low to medium irradiance level on a given day, the curve library contains three typical normalized harmonic curves. User X has 102 sampling points within its effective illumination window from 08:15 to 16:45. The net photovoltaic harmonic amplitude at each point ranges from 0.05A to 0.37A. The amplitude at each point is used as a weight; for example, if the amplitude at 12:00 is 0.37A, the weight is 0.37; if the amplitude at 08:15 is 0.15A, the weight is 0.15.
[0077] Next, for each typical normalized harmonic curve in the curve library, the weighted correlation coefficient between it and the normalized harmonic characteristics of the user to be identified is calculated. During the calculation, the weight of each sampling time is incorporated into the correlation formula, so that high-amplitude sampling points contribute more to the correlation results. After traversing all typical curves under this irradiance level, the maximum value among the weighted correlation coefficients corresponding to each curve is taken as the user's weighted correlation coefficient for that day.
[0078] Preferably, the weighted correlation coefficient is the weighted Pearson correlation coefficient. The Pearson correlation coefficient measures the degree of linear correlation between two variables, with a value ranging from -1 to 1; a value closer to 1 indicates a stronger positive correlation. Extending it to a weighted form scales the contribution of each sampling point to the correlation coefficient according to its weight, giving higher-amplitude, higher-signal-to-noise-ratio sampling points a greater weight in the correlation calculation and reducing correlation distortion caused by noise interference in low-amplitude segments. Compared to the Spearman rank correlation coefficient, which only focuses on ordinal consistency, or cosine similarity, which only measures directional consistency, the weighted Pearson correlation coefficient considers both the matching degree of numerical magnitude and trend, and is computationally efficient and has strong interpretability, making it more suitable for scenarios involving time-domain curve shape matching.
[0079] For example, the normalized harmonic characteristics of user X are a 102×2 matrix. Weighted Pearson correlation coefficients were calculated with three typical curves. Taking the time domain as an example, the weighted correlation coefficient with the first type of curve is 0.52, with the second type of flat curve is 0.78, and with the third type of curve is 0.61. Taking the maximum value of 0.78 as the weighted correlation coefficient for user X on that day indicates that the harmonic pattern of user X is most similar to the second type of typical photovoltaic characteristics under low-to-medium irradiance in the region.
[0080] Finally, with a preset observation window length K, the weighted correlation coefficient calculation is performed daily for K consecutive days for the user to be identified, resulting in K daily weighted correlation coefficients. A first preset threshold is set, and the number of days within K days where the daily weighted correlation coefficient exceeds the first preset threshold is counted. The ratio of the number of days exceeding the first preset threshold to the total number of days K is calculated as the cumulative confidence level, reflecting the degree to which the user's harmonic characteristics and typical photovoltaic characteristics continuously match over multiple days. A single day's occasional high correlation is insufficient to form a high confidence level.
[0081] The first preset threshold is used to determine whether the daily weighted correlation coefficient meets the standard, and a value range of 0.60 to 0.75 is recommended. A threshold that is too low will introduce more false positives, while a threshold that is too high may lead to false negatives. An example of a threshold of 0.65 can effectively filter out most random fluctuations while ensuring a certain level of recognition sensitivity.
[0082] The recommended observation window length K is 7 to 14 days. If K is too short, such as 3 days, it is easily affected by short-term weather fluctuations, leading to unstable confidence levels; if K is too long, such as 30 days, the identification period is prolonged, which is not conducive to timely detection of users who have not registered. Taking 7 days as an example achieves a balance between identification timeliness and confidence stability.
[0083] For example, setting K=7 and the first preset threshold to 0.65, the weighted correlation coefficient for user X is calculated for 7 consecutive days, with results of 0.78, 0.72, 0.69, 0.58, 0.74, 0.81, and 0.66. Six days have a correlation coefficient exceeding 0.65, with only the 4th day at 0.58 failing to meet the threshold. Therefore, the cumulative confidence level = 6 / 7 ≈ 0.857, or 85.7%, indicating that user X's harmonic characteristics consistently and highly match the typical characteristics of users who have already registered for photovoltaic installations over multiple days, increasing the likelihood that user X will be classified as someone who has not registered for photovoltaic installations.
[0084] It should be noted that the above values are for illustrative purposes only and do not constitute a limitation on the present invention.
[0085] In summary, this step effectively reduces the interference of users' own electricity load on photovoltaic feature identification by stripping the load-side harmonic background and integrating time-frequency domain normalized features, performing weighted correlation comparison with a library of typical photovoltaic harmonic curves in the region, and comprehensively evaluating the cumulative confidence over multiple consecutive days. This improves the accuracy and robustness of identifying users who have not registered for photovoltaic installations.
[0086] S300: If the cumulative confidence level is greater than or equal to the preset threshold, the user to be identified is determined to be a suspected unregistered photovoltaic user. The measured harmonic content amplitude of the user to be identified under the set illumination conditions is extracted and input into the harmonic-capacity relationship model, and the estimated installed capacity is output.
[0087] When the calculated cumulative confidence level reaches or exceeds the preset judgment threshold in this step, it indicates that the harmonic characteristics of the user to be identified have consistently and highly matched the typical characteristics of registered photovoltaic users in the region over multiple days, providing sufficient basis to determine it as a suspected unregistered photovoltaic user. After the determination, the measured harmonic content amplitude of the user under set illumination conditions is further extracted, input into the harmonic-capacity relationship model, and the quantitative result of the user's estimated installed capacity is determined.
[0088] In this step, if the cumulative confidence level is greater than or equal to the preset threshold, the user to be identified is determined to be a suspected user who has not registered for photovoltaic installation. The measured harmonic content amplitude of the user to be identified under the set illumination conditions is extracted and input into the harmonic-capacity relationship model, and the estimated installed capacity is output.
[0089] The preset threshold is the decision threshold for cumulative confidence, used to define the minimum compliance rate of a user's harmonic characteristics with the typical characteristics of regional photovoltaic systems over multiple days. When the percentage of days with an excessive weighted correlation coefficient within K consecutive days reaches or exceeds this threshold, the user is judged as a suspected non-registered photovoltaic user; otherwise, the user is excluded from suspicion or marked for observation. Setting the preset threshold too low can easily lead to false positives, while setting it too high increases the risk of false negatives. A value range of 0.6 to 0.8 is recommended, with an example of 0.7, which strikes a balance between identification accuracy and engineering tolerance requirements. The preset threshold can be dynamically adjusted according to the target area's sunlight and climate conditions and data quality.
[0090] For example, with an observation window length K=7 days and a preset threshold of 0.7, if user X's weighted correlation coefficient exceeds the first preset threshold for 6 out of 7 consecutive days, and the cumulative confidence level 6 / 7≈0.857≥0.7, then the user is judged as a suspected user who has not registered for photovoltaic installation.
[0091] In this step, the measured harmonic content amplitude of the user to be identified under set illumination conditions is extracted and input into the harmonic-capacity relationship model, and the estimated installed capacity is output, including: The midday period is used as the feature extraction period corresponding to the set illumination conditions. The method for determining the feature extraction period is consistent with the period for extracting the characteristic amplitude of each registered photovoltaic user when constructing the harmonic-capacity relationship model. Extract the measured harmonic content amplitude of the user to be identified during the feature extraction period, input the measured harmonic content amplitude into the mapping function in the harmonic-capacity relationship model, and output the estimated installed capacity of the user to be identified.
[0092] Specifically, the feature extraction time period corresponding to the illumination conditions is set to be consistent with the time period used to extract the feature amplitude of registered photovoltaic users when constructing the harmonic-capacity relationship model, i.e., 11:45 to 12:15 every day. Maintaining consistency in the time period ensures that the harmonic content extraction conditions for the users to be identified are the same as those for the model training data, and the irradiance levels are similar, thus making the model input and training samples comparable in terms of illumination conditions and ensuring the reliability of the capacity estimation results. For example, if user X has been identified as a suspected unregistered photovoltaic user through cumulative confidence, the same 11:45 to 12:15 time period is used as the feature extraction window for user X.
[0093] Furthermore, after identifying a user as a suspected non-registered photovoltaic user, the measured harmonic content amplitude of that user within the feature extraction period is extracted. The measured harmonic content of the user during the corresponding period of one or more recent sunny days can be extracted, and the arithmetic mean is used as the input value to reduce the impact of occasional daily fluctuations. The measured harmonic content amplitude is then input into the pre-trained harmonic-capacity relationship model, and after conversion by the mapping function in the model, the estimated installed capacity of the user is output.
[0094] For example, user X measured harmonic amplitudes of 0.58A, 0.62A, and 0.55A respectively during the period from 11:45 to 12:15 on three consecutive sunny days, with an average of 0.583A. Inputting 0.583A into a pre-trained harmonic-capacity relationship model, the model's mapping function outputs an estimated installed capacity of 4.2kW.
[0095] This step, after outputting the estimated installed capacity, also includes: Obtain the installed capacity of the user to be identified, wherein the installed capacity is zero for users who have not completed any installation procedures; Calculate the deviation rate between the estimated installed capacity and the reported installed capacity. If the deviation rate exceeds a preset deviation threshold, the user to be identified is determined as a high-risk user who has not reported photovoltaic installation, and an on-site verification work order is generated.
[0096] Specifically, the process begins by querying the user's registration information with the power supply department to obtain their registered photovoltaic capacity. If the user has never applied for any photovoltaic installation, their registered capacity is recorded as zero; if the user has applied for a portion of the capacity, their registered capacity is recorded.
[0097] For example, user X has no photovoltaic installation record in the power supply department's registration system, and the installed capacity is 0kW. If another user Y has registered a 3kW photovoltaic system with the power supply department, but may have actually increased the capacity without authorization, the installed capacity is 3kW. If another user Z has legally installed a 5kW photovoltaic system, the installed capacity is 5kW.
[0098] Furthermore, the deviation rate between the estimated installed capacity and the reported installed capacity is calculated. The deviation rate is calculated as follows: Deviation rate = |Estimated installed capacity - Reported installed capacity| / Estimated installed capacity × 100%. If the estimated installed capacity is zero, the deviation rate is defined as zero. The calculated deviation rate is compared with a preset deviation threshold. If the deviation rate exceeds the preset deviation threshold, it indicates a significant difference between the estimated capacity and the reported capacity, suggesting a strong suspicion of non-reporting or unauthorized capacity expansion, and the user is classified as a high-risk non-reported photovoltaic user. If the deviation rate does not exceed the threshold, the difference is within a reasonable fluctuation range, and the user's risk level is low.
[0099] The preset deviation threshold is a deviation rate threshold used to determine whether a user belongs to the high-risk category of those who have not applied for photovoltaic installation. It defines the tolerable range of difference between the estimated installed capacity and the actually applied capacity: if the deviation rate exceeds the threshold, it is judged as high-risk and an on-site verification work order is generated; if it does not exceed the threshold, it is considered a reasonable fluctuation. If the threshold is too high, a large number of actual users who have not applied for installation will be missed; if it is too low, too many invalid verification work orders will be generated due to normal estimation errors in the model. It is recommended that the value be between 20% and 50%, with an example of 30%, which can strike a balance between controlling false positives and false negatives. The preset deviation threshold can be dynamically adjusted according to the adequacy of verification resources and the level of cumulative confidence, so as to achieve refined management of risk classification.
[0100] For example, the preset deviation threshold is 30%. User X estimates an installed capacity of 4.2kW, with 0kW already registered. The deviation rate = |4.2-0| / 4.2×100% = 100% > 30%, classifying them as a high-risk user who has not registered for photovoltaic installation. User Y estimates an installed capacity of 7.8kW, with 3kW already registered. The deviation rate = |7.8-3| / 7.8×100% ≈ 61.5% > 30%, classifying them as high-risk and suspected of unauthorized capacity expansion. User Z estimates an installed capacity of 5.3kW, with 5kW already registered. The deviation rate = |5.3-5| / 5.3×100% ≈ 5.7% ≤ 30%, classifying them as a normal user.
[0101] Finally, for users identified as high-risk for not having applied for photovoltaic (PV) installation, an on-site verification work order is automatically generated. The work order includes basic user information such as account number, address, account name, applied-for capacity, estimated installed capacity, deviation rate, cumulative confidence level, and weighted correlation coefficient record of the nearest identification date. This information is then pushed to the power supply department's verification personnel's terminal as the basis for on-site verification. On-site verification personnel conduct on-site verification based on the work order information to confirm whether there is any failure to apply for PV installation or unauthorized capacity expansion, and provide feedback on the verification results.
[0102] For example, an automatic verification work order is generated for user X, containing: household number, address (Building 12, Unit 301, a certain community), reported installed capacity (0kW), estimated capacity (4.2kW), deviation rate (100%), cumulative confidence level (85.7%), and weighted correlation coefficient records for the past 7 days. The work order is pushed to the mobile terminal of the verification personnel in the relevant area. The verification personnel use the work order to verify the information on-site, confirming that user X has privately installed approximately 4kW of photovoltaic panels on their roof. They take photos as evidence on-site and upload them to the system, completing the verification loop. Simultaneously, the verification feedback results can be used to subsequently optimize the harmonic-capacity relationship model, forming a data loop.
[0103] It should be noted that the above values are for illustrative purposes only and do not constitute a limitation on the present invention.
[0104] In summary, by using the cumulative confidence threshold to identify suspected unregistered photovoltaic users, and extracting measured harmonic content under a set illumination period consistent with the model training conditions, the installed capacity is estimated using a regional harmonic-capacity relationship model. This achieves a complete closed loop from qualitative identification to quantitative estimation, providing power supply departments with accurate verification basis and capacity data support.
[0105] In summary, this invention effectively overcomes the shortcomings of traditional net load curve analysis methods, such as susceptibility to fluctuations in user electricity consumption and difficulty in quantitatively estimating capacity. It improves the accuracy of identifying unregistered photovoltaic users while providing reliable capacity data support.
[0106] Example 2, as Figure 3 As shown, this invention provides a system for identifying and estimating the capacity of unregistered photovoltaic systems based on harmonic characteristics. The system includes: Model building module 11 is used to build a harmonic-capacity relationship model and a normalized harmonic curve library for the target area based on harmonic monitoring data and registered installed capacity of multiple photovoltaic users who have applied for installation within the target area.
[0107] The construction of the harmonic-capacity relationship model for the target region includes: The harmonic content amplitude of each photovoltaic user who has applied for installation is extracted at noon as the characteristic amplitude when the irradiation conditions are similar; Using the registered installed capacity of each photovoltaic user as the independent variable and the characteristic amplitude as the dependent variable, a training sample set is formed by combining the registered installed capacity of all photovoltaic users with the corresponding characteristic amplitude. A machine learning regression model is trained using the training sample set, and the trained model is validated using cross-validation. The model parameters with the smallest root mean square error are selected as the final model parameters. The mapping function between installed capacity and harmonic content amplitude is determined using the final model parameters, thus forming the harmonic-capacity relationship model.
[0108] Construct a normalized harmonic curve library for the target region, including: Calculate the daily average value of harmonic content amplitude for each photovoltaic user registered in the target area on each monitoring day, and divide all monitoring days into several irradiance levels according to the preset quantile interval based on the level of the daily average value. Within each irradiance level, the harmonic content data of all photovoltaic users who have registered for installation under that irradiance level are extracted on the corresponding monitoring day. The harmonic content data is then normalized to obtain a set of normalized harmonic curves under that irradiance level. Unsupervised clustering is performed on the normalized harmonic curve set, and the center curve of each cluster is used as a typical normalized harmonic curve under the irradiance level. All irradiance levels and their corresponding typical normalized harmonic curves are organized into a two-dimensional typical sub-library matrix indexed by irradiance level and curve category, thus forming the normalized harmonic curve library.
[0109] The unsupervised clustering of the normalized harmonic curve set includes: Within a preset range of cluster numbers, the cluster numbers are set to different values in turn, and cluster calculations are performed on the normalized harmonic curve set respectively. For each clustering number value, the silhouette coefficient of the clustering result is calculated, and the clustering number corresponding to the maximum silhouette coefficient is determined as the optimal clustering number. The normalized harmonic curve set is then divided into final clusters using the optimal clustering number. The mean curve of all normalized harmonic curves within each cluster is calculated as the center curve of the cluster, and is determined as a typical normalized harmonic curve under the irradiance level.
[0110] The identification and determination module 12 is used to perform load-side harmonic stripping and normalization processing on the harmonic monitoring data of the user to be identified to obtain normalized harmonic features, calculate the weighted correlation coefficient between the normalized harmonic features and the curves in the normalized harmonic curve library, and evaluate and determine the cumulative confidence level.
[0111] Among them, the harmonic monitoring data of the user to be identified is subjected to load-side harmonic stripping and normalization processing to obtain normalized harmonic characteristics, including: Harmonic content data of the user to be identified during the nighttime hours when there is no photovoltaic power generation is extracted each day, and the average value of the harmonic content data during the nighttime hours of multiple consecutive days is used as the load baseline value of the user to be identified. The net photovoltaic harmonic component sequence is obtained by subtracting the load baseline value from the measured harmonic content at each sampling time of the user to be identified during the day. Read the continuous period in the net photovoltaic harmonic component sequence where the amplitude is continuously higher than the load baseline value, and determine the continuous period as the effective sunshine window for the day; Within the effective illumination window of the day, the maximum value of the net photovoltaic harmonic component sequence is extracted as the normalization reference value. Each data point in the net photovoltaic harmonic component sequence within the effective illumination window of the day is divided by the normalization reference value to complete the normalization process and obtain the time-domain normalized harmonic curve, which serves as the normalized harmonic feature of the user to be identified.
[0112] The normalized harmonic features further include frequency domain features, and the steps for obtaining the frequency domain features include: At each sampling time, the amplitudes of the 5th and 7th harmonics are extracted, and the ratio between the amplitudes of the 5th and 7th harmonics is calculated, as well as the ratios between the amplitudes of the 5th and 7th harmonics and the total harmonic distortion rate, to obtain the proportional characteristics between the harmonic orders at the sampling time. Normalization is performed on the time series based on the proportional characteristics between harmonic orders at all sampling times to obtain a frequency domain normalized sequence; Determine whether the amplitude of the 5th harmonic and the amplitude of the 7th harmonic at each sampling time exceed the product of the load baseline value and the preset gate multiple. If they do, the frequency domain normalized feature at the sampling time is valid. The frequency domain normalized feature is then concatenated with the corresponding values of the time domain normalized harmonic curve to form a two-dimensional feature vector at the sampling time. If the sampling time does not exceed the limit, the frequency domain dimension of the sampling time is set to zero, and the time domain normalized feature of the sampling time is used as the feature vector of the sampling time. The feature vectors of all sampling times are combined in chronological order to form the normalized harmonic features of the user to be identified.
[0113] Calculate the weighted correlation coefficient between the normalized harmonic characteristics and the curves in the normalized harmonic curve library, and evaluate and determine the cumulative confidence level, including: During the period when the net photovoltaic harmonic amplitude of the user to be identified exceeds the load baseline value, the net photovoltaic harmonic amplitude corresponding to each sampling time is used as the calculation weight of the sampling time. The correlation between the normalized harmonic characteristics of the user to be identified and each typical normalized harmonic curve in the normalized harmonic curve library is weighted and calculated. The maximum value of the weighted correlation coefficient is taken as the weighted correlation coefficient of the user to be identified on that day. The weighted correlation coefficient is calculated for the user to be identified for K consecutive days. The number of days in which the daily weighted correlation coefficient exceeds the first preset threshold is counted, and the ratio of the number of days to K is calculated as the cumulative confidence of the user to be identified, where K is a preset constant.
[0114] The capacity estimation module 13 is used to determine that the user to be identified is a suspected unregistered photovoltaic user if the cumulative confidence level is greater than or equal to a preset threshold, extract the measured harmonic content amplitude of the user to be identified under set illumination conditions and input it into the harmonic-capacity relationship model, and output the estimated installed capacity.
[0115] Specifically, the measured harmonic content amplitude of the user to be identified under set illumination conditions is extracted and input into the harmonic-capacity relationship model, and the estimated installed capacity is output, including: The midday period is used as the feature extraction period corresponding to the set illumination conditions. The method for determining the feature extraction period is consistent with the period for extracting the characteristic amplitude of each registered photovoltaic user when constructing the harmonic-capacity relationship model. Extract the measured harmonic content amplitude of the user to be identified during the feature extraction period, input the measured harmonic content amplitude into the mapping function in the harmonic-capacity relationship model, and output the estimated installed capacity of the user to be identified.
[0116] Extract the measured harmonic content amplitude of the user to be identified under set illumination conditions and input it into the harmonic-capacity relationship model to output an estimated installed capacity, including: The midday period is used as the feature extraction period corresponding to the set illumination conditions. The method for determining the feature extraction period is consistent with the period for extracting the characteristic amplitude of each registered photovoltaic user when constructing the harmonic-capacity relationship model. Extract the measured harmonic content amplitude of the user to be identified during the feature extraction period, input the measured harmonic content amplitude into the mapping function in the harmonic-capacity relationship model, and output the estimated installed capacity of the user to be identified.
[0117] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
[0118] This specification and accompanying drawings are merely illustrative examples of the invention and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of the invention. Clearly, those skilled in the art can make various alterations and modifications to the invention without departing from its scope. Therefore, if such modifications and modifications fall within the scope of the invention and its equivalents, the invention is intended to include these modifications and modifications.
Claims
1. A method for identifying and estimating the capacity of unregistered photovoltaic systems based on harmonic characteristics, characterized in that, include: Based on harmonic monitoring data and registered installed capacity of multiple photovoltaic users in the target area, a harmonic-capacity relationship model and a normalized harmonic curve library for the target area are constructed. The load-side harmonic stripping and normalization processing of the harmonic monitoring data of the user to be identified is performed to obtain normalized harmonic features. The weighted correlation coefficient between the normalized harmonic features and the curves in the normalized harmonic curve library is calculated, and the cumulative confidence level is evaluated and determined. If the cumulative confidence level is greater than or equal to a preset threshold, the user to be identified is determined to be a suspected user who has not registered for photovoltaic installation. The measured harmonic content amplitude of the user to be identified under the set illumination conditions is extracted and input into the harmonic-capacity relationship model, and the estimated installed capacity is output.
2. The method for identifying and estimating the capacity of unregistered photovoltaic systems based on harmonic characteristics according to claim 1, characterized in that, Constructing a harmonic-capacity relationship model for the target region includes: The harmonic content amplitude of each photovoltaic user who has applied for installation is extracted at noon as the characteristic amplitude when the irradiation conditions are similar; Using the registered installed capacity of each photovoltaic user as the independent variable and the characteristic amplitude as the dependent variable, a training sample set is formed by combining the registered installed capacity of all photovoltaic users with the corresponding characteristic amplitude. A machine learning regression model is trained using the training sample set, and the trained model is validated using cross-validation. The model parameters with the smallest root mean square error are selected as the final model parameters. The mapping function between installed capacity and harmonic content amplitude is determined using the final model parameters, thus forming the harmonic-capacity relationship model.
3. The method for identifying and estimating the capacity of unregistered photovoltaic systems based on harmonic characteristics according to claim 1, characterized in that, Construct a normalized harmonic curve library for the target region, including: Calculate the daily average value of harmonic content amplitude for each photovoltaic user registered in the target area on each monitoring day, and divide all monitoring days into several irradiance levels according to the preset quantile interval based on the level of the daily average value. Within each irradiance level, the harmonic content data of all photovoltaic users who have registered for installation under that irradiance level are extracted on the corresponding monitoring day. The harmonic content data is then normalized to obtain a set of normalized harmonic curves under that irradiance level. Unsupervised clustering is performed on the normalized harmonic curve set, and the center curve of each cluster is used as a typical normalized harmonic curve under the irradiance level. All irradiance levels and their corresponding typical normalized harmonic curves are organized into a two-dimensional typical sub-library matrix indexed by irradiance level and curve category, thus forming the normalized harmonic curve library.
4. The method for identifying and estimating the capacity of unregistered photovoltaic systems based on harmonic characteristics according to claim 3, characterized in that, Unsupervised clustering is performed on the normalized harmonic curve set, including: Within a preset range of cluster numbers, the cluster numbers are set to different values in turn, and cluster calculations are performed on the normalized harmonic curve set respectively. For each clustering number value, the silhouette coefficient of the clustering result is calculated, and the clustering number corresponding to the maximum silhouette coefficient is determined as the optimal clustering number. The normalized harmonic curve set is then divided into final clusters using the optimal clustering number. The mean curve of all normalized harmonic curves within each cluster is calculated as the center curve of the cluster, and is determined as a typical normalized harmonic curve under the irradiance level.
5. The method for identifying and estimating the capacity of unregistered photovoltaic systems based on harmonic characteristics according to claim 1, characterized in that, The harmonic monitoring data of the user to be identified is subjected to load-side harmonic stripping and normalization to obtain normalized harmonic characteristics, including: Harmonic content data of the user to be identified during the nighttime hours when there is no photovoltaic power generation is extracted each day, and the average value of the harmonic content data during the nighttime hours of multiple consecutive days is used as the load baseline value of the user to be identified. The net photovoltaic harmonic component sequence is obtained by subtracting the load baseline value from the measured harmonic content at each sampling time of the user to be identified during the day. Read the continuous period in the net photovoltaic harmonic component sequence where the amplitude is continuously higher than the load baseline value, and determine the continuous period as the effective sunshine window for the day; Within the effective illumination window of the day, the maximum value of the net photovoltaic harmonic component sequence is extracted as the normalization reference value. Each data point in the net photovoltaic harmonic component sequence within the effective illumination window of the day is divided by the normalization reference value to complete the normalization process and obtain the time-domain normalized harmonic curve, which serves as the normalized harmonic feature of the user to be identified.
6. The method for identifying and estimating the capacity of unregistered photovoltaic systems based on harmonic characteristics according to claim 5, characterized in that, The normalized harmonic features also include frequency domain features, and the steps for obtaining the frequency domain features include: At each sampling time, the amplitudes of the 5th and 7th harmonics are extracted, and the ratio between the amplitudes of the 5th and 7th harmonics is calculated, as well as the ratios between the amplitudes of the 5th and 7th harmonics and the total harmonic distortion rate, to obtain the proportional characteristics between the harmonic orders at the sampling time. Normalization is performed on the time series based on the proportional characteristics between harmonic orders at all sampling times to obtain a frequency domain normalized sequence; Determine whether the amplitude of the 5th harmonic and the amplitude of the 7th harmonic at each sampling time exceed the product of the load baseline value and the preset gate multiple. If they do, the frequency domain normalized feature at the sampling time is valid. The frequency domain normalized feature is then concatenated with the corresponding values of the time domain normalized harmonic curve to form a two-dimensional feature vector at the sampling time. If the sampling time does not exceed the limit, the frequency domain dimension of the sampling time is set to zero, and the time domain normalized feature of the sampling time is used as the feature vector of the sampling time. The feature vectors of all sampling times are combined in chronological order to form the normalized harmonic features of the user to be identified.
7. The method for identifying and estimating the capacity of unregistered photovoltaic systems based on harmonic characteristics according to claim 5, characterized in that, Calculate the weighted correlation coefficient between the normalized harmonic characteristics and the curves in the normalized harmonic curve library, and evaluate and determine the cumulative confidence level, including: During the period when the net photovoltaic harmonic amplitude of the user to be identified exceeds the load baseline value, the net photovoltaic harmonic amplitude corresponding to each sampling time is used as the calculation weight of the sampling time. The correlation between the normalized harmonic characteristics of the user to be identified and each typical normalized harmonic curve in the normalized harmonic curve library is weighted and calculated. The maximum value of the weighted correlation coefficient is taken as the weighted correlation coefficient of the user to be identified on that day. The weighted correlation coefficient is calculated for the user to be identified for K consecutive days. The number of days in which the daily weighted correlation coefficient exceeds the first preset threshold is counted, and the ratio of the number of days to K is calculated as the cumulative confidence of the user to be identified, where K is a preset constant.
8. The method for identifying and estimating the capacity of unregistered photovoltaic systems based on harmonic characteristics according to claim 1, characterized in that, Extract the measured harmonic content amplitude of the user to be identified under set illumination conditions and input it into the harmonic-capacity relationship model to output an estimated installed capacity, including: The midday period is used as the feature extraction period corresponding to the set illumination conditions. The method for determining the feature extraction period is consistent with the period for extracting the characteristic amplitude of each registered photovoltaic user when constructing the harmonic-capacity relationship model. Extract the measured harmonic content amplitude of the user to be identified during the feature extraction period, input the measured harmonic content amplitude into the mapping function in the harmonic-capacity relationship model, and output the estimated installed capacity of the user to be identified.
9. The method for identifying and estimating the capacity of unregistered photovoltaic systems based on harmonic characteristics according to claim 1, characterized in that, After outputting the estimated installed capacity, it also includes: Obtain the installed capacity of the user to be identified, wherein the installed capacity is zero for users who have not completed any installation procedures; Calculate the deviation rate between the estimated installed capacity and the reported installed capacity. If the deviation rate exceeds a preset deviation threshold, the user to be identified is determined as a high-risk user who has not reported photovoltaic installation, and an on-site verification work order is generated.
10. A system for identifying and estimating the capacity of unregistered photovoltaic systems based on harmonic characteristics, characterized in that: The method for identifying and estimating the capacity of unregistered photovoltaic systems based on harmonic characteristics, as described in any one of claims 1 to 9, includes: The model building module is used to build a harmonic-capacity relationship model and a normalized harmonic curve library for the target area based on harmonic monitoring data and registered installed capacity of multiple photovoltaic users who have registered within the target area. The identification and determination module is used to perform load-side harmonic stripping and normalization processing on the harmonic monitoring data of the user to be identified to obtain normalized harmonic features, calculate the weighted correlation coefficient between the normalized harmonic features and the curves in the normalized harmonic curve library, and evaluate and determine the cumulative confidence level. The capacity estimation module is used to determine that the user to be identified is a suspected unregistered photovoltaic user if the cumulative confidence level is greater than or equal to a preset threshold, extract the measured harmonic content amplitude of the user to be identified under set illumination conditions and input it into the harmonic-capacity relationship model, and output the estimated installed capacity.