A distributed photovoltaic power generation anomaly positioning optimization method

By comparing data slicing with dual-domain clustering and modeling pollution diffusion, the verification errors caused by meteorological fluctuations and component differences in distributed photovoltaic power generation systems were resolved, enabling accurate location and cause classification of power generation losses, and improving verification efficiency and accuracy.

CN120892854BActive Publication Date: 2026-03-24INFORMATION CENT OF YUNNAN POWER GRID CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Existing verification methods for distributed photovoltaic power generation systems cannot effectively eliminate errors caused by weather fluctuations and component performance heterogeneity, leading to waste of operation and maintenance resources, deviations in power generation revenue, and settlement disputes, especially in mountainous areas or contiguous rooftop scenarios where false alarms and omissions occur frequently.

Method used

By comparing data slices with dual-domain clustering, combined with spatiotemporal hotspot identification, pollution diffusion modeling, and multi-scale noise removal, the system achieves accurate location and diagnosis of power generation losses. It uses K-means and DBSCAN clustering to generate theoretical and measured clusters, constructs a cluster difference matrix, identifies and eliminates noise, reconstructs abnormal transmission paths, and diagnoses line anomalies and equipment performance anomalies.

Benefits of technology

It achieves macroscopic error shielding and microscopic efficiency difference capture under different design parameters and climatic conditions, quickly locates and classifies the causes of anomalies, improves the accuracy of verification and the precision of priority determination of operation and maintenance work orders, and reduces false alarm rate and false alarm rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120892854B_ABST
    Figure CN120892854B_ABST
Patent Text Reader

Abstract

The application discloses a power generation abnormality positioning optimization method for distributed photovoltaics, and particularly relates to the technical field of power generation abnormality positioning. The method is characterized in that the static information, environmental parameters and dynamic power generation data of multiple stations in a target area are spatio-temporally aligned, the spatial density is clustered based on the three-dimensional distance of the stations, and the spatial sub-clusters are divided. Then, the decoupling model of historical loss rate and irradiance is combined to identify the spatio-temporal hotspot area. For the hotspot area, a spatial loss gradient field is constructed, and the pollution propagation path is reconstructed by reverse tracing to the normal threshold boundary point. For the non-hotspot area, the multi-scale attenuation and noise characteristics are extracted by wavelet packet decomposition, variational mode decomposition and short-time Fourier transform, and the artifacts are removed by neighborhood slope difference, so that the macroscopic error shielding and microscopic abnormality positioning of meteorological and component differences are realized, and the checking accuracy and operation and maintenance efficiency are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power generation anomaly location technology, and more specifically, to an optimization method for power generation anomaly location in distributed photovoltaic systems. Background Technology

[0002] Distributed photovoltaic (PV) power generation systems have been widely applied in various scenarios, including urban and rural rooftops, industrial and commercial plantations, and mountainous farmlands. Due to variations in terrain, sunlight intensity, and module nominal power, the output power of different sites is highly coupled with environmental parameters. Grid operators and asset managers need to accurately verify the power generation of numerous sites within the target area to meet the needs of revenue settlement, energy efficiency assessment, and operation and maintenance decisions.

[0003] However, existing verification methods mostly rely on single-site comparisons or empirical thresholds, which cannot simultaneously eliminate errors caused by large-scale weather fluctuations and heterogeneity in component performance. In mountainous areas or on contiguous rooftops, sudden shading, post-rain dust accumulation, or differences in inverter efficiency can easily trigger false alarms, while slight cluster-level power generation deficits are difficult to detect, resulting in wasted operation and maintenance resources, discrepancies in power generation revenue, and settlement disputes. This technical bottleneck severely restricts the refined management and efficient maintenance of distributed photovoltaic assets. Summary of the Invention

[0004] To overcome the aforementioned deficiencies in existing technologies, this invention provides an optimization method for locating power generation anomalies in distributed photovoltaic systems. By comparing data slicing and dual-domain clustering, combined with spatiotemporal hotspot identification, pollution diffusion modeling, and multi-scale noise removal, the method achieves accurate location and diagnosis of power generation losses, thereby solving the problems of false alarms and missed alarms caused by meteorological fluctuations and component differences.

[0005] To achieve the above objectives, the present invention provides the following technical solution: a method for optimizing the location of power generation anomalies in distributed photovoltaic systems, characterized in that it includes:

[0006] Step 1: Asset Identification and Data Collection: Collect static information of all distributed photovoltaic sites within the target verification spatial area and label the site index; collect environmental parameters within the target verification time area and dynamic power generation information of all distributed photovoltaic sites, including inverter power generation curves and module temperatures; upload the spatiotemporally aligned static information, environmental parameters, and dynamic power generation information to the data analysis center;

[0007] Step 2: Constructing a station-level theoretical benchmark: Obtain the theoretical performance ratio and the theoretical expected power of each distributed photovoltaic site. Using environmental parameters as input, output the ideal power generation efficiency under the current environmental parameters at the same time resolution. Call up static information related to power generation efficiency and use environmental parameters as boundary conditions to obtain the theoretical expected power.

[0008] Step 3: Perform dual-domain clustering comparison: Perform K-means clustering on the theoretical expected power set to generate theoretical clusters CT; the ratio of actual power generation to expected power generation is the measured performance ratio, and perform DBSCAN clustering on the measured performance ratio set to generate measured clusters CP; output an abnormal candidate cluster list based on the cluster difference matrix;

[0009] Step 4: Photovoltaic Anomaly Location: Summarize the time-series data of the anomaly candidate cluster list, extract the spatiotemporal characteristics of the anomaly candidate clusters, and based on the spatiotemporal distribution characteristics, reveal the anomaly transmission path through pollution diffusion vector modeling, identify the anomaly as line anomaly, pollution deposition or component aging, and trigger the corresponding cleaning or replacement maintenance work order.

[0010] Preferably, the method for outputting the list of anomalous candidate clusters based on the cluster difference matrix is ​​as follows:

[0011] Record dual-domain cluster labels: acquire theoretical cluster labels and measured cluster labels for each distributed photovoltaic site to form dual-domain cluster label data;

[0012] Construct a cluster difference statistics table: Using theoretical cluster labels as row indexes and measured cluster labels as column indexes, count the number of sites at the intersection of rows and columns to generate a cluster difference statistics table, thereby quantifying the correspondence between two-domain clustering;

[0013] Determine the main performance cluster: For each theoretical cluster in the cluster difference statistics table, select the measured cluster with the largest number of cross-sites as the main performance cluster and establish an intra-cluster reference performance benchmark;

[0014] Identify off-site sites: Compare the measured cluster labels of each site within the theoretical cluster with the main representation cluster labels, identify sites with inconsistent labels as off-site sites, and generate a set of off-site sites;

[0015] Generate abnormal candidate clusters: Calculate the cluster deviation rate. When the cluster deviation rate is not lower than the first preset threshold, write the corresponding theoretical cluster into the abnormal candidate cluster list and record the cluster identifier, cluster capacity, deviation site index and cluster deviation rate to complete the abnormal candidate cluster registration.

[0016] Sorted Output List: Sort the list of abnormal candidate clusters from high to low according to the cluster deviation rate, and output the sorted list of abnormal candidate clusters to improve the accuracy of subsequent diagnostic priority determination.

[0017] Preferably, the first preset threshold is a dynamic value that changes with the environment, and it is obtained as follows:

[0018] Within the most recent W assessment cycles, the median of the deviation rate and the median of the absolute deviation of each theoretical cluster are calculated and added together to obtain the steady-state benchmark representing the background fluctuation of the equipment.

[0019] Using the irradiance and component temperature within the same time window, the ratio of the standard deviation to the mean is calculated to obtain the irradiance fluctuation coefficient and the temperature rise fluctuation coefficient. The two are then weighted and summed to form the environmental fluctuation index.

[0020] The environmental fluctuation index is mapped to the 0–1 interval using a linear normalization method to obtain a dimensionless environmental scale.

[0021] The steady-state benchmark is processed by multiplication and scaling to generate a dynamic threshold, which is then clipped using preset upper and lower limits to obtain a first preset threshold that adapts to weather fluctuations.

[0022] Preferably, step four involves parallel execution of line anomaly diagnosis and equipment performance anomaly location. The line anomaly diagnosis is based on electrical topology reconstruction and line loss mutation monitoring, injecting characteristic harmonic currents into alarm sites, inverting the equivalent impedance value of the line, and diagnosing the fault type and location coordinates.

[0023] The equipment performance anomaly location is achieved by generating anomaly spatial subclusters through spatial density clustering; constructing a loss rate-irradiation decoupling model to screen spatiotemporal hotspot areas; tracing back to the pollution boundary through pollution diffusion vector modeling; performing aging feature temporal decomposition on unmarked line anomaly sites; dynamically quantifying coupling effects and dynamically adjusting the aging tolerance threshold by integrating pollution acceleration factors.

[0024] Preferably, the line anomaly diagnosis process includes the following steps:

[0025] The regional electrical connection topology map is reconstructed based on the geographical coordinates of photovoltaic sites, and the hierarchical relationship of each site in the distribution network is marked. The daily line loss rate change trend is calculated by continuously monitoring the difference between the inverter output power and the metered power at the grid connection point. When the line loss change amplitude exceeds the threshold for several consecutive days, the line anomaly marker is triggered.

[0026] Harmonic signals are injected into alarm sites, and the equivalent impedance of the line is retrieved based on the voltage response signal.

[0027] The fault type is diagnosed based on the degree of deviation of the impedance value from the standard value: a significant increase in impedance indicates poor contact or oxidation of the connector; a significant decrease in impedance indicates insulation damage or partial short circuit; a local abnormal increase in impedance indicates capacitive load fault; and abnormal output line location coordinates and fault type labels are used.

[0028] Preferably, the device performance anomaly localization includes the following steps:

[0029] Step S11: Spatial density clustering segmentation: Calculate the three-dimensional geographic distance matrix based on the latitude, longitude and altitude information of the stations in the abnormal candidate cluster list, set the median of the nearest neighbor distance of all stations as the spatial truncation threshold, divide the spatial sub-clusters according to the spatial clustering intensity and determine the geographic boundary range.

[0030] Step S12: Spatiotemporal hotspot region identification: Calculate the distribution density of abnormal sites using spatial subclusters, construct a loss rate-irradiance decoupling model by combining historical loss rate sequences and environmental irradiance sequences, and screen out spatiotemporal hotspot regions that simultaneously satisfy high distribution density and strong decoupling characteristics;

[0031] Step S13: Pollution diffusion vector modeling: Construct the spatial loss rate gradient field of the spatiotemporal hotspot area, establish a distance inverse weight matrix to calculate the spatial change intensity and direction vector of the loss rate, aggregate to generate the main propagation direction vector and trace back to the normal threshold boundary point where the loss rate is less than the threshold.

[0032] Step S14: For the monthly and weekly average power generation efficiency sequences of non-spatiotemporal hotspot areas, wavelet packet decomposition is used to extract multi-scale modes; variational mode decomposition is applied to separate the long-term decay curve on the low-frequency trend mode, and the decay slope is calculated; short-time Fourier transform is applied to extract transient fluctuation features on the high-frequency noise mode; for each non-spatiotemporal hotspot area, the trend slope is differiated from the slope of neighboring stations in the same cluster and normalized. If the difference exceeds the product of the local residual standard deviation and the confidence factor, it is marked as a structural anomaly; otherwise, it is removed as random noise; calculate and output the noise confidence level P after artifact removal.

[0033] Step S15: Multi-source anomaly collaborative determination: Determine the cause of anomalies in spatiotemporal hotspot areas and non-spatiotemporal hotspot areas respectively; output the anomaly category and classification confidence level; output a maintenance work order integrating geographic coordinates and anomaly category labels.

[0034] Preferably, based on the initial risk score of the anomaly category (pollution 0.5, aging 0.6, line 0.7), the classification confidence and the maintenance cost depreciation factor, the work order priority score is calculated; maintenance work orders are generated in descending order of priority score, and the work order execution results are recorded in real time for subsequent online model updates.

[0035] Preferably, the noise artifact removal process includes the following steps:

[0036] The system monitors the rate of change in power generation efficiency in real time. When the rate of change increases over time and exceeds the judgment threshold, it outputs an abnormal fluctuation indicator signal. The judgment threshold for the rate of change in power generation efficiency is set based on the long-term decay slope characteristics.

[0037] The difference between the timing reference phase of environmental parameters and the phase of power response is calculated as the mismatch. When the absolute value of the mismatch is greater than the judgment threshold, a climate interference identification signal is output. The judgment threshold of the mismatch is set based on the characteristics of the climate reference phase mismatch.

[0038] When both an abnormal fluctuation indicator signal and a climate interference indicator signal are received simultaneously, a climate noise filtering instruction is generated. Abnormal data segments that trigger both types of signals are marked as climate interference artifacts, removed from the performance degradation curve, and reconstructed by interpolation. Furthermore, fault root cause classification is performed based on the sign of the mismatch.

[0039] When efficiency continues to decrease but the phase mismatch exceeds the secondary threshold for a long period of time, it is judged as a progressive aging fault; when the mismatch is positive, a heat dissipation system failure alarm is triggered, and when the mismatch is negative, a surface contamination alarm is triggered.

[0040] Preferably, the spatial truncation threshold is obtained as follows:

[0041] The target verification area is divided into L×L grids of equal area;

[0042] Calculate the number of stations D in each grid and the difference in number between each grid and its adjacent grids;

[0043] The spatial truncation threshold is based on the median nearest neighbor distance of the entire network. When the difference in the number of neighbors exceeds the first preset difference, the spatial truncation threshold is reduced proportionally by η1.

[0044] When the quantity difference is lower than the second preset difference, the space truncation threshold is expanded proportionally by η2.

[0045] The obtained spatial truncation threshold is used for spatial sub-cluster partitioning of the three-dimensional distance matrix.

[0046] Preferably, the method further includes a data slicing management step, which involves extracting steady-state segments using irradiance-temperature dual thresholds; dividing the steady-state segments into boxes according to preset irradiance intensity closed intervals and then calculating a normalized power sequence based on the average irradiance value within each box; and storing the non-steady-state segments independently for subsequent fault tracing based on the law of constant irradiance due to component temperature drop and the mandatory retention event of inverter power collapse symptom markers.

[0047] Preferably, based on the stability of environmental parameters, data slicing is performed, and the data to be analyzed is filtered and output, including the following sub-steps:

[0048] Triggering data slicing: When the data cleaning module outputs the environmental parameter sequence and the power generation information sequence, the intensity of irradiance change and the intensity of component temperature change are calculated to generate a set of continuous time segments that meet the meteorological stability constraints, which are denoted as steady-state segments. The minimum segment length and maximum jump rate of the steady-state segments are limited. The meteorological stability constraints are dynamically adjusted based on the characteristics of atmospheric transmittance abrupt change and the difference in thermal inertia of photovoltaic modules.

[0049] Extract key events: Identify equipment operation state transition events in historical operation based on irradiance-temperature response decoupling coefficient, and calibrate restart symptom periods by associating them with inverter power collapse characteristics, and summarize and output a dataset that is forced to be retained;

[0050] Reconstruct equivalent environmental data: Apply steady-state fragments to steps two through four to reduce the data size of subsequent processing and force the retention of the dataset as training samples for identifying anomaly types.

[0051] The technical effects and advantages of this invention are as follows:

[0052] This invention achieves spatiotemporal alignment of static information, environmental parameters, and dynamic power generation information of all distributed photovoltaic sites within the target verification area. Based on the theoretical performance ratio, it generates theoretical expected power and forms theoretical clusters (CT) through K-means clustering and measured clusters (CP) through DBSCAN clustering. The cluster difference matrix is ​​then used for cross-comparison. This invention realizes macroscopic error shielding and microscopic efficiency difference capture under different design parameters and climatic conditions, effectively solving the problem of false alarms or omissions in traditional site-by-site comparisons that are easily affected by meteorological fluctuations and component differences.

[0053] This invention constructs spatial subclusters using a three-dimensional geographic distance matrix and identifies spatiotemporal hotspots. It combines a loss rate-irradiance decoupling model and pollution diffusion vector modeling to reconstruct anomaly propagation paths. For non-hotspot areas, it employs wavelet packet decomposition, VMD, and STFT multi-scale noise removal, as well as a random forest classification model to accurately determine pollution-type and aging-type anomalies. This achieves a closed-loop diagnostic process from rapid cluster-level localization to cause classification, effectively solving the problems of low verification efficiency, insufficient anomaly localization accuracy, and inaccurate priority determination of subsequent maintenance work orders. Attached Figure Description

[0054] Figure 1 This is a simplified flowchart of the distributed photovoltaic power generation anomaly location optimization method of the present invention.

[0055] Figure 2 This is a simplified flowchart of the equipment performance anomaly localization process of the present invention. Detailed Implementation

[0056] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0057] At the same time, it should be understood that, for ease of description, the dimensions of the various parts shown in the accompanying drawings are not drawn according to actual scale.

[0058] The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit the scope of this application and its application or use.

[0059] Techniques, methods, and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and equipment should be considered part of the specification.

[0060] Example 1, see Figure 1 A simplified flowchart of the distributed photovoltaic power generation anomaly location optimization method is provided in this invention. Figure 1 The method for optimizing the location of power generation anomalies in distributed photovoltaic systems, as shown, includes:

[0061] Step 1: Asset Identification and Data Acquisition: Collect static information of all distributed photovoltaic sites within the target verification spatial area, including spatial coordinates, installed capacity, module / inverter model, and metering device accuracy, and label the site index; collect environmental parameters (at least including irradiance and temperature) and dynamic power generation information of all distributed photovoltaic sites, including inverter power generation curves and module temperature, within the target verification time area; after performing missing packet interpolation, clock alignment, and 3σ outlier shielding preprocessing, map the dynamic power generation information to the corresponding site index based on coordinates, and upload the spatiotemporally aligned static information, environmental parameters, and dynamic power generation information to the data analysis center;

[0062] Step 2: Constructing a station-level theoretical benchmark: Obtain the theoretical performance ratio and the theoretical expected power of each distributed photovoltaic site. Using environmental parameters (illuminance and module temperature) as input, output the ideal power generation efficiency under the current environmental parameters at the same time resolution. Call up static information related to power generation efficiency, such as tilt angle, azimuth angle, nominal module power, and inverter efficiency, and use the environmental parameters as boundary conditions to obtain the theoretical expected power.

[0063] In this invention, ideal power generation efficiency refers to the maximum energy conversion capacity that a photovoltaic module can achieve per unit of incident irradiance under ideal operating conditions of no faults, no pollution, and no shading, given the real-time values ​​of current environmental parameters—planar incident irradiance (G_POA, unit W / m²), module surface temperature (T_mod, unit °C), and ambient wind speed (v_wind, unit m / s). It is expressed as a percentage. The theoretical power generation is calculated by combining the ideal power generation efficiency with the installed capacity of the distributed photovoltaic site, the nominal power of the module, the module installation tilt angle, azimuth angle, inverter efficiency, shading correction factor, and other system structural parameters, using a normalized model to estimate the standard power generation under the above environmental conditions. It is expressed in kilowatt-hours. Specifically, firstly, the plane incident irradiance, module surface temperature, and ambient wind speed are used as input variables for the mapping model. Simultaneously, their respective effective value ranges (plane incident irradiance 200–1200 W / m², module surface temperature –10–70℃, ambient wind speed 0–10 m / s) are defined as boundary conditions in the model for dynamically estimating the ideal power generation efficiency. Subsequently, a directional correction coefficient is determined based on the module installation tilt angle, azimuth angle, and on-site shading conditions. This coefficient is set to 1 when there is no shading; if shading needs to be considered, then… The correction value is less than 1 obtained from actual measurement or model calculation. Finally, the estimated ideal power generation efficiency is multiplied by the site installed capacity, the directional correction factor and the inverter efficiency in sequence, and the product is multiplied by the ratio of the current plane incident irradiance to the standard irradiance (1000W / m²) to obtain the theoretical power generation at the current moment. The above calculation assumes that the system is fault-free, pollution-free, unobstructed and other external interferences, and uses the ideal optimal state as the boundary condition to provide a unified and traceable benchmark reference for subsequent actual power generation deviation judgment and anomaly identification.

[0064] Step 3: Perform dual-domain clustering comparison: Perform K-means clustering on the theoretical expected power set to generate theoretical clusters CT; the ratio of actual power generation to expected power generation is the measured performance ratio, and perform DBSCAN clustering on the measured performance ratio set to generate measured clusters CP; output an abnormal candidate cluster list based on the cluster difference matrix;

[0065] The explanation is as follows: In the theoretical domain, K-means clustering is used to screen distributed photovoltaic sites with similar design parameters and climate to eliminate the influence of macro-meteorological differences. The K value is adaptively determined based on the profile coefficient or elbow rule. In the measured domain, DBSCAN clustering is used to normalize the daily performance ratio, capturing the micro-differences in actual conversion efficiency under the same climate. The parameters of DBSCAN clustering are set based on the standard deviation of the daily performance ratio of adjacent sites to ensure the consistency of density within the cluster. The advantage is that the results of the two domains are cross-compared, and only sites with inconsistencies between the two domains are retained, making the power generation verification immune to errors and directly locking the real power generation loss. This is used for the rapid location of anomalies in the spatiotemporal area of ​​the target verification.

[0066] Step 4: Photovoltaic Anomaly Location: Summarize the time-series data of the anomaly candidate cluster list, extract the spatiotemporal characteristics of the anomaly candidate clusters, and based on the spatiotemporal distribution characteristics, reveal the anomaly transmission path through pollution diffusion vector modeling, identify the anomaly as line anomaly, pollution deposition or component aging, and trigger the corresponding cleaning or replacement maintenance work order.

[0067] Furthermore, the method for outputting the list of anomalous candidate clusters based on the cluster difference matrix is ​​as follows:

[0068] Record dual-domain cluster labels: acquire theoretical cluster labels and measured cluster labels for each distributed photovoltaic site to form dual-domain cluster label data;

[0069] Construct a cluster difference statistics table: Using theoretical cluster labels as row indexes and measured cluster labels as column indexes, count the number of sites at the intersection of rows and columns to generate a cluster difference statistics table, thereby quantifying the correspondence between two-domain clustering;

[0070] Determine the main performance cluster: For each theoretical cluster in the cluster difference statistics table, select the measured cluster with the largest number of cross-sites as the main performance cluster and establish an intra-cluster reference performance benchmark;

[0071] Identify off-site sites: Compare the measured cluster labels of each site within the theoretical cluster with the main representation cluster labels, identify sites with inconsistent labels as off-site sites, and generate a set of off-site sites;

[0072] Generate abnormal candidate clusters: Calculate the cluster deviation rate (i.e., the cluster deviation rate of the set of deviating sites relative to the total number of sites in the theoretical cluster). When the cluster deviation rate is not lower than the first preset threshold, write the corresponding theoretical cluster into the abnormal candidate cluster list and record the cluster identifier, cluster capacity, deviating site index and cluster deviation rate to complete the registration of abnormal candidate clusters.

[0073] Sorted Output List: Sort the list of abnormal candidate clusters from high to low according to the cluster deviation rate, and output the sorted list of abnormal candidate clusters to improve the accuracy of subsequent diagnostic priority determination.

[0074] Furthermore, when the cluster deviation rate is lower than the first preset threshold, it indicates that the measured performance of the vast majority of stations within the theoretical cluster is consistent with that of the main cluster, the cluster-level performance is within the random fluctuation range allowed by the model, and there is no systemic power generation gap; the appropriate measures at this time are:

[0075] Cluster-level status marking: Mark the theoretical cluster as a normal cluster and do not include it in the list of abnormal candidate clusters;

[0076] Routine monitoring continues: Maintain the original daily monitoring frequency, continue to collect and update the measured daily performance ratio for rolling cluster deviation rate calculation;

[0077] Single-point-of-care threshold: If an individual site within the cluster still triggers an inverter-level or metering-level alarm due to a sudden failure, it will be handled according to the single-site strategy and will not trigger a cluster-level work order.

[0078] Threshold adaptive update (optional): If the cluster deviation rate is consistently much lower than the threshold for multiple consecutive assessment cycles, the cluster-level threshold can be tightened or the model parameters optimized using statistical methods to improve the sensitivity of subsequent screening.

[0079] The explanation is as follows: First, two labels are recorded for each distributed photovoltaic site: one is a theoretical cluster label, indicating which sites the site is most similar to in the design-climate dimension; the other is a measured cluster label, indicating which sites the site is most similar to in the actual daily performance ratio dimension. Using all theoretical clusters as rows and all measured clusters as columns, the total number of sites at the intersection of rows and columns is counted to obtain a cluster difference statistics table. For each row (i.e., the same theoretical cluster), the column with the most occurrences is found, and the corresponding measured cluster is considered the main performance cluster of that theoretical cluster. Sites belonging to the same theoretical cluster but falling into other columns (i.e., other measured clusters) are collectively referred to as deviation sites. The proportion of deviation sites in a row to the total number of sites in that row is calculated. If the proportion is not lower than a preset threshold, the theoretical cluster corresponding to this row is recorded as an abnormal candidate cluster, and its entire deviation site index, cluster capacity, and deviation proportion are written into the abnormal candidate cluster list. Finally, this list is sorted from high to low according to the deviation proportion for subsequent detailed diagnosis.

[0080] Furthermore, the first preset threshold is a dynamic value that changes with the environment, and it is obtained as follows:

[0081] Within the most recent W assessment cycles (W is the most recent 30 daily assessment cycles, which can be adjusted by the system administrator, but not less than 7 daily assessment cycles), first calculate the median of the deviation rate of each theoretical cluster and the median of the absolute deviation, and add them together to obtain the steady-state benchmark representing the background fluctuation of the equipment (under the steady-state fluctuation of equipment operation);

[0082] Using the irradiance and component temperature within the same time window, the ratio of the standard deviation to the mean is calculated to obtain the irradiance fluctuation coefficient and the temperature rise fluctuation coefficient. The two are then weighted and summed to form the environmental fluctuation index.

[0083] The environmental fluctuation index is mapped to the 0–1 interval using a linear normalization method to obtain a dimensionless environmental scale.

[0084] The steady-state benchmark is processed by multiplication scaling, that is, multiplied by [1 + environmental scale × environmental sensitivity coefficient], to generate a dynamic threshold. Then, it is clipped with preset upper and lower limits to obtain a first preset threshold that adapts to meteorological fluctuations.

[0085] The environmental sensitivity coefficient is a dimensionless quantity used to quantify the impact of fluctuations in environmental parameters (such as irradiance and module temperature) within a target area on changes in the daily performance ratio or power generation efficiency of a photovoltaic site. The larger the coefficient, the more significant the disturbance to the equipment output caused by meteorological changes; conversely, the smaller the coefficient, the more stable the equipment performance and the weaker the response to environmental fluctuations.

[0086] The environmental sensitivity coefficient is obtained by extracting the daily performance ratio sequence corresponding to the environmental fluctuation index from historical operational data, calculating the time series correlation coefficient or regression slope of the two, and obtaining a set of candidate sensitivity indicators; by using the sliding window cross-validation method, the prediction accuracy of different candidate indicators under various meteorological conditions is evaluated, and the indicator value that best balances the false alarm rate and the missed alarm rate is selected as the final environmental sensitivity coefficient.

[0087] The environmental sensitivity coefficient is set based on the following: The initial value of the environmental sensitivity coefficient is recommended to be determined based on the historical data statistical characteristics of typical scenarios (mountainous areas, plains, urban areas). Its range is usually set between 0.2 and 0.8 to balance the dynamic response to meteorological fluctuations and detection stability. After actual deployment, the coefficient can be continuously fine-tuned in combination with the performance feedback of the online model to ensure that efficient fault differentiation capability can be maintained under different seasons and climatic conditions.

[0088] Furthermore, step four involves parallel execution of line anomaly diagnosis and equipment performance anomaly location. The line anomaly diagnosis is based on electrical topology reconstruction and line loss mutation monitoring, injecting characteristic harmonic currents into alarm sites, inverting the equivalent impedance value of the line, and diagnosing the fault type and location coordinates. The equipment performance anomaly location generates anomaly spatial subclusters through spatial density clustering; constructs a loss rate-irradiance decoupling model to screen spatiotemporal hotspot areas; uses pollution diffusion vector modeling to trace back to the pollution boundary; performs aging feature temporal decomposition on sites without marked line anomalies; dynamically quantifies coupling effects, integrates pollution acceleration factors to dynamically adjust aging tolerance thresholds; and prioritizes anomaly types, prioritizing line anomalies by outputting maintenance work orders, otherwise determining pollution / aging anomalies based on the main propagation direction vector and performance decay slope. The line anomaly diagnosis results are input into the equipment performance anomaly location process in real time: if a site is marked as a line anomaly, pollution diffusion modeling and aging feature temporal decomposition are skipped.

[0089] Furthermore, the line anomaly diagnosis process includes the following steps:

[0090] The regional electrical connection topology map is reconstructed based on the geographical coordinates of photovoltaic sites, and the hierarchical relationship of each site in the distribution network is marked. The daily line loss rate change trend is calculated by continuously monitoring the difference between the inverter output power and the metered power at the grid connection point. When the line loss change exceeds the threshold for several consecutive days, a line anomaly marker is triggered. (For example, the line anomaly marker is triggered when the line loss rate change exceeds the historical average ±2σ for three consecutive days.)

[0091] Harmonic signals are injected into the alarm site (when injecting harmonic signals, the main frequency range of power frequency harmonics and the high frequency switching noise range should be avoided to eliminate the interference of background noise on signal integrity; select the phase-sensitive frequency band of the power grid line impedance so that the voltage response signal has high resolution for small impedance changes; limit the injected power to not exceed the equipment tolerance threshold to ensure system stability and standard compliance, and dynamically adjust according to local power grid access regulations), and the equivalent impedance value of the line is inverted based on the voltage response signal;

[0092] The fault type is diagnosed based on the degree of deviation of the impedance value from the standard value: a significant increase in impedance indicates poor contact or oxidation of the connector; a significant decrease in impedance indicates insulation damage or partial short circuit; a local abnormal increase in impedance indicates capacitive load fault; and abnormal output line location coordinates and fault type labels are used.

[0093] For further details, please refer to [link / reference]. Figure 2 A simplified flowchart for locating equipment performance anomalies, which includes the following steps:

[0094] Step S11: Spatial density clustering segmentation: Calculate the three-dimensional geographic distance matrix based on the latitude, longitude and altitude information of the stations in the abnormal candidate cluster list, set the median of the nearest neighbor distance of all stations as the spatial truncation threshold, divide the spatial sub-clusters according to the spatial clustering intensity and determine the geographic boundary range.

[0095] Step S12: Spatiotemporal hotspot region identification: Calculate the distribution density of abnormal sites using spatial subclusters, construct a loss rate-irradiance decoupling model by combining historical loss rate sequences and environmental irradiance sequences, and screen out spatiotemporal hotspot regions that simultaneously satisfy high distribution density and strong decoupling characteristics;

[0096] The explanation is that the loss rate-irradiance decoupling model is used to quantify the independence of photovoltaic power generation losses from meteorological fluctuations. By analyzing the conditional correlation between historical loss rate sequences and irradiance sequences, the proportion of losses that can be explained by meteorological factors is calculated. When the decoupling value output by the loss rate-irradiance decoupling model exceeds >0.85, it indicates that more than 85% of the loss fluctuations cannot be explained by irradiance changes, that is, there are abnormal losses that are unrelated to meteorology.

[0097] Step S13: Pollution diffusion vector modeling: Construct the spatial loss rate gradient field of the spatiotemporal hotspot area, establish a distance inverse weight matrix to calculate the spatial change intensity and direction vector of the loss rate, aggregate to generate the main propagation direction vector and trace back to the normal threshold boundary point where the loss rate is less than the threshold.

[0098] The explanation is as follows: the spatial loss rate gradient field refers to mapping the loss rate values ​​of each photovoltaic site in the region onto the geographic space, forming a loss rate distribution map similar to terrain undulations; the loss rate of discrete sites is transformed into a continuous gradient field (similar to a meteorological contour map), revealing the direction of loss rate diffusion; the main propagation direction vector refers to the dominant path of loss rate diffusion throughout the region; in one possible embodiment, a normal threshold boundary point with a loss rate ≤ 5% is set. The reason for setting it to 5% is that, in actual situations, if the inherent losses of photovoltaic modules (2-3%), inverter conversion losses (1-2%), and normal line losses (<1%) exceed 5%, there are unavoidable unnatural losses. Therefore, the location with a loss rate ≤ 5% is marked as the normal threshold boundary point.

[0099] Step S14: For the monthly and weekly average power generation efficiency sequences of non-spatiotemporal hotspot areas, wavelet packet decomposition (3 layers) is used to extract multi-scale modes; variational mode decomposition is applied to separate the long-term decay curve on the low-frequency trend mode, and the decay slope is calculated; short-time Fourier transform is applied to extract transient fluctuation features on the high-frequency noise mode; for each non-spatiotemporal hotspot area, the trend slope is differed from the slope of neighboring stations in the same cluster and normalized. If the difference exceeds the product of the local residual standard deviation and the confidence factor, it is marked as a structural anomaly; otherwise, it is removed as random noise; calculate and output the noise confidence level P after artifact removal.

[0100] Step S15: Multi-source anomaly collaborative determination: Determine the causes of anomalies in spatiotemporal hotspot areas and non-spatiotemporal hotspot areas respectively;

[0101] The method for determining the cause of anomalies in spatiotemporal hotspot areas is as follows: Pollution diffusion vector modeling is performed. If the vector magnitude in the main propagation direction is significantly large (e.g., main propagation direction vector magnitude > 0.5) and the attenuation slope within the region changes slowly (e.g., attenuation slope within the region ≤ 0.001), it is determined to be a pollution diffusion type anomaly. Conversely, if the vector magnitude is low (e.g., main propagation direction vector magnitude ≤ 0.1) and the attenuation slope is significantly increased (e.g., attenuation slope > 0.1), it indicates that the equipment performance shows a monotonically decreasing trend and there is no obvious spatial diffusion effect, thus it is determined to be an aging type anomaly.

[0102] The method for determining the cause of anomalies in non-spatiotemporal hotspot areas is as follows: construct a five-dimensional feature vector by taking the main propagation direction vector magnitude, performance attenuation slope, noise confidence P after artifact removal, environmental sensitivity coefficient, and attenuation slope standard deviation index; input the trained random forest classification model, output the anomaly category and classification confidence; and output a maintenance work order integrating geographic coordinates and anomaly category labels.

[0103] Furthermore, in the method for determining the cause of anomalies in spatiotemporal hotspot areas, to further identify pollution deposition anomalies, this embodiment of the invention combines the pollution diffusion vector and the characteristics of equipment surface temperature changes to construct a pollution deposition assessment step: First, within the spatiotemporal hotspot area, the cumulative increment of the loss rate gradient along the line is calculated based on the main propagation direction of the pollution diffusion vector; simultaneously, short-term abrupt changes in the difference between component surface temperature and ambient temperature are monitored; when the increment of the loss rate gradient in the same direction continuously exceeds the static threshold and is accompanied by a constant increase in surface temperature difference, the system determines the area as a pollution deposition anomaly and generates a visual map of pollution deposition depth and propagation path, supporting targeted cleaning and maintenance.

[0104] In one possible implementation, the work order priority score is calculated based on the initial risk score of the anomaly category (pollution 0.5, aging 0.6, line 0.7), the classification confidence score, and the maintenance cost depreciation factor; maintenance work orders are generated in descending order of priority score, and the work order execution results are recorded in real time for subsequent online model updates.

[0105] In one possible embodiment, the noise artifact removal process includes the following steps:

[0106] Background explanation: The monotonicity of equipment performance degradation refers to the unidirectional decrease in equipment performance (photovoltaic power generation efficiency) over time. Ideally, the degradation of photovoltaic power generation efficiency should always remain monotonically decreasing over time. If a rebound in photovoltaic power generation efficiency is detected (such as abnormally higher power generation at a certain time than before), it is considered a violation of monotonicity, which may be caused by data acquisition errors, temporary environmental interference (such as dust falling off), or equipment failure, and an alarm needs to be triggered for verification.

[0107] The degree of climate reference phase mismatch is used to describe the time delay difference between climate fluctuations (irradiance / temperature) and equipment power response. If the time delay difference is too large, it indicates that the component response is lagging (aging leads to thermal conductivity deterioration) or leading (dust buffering effect), reflecting abnormal equipment status, which requires triggering an alarm for verification.

[0108] The system monitors the rate of change in power generation efficiency in real time. When the rate of change increases over time and exceeds the judgment threshold, it outputs an abnormal fluctuation indicator signal. The judgment threshold for the rate of change in power generation efficiency is set based on the long-term decay slope characteristics.

[0109] The difference between the timing reference phase of environmental parameters and the phase of power response is calculated as the mismatch. When the absolute value of the mismatch is greater than the judgment threshold, a climate interference identification signal is output. The judgment threshold of the mismatch is set based on the characteristics of the climate reference phase mismatch.

[0110] When both abnormal fluctuation and climate interference signals are received simultaneously, a climate noise filtering command is generated. Abnormal data segments that trigger both types of signals are marked as climate interference artifacts, removed from the performance degradation curve, and reconstructed by interpolation. Fault root cause classification is performed based on the sign of the mismatch: when efficiency continues to decrease but the phase mismatch exceeds the secondary threshold for a long time, it is determined to be a progressive aging fault; when the mismatch is positive, a heat dissipation system failure alarm is triggered; when the mismatch is negative, a surface contamination alarm is triggered.

[0111] Furthermore, the spatial truncation threshold is obtained as follows:

[0112] The target verification area is divided into L×L grids of equal area;

[0113] Calculate the number of stations D in each grid and the difference in number between each grid and its adjacent grids;

[0114] The spatial truncation threshold is based on the median nearest neighbor distance of the entire network. When the difference in the number of neighbors exceeds the first preset difference, the spatial truncation threshold is reduced proportionally by η1.

[0115] When the quantity difference is lower than the second preset difference, the space truncation threshold is expanded proportionally by η2.

[0116] The obtained spatial truncation threshold is used for spatial sub-cluster partitioning of the three-dimensional distance matrix.

[0117] In this embodiment, η1 and η2 are dimensionless coefficients for shrinking and expanding the basic spatial truncation threshold, respectively: when the difference in the number of adjacent grid stations ΔD exceeds the historical 95th percentile, the threshold is shrunk by multiplying η1 with the basic spatial truncation threshold to avoid excessive cross-regional aggregation; when the difference in the number of stations is lower than the historical 5th percentile, the threshold is expanded by multiplying η2 with the basic spatial truncation threshold to prevent excessive splitting within the cluster; the optimal values ​​of η1 and η2 are obtained through offline sensitivity analysis and multivariate regression calibration. Historical station distribution and fault location results are collected in typical mountainous areas, plains and urban areas. Combined with the clustering profile coefficient and false alarm rate, while ensuring clustering quality and verification accuracy, and taking into account computational performance, it is finally determined that η1 should be 0.8–0.9 and η2 should be 1.1–1.2. Fine-tuning can be carried out in conjunction with field verification if necessary.

[0118] Furthermore, the method for obtaining the threshold for determining the rate of change in power generation efficiency based on the long-term decay slope characteristics is as follows: In the labeled historical dataset, VMD separation and fitting are performed on the long-term power generation efficiency decay slope of each non-hotspot area site, the daily average efficiency change rate sequence within the corresponding time window is calculated, and the ratio distribution of this to the absolute value of the slope is statistically analyzed; then, the 95th percentile on this ratio distribution is selected as the benchmark sensitivity factor, and the real-time daily average efficiency change rate exceeding "benchmark sensitivity factor × current long-term decay slope absolute value" is considered an abnormal fluctuation; this sensitivity factor can be optimized through offline cross-validation to control the false alarm rate while ensuring high recall; The method for obtaining the threshold for determining the amount of mismatch based on the climate benchmark phase mismatch characteristics is as follows: In the steady-state operation data without anomalies, the time-series phase difference sequence of environmental parameters (irradiance / temperature) and power response is first extracted, and the phase mismatch characteristics are labeled in combination with equipment model and meteorological conditions; the mismatch characteristics are subjected to normal fitting or kernel density estimation, and the average value ± 2 standard deviations are taken as the upper and lower bidirectional threshold intervals; when the real-time phase difference exceeds this interval, a climate interference label is output. This threshold range can be periodically updated based on the latest operational data to adapt to seasonal and component differences.

[0119] In summary, Embodiment 1 of this invention combines design-climate similarity with measured performance differences through a dual-domain clustering comparison method to identify anomalous candidate clusters; it constructs a spatial loss rate gradient field for hotspot areas and traces the source in reverse, combined with aging time-series decomposition, to achieve accurate classification of pollution diffusion and aging faults. This enables rapid location and cause determination of power generation losses under multiple sites and meteorological conditions, effectively solving the problems of false alarms, missed alarms, and inaccurate location caused by meteorological interference and component differences in traditional methods.

[0120] Example 2: The difference between this embodiment and Example 1 is that...

[0121] The process of locating equipment performance anomalies includes the following steps:

[0122] Digital twin simulation calibration: A lightweight digital twin model is constructed based on the parameters of each photovoltaic site and its surrounding environment in the list of anomaly candidate clusters. The model simulates the process of pollutant deposition and component aging online, generates a gradient field for predicting loss rate, and uses the Kalman filter algorithm to perform closed-loop calibration between the simulation output and historical observations.

[0123] Closed-loop threshold calibration spatial density clustering: The gradient of the predicted loss rate of digital twins is combined with the latitude, longitude and altitude information of the site as clustering input. The spatial cutoff threshold is dynamically adjusted to divide spatial subclusters and determine geographical boundaries by weighted distance that integrates spatial distribution and simulation risk.

[0124] Spatiotemporal hotspot region identification: The distribution density of subclusters is calculated from the fusion clustering results. By combining the decoupling model of historical power generation loss rate sequence and environmental irradiance sequence, spatiotemporal hotspot regions that simultaneously meet the characteristics of high spatial density and high simulation risk are screened.

[0125] Simulation: Measured closed-loop verification and multi-scale aging feature extraction—In each hot spot area, the deviation between the loss rate predicted by the digital twin and the actual monthly and weekly average power generation efficiency is fed back to the digital twin model to correct the simulation parameters; For the power generation efficiency sequence in non-hot spot areas, wavelet packet decomposition, variational mode decomposition and short-time Fourier transform are applied to extract multi-scale aging trends and transient fluctuation features respectively.

[0126] Multi-source anomaly collaborative judgment: Based on closed-loop calibrated digital twin risk, spatial clustering features and multi-scale aging indicators, the causes of anomalies in hot and non-hot areas are judged separately, the anomaly category and its confidence level are output, and maintenance work orders integrating geographic coordinates and anomaly labels are generated.

[0127] Furthermore, the method also includes a data slicing management step, which involves extracting steady-state segments using irradiance-temperature dual thresholds; dividing the steady-state segments into bins according to preset irradiance intensity closed intervals and then calculating a normalized power sequence based on the average irradiance value within each bin; and storing the non-steady-state segments independently for subsequent fault tracing based on the component temperature drop irradiance invariance law and the inverter power collapse symptom marker forced retention event.

[0128] In summary, by adding a data slicing step in step two, the resource consumption of analysis is reduced, and excessive data consumption is avoided.

[0129] Furthermore, based on the stability of environmental parameters, data slicing is performed, and the data to be analyzed is filtered and output, including the following sub-steps:

[0130] Triggering data slicing: When the data cleaning module outputs the environmental parameter sequence and the power generation information sequence, the intensity of irradiance change and the intensity of component temperature change are calculated to generate a set of continuous time segments that meet the meteorological stability constraints, which are denoted as steady-state segments. The minimum segment length and maximum jump rate of the steady-state segments are limited. The meteorological stability constraints are dynamically adjusted based on the characteristics of atmospheric transmittance abrupt change and the difference in thermal inertia of photovoltaic modules.

[0131] Extract key events: Identify equipment operation state transition events in historical operation based on irradiance-temperature response decoupling coefficient, and calibrate restart symptom periods by associating them with inverter power collapse characteristics, and summarize and output a dataset that is forced to be retained;

[0132] Reconstruct equivalent environmental data: Apply steady-state fragments to steps two through four to reduce the data scale of subsequent processing, and force the retention of the dataset as training samples for identifying anomaly types (random forest classification model).

[0133] In summary, Embodiment 2 of this invention first extracts steady-state and key event fragments based on irradiation-temperature stability, combines dynamic gridded spatial truncation thresholds to screen high-quality input data and reduce computational load, and then follows the clustering and vector modeling process of Embodiment 1 to achieve efficient processing and high-precision clustering of core abnormal data in a large-scale data environment. This effectively solves the problems of wasted computational resources and spatial sub-clustering distortion caused by improper threshold selection in existing systems due to full data processing.

[0134] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for optimizing the location of power generation anomalies in distributed photovoltaic systems, characterized in that, include: Step 1: Asset Identification and Data Collection: Collect static information of all distributed photovoltaic sites within the target verification area and label the site index; Collect environmental parameters within the target verification time area and dynamic power generation information of all distributed photovoltaic sites, including inverter power generation curves and module temperatures; upload the spatiotemporally aligned static information, environmental parameters, and dynamic power generation information to the data analysis center; Step 2: Constructing a station-level theoretical benchmark: Obtain the theoretical performance ratio and the theoretical expected power of each distributed photovoltaic site. Using environmental parameters as input, output the ideal power generation efficiency under the current environmental parameters at the same time resolution. Call upon static information related to power generation efficiency, and use environmental parameters as boundary conditions to obtain the theoretical expected power. The ideal power generation efficiency refers to the energy conversion capability that a photovoltaic module should possess under a unit of solar irradiance under the current environmental parameters; it is a performance indicator reflecting the module's response to environmental conditions under standard conditions. The theoretical expected power refers to the power generation that should be output under ideal conditions under the current environmental parameters and static information. Step 3: Perform dual-domain clustering comparison: Perform K-means clustering on the theoretical expected power set to generate theoretical clusters CT; the ratio of actual power generation to expected power generation is the measured performance ratio, and perform DBSCAN clustering on the measured performance ratio set to generate measured clusters CP; Output a list of anomalous candidate clusters based on the cluster difference matrix; The method for outputting the list of abnormal candidate clusters based on the cluster difference matrix is ​​as follows: Record dual-domain cluster labels: acquire theoretical cluster labels and measured cluster labels for each distributed photovoltaic site to form dual-domain cluster label data; Construct a cluster difference statistics table: Using theoretical cluster labels as row indexes and measured cluster labels as column indexes, count the number of sites at the intersection of rows and columns to generate a cluster difference statistics table, thereby quantifying the correspondence between two-domain clustering; Determine the main performance cluster: For each theoretical cluster in the cluster difference statistics table, select the measured cluster with the largest number of cross-sites as the main performance cluster and establish an intra-cluster reference performance benchmark; Identify off-site sites: Compare the measured cluster labels of each site within the theoretical cluster with the main representation cluster labels, identify sites with inconsistent labels as off-site sites, and generate a set of off-site sites; Generate abnormal candidate clusters: Calculate the cluster deviation rate. When the cluster deviation rate is not lower than the first preset threshold, write the corresponding theoretical cluster into the abnormal candidate cluster list and record the cluster identifier, cluster capacity, deviation site index and cluster deviation rate to complete the abnormal candidate cluster registration. Sorting output list: Sort the list of abnormal candidate clusters from high to low according to the cluster deviation rate, and output the sorted list of abnormal candidate clusters to improve the accuracy of subsequent diagnostic priority determination; Step 4: Photovoltaic anomaly location: Summarize the time series data of the anomaly candidate cluster list, extract the spatiotemporal characteristics of the anomaly candidate clusters, and based on the spatiotemporal distribution characteristics, reveal the anomaly transmission path through pollution diffusion vector modeling, identify the anomaly as line anomaly, pollution deposition or component aging, and trigger the corresponding cleaning or replacement maintenance work order. Step four involves parallel execution of line anomaly diagnosis and equipment performance anomaly location. The line anomaly diagnosis is based on electrical topology reconstruction and line loss mutation monitoring, injecting characteristic harmonic currents into alarm sites, inverting the equivalent impedance value of the line, and diagnosing the fault type and location coordinates. The equipment performance anomaly location is achieved by generating anomaly spatial subclusters through spatial density clustering; constructing a loss rate-irradiation decoupling model to screen spatiotemporal hotspot areas; tracing back to the pollution boundary through pollution diffusion vector modeling; performing aging feature temporal decomposition on unmarked line anomaly sites; dynamically quantifying coupling effects and dynamically adjusting the aging tolerance threshold by integrating pollution acceleration factors.

2. The method for optimizing the location of power generation anomalies in distributed photovoltaic systems according to claim 1, characterized in that, The first preset threshold is a dynamic value that changes with the environment, and it is obtained as follows: Within the most recent W assessment cycles, the median of the deviation rate and the median of the absolute deviation of each theoretical cluster are calculated and added together to obtain the steady-state benchmark representing the background fluctuation of the equipment. Using the irradiance and component temperature within the same time window, the ratio of the standard deviation to the mean is calculated to obtain the irradiance fluctuation coefficient and the temperature rise fluctuation coefficient. The two are then weighted and summed to form the environmental fluctuation index. The environmental fluctuation index is mapped to the 0–1 interval using a linear normalization method to obtain a dimensionless environmental scale. The steady-state benchmark is processed by multiplication and scaling to generate a dynamic threshold, which is then clipped using preset upper and lower limits to obtain a first preset threshold that adapts to weather fluctuations.

3. The method for optimizing the location of power generation anomalies in distributed photovoltaic systems according to claim 1, characterized in that, The line anomaly diagnosis process includes the following steps: The regional electrical connection topology map is reconstructed based on the geographical coordinates of photovoltaic sites, and the hierarchical relationship of each site in the distribution network is marked. The daily line loss rate change trend is calculated by continuously monitoring the difference between the inverter output power and the metered power at the grid connection point. When the line loss change amplitude exceeds the threshold for several consecutive days, the line anomaly marker is triggered. Harmonic signals are injected into alarm sites, and the equivalent impedance of the line is retrieved based on the voltage response signal. The fault type is diagnosed based on the degree of deviation of the impedance value from the standard value: a significant increase in impedance indicates poor contact or oxidation of the connector; a significant decrease in impedance indicates insulation damage or partial short circuit; a local abnormal increase in impedance indicates capacitive load fault; and abnormal output line location coordinates and fault type labels are used.

4. The method for optimizing the location of power generation anomalies in distributed photovoltaic systems according to claim 1, characterized in that, The process of locating equipment performance anomalies includes the following steps: Step S11: Spatial density clustering segmentation: Calculate the three-dimensional geographic distance matrix based on the latitude, longitude and altitude information of the stations in the abnormal candidate cluster list, set the median of the nearest neighbor distance of all stations as the spatial truncation threshold, divide the spatial sub-clusters according to the spatial clustering intensity and determine the geographic boundary range. Step S12: Spatiotemporal hotspot region identification: Calculate the distribution density of abnormal sites using spatial subclusters, construct a loss rate-irradiance decoupling model by combining historical loss rate sequences and environmental irradiance sequences, and screen out spatiotemporal hotspot regions that simultaneously satisfy high distribution density and strong decoupling characteristics; Step S13: Pollution diffusion vector modeling: Construct the spatial loss rate gradient field of the spatiotemporal hotspot area, establish a distance inverse weight matrix to calculate the spatial change intensity and direction vector of the loss rate, aggregate to generate the main propagation direction vector and trace back to the normal threshold boundary point where the loss rate is less than the threshold. Step S14: For the monthly and weekly average power generation efficiency sequences of non-spatiotemporal hotspot areas, wavelet packet decomposition is used to extract multi-scale modes. Variational mode decomposition is applied to separate the long-term decay curve in the low-frequency trend mode, and the decay slope is calculated; short-time Fourier transform is applied to extract transient fluctuation characteristics in the high-frequency noise mode. For each non-spatiotemporal hotspot region, the trend slope is differentiated from the slope of neighboring stations in the same cluster and normalized. If the difference exceeds the product of the local residual standard deviation and the confidence factor, it is marked as a structural anomaly; otherwise, it is removed as random noise. The noise confidence level P after removing artifacts is calculated and output. Step S15: Multi-source anomaly collaborative determination: Determine the cause of anomalies in spatiotemporal hotspot areas and non-spatiotemporal hotspot areas respectively; output the anomaly category and classification confidence level; output a maintenance work order integrating geographic coordinates and anomaly category labels.

5. The method for optimizing the location of power generation anomalies in distributed photovoltaic systems according to claim 4, characterized in that, The noise artifact removal process includes the following steps: The system monitors the rate of change in power generation efficiency in real time. When the rate of change increases over time and exceeds the judgment threshold, it outputs an abnormal fluctuation indicator signal. The judgment threshold for the rate of change in power generation efficiency is set based on the long-term decay slope characteristics. The difference between the timing reference phase of environmental parameters and the phase of power response is calculated as the mismatch. When the absolute value of the mismatch is greater than the judgment threshold, a climate interference identification signal is output. The judgment threshold of the mismatch is set based on the characteristics of the climate reference phase mismatch. When both an abnormal fluctuation indicator signal and a climate interference indicator signal are received simultaneously, a climate noise filtering instruction is generated. Abnormal data segments that trigger both types of signals are marked as climate interference artifacts, removed from the performance degradation curve, and reconstructed by interpolation. Furthermore, fault root cause classification is performed based on the sign of the mismatch. When efficiency continues to decrease but the phase mismatch exceeds the secondary threshold for a long period of time, it is judged as a progressive aging fault; when the mismatch is positive, a heat dissipation system failure alarm is triggered, and when the mismatch is negative, a surface contamination alarm is triggered.

6. The method for optimizing the location of power generation anomalies in distributed photovoltaic systems according to claim 1, characterized in that, The process of locating equipment performance anomalies includes the following steps: Digital twin simulation calibration: A lightweight digital twin model is constructed based on the parameters of each photovoltaic site and its surrounding environment in the list of anomaly candidate clusters. The model simulates the process of pollutant deposition and component aging online, generates a gradient field for predicting loss rate, and uses the Kalman filter algorithm to perform closed-loop calibration between the simulation output and historical observations. Closed-loop threshold calibration spatial density clustering: The gradient of the predicted loss rate of digital twins is combined with the latitude, longitude and altitude information of the site as clustering input. The spatial cutoff threshold is dynamically adjusted to divide spatial subclusters and determine geographical boundaries by weighted distance that integrates spatial distribution and simulation risk. Spatiotemporal hotspot region identification: The distribution density of subclusters is calculated from the fusion clustering results. By combining the decoupling model of historical power generation loss rate sequence and environmental irradiance sequence, spatiotemporal hotspot regions that simultaneously meet the characteristics of high spatial density and high simulation risk are screened. Simulation: Measured closed-loop verification and multi-scale aging feature extraction—In each hot spot area, the deviation between the loss rate predicted by the digital twin and the actual monthly and weekly average power generation efficiency is fed back to the digital twin model to correct the simulation parameters; For the power generation efficiency sequence in non-hot spot areas, wavelet packet decomposition, variational mode decomposition and short-time Fourier transform are applied to extract multi-scale aging trends and transient fluctuation features respectively. Multi-source anomaly collaborative judgment: Based on closed-loop calibrated digital twin risk, spatial clustering features and multi-scale aging indicators, the causes of anomalies in hot and non-hot areas are judged separately, the anomaly category and its confidence level are output, and maintenance work orders integrating geographic coordinates and anomaly labels are generated.

7. The method for optimizing the location of power generation anomalies in distributed photovoltaic systems according to claim 1, characterized in that, The method also includes a data slicing management step, which involves extracting steady-state segments using irradiance-temperature dual thresholds; dividing the steady-state segments into boxes according to preset irradiance intensity closed intervals and then calculating a normalized power sequence based on the average irradiance value within each box; and storing the non-steady-state segments independently for subsequent fault tracing based on the law of constant irradiance due to component temperature drop and the mandatory retention event of inverter power collapse symptom markers.

8. The method for optimizing the location of power generation anomalies in distributed photovoltaic systems according to claim 7, characterized in that, Based on the stability of environmental parameters, data slicing and filtering are performed to output the data to be analyzed, including the following sub-steps: Triggering data slicing: When the data cleaning module outputs the environmental parameter sequence and the power generation information sequence, the intensity of irradiance change and the intensity of component temperature change are calculated to generate a set of continuous time segments that meet the meteorological stability constraints, which are denoted as steady-state segments. The minimum segment length and maximum jump rate of the steady-state segments are limited. The meteorological stability constraints are dynamically adjusted based on the characteristics of atmospheric transmittance abrupt change and the difference in thermal inertia of photovoltaic modules. Extract key events: Identify equipment operation state transition events in historical operation based on irradiance-temperature response decoupling coefficient, and calibrate restart symptom periods by associating them with inverter power collapse characteristics, and summarize and output a dataset that is forced to be retained; Reconstructing equivalent environmental data: Applying steady-state fragments to power consumption verification reduces the data scale of subsequent processing and forces the retention of the dataset as training samples for identifying anomaly types.

Citation Information

Patent Citations

  • Method for evaluating state of photovoltaic array in large photovoltaic power station

    CN112085350A

  • Method for diagnosing aging fault of photovoltaic module

    CN114519310A