A method and system for locating photovoltaic modules in photovoltaic power plants based on average power change rate
By constructing a personalized power prediction model and using wavelet transform and dynamic time warping algorithms, the system can accurately identify photovoltaic module faults, solving the problem of ambiguous fault location in photovoltaic power plants. This achieves automation and precision in fault diagnosis, improving operation and maintenance efficiency and power generation benefits.
Patent Information
- Application Number
- CN202510941458.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2045-07-09
AI Technical Summary
Existing photovoltaic power plant monitoring and diagnostic technologies have bottlenecks in the accuracy of fault location, making it impossible to accurately locate photovoltaic module faults. This forces maintenance personnel to spend a lot of time and manpower to troubleshoot one by one, and the delay in diagnosis may lead to the deterioration of the fault.
By constructing a personalized power prediction model based on health datasets and sky cloud images, and using wavelet transform and dynamic time warping algorithms, multi-dimensional fault feature vectors are extracted to accurately determine the independence or synergy of photovoltaic module faults, thereby achieving scientific judgment and precise location of fault sources.
It has achieved full automation and precision in the fault diagnosis process of photovoltaic power plants, reduced the diagnostic difficulty and workload of operation and maintenance personnel, and significantly improved operation and maintenance efficiency and power generation benefits.
Smart Images

Figure CN120781695B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of photovoltaic power plant technology, specifically relating to a method and system for positioning photovoltaic modules in a photovoltaic power plant based on the average power change rate. Background Technology
[0002] As a key technological pathway to achieving carbon neutrality, photovoltaic (PV) power generation is experiencing rapid growth in both the scale and complexity of its power plants. Ensuring the long-term, stable, and efficient operation of PV power plants is crucial for guaranteeing their economic and social benefits. Throughout the lifecycle of a PV power plant, PV modules and related electrical equipment can experience various failures due to aging, environmental corrosion, manufacturing defects, and other reasons. Timely detection and handling of these failures are essential for reducing power generation losses and preventing the escalation of safety hazards. Currently, large-scale PV power plants generally rely on Supervisory Control and Data Acquisition (SCADA) systems for macroscopic monitoring of the power generation units.
[0003] However, existing monitoring and diagnostic technologies have significant limitations in the accuracy of fault location. SCADA systems typically monitor data at the string or inverter level. When a system issues a power anomaly alarm, it can only indicate a relatively large fault range encompassing dozens or even hundreds of photovoltaic modules. Upon receiving the alarm, maintenance personnel cannot directly identify which module(s) are malfunctioning, nor can they determine whether the root cause lies with the module itself or with shared electrical equipment such as combiner boxes or inverters. This ambiguity in fault location makes subsequent troubleshooting extremely difficult, requiring maintenance personnel to spend considerable time and manpower checking each module individually. This is not only inefficient and costly but can also lead to worsening of the fault due to diagnostic delays. Summary of the Invention
[0004] This invention provides a method and system for locating photovoltaic modules in a photovoltaic power plant based on the average power change rate, in order to solve the above-mentioned technical problems.
[0005] In a first aspect, the present invention provides a method for locating photovoltaic modules in a photovoltaic power plant based on the average power change rate, the method comprising the following steps:
[0006] Acquire historical environmental data of the photovoltaic power plant and historical output power data of each photovoltaic string in the photovoltaic power plant;
[0007] The health status recognition algorithm identifies and marks abnormal data segments caused by known historical fault events in the historical output power data, and removes the abnormal data segments from the historical output power data to form a health status dataset.
[0008] By combining health status datasets and historical environmental data, machine learning algorithms are used to train and generate corresponding power prediction models for each photovoltaic string. The inputs to the power prediction models are real-time environmental data of the photovoltaic power station and cloud image data from the sky imager.
[0009] The actual output power of the photovoltaic string is acquired in real time, and the actual output power is subtracted from the theoretical output power output by the power prediction model to generate a performance deviation characteristic spectrum.
[0010] Wavelet transform is applied to decompose the performance deviation feature spectrum into multiple coefficient components containing time and frequency domain information, and multi-dimensional fault feature vectors are extracted based on the statistical indicators of the coefficient components.
[0011] Multiple photovoltaic strings under the same inverter are taken as an analysis unit. The dynamic time warping algorithm is used to measure the morphological similarity between the performance deviation characteristic spectra of each photovoltaic string in the analysis unit, and the fault anomaly of the photovoltaic string is judged as an independent fault or a coordinated fault based on the morphological similarity.
[0012] If it is an independent fault, the fault feature vector is matched with the preset fault mode library to determine the fault location and fault type of the faulty photovoltaic string; if it is a coordinated fault, the fault is located in the area where the inverter corresponding to the photovoltaic string is located.
[0013] Optionally, the step of identifying and marking abnormal data segments caused by known historical fault events in the historical output power data through the health status recognition algorithm includes the following steps:
[0014] Establish a historical maintenance work order database, which is used to record the occurrence time, end time and corresponding photovoltaic string location of each fault in the photovoltaic power station;
[0015] Align historical output power data with the historical maintenance work order database using timestamps;
[0016] For any photovoltaic string, if the historical output power data of the photovoltaic string within a certain time period coincides with the fault time period of the string recorded in the database;
[0017] Then all historical output power data points within the overlapping time period will be automatically identified and marked as abnormal data segments.
[0018] Optionally, after combining the health status dataset and historical environmental data to train and generate a corresponding power prediction model for each photovoltaic string using a machine learning algorithm, the following steps are also included:
[0019] Sky imagers deployed within photovoltaic power plants capture continuous frames of full-sky cloud images in real time.
[0020] Image analysis methods are used to analyze the entire sky cloud image to generate cloud condition prediction data for photovoltaic power plants;
[0021] Cloud condition prediction data is fused with real-time environmental data obtained through a sensor network deployed in photovoltaic power plants to obtain fused environmental data.
[0022] By using environmental fusion data as real-time input to the power prediction model, the theoretical output power of the power prediction model has been dynamically compensated for in advance for the impact of future cloud cover.
[0023] Optionally, the step of using image analysis methods to analyze the entire sky cloud image and generate cloud condition prediction data for photovoltaic power plants includes the following steps:
[0024] Image segmentation is performed on continuous frames of full-sky cloud images to identify the outline and extent of the cloud region above the photovoltaic power station;
[0025] Calculate the centroid position, area, and average pixel grayscale value of the cloud region in each frame of the full-sky cloud image, where the average pixel grayscale value is used to characterize the optical thickness of the cloud.
[0026] By comparing the changes in the centroid positions of various cloud clusters in continuous frames of full-sky cloud images, a cloud movement vector field covering the entire sky cloud image is calculated and generated.
[0027] The cloud movement vector field is spatially projected onto the physical topology model of the photovoltaic power station to predict the number of photovoltaic strings that each cloud cluster will cover at a predetermined time in the future.
[0028] By combining the optical thickness of the cloud layer and the rate of change of the centroid position, the irradiance shading coefficient that varies with time for each photovoltaic string is calculated as cloud condition prediction data.
[0029] Optionally, the step of applying wavelet transform to decompose the performance deviation feature spectrum into multiple coefficient components containing time and frequency domain information, and extracting a multi-dimensional fault feature vector based on the statistical indicators of the coefficient components, includes the following steps:
[0030] Extract the performance deviation characteristic spectrum over a preset time period;
[0031] Define the wavelet basis functions and the number of decomposition levels for the wavelet transform;
[0032] Based on the wavelet basis function and the number of decomposition levels, a discrete wavelet transform is performed on the truncated performance deviation feature spectrum to decompose the truncated performance deviation feature spectrum into a set of approximate coefficients and multiple sets of detail coefficients.
[0033] The approximation coefficient is used as a slowly varying component reflecting the low-frequency trend of the performance deviation characteristic spectrum;
[0034] The detail coefficients are used as abrupt components reflecting high-frequency perturbations in the performance deviation characteristic spectrum;
[0035] A multi-dimensional fault feature vector is constructed based on the slowly varying component and the abrupt change component.
[0036] Optionally, constructing a multi-dimensional fault feature vector based on slowly varying components and abruptly changing components includes the following steps:
[0037] Calculate the signal energy of the slowly varying component and the abruptly changing component in their respective frequency bands;
[0038] Calculate the signal kurtosis of the slowly varying component and the abruptly changing component respectively;
[0039] Calculate the signal entropy values for the slowly varying component and the abruptly changing component respectively;
[0040] The calculated signal energy, signal kurtosis, and signal entropy values are normalized.
[0041] The normalized signal energy, signal kurtosis, and signal entropy values are arranged in a predetermined order and combined to form a multi-dimensional fault feature vector.
[0042] Optionally, the step of taking multiple photovoltaic strings under the same inverter as an analysis unit, using a dynamic time warping algorithm to measure the morphological similarity between the performance deviation characteristic spectra of each photovoltaic string within the analysis unit, and judging whether the fault anomaly of the photovoltaic string is an independent fault or a coordinated fault based on the morphological similarity includes the following steps:
[0043] Taking multiple photovoltaic strings under the same inverter as an analysis unit, the performance deviation characteristic spectrum of all photovoltaic strings under the analysis unit is obtained;
[0044] Pair the performance deviation characteristic spectra of any two photovoltaic strings within the analysis unit;
[0045] The dynamic time warping algorithm is used to calculate the time warping distance between each pair of performance deviation feature spectra. The time warping distance characterizes the morphological similarity between the two performance deviation feature spectra.
[0046] Based on the time warped distance of all pairs, a similarity matrix is constructed to characterize the similarity relationship between all photovoltaic strings within the analysis unit;
[0047] Based on the clustering characteristics of the similarity matrix, it can be determined whether the fault or anomaly of the photovoltaic string is an independent fault or a collaborative fault.
[0048] Optionally, determining whether a fault or a coordinated fault in a photovoltaic string is an independent fault based on the clustering characteristics of the similarity matrix includes the following steps:
[0049] Preset time-normalization distance threshold and collaborative quantity ratio threshold;
[0050] Traverse the similarity matrix and count the number of photovoltaic strings whose morphological similarity to any photovoltaic string is higher than the time warp distance threshold;
[0051] Determine whether the number of photovoltaic strings exceeds the threshold for the proportion of coordinated units;
[0052] If the number of photovoltaic strings exceeds the threshold for the proportion of coordinated units, the fault or abnormality occurring in the corresponding photovoltaic string will be judged as a coordinated fault.
[0053] If the number of photovoltaic strings does not exceed the threshold for the proportion of coordinated quantities, the fault or abnormality occurring in the corresponding photovoltaic string will be judged as an independent fault.
[0054] In a second aspect, the present invention also provides a photovoltaic module positioning system for a photovoltaic power station based on average power change rate, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the photovoltaic module positioning method for a photovoltaic power station based on average power change rate as described in any one of the first aspects.
[0055] Thirdly, the present invention also provides a computer-readable storage medium storing instructions that, when executed by a processor, cause the processor to perform a photovoltaic module positioning method for a photovoltaic power plant based on the average power change rate according to any one of the first aspects.
[0056] The beneficial effects of this invention are:
[0057] This invention establishes an extremely accurate theoretical performance benchmark for each photovoltaic string by constructing a personalized power prediction model based on a health dataset and sky cloud images, thereby enabling it to keenly detect subtle performance deviations caused by faults. Its core innovation lies not only in utilizing wavelet transform for deep time-frequency analysis of the performance deviation feature spectrum to extract multi-dimensional feature vectors that finely characterize the essence of the fault, but also in uniquely introducing a dynamic time warping algorithm. By measuring the morphological similarity of fault characteristics across different strings, it scientifically judges the level of the fault source at the initial diagnostic stage, effectively distinguishing between independent faults of the component itself and regional collaborative faults such as those of the inverter. This differentiation mechanism fundamentally solves the problems of ambiguous fault location and unclear responsibility attribution in traditional methods, avoiding ineffective investigation of a large number of normal components. This invention elevates fault diagnosis from simple threshold judgment to intelligent reasoning based on morphology and feature matching, achieving full automation and precision from fault detection and classification to fault level location. This greatly reduces the diagnostic difficulty and workload for maintenance personnel, significantly improving the operation and maintenance efficiency and power generation benefits of photovoltaic power plants. Attached Figure Description
[0058] Figure 1 This is a flowchart illustrating a photovoltaic module positioning method for a photovoltaic power station based on average power change rate in one embodiment of this application.
[0059] Figure 2 This is a schematic diagram of the process for generating cloud condition prediction data in one embodiment of this application.
[0060] Figure 3 This is a schematic diagram of the process for extracting fault feature vectors in one embodiment of this application. Detailed Implementation
[0061] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0062] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0063] Figure 1 This is a flowchart illustrating a photovoltaic module positioning method for a photovoltaic power plant based on average power change rate in one embodiment. It should be understood that, although... Figure 1 The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order in which these steps are executed, and they can be performed in other orders. Figure 1 At least some steps in the process may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in turn with other steps or at least a portion of the sub-steps or stages of other steps. For example Figure 1 As shown, the photovoltaic module positioning method for photovoltaic power plants based on average power change rate disclosed in this invention specifically includes the following steps:
[0064] S101. Obtain historical environmental data of the photovoltaic power station and historical output power data of each photovoltaic string in the photovoltaic power station.
[0065] The core of this step lies in comprehensively collecting various historical information reflecting the operational status of the photovoltaic power plant. Specifically, a sensor network deployed at the power plant continuously records key environmental parameters, such as total solar irradiance, ambient temperature, and wind speed. Simultaneously, DC-side output power data for each photovoltaic string within the same time period is retrieved from the power plant's monitoring system. These two types of data are strictly aligned in time, forming a massive time-series dataset encompassing both environmental inputs and power outputs. The ultimate result of this process is the construction of a complete historical information database. This database not only records the power plant's power generation performance but, more importantly, captures the intrinsic correlation between power generation performance and the external environment, providing essential raw materials for subsequent data cleaning and model training.
[0066] S102. Identify and mark abnormal data segments caused by known historical fault events in the historical output power data through a health status recognition algorithm, remove abnormal data segments from the historical output power data, and form a health status dataset.
[0067] To ensure that the subsequent machine learning model learns the true power generation patterns of the photovoltaic strings under healthy conditions, the original historical data must be purified. This step involves using known maintenance records to identify and remove fault segments from the data. The implementation method is to first establish a structured historical maintenance work order database, which records in detail the occurrence time, end time, and corresponding photovoltaic string location of each fault. Then, this database is timestamped with the collected historical output power data. When the time interval of a power data segment for a particular string completely overlaps with the fault time recorded in the database for that string, this power data segment is automatically marked as an abnormal data segment and removed. In this way, all known data pollution caused by faults is accurately removed, ultimately forming a clean healthy state dataset. This significantly improves the quality of subsequent model training, ensuring that the model accurately reflects the power generation characteristics under normal operating conditions.
[0068] S103. Combining the health status dataset and historical environmental data, a corresponding power prediction model is trained and generated for each photovoltaic string using machine learning algorithms. The input to the power prediction model is the real-time environmental data of the photovoltaic power station and cloud image data from the sky imager.
[0069] Based on the purified health status dataset, a dedicated, high-precision power prediction model is constructed for each photovoltaic string. This step aims to establish a digital benchmark capable of accurately simulating the theoretical power generation of a healthy string under arbitrary environmental conditions. Specifically, advanced machine learning algorithms, such as gradient boosting decision trees or long short-term memory networks, are employed for independent training of each string. The model's input integrates real-time environmental data from the power plant, such as irradiance G and module temperature T, as well as real-time cloud imagery data C provided by a sky imager. By learning from historical health data, the model establishes a correlation between these input variables and the output power under healthy conditions. The complex nonlinear mapping relationship between them can be expressed as: After training, a power prediction model corresponding to each photovoltaic string in the power plant will be obtained. Its effect is that, under any real-time weather conditions, it can provide a dynamic and accurate theoretical output power reference value for each string, serving as the gold standard for judging whether it deviates from its normal operating state.
[0070] S104. Obtain the actual output power of the photovoltaic string in real time, and subtract the actual output power from the theoretical output power output by the power prediction model to generate a performance deviation characteristic spectrum.
[0071] After obtaining the theoretical power benchmark, the real-time fault detection phase can begin, the core of which lies in quantifying the deviation between actual power generation performance and theoretical values. This process is accomplished by comparing the actual output power of the photovoltaic strings with the theoretical output power predicted by the model in real time. In practice, the system continuously collects the actual output power of each string from the field and simultaneously inputs real-time environmental data into the power prediction model of the corresponding string to obtain the theoretical output power. Subsequently, the two are subtracted to generate a performance deviation sequence that varies over time. This sequence is the performance deviation characteristic spectrum. The effect of this step is to transform the abstract concept of fault into a concrete and measurable mathematical signal. When the string is running healthily, the value of this characteristic spectrum fluctuates slightly around zero; once a fault occurs and causes a power drop, the characteristic spectrum will show a significant positive value, the magnitude of which directly reflects the severity of the power loss, providing a direct basis for subsequent fault feature extraction.
[0072] S105. Wavelet transform is applied to decompose the performance deviation feature spectrum into multiple coefficient components containing time and frequency domain information, and multi-dimensional fault feature vectors are extracted based on the statistical indicators of the coefficient components.
[0073] To extract deeper and more discriminative fault information from the original performance deviation feature spectrum, wavelet transform technology was applied for signal decomposition and feature extraction. The principle of wavelet transform lies in its ability to simultaneously provide localized information about the signal in both the time and frequency domains, thereby effectively distinguishing different types of fault signals. In practice, a performance deviation feature spectrum of a predetermined time length was first extracted, and appropriate wavelet basis functions and decomposition levels were selected. By performing discrete wavelet transform, the signal was decomposed into approximate coefficients representing its low-frequency trends and multi-layer detail coefficients representing its high-frequency disturbances. The former reflects slowly changing fault characteristics such as aging and dust accumulation, while the latter captures details of sudden faults such as obstruction and open circuits. The effectiveness of this step lies in transforming the one-dimensional deviation signal into a multi-dimensional feature set. These coefficient components characterize the dynamic behavior of the fault at different scales and frequencies, significantly amplifying the inherent differences between different fault modes and laying a solid foundation for subsequent accurate fault identification and classification.
[0074] S106. Taking multiple photovoltaic strings under the same inverter as an analysis unit, the dynamic time warping algorithm is used to measure the morphological similarity between the performance deviation characteristic spectra of each photovoltaic string in the analysis unit, and the fault anomaly of the photovoltaic string is determined as an independent fault or a coordinated fault based on the morphological similarity.
[0075] After obtaining the multi-dimensional features of the fault signal, it is necessary to determine whether the fault is an independent event affecting a single string or a coordinated event affecting a region. This step utilizes the Dynamic Time Warping (DTW) algorithm to measure the morphological similarity of the performance deviation feature spectra of each string under the same inverter. The DTW algorithm can calculate the optimal matching path between two time series of unequal durations, thereby obtaining their similarity distance, even if there is some scaling or translation on the time axis. In specific implementation, all photovoltaic strings under the same Verilog inverter are taken as an analysis unit, and the DTW distance between their performance deviation feature spectra is calculated pairwise. If the deviation spectrum morphology of most strings is highly similar, that is, their DTW distances are generally small, it indicates that there is a common source of influence. The effectiveness of this method lies in its ability to effectively distinguish the root cause of the fault: if only one or a few strings show abnormalities, it is determined to be an independent fault, and the problem may lie in the string itself; if most strings in an analysis unit show similar performance degradation at the same time, it is determined to be a coordinated fault, and the problem is likely to originate from the inverter they share or a large area of external shading.
[0076] S107. If it is an independent fault, the fault feature vector is matched with the preset fault mode library to determine the fault location and fault type of the faulty photovoltaic string; if it is a coordinated fault, the fault is located in the area where the inverter corresponding to the photovoltaic string is located.
[0077] Finally, based on the determination of whether a fault is independent or collaborative, the system performs a precise fault location. This final step aims to provide maintenance personnel with clear and actionable troubleshooting instructions. When a fault is determined to be independent, the system matches the previously extracted multi-dimensional fault feature vector with a pre-established fault pattern library. This library stores standard feature vectors corresponding to various known faults (such as module hotspots, diode short circuits, and partial shading). By calculating the similarity between the feature vector of the fault to be diagnosed and each pattern in the library, the specific type and location of the fault can be determined with the highest probability, down to a specific photovoltaic string. When a fault is determined to be collaborative, since the problem originates from common equipment, the system directly locates the fault to the area where the inverters connected to these strings are located. The ultimate effect of this step is to achieve a closed loop in fault diagnosis, not only identifying problems but also providing highly accurate location and preliminary type judgment, thereby greatly improving the operation and maintenance efficiency of photovoltaic power plants and shortening fault repair time.
[0078] In one embodiment, identifying and marking abnormal data segments caused by known historical fault events in historical output power data using a health status recognition algorithm includes the following steps:
[0079] Establish a historical maintenance work order database, which is used to record the occurrence time, end time and corresponding photovoltaic string location of each fault in the photovoltaic power station;
[0080] Align historical output power data with the historical maintenance work order database using timestamps;
[0081] For any photovoltaic string, if the historical output power data of the photovoltaic string within a certain time period coincides with the fault time period of the string recorded in the database;
[0082] Then all historical output power data points within the overlapping time period will be automatically identified and marked as abnormal data segments.
[0083] In this implementation, to ensure the data quality for subsequent model training, the primary task is to construct a structured and standardized historical maintenance work order database. The principle behind this step is to transform fragmented, unstructured maintenance records into a machine-readable format, serving as an authoritative basis for data cleaning. Specifically, this requires creating a database table containing key fields, such as assigning a unique ID to each fault and recording its corresponding photovoltaic string location code, the exact timestamp of the fault's start, and the timestamp of the fault's repair. A fault record R can be formally represented as a tuple. ,in It is the unique identifier of the string. and These are the start and end times of the fault, respectively. By entering all historical paper or electronic work order information into this database, a comprehensive, accurate, and quickly searchable panoramic archive of historical fault events is ultimately formed, providing a solid foundation for subsequent automated data labeling.
[0084] After obtaining historical fault databases and historical power data, it is crucial to ensure that they correspond precisely in the time dimension; this is a prerequisite for automatic tagging. The principle behind this step is to unify the time base of the two heterogeneous datasets, eliminating deviations caused by differences in time formats, time zones, or recording frequencies. In practice, this requires standardizing the timestamps of historical output power data and the start and end times in the maintenance work order database, for example, by converting them all to a unified International Standard Time (UTC) or UNIX timestamp. Each power data point inherently carries a measurement timestamp. Each fault record defines a time interval. Timestamp alignment ensures , and These three can be directly compared in magnitude. The effect of this process is to establish a reliable time bridge connecting maintenance events and performance data, allowing the program to unambiguously determine whether any power sampling point falls within a known fault time window.
[0085] Finally, based on the aligned data, the core process of automatic anomaly data identification and labeling is executed. The fundamental principle of this step is to use established historical fault facts to precisely filter and purify massive amounts of historical power data. Specifically, this involves systematically traversing all historical output power data for each photovoltaic string. For any given power data point, its string location is determined. and collection time The program will automatically query the historical maintenance work order database to check if there is a matching work order. Matches and satisfies the conditions The system records fault data. Once a record meeting the criteria is found, the power data point is immediately marked as an anomaly. This process is repeated until all historical data has been reviewed. The ultimate result of this step is a cleaned and labeled power dataset, in which all distorted data caused by known equipment faults are accurately identified and can be easily removed from subsequent model training datasets, ensuring the purity of the training data.
[0086] In one implementation, after combining health status datasets and historical environmental data to train and generate a corresponding power prediction model for each photovoltaic string using machine learning algorithms, the following steps are also included:
[0087] Sky imagers deployed within photovoltaic power plants capture continuous frames of full-sky cloud images in real time.
[0088] Image analysis methods are used to analyze the entire sky cloud image to generate cloud condition prediction data for photovoltaic power plants;
[0089] Cloud condition prediction data is fused with real-time environmental data obtained through a sensor network deployed in photovoltaic power plants to obtain fused environmental data.
[0090] By using environmental fusion data as real-time input to the power prediction model, the theoretical output power of the power prediction model has been dynamically compensated for in advance for the impact of future cloud cover.
[0091] In this implementation, firstly, to anticipate the impact of cloud movement on power generation, it is necessary to acquire real-time visual information of the sky above the photovoltaic power station. This step utilizes a wide-angle imaging device to continuously capture images of the entire sky at a fixed high frequency, thus forming a dynamic visual data stream. Specifically, one or more all-sky imagers equipped with fisheye lenses will be deployed in an open area within the photovoltaic power station site. These devices will face the zenith, enabling them to capture a 180-degree view of the sky without blind spots. The camera will automatically capture a high-resolution color cloud image at preset time intervals, such as every 30 seconds or 1 minute, and these image sequences will be uploaded. The data is transmitted in real time to the data processing center. The direct result of this process is a continuous sequence of full-sky images with precise timestamps, which acts like a live video of the sky, providing the most original and intuitive data foundation for subsequent cloud identification, motion analysis, and occlusion prediction.
[0092] After obtaining continuous cloud imagery, the next step is to utilize advanced image analysis techniques to transform this pixel information into quantifiable, predictive cloud condition data. The core principle of this step is computer vision, which uses algorithms to automatically identify cloud clusters in the image and analyze their physical properties and movement trends. One implementation method involves first processing each frame of the cloud image... Image segmentation is performed, using pixel color (such as red-to-blue ratio) or brightness thresholds to precisely separate the sky background from the cloud region, determining the outline of each individual cloud. Next, key features, including the centroid location, are calculated for each identified cloud. To determine its spatial coordinates, pixel area to measure its size, and average pixel grayscale value to characterize its optical thickness or light-blocking ability. This is achieved by comparing the same cloud cluster in consecutive frames. and By observing the change in the position of the center of mass, its velocity vector can be calculated. Its formula can be simplified to The effect of this step is to convert the raw visual image into structured digital information, that is, to generate a precise description of the location, size, thickness and future trajectory of each cloud over the power station.
[0093] Subsequently, to construct a comprehensive environmental dataset that reflects both the current macroscopic environment and predicts future local changes, it is necessary to effectively integrate the cloud condition prediction data generated in the previous step with traditional sensor data. This leverages complementary strengths: sensor networks (such as irradiance meters and thermometers) provide extremely accurate single-point, real-time environmental parameters, while cloud condition prediction data provides trend information on large-scale changes in illumination over a short period. In implementation, cloud condition prediction data (such as the predicted cloud cover coefficient) will be integrated... And cloud speed (s) and environmental data collected in real time by sensor networks (such as total horizontal irradiance G and solar panel temperature) This integration can be done at the data level, combining these different data sources into a higher-dimensional feature vector. ,For example The result of this approach is the generation of a fused environmental data set that not only contains the precise environmental conditions at the current moment but also includes predictions of changes in illumination caused by cloud movement in the coming minutes, providing unprecedented depth of information for power prediction models.
[0094] Finally, this information-rich environmental fusion data is used as input to a pre-trained machine learning power prediction model to generate a theoretical output power that can dynamically compensate for the impact of future cloud cover. By providing the model with forward-looking input, its predictive ability shifts from passive response to proactive prediction. In practice, this multi-dimensional environmental feature vector, containing both real-time and predictive information, is used as the input variable for the power prediction function. The model calculates the expected theoretical output power based on the complex mapping relationships learned from health status data. Since the input already includes information about the impending cloud cover, the theoretical power curve output by the model will reflect the power decrease caused by cloud cover in advance. The final result is that the obtained theoretical output power benchmark is no longer a smooth, ideal curve, but a dynamic curve that has taken into account cloud disturbances and is closer to the actual health status. This allows for effective avoidance of false alarms caused by sudden weather changes when identifying performance deviations by comparing actual power with theoretical power, significantly improving the accuracy of fault detection.
[0095] In one implementation, the process of analyzing a full-sky cloud image using image analysis methods to generate cloud condition prediction data for a photovoltaic power station includes the following steps:
[0096] Image segmentation is performed on continuous frames of full-sky cloud images to identify the outline and extent of the cloud region above the photovoltaic power station;
[0097] Calculate the centroid position, area, and average pixel grayscale value of the cloud region in each frame of the full-sky cloud image, where the average pixel grayscale value is used to characterize the optical thickness of the cloud.
[0098] By comparing the changes in the centroid positions of various cloud clusters in continuous frames of full-sky cloud images, a cloud movement vector field covering the entire sky cloud image is calculated and generated.
[0099] The cloud movement vector field is spatially projected onto the physical topology model of the photovoltaic power station to predict the number of photovoltaic strings that each cloud cluster will cover at a predetermined time in the future.
[0100] By combining the optical thickness of the cloud layer and the rate of change of the centroid position, the irradiance shading coefficient that varies with time for each photovoltaic string is calculated as cloud condition prediction data.
[0101] In this embodiment, the significant differences in optical properties between clouds and clear sky are first utilized to separate them at the pixel level. One specific implementation involves converting the real-time captured color cloud image to a more discriminative color space, and then identifying each pixel based on the red-to-blue ratio. Since the ratio of red to blue light components in cloud pixels is typically much larger than that in clear sky backgrounds, a discrimination threshold can be set. When the red-to-blue ratio of any pixel is greater than this threshold, it is classified as a cloud region; otherwise, it is classified as sky background. By performing this operation on all pixels in the image, a binarized mask image is ultimately generated. This image clearly outlines the precise contours and coverage of all cloud clusters, transforming unstructured visual information into structured region data that can be analyzed by computers.
[0102] After accurately identifying the cloud regions, their geometric and physical characteristics need to be further quantified for dynamic analysis. The visual morphology of each segmented individual cloud cluster is converted into a set of descriptive numerical parameters. Specifically, for each connected region (i.e., a single cloud cluster) in the binary mask image generated in the previous step, its key features are calculated using image processing algorithms. This includes counting the total number of pixels in the region to obtain the cloud cluster's area A; and calculating the arithmetic mean of all pixel coordinates to determine its centroid position in the image coordinate system. Simultaneously, in the original grayscale image, the average grayscale value of all pixels within the area covered by the cloud is calculated. This average grayscale value is used as a key indicator of the optical thickness of the cloud layer, because thicker clouds with higher water content typically appear brighter in the image. The effect of this step is to transform each cloud from a uniform graphic into a digital object defined by parameters such as location, size, and thickness.
[0103] Next, by continuously tracking these digitized cloud objects, their motion patterns can be accurately deduced. This step is based on a fundamental physical assumption: that, over a short period of time, the motion of clouds can be considered uniform linear motion. Specifically, a matching algorithm is used to identify clouds in two consecutive frames (with a time interval of...). Find the same cloud cluster and obtain its centroid position. and The velocity vector of the cloud cluster. It can be calculated by dividing the difference in the positions of the centroids at two moments by the time interval, i.e. Repeating this process for all traceable clouds within the field of view generates a cloud movement vector field covering the entire sky. This vector field visually depicts the real-time movement direction and speed of clouds in different areas above the power station using a series of arrows. Essentially, it creates an instant wind direction and speed map of the sky, providing a dynamic basis for forecasting.
[0104] After obtaining the cloud movement vector field, it needs to be combined with the actual geographical layout of the photovoltaic power station to achieve accurate prediction from the sky to the ground. The principle of this step is to map the cloud movement in the image coordinate system onto the photovoltaic string array in the ground physical coordinate system through geometric projection relationships. In implementation, a precise physical topology model of the photovoltaic power station needs to be established in advance, which includes the geographical coordinates of each photovoltaic string. Simultaneously, a projection transformation function from all-sky image pixels to ground coordinates needs to be established. Using the current position and velocity vector of the cloud calculated in the previous step, its position at any predetermined time in the future can be predicted. The predicted cloud formation is then projected onto the physical topology model of the power plant using a projection transformation function. This allows us to determine which specific photovoltaic strings will be covered by the cloud's shadow. This process establishes a link between macroscopic cloud images and microscopic strings, enabling the prediction of which cloud formation will affect which strings at a specific future point in time.
[0105] Finally, to ensure the prediction results can be directly used in the power model, qualitative shading events need to be quantified into specific irradiance impact coefficients. The principle behind this step is that the degree to which clouds reduce solar irradiance is not only related to their presence but also closely related to their physical properties (thickness) and movement speed. In practice, for each photovoltaic string predicted to be shaded in the future, calculations are performed using two key characteristics of the cloud cluster: one is the average pixel grayscale value representing optical thickness. Secondly, the magnitude of the rate of change of its center of mass position. A thicker one ( Higher value), slower movement ( Clouds with smaller irradiance values will result in more severe and longer-lasting irradiance reduction. Therefore, an irradiance shading coefficient that varies over time can be constructed for each photovoltaic string. and The function generates a refined occlusion prediction curve for each string, covering the next few minutes to tens of minutes. This data is the final prediction product that incorporates the impact of dynamic cloud conditions.
[0106] In one implementation, the process of applying wavelet transform to decompose the performance deviation feature spectrum into multiple coefficient components containing time and frequency domain information, and extracting a multi-dimensional fault feature vector based on the statistical indices of the coefficient components, includes the following steps:
[0107] Extract the performance deviation characteristic spectrum over a preset time period;
[0108] Define the wavelet basis functions and the number of decomposition levels for the wavelet transform;
[0109] Based on the wavelet basis function and the number of decomposition levels, a discrete wavelet transform is performed on the truncated performance deviation feature spectrum to decompose the truncated performance deviation feature spectrum into a set of approximate coefficients and multiple sets of detail coefficients.
[0110] The approximation coefficient is used as a slowly varying component reflecting the low-frequency trend of the performance deviation characteristic spectrum;
[0111] The detail coefficients are used as abrupt components reflecting high-frequency perturbations in the performance deviation characteristic spectrum;
[0112] A multi-dimensional fault feature vector is constructed based on the slowly varying component and the abrupt change component.
[0113] In this implementation, to ensure the real-time nature and operability of the analysis, fixed-length segments need to be extracted from the continuous performance deviation data stream for processing. By using a windowing method, an infinitely long time-series signal is transformed into a series of finite-length, computationally efficient analysis units, thereby achieving periodic snapshot evaluation of the system state. Specifically, a preset time window length L is set, the selection of which must balance computational efficiency with the need to capture the complete dynamic characteristics of the fault. Subsequently, data from the most recent L time points are extracted from the real-time generated performance deviation feature spectrum to form a discrete time-series sample. The effect of this operation is to discretize the continuous monitoring process into independent analysis tasks, providing standardized and regular data objects for subsequent applications such as discrete wavelet transform, which require algorithms with finite input lengths.
[0114] Next, to ensure the wavelet transform decomposes the signal most effectively, its core parameters must be pre-defined: the wavelet basis function and the number of decomposition levels. The principle behind this step is that different fault types exhibit different waveform characteristics in the signal; selecting a wavelet basis function similar to the fault waveform shape can achieve optimal energy matching and feature extraction. Simultaneously, the number of decomposition levels determines the level of detail in the signal analysis in the frequency domain. In practice, based on empirical knowledge or prior experiments, a suitable wavelet basis function is selected for common fault characteristics of photovoltaic systems (such as step faults, gradual changes, and oscillations). For example, the Daubechies wavelet, which has good tight support. Then, an integer decomposition level J is set, which determines how low the frequency the signal will be decomposed to. This step sets a precise scale and benchmark for subsequent transform analysis, much like selecting the appropriate objective lens and magnification for a microscope, ensuring the targeted and effective decomposition process. After setting the parameters, the core decomposition operation is performed on the extracted performance deviation characteristic spectrum samples.
[0115] By utilizing the discrete wavelet transform, a one-dimensional time-domain signal is mapped to a two-dimensional time-frequency plane, thus simultaneously revealing the signal's frequency composition and its variation over time. In practice, the truncated performance deviation spectrum is used as input, and a selected wavelet basis function and decomposition level J are applied through an iterative process of multi-level filtering and downsampling. The mathematical essence of this process can be expressed as follows: After J-level decomposition, the original signal is decomposed into a set of approximate coefficients representing the low-frequency profile of the signal. Group J represents the detail coefficients of high-frequency information in different frequency bands. The effectiveness of this approach lies in its ability to successfully separate the various fault information that are mixed together into different coefficient components according to their physical characteristics of changing speed, thus realizing the transformation from the original mixed signal to a multi-channel characteristic signal.
[0116] After obtaining the decomposed coefficients, the focus is first on the slowly varying components representing the low-frequency trend of the signal. The principle behind this information is that the approximate coefficients are obtained by repeatedly low-pass filtering the signal, which removes all high-frequency noise and abrupt changes, preserving the core and smoothest overall contour of the signal. In practice, the only set of approximate coefficients obtained after J-level decomposition is directly extracted. This coefficient sequence physically corresponds to the power loss in performance deviations caused by factors such as equipment aging, seasonal lighting changes, and large-area uniform dust accumulation, characterized by long periods and gradual amplitude. The effect of this step is to successfully extract the long-term, trend-based variation characteristics from the complex deviation signal, providing a clear and interference-free basis for diagnosing those slowly developing but far-reaching gradual faults. Correspondingly, another crucial piece of information is contained in the abrupt components reflecting high-frequency disturbances. The principle behind this is that the detail coefficients... It is the product of a signal passing through a series of bandpass filters, with each set of coefficients precisely corresponding to the instantaneous changes and local singularities of the original signal within a specific frequency band. In practice, all J sets of detail coefficients obtained from the J-layer decomposition are... All of these parameters are extracted. The energy concentration regions of these coefficient sequences precisely pinpoint the instantaneous events of the fault in the time domain, such as sudden open circuits in components, rapid formation of hot spots, or rapid obstruction by small clouds. These sudden events cause spikes, glitches, or steps in the power curve. The effect of this step is to successfully capture all sudden events and high-frequency details in the performance deviation signal, providing crucial location and qualitative information for diagnosing transient faults that occur suddenly and change drastically.
[0117] Finally, by calculating the statistical properties of each coefficient component, the original coefficient sequences of varying lengths and physical meanings are compressed into a fixed-length, highly condensed numerical vector—the digital fingerprint of the fault. In practice, statistical indicators such as signal energy, kurtosis, and entropy are calculated for both the approximation coefficients and each set of detail coefficients. For example, signal energy can be expressed by the formula... The calculations show that 'c' represents a set of coefficients. Arranging all the calculated statistical indicators in a predetermined order forms a multi-dimensional fault feature vector. This approach represents a crucial step from signal processing to pattern recognition, transforming the complex waveform analysis problem into a standardized vector classification problem, laying the foundation for subsequent automatic matching with a fault database.
[0118] In one implementation, constructing a multi-dimensional fault feature vector based on slowly varying and abruptly changing components includes the following steps:
[0119] Calculate the signal energy of the slowly varying component and the abruptly changing component in their respective frequency bands;
[0120] Calculate the signal kurtosis of the slowly varying component and the abruptly changing component respectively;
[0121] Calculate the signal entropy values for the slowly varying component and the abruptly changing component respectively;
[0122] The calculated signal energy, signal kurtosis, and signal entropy values are normalized.
[0123] The normalized signal energy, signal kurtosis, and signal entropy values are arranged in a predetermined order and combined to form a multi-dimensional fault feature vector.
[0124] In this implementation, firstly, to quantify the intensity of different fault components, it is necessary to calculate the signal energy of the slowly varying component and the abruptly changing component within their respective frequency bands. The principle behind this step is that signal energy is physically proportional to the square of the signal amplitude, which can intuitively reflect the magnitude of the fault disturbance. A component with a large energy value means that its corresponding fault phenomenon (whether slowly varying or abruptly changing) contributes more to the system performance deviation. Specifically, for the slowly varying component (approximate coefficient) obtained from the discrete wavelet transform... ) and mutation components at all levels (details coefficients) The energy of the signal is calculated using the formulas for each component. For example, the energy of any coefficient component c is calculated using these formulas. It can be derived from the sum of squares of all its coefficient points, i.e. The effect of this step is to assign a clear intensity index to the fault signal of each frequency band, so that it is possible to quantitatively compare whether long-term, slowly varying faults are more significant or instantaneous, sudden faults are more severe.
[0125] Next, to capture the waveform characteristics of the fault signal, it is necessary to calculate the kurtosis of the slowly varying and abruptly changing components. Kurtosis is a statistical indicator that measures the sharpness or impact of the signal amplitude distribution. A signal with a high kurtosis value typically contains numerous spikes or outliers, which often corresponds to impactful events in fault diagnosis, such as electric arcs or transient contact failures. Signals with a near-Gaussian distribution have lower kurtosis values. In practice, the kurtosis value K is calculated for the approximation coefficients and each set of detail coefficients. Kurtosis measures the peak-ratio of the signal distribution relative to a normal distribution and can help distinguish between impactful faults and non-impactful disturbances. This step adds a dimension describing the waveform morphology to the fault feature vector, allowing even two fault signals with similar energy to be effectively distinguished by their difference in impact.
[0126] To assess the complexity or uncertainty of the fault signal, it's necessary to calculate the signal entropy values of the slowly varying and abruptly changing components. Signal entropy, derived from information theory, measures the amount of information contained in a signal or its inherent degree of disorder. A highly regular, predictable signal (such as a periodic oscillation) has a low entropy value, while a chaotic, highly random signal has a high entropy value. In practice, for each set of coefficient components, its probability distribution is first calculated, and then the formula for Shannon entropy is used. To calculate its entropy value, where It represents the probability of a signal taking a specific value. The benefit of this step is that it provides a quantitative description of the inherent structure of the fault signal, helps to distinguish between deterministic fault modes and random noise interference, and adds a criterion regarding the regularity of the signal for fault identification.
[0127] After calculating the raw features such as signal energy, kurtosis, and entropy, normalization is necessary to eliminate differences in dimensions and numerical ranges between different features. The principle behind this step is that the originally calculated energy value may be very large, while kurtosis and entropy values are usually within a smaller range. Without processing, the larger features will dominate in subsequent pattern matching, overshadowing the information from other features. A common implementation method is min-max normalization, which normalizes each feature value X using the formula... Mapped to the interval between 0 and 1. Wherein and These are the minimum and maximum values of the feature observed across a large number of historical samples. This step ensures that all features are on the same scale, guaranteeing they have equal weight when constructing the final fault vector, thereby improving the fairness and accuracy of subsequent classification and matching algorithms.
[0128] Finally, all normalized feature values are arranged in a pre-defined order to form a single, fixed-length, multi-dimensional fault feature vector. The principle behind this step is to integrate scattered, independent feature indicators into a structured data entity, making it a standard input for pattern recognition algorithms. In practice, a strict arrangement order is defined; for example, the energy, kurtosis, and entropy values of slowly varying components are placed first, followed by the energy, kurtosis, and entropy values of the first-level abruptly changing components, then the second level, and so on, until all component feature values are included. The final result is a vector of the form shown below. The ultimate result of this approach is the creation of a highly condensed digital fingerprint of the fault. This vector comprehensively describes the combined characteristics of the fault event in terms of intensity, form, and complexity in a compact and standardized manner, laying the foundation for accurate and automated fault diagnosis.
[0129] In one implementation, multiple photovoltaic strings under the same inverter are considered as one analysis unit. The dynamic time warping algorithm is used to measure the morphological similarity between the performance deviation characteristic spectra of each photovoltaic string within the analysis unit. Based on the morphological similarity, the method for determining whether the fault anomaly of the photovoltaic string is an independent fault or a coordinated fault includes the following steps:
[0130] Taking multiple photovoltaic strings under the same inverter as an analysis unit, the performance deviation characteristic spectrum of all photovoltaic strings under the analysis unit is obtained;
[0131] Pair the performance deviation characteristic spectra of any two photovoltaic strings within the analysis unit;
[0132] The dynamic time warping algorithm is used to calculate the time warping distance between each pair of performance deviation feature spectra. The time warping distance characterizes the morphological similarity between the two performance deviation feature spectra.
[0133] Based on the time warped distance of all pairs, a similarity matrix is constructed to characterize the similarity relationship between all photovoltaic strings within the analysis unit;
[0134] Based on the clustering characteristics of the similarity matrix, it can be determined whether the fault or anomaly of the photovoltaic string is an independent fault or a collaborative fault.
[0135] In this implementation, to effectively distinguish the root cause of the fault, the focus of the analysis needs to be shifted from individual photovoltaic strings to a logical unit with common electrical connections. The principle behind this step is that all photovoltaic strings connected to the same inverter constitute a natural common-cause failure analysis group, because a fault in the inverter itself or external factors affecting the inverter region (such as large cloud shadows) will simultaneously affect all its subordinate strings. Specifically, an inverter is used as the core, and all N photovoltaic strings connected to it are defined as an analysis unit. Then, the performance deviation characteristic spectrum of each string in this unit is obtained within the same time window, thus obtaining a set of time-series data. This approach effectively defines a clear scope for subsequent correlation analysis, shifting the focus from judging the quality of individual strings to determining whether a group of closely related strings exhibits consistent behavior.
[0136] After determining the analysis unit, in order to comprehensively explore the relationships between the various strings within the unit, it is necessary to perform pairwise comparisons. The principle behind this step is that only by exhausting all possible pairings can a complete relationship network be constructed, thus ensuring that no potential cooperative behavior patterns are missed. The implementation method is very straightforward: within an analysis unit containing N strings, systematically generate all unique string pairs. For example, pair the first string with the remaining N-1 strings, then pair the second string with the remaining N-2 strings, and so on, until all unique combinations are created. The final result of this process is the generation of a network containing... The list of tasks, each of which is a preparatory step for calculating the similarity of a pair of specific performance deviation feature spectra, ensures the completeness of subsequent analysis.
[0137] Next, Dynamic Time Warping (DTW) is used to accurately measure the morphological similarity of each pair of performance deviation feature spectra. The principle behind this step is that DTW is a powerful algorithm specifically designed for calculating the similarity between two time series. It can effectively handle local distortions, stretching, or compression of the sequences along the time axis, which perfectly matches the scenario where fault manifestations caused by factors such as moving cloud shadows show slight time differences across different sequences. In specific implementation, for any pair of performance deviation spectra... and The DTW algorithm calculates the minimum cumulative distance between two points by constructing a cost matrix and finding an optimal regularized path. This distance value is the time regularized distance. Its calculation can be expressed as: This step quantifies the morphological similarity between two complex waveforms into a single, intuitive numerical value. The smaller the distance value, the more similar the performance deviation patterns of the two strings are.
[0138] After calculating the time-normalized distances of all pairs within a cell, these discrete distance values need to be integrated into a unified structure for macroscopic pattern analysis. This step utilizes matrices as a mathematical tool to organize and present the pairwise similarities between all strings in a compact and visual manner. Specifically, an N×N square matrix, called the similarity matrix M, is constructed. The element in the i-th row and j-th column of the matrix... The value is assigned to the time-warped distance between string i and string j. Since the distance between a string and itself is 0, and the distance between strings i and j is equal to the distance between j and i, all diagonal elements of the matrix are 0, and the matrix is symmetric along its diagonal. The effect of this approach is to create a comprehensive view that reflects the correlation between all strings within the analysis unit, providing a structured data foundation for the final fault nature determination.
[0139] Finally, based on the clustering characteristics presented by the similarity matrix, a final judgment is made on the nature of the fault. The principle behind this step is that cooperative faults cause most strings within the analysis unit to exhibit highly similar performance deviations, which will be reflected in the similarity matrix as a large cluster or dense block composed of many extremely small distance values. Conversely, independent faults only affect a single string, whose deviation spectrum is vastly different from all other strings, resulting in its corresponding rows and columns in the matrix being composed of large distance values, making clustering impossible. In practice, the numerical distribution of matrix M can be analyzed, or clustering algorithms can be applied to identify whether a sufficiently large, dense cluster exists. If a cluster containing multiple strings is found, and the average time-normalized distance between its members is much smaller than their distance from members outside the cluster, it is determined to be a cooperative fault. Conversely, if only a few outlier strings are found, it is determined to be an independent fault. This step effectively distinguishes the fault source, pointing the way for subsequent maintenance and localization.
[0140] In one implementation, determining whether a fault in a photovoltaic string is an independent or coordinated fault based on the clustering characteristics of the similarity matrix includes the following steps:
[0141] Preset time-normalization distance threshold and collaborative quantity ratio threshold;
[0142] Traverse the similarity matrix and count the number of photovoltaic strings whose morphological similarity to any photovoltaic string is higher than the time warp distance threshold;
[0143] Determine whether the number of photovoltaic strings exceeds the threshold for the proportion of coordinated units;
[0144] If the number of photovoltaic strings exceeds the threshold for the proportion of coordinated units, the fault or abnormality occurring in the corresponding photovoltaic string will be judged as a coordinated fault.
[0145] If the number of photovoltaic strings does not exceed the threshold for the proportion of coordinated quantities, the fault or abnormality occurring in the corresponding photovoltaic string will be judged as an independent fault.
[0146] In this implementation, firstly, in order to make black-and-white judgments from the numerical similarity matrix, a clear and quantifiable set of decision rules must be established. The principle behind this step is to transform subjective judgments (e.g., whether something looks like a cluster) into objective and repeatable algorithmic logic. Specifically, two key threshold parameters need to be pre-set: one is the time warping distance threshold. It defines two upper limits for the morphological similarity of performance deviation spectra, i.e., similarity is defined as a distance less than a certain value; the other is the threshold for the proportion of cooperative quantities. It defines what percentage of strings must exhibit similar behavior to be classified as a collaborative event; for example, 0.75 represents 75%. These two thresholds are typically determined based on historical data analysis and expert experience. The effect of this is to provide a specific, unambiguous benchmark for subsequent automated judgments, enabling machines to simulate expert decision-making.
[0147] Next, based on the established criteria, an individualized correlation assessment needs to be performed on each photovoltaic string within the analysis unit. The principle behind this step is to systematically examine the similarity matrix and count the number of similar partners for each string. In practice, the similarity matrix M is traversed row by row. For the i-th row, representing the i-th photovoltaic string, the program checks all elements in that row except for the diagonal elements. When a certain distance value When the distance is less than the preset time warping distance threshold, it means that the behavior of the j-th string is similar to that of the i-th string. This condition is statistically satisfied. By counting the number of substrings j, we can obtain the number of collaborating partners for the i-th substring. The effect of this process is to liberate each substring from complex network relationships and give it a simple and clear indicator: the degree of collaboration.
[0148] After obtaining the number of cooperating partners for each string, it is necessary to determine whether this number reaches the scale required to constitute a cooperative event. The principle behind this step is to compare the local partner count with a global criterion to determine whether the string belongs to a large group or is merely a small-scale coincidence. In practice, the number of cooperating partners for each string obtained in the previous step is compared with an absolute quantity standard calculated based on a cooperative quantity ratio threshold. If an analysis unit has a total of N photovoltaic strings, then the threshold for determining the cooperative quantity is... Subsequently, for each string i, a logical judgment is performed: whether the number of its cooperating partners is strictly greater than the judgment threshold. The effect of this step is to assign a temporary logical label to each string experiencing an anomaly, that is, whether its behavior belongs to a large-scale collective behavior, thus paving the way for the final qualitative classification.
[0149] Subsequently, based on the above judgment result being true (i.e., the number exceeds the threshold), a clear diagnosis of the fault nature is made. The logic of this step is that if the performance deviation of a string is highly similar to the performance deviation of the vast majority of other strings within the analysis unit, then the root cause of this deviation is highly unlikely to be an individual problem of that string itself, but more likely a common factor affecting the entire unit, such as inverter failure or large-area cloud cover. Therefore, in implementation, as long as the number of cooperating partners of a string i is determined to exceed the cooperating quantity threshold, the fault anomaly occurring therein will be automatically and definitively marked as a cooperating fault by the system. The ultimate effect of this is that for faults caused by common reasons, the diagnostic scope can be quickly elevated from the specific string to a higher level of equipment or area, thereby guiding maintenance personnel to directly troubleshoot common fault points.
[0150] Finally, as another aspect of the logic, supplementary diagnosis is made for cases where the judgment result is false (i.e., the number does not exceed the threshold). The principle of this step is the process of elimination: if the performance deviation behavior of a string is abnormal, but the number of partners with similar behavior is very small and fails to meet the judgment criteria for a collaborative event, then its cause of failure is likely endogenous and localized. In specific implementation, for any string where the number of collaborative partners does not exceed the collaborative quantity threshold, the system will automatically classify its abnormal failure as an independent failure. The effect of this is to complete the classification of all abnormal situations, providing precise directions for failures caused by their own reasons (such as component damage, line problems, local contamination), guiding maintenance personnel to skip the inspection of common equipment and go directly to the problematic string itself for detailed troubleshooting, thereby achieving a closed loop in diagnosis and maximizing efficiency.
[0151] The present invention also discloses a photovoltaic module positioning system for a photovoltaic power station based on average power change rate, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the photovoltaic module positioning method for a photovoltaic power station based on average power change rate as described in any of the above embodiments.
[0152] The processor can be a central processing unit (CPU). Of course, depending on the actual use, it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc., and this application does not limit it.
[0153] The memory can be an internal storage unit of a computer device, such as a hard disk or RAM, or an external storage device, such as a plug-in hard disk, smart memory card (SMC), secure digital card (SD), or flash memory card (FC) provided on the computer device. Furthermore, the memory can be a combination of internal storage units and external storage devices of a computer device. The memory is used to store computer programs and other programs and data required by the computer device. The memory can also be used to temporarily store data that has been output or will be output. This application does not limit this.
[0154] The present invention also discloses a computer-readable storage medium storing instructions that, when executed by a processor, cause the processor to perform the photovoltaic module positioning method and system method for photovoltaic power plants based on average power change rate described in any of the above embodiments.
[0155] The computer program can be stored in a machine-readable medium. The computer program includes computer program code, which can be in the form of source code, object code, executable file, or certain middleware. The machine-readable medium includes any entity or device capable of carrying computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the machine-readable medium includes, but is not limited to, the above-mentioned components.
[0156] The photovoltaic module positioning method and system method based on average power change rate in the above embodiments are stored in the computer-readable storage medium and loaded and executed on the processor to facilitate the storage and application of the above methods.
[0157] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of protection of this application is limited to these examples; within the framework of this application, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of one or more embodiments of this application as described above, which are not provided in detail for the sake of brevity.
[0158] One or more embodiments in this application are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of this application. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of one or more embodiments in this application should be included within the protection scope of this application.
Claims
1. A method for positioning photovoltaic modules of a photovoltaic power plant based on average power rate of change, characterized in that, The method comprises the following steps: Obtain historical environmental data of the photovoltaic power station and historical output power data of each photovoltaic string in the photovoltaic power station; Identify and mark abnormal data segments caused by known historical fault events in the historical output power data through a health state recognition algorithm, remove the abnormal data segments from the historical output power data, and form a health state data set; Combine the health state data set and the historical environmental data, and train and generate a corresponding power prediction model for each photovoltaic string through a machine learning algorithm, wherein the model input of the power prediction model is real-time environmental data and sky imager cloud image data of the photovoltaic power station; Real-time obtain actual output power of the photovoltaic string, and generate a performance deviation feature spectrum by subtracting the actual output power from the theoretical output power output by the power prediction model; Apply wavelet transform to decompose the performance deviation feature spectrum into multiple coefficient components containing time domain and frequency domain information, and extract a multi-dimensional fault feature vector based on statistical indicators of the coefficient components; Take multiple photovoltaic strings under the same inverter as an analysis unit, measure the morphological similarity between the performance deviation feature spectrums of the photovoltaic strings in the analysis unit by using a dynamic time warping algorithm, and determine whether the fault anomaly of the photovoltaic string is an independent fault or a cooperative fault according to the morphological similarity; If it is an independent fault, match the fault feature vector with a preset fault mode library to determine the fault position and fault type of the fault photovoltaic string; If it is a cooperative fault, locate the fault to the area where the inverter corresponding to the photovoltaic string is located.
2. The method for positioning photovoltaic modules of a photovoltaic plant based on the average power variation rate according to claim 1, characterized in that, The step of identifying and marking abnormal data segments caused by known historical fault events in the historical output power data through a health state recognition algorithm comprises the following steps: Establish a historical operation and maintenance work order database, which is used to record the occurrence time, end time and corresponding photovoltaic string position of each fault in the photovoltaic power station; Align the historical output power data with the historical operation and maintenance work order database by time stamp; For any photovoltaic string, if the historical output power data of the photovoltaic string in a certain time period coincides with the fault time period of the string recorded in the database; All historical output power data points in the coincident time period are automatically identified and marked as abnormal data segments.
3. The method for positioning photovoltaic modules of a photovoltaic plant based on the average rate of power variation according to claim 2, characterized by the fact that, The method further comprises the following steps after the step of combining the health state data set and the historical environmental data, training and generating a corresponding power prediction model for each photovoltaic string through a machine learning algorithm: Real-time capture continuous frames of full-sky cloud images through a sky imager deployed in the photovoltaic power station; Analyze the full-sky cloud images by using an image analysis method to generate cloud condition prediction data of the photovoltaic power station; Fuse the cloud condition prediction data with real-time environmental data obtained through a sensor network deployed in the photovoltaic power station to obtain environmental fusion data; Use the environmental fusion data as real-time input of the power prediction model, so that the theoretical output power output by the power prediction model has been dynamically compensated for the future cloud layer shading influence in advance.
4. The method for positioning photovoltaic modules of a photovoltaic plant based on the average rate of power variation according to claim 3, characterized by the fact that, The step of analyzing the full-sky cloud images by using an image analysis method to generate cloud condition prediction data of the photovoltaic power station comprises the following steps: Image segmentation is performed on the full-sky all-sky images of continuous frames to identify the outline and range of the cloud region above the photovoltaic power station; The centroid position, area and pixel gray mean value of the cloud region in each frame of the full-sky all-sky image are calculated, wherein the pixel gray mean value is used to represent the optical thickness of the cloud layer; By comparing the centroid position changes of each cloud cluster in the continuous frame full-sky all-sky image, the cloud movement vector field covering the full-sky all-sky image is calculated and generated; The cloud movement vector field is projected in space with the physical topology model of the photovoltaic power station to predict each cloud cluster that will cover the photovoltaic string at a future preset time point; The irradiation shielding coefficient of each photovoltaic string is calculated as the cloud condition prediction data by combining the optical thickness of the cloud layer and the centroid position change speed.
5. The method for positioning photovoltaic modules of a photovoltaic plant based on the average power variation rate according to claim 1, characterized in that, The application wavelet transform decomposes the performance deviation feature spectrum into a plurality of coefficient components containing time domain and frequency domain information, and extracts a multi-dimensional fault feature vector based on statistical indicators of the coefficient components, including the following steps: Cutting the performance deviation feature spectrum of a preset time length; Setting the wavelet basis function and the decomposition level of the wavelet transform; Performing discrete wavelet transform on the cut performance deviation feature spectrum based on the wavelet basis function and the decomposition level, and decomposing the cut performance deviation feature spectrum into a group of approximation coefficients and a plurality of groups of detail coefficients; The approximation coefficients are used as slowly varying components reflecting the low frequency trend of the performance deviation feature spectrum; The detail coefficients are used as sudden change components reflecting the high frequency disturbance of the performance deviation feature spectrum; Based on the slowly varying components and the sudden change components, a multi-dimensional fault feature vector is constructed.
6. The method for positioning photovoltaic modules of a photovoltaic plant based on the average rate of power variation according to claim 5, characterized by the fact that, The construction of the multi-dimensional fault feature vector based on the slowly varying components and the sudden change components includes the following steps: Respectively calculating the signal energy of the slowly varying components and the sudden change components in their respective frequency bands; Respectively calculating the signal kurtosis of the slowly varying components and the sudden change components; Respectively calculating the signal entropy of the slowly varying components and the sudden change components; Normalizing the calculated signal energy, signal kurtosis and signal entropy; Arranging the normalized signal energy, signal kurtosis and signal entropy in a predetermined order to form a multi-dimensional fault feature vector.
7. The method for positioning photovoltaic modules of a photovoltaic plant based on the average power variation rate according to claim 1, characterized by the fact that, The plurality of photovoltaic strings under the same inverter are taken as an analysis unit, the dynamic time warping algorithm is used to measure the morphological similarity between the performance deviation feature spectra of the photovoltaic strings in the analysis unit, and whether the fault anomaly of the photovoltaic strings is an independent fault or a cooperative fault is determined according to the morphological similarity, including the following steps: Taking the plurality of photovoltaic strings under the same inverter as an analysis unit, obtaining the performance deviation feature spectra of all photovoltaic strings under the analysis unit; Pairing the performance deviation feature spectra of any two photovoltaic strings in the analysis unit; Using the dynamic time warping algorithm to calculate the time warping distance between each pair of performance deviation feature spectra, which represents the morphological similarity of the two performance deviation feature spectra; Based on all the paired time warping distances, a similarity matrix representing the similarity relationship between all photovoltaic strings in the analysis unit is constructed; According to the clustering characteristics of the similarity matrix, whether the fault anomaly of the photovoltaic strings is an independent fault or a cooperative fault is determined.
8. The method for positioning photovoltaic modules of a photovoltaic plant based on the average rate of power variation according to claim 7, characterized by the fact that, The determination of whether the fault anomaly of the photovoltaic strings is an independent fault or a cooperative fault according to the clustering characteristics of the similarity matrix includes the following steps: The preset time regularized distance threshold and the cooperative number proportion threshold; traversing the similarity matrix to count the number of photovoltaic strings with a similarity higher than the time regularized distance threshold; judging whether the number of photovoltaic strings exceeds the cooperative number proportion threshold; if the number of photovoltaic strings exceeds the cooperative number proportion threshold, determining that the fault abnormality of the corresponding photovoltaic string is a cooperative fault; if the number of photovoltaic strings does not exceed the cooperative number proportion threshold, determining that the fault abnormality of the corresponding photovoltaic string is an independent fault.
9. A photovoltaic plant photovoltaic module positioning system based on average power rate of change, comprising a memory, a processor and a computer program stored on the memory and executable on the processor, characterized in that, The processor implements the photovoltaic module positioning method based on average power change rate of a photovoltaic power station according to any one of claims 1 to 8 when executing the computer program.
10. A computer-readable storage medium having stored thereon instructions, the computer-readable storage medium comprising: The instructions cause the processor to be configured to implement the photovoltaic module positioning method based on average power change rate of a photovoltaic power station according to any one of claims 1 to 8 when executed by the processor.
Citation Information
Patent Citations
Regional photovoltaic power generation power prediction method and device, equipment and medium
CN117709482A
Vulnerability assessment method for electric power system containing wind power
CN118278812A