Distributed photovoltaic equipment detection method and system based on big data analysis
Through big data analysis methods, the fault location problems caused by the dispersion of photovoltaic power station equipment are solved, and fast and accurate fault diagnosis and prevention are achieved, which improves equipment management efficiency.
Patent Information
- Application Number
- CN202510732987.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-04
AI Technical Summary
The existing photovoltaic power stations have a large number of equipment, a large number of components and a scattered deployment, which makes it difficult for manual operation and maintenance to achieve accurate positioning of faults, affecting equipment efficiency and reliability.
A distributed photovoltaic equipment detection method based on big data analysis is adopted, and through multi-source data acquisition, cleaning and standardized processing, combined with LSTM prediction model and multi-model fusion diagnosis, a fault knowledge base is built to realize real-time detection and abnormal alarms.
Improves the accuracy and efficiency of fault diagnosis, enables rapid locating of potential problems, and provides data support for future fault prevention and maintenance.
Smart Images

Figure CN120263102A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of electromechanical equipment, and particularly to a distributed photovoltaic equipment detection method and system based on big data analysis. Background Art
[0002] At present, photovoltaic power stations are characterized by a large variety of equipment, a huge number of components, and scattered deployment. Photovoltaic modules are the smallest power generation units of photovoltaic power stations. The unique series-parallel structure of photovoltaic power stations causes a single module failure to result in hundreds of times of power loss. It is very difficult to accurately locate faults relying on manual operation and maintenance.
[0003] In summary, there is a need for a distributed photovoltaic equipment detection method and system based on big data analysis to solve the deficiencies in the prior art. Summary of the Invention
[0004] In view of the deficiencies of the prior art, the present invention provides a distributed photovoltaic equipment detection method and system based on big data analysis, aiming to solve the above problems.
[0005] To achieve the above object, the present invention provides the following technical solution: A distributed photovoltaic equipment detection method based on big data analysis, comprising the following steps: Step S1: Multi-source data collection, collecting equipment operation data, status data, power quality data, and environmental data, and transmitting them to the data processing module; Step S2: Data processing, cleaning and standardizing the data, and extracting basic features, derived features, and time series features; Step S3: Using a time series database to store real-time data, managing structured data through a cluster, and combining HDFS and object storage to process unstructured data; implementing a data partitioning strategy based on geographical location, photovoltaic power station scale, and time range; Step S4: Real-time detection and anomaly detection, training an LSTM prediction model based on historical normal data, outputting the expected ranges of power and voltage, and combining quantile regression for dynamic threshold warning; Step S5: Adopting multi-model fusion diagnosis, combining feature indicators and detection methods to give the fault type, and constructing a fault knowledge base.
[0006] Optionally, the equipment operation data in step S1 includes but is not limited to: DC voltage, current, and power of photovoltaic modules; AC voltage, current, power, frequency, efficiency, and operating mode of inverters; Branch current and insulation impedance of busbar collectors.
[0007] Optionally, the status data in step S1 includes but is not limited to component status, inverter status, and bracket status; Component status: hot spot temperature and hidden crack; Inverter status: cooling fan speed, capacitor aging index, and fault code; Bracket status: inclination angle and vibration.
[0008] Optionally, the power quality data and environmental data in step S1 include but are not limited to: Power quality data: harmonic distortion rate, voltage fluctuation, frequency deviation, and three-phase unbalance degree; Environmental data: irradiance, environmental temperature, humidity, and wind speed.
[0009] Optionally, the data cleaning in step S2 is performed in the following ways: Missing value processing: By statistically calculating the missing rate of each field, identifying the missing fields, and filling them through linear interpolation or historical data based on similar dates; Outlier processing: By statistically calculating outliers based on physical rules' hard boundaries, deleting them or replacing them with the moving window mean.
[0010] Optionally, the data normalization in step S2 is performed in the following ways: Associate data from different sources through device ID and timestamp, and perform linear resampling on non-uniformly sampled data.
[0011] Optionally, the basic features, derived features, and time series features in step S2 are: Basic features: directly extract fields, including original sensor data and device metadata; Derived features include inverter conversion efficiency, component performance ratio, and physical relationship features of temperature loss coefficient; Time series features include time dimension features, lag features, and periodic features.
[0012] Optionally, the multi-model fault diagnosis in step S5 is performed in the following ways: Step A1: According to different fault types and characteristic indicators, select appropriate single models for training and testing; Step A2: Integrate the results of multiple single models, adopt the weighted average method and / or voting method, and combine design rules to prioritize or conditionally filter the outputs of different models; Step A3: Set thresholds for each fault type, and combine with the probability values output by the model to determine the fault type.
[0013] A distributed photovoltaic equipment detection system based on big data analysis, adopting the distributed photovoltaic equipment detection method based on big data analysis, includes: a multi-source data acquisition module, a data processing module, a data storage and management module, a real-time detection and anomaly detection module, and a fault diagnosis and knowledge base module; The multi-source data acquisition module is responsible for collecting operation data, status data, power quality data, and environmental data of equipment such as photovoltaic modules, inverters, and busbar boxes, and transmitting the data to the data processing module; The data processing module is used for data cleaning, standardization, and feature extraction; The data storage and management module is used to store real-time data using a time-series database, manage structured data through a cluster, and process unstructured data in combination with HDFS and object storage technologies, and implement a data partitioning strategy according to geographical location, power station scale, and time range; The real-time detection and anomaly detection module is used to dynamically set thresholds for alarm based on the LSTM prediction model and in combination with the quantile regression method to monitor whether the output power and voltage parameters of photovoltaic equipment are within the expected range; The fault diagnosis and knowledge base module is used for fusion diagnosis using multiple models. First, select appropriate single models for training and testing, then fuse the results of multiple models by means of weighted average method or voting method, and finally determine the specific fault type according to the set threshold and update the fault knowledge base.
[0014] Advantages of the present invention: 1. In the present invention, through multi-source data acquisition, including equipment operation data, status data, power quality data, and environmental data, comprehensive working information of the photovoltaic power station can be obtained, providing rich data support for subsequent analysis. By using data cleaning and standardization technologies, the quality and consistency of the data are ensured, which helps to improve the accuracy of subsequent model training and fault diagnosis; 2. In the present invention, a time-series database, a cluster is used to manage structured data, HDFS and object storage are used to process unstructured data, and a data partitioning strategy is implemented, enabling the system to effectively manage massive data, improving data query and analysis efficiency. Based on the LSTM prediction model and the quantile regression method, dynamic threshold alarm is realized, which can timely detect whether the output power and voltage parameters of photovoltaic equipment deviate from the expected range, thus quickly locating potential problems; 3. In the present invention, through the method of multi-model fusion diagnosis, combining feature indicators and detection methods to give the fault type and constructing a fault knowledge base not only improves the accuracy of fault diagnosis, but also provides valuable experience and data support for future fault prevention and maintenance. Description of the Drawings
[0015] Figure 1Schematic diagram of a method flow according to the present invention.
[0016] Figure 2 Internal flowchart of step S5 according to the present invention.
[0017] Figure 3 Schematic diagram of a system structure according to the present invention. Detailed implementation manners
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the invention or the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0019] As Figure 1 shown, a distributed photovoltaic device detection method based on big data analysis includes the following steps: Step S1: Multi-source data collection, collecting device operation data, status data, power quality data, and environmental data, and transmitting them to the data processing module; At the photovoltaic module level: DC sensors are deployed for each string of modules, such as Hall effect current sensors.
[0020] Temperature sensors are installed on the back of key components, such as surface-mounted PT100.
[0021] Data is read through the built-in communication interface (RS485 / CAN) of the inverter.
[0022] A weather station is installed near the photovoltaic array to avoid shadow occlusion.
[0023] Device operation data is transmitted through an intelligent electricity meter or the built-in communication module of the inverter, including but not limited to: DC voltage, current, and power of each group or string of photovoltaic modules; AC voltage, current, power, frequency, efficiency, and operating mode of the inverter; Branch current and insulation impedance of the busbar box; Status data includes but not limited to module status, inverter status, and support status; Module status: hot spot temperature and hidden crack. The hot spot temperature is collected by an infrared sensor, and the hidden crack is detected by an EL detector, which needs to be collected offline regularly; Inverter status: rotational speed of the cooling fan, capacitor aging index, and fault code; Support status: inclination angle and vibration, which are collected by an acceleration sensor; Power quality data and environmental data include but not limited to: Power quality data: harmonic distortion rate, voltage fluctuation, frequency deviation, and three-phase unbalance degree, which are collected by a power quality analyzer or a high-precision electricity meter; Environmental data: irradiance, ambient temperature, humidity, and wind speed.
[0024] Step S2: Data processing, cleaning and normalizing the data, and extracting basic features, derived features, and time series features; Detection methods for missing value processing: Statistical missing rates for each field, such as df.isnull().sum(), to identify high-frequency missing fields.
[0025] Mark continuous missing values caused by sensor failures, such as no data for an entire hour.
[0026] For numerical data, linear interpolation, which is applicable to steadily changing temperature and irradiance.
[0027] Filling based on historical data of similar dates, such as using the power value at the same moment in the previous week.
[0028] For categorical data, such as fault codes: marked as the Unknown category.
[0029] Detection methods for outlier processing: Threshold method: Set a hard boundary according to physical rules, such as the current of a photovoltaic module cannot be negative; Deletion, such as transient noise.
[0030] Replace with the moving window mean, such as for mutated temperature data; Data normalization is achieved through the following methods: Associate data from different sources through device ID and timestamp, and perform linear resampling on non-uniformly sampled data; Basic features, derived features, and time series features are as follows: Basic features directly extract fields, including original sensor data and device metadata; Derived features include physical relationship features such as inverter conversion efficiency, component performance ratio, and temperature loss coefficient; Time series features include time dimension features, lag features, and periodic features; Step S3: Use a time series database to store real-time data, manage structured data through a cluster, and combine HDFS and object storage to process unstructured data; Implement a data partitioning strategy based on geographical location, photovoltaic power station scale, and time range; Step S4: Real-time detection and anomaly detection, train an LSTM prediction model based on historical normal data, output the expected ranges of power and voltage, and combine quantile regression for dynamic threshold warning; Step S5: Perform fusion diagnosis using multiple models, combine feature indicators and detection methods, give the fault type, and construct a fault knowledge base; As Figure 2 shown, multi-model fault diagnosis is carried out in the following ways: Step A1: Select appropriate single models for training and testing according to different fault types and feature indicators; Step A2: Integrate the results of multiple single models, use the weighted average method and / or the voting method, and combine design rules to prioritize or conditionally filter the outputs of different models; Step A3: Set a threshold for each fault type, and determine the fault type in combination with the probability value output by the model.
[0031] As Figure 3 shown, a distributed photovoltaic device detection system based on big data analysis, adopting the above-mentioned distributed photovoltaic device detection method based on big data analysis, includes: a multi-source data acquisition module, a data processing module, a data storage and management module, a real-time detection and anomaly detection module, and a fault diagnosis and knowledge base module; The multi-source data acquisition module is responsible for collecting the operation data, status data, power quality data, and environmental data of devices such as photovoltaic modules, inverters, and busbar boxes, and transmitting the data to the data processing module; The data processing module is used for data cleaning, standardization, and feature extraction; The data storage and management module is used to store real-time data using a time series database, manage structured data through a cluster, and process unstructured data in combination with HDFS and object storage technologies, and implement a data partitioning strategy according to geographical location, power station scale, and time range; The real-time detection and anomaly detection module is used to dynamically set a threshold for alarm based on the LSTM prediction model and in combination with the quantile regression method to monitor whether the output power and voltage parameters of photovoltaic devices are within the expected range; The fault diagnosis and knowledge base module is used to perform fusion diagnosis using multiple models. First, select appropriate single models for training and testing, then fuse the results of multiple models by the weighted average method or the voting method, and finally determine the specific fault type according to the set threshold and update the fault knowledge base.
[0032] Through multi-source data acquisition of the present invention, including device operation data, status data, power quality data, and environmental data, comprehensive working information of a photovoltaic power station can be obtained, providing rich data support for subsequent analysis. By adopting data cleaning and standardization technologies, the quality and consistency of the data are ensured, which helps to improve the accuracy of subsequent model training and fault diagnosis; Use time-series databases, cluster management of structured data, HDFS, and object storage to process unstructured data, and implement a data partitioning strategy so that the system can effectively manage massive amounts of data, improve data query and analysis efficiency, and implement dynamic threshold alarms based on the LSTM prediction model and quantile regression method, enabling timely detection of whether the output power and voltage parameters of photovoltaic devices deviate from the expected range, thus quickly locating potential problems; Through a multi-model fusion diagnosis method, combine feature indicators and detection methods to give the fault type, and build a fault knowledge base, which not only improves the accuracy of fault diagnosis, but also provides valuable experience and data support for future fault prevention and maintenance.
[0033] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, or improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A distributed photovoltaic device detection method based on big data analysis, characterized in that It includes the following steps: Step S1: Multi-source data collection. Collect device operation data, status data, power quality data, and environmental data, and transmit them to the data processing module; Step S2: Data processing. Clean and standardize the data, and extract basic features, derived features, and time series features; Step S3: Use a time series database to store real-time data, manage structured data through a cluster, and combine HDFS and object storage to process unstructured data; Implement a data partitioning strategy based on geographical location, photovoltaic power station scale, and time range; Step S4: Real-time detection and anomaly detection. Train an LSTM prediction model based on historical normal data, output the expected ranges of power and voltage, and combine quantile regression for dynamic threshold warning; Step S5: Adopt multi-model fusion diagnosis. Combine feature indicators and detection methods to give the fault type and build a fault knowledge base.
2. The distributed photovoltaic device detection method based on big data analysis according to claim 1, wherein In the said step S1, the device operation data includes but is not limited to: DC voltage, current, and power of photovoltaic modules; AC voltage, current, power, frequency, efficiency, and operation mode of inverters; Branch current and insulation impedance of busbar boxes.
3. The distributed photovoltaic device detection method based on big data analysis according to claim 1, wherein, In the said step S1, the status data includes but is not limited to component status, inverter status, and bracket status; Component status: Hot spot temperature and hidden crack; Inverter status: Cooling fan speed, capacitor aging index, and fault code; Bracket status: Tilt angle and vibration.
4. The distributed photovoltaic device detection method based on big data analysis according to claim 1, wherein, In the said step S1, the power quality data and environmental data includes but is not limited to: Power quality data: Harmonic distortion rate, voltage fluctuation, frequency deviation, and three-phase unbalance degree; Environmental data: Irradiance, ambient temperature, humidity, and wind speed.
5. The distributed photovoltaic device detection method based on big data analysis according to claim 1, characterized in that In the said step S2, the data cleaning is carried out in the following ways: Missing value processing: By statistically calculating the missing rate of each field, identifying the missing fields, and filling them through linear interpolation or historical data based on similar dates; Outlier processing: Statistically calculate outliers through physical rule hard boundaries, and delete them or replace them with the moving window mean.
6. The distributed photovoltaic equipment detection method based on big data analysis according to claim 1, characterized in that, In the said step S2, the data standardization is carried out in the following ways: Associate data from different sources through device ID and timestamp, and perform linear resampling on non-uniformly sampled data.
7. The distributed photovoltaic device detection method based on big data analysis according to claim 1, wherein, In the said step S2, the basic features, derived features, and time series features are: Basic features directly extract fields, including original sensor data and device metadata; Derived features include physical relationship features such as inverter conversion efficiency, component performance ratio, and temperature loss coefficient; Time series features include time dimension features, lag features, and periodic features.
8. The distributed photovoltaic equipment detection method based on big data analysis according to claim 1, characterized in that In the said step S5, the multi-model fault diagnosis is carried out in the following ways: Step A1: According to different fault types and feature indicators, select appropriate single models for training and testing; Step A2: Integrate the results of multiple single models, adopt the weighted average method and / or voting method, and combine design rules to prioritize or conditionally filter the outputs of different models; Step A3: Set a threshold for each fault type, and combine the probability values output by the model to determine the fault type.
9. A distributed photovoltaic device detection system based on big data analysis, which adopts the distributed photovoltaic device detection method based on big data analysis according to any one of claims 1-8, characterized in that, It includes: Multi-source data collection module, data processing module, data storage and management module, real-time detection and anomaly detection module, and fault diagnosis and knowledge base module; The multi-source data acquisition module is responsible for collecting the operation data, status data, power quality data, and environmental data of photovoltaic modules, inverters, and busbar boxes, and transmitting the data to the data processing module; The data processing module is used for data cleaning, standardization, and feature extraction; The data storage and management module is used to store real-time data using a time series database, manage structured data through a cluster, and process unstructured data by combining HDFS and object storage technologies, and implement a data partitioning strategy based on geographical location, power station scale, and time range; The real-time detection and anomaly detection module is used to dynamically set thresholds for alarms based on the LSTM prediction model and combined with the quantile regression method to monitor whether the output power and voltage parameters of photovoltaic devices are within the expected range; The fault diagnosis and knowledge base module is used to perform fusion diagnosis using multiple models. First, select appropriate single models for training and testing, then fuse the results of multiple models by weighted average method or voting method, and finally determine the specific fault type according to the set threshold and update the fault knowledge base.
Citation Information
Patent Citations
System state evaluation method and device based on photovoltaic power generation data, equipment and medium
CN117171410A
Abnormity detection method based on real-time data of distributed photovoltaic power station
CN118673456A
Remote monitoring and fault diagnosis method for distributed photovoltaic power generation system
CN119483506A
Distributed photovoltaic operation and maintenance data intelligent processing method and system
CN119944940A
Fault determination system, fault determination program, and fault determination method
JP7177238B1
Cited By
Photovoltaic smart station early warning method, device and equipment
CN121117674A
A photovoltaic intelligent station early warning method, device and equipment
CN121117674B
Photovoltaic equipment cluster benchmarking analysis and hidden fault early warning method and system
CN121261330A
A method and system for photovoltaic device cluster benchmarking and latent fault early warning
CN121261330B