Ground centralized photovoltaic module fault diagnosis method and system based on machine learning

Through machine learning-based fault diagnosis methods, multiple sub-fault diagnosis models are constructed using historical sensing data and group learning algorithms, and the problems of low fault diagnosis efficiency and high misjudgment rate of ground centralized photovoltaic modules in the prior art are solved, achieving higher diagnostic accuracy and efficiency.

CN120217145APending Publication Date: 2025-06-27TUNGHSU AZURE RENEWABLE ENERGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510235183.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art has low diagnostic efficiency and high fault misjudgment rate in ground centralized photovoltaic module fault diagnosis, making it difficult to adapt to complex fault types.

Method used

The fault diagnosis method based on machine learning is adopted, and the training data is preprocessed and annotated by obtaining historical sensing data, and a group learning algorithm is used to train multiple sub-fault diagnosis models in parallel, and finally the output results are counted through the voting method to obtain the final fault diagnosis results.

Benefits of technology

It improves the accuracy and efficiency of fault diagnosis of ground centralized photovoltaic modules, reduces the fault misjudgment rate, and is suitable for complex fault diagnosis scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120217145A_ABST
    Figure CN120217145A_ABST
Patent Text Reader

Abstract

The invention discloses a ground centralized photovoltaic module fault diagnosis method and system based on machine learning, and belongs to the technical field of photovoltaic module fault diagnosis, and the method comprises the steps: obtaining historical sensing data stored in a ground centralized photovoltaic module, and carrying out the data preprocessing of the historical sensing data, so as to form a training data set; marking the data in the training data set as normal data and fault data according to the fault occurrence time; performing random sampling on the training data set, screening out normal data and fault data, and dynamically adjusting a data proportion by using a random undersampling method to form a plurality of sub training data sets; carrying out parallel training on the plurality of sub training data sets to obtain a plurality of sub fault diagnosis models; and finally, acquiring real-time sensing data, inputting the real-time sensing data into each sub-fault diagnosis model, and counting an output result to obtain a final fault diagnosis result. Therefore, the fault diagnosis efficiency of the ground centralized photovoltaic module is improved, the final diagnosis effect is effectively improved, and the fault diagnosis misjudgment rate is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of photovoltaic module fault diagnosis and machine learning, and particularly relates to a method for diagnosing faults of ground centralized photovoltaic modules based on machine learning. Background Art

[0002] Ground centralized photovoltaic modules are an important part of centralized photovoltaic power generation systems. They are mainly installed on the ground, converting sunlight into electrical energy through large-scale photovoltaic power generation arrays and directly connecting to the national power grid for long-distance power supply. Among them, the photovoltaic power generation array is composed of multiple photovoltaic panels connected in series and parallel, and is connected in parallel with the power grid through a busbar box, an inverter, and a transformer. Ground centralized photovoltaic power stations are mainly built in areas with rich land resources and good lighting conditions, such as deserts, gobi, and mountains. These areas usually have vast land and sufficient lighting resources, suitable for large-scale photovoltaic power generation. However, these areas often also have harsh environmental conditions, which will have a great impact on the ground centralized photovoltaic modules exposed outdoors, resulting in their failures. Among them, the probability of failure of photovoltaic panels is the highest, for example: hail and gravel hitting the photovoltaic panels cause short circuits and open circuits of the panels, long-term exposure to sunlight or rain erosion causes the power generation components to age, and dust and dirt blocking cause hot spots on the panels.

[0003] For such faults, it is often necessary for staff to conduct inspections and repairs in a timely manner. In recent years, some new fault diagnosis methods have also emerged, such as using infrared images for fault diagnosis. Due to its excessive dependence on the precision of infrared devices, it not only has a high cost, but also can often only diagnose a single type of fault. Therefore, it is not applicable to the fault diagnosis of complex ground centralized photovoltaic modules. Then, some fault diagnosis methods based on deep learning and / or machine learning have emerged, but most of these methods face the problems of poor diagnosis efficiency and high fault misjudgment rate.

[0004] Therefore, there is an urgent need to provide a method for diagnosing faults of ground centralized photovoltaic modules based on machine learning that can effectively improve the diagnosis efficiency and reduce the fault misjudgment rate. Summary of the Invention

[0005] The purpose of the present invention is to provide a method and system for diagnosing faults of ground centralized photovoltaic modules based on machine learning to solve the above problems existing in the prior art.

[0006] To achieve the above purpose, the present invention adopts the following technical solutions:

[0007] In the first aspect, the present invention provides a method for diagnosing faults of ground centralized photovoltaic modules based on machine learning, which includes:

[0008] Obtain the historical sensing data stored in the ground centralized photovoltaic module, and perform data preprocessing on the historical sensing data to form a training data set. Among them, the historical sensing data includes historical data of direct irradiance, diffuse irradiance, temperature, humidity, DC side voltage, DC side current, and output active power;

[0009] According to the fault occurrence time recorded by the ground centralized photovoltaic module, label each historical sensing data in the training data set as normal data and fault data;

[0010] Perform random sampling on the training data set to form multiple pre-sub-training data sets, screen out the normal data and fault data in each pre-sub-training data set, and adjust the ratio of normal data and fault data in each pre-sub-training data set using the random undersampling method according to the preset sample ratio to form multiple sub-training data sets;

[0011] Based on the group learning algorithm, perform fault diagnosis training on multiple sub-training data sets in parallel to obtain multiple sub-fault diagnosis models;

[0012] Obtain the real-time sensing data fed back by the ground centralized photovoltaic module, input the real-time sensing data into each sub-fault diagnosis model, and based on the voting method, statistically analyze the output results of each sub-fault diagnosis model to obtain the final fault diagnosis result.

[0013] In a possible design, the data preprocessing of the historical sensing data includes:

[0014] Calculate the data average value of the historical sensing data;

[0015] Perform data identification on the historical sensing data to screen out missing values and duplicate values in the historical sensing data;

[0016] According to the data average value of the historical sensing data, fill in the missing values in the historical sensing data with the mean value, and delete the duplicate values in the historical sensing data to obtain pre-training data;

[0017] Perform normalization processing on the pre-training data to obtain training data, and form a training data set based on the training data.

[0018] In a possible design, according to the fault occurrence time recorded by the ground centralized photovoltaic module, labeling each historical sensing data in the training data set as normal data and fault data includes:

[0019] Arrange each historical sensing data in the training data set according to the time information to transform each historical sensing data into historical sensing time series data;

[0020] According to the storage records of the ground centralized photovoltaic modules, obtain the occurrence time of historical fault events and the types of fault events;

[0021] According to the occurrence time of historical fault events, perform data annotation on historical sensing time series data. Among them, the historical sensing time series data corresponding to the occurrence time of historical fault events is annotated as fault data, and the type of fault event is correspondingly annotated, and the remaining historical sensing time series data in the training data set is annotated as normal data.

[0022] In a possible design, randomly sample the training data set to form multiple pre-sub-training data sets, including:

[0023] According to the preset sampling parameters, randomly sample historical sensing data from the training data set to form a pre-sub-training data set, and put the randomly sampled historical sensing data back into the training data set, where the preset sampling parameters include the size of the preset pre-sub-training data set;

[0024] Repeat random sampling and putting back according to the preset sampling parameters again to form multiple pre-sub-training data sets;

[0025] Correspondingly, according to the preset sample ratio, use the random undersampling method to adjust the ratio of normal data to fault data in each pre-sub-training data set to form multiple sub-training data sets, including:

[0026] Define the normal data in each pre-sub-training data set as the majority class samples, and define the fault data in each pre-sub-training data set as the minority class samples;

[0027] According to the preset sample ratio, calculate the number of majority class samples to be deleted for each pre-sub-training data set respectively, and randomly sample and delete the majority class samples according to the calculation results;

[0028] Define each pre-sub-training data set after random sampling and deletion of majority class samples as a sub-training data set.

[0029] In a possible design, based on the ensemble learning algorithm, perform fault diagnosis training on multiple sub-training data sets in parallel to obtain multiple sub-fault diagnosis models, including:

[0030] Use the ensemble learning algorithm to allocate multiple sub-training data sets to different computing nodes respectively;

[0031] Independently process the sub-training data sets assigned to them by each different computing node, and independently calculate the sub-fault diagnosis model parameters;

[0032] Optimize the sub-fault diagnosis model parameters through the backpropagation algorithm, and introduce the L2 regularization term to suppress overfitting;

[0033] Generate multiple sub - fault diagnosis models based on the regularized parameters of the sub - fault diagnosis models.

[0034] In a possible design, obtain the real - time sensing data fed back by the ground - mounted centralized photovoltaic module, input the real - time sensing data into each sub - fault diagnosis model, and based on the voting method, count the output results of each sub - fault diagnosis model to obtain the final fault diagnosis result, including:

[0035] Obtain the real - time sensing data fed back from the ground - mounted centralized photovoltaic module and input it into each sub - fault diagnosis model respectively to obtain the output results of each sub - fault diagnosis model;

[0036] Statistically analyze the output results of each sub - fault diagnosis model and draw the probability distribution diagram of each fault type according to the statistical results;

[0037] Based on the probability distribution diagram of each fault type, calculate the occurrence probability values of each fault type respectively;

[0038] Using the voting method, compare the occurrence probability values of each fault type. If the fault type with the highest occurrence probability value is not unique, introduce a dynamic resampling mechanism, re - obtain the historical sensing data stored in the ground - mounted centralized photovoltaic module to form a secondary training dataset and complete secondary training. According to the secondary training, construct multiple secondary sub - fault diagnosis models, and use the voting method again to determine whether the fault type with the highest occurrence probability value is unique until the result is yes. If the fault type with the highest occurrence probability value is unique, use the fault type with the highest occurrence probability value as the final fault diagnosis result.

[0039] In a possible design, after obtaining the final fault diagnosis result, it further includes:

[0040] According to the final fault diagnosis result, combined with the actual topological structure of the ground - mounted centralized photovoltaic module, locate multiple suspected fault points through voltage and current outliers;

[0041] Obtain the infrared image of the ground - mounted centralized photovoltaic module. According to the infrared image of the ground - mounted centralized photovoltaic module, through secondary verification and interference elimination of multiple suspected fault points, obtain the fault location information;

[0042] Send an alarm message through communication, where the alarm message includes the final fault diagnosis result and the fault location information.

[0043] In a second aspect, the present invention provides a ground - mounted centralized photovoltaic module fault diagnosis system based on machine learning, which includes:

[0044] The photovoltaic module data acquisition module is used to obtain the historical sensing data stored in the ground centralized photovoltaic module, and perform data preprocessing on the historical sensing data to form a training data set. Among them, the historical sensing data includes historical data of direct irradiance, diffuse irradiance, temperature, humidity, DC side voltage, DC side current, and output active power;

[0045] The data processing module is used to label each historical sensing data as normal data and fault data according to the fault occurrence time recorded by the ground centralized photovoltaic module; it is also used to randomly sample the training data set to form multiple pre-sub-training data sets, screen out the normal data and fault data in each pre-sub-training data set, and adjust the ratio of normal data and fault data in each pre-sub-training data set by using the random undersampling method according to the preset sample ratio to form multiple sub-training data sets;

[0046] The model training module is used to perform fault diagnosis training on multiple sub-training data sets in parallel based on the group learning algorithm to obtain sub-fault diagnosis models;

[0047] The fault diagnosis module is used to obtain the real-time sensing data fed back by the ground centralized photovoltaic module, input the real-time sensing data into each sub-fault diagnosis model, and statistically obtain the output results of each sub-fault diagnosis model based on the voting method to obtain the final fault diagnosis result.

[0048] In a third aspect, the present invention provides an electronic device, including a memory, a processor, and a transceiver that are communicatively connected in sequence. Among them, the memory is used to store a computer program, the transceiver is used to send and receive messages, and the processor is used to read the computer program and execute the fault diagnosis method for the ground centralized photovoltaic module based on machine learning as described in the first aspect or any possible design of the first aspect;

[0049] In a fourth aspect, the present invention provides a computer program product containing instructions, which when the instructions run on a computer, cause the computer to execute the fault diagnosis method for the ground centralized photovoltaic module based on machine learning as described in the first aspect or any possible design of the first aspect.

[0050] Beneficial effects: The present invention provides a fault diagnosis method and system for ground centralized photovoltaic modules based on machine learning. By obtaining the historical sensing data stored in the ground centralized photovoltaic modules and performing data preprocessing on the historical sensing data to form a training data set; and according to the fault occurrence time recorded by the ground centralized photovoltaic modules, labeling each historical sensing data in the training data set as normal data and fault data; then randomly sampling the training data set to form multiple pre-sub-training data sets, screening out the normal data and fault data in each pre-sub-training data set, and adjusting the ratio of normal data and fault data in each pre-sub-training data set by using the random undersampling method according to a preset sample ratio to form multiple sub-training data sets; then based on the ensemble learning algorithm, performing fault diagnosis training on the multiple sub-training data sets in parallel to obtain multiple sub-fault diagnosis models; finally, obtaining the real-time sensing data fed back by the ground centralized photovoltaic modules, inputting the real-time sensing data into each sub-fault diagnosis model, and statistically obtaining the output results of each sub-fault diagnosis model based on the voting method to obtain the final fault diagnosis result. This diagnosis method constructs multiple sub-fault diagnosis models through parallel training of multiple sub-training data sets, greatly improving the fault diagnosis accuracy and diagnosis efficiency for ground centralized photovoltaic modules, and dynamically adjusting the data ratio during the training process to prevent misjudgment of the final diagnosis result due to unbalanced training data, effectively improving the final diagnosis effect and reducing the fault diagnosis misjudgment rate. Description of the Drawings

[0051] Figure 1 It is a flowchart of the fault diagnosis method for ground centralized photovoltaic modules based on machine learning in the embodiment of the present invention. Detailed Embodiments

[0052] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the present invention in combination with the drawings and the descriptions of the embodiments or the prior art. Obviously, the following descriptions of the structures of the drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. It should be noted here that the descriptions of these embodiment modes are used to help understand the present invention, but do not constitute a limitation to the present invention.

[0053] It should be understood that although terms such as first and second may be used herein to describe various units, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, the first unit can be called the second unit, and similarly, the second unit can be called the first unit, without departing from the scope of the exemplary embodiments of the present invention.

[0054] It should be understood that for the term "and / or" that may appear in this text, it is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, B exists alone, and both A and B exist simultaneously. For the term " / and" that may appear in this text, it describes another association object relationship, indicating that there can be two relationships. For example, A / and B can represent: A exists alone, and both A and B exist. Additionally, for the character " / " that may appear in this text, it generally indicates that the associated objects before and after are in an "or" relationship.

[0055] Embodiment:

[0056] As Figure 1 shown, this embodiment provides a method for fault diagnosis of a ground centralized photovoltaic module based on machine learning, which includes:

[0057] S100. Obtain the historical sensing data stored in the ground centralized photovoltaic module, and perform data preprocessing on the historical sensing data to form a training data set. Among them, the historical sensing data includes historical data of direct irradiance, diffuse irradiance, temperature, humidity, DC side voltage, DC side current, and output active power;

[0058] In a possible implementation manner, in step S100, performing data preprocessing on the historical sensing data includes:

[0059] S1001. Calculate the data average value of the historical sensing data;

[0060] S1002. Perform data identification on the historical sensing data to screen out missing values and duplicate values in the historical sensing data;

[0061] S1003. Fill in the missing values in the historical sensing data with the average value according to the data average value of the historical sensing data, and delete the duplicate values in the historical sensing data to obtain pre-training data;

[0062] S1004. Perform normalization processing on the pre-training data to obtain training data, and form a training data set based on the training data.

[0063] Among them, ground-mounted centralized photovoltaic modules generally have corresponding data acquisition devices, which can collect data related to photovoltaic power generation, such as light intensity, temperature, air pressure, air humidity, DC-side voltage and current of each photovoltaic string, inverter-side voltage and current, etc., and store them as historical sensing data through corresponding data storage devices. Data preprocessing of historical sensing data is to avoid the influence of possible missing values and duplicate values in the original historical sensing data on subsequent training. Therefore, it is necessary to clean the data in advance. And since there are differences in the dimensions of the cleaned data, it is necessary to eliminate the dimensional differences of the data by means of normalization processing.

[0064] S200. According to the fault occurrence time recorded by the ground-mounted centralized photovoltaic module, label each historical sensing data in the training data set as normal data and fault data;

[0065] In a possible implementation manner, in step S200, according to the fault occurrence time recorded by the ground-mounted centralized photovoltaic module, labeling each historical sensing data in the training data set as normal data and fault data includes:

[0066] S2001. Arrange each historical sensing data in the training data set according to the time information to transform each historical sensing data into historical sensing time-series data;

[0067] S2002. Obtain the occurrence time and fault event type of the historical fault event according to the storage record of the ground-mounted centralized photovoltaic module;

[0068] S2003. Perform data labeling on the historical sensing time-series data according to the occurrence time of the historical fault event. Among them, label the historical sensing time-series data corresponding to the occurrence time of the historical fault event as fault data, and correspondingly label the fault event type, and label the remaining historical sensing time-series data in the training data set as normal data.

[0069] Among them, the historical data stored in the ground-mounted centralized photovoltaic module all have time attributes and carry time information, and the occurrence time of the historical fault event is a time interval. The historical sensing data (the historical sensing time-series data corresponding to the occurrence time of the historical fault event) within this time interval are all considered as fault data. Correspondingly, labeling the fault event type for these fault data is for the subsequent model training to accurately diagnose various different fault types.

[0070] S300. Randomly sample the training data set to form multiple pre-sub-training data sets, screen out the normal data and fault data in each pre-sub-training data set, and adjust the ratio of normal data and fault data in each pre-sub-training data set by using the random undersampling method according to the preset sample ratio to form multiple sub-training data sets;

[0071] In a possible implementation, in step S300, the training data set is randomly sampled to form a plurality of pre-sub-training data sets, including:

[0072] S3001. According to the preset sampling parameters, historical sensing data is randomly sampled from the training data set to form a pre-sub-training data set, and the randomly sampled historical sensing data is put back into the training data set, where the preset sampling parameters include the size of the preset pre-sub-training data set;

[0073] S3002. Random sampling and putting back are repeatedly performed multiple times according to the preset sampling parameters to form a plurality of pre-sub-training data sets;

[0074] Correspondingly, according to the preset sample ratio, the random undersampling method is used to adjust the ratio of normal data to fault data in each pre-sub-training data set to form a plurality of sub-training data sets, including:

[0075] S3003. Define the normal data in each pre-sub-training data set as the majority class samples, and define the fault data in each pre-sub-training data set as the minority class samples;

[0076] S3004. According to the preset sample ratio, calculate the number of majority class samples to be deleted for each pre-sub-training data set respectively, and randomly sample and delete the majority class samples according to the calculation results;

[0077] S3005. Define each pre-sub-training data set after random sampling and deletion of the majority class samples as a sub-training data set.

[0078] Among them, the training data set is randomly sampled with replacement to form multiple pre-sub-training data sets according to the preset sampling parameters. And according to actual needs, the number of pre-sub-training data sets is set to at least 51. Each normal data in the pre-sub-training data set is defined as the majority class sample, and each faulty data in the pre-sub-training data set is defined as the minority class sample. Then, according to the preset sample ratio, random sampling and deletion are performed on the majority class samples according to the calculation results. This is because the ratio of faulty data to normal data in the pre-sub-training data set is seriously unbalanced, and its data distribution has obvious class imbalance characteristics. Samples in the normal operating state (normal data) always account for the vast majority of the total data, while the number of samples in various faulty states (faulty data) is very small. Therefore, if the pre-sub-training data set is directly used to train the fault diagnosis model, to a large extent, it will prompt the model to over-learn the characteristics of normal data, so that the constructed fault diagnosis model will regard some faulty data as normal data. When using this model for diagnosis, it is easy to misjudge "fault" as "normal". Therefore, the method of random undersampling is used to adjust the ratio of faulty data to normal data and reach the preset ratio, which can well solve this problem, greatly improve its diagnosis accuracy, and reduce the fault misjudgment rate. To achieve this goal, in actual training, the preset ratio can be set to: the amount of faulty data: the amount of normal data = 1:1 to ensure that the trained model can achieve the goal of reducing the fault misjudgment rate.

[0079] S400. Based on the ensemble learning algorithm, perform fault diagnosis training on multiple sub-training data sets in parallel to obtain multiple sub-fault diagnosis models;

[0080] In a possible implementation manner, in step S400, based on the ensemble learning algorithm, performing fault diagnosis training on multiple sub-training data sets in parallel to obtain multiple sub-fault diagnosis models includes:

[0081] S4001. Use the ensemble learning algorithm to allocate multiple sub-training data sets to different computing nodes respectively;

[0082] S4002. Independently process the sub-training data sets assigned to each by different computing nodes, and independently calculate the sub-fault diagnosis model parameters;

[0083] S4003. Optimize the sub-fault diagnosis model parameters through the backpropagation algorithm, and introduce the L2 regularization term to suppress overfitting;

[0084] S4004. Generate multiple sub-fault diagnosis models based on the regularized sub-fault diagnosis model parameters.

[0085] Among them, the group learning algorithm is a method of improving the overall prediction performance by combining multiple learning models. The basic idea of this method is to use multiple weak learners for training and integrate the prediction results of these weak learners through a method such as voting, so as to obtain a more stable and accurate prediction. The weak learners mentioned here are actually the sub-fault diagnosis models constructed in this embodiment. Through parallel data processing by multiple sub-fault diagnosis models, they can capture different features in the data, thereby improving the accuracy of the overall diagnosis and greatly improving the diagnosis speed.

[0086] It should be noted that during the training of the sub-fault diagnosis model, the model over-learns the details and noises in the training data, so that it cannot generalize well to unseen data, that is, overfitting occurs. The overfitting model may learn each data in the sub-training dataset too precisely, so that during actual diagnosis, as long as the data that appears is slightly different, correct diagnosis cannot be performed. Therefore, it is necessary to optimize the parameters of the sub-fault diagnosis model through the backpropagation algorithm and introduce the L2 regularization term to suppress the overfitting phenomenon.

[0087] S500. Obtain the real-time sensing data fed back by the ground centralized photovoltaic module, input the real-time sensing data into each sub-fault diagnosis model, and based on the voting method, count the output results of each sub-fault diagnosis model to obtain the final fault diagnosis result.

[0088] In a possible implementation manner, in step S500, obtaining the real-time sensing data fed back by the ground centralized photovoltaic module, inputting the real-time sensing data into each sub-fault diagnosis model, and based on the voting method, counting the output results of each sub-fault diagnosis model to obtain the final fault diagnosis result includes:

[0089] S5001. Obtain the real-time sensing data fed back from the ground centralized photovoltaic module and input it into each sub-fault diagnosis model respectively to obtain the output results of each sub-fault diagnosis model;

[0090] S5002. Statistically analyze the output results of each sub-fault diagnosis model and draw a probability distribution diagram of each fault type according to the statistical results;

[0091] S5003. Based on the probability distribution diagram of each fault type, calculate the occurrence probability value of each fault type respectively;

[0092] S5004. Using the voting method, compare the occurrence probability values of each fault type. If the fault type with the highest occurrence probability value is not unique, introduce a dynamic resampling mechanism to re-obtain the historical sensing data stored in the ground centralized photovoltaic module to form a secondary training dataset and complete secondary training. Construct multiple secondary sub-fault diagnosis models based on the secondary training, and use the voting method again to determine whether the fault type with the highest occurrence probability value is unique until the result is yes. If the fault type with the highest occurrence probability value is unique, use the fault type with the highest occurrence probability value as the final fault diagnosis result.

[0093] Among them, the voting method actually statistically analyzes the output results of each sub-fault diagnosis model and sums the occurrence probabilities of each fault type corresponding to the results according to the type to obtain the occurrence probability values of each fault type, and compares them to select the fault type with the highest probability value as the final fault diagnosis result. If the fault type with the highest probability value is not unique, a dynamic resampling mechanism can be introduced to achieve secondary training and finally obtain a unique fault type as the final fault diagnosis result for output.

[0094] In a possible implementation manner, after obtaining the final fault diagnosis result, it further includes:

[0095] According to the final fault diagnosis result, combined with the actual topological structure of the ground centralized photovoltaic module, locate multiple suspected fault points through voltage and current outliers;

[0096] Obtain the infrared image of the ground centralized photovoltaic module. According to the infrared image of the ground centralized photovoltaic module, through secondary verification and interference elimination of multiple suspected fault points, obtain the fault location information;

[0097] Send an alarm message through communication, where the alarm message includes the final fault diagnosis result and the fault location information.

[0098] Among them, the suspected fault points are judged by combining the fault type with the actual topological structure of the ground centralized photovoltaic module, and the infrared image of the ground centralized photovoltaic module can be obtained through an external infrared imaging device.

[0099] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for fault diagnosis of ground-based centralized photovoltaic modules based on machine learning, characterized in that: include: Acquire historical sensor data stored in the ground centralized photovoltaic module, and perform data preprocessing on the historical sensor data to form a training data set, wherein the historical sensor data includes historical data of direct irradiance, diffuse irradiance, temperature, humidity, DC side voltage, DC side current and output active power; According to the fault occurrence time recorded by the ground-based centralized photovoltaic modules, each historical sensor data in the training data set is labeled as normal data and fault data; Randomly sampling the training data set to form multiple pre-sub-training data sets, screening out normal data and fault data in each pre-sub-training data set, and adjusting the ratio of normal data to fault data in each pre-sub-training data set by using a random under-sampling method according to a preset sample ratio to form multiple sub-training data sets; Based on the group learning algorithm, multiple sub-training data sets are trained in parallel for fault diagnosis to obtain multiple sub-fault diagnosis models; The real-time sensor data fed back by the ground centralized photovoltaic modules is obtained, the real-time sensor data is input into each sub-fault diagnosis model, and the output results of each sub-fault diagnosis model are counted based on the voting method to obtain the final fault diagnosis result.

2. The method for diagnosing faults of ground-based centralized photovoltaic modules based on machine learning according to claim 1, characterized in that: Perform data preprocessing on historical sensor data, including: Calculating a data average of the historical sensor data; Performing data recognition on the historical sensor data to screen out missing values ​​and duplicate values ​​in the historical sensor data; According to the data average value of the historical sensor data, the missing values ​​in the historical sensor data are mean-filled, and the duplicate values ​​in the historical sensor data are deleted to obtain pre-training data; The pre-training data is normalized to obtain training data, and a training data set is formed based on the training data.

3. The method for diagnosing faults of ground-based centralized photovoltaic modules based on machine learning according to claim 1, characterized in that: According to the fault occurrence time recorded by the ground-based centralized photovoltaic module, each historical sensor data in the training data set is labeled as normal data and fault data, including: Arrange each historical sensor data in the training data set according to time information to transform each historical sensor data into historical sensor time series data; According to the storage records of the ground-based centralized photovoltaic modules, the occurrence time and type of historical fault events are obtained; According to the occurrence time of historical fault events, the historical sensor time series data are labeled, wherein the historical sensor time series data corresponding to the occurrence time of the historical fault event is labeled as fault data, and the fault event type is labeled accordingly, and the remaining historical sensor time series data in the training data set are labeled as normal data.

4. The method for diagnosing faults of ground-based centralized photovoltaic modules based on machine learning according to claim 1, characterized in that: The training data set is randomly sampled to form multiple pre-training data sets, including: According to preset sampling parameters, randomly sampling historical sensor data from the training data set to form a pre-sub-training data set, and putting the randomly sampled historical sensor data back into the training data set, wherein the preset sampling parameters include a preset size of the pre-sub-training data set; Repeat the random sampling and replacement according to the preset sampling parameters to form multiple pre-training data sets; Accordingly, according to the preset sample ratio, the ratio of normal data to fault data in each pre-sub-training data set is adjusted by using a random under-sampling method to form multiple sub-training data sets, which includes: The normal data in each pre-sub-training dataset is defined as the majority class sample, and the fault data in each pre-sub-training dataset is defined as the minority class sample; According to the preset sample ratio, the number of majority class samples to be deleted is calculated for each pre-training data set, and the majority class samples are randomly sampled and deleted according to the calculation results; Each pre-sub-training data set that has been randomly sampled and deleted from the majority class samples is defined as a sub-training data set.

5. The method for diagnosing faults of ground-based centralized photovoltaic modules based on machine learning according to claim 1, characterized in that: Based on the group learning algorithm, fault diagnosis training is performed on multiple sub-training data sets in parallel to obtain multiple sub-fault diagnosis models, including: Using group learning algorithm, multiple sub-training data sets are assigned to different computing nodes; The sub-training data sets assigned to them are processed independently by different computing nodes, and the sub-fault diagnosis model parameters are calculated independently; The sub-fault diagnosis model parameters are optimized by back-propagation algorithm, and L2 regularization term is introduced to suppress overfitting; Based on the regularized sub-fault diagnosis model parameters, multiple sub-fault diagnosis models are generated.

6. The method for diagnosing faults of ground-based centralized photovoltaic modules based on machine learning according to claim 1, characterized in that: The real-time sensor data fed back by the ground-based centralized photovoltaic modules is obtained, and the real-time sensor data is input into each sub-fault diagnosis model. The output results of each sub-fault diagnosis model are counted based on the voting method, and the final fault diagnosis results are obtained, including: Acquire the real-time sensor data fed back from the ground-based centralized photovoltaic modules, and input them into each sub-fault diagnosis model to obtain the output results of each sub-fault diagnosis model; The output results of each sub-fault diagnosis model are counted, and the probability distribution diagram of each fault type is drawn according to the statistical results; Based on the probability distribution diagram of each fault type, the occurrence probability value of each fault type is calculated respectively; The voting method is used to compare the occurrence probability values ​​of each fault type. If the fault type with the highest occurrence probability value is not unique, a dynamic resampling mechanism is introduced to re-acquire the historical sensor data stored in the ground centralized photovoltaic module to form a secondary training data set and complete the secondary training. Based on the secondary training, multiple secondary sub-fault diagnosis models are constructed, and the voting method is used again to determine whether the fault type with the highest occurrence probability value is unique until the result is yes. If the fault type with the highest occurrence probability value is unique, the fault type with the highest occurrence probability value will be used as the final fault diagnosis result.

7. The method for diagnosing faults of ground-based centralized photovoltaic modules based on machine learning according to claim 1, characterized in that: After obtaining the final fault diagnosis results, it also includes: According to the final fault diagnosis results, combined with the actual topological structure of the ground-based centralized photovoltaic modules, multiple suspected fault points were located through abnormal voltage and current values; Obtain infrared images of ground-based centralized photovoltaic modules, and obtain fault location information by performing secondary verification and interference elimination on multiple suspected fault points based on the infrared images of the ground-based centralized photovoltaic modules; The communication sends out an alarm message, wherein the alarm message includes the final fault diagnosis result and the fault location information.

8. A ground-based centralized photovoltaic module fault diagnosis system based on machine learning, characterized in that: include: A photovoltaic module data acquisition module is used to obtain historical sensor data stored in the ground centralized photovoltaic module and perform data preprocessing on the historical sensor data to form a training data set, wherein the historical sensor data includes historical data of direct irradiance, diffuse irradiance, temperature, humidity, DC side voltage, DC side current and output active power; The data processing module is used to mark each historical sensor data as normal data and fault data according to the fault occurrence time recorded by the ground centralized photovoltaic module; it is also used to randomly sample the training data set to form multiple pre-sub-training data sets, screen out the normal data and fault data in each pre-sub-training data set, and adjust the ratio of normal data to fault data in each pre-sub-training data set by using a random under-sampling method according to a preset sample ratio to form multiple sub-training data sets; A model training module is used to perform fault diagnosis training on multiple sub-training data sets in parallel based on a group learning algorithm to obtain a sub-fault diagnosis model; The fault diagnosis module is used to obtain the real-time sensor data fed back by the ground centralized photovoltaic module, input the real-time sensor data into each sub-fault diagnosis model, and count the output results of each sub-fault diagnosis model based on the voting method to obtain the final fault diagnosis result.

9. An electronic device, characterized in that: It includes a memory, a processor and a transceiver which are communicatively connected in sequence, wherein the memory is used to store a computer program, the transceiver is used to send and receive messages, and the processor is used to read the computer program to execute the ground centralized photovoltaic module fault diagnosis method based on machine learning as described in any one of claims 1 to 7.

10. A computer program product comprising a computer program or instructions, characterized in that When the computer program or the instruction is executed by a computer, the method for diagnosing faults of ground-based centralized photovoltaic modules based on machine learning is implemented as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Distributed photovoltaic grid-connected area protection method and system

    CN120765208A