Photovoltaic string shielding identification method and device based on data mining and electronic equipment

By using an improved random forest algorithm to identify photovoltaic string shading, and utilizing feature importance weights and weighted information gain, the problems of difficulty in distinguishing shading types and high misjudgment rate in traditional methods are solved, achieving more efficient shading identification and power generation efficiency.

CN120744696APending Publication Date: 2025-10-03HUNAN WULING POWER TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510767592.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Traditional photovoltaic string shading recognition methods cannot effectively distinguish the types of shading, resulting in a high misjudgment rate, which affects the power generation efficiency of the photovoltaic system.

Method used

An occlusion recognition model based on an improved random forest algorithm is adopted. By obtaining the operating data of photovoltaic strings and surrounding weather data, feature extraction and feature importance weight adjustment are performed. The weighted information gain is combined as the decision tree splitting criterion to identify the occlusion type of photovoltaic strings.

Benefits of technology

The classification accuracy of shading types is improved, the misjudgment rate is reduced, and the accuracy and efficiency of photovoltaic string shading identification are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120744696A_ABST
    Figure CN120744696A_ABST
Patent Text Reader

Abstract

The invention provides a photovoltaic string shielding recognition method and device based on data mining and electronic equipment. The method comprises the steps that operation data and surrounding weather data of a photovoltaic string are acquired, the operation data comprise real-time current and backboard temperature, and the weather data comprise illumination intensity, environment temperature, wind speed and humidity; performing feature extraction on the operation data to obtain operation features, and performing feature extraction on the surrounding weather data to obtain weather features; inputting the operation features and the weather features into a pre-trained shielding recognition model to obtain the shielding type of the photovoltaic string; wherein the shielding recognition model is constructed based on an improved random forest algorithm, feature importance weights and weighted information gains are introduced into the improved random forest algorithm to serve as decision tree splitting standards, and the feature importance weights are adjusted based on shielding types. According to the invention, the classification precision of the shielding types can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to technical fields such as photovoltaic power generation and machine learning, and in particular to a photovoltaic string shading identification method, device and electronic equipment based on data mining, which are used to identify the shading conditions of photovoltaic strings. Background Art

[0002] Photovoltaic strings are susceptible to shading during operation, which can cause a drop in string output power, impacting the system's power generation efficiency. Traditional methods for identifying PV string shading rely on manual inspections or simple threshold judgments, often resulting in an inability to distinguish between shading types and a high rate of false positives. Summary of the Invention

[0003] The embodiments of the present application provide a photovoltaic string shading identification method, device, and electronic device based on data mining, which can solve the problems of traditional photovoltaic string shading identification technology such as the inability to distinguish shading types and high misjudgment rate.

[0004] According to a first aspect of an embodiment of the present application, a method for identifying photovoltaic string obstructions based on data mining is provided, comprising:

[0005] Acquire operating data of the photovoltaic strings and surrounding weather data, the operating data including real-time current and backplane temperature, and the weather data including light intensity, ambient temperature, wind speed, and humidity;

[0006] performing feature extraction on the operation data to obtain operation features, and performing feature extraction on the surrounding weather data to obtain weather features;

[0007] The operating characteristics and weather characteristics are input into a pre-trained shading recognition model to obtain the shading type of the photovoltaic string; wherein the shading recognition model is constructed based on an improved random forest algorithm, and the improved random forest algorithm introduces feature importance weights and weighted information gain as decision tree splitting criteria, and the feature importance weights are adjusted based on the shading type.

[0008] According to a second aspect of an embodiment of the present application, a photovoltaic string shading identification device based on data mining is provided, comprising: a data acquisition module for acquiring operating data of the photovoltaic string and surrounding weather data, the operating data including real-time current and backplane temperature, and the weather data including light intensity, ambient temperature, wind speed, and humidity;

[0009] a feature extraction module, configured to perform feature extraction on the operation data to obtain operation features, and perform feature extraction on the surrounding weather data to obtain weather features;

[0010] An identification module is configured to input the operating characteristics and weather characteristics into a pre-trained shading identification model to obtain the shading type of the photovoltaic string; wherein the shading identification model is constructed based on an improved random forest algorithm, and the improved random forest algorithm introduces feature importance weights and weighted information gain as decision tree splitting criteria, and the feature importance weights are adjusted based on the shading type.

[0011] According to a third aspect of the embodiments of the present application, there is provided an electronic device, including:

[0012] at least one processor; and

[0013] a memory communicatively connected to the at least one processor; wherein,

[0014] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the method described in the first aspect.

[0015] According to a fourth aspect of an embodiment of the present application, a storage medium is provided, wherein the storage medium stores instructions. When the instructions are executed on an electronic device, the electronic device executes the method described in the first aspect above.

[0016] According to a fifth aspect of an embodiment of the present application, a program product is provided, comprising at least one of a program and an instruction, wherein when the at least one of the program and the instruction is executed by an electronic device, the steps of the method described in the first aspect are implemented.

[0017] According to the technical solution of the present application, by introducing feature importance weights (i.e., taking into account the adaptability of data characteristics, the importance of features under different occlusion types is different, such as fixed occlusion is more dependent on backplane temperature, and temporary occlusion is more dependent on weather characteristics) and weighted information gain as decision tree splitting criteria, random forests can maintain high efficiency while improving model performance, thereby improving the classification accuracy of occlusion types and reducing the misjudgment rate.

[0018] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:

[0020] Figure 1 A schematic flow chart of a photovoltaic string shading identification method based on data mining provided in an embodiment of the present application;

[0021] Figure 2 A schematic flow chart of a photovoltaic string shading identification method based on data mining provided in an embodiment of the present application;

[0022] Figure 3 A flowchart of a method for training an occlusion recognition model provided in an embodiment of the present application;

[0023] Figure 4 A block diagram of a photovoltaic string shading identification device based on data mining provided in an embodiment of the present application;

[0024] Figure 5 is a block diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0025] The following describes in detail embodiments of the present application, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.

[0026] The following describes a photovoltaic string shading identification method, device, and electronic device based on data mining according to an embodiment of the present application with reference to the accompanying drawings.

[0027] It should be noted that the executor of the photovoltaic string shading identification method based on data mining in the embodiment of the present application can be a photovoltaic string shading identification device based on data mining, which can be implemented by software and / or hardware. The device can be configured in an electronic device, and the electronic device may include but is not limited to a terminal, a server, an edge computing device, etc.

[0028] Figure 1 This is a flow chart of a photovoltaic string shading identification method based on data mining provided in an embodiment of the present application. Figure 1 As shown, the photovoltaic string shading identification method based on data mining may include but is not limited to the following steps.

[0029] In step 101 , the operation data of the photovoltaic strings and the surrounding weather data are obtained.

[0030] In an embodiment of the present application, the operating data may include but is not limited to real-time current and backplane temperature; the weather data may include but is not limited to light intensity, ambient temperature, wind speed and humidity.

[0031] For example, sensors can be used to collect real-time operating data such as the current and backplane temperature of the photovoltaic strings, and weather observation equipment at the site where the photovoltaic strings are located can be used to obtain weather data surrounding the photovoltaic strings. Alternatively, the weather data surrounding the photovoltaic strings can be obtained from a weather station, but is not limited thereto.

[0032] In step 102 , feature extraction is performed on the operation data to obtain operation features, and feature extraction is performed on the surrounding weather data to obtain weather features.

[0033] In some embodiments, the operating characteristics may include, but are not limited to, current fluctuation rate and backplate temperature gradient. Optionally, the operating characteristics may also include, but are not limited to, current, backplate temperature, etc. In one possible implementation, the current dispersion and mean value may be determined based on the real-time current in the operating data, and the current fluctuation rate may be determined based on the current dispersion and mean value; and the backplate temperature gradient may be determined based on the backplate temperature in the operating data.

[0034] For example, the standard deviation of the current can be determined based on the real-time current in the operating data. The standard deviation of the current can be expressed as the degree of dispersion of the current, reflecting the absolute fluctuation amplitude of the current from the average value. The mean value of the current can be determined based on the real-time current in the operating data, representing the typical level of the current. The ratio of the degree of dispersion of the current to the mean value of the current is determined as the current fluctuation rate, and the formula is expressed as follows: current fluctuation rate Cvar = standard deviation (I) / mean (I), which means capturing abnormal current fluctuations caused by obstruction.

[0035] For example, a backplane temperature gradient can be calculated based on the backplane temperature in the operating data. Optionally, the backplane temperature gradient can be calculated using the following formula: Tgrad = ΔT / Δt, where Tgrad is the backplane temperature gradient and ΔT is the change in backplane temperature within the time period Δt. The backplane temperature gradient is selected as a feature because the backplane temperature change rate is faster than that of a normal PV string due to the obstruction of heat dissipation in the shaded PV string.

[0036] In some embodiments, the weather characteristics may include but are not limited to a weather mutation index. Optionally, the weather characteristics may also include light intensity, ambient temperature, wind speed, humidity, etc., but are not limited thereto. In one possible implementation, the instantaneous rate of change of light intensity may be determined based on the light intensity in the surrounding weather data; the weather mutation index may be determined based on the instantaneous rate of change of light intensity and the inverse ratio of humidity. Exemplarily, the instantaneous rate of change of light intensity and the inverse ratio of humidity may be multiplied to obtain the weather mutation index. As an example, the calculation formula of the weather mutation index may be expressed as follows: Among them, Windex is the weather mutation index, is the instantaneous rate of change of light intensity, H is the humidity, and the weather mutation index feature can be used to distinguish cloud cover from other obstructions.

[0037] In step 103 , the operating characteristics and weather characteristics are input into a pre-trained shading recognition model to obtain the shading type of the photovoltaic string.

[0038] Among them, in an embodiment of the present application, the occlusion recognition model can be constructed based on an improved random forest algorithm, and the occlusion recognition model based on the improved random forest algorithm can be pre-trained based on a training data set. In an embodiment of the present application, the improved random forest algorithm introduces feature importance weights and weighted information gain as decision tree splitting criteria, and the feature importance weights can be dynamically adjusted based on the occlusion type. Exemplarily, the above-mentioned feature importance weights and weighted information gain can be used to select the best splitting feature, that is, the decision tree can use feature importance weights and weighted information gain as core indicators when selecting the best splitting feature to measure the ability of a feature to distinguish classification results, with the aim of finding the best splitting point so that the splitted data is as pure as possible (similar samples are clustered).

[0039] In the above embodiment, by introducing feature importance weights (i.e., taking into account the adaptability of data characteristics, the importance of features under different occlusion types is different, such as fixed occlusion is more dependent on backplane temperature, and temporary occlusion is more dependent on weather characteristics) and weighted information gain as decision tree splitting criteria, random forests can maintain high efficiency while improving model performance, thereby improving the classification accuracy of occlusion types and reducing the misjudgment rate.

[0040] Figure 2 This is a flow chart of a photovoltaic string shading identification method based on data mining provided in an embodiment of the present application. Figure 2 As shown, the photovoltaic string shading identification method based on data mining may include but is not limited to the following steps.

[0041] In step 201 , the operation data of the photovoltaic strings and the surrounding weather data are obtained.

[0042] In an embodiment of the present application, the operating data may include but is not limited to real-time current and backplane temperature; the weather data may include but is not limited to light intensity, ambient temperature, wind speed and humidity.

[0043] Optionally, step 201 may be implemented using any one of the implementation methods in the embodiments of the present application. The embodiments of the present application do not limit this and will not be described in detail.

[0044] In step 202 , feature extraction is performed on the operation data to obtain operation features, and feature extraction is performed on the surrounding weather data to obtain weather features.

[0045] Optionally, step 202 may be implemented using any one of the implementation methods in the embodiments of the present application. The embodiments of the present application do not limit this and will not be described in detail.

[0046] In step 203 , the operating characteristics and weather characteristics are input into a pre-trained shading recognition model to obtain the shading type of the photovoltaic string.

[0047] Optionally, step 203 may be implemented using any implementation method in each embodiment of the present application. The embodiments of the present application do not limit this and will not be described in detail.

[0048] It should be noted that, in some embodiments, the shading type of the photovoltaic string can be divided into three categories: fixed shading, temporary shading, and no shading. The subsequent processing flow will be different depending on the shading type of the photovoltaic string. In some embodiments, when the shading type of the photovoltaic string is no shading, the operation data and surrounding weather data of the photovoltaic string can continue to be obtained without the need for post-shading processing and operation and maintenance decisions. In some embodiments, when the shading type of the photovoltaic string is fixed shading, the following step 204 can be executed. In some embodiments, when the shading type of the photovoltaic string is temporary shading, the following step 205 can be executed.

[0049] In step 204, when the shading type of the photovoltaic string is fixed shading, the recoverable loss of electricity is determined based on the real-time power of the non-shading photovoltaic string with the same capacity and the same inverter, and the measured power of the photovoltaic string, and the first push information is visually displayed. The first push information may at least include but is not limited to the recoverable loss of electricity, the shading type and the time when the shading occurs.

[0050] For example, when the shading type of a photovoltaic string is fixed shading, the real-time power of an unshaded photovoltaic string with the same capacity and the same inverter as the photovoltaic string can be obtained, and the measured power of the photovoltaic string can be obtained. Based on the real-time power of the unshaded photovoltaic string and the measured power of the photovoltaic string, the recoverable lost electricity can be calculated. As an example, the calculation formula for the recoverable lost electricity can be expressed as follows:

[0051]

[0052] Among them, E loss P is the recoverable lost electricity; ref (t) is the real-time power of the non-shaded PV string with the same capacity and the same inverter at time t; P obs (t) is the measured power of the PV string at time t; T is the time period.

[0053] In the embodiments of the present application, the time of occurrence of fixed obstructions can also be recorded, and the recoverable power loss, obstruction type (i.e., fixed obstruction), and obstruction occurrence time can be visualized. Early warnings can also be issued to facilitate corresponding operation and maintenance decisions based on the recoverable power loss and / or obstruction occurrence time, such as whether the fixed obstruction needs to be removed or whether the PV strings need to be moved. Therefore, by combining the real-time power comparison of the same capacity and the same inverter strings, the recoverable power loss can be accurately quantified, improving the accuracy of the lost power.

[0054] In step 205, when the shading type of the photovoltaic string is temporary shading, the shading type and the shading occurrence time are recorded, and the second push information is visually displayed. The second push information may at least include but is not limited to the shading type and the shading occurrence time.

[0055] For example, when the shading type of the photovoltaic string is temporary, only the shading type (i.e., temporary shading) and the time when the shading occurred can be recorded. Since cloud shading cannot be eliminated through operation and maintenance, the lost electricity will not be accumulated.

[0056] Optionally, in some embodiments, a rule engine can be pre-placed to exclude non-blocking factors such as photovoltaic string failures (such as inverter alarms), maintenance status (operation and maintenance marks), etc. In one possible implementation, the rule engine may include inverter alarm information and / or the maintenance status of the photovoltaic string. When obtaining the operating data and surrounding weather data of the photovoltaic string, it can be determined based on the rule engine whether the photovoltaic string has non-blocking factors. If it is determined based on the rule engine that the photovoltaic string does not have non-blocking factors, occlusion identification can be performed based on the operating data and surrounding weather data of the photovoltaic string. If it is determined based on the rule engine that the photovoltaic string has non-blocking factors, the photovoltaic string occlusion identification process can be exited, which means that the operating data abnormality of the photovoltaic string is not caused by occlusion. Therefore, by integrating the rule engine with the machine learning model (i.e., the occlusion identification model), the misjudgment rate of occlusion type identification can be further reduced.

[0057] It should be noted that the occlusion recognition model based on the improved random forest algorithm can be pre-trained based on a training data set. In some embodiments, Figure 3 As shown, the training method of the occlusion recognition model may include but is not limited to the following steps.

[0058] In step 301 , a training data set is obtained, where the training data set includes operation characteristics and weather characteristics of a plurality of photovoltaic string samples, as well as labels.

[0059] In some embodiments, the operating characteristics may include, but are not limited to, current fluctuation rate and backplate temperature gradient. Optionally, the operating characteristics may also include, but are not limited to, current, backplate temperature, etc. The method for determining the current fluctuation rate and backplate temperature gradient can be found in the relevant description of step 102 above and will not be repeated here.

[0060] In some embodiments, the weather characteristics may include, but are not limited to, a weather mutation index. Optionally, the weather characteristics may also include, but are not limited to, light intensity, ambient temperature, wind speed, humidity, etc. The method for determining the weather mutation index can be found in the description of step 102 above and will not be repeated here.

[0061] In some embodiments, the label may include a fixed occlusion type label, a temporary occlusion type label, and a non-occlusion type label. For example, the label may be verified in conjunction with surveillance video or operation and maintenance records.

[0062] In step 302, random sampling is performed from the training data set to obtain training sets for multiple decision trees.

[0063] In some embodiments, for each decision tree, the training data of the minority class (such as the fixed occlusion type) can be randomly sampled from the training data set using an oversampling technique (such as SMOTE (Synthetic Minority Oversampling Technique)), and the training data of the majority class (such as the non-occlusion type) can be randomly sampled from the training data set using an undersampling technique (such as Tomek Links undersampling), so as to obtain the training set for training the decision tree. The minority class is sampled by SMOTE oversampling and the majority class is sampled by Tomek Links undersampling, so that the data set can be balanced. Exemplarily, when randomly sampling, the fixed occlusion samples can be oversampled to 30%, and the non-occlusion samples can be undersampled to 50%.

[0064] In step 303, for each decision node in each decision tree, some features are randomly selected from the operating features and weather features in the training set of the decision tree as candidate features, and the best splitting feature is selected from the candidate features according to the feature importance weight and the weighted information gain.

[0065] For example, for each decision node in each decision tree, some features can be randomly selected from the operational features and weather features of the training set of the decision tree (the number of these features is, for example, For example, if the total number of operating features and weather features is n, the number of randomly selected features can be Then, the best split feature can be selected from the candidate features according to the feature importance weight and the weighted information gain, and the best split feature is used for node splitting.

[0066] In some embodiments, the formula for the feature importance weight can be expressed as follows:

[0067]

[0068] Among them, w i,k Importance is the weight of the i-th feature in the k-th occlusion; i,k Importance is the importance of the i-th feature in the k-th occlusion. i,k It is measured by the sum of the Gini impurities reduced when the i-th feature splits nodes in all decision trees; n is the total number of operational and weather features. For example, the backplane temperature gradient is weighted more heavily under fixed shading, while the weather mutation index is weighted more heavily under temporary shading.

[0069] For example, the Gini reduction of a single split of the i-th feature can be calculated, and each decision tree in the random forest is traversed. The Gini reduction of the i-th feature when splitting at all nodes is recorded and added to the Gini reduction of the i-th feature, so that the sum of the Gini impurities reduced when the i-th feature splits nodes in all decision trees (or the total Gini reduction) can be obtained, and the total Gini reduction of all features is normalized to a weight sum of 1 to obtain the final importance of the i-th feature in the k-th occlusion. i,k Utilize this Importance i,k Combining the above formula of feature importance weight, we can get the weight w of the i-th feature in the k-th occlusion i,k .

[0070] For example, the steps for calculating the Gini reduction of a single split of the i-th feature may be as follows: 1) Calculate the Gini coefficient of the parent node using the Gini coefficient formula, which can be expressed as follows: K is the total number of occlusion types, p k is the number of samples blocked by the kth class, and N is the total number of samples of the parent node. 2) Split the data and calculate the Gini coefficient of the child nodes: Split the parent node D into m child nodes D1, D2, ..., D according to the i-th feature. m , the Gini coefficient of each child node is: p kj For child node D j The number of samples of the kth type of occlusion, N j For child node D j3) Calculate the weighted Gini coefficient of the child node: The Gini coefficient of each child node is weighted and summed according to the sample ratio. The calculation formula is as follows: 4) Subtract the weighted Gini coefficient of the child node from the Gini coefficient of the parent node to get the Gini reduction of the single split of the i-th feature. After obtaining the Gini reduction of the single split of the i-th feature, you can traverse each decision tree in the random forest, record the Gini reduction of the i-th feature when all nodes are split, and accumulate these Gini reductions to get the sum of the Gini impurities reduced when the i-th feature splits nodes in all decision trees (or the total Gini reduction), and normalize the total Gini reduction of all features to a weight sum of 1 to get the final importance of the i-th feature in the k-th category occlusion. i,k .

[0071] In some embodiments, the weighted information gain formula is as follows:

[0072]

[0073] Among them, WIG is the weighted information gain, a k is the penalty coefficient for the kth type of occlusion, IG k is the information gain of the kth type of occlusion. As an example, for the fixed occlusion type, the penalty coefficient a=1.5; for the temporary occlusion type, the penalty coefficient a=1.2; for the no occlusion type, the penalty coefficient a=1.

[0074] For example, information gain is based on information entropy, which measures the reduction in uncertainty before and after splitting. The larger the value, the better the splitting effect. The formula for information entropy is as follows: Among them, p k is the proportion of samples belonging to the kth type of occlusion in the current node. When splitting the node according to the i-th feature, the calculation steps of the information gain of the kth type of occlusion can be as follows: 1) Calculate the information entropy of the current node using the information entropy formula 父节点 . 2) Calculate the weighted information entropy after splitting by the i-th feature: 3) Based on the information entropy and weighted information entropy of the current node, the information gain IG of the kth type of occlusion can be calculated k , the calculation formula is as follows: The information gain of each occlusion type can be calculated using the above method, and then the weighted information gain WIG can be obtained using the above formula (2).

[0075] In the embodiment of the present application, for each decision node in each decision tree, after randomly selecting some features from the operational features and weather features of the training set of the decision tree as candidate features, the weight w of the i-th feature in the k-th occlusion can be used. i,k And the weighted information gain WIG of the i-th feature, the best split feature is selected from the candidate features.

[0076] In step 304 , for each decision tree, the decision tree is trained using the best splitting feature associated with the decision tree and the training set.

[0077] In step 305 , the trained multiple decision trees are combined into a random forest, and grid search is used to optimize the parameters of the random forest to construct an occlusion recognition model.

[0078] In an embodiment of the present application, in a classification task, each decision tree in the random forest classifies a sample, and ultimately determines the sample category through a voting mechanism. In an embodiment of the present application, a grid search can be used to optimize the random forest parameters, such as a maximum depth of 15 and a number of trees of 200.

[0079] In the above embodiment, by introducing feature importance weights (i.e., taking into account the adaptability of data characteristics, the importance of features under different occlusion types is different, such as fixed occlusion is more dependent on backplane temperature, and temporary occlusion is more dependent on weather characteristics) and weighted information gain as decision tree splitting criteria, random forests can maintain high efficiency while improving model performance, thereby improving the classification accuracy of occlusion types and reducing the misjudgment rate.

[0080] Figure 4 This is a block diagram of a photovoltaic string shading identification device based on data mining provided in an embodiment of the present application. Figure 4 As shown, the photovoltaic string shading identification device based on data mining may include: a data acquisition module 401 , a feature extraction module 402 and an identification module 403 .

[0081] The data acquisition module 401 is used to acquire the operation data of the photovoltaic string and the surrounding weather data. The operation data includes real-time current and backplane temperature, and the weather data includes light intensity, ambient temperature, wind speed and humidity.

[0082] Feature extraction module 402 is configured to extract features from the operating data to obtain operating characteristics, and to extract features from the surrounding weather data to obtain weather characteristics. In some embodiments, the operating characteristics include at least current fluctuation rate and backplane temperature gradient. In one possible implementation, feature extraction module 402 is configured to: determine the current dispersion and mean value based on the real-time current in the operating data, and determine the current fluctuation rate based on the current dispersion and mean value; and determine the backplane temperature gradient based on the backplane temperature in the operating data.

[0083] In some embodiments, the weather characteristics include at least a weather mutation index. In one possible implementation, the feature extraction module 402 is configured to: determine the instantaneous rate of change of light intensity based on the light intensity in the surrounding weather data; and determine the weather mutation index based on the instantaneous rate of change of light intensity and the inverse ratio of humidity.

[0084] Identification module 403 is used to input operating characteristics and weather characteristics into a pre-trained shading identification model to obtain the shading type of the photovoltaic string; wherein the shading identification model is constructed based on an improved random forest algorithm, and the improved random forest algorithm introduces feature importance weights and weighted information gain as decision tree splitting criteria, and the feature importance weights are adjusted based on the shading type.

[0085] In some embodiments, the data mining-based photovoltaic string obstruction identification device may further include a training module. The training module is configured to: obtain a training data set, comprising the operating characteristics and weather characteristics of a plurality of photovoltaic string samples, as well as labels; perform random sampling from the training data set to obtain training sets for a plurality of decision trees; for each decision node in each decision tree, randomly select a portion of features from the operating characteristics and weather characteristics of the decision tree's training set as candidate features, and select the best splitting feature from the candidate features based on feature importance weights and weighted information gain; for each decision tree, train the decision tree using the best splitting feature associated with the decision tree and the training set; combine the trained decision trees into a random forest, and optimize the random forest parameters using grid search to construct an obstruction identification model.

[0086] In some embodiments, the feature importance weight is formulated as follows:

[0087]

[0088] Among them, w i,k Importance is the weight of the i-th feature in the k-th occlusion; i,k Importance is the importance of the i-th feature in the k-th occlusion. i,kIt is measured by the sum of the Gini impurities reduced when the i-th feature splits the nodes in all decision trees; n is the total number of operating features and weather features.

[0089] In some embodiments, the weighted information gain is formulated as follows:

[0090]

[0091] Among them, WIG is the weighted information gain, a k is the penalty coefficient for the kth type of occlusion, IG k is the information gain of the k-th occlusion.

[0092] In some embodiments, the photovoltaic string obstruction identification device based on data mining may further include a post-processing module. The post-processing module is configured to: when the photovoltaic string obstruction type is fixed obstruction, determine the recoverable power loss based on the real-time power of non-obstructed photovoltaic strings with the same capacity and inverter, and the measured power of the photovoltaic strings, and visually display a first push message, the first push message including at least the recoverable power loss, the obstruction type, and the time when the obstruction occurred; or, when the photovoltaic string obstruction type is temporary obstruction, record the obstruction type and the time when the obstruction occurred, and visually display a second push message, the second push message including at least the obstruction type and the time when the obstruction occurred.

[0093] It should be noted that the aforementioned explanation of the embodiment of the photovoltaic string obstruction identification method based on data mining is also applicable to the photovoltaic string obstruction identification device based on data mining in this embodiment, and will not be repeated here.

[0094] According to an embodiment of the present application, the present application also provides an electronic device and a readable storage medium.

[0095] like Figure 5 , is a block diagram of an electronic device according to an embodiment of the present application. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or required herein.

[0096] like Figure 5As shown, the electronic device includes: one or more processors 501, a memory 502, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. The various components are connected to each other using different buses and can be installed on a common mainboard or installed in other ways as needed. The processor can process instructions executed in the electronic device, including instructions stored in or on the memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In other embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple electronic devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 5 A processor 501 is taken as an example.

[0097] Memory 502 is the non-transitory computer-readable storage medium provided in this application. The memory stores instructions executable by at least one processor, causing the at least one processor to execute the data mining-based photovoltaic string shading identification method provided in this application. The non-transitory computer-readable storage medium of this application stores computer instructions for causing a computer to execute the data mining-based photovoltaic string shading identification method provided in this application.

[0098] The memory 502 is a non-transient computer-readable storage medium that can be used to store non-transient software programs, non-transient computer executable programs and modules, such as the program instructions / modules corresponding to the photovoltaic string shading identification method based on data mining in the embodiment of the present application (for example, the attached Figure 4 The processor 501 executes the non-transient software programs, instructions, and modules stored in the memory 502 to execute various functional applications and data processing of the server, thereby implementing the photovoltaic string shading identification method based on data mining in the above method embodiment.

[0099] The memory 502 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and applications required for at least one function; the data storage area may store data created based on the use of the electronic device, etc. In addition, the memory 502 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory 502 may optionally include a memory remotely located relative to the processor 501, and these remote memories may be connected to the electronic device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0100] The electronic device may further include: an input device 503 and an output device 504. The processor 501, the memory 502, the input device 503 and the output device 504 may be connected via a bus or other means. Figure 5 The bus connection is taken as an example.

[0101] The input device 503 can receive input digital or character information and generate key signal input related to user settings and function control of the electronic device, such as input devices such as a touch screen, a keypad, a mouse, a trackpad, a touch pad, an indicator stick, one or more mouse buttons, a trackball, and a joystick. The output device 504 can include a display device, an auxiliary lighting device (e.g., an LED), and a tactile feedback device (e.g., a vibration motor). The display device can include, but is not limited to, a liquid crystal display (LCD), a light emitting diode (LED) display, and a plasma display. In some embodiments, the display device can be a touch screen.

[0102] Various implementations of the systems and techniques described herein can be realized in digital electronic circuit systems, integrated circuit systems, dedicated ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0103] These computer programs (also referred to as programs, software, software applications, or code) include machine instructions for a programmable processor and can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device (e.g., a magnetic disk, an optical disk, a memory, a programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.

[0104] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0105] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.

[0106] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact via a communication network. This client-server relationship is established by computer programs running on the respective computers, establishing a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host, a host product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosts and VPS services ("Virtual Private Servers" or simply "VPS"). The server may also be a server in a distributed system or a server integrated with blockchain.

[0107] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this application can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of this application can be achieved. This is not limited herein.

[0108] The above specific embodiments do not constitute a limitation on the scope of protection of this application. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application shall be included within the scope of protection of this application.

Claims

1. A photovoltaic string shading identification method based on data mining, characterized in that: include: Acquire operating data of the photovoltaic strings and surrounding weather data, the operating data including real-time current and backplane temperature, and the weather data including light intensity, ambient temperature, wind speed, and humidity; performing feature extraction on the operation data to obtain operation features, and performing feature extraction on the surrounding weather data to obtain weather features; The operating characteristics and weather characteristics are input into a pre-trained shading recognition model to obtain the shading type of the photovoltaic string; wherein the shading recognition model is constructed based on an improved random forest algorithm, and the improved random forest algorithm introduces feature importance weights and weighted information gain as decision tree splitting criteria, and the feature importance weights are adjusted based on the shading type.

2. The method according to claim 1, characterized in that The operating characteristics include at least a current fluctuation rate and a backplate temperature gradient; and extracting characteristics from the operating data to obtain the operating characteristics includes: Determining the current dispersion and mean value according to the real-time current in the operating data, and determining the current fluctuation rate according to the current dispersion and mean value; A backplate temperature gradient is determined according to the backplate temperature in the operating data.

3. The method according to claim 1, characterized in that The weather characteristics include at least a weather mutation index; and extracting features from the surrounding weather data to obtain weather characteristics includes: Determining an instantaneous rate of change of light intensity based on the light intensity in the ambient weather data; The weather mutation index is determined according to the instantaneous change rate of the light intensity and the inverse ratio of the humidity.

4. The method according to claim 1, wherein The occlusion recognition model is pre-trained in the following way: Obtaining a training data set, the training data set including operation characteristics and weather characteristics of a plurality of photovoltaic string samples, and labels; Randomly sampling from the training data set to obtain training sets for multiple decision trees; For each decision node in each decision tree, randomly select some features from the operating features and weather features in the training set of the decision tree as candidate features, and select the best splitting feature from the candidate features according to feature importance weights and weighted information gains; For each decision tree, training the decision tree using the best splitting feature associated with the decision tree and the training set; The trained multiple decision trees are combined into a random forest, and grid search is used to optimize the parameters of the random forest to construct the occlusion recognition model.

5. The method according to claim 1 or 4, characterized in that The formula for the feature importance weight is as follows: Among them, w i,k Importance is the weight of the i-th feature in the k-th occlusion; i,k is the importance of the i-th feature in the k-th occlusion, the Importance i,k It is measured by the sum of the Gini impurities reduced when the i-th feature splits nodes in all decision trees; n is the total number of the operating features and the weather features.

6. The method according to claim 1 or 4, characterized in that The formula of the weighted information gain is as follows: Wherein, WIG is the weighted information gain, a k is the penalty coefficient for the kth type of occlusion, IG k is the information gain of the k-th type of occlusion.

7. The method according to any one of claims 1 to 6, characterized in that The method further comprises: In the case where the shading type of the photovoltaic string is fixed shading, based on the real-time power of non-shaded photovoltaic strings with the same capacity and the same inverter and the measured power of the photovoltaic strings, the recoverable lost electricity is determined, and a first push information is visually displayed, where the first push information includes at least the recoverable lost electricity, the shading type, and the time when the shading occurred; or When the shading type of the photovoltaic string is temporary shading, the shading type and the shading occurrence time are recorded, and second push information is visually displayed, where the second push information at least includes the shading type and the shading occurrence time.

8. A photovoltaic string shading identification device based on data mining, characterized in that: include: A data acquisition module is used to acquire the operating data of the photovoltaic strings and the surrounding weather data. The operating data includes real-time current and backplane temperature, and the weather data includes light intensity, ambient temperature, wind speed and humidity. a feature extraction module, configured to perform feature extraction on the operation data to obtain operation features, and perform feature extraction on the surrounding weather data to obtain weather features; An identification module is configured to input the operating characteristics and weather characteristics into a pre-trained shading identification model to obtain the shading type of the photovoltaic string; wherein the shading identification model is constructed based on an improved random forest algorithm, and the improved random forest algorithm introduces feature importance weights and weighted information gain as decision tree splitting criteria, and the feature importance weights are adjusted based on the shading type.

9. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.

10. A storage medium storing instructions, characterized in that: When the instructions are executed on an electronic device, the electronic device is caused to execute the method according to any one of claims 1 to 7.