Photovoltaic power generation fault risk prediction method and system based on big data processing

The photovoltaic power generation fault risk prediction method based on big data processing analyzes the equipment operating status and environmental parameters, identifies the correlation and risk level between equipment, optimizes risk classification, and generates intelligent matching detection groups. This solves the problems of insufficient flexibility and accuracy in traditional methods and realizes efficient fault monitoring and prediction of photovoltaic power generation systems.

CN121329156BActive Publication Date: 2026-03-20SICHUAN HUADIAN JINCHUAN HYDROPOWER DEV CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-09
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

Traditional photovoltaic power generation fault risk prediction methods based on big data processing lack flexibility and accuracy, making it difficult to make dynamic adjustments in complex and ever-changing operating environments. They are unable to monitor and predict the fault risks of large-scale complex systems in real time, resulting in large deviations in prediction results and affecting the timeliness and effectiveness of early warnings.

Method used

The photovoltaic power generation fault risk prediction method based on big data processing extracts equipment operating status, environmental parameters and original fault data, identifies the spatial distribution characteristics of equipment operation, generates a standardized equipment operation dataset, analyzes the correlation strength and risk level between equipment, screens potential risk areas, assesses the probability of risk conflict between areas, optimizes risk classification, extracts core equipment, forms intelligent matching detection groups, and generates a photovoltaic power generation fault risk prediction table.

Benefits of technology

It improves the predictive flexibility and accuracy of photovoltaic power generation systems, enhances the timeliness and effectiveness of fault early warning, optimizes the coordination and consistency analysis between equipment, and ensures accurate monitoring and risk prediction capabilities in large-scale complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121329156B_ABST
    Figure CN121329156B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of big data analysis, in particular to a photovoltaic power generation fault risk prediction method and system based on big data processing, comprising the following steps: based on the basic information of photovoltaic power generation equipment, identifying the spatial distribution characteristics of equipment operation and labeling the risk level, analyzing the correlation strength between equipment, analyzing the dependency relationship between regions, evaluating the risk conflict probability and adjusting the risk classification, analyzing equipment consistency, analyzing group coverage efficiency, screening the optimal efficiency group into the execution plan, and obtaining a photovoltaic power generation fault risk prediction table; in the present application, through deep analysis of the operating state, environmental parameters and fault data, the limitations of relying on fixed rules are reduced, the prediction flexibility and accuracy are improved, the timeliness and effectiveness of fault early warning are enhanced, the coverage efficiency is improved through matching and optimizing the detection group, the accurate monitoring and risk prediction ability under complex environment are ensured, and the operation management effect of the photovoltaic power generation system is comprehensively improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of big data analysis, in particular to a photovoltaic power generation fault risk prediction method and system based on big data processing. BACKGROUND

[0002] Big data analysis technology is widely used in various industries, including finance, healthcare, energy, etc. Its core task is to collect, store, process and analyze massive data to extract valuable information and insights, helping enterprises and organizations make data-driven decisions. In the technical field, data processing technology is the most critical component, mainly involving data collection, data storage, data cleaning, data analysis, etc. Through big data analysis technology, complex business scenarios can be deeply understood, potential problems can be found, operational efficiency can be improved, and resource allocation can be optimized. Big data analysis technology also covers machine learning, artificial intelligence and cloud computing techniques to more efficiently process and analyze large-scale data sets, thereby maximizing the value of data.

[0003] Among them, the traditional photovoltaic power generation fault risk prediction method based on big data processing refers to analyzing the historical operation data of the photovoltaic power generation system, combining the operating status of the equipment, environmental factors and other related data, and using statistical analysis or empirical models to predict the occurrence of faults or risks. The traditional method relies on manually set thresholds and rules, and makes empirical judgments based on historical data to assess the fault risk of the photovoltaic power generation system. The traditional method relies on fixed models, lacks flexibility and precision, and is difficult to dynamically adjust in different operating environments, and has limited real-time monitoring and prediction capabilities for large-scale complex systems. Therefore, the accuracy and applicability of the traditional method have certain limitations.

[0004] The existing technology relies on statistical analysis or empirical models based on historical data to predict the fault risk of photovoltaic power generation systems using manually set thresholds and rules, which lacks flexibility and precision. Especially in the face of complex and variable operating environments, it is difficult to dynamically adjust and cannot meet the real-time monitoring and prediction needs of large-scale systems, resulting in a large deviation in the prediction results and the inability to reflect the system state and changes in real time. In addition, it is difficult to conduct in-depth analysis of the interrelationships between different regions and equipment, which limits the accuracy and applicability of fault prediction and affects the timeliness and effectiveness of early warning, making it difficult to meet the precise management and scheduling needs of complex systems. SUMMARY

[0005] To solve the technical problems existing in the prior art, the present application provides a photovoltaic power generation fault risk prediction method based on big data processing, comprising the following steps:

[0006] In order to achieve the above object, the present application adopts the following technical scheme: a photovoltaic power generation fault risk prediction method based on big data processing, comprising the following steps:

[0007] S1: based on the basic information of the photovoltaic power generation equipment, extracting the equipment running state, environmental parameters and original fault data, identifying the spatial distribution characteristics of equipment operation, and labeling risk level labels, generating a standardized equipment operation data set;

[0008] S2: based on the standardized equipment operation data set, analyzing the correlation strength between equipment, combining the risk level label and the environmental parameter, screening the potential risk area, and generating a risk area data set;

[0009] S3: calling the risk area data set, analyzing the dependency relationship between regions, evaluating the risk conflict probability between regions, adjusting the risk classification, and obtaining the risk classification optimization result;

[0010] S4: based on the risk classification optimization result, extracting the core equipment, analyzing the consistency between equipment, screening the consistency equipment into the same detection group, and outputting the intelligent matched detection group;

[0011] S5: calling the intelligent matched detection group, analyzing the group coverage efficiency, sorting the group priority according to the coverage efficiency, screening the optimal efficiency group into the execution plan, and obtaining the photovoltaic power generation fault risk prediction table.

[0012] As a further scheme of the present application, the equipment operation spatial distribution diagram includes an equipment operation state layer, an environmental parameter layer, an original fault data layer and a risk level identification layer, the standardized equipment operation data set includes equipment identification number, environmental parameter information, fault data label and risk level code, the risk area data set includes region execution sequence table, inter-regional correlation matrix and risk priority parameter, the risk classification optimization result includes risk classification list, risk conflict probability table and execution time period distribution diagram, the intelligent matched detection group includes equipment sequence, equipment consistency index and group aggregation grouping result, and the photovoltaic power generation fault risk prediction table includes detection group number, group efficiency score and execution priority sequence.

[0013] As a further scheme of the present application, the step of the standardized equipment operation data set is specifically:

[0014] S101: based on the basic information of the photovoltaic power generation equipment, extracting the equipment running state, environmental parameters and original fault data, removing redundant information, and generating equipment operation basic data;

[0015] S102: according to the equipment operation basic data, analyzing the spatial relationship between the equipment operation state and the environmental parameter, labeling the risk level label, combining the fault data and the equipment attribute, and generating the risk distribution interval;

[0016] S103: According to the risk distribution interval, the environmental parameter characteristics are analyzed by equipment classification, the environmental types, load states and fluctuations are grouped, the key operation behaviors are summarized, and the standardized equipment operation data set is generated.

[0017] As a further scheme of the present application, the step of the risk area data set is specifically:

[0018] S201: Based on the equipment distribution of the standardized equipment operation data set, the correlation strength between equipment is analyzed, the potential risk area is filtered according to the correlation strength value, and the equipment correlation strength matrix is generated;

[0019] S202: Based on the equipment correlation strength matrix, the dependence between equipment is analyzed by combining risk level label and environmental parameter, the area meeting the risk priority requirement is filtered, the equipment quantity, area span and risk distribution data are counted, and the equipment distribution data set is obtained;

[0020] S203: The equipment distribution data set is called, the area optimization factor is calculated, and the risk area data set is divided according to the equipment quantity, area span and risk distribution.

[0021] As a further scheme of the present application, the step of the risk classification optimization result is specifically:

[0022] S301: The risk area data set is called, the dependence between areas is analyzed, the area type and execution time window are classified, and the area dependence distribution data is obtained;

[0023] S302: Based on the area dependence distribution data, the risk conflict probability between areas is identified, the risk overlap part in the differentiated area is extracted, the overlap times and risk conflict probability are counted, and the area risk conflict matrix is generated;

[0024] S303: The area risk conflict matrix is called, the area is divided according to the conflict probability, the low conflict area is merged, the high conflict area is split, and the risk classification optimization result is obtained.

[0025] As a further scheme of the present application, the risk overlap part refers to the intersection of two or more differentiated areas in risk characteristics, influence range or trigger condition when performing area risk analysis;

[0026] The risk conflict probability refers to the estimation that the risk events occur simultaneously or continuously due to the mutual restriction of resources, time sequence or logic between differentiated areas, so as to obtain a quantitative index;

[0027] The low conflict area refers to an area in which the risk conflict probability between the area and other areas is generally lower than a preset merging threshold in the area risk conflict matrix.

[0028] The high conflict area refers to an area in which the risk conflict probability between the area and other areas is higher than a preset splitting threshold in the area risk conflict matrix.

[0029] As a further scheme of the present application, the step of detecting the group after intelligent matching is specifically:

[0030] S401: Based on the risk classification optimization result, extract the core equipment, filter the consistent equipment in the group according to the consistency between the equipment, and calculate the equipment consistency value;

[0031] S402: According to the equipment consistency value, analyze the consistency characteristics between the equipment, identify the behavior mode between the adjacent equipment, filter the consistent equipment combination according to the consistency and the distribution characteristics of the equipment, and generate the equipment consistency data set;

[0032] S403: Based on the equipment consistency data set, the consistent equipment is used to cluster the equipment belonging to the same power generation area, merge the same equipment and divide the attribution group, and generate the detection group after intelligent matching.

[0033] As a further scheme of the present application, the consistent equipment in the group refers to that when the consistency index between the core equipment is higher than a preset consistency threshold, the core equipment is determined as a consistent equipment;

[0034] The equipment consistency value refers to the weighted average result of the average running efficiency, fault frequency and environmental adaptability of the core equipment within a preset monitoring period;

[0035] The behavior mode between the adjacent equipment refers to constructing the behavior mode by analyzing the running state change, load fluctuation and environmental response characteristics of the adjacent equipment within a fixed time window.

[0036] As a further scheme of the present application, the step of the photovoltaic power generation fault risk prediction table is specifically:

[0037] S501: Call the power generation subarea group information in the detection group after intelligent matching, calculate the coverage efficiency of the group according to the ratio of the number of group coverage power generation key points to the number of subarea planning points, combine the equipment stay time and the interval between segments, generate the power generation subarea coverage efficiency value set;

[0038] S502: Based on the power generation subarea coverage efficiency value set, extract the group with high coverage efficiency, calculate the coincidence rate of the work section and the equipment, and generate the group priority sorting coefficient set;

[0039] S503: According to the group priority ranking coefficient set, the group without job label coverage integrity is eliminated, the detection task set is formed by section group, and the photovoltaic power generation fault risk prediction table is obtained.

[0040] The photovoltaic power generation fault risk prediction system based on big data processing comprises:

[0041] The equipment operation distribution module extracts equipment operation state, environmental parameters and original fault data based on the basic information of the photovoltaic power generation equipment, labels risk level labels, and generates a standardized equipment operation data set.

[0042] The risk area optimization module analyzes the device correlation strength based on the standardized equipment operation data set, combines the risk level label and the environmental parameter, filters the potential risk area, and generates a risk area data set.

[0043] The risk classification optimization module analyzes the dependency relationship between areas based on the risk area data set, estimates the risk conflict probability between areas, optimizes the risk classification, and generates a risk classification optimization result.

[0044] The detection group matching module extracts core equipment based on the risk classification optimization result, analyzes the consistency between devices, filters the consistent devices into the same detection group, and outputs the intelligent matched detection group.

[0045] The intelligent prediction execution module analyzes the group coverage efficiency based on the intelligent matched detection group, sorts the group priority according to the coverage efficiency, filters the optimal efficiency group into the execution plan, and generates a photovoltaic power generation fault risk prediction table.

[0046] Compared with the prior art, the advantages and positive effects of the present application are that:

[0047] In the present application, by deeply analyzing the operation state, environmental parameters and fault data of the equipment, combining the correlation between the equipment and the risk level label, the potential risk area is effectively identified and dynamically optimized, the limitation of relying on fixed rules is reduced, the flexibility and accuracy of the prediction are improved, the risk classification can be adjusted in real time and the change of different operating environments is responded to, the cooperation and consistency analysis between the equipment are optimized, the timeliness and effectiveness of the fault warning are enhanced, the detection group is intelligently matched and optimized, the coverage efficiency of the fault detection is improved, the accurate monitoring and risk prediction ability under large-scale complex environment are ensured, and the operation management effect and risk prediction ability of the photovoltaic power generation system are greatly improved. BRIEF DESCRIPTION OF DRAWINGS

[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed in the embodiments description. Obviously, the drawings in the following description only show some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained based on these drawings without any creative effort.

[0049] Figure 1 Schematic diagram of the step flow of the present application;

[0050] Figure 2 Schematic diagram of S1 refinement of the present application;

[0051] Figure 3 Schematic diagram of S2 refinement of the present application;

[0052] Figure 4 Schematic diagram of S3 refinement of the present application;

[0053] Figure 5 Schematic diagram of S4 refinement of the present application;

[0054] Figure 6 Schematic diagram of S5 refinement of the present application;

[0055] Figure 7 System module diagram of the present application. DETAILED DESCRIPTION

[0056] The technical solutions in the present application will be described below with reference to the drawings.

[0057] In the embodiments of the present application, the words such as "example", "for example" and the like are used to represent an example, illustration or description. Any embodiment or design scheme described as "example" in the present application should not be interpreted as more preferred or more advantageous than other embodiments or design schemes. Rather, the word "example" is intended to present the concept in a specific manner. In addition, in the embodiments of the present application, the meaning expressed by "and / or" can be both, or can be one of the two.

[0058] In the embodiments of the present application, "image" and "picture" can be used interchangeably at times. It should be pointed out that when the distinction is not emphasized, the meanings expressed are consistent. "Of", "corresponding" and "relevant" can be used interchangeably at times. It should be pointed out that when the distinction is not emphasized, the meanings expressed are consistent.

[0059] In the embodiments of the present application, sometimes the subscript such as W 1预估 is written in the form of non-subscript such as W1, and when the distinction is not emphasized, the meanings expressed are consistent.

[0060] In order to make the technical problems, technical solutions and advantages to be solved by the present application more clear, the following will be described in detail in combination with the drawings and specific embodiments.

[0061] Please refer to Figure 1 The embodiment of the present application provides a photovoltaic power generation fault risk prediction method based on big data processing, which comprises the following steps:

[0062] S1: Based on the basic information of the photovoltaic power generation equipment, the equipment running state, the environmental parameters and the original fault data are extracted, the spatial distribution characteristics of the equipment running are identified, and the risk level label is labeled, and the standardized equipment running data set is generated;

[0063] S2: Based on the standardized equipment running data set, the correlation strength between the equipment is analyzed, the risk level label and the environmental parameters are combined, the potential risk area is screened, and the risk area data set is generated;

[0064] S3: The risk area data set is called, the dependence between the areas is analyzed, the risk conflict probability between the areas is evaluated, the risk classification is adjusted, and the risk classification optimization result is obtained;

[0065] S4: Based on the risk classification optimization result, the core equipment is extracted, the consistency between the equipment is analyzed, the consistent equipment is screened into the same detection group, and the intelligent matched detection group is output;

[0066] S5: The intelligent matched detection group is called, the group coverage efficiency is analyzed, the group priority is sorted according to the coverage efficiency, the optimal efficiency group is screened into the execution plan, and the photovoltaic power generation fault risk prediction table is obtained.

[0067] The equipment running spatial distribution diagram comprises an equipment running state layer, an environmental parameter layer, an original fault data layer and a risk level identification layer, the standardized equipment running data set comprises equipment identification number, environmental parameter information, fault data label and risk level code, the risk area data set comprises area execution sequence table, inter-area correlation matrix and risk priority parameter, the risk classification optimization result comprises risk classification list, risk conflict probability table and execution time period distribution diagram, the intelligent matched detection group comprises equipment sequence, equipment consistency index and group aggregation grouping result, and the photovoltaic power generation fault risk prediction table comprises detection group number, group efficiency score and execution priority sequence.

[0068] Please refer to Figure 2 The step of the standardized equipment running data set is specifically:

[0069] S101: Based on the basic information of the photovoltaic power generation equipment, the equipment running state, the environmental parameters and the original fault data are extracted, the redundant information is removed, and the equipment running basic data is generated;

[0070] According to the photovoltaic power generation equipment, assuming that the basic information of array No. PV-A01 is clear, the equipment operation state data is clear that the direct current input voltage of inverter INV-001 is 550.2V, the direct current input current is 20.1A, the alternating current output power is 10.51kW, the conversion efficiency is 98.53%, the environmental parameter is clear that the light intensity meter reading is 953W / m², the component surface thermocouple temperature is 45.2℃, and the louver box environmental temperature is 32.1℃, the original fault data is clear that the alarm code “0x002A” recorded in the controller log corresponds to the grid voltage out-of-limit event, then, the redundant information removal processing is performed, the specific operation is that the same physical quantity collected by different sensors in the same collection period (the time stamp error is less than 10 milliseconds) is compared, the component surface temperature measured by sensor T-01 is 45.2℃, sensor T-02 measures the temperature at the same position as 45.3℃, the absolute difference between the two is 0.1℃, which is less than the set redundant judgment threshold value 0.5℃, the setting of the threshold value refers to the sensor accuracy level and historical data stability analysis, and the maximum normal fluctuation value in the 99% confidence interval is taken as the setting basis, after the judgment is passed, the collection data of T-01 is retained and the data of T-02 is marked as redundant, after the traversal processing of all collection points is completed, all non-redundant data is integrated to generate equipment operation basic data, this data contains fields: equipment ID, accurate time stamp, direct current voltage, direct current, alternating current power, conversion efficiency, light intensity, component temperature, environmental temperature, fault code.

[0071] S102: According to the equipment operation basic data, the spatial relationship between the equipment operation state and the environmental parameter is analyzed, the risk level label is marked, the risk distribution interval is generated combined with the fault data and the equipment attribute;

[0072] According to the device operation basic data, the operating parameters of the inverter INV-001 under the light intensity of 953 W / m² and the component temperature of 45.2℃ are analyzed and compared with the operating parameters of the physically adjacent inverter INV-002 under the light intensity of 948 W / m² and the component temperature of 46.1℃, it is found that the conversion efficiency (98.53%) of INV-001 decreases by 0.47% compared with the design nominal efficiency (99.00%), while the efficiency of INV-002 decreases by 0.38%, it is identified that there is a difference in performance degradation of the two under similar environmental input, then, the risk level label is marked, and the risk level determination rule is quantified as: when the relative attenuation rate of device conversion efficiency to nominal value is in the interval [0, 0.5%), the risk level is 1; in the interval [0.5%, 1.0%), the risk level is 2; in the interval [1.0%, ∞), the risk level is 3, since the efficiency attenuation rate of INV-001 is 0.47% in the interval [0, 0.5%), its risk level is marked as 1, and INV-002 is also 1, then, combined with the fault data “0x002A” (grid voltage abnormality) and the device attribute (inverter brand A is lower than brand B in the tolerance of grid harmonic), the risk level 1 state is strongly associated with the grid disturbance event and the specific model of the device, and a risk distribution interval containing environmental parameter range, efficiency attenuation interval, risk level, associated fault type and device model is generated.

[0073] S103: According to the risk distribution interval, analyze the environmental parameter characteristics according to the device classification, group the environment type, load state and fluctuation, summarize the key operation behavior, and generate a standardized device operation data set;

[0074] According to the risk distribution interval, for the inverter equipment category, the corresponding all environmental parameter records when the risk level is 2 are counted, the statistical distribution range of light intensity is [900, 1050] W / m², the component temperature distribution is [42, 55] ℃, then, the environment type, load state and fluctuation are grouped, the environment type is divided according to two-dimensional parameters, the environment with light intensity greater than 800 W / m² and component temperature greater than 40 ℃ is classified as “A type environment”, the load state is divided according to the ratio of real-time output power of inverter to its rated power (load rate), the state with load rate in the interval [0.9, 1.0] is classified as “heavy load”, the current 10.51kW output corresponds to 11kW rated power, the load rate is 95.5%, which belongs to “heavy load” state, the fluctuation grouping is realized by calculating the standard deviation of output power in 1 minute time window, if the standard deviation is greater than 1.5% of rated power (the value is set based on the 95th percentile of historical stable operation data standard deviation), it is recorded as “strong fluctuation”, then, the key operation behavior is summarized, specifically, “A type environment”, “heavy load” state, “strong fluctuation” grouping and risk level 2 are mapped to form a standardized key operation behavior record, finally, all the key operation behavior records of the equipment under different working conditions are collected to form a structured standardized equipment operation data set containing equipment ID, behavior mode code, environment classification code, load classification code, fluctuation classification code and risk level.

[0075] Please refer to Figure 3 , the steps of the risk area data set are as follows:

[0076] S201: Based on the device distribution of the standardized device operation data set, analyze the association strength between devices, filter potential risk areas according to the association strength value, and generate a device association strength matrix;

[0077] Based on the geographical distribution information of the equipment in the standardized equipment operation data set, the inverter INV-001 and INV-002 in the data set are selected, the standardized operation data record sequences in the past 48 hours are extracted, the frequencies of the two sequences taking the same values in four discrete dimensions of "environment classification code", "load classification code", "fluctuation classification code" and "risk level" are calculated, among the total of 2880 (one per minute) sampling points, 2304 sampling points of the two are in the same classification combination, the consistency frequency is calculated as 2304, the correlation strength value is defined as the consistency frequency divided by the total sampling points, that is, 2304 / 2880=0.8, the calculation is performed on all equipment pairs (N×(N-1) / 2 pairs) in the region, the potential risk area is screened according to the correlation strength value, and the correlation strength screening threshold is set to 0.75. The threshold is the minimum correlation strength value that can cover 90% of the cases determined by backtracking analysis of 50 historical linkage failure cases. Since the correlation strength value of INV-001 and INV-002 is 0.8, which is higher than 0.75, they are preliminarily included in the same potential risk area. Finally, the correlation strength values of all equipment pairs are organized into an N×N symmetric matrix, and the element in the (i, j) position of the matrix is the correlation strength value of equipment i and equipment j. The equipment correlation strength matrix is generated.

[0078] S202: Based on the equipment correlation strength matrix, the dependence between equipment is analyzed in combination with the risk level label and the environment parameter, the area meeting the risk priority requirement is screened, the number of equipment, the area span and the risk distribution data are counted, and the equipment distribution data set is obtained;

[0079] Based on the device correlation strength matrix, the device pair INV-001 (risk level 1) and INV-002 (risk level 1) with a correlation strength of 0.8 are extracted, and the high-resolution (1 second sampling rate) data stream in the "class A environment" is analyzed. Through cross-correlation analysis, it is found that when the DC side voltage of INV-001 appears a transient disturbance with a peak value of 5V and a duration of 2 seconds, the AC output current of INV-002 also presents a same trend fluctuation with a correlation coefficient of 0.6 after a delay of 1.5 seconds. This response relationship is determined as the operating dependency of INV-002 on INV-001. Then, the region that meets the risk priority requirement is screened out. The risk priority score is calculated by weighted summation, and the weights are set based on the analysis results of a database containing 5000 fault records. The weight of risk level is 0.6, and the weight of environmental parameter risk (1 if the component temperature is higher than 50°C, otherwise 0) is 0.4. The risk priority score of this region is (10.6+10.6) / 2+0×0.4=0.6. The priority screening threshold is set to 1.0, which corresponds to the average value of the priority score of the region that causes a loss of more than 5% of the power generation in the historical data. Since 0.6<1.0, this region is not screened out. Until a region with a priority score greater than 1.0 is found, then the number of devices in the screened region, the region span (the maximum Euclidean distance of the physical coordinates of the devices) and the risk distribution data (the count of devices of each risk level) are counted to obtain the device distribution data set.

[0080] S203: Call the device distribution data set, calculate the region optimization factor, and divide to obtain the risk region data set according to the number of devices, the region span and the risk distribution;

[0081] The device distribution dataset is called, and the area optimization factor is composed of the number of devices, the area span, and the risk distribution diversity, and the weight coefficient is set according to the cost-benefit analysis of the operation and maintenance simulation model. The weight of the number of devices is 0.4, the weight of the area span is -0.25 (the span increases to cause the factor to decrease), and the weight of the risk level diversity (the number of different risk levels in the area) is 0.35. For a region A containing 5 devices, the region span is 30 meters, and the risk level has two levels of 1 and 2, the normalized reference value is set as: the number of devices is 20, the span is 100 meters, and the number of risk levels is 3. The area optimization factor=(5 / 20)0.4+(1-30 / 100)(-0.25)+(2 / 3)×0.35=0.158. Then, according to the number of devices, the area span, and the risk distribution, the optimization factors of all areas are iteratively divided. The areas with an optimization factor difference less than 0.05 (the difference is set based on the inter-class distance of the clustering algorithm) and adjacent physical boundaries are merged. For example, the optimization factor of region B is 0.145, which is similar to region A and adjacent. Then, A and B are merged into a new region C. Conversely, if a region can be divided into two sub-regions, and the optimization factor of the sub-region is greater than 0.2 different from the factor of the parent region, then the splitting is performed. By performing the above merging and splitting operations on all areas, the optimization factors of all areas are stable in their respective convergence intervals, and the risk area dataset is obtained by division.

[0082] Please refer to Figure 4 The steps of the risk classification optimization result are as follows:

[0083] S301: Call the risk area dataset, analyze the dependency relationship between areas, classify according to the area type and the execution time window, and obtain the area dependency distribution data;

[0084] The risk area data set is called, two divided risk areas A (identified as "inverter high-frequency disturbance area") and area B (identified as "bus communication abnormal area") are selected, synchronous operation data in the same 24-hour cycle is extracted, time series of two area core indicators (inverter DC voltage standard deviation of area A and bus data reporting success rate of area B) are analyzed through Granger causality test, the test result shows that the change of the area A indicator leads the change of the area B indicator in time, and the P value is less than 0.05, this result shows that the running state of the area B has a statistical dependence on the area A, then, according to the area type and the execution time window, the area type is divided into "inverter high-frequency disturbance area" and "bus communication abnormal area", the execution time window is divided into "power increasing period" (07:00-10:00), "peak power period" (10:00-14:00) and "power decreasing period" (14:00-17:00) according to the typical sunshine curve of the power station, the dependence relationship obtained through the test is classified as the influence of "inverter high-frequency disturbance area" on "bus communication abnormal area", and it is marked that the influence mainly occurs in the "peak power period", causality test and classification are repeatedly performed on all areas, and a region dependence distribution data including a source area, a target area, a dependence strength, a lag time and an occurrence time window is obtained.

[0085] S302: Based on the region dependence distribution data, the risk conflict probability between regions is identified, the risk overlap part in the differentiated region is extracted, the overlap times and the risk conflict probability are counted, and a region risk conflict matrix is generated;

[0086] The risk overlap part refers to the intersection of two or more differentiated regions in risk characteristics, influence range or trigger condition when performing regional risk analysis;

[0087] The risk conflict probability refers to a quantitative index obtained by calculating the estimated risk event occurrence caused by mutual restriction in resources, time sequence or logic between differentiated regions;

[0088] Based on the regional dependence distribution data, the dependence relationship of region A (inverter high-frequency disturbance area) and region B (busbar communication abnormal area) is analyzed, and it is found that in the "peak power period", when the standard deviation of the DC voltage of the inverter in region A exceeds the threshold value 0.8V, the success rate of the busbar data reporting in region B in the next 5 minutes is less than 90%, and the conditional probability is 0.4, which is the initial risk conflict probability. Then, the risk overlap part in the differentiated region is extracted, and by analyzing the electrical wiring diagram and communication network topology of the power station, it is found that the inverter in region A and the busbar in region B are both powered by the same secondary distribution cabinet (number PDB-02). Therefore, "PDB-02 power supply abnormality" is identified as a risk overlap part of the two differentiated regions. Then, the overlap times and risk conflict probability are counted, and by querying the operation and maintenance log, "PDB-02 power supply abnormality" has been recorded for 10 times in the past 12 months, of which 7 times are accompanied by voltage disturbance in region A and communication interruption in region B. Therefore, the overlap times are 7, and the risk conflict probability based on this overlap event is corrected to 7 / 10=0.7. The conflict probabilities between all regions based on different overlap parts are calculated and integrated to generate a N×N regional risk conflict matrix, where N is the total number of regions.

[0089] S303: Call the regional risk conflict matrix, divide the regions according to the conflict probability, merge the low conflict regions, split the high conflict regions, and obtain the risk classification optimization result;

[0090] Low conflict region, refers to the region in the regional risk conflict matrix, whose risk conflict probability with other regions is generally lower than the preset merging threshold;

[0091] High conflict region, refers to the region in the regional risk conflict matrix, whose risk conflict probability inside and outside is higher than the preset splitting threshold;

[0092] The risk conflict matrix of the calling area is called, the low conflict merging threshold is set to 0.15, and the high conflict splitting threshold is set to 0.65. The setting of the threshold is based on a sensitivity analysis experiment on the efficiency of operation and maintenance resource allocation. The experimental data show that the comprehensive operation and maintenance cost of the area with a merging conflict probability lower than 0.15 can be reduced by 18%, and the risk omission rate increases by no more than 1%. The occurrence rate of cascading failures of the area with a splitting internal or external conflict probability higher than 0.65 can be reduced by 35%. The low conflict areas are merged. If the risk conflict probability value between area D and area E in the matrix is 0.12, which is lower than 0.15, and the physical inspection paths of the two areas can be continuous, then area D and area E are merged into a unified work unit F at the operation planning level. Then, the high conflict areas are split. If the average conflict probability between the devices in area A, or the conflict probability between area A and area B is 0.7, which is higher than 0.65, then the inspection tasks of area A are divided into two or more independent sub-tasks when the inspection tasks are formulated, and it is ensured that the tasks of area B do not overlap in execution time or required resources (such as specific tools). By judging all elements in the matrix and performing corresponding merging and splitting operations, the risk classification optimization result is obtained.

[0093] Please refer to Figure 5 The steps of the intelligent matching and post-detection group are as follows:

[0094] S401: Based on the risk classification optimization result, extract the core equipment, filter the consistent equipment in the group according to the consistency between the devices, and calculate the device consistency value;

[0095] The consistent equipment in the group refers to when the consistency index between the core equipment is higher than the preset consistency threshold, the core equipment is determined as a consistent equipment;

[0096] The device consistency value refers to the weighted average result of the average running efficiency, fault frequency and environmental adaptability of the core equipment within the preset monitoring period;

[0097] Based on the risk classification optimization results, first extract the core equipment in the optimized risk area G, the core equipment is determined according to its risk contribution, the risk contribution is weighted calculated by the risk level of the equipment itself and its centrality (such as PageRank score) in the regional risk propagation network, the weight is obtained by regression analysis of the original fault data, the risk level weight is 0.7, the centrality weight is 0.3, the highest score of the inverter INV-005 is calculated after the score, then, according to the consistency between devices, screen the consistency equipment in the group, set the consistency threshold to 0.90, the threshold is determined based on the distribution of 1000 same batch, same type equipment in the first year of operation in the power station The performance index of each is the 85th percentile, to ensure that the performance of the selected equipment is representative, then, calculate the device consistency value, specifically, calculate the consistency value of the other devices in the area G and the core equipment INV-005, the value is obtained by weighted average of three normalized indexes: average operating efficiency (weight 0.5), fault-free operation time (weight 0.3), environmental adaptability index (weight 0.2), such as the calculation result of device INV-006 is 0.93, since 0.93 is higher than the threshold 0.90, INV-006 is determined as the consistency equipment in the group.

[0098] S402: According to the device consistency value, analyze the consistency characteristics between devices, identify the behavior mode between adjacent devices, screen the consistency equipment combination according to the consistency and device distribution characteristics, and generate the device consistency data set;

[0099] The behavior mode between adjacent devices refers to constructing the behavior mode by analyzing the running state change, load fluctuation and environmental response characteristics of adjacent devices in a fixed time window.

[0100] According to the device consistency value, the high-frequency (1 Hz) operation data of the consistent devices INV-005 and INV-006 within a fixed time window (for example, 30 minutes) is extracted, the time sequence similarity of the two devices in the load fluctuation response is calculated by using a dynamic time warping (DTW) algorithm, when the light intensity is blocked by a cloud shadow event, the DTW distance of the output power curves of INV-005 and INV-006 is less than a preset distance threshold 0.2 (the threshold represents the maximum normal difference of the response curves of the same type of healthy devices), and the highly similar dynamic response is constructed as a specific behavior mode, named "cloud shadow synchronous response mode", then, according to the consistency and device distribution characteristics, the consistent device combination is screened, in the region G, it is identified that INV-008 (consistency value 0.91) also has the "cloud shadow synchronous response mode" with INV-005, and the three devices (INV-005, INV-006, and INV-008) form an adjacent triangular cluster (the distance between any two devices is less than 15 meters) in the physical layout, satisfying the compactness requirement of the device distribution, therefore, {INV-005, INV-006, and INV-008} are screened as an effective consistent device combination, and finally all the combinations meeting the conditions are collected to generate a device consistent data set.

[0101] S403: Based on the device consistent data set, the devices belonging to the same power generation region are clustered by the consistent devices, the same type of devices are merged, and the attribution groups are divided, to generate a detection group after intelligent matching;

[0102] Based on the device consistent data set, the consistent device combination {INV-005, INV-006, INV-008} is combined as an initial cluster core, and the "behavior-space distance" of the remaining devices (such as INV-009) in the power generation area from the cluster core is calculated, which is a comprehensive measure calculated as: 0.6 x (DTW distance of INV-009 and core behavior mode) + 0.4 x (normalized physical distance of INV-009 to cluster geometric center), the weights 0.6 and 0.4 are determined according to expert scoring method, emphasizing the dominant role of behavior similarity, if the "behavior-space distance" calculated by INV-009 is less than the set cluster radius threshold 0.3 (the threshold is optimized by the contour coefficient method), then INV-009 is absorbed into the cluster, then the same type of devices are combined and divided into attribution groups, after the iterative absorption process is completed, all devices belonging to the cluster of {INV-005, INV-006, INV-008} (assuming that it finally contains INV-009, INV-011, INV-012) are combined to form a detection group, and a unique identifier "G1-Alpha" is assigned, that is, an attribution group, repeat the clustering and merging process for all consistent device combinations until all devices in the area are assigned to an attribution group, generating the final intelligent matching detection group set.

[0103] Please refer to Figure 6 , the steps of the photovoltaic power generation fault risk prediction table are:

[0104] S501: Call the power generation subarea group information in the intelligent matching detection group, calculate the coverage efficiency of the group according to the ratio of the number of key points covered by the group to the number of subarea planning points, combined with the device stay time and the interval between segments, to generate a set of power generation subarea coverage efficiency values;

[0105] After calling the intelligent matching, the power generation sub-area group information in the detection group is detected. First, the configuration information of the "G1-Alpha" group (containing 6 inverters) is extracted, and its coverage efficiency is calculated. The power generation sub-area has a total of 15 key inspection points, which are determined according to the fault frequency and criticality evaluation. The 6 devices in the "G1-Alpha" group cover 6 key points, so the number of key points covered by the group is 6, and the ratio of the number of key points covered by the group to the number of points planned by the sub-area is 6 / 15=0.4. According to the inspection SOP, the standard inspection time (device stay time) of a single inverter is 4 minutes, and the average navigation and preparation time (segment interval) between devices in the group is 1.5 minutes. Therefore, the total inspection time of the group is 6x4+(6-1)x1.5=24+7.5=31.5 minutes. The coverage efficiency of the group is defined as the ratio of the key point coverage ratio to the total time, that is, 0.4 / 31.5≈0.0127. This index quantifies the proportion of key points covered in unit time. Repeat this calculation process for each detection group (such as G1-Beta, G2-Alpha, etc.). Collect all the calculation results (group ID and its corresponding coverage efficiency value) to generate a set of power generation sub-area coverage efficiency values.

[0106] S502: Based on the power generation sub-area coverage efficiency value set, extract the group with high coverage efficiency, calculate the coincidence rate of the work section and the device, and generate a group priority ranking coefficient set;

[0107] Based on the power generation sub-area coverage efficiency value set, all groups in the value set are arranged in descending order according to the coverage efficiency value. Set an extraction ratio, for example, the top 20%. If there are 30 groups, extract the top 6 groups with the highest efficiency into the next round of calculation. Assume that the "G1-Alpha" group (efficiency 0.0127) and the "G2-Beta" group (efficiency 0.0151) are both in this list. Then, the coincidence rate of the work section and the device is calculated. According to the risk level, the power station is divided into several work sections, such as "high-risk work section R1" containing 8 device points. There are 5 devices in the "G2-Beta" group located in the R1 section, so the coincidence rate is 5 / 8=0.625. Then calculate the group priority ranking coefficient. The coefficient is obtained by weighted sum of coverage efficiency and coincidence rate. The weight is dynamically adjusted according to the current operation and maintenance strategy. Under the "emergency response" strategy, the coincidence rate weight (w_r) is set to 0.7, and the coverage efficiency weight (w_e) is set to 0.3. The priority ranking coefficient of the "G2-Beta" group is 0.0151x0.3+0.625x0.7=0.44203. Perform this calculation on all extracted top groups to generate a group priority ranking coefficient set containing group ID and its corresponding priority ranking coefficient.

[0108] S503: According to the set of group priority ranking coefficients, eliminate groups without complete coverage of the job label, build a detection task set by section, and obtain a photovoltaic power generation fault risk prediction table;

[0109] According to the set of group priority ranking coefficients, the job label defines the necessary combination of equipment types for performing a specific inspection task, for example, the job label of the "DC side special inspection" task is {photovoltaic module, combiner box, DC cable}. If a detection group, even if its priority ranking coefficient is high, but its members only include inverters, then this group does not have complete coverage for the "DC side special inspection" task, and will be eliminated from the candidate groups for this task. Subsequently, among the groups that pass the integrity screening, according to the priority ranking coefficients from high to low, the detection task set is built by section, and the "G2-Beta" group with the highest ranking (coefficient 0.44203) is selected, which is assigned to the "high-risk operation section R1", generating the first detection task: "Task ID001 | Section R1 | Detection Group G2-Beta | Execution Window 09:00-10:00 | Responsible Person A". Then, the group with the next highest coefficient is selected and assigned to the section with the highest overlap rate corresponding to it, and so on, until all high-risk sections are assigned tasks or all eligible groups are scheduled. Finally, all generated detection tasks are arranged in chronological order and priority, forming a structured and directly executable inspection work plan, which is the photovoltaic power generation fault risk prediction table.

[0110] Please refer to Figure 7 The photovoltaic power generation fault risk prediction system based on big data processing includes:

[0111] The device operation distribution module extracts device operation status, environmental parameters, and original fault data based on the basic information of photovoltaic power generation equipment, labels risk level tags, and generates a standardized device operation data set.

[0112] The risk area optimization module analyzes the device correlation strength based on the standardized device operation data set, combines the risk level label and the environmental parameter, filters the potential risk area, and generates a risk area data set.

[0113] The risk classification optimization module analyzes the dependency relationship between areas based on the risk area data set, evaluates the risk conflict probability between areas, optimizes the risk classification, and generates a risk classification optimization result.

[0114] The detection group matching module extracts core equipment based on the risk classification optimization result, analyzes the consistency between devices, filters consistent devices into the same detection group, and outputs the intelligent matched detection group.

[0115] The intelligent prediction execution module analyzes the group coverage efficiency based on the intelligent matched detection group, sorts the group priority according to the coverage efficiency, filters the optimal efficiency group into the execution plan, and generates a photovoltaic power generation fault risk prediction table.

[0116] The above merely illustrates the specific embodiments of the present application, but the protection scope of the present application is not limited thereto, any person skilled in the art can easily think of the changes or replacements within the technical range disclosed by the present application, which should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A photovoltaic power generation fault risk prediction method based on big data processing, characterized in that, Includes the following steps: S1: Based on the basic information of photovoltaic power generation equipment, extract the equipment operating status, environmental parameters and original fault data, identify the spatial distribution characteristics of equipment operation, and label the risk level to generate a standardized equipment operation dataset; S2: Based on the standardized equipment operation dataset, analyze the correlation strength between equipment, combine risk level labels and environmental parameters to screen potential risk areas, and generate a risk area dataset; S3: Call the risk area dataset, analyze the dependencies between areas, assess the probability of risk conflicts between areas, adjust the risk classification, and obtain the risk classification optimization result; S4: Based on the risk classification optimization results, extract the core equipment, analyze the consistency between the equipment, screen the consistent equipment and classify it into the same detection group, and output the intelligent matching detection group; S5: After calling the intelligent matching detection group, analyze the group coverage efficiency, sort the group priority according to the coverage efficiency, select the optimal efficiency group and include it in the execution plan to obtain the photovoltaic power generation fault risk prediction table.

2. The photovoltaic power generation fault risk prediction method based on big data processing according to claim 1, characterized in that, The equipment operation space distribution map includes an equipment operation status layer, an environmental parameter layer, a raw fault data layer, and a risk level identification layer. The standardized equipment operation dataset includes equipment identification number, environmental parameter information, fault data labels, and risk level codes. The risk area dataset includes a regional execution sequence table, an inter-regional correlation matrix, and risk priority parameters. The risk classification optimization results include a risk classification list, a risk conflict probability table, and an execution time distribution map. The intelligent matching detection group includes equipment sequence, equipment consistency index, and group aggregation results. The photovoltaic power generation fault risk prediction table includes a detection group number, group efficiency score, and execution priority sequence.

3. The photovoltaic power generation fault risk prediction method based on big data processing according to claim 1, characterized in that, The steps for creating the standardized equipment operation dataset are as follows: S101: Based on the basic information of photovoltaic power generation equipment, extract the equipment operating status, environmental parameters and original fault data, remove redundant information and generate basic equipment operating data; S102: Based on the basic equipment operation data, analyze the spatial relationship between equipment operation status and environmental parameters, label risk level, and generate risk distribution intervals by combining fault data and equipment attributes; S103: Based on the risk distribution range, analyze the environmental parameter characteristics according to equipment classification, group the environmental type, load status and fluctuation, summarize key operating behaviors, and generate a standardized equipment operation dataset.

4. The photovoltaic power generation fault risk prediction method based on big data processing according to claim 3, characterized in that, The specific steps for creating the risk area dataset are as follows: S201: Based on the equipment distribution of the standardized equipment operation dataset, analyze the correlation strength between equipment, filter potential risk areas according to the correlation strength value, and generate an equipment correlation strength matrix; S202: Based on the device association strength matrix, combined with risk level labels and environmental parameters, analyze the dependencies between devices, screen areas that meet the risk priority requirements, and statistically analyze the number of devices, area span, and risk distribution data to obtain the device distribution dataset; S203: Call the device distribution dataset, calculate the regional optimization factor, and divide the risk area dataset according to the number of devices, regional span and risk distribution.

5. The photovoltaic power generation fault risk prediction method based on big data processing according to claim 4, characterized in that, The specific steps for optimizing the risk classification results are as follows: S301: Call the risk area dataset, analyze the dependencies between areas, classify by area type and execution time window, and obtain area dependency distribution data; S302: Based on the regional dependent distribution data, identify the probability of risk conflict between regions, extract the risk overlap part in the differentiated regions, count the number of overlaps and the probability of risk conflict, and generate a regional risk conflict matrix; S303: Call the aforementioned regional risk conflict matrix, divide the regions according to the conflict probability, merge low-conflict regions, split high-conflict regions, and obtain the risk classification optimization results.

6. The photovoltaic power generation fault risk prediction method based on big data processing according to claim 5, characterized in that, The overlapping risk portion refers to the intersection of two or more differentiated regions in terms of risk characteristics, scope of impact, or triggering conditions when conducting regional risk analysis. The risk conflict probability refers to a quantitative indicator obtained by calculating the predictability of risk events occurring simultaneously or consecutively between differentiated regions due to mutual constraints in resources, timing, or logic. The low-conflict area refers to the area in the regional risk conflict matrix whose probability of risk conflict with other areas is generally lower than the preset merging threshold. The high-conflict region refers to a region in the regional risk conflict matrix where the probability of risk conflict between itself and other regions is higher than a preset splitting threshold.

7. The photovoltaic power generation fault risk prediction method based on big data processing according to claim 5, characterized in that, The specific steps for the intelligent matching and detection group are as follows: S401: Based on the risk classification optimization results, extract the core equipment, and based on the consistency between equipment, screen the consistent equipment in the group and calculate the equipment consistency value; S402: Based on the device consistency value, analyze the consistency characteristics between devices, identify the behavior patterns between adjacent devices, and based on consistency and device distribution characteristics, filter consistent device combinations to generate a device consistency dataset; S403: Based on the device consistency dataset, cluster devices belonging to the same power generation area through consistent devices, merge similar devices and divide them into groups, and generate intelligent matching detection groups.

8. The photovoltaic power generation fault risk prediction method based on big data processing according to claim 7, characterized in that, The term "consistent device in the group" refers to a core device being identified as a consistent device when the consistency index among core devices is higher than a preset consistency threshold. The equipment consistency value refers to the weighted average result of the core equipment's average operating efficiency, failure frequency, and environmental adaptability within a preset monitoring period. The behavioral patterns between adjacent devices refer to the behavioral patterns constructed by analyzing the changes in the operating status, load fluctuations, and environmental response characteristics of adjacent devices within a fixed time window.

9. The photovoltaic power generation fault risk prediction method based on big data processing according to claim 7, characterized in that, The specific steps for creating the photovoltaic power generation fault risk prediction table are as follows: S501: Call the power generation zone group information in the intelligent matching detection group, calculate the group's coverage efficiency based on the ratio of the number of key power generation points covered by the group to the number of planned points in the zone, and combine the equipment dwell time and inter-segment interval, and generate a set of power generation zone coverage efficiency values. S502: Based on the set of coverage efficiency values ​​of the power generation zone, extract the top group of coverage efficiency, calculate the overlap rate between the working section and the equipment, and generate a set of group priority ranking coefficients. S503: Based on the group priority sorting coefficient set, eliminate groups that do not have complete coverage of operation labels, build detection task sets by segment, and obtain the photovoltaic power generation fault risk prediction table.

10. A photovoltaic power generation fault risk prediction system based on big data processing, characterized in that, The system is used to implement the photovoltaic power generation fault risk prediction method based on big data processing as described in any one of claims 1-9, and the system includes: The equipment operation distribution module extracts equipment operating status, environmental parameters and raw fault data based on the basic information of photovoltaic power generation equipment, labels risk level, and generates a standardized equipment operation dataset. The risk area optimization module analyzes the correlation strength of equipment based on the standardized equipment operation dataset, combines risk level labels and environmental parameters to screen potential risk areas, and generates a risk area dataset. The risk classification optimization module analyzes the dependencies between regions based on the risk region dataset, assesses the probability of risk conflicts between regions, optimizes risk classification, and generates risk classification optimization results. Based on the risk classification optimization results, the detection group matching module extracts core equipment, analyzes the consistency between equipment, filters consistent equipment and assigns it to the same detection group, and outputs the intelligently matched detection group. The intelligent prediction and execution module analyzes the coverage efficiency of the intelligent matching detection group, prioritizes the groups according to their coverage efficiency, selects the group with the best efficiency and includes it in the execution plan, and generates a photovoltaic power generation fault risk prediction table.

Citation Information

Patent Citations

  • Photovoltaic power station inspection operation and maintenance method, device, equipment and medium

    CN118365305A

  • Plateau area well logging operation automation system

    CN119102611A