A big data-based distributed photovoltaic data analysis method and system
Patent Information
- Application Number
- CN202611039018.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-14
- Publication Date
- 2026-09-29
AI Technical Summary
[0007]本发明提供了一种基于大数据的分布式光伏数据分析方法及系统,以至少解决因传统数据清洗和特征工程局限性导致“伪正常数据”被误判为正常数据,进而使控制策略响应滞后的问题
通过根据运行数据和物理特性建立随运行条件动态变化的运行表现参考,可精确识别并量化被传统数据清洗方法误判为正常的系统性偏差(包括光伏组件的非线性衰减模式、逆变器固件升级引入的周期性功率波动以及传感器系统性测量漂移),根据偏差的类型、幅度及影响范围自适应调整控制参数,并通过效果验证和优化形成闭环反馈机制,从而避免了因“伪正常数据”导致的控制策略响应滞后问题,显著提升了分布式光伏系统的发电效率、电网适应性以及储能配合策略的经济性。
Smart Images

Figure CN122844271A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of distributed photovoltaic data analysis technology, and in particular to a distributed photovoltaic data analysis method and system based on big data. Background Technology
[0002] The daily operation of distributed photovoltaic power stations involves multi-dimensional data such as photovoltaic module output power, inverter conversion efficiency, ambient irradiance, module backsheet temperature, grid connection point voltage and frequency. The accurate analysis and utilization of this data is the foundation for evaluating power generation efficiency, maintaining grid adaptability, and scheduling energy storage equipment.
[0003] Before data enters the analysis platform, it typically undergoes a data cleaning process. This process uses statistical characteristics calibrated under long-term stable operating conditions as a benchmark, sets outlier removal thresholds, and is mainly used to filter out outlier data points caused by occasional acquisition noise, instantaneous communication errors, or random sensor failures. Data that has undergone this process is considered "clean," and subsequent performance evaluations and control strategies are based on this foundation.
[0004] However, the effectiveness of the aforementioned data cleaning process heavily relies on the premise that the removed deviations must be sporadic, random, and significant. In actual operation, three types of deviations do not meet this premise. First, some batches of photovoltaic modules use new encapsulation materials. Under long-term outdoor high temperatures, temperature cycling, and ultraviolet radiation, the interface characteristics between the materials and the cells undergo slow and non-linear changes, resulting in small but persistent power deviations under specific operating conditions. Second, the grid support function introduced by inverter firmware upgrades can generate periodic small power fluctuations under certain load changes, lasting only tens to hundreds of milliseconds. Third, after long-term outdoor operation, a large number of deployed temperature sensors experience consistent parameter drift in their internal thermistor materials, forming systematic measurement deviations. These three types of deviations share the following characteristics: their amplitude is much smaller than the traditional cleaning threshold, and they are systematic, not random but strongly correlated with operating conditions, time, or equipment batches. Under these characteristics, they are misjudged as normal data by the data cleaning process and mixed into the data stream for subsequent analysis.
[0005] When the aforementioned "pseudo-normal data" enters the input layer of the power generation prediction model and energy storage dispatch strategy, the system behavior patterns learned by the model actually include these unrecognized systematic biases. The deviation between the prediction results and the actual system state accumulates over time, eventually leading to lag in control strategy response, inaccurate timing of energy storage charging and discharging, deviations in reactive power compensation, and unreasonable grid-connected power limits. How to perceive and quantify these hidden biases at the data level, under the constraint that data cleaning processes cannot effectively identify them, and then adaptively adjust control parameters to correct their impact on upper-level strategies, is a practical problem faced by the operation and maintenance of distributed photovoltaic systems.
[0006] To address the aforementioned technical problems, this invention provides a distributed photovoltaic data analysis method and system based on big data. Summary of the Invention
[0007] This invention provides a distributed photovoltaic data analysis method and system based on big data, which at least solves the problem that "pseudo-normal data" is misjudged as normal data due to the limitations of traditional data cleaning and feature engineering, thus causing the control strategy response to lag.
[0008] In a first aspect, the present invention provides a distributed photovoltaic data analysis method based on big data, comprising the following steps: Obtain operational data from distributed photovoltaic systems; Based on the operational data and the physical characteristics of the distributed photovoltaic system, an operational performance reference is established that varies with operating conditions, and operational deviations that deviate from the operational performance reference in a specific pattern are identified. The operational deviations include type, magnitude, and scope of impact. Based on the type, magnitude, and scope of the operational deviation, adjust the control parameters in the distributed photovoltaic system; The effects of the adjusted control parameters are verified, and the control parameters are optimized based on the results of the verification.
[0009] Optionally, the step of establishing an operational performance reference that varies with operating conditions based on the operational data and the physical characteristics of the distributed photovoltaic system includes: Acquire environmental irradiance data and environmental temperature data; Based on the ambient temperature data, the ambient irradiance data is corrected to obtain the corrected irradiance data; Based on the corrected irradiance data and the physical characteristics of the distributed photovoltaic system, an operational performance reference is established that varies with operating conditions.
[0010] Optionally, the step of verifying the effect of the adjusted control parameters and optimizing the control parameters based on the results of the verification includes: Identify the low-impact periods of the distributed photovoltaic power station, generate and send control pulses during the low-impact periods; During the application of the control pulse, temperature sensor readings in the area affected by the control pulse are acquired at high frequency to obtain the actual response curve of the sensor. Based on the current environmental conditions, the change in charging and discharging power of the energy storage system, and the known batch characteristics and aging trend of the target sensor, the expected response curve of the sensor under ideal conditions is generated, and the actual response curve of the sensor is compared with the expected response curve to quantify the deviation between the two. The calibration factor of the sensor under the current operating conditions is corrected, the corrected sensor data is fed back to the upper control system, and the control parameters of the upper control system are adaptively adjusted according to the corrected sensor data. The adjusted control parameters are verified in two layers. The two-layer verification includes continuously monitoring the response of the sensor in subsequent control pulse detection tests to confirm the continued reliability of the calibration factor, and comparing the overall operating performance of the adjusted system with the expected target.
[0011] Optionally, the step of identifying operational deviations that deviate from the operational performance reference in a specific pattern includes: Collect the deviation sequence between actual operating data and the operating performance reference; The evolution characteristics of the deviation sequence are extracted by combining environmental conditions, system operation mode, and equipment operating time; Update the preset deviation pattern feature library based on the evolution characteristics of the deviation sequence; The deviation sequence is matched with the updated deviation pattern feature library to identify the pattern of the operational deviation, determine the type of the operational deviation, quantify the magnitude of the operational deviation, and assess the scope of the impact of the operational deviation.
[0012] Optionally, the step of identifying operational deviations that deviate from the operational performance reference in a specific pattern includes: Collect the deviation sequence between actual operating data and the operating performance reference; Pattern analysis is performed on the deviation sequence to identify periodic, trend, or transient characteristics; Based on the physical characteristics and interaction relationships of the components within the distributed photovoltaic system, as well as the current environmental conditions, the deviation is decomposed into multiple independent sub-deviations; Quantify the magnitude of the independent sub-biases; The influence range of the independent sub-bias is assessed based on the type and magnitude of the independent sub-bias.
[0013] Optionally, the step of identifying operational deviations that deviate from the operational performance reference in a specific pattern includes: Collect the deviation sequence between actual operating data and the operating performance reference, and detect the presence of patterns in the deviation sequence that do not match the current pattern library; From the historical data of the long-term operation of distributed photovoltaic systems, extract historical deviation sequences and corresponding historical deviation characteristics that are similar to the mismatch patterns in terms of time, environmental conditions or equipment type; The historical deviation features are compared with the features of the mismatch patterns to identify candidate deviation patterns; Add the key parameters of the candidate deviation patterns to the deviation pattern feature library; The historical data was re-analyzed to verify the accuracy and coverage of the new integrated pattern.
[0014] Optionally, the step of performing two-level verification on the adjusted control parameters includes: Record the sensor calibration factor and its changing trend over time; Record the long-term trend of the overall system performance; By comparing the changing trend of the sensor calibration factor with the long-term trend of the overall system performance, the difference between the two is quantified. Identify the potential factors that cause the discrepancies and generate an inconsistency report.
[0015] Optionally, the step of identifying potential factors leading to the difference includes: Perform pattern analysis on the differential data to identify periodic, trend, or transient patterns; Based on the physical characteristics, interaction relationships, and current environmental conditions of the components within the distributed photovoltaic system, a physical interconnection structure is constructed. Based on the physical correlation structure, the independent influence components of different potential factors are extracted from the differential data; Quantify the contribution of each independent influencing component to the difference; When nonlinear coupling is identified among multiple potential factors, the interaction parameters in the physical association structure are adjusted, and the type, contribution, and coupling relationship of each potential factor are output.
[0016] Optionally, the step of quantifying the difference between the trend of the sensor calibration factor change and the long-term trend of the overall system performance includes: The system acquires environmental event records for a distributed photovoltaic system within a specific time period, including local power grid failures and extreme weather events. A preliminary analysis is conducted on the trend data of the sensor calibration factor change and the long-term trend data of the overall system performance to identify whether there are any instantaneous, non-periodic data fluctuations that coincide with the environmental event records in time. When the transient, non-periodic data fluctuations are identified, it is determined whether the data fluctuations are caused by an external sudden event, based on the type and occurrence time of the environmental event record. When the data fluctuations are caused by external unforeseen events, the data fluctuations should be isolated or corrected before differential quantification. The differences between the isolated or corrected sensor calibration factor change trend data and the long-term trend data of the overall system performance are quantified by comparing the two.
[0017] Secondly, the present invention provides a distributed photovoltaic data analysis system based on big data, the system comprising: The data acquisition module is used to acquire the operating data of the distributed photovoltaic system; The deviation identification module is used to establish an operational performance reference that varies with operating conditions based on the operational data and the physical characteristics of the distributed photovoltaic system, and to identify operational deviations that deviate from the operational performance reference in a specific pattern. The operational deviations include type, magnitude, and scope of influence. The parameter adjustment module is used to adjust the control parameters in the distributed photovoltaic system according to the type, magnitude, and scope of influence of the operating deviation. The effect verification and optimization module is used to verify the effect of the adjusted control parameters and optimize the control parameters based on the results of the effect verification.
[0018] Compared with existing technologies, the distributed photovoltaic data analysis method and system based on big data provided by this invention has at least the following beneficial technical effects: By establishing a dynamic performance reference based on operational data and physical characteristics, which changes with operating conditions, systematic deviations that are misjudged as normal by traditional data cleaning methods can be accurately identified and quantified (including nonlinear degradation modes of photovoltaic modules, periodic power fluctuations introduced by inverter firmware upgrades, and systematic measurement drift of sensors). Control parameters are adaptively adjusted according to the type, magnitude, and scope of the deviation, and a closed-loop feedback mechanism is formed through effect verification and optimization. This avoids the problem of control strategy response lag caused by "pseudo-normal data," and significantly improves the power generation efficiency, grid adaptability, and economic efficiency of energy storage coordination strategies of distributed photovoltaic systems.
[0019] The above and other features, objects and advantages of the present invention will be more clearly set forth in the following detailed description taken in conjunction with the accompanying drawings. Attached Figure Description
[0020] Figure 1 This is a flowchart illustrating a distributed photovoltaic data analysis method based on big data, according to an exemplary embodiment.
[0021] Figure 2 This is a flowchart illustrating step S21 according to an exemplary embodiment.
[0022] Figure 3 This is a flowchart illustrating step S41 according to an exemplary embodiment.
[0023] Figure 4 This is a flowchart illustrating step S241 according to an exemplary embodiment.
[0024] Figure 5 This is a flowchart illustrating step S251 according to an exemplary embodiment.
[0025] Figure 6 This is a flowchart illustrating step S261 according to an exemplary embodiment.
[0026] Figure 7 This is a block diagram illustrating a distributed photovoltaic data analysis system based on big data, according to an exemplary embodiment. Detailed Implementation
[0027] To make the objectives, technical solutions, and advantages of this application clearer, the application is described and illustrated below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application. All other embodiments obtained by those skilled in the art based on the embodiments provided in this application without inventive effort are within the scope of protection of this application.
[0028] Obviously, the accompanying drawings described below are merely some examples or embodiments of this application. Those skilled in the art can apply this application to other similar scenarios based on these drawings without any inventive effort. Furthermore, it is understood that although the efforts made in this development process may be complex and lengthy, for those skilled in the art related to the content disclosed in this application, any changes to design, manufacturing, or production based on the technical content disclosed in this application are merely conventional technical means and should not be construed as insufficient disclosure of this application.
[0029] In this application, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment that is mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described in this application may be combined with other embodiments without conflict.
[0030] Unless otherwise defined, the technical or scientific terms used in this application shall have the ordinary meaning understood by one of ordinary skill in the art to which this application pertains. The terms “a,” “an,” “an,” “the,” and similar words used in this application do not indicate quantity limitation and may indicate singular or plural. The terms “comprising,” “including,” “having,” and any variations thereof used in this application are intended to cover non-exclusive inclusion.
[0031] Example 1 Embodiment 1 of the present invention provides a distributed photovoltaic data analysis method based on big data. Figure 1 This is a flowchart illustrating a distributed photovoltaic data analysis method based on big data, according to an exemplary embodiment. Figure 1 As shown, the method includes the following steps: S10, Obtain operational data from distributed photovoltaic power stations; In this step, operational data is acquired through various sensors and data acquisition devices deployed in the distributed photovoltaic power station. This operational data includes, but is not limited to, real-time output power of the photovoltaic modules, ambient irradiance (obtained via a standard radiometer installed near the photovoltaic array), module backsheet temperature (obtained via a temperature sensor on the module backsheet), inverter output voltage and current, and voltage and frequency at the grid interface. The data is transmitted in real-time to the central data processing unit via wired or wireless communication networks, or it can be acquired through data interface integration with existing monitoring or SCADA systems. The system continuously receives operational data from the distributed photovoltaic power station, providing a data foundation for subsequent deviation identification and control parameter adjustments.
[0032] S20, Based on the operating data and the physical characteristics of the distributed photovoltaic power station, establish an operating performance reference that varies with operating conditions, and identify operating deviations that deviate from the operating performance reference in a specific pattern, wherein the operating deviations include type, magnitude and scope of influence; In this step, the system establishes a dynamic "ideal physical behavior reference" for each photovoltaic module. This reference is not a fixed value, but rather a calculation based on real-time environmental irradiance, module backsheet temperature, and other operating conditions. It utilizes the module's physical characteristic parameters, including the maximum power, temperature coefficient, and irradiance coefficient under standard test conditions, to determine the ideal output power the module should have under the current operating conditions. In this embodiment, the calculation is performed using the following formula: in, This represents the ideal output power (W) of the component under current operating conditions. This represents the maximum power (W) of the component under standard test conditions. Real-time ambient irradiance (W / m²) 2 ), The standard test irradiance is 1000W / m. 2 , The power temperature coefficient (% / ℃) of the component. The actual temperature (°C) of the module's solar cells. The standard test temperature is 25℃.
[0033] The system will collect the actual component output power. and By comparison, the real-time deviation is obtained: in, The real-time deviation value (W) is the value of the deviation. The actual output power (W) of the component. This represents the ideal output power (W) of the component under current operating conditions. When the deviation... When a deviation persists and, although its magnitude is small, exhibits a specific time or operating condition correlation pattern, even if its magnitude is within the traditional outlier removal threshold of ±5%, the system will still identify it as a specific pattern of operational deviation signal. The system further analyzes the characteristics of the deviation signal, including the duration of the deviation, its frequency of occurrence, its correlation with specific environmental conditions, and its distribution on different components or inverters. It categorizes these deviations into specific types such as "component degradation mode," "inverter fluctuation mode," or "sensor drift mode," and quantifies the deviation magnitude and assesses the potential impact on power generation efficiency.
[0034] S30, adjust the control parameters in the distributed photovoltaic power station according to the type, magnitude and scope of the operational deviation; In this step, the system triggers control parameter reconfiguration logic based on the categorized operational deviation information. For example, when a systemic underreporting phenomenon is detected in a certain area's temperature sensors, i.e., a continuous underreporting of approximately 2°C, the temperature-related charging limit threshold in the energy storage dispatch strategy will be automatically adjusted. If the original threshold was set at 50°C, the system will adjust it to 48°C, ensuring that the charging limit takes effect promptly when the actual component temperature reaches 50°C. When it is found that the power generation prediction model continuously overestimates the actual power generation potential due to the nonlinear degradation of components, the reactive power compensation benchmark or grid-connected power limit in grid adaptability adjustment is lowered accordingly. When a periodic small power fluctuation caused by an inverter firmware upgrade is detected, the voltage compliance judgment boundary value is fine-tuned to distinguish between internal fluctuations and genuine grid anomalies. This adjustment is direct and targeted, avoiding the problem of control strategy response lag due to prediction value deviations.
[0035] S40, verify the effect of the adjusted control parameters, and optimize the control parameters based on the results of the effect verification; In this step, the adjusted control parameters do not take effect immediately. Instead, short-term simulations and monitoring of actual operation are conducted through real-time effect verification logic. The verification logic continuously collects adjusted system operation data and compares it with historical data before adjustment and the expected target based on the corrected predicted values. For example, after adjusting the energy storage charging limit, the actual battery temperature, charging power, and health status are monitored; after adjusting the reactive power compensation benchmark, voltage stability and power factor at the grid interface are monitored. If the adjustment effect is found to be unsatisfactory or new problems are introduced, the verification logic immediately feeds back to the control parameter reconfiguration logic for further fine-tuning until the optimal control effect is achieved. This closed-loop mechanism of "perception-decision-verification-optimization" ensures the reliability and safety of control strategy adjustments.
[0036] In the technical solution of the above embodiments, by acquiring the operating data of the distributed photovoltaic power station and establishing an operating performance reference that dynamically changes with operating conditions based on physical characteristics, it is possible to accurately identify and quantify systematic deviations that are misjudged as normal by traditional data cleaning methods, including nonlinear degradation of components, periodic power fluctuations of inverters, and systematic measurement drift of sensors. The control parameters are adaptively adjusted according to the type, magnitude, and scope of influence of the deviations, and a closed-loop feedback mechanism is formed through effect verification and optimization. This overcomes the fundamental problem that traditional fixed threshold cleaning schemes cannot cope with "pseudo-normal data" and thus cause the control strategy to lag in response.
[0037] In one possible design, Figure 2 This is a flowchart illustrating step S20, which establishes an operational performance reference that varies with operating conditions, according to an exemplary embodiment. (Refer to...) Figure 2In step S20, the step of establishing an operational performance reference that varies with operating conditions based on the operational data and the physical characteristics of the distributed photovoltaic power station includes: S21, acquire ambient irradiance data and ambient temperature data; In this step, ambient irradiance and ambient temperature data are acquired using a standard radiometer and temperature sensor installed near the photovoltaic array. Irradiance data reflects the actual intensity of solar radiation reaching the module surface, while temperature data reflects the temperature level of the surrounding atmosphere; both are fundamental physical quantities for calculating the ideal output power of the module.
[0038] S22, Based on the ambient temperature data, correct the ambient irradiance data to obtain the corrected irradiance data; In this step, the actual output power of the photovoltaic module is affected not only by irradiance but also by the module's operating temperature. The irradiance measured by the standard radiometer is full-spectrum irradiance, while the absorption efficiency of photovoltaic modules for different wavelengths of light varies under different temperature conditions. The system performs temperature compensation correction on the original irradiance data based on ambient temperature data. In this embodiment, the calculation is performed using the following formula: in, For the corrected irradiance data (W / m 2 ), The original measured irradiance data (W / m 2 ), This is the temperature correction factor for irradiance (typically ranging from −0.05% / ℃ to −0.15% / ℃). The ambient temperature is (°C). The reference temperature is 25°C. Correction is applied to eliminate the influence of temperature on the accuracy of irradiance measurements, resulting in more accurate corrected irradiance data.
[0039] S23. Based on the corrected irradiance data and the physical characteristics of the distributed photovoltaic power station, establish an operational performance reference that varies with operating conditions.
[0040] In this step, based on the corrected irradiance data And the physical characteristics of the components, such as the maximum power under standard test conditions. Power temperature coefficient Using the aforementioned formula for the ideal output power of the components: The system calculates a performance reference for the computing components. This reference is dynamically updated in real time as operating conditions such as irradiance and component temperature change, providing an accurate baseline for deviation identification.
[0041] In the technical solution of the above embodiments, by acquiring environmental irradiance and temperature data, temperature compensation correction is performed on the irradiance data, and then combined with the physical characteristics of the components to establish an operating performance reference that changes with operating conditions, the operating performance reference can more accurately reflect the expected performance of the components under the current real operating conditions, thereby improving the accuracy of subsequent deviation identification.
[0042] In one example, suppose a distributed photovoltaic (PV) power station is located in an area with large diurnal and seasonal temperature variations. At midday in summer, despite high irradiance, the surface temperature of the modules can reach 60°C to 70°C, leading to a decrease in module efficiency. If only the raw irradiance data is used to establish an operational performance reference, the reference model will overestimate the system's expected power generation at high temperatures. According to this solution, firstly, ambient irradiance and ambient temperature data are acquired. At midday in summer, an irradiance meter measures 1000 W / m², while a temperature sensor measures the module surface temperature at 65°C. Based on a pre-set temperature correction model (based on the module's temperature coefficient, power decreases by approximately 0.4% for every 1°C increase), the raw irradiance data is corrected. The corrected irradiance data more accurately reflects the energy input that the PV modules can actually effectively utilize at 65°C; the corrected effective irradiance is approximately 900 W / m². Ultimately, the operational performance reference established based on the corrected irradiance improves the accuracy of power generation prediction and avoids the overestimation problem caused by high temperatures.
[0043] In one possible design, Figure 3 This is a flowchart illustrating the effect verification and optimization in step S40 according to an exemplary embodiment. (Refer to...) Figure 3 In step S40, the step of verifying the effect of the adjusted control parameters and optimizing the control parameters based on the results of the effect verification includes: S41, identify the low-impact period of the distributed photovoltaic power station, generate a control pulse during the low-impact period and send the control pulse; In this step, the low-impact period refers to the period with the least impact on power generation, such as nighttime periods without sunlight, periods of extremely low irradiance under cloudy weather, or periods of low grid load. Performing control pulse testing during the low-impact period avoids any substantial impact on normal power generation and grid operation during the testing process. A control pulse is a predefined, short-duration control command signal, such as an adjustment of energy storage charging / discharging power or a fine-tuning of inverter reactive power compensation lasting from several seconds to several minutes, used to trigger a controllable response in the system.
[0044] S42, During the application of the control pulse, the temperature sensor readings of the area affected by the control pulse are acquired at high frequency to obtain the actual response curve of the sensor; In this step, temperature sensor readings within the target area are acquired at a higher sampling frequency than usual, such as 1Hz to 10Hz. The temperature change process of the sensor before and after the application of the control pulse is recorded, and the actual response curve of the sensor is obtained. High-frequency acquisition can capture minute temperature changes caused by control pulses.
[0045] S43. Based on the current environmental conditions, the change in charging and discharging power of the energy storage system, and the known batch characteristics and aging trend of the target sensor, generate the expected response curve that the sensor should have under ideal conditions, and compare the actual response curve of the sensor with the expected response curve to quantify the deviation between the two. In this step, the system utilizes the known batch characteristics of the sensor, including initial calibration parameters, response time constant, and aging trend model—that is, the sensor performance degradation curve established based on historical data—combined with current environmental conditions such as ambient temperature, wind speed, and the energy change caused by the control pulse, to calculate the sensor's ideal temperature response curve. The actual response curve Compared with the expected response curve The comparison is performed in the time domain to quantify the deviations between the two in dimensions such as response amplitude, response delay time, and steady-state error. In this embodiment, the calculation is performed using the following formula: in, Root mean square deviation (°C) The total number of sampling points. For the first The actual temperature readings (°C) at each sampling point. Let be the expected temperature reading (°C) at the i-th sampling point. The overall deviation between the actual response and the expected response is quantified by the root mean square deviation.
[0046] S44, correct the calibration factor of the sensor under the current operating conditions, feed back the corrected sensor data to the upper-level control data processing device, and adaptively adjust the control parameters of the upper-level control system according to the corrected sensor data. In this step, based on the quantified deviation... Based on its evolution over time, the correction calibration factor of the sensor is calculated. In this embodiment, it is calculated using the following formula: in, This is the corrected calibration factor. The calibration factor before correction. Root mean square deviation (°C) The reference temperature is ℃. The corrected sensor data, i.e., the data after applying the new calibration factor, is fed back to the upper-level control data processing device. Based on this, the upper-level control system reassesses the system's operating status and adaptively adjusts control parameters such as energy storage scheduling and reactive power compensation.
[0047] S45, Perform two-layer verification on the adjusted control parameters. The two-layer verification includes continuously monitoring the response of the sensor in the subsequent control pulse detection test to confirm the continuous reliability of the calibration factor, and comparing the overall operating performance of the adjusted system with the expected target. In this step, the first layer of the two-layer verification is sensor-level verification, which involves continuously monitoring the sensor response during subsequent control pulse tests to confirm the stability of the calibration factor correction effect. The second layer is system-level verification, which compares the overall power generation efficiency, grid adaptability, and energy storage operation economy of the adjusted system with the preset expected targets to ensure that the improvement brought about by parameter adjustment is system-wide rather than localized. If the results of both layers of verification meet the preset standards, the calibration factor and control parameters are marked as confirmed and effective; if either layer of verification fails, the process returns to S43 to re-quantify the deviation and correct the calibration factor.
[0048] In the technical solution of the above embodiments, the actual response curve of the sensor is obtained by performing control pulse testing during low-impact periods. The expected response curve is generated and the deviation is quantified by combining the sensor batch characteristics and aging trends. Then, the calibration factor is corrected and the control parameters are adjusted. Finally, the reliability of the adjustment is ensured through sensor-level and system-level dual-layer verification. This method optimizes control parameters based on actual physical response data, achieving accurate calibration of sensor drift and adaptive optimization of the control strategy, avoiding control parameter setting deviations caused by inaccurate sensor data in traditional solutions.
[0049] In one example, during the high-temperature period of summer, the operator of a large-scale distributed photovoltaic power station found that the charging limitation of the energy storage system seemed untimely, suspecting that the module backsheet temperature sensor might be experiencing systematic underreporting. Using this solution, the system generated a series of control pulses during the low-impact period from 2 AM to 4 AM, such as slightly adjusting the energy storage charging power and collecting high-frequency temperature sensor data from the affected area. Based on the known characteristics of the sensor batch (the batch's factory-calibrated response time constant was 15 seconds) and historical aging data, the system calculated that after three years of operation, the response time constant would increase by approximately 5% annually, generating the expected response curve. Comparing the actual response curve with the expected response curve revealed a root mean square deviation. The reading was 1.8℃, confirming that the sensor did indeed have a systematic underreporting issue. After correcting the calibration factor accordingly, subsequent control pulse tests showed... The temperature dropped to 0.3℃, while the overall power generation efficiency of the system improved by about 0.3% within a week after the correction, and the response timing of the energy storage charging and discharging strategy was also more accurate.
[0050] In one possible design, Figure 4 This is a flowchart illustrating the deviation pattern recognition in step S241 according to an exemplary embodiment. (Refer to...) Figure 4 In step S20, the step of identifying operational deviations that deviate from the operational performance reference in a specific pattern includes: S241, Collect the deviation sequence between actual operating data and the operating performance reference; In this step, the system continuously calculates the actual output power of the components within each sampling period. Performance Reference Deviation between The deviation values from multiple consecutive sampling periods are organized into a deviation sequence in chronological order. .
[0051] S242, Combining environmental conditions, system operating mode, and equipment operating time, extract the evolution characteristics of the deviation sequence; In this step, the system performs correlation analysis with the deviation sequence and environmental conditions during the same period, such as irradiance, temperature, humidity, etc., system operation mode, such as grid-connected mode, islanded mode, power-limited mode, etc., as well as the cumulative operating time of the equipment, to extract the key evolution characteristics of the deviation sequence, including the mean, variance, trend direction of the deviation (e.g., whether it increases or decreases over time), the correlation coefficient with irradiance or temperature, and the daily or seasonal cycle pattern of the deviation.
[0052] S243, Update the preset deviation pattern feature library according to the evolution characteristics of the deviation sequence; In this step, the deviation mode feature library pre-stores feature templates for various known deviation modes. For example, the "component high-temperature degradation mode" is characterized by a deviation amplitude that is positively correlated with irradiance and temperature, and the trend is that the deviation increases year by year. The "inverter periodic fluctuation mode" is characterized by a deviation exhibiting periodic oscillations at a fixed frequency, with the amplitude related to changes in grid load. The system compares the currently extracted deviation evolution feature with the templates in the feature library. If it highly matches a template, the parameters of that template are updated, such as updating the average degradation rate. If the current feature does not match any templates in the library, a new candidate deviation mode is created in the feature library.
[0053] S244, Match the deviation sequence with the updated deviation pattern feature library to identify the pattern of the operational deviation, determine the type of the operational deviation, quantify the magnitude of the operational deviation, and evaluate the scope of the impact of the operational deviation.
[0054] In this step, the updated deviation pattern feature library is used to perform final pattern matching on the current deviation sequence. The matching comprehensively considers the time-domain characteristics, frequency-domain characteristics, correlation with operating conditions, and evolution trend of the deviation. Upon successful matching, the deviation type is determined, such as component degradation mode, inverter fluctuation mode, or sensor drift mode, and the deviation magnitude is quantified, such as the average power deviation percentage, the sensor temperature low alarm magnitude, and the assessment impact range, such as the number of affected components, their location, and the percentage impact on overall power generation. The assessment results are output in structured data format for use by the subsequent control parameter adjustment module.
[0055] In the technical solution of the above embodiments, a complete analysis link from raw deviation data to structured deviation diagnosis results is established through four steps: deviation sequence collection, evolution feature extraction, feature library update, and pattern matching. This enables accurate pattern recognition, type determination, amplitude quantification, and impact range assessment of operational deviations. In particular, the dynamic update mechanism of the deviation pattern feature library allows the system to adapt to the emergence of new deviation patterns without the need for manual model retraining.
[0056] In one possible design, Figure 5 This is a flowchart illustrating the deviation decomposition in step S251 according to an exemplary embodiment. (Refer to...) Figure 5 In step S20, the step of identifying operational deviations that deviate from the operational performance reference in a specific pattern includes: S251, Collect the deviation sequence between actual operating data and the operating performance reference; In this step, the deviation sequence is continuously acquired in the same way as in S241 above.
[0057] S252, Perform pattern analysis on the deviation sequence to identify periodic, trend, or instantaneous characteristics; In this step, the system performs time-frequency analysis on the deviation sequence, detecting periodic components, such as power fluctuations at specific frequencies introduced by inverter firmware upgrades, through Fast Fourier Transform; trend components, such as the continuous degradation trend of component performance, through Linear Regression or Moving Average; and transient components, such as deviation spikes caused by occasional grid disturbances, through outlier detection. These three characteristics are labeled and recorded separately.
[0058] S253, based on the physical characteristics and interaction relationships of the components inside the distributed photovoltaic power station, as well as the current environmental conditions, the deviation is decomposed into multiple independent sub-deviations; In this step, the total deviation ΔP_total of the distributed photovoltaic power station originates from multiple independent deviation sources. The system establishes a deviation decomposition model based on the physical characteristics of its internal components: in, The total power deviation (W) of the distributed photovoltaic system. The deviation component (W) caused by component attenuation. The deviation component (W) caused by inverter fluctuations. The deviation component (W) caused by sensor drift. This refers to the deviation component (W) caused by changes in environmental conditions, i.e., the deviation component caused by non-systematic changes. Each component is separated using independent component analysis or parameter estimation methods based on physical models. For example, using component characteristic parameters, where the temperature coefficient... Calculate the theoretical contribution of temperature change to power for known parameters, and isolate the portion caused by environmental factors from the total deviation.
[0059] S254, quantify the magnitude of the independent sub-bias, and assess the influence range of the independent sub-bias based on the type and magnitude of the independent sub-bias; In this step, the deviation of each separated independent sub-item is quantified. For example, The magnitude is expressed as the average annual power decay rate (% / year). The amplitude is expressed by the peak value and frequency of the fluctuation. The magnitude is expressed as an equivalent temperature offset. The quantification results are recorded as amplitude parameters for each deviation source. For each individual sub-deviation type and magnitude, its impact on different functional levels of the system is assessed. For example, the impact of component attenuation sub-deviation on power generation prediction accuracy is assessed, the impact of sensor drift sub-deviation on the accuracy of energy storage dispatch strategies is assessed, and the impact of inverter fluctuation sub-deviation on grid compliance is assessed. The assessment results include the affected system functional modules, the quantification of the impact, and recommended processing priorities.
[0060] In the technical solution of the above embodiments, the composite total deviation is decomposed into independent deviation source components through a four-step process of pattern analysis, deviation decomposition, amplitude quantification, and impact assessment, providing precise amplitude quantification and impact range assessment for each deviation source. This method allows for precise intervention of control parameters targeting specific deviation sources, rather than general adjustments to global parameters.
[0061] In one possible design, Figure 6 This is a flowchart illustrating step S261, the update of the deviation pattern library, according to an exemplary embodiment. (Refer to...) Figure 6 In step S20, the step of identifying operational deviations that deviate from the operational performance reference in a specific pattern includes: S261, Collect the deviation sequence between the actual operating data and the operating performance reference, and detect the presence of a pattern in the deviation sequence that does not match the current pattern library; In this step, the system compares the features of the current deviation sequence with all existing templates in the deviation pattern feature library one by one. When the matching degree of a certain sub-pattern in the deviation sequence is lower than the preset threshold of all existing templates, it is marked as a "non-matching pattern".
[0062] S262, extract historical deviation sequences and corresponding historical deviation features that are similar to the mismatch pattern in terms of time, environmental conditions or equipment type from the historical data of the long-term operation of distributed photovoltaic power stations; In this step, the system searches the historical database for historical deviation sequences that share similar contextual conditions with the current mismatch pattern, such as similar equipment, similar environmental conditions, and similar time periods. The deviation characteristics of these historical sequences are extracted, including the temporal evolution pattern of the deviation, its correlation with operating conditions, and its amplitude variation characteristics.
[0063] S263, compare the historical deviation features with the features of the mismatch pattern to identify candidate deviation patterns; In this step, feature comparison is used to determine whether the mismatched pattern is essentially the same type as a certain historical deviation pattern, differing only in parameters or manifestation. If a match is successful, the historical pattern is marked as a candidate deviation pattern.
[0064] S264, Add the key parameters of the candidate deviation pattern to the deviation pattern feature library; In this step, key feature parameters of candidate deviation patterns are extracted, such as the typical deviation amplitude range, correlation coefficient with the environment, and rate of temporal evolution. These are then added to the feature library as new deviation pattern templates. During addition, the creation time, source data range, and contextual conditions at the time of initial identification are recorded.
[0065] S265, The historical data is re-analyzed to verify the accuracy and coverage of the new integrated pattern.
[0066] In this step, the system uses the updated feature library to reanalyze the deviation sequences in the historical data, verifying whether the newly added pattern templates can correctly identify the corresponding deviation events that were previously missed or falsely reported in the historical data. The system calculates the recognition accuracy of the new patterns (the proportion of correctly identified events to the total number that should be identified) and the coverage (the proportion of covered deviation events to the total number of historical deviation events), ensuring that the update of the pattern library is effective and reliable.
[0067] In the technical solution of the above embodiments, the automatic expansion and verification of the deviation pattern feature library is achieved through a five-step process: mismatch pattern detection, historical data mining, candidate pattern recognition, feature library update, and historical verification. This mechanism enables the system to continuously learn and adapt to newly emerging deviation pattern types, improving the system's ability to identify unknown or novel deviations and avoiding the problem of insufficient recognition capability of traditional static pattern libraries when facing new component or new inverter behavior patterns.
[0068] In one example, a large-scale distributed photovoltaic (PV) power plant comprises approximately 3,000 PV modules and associated inverters, equipped with an energy storage system to optimize grid connection. The operator found that although the overall system appeared to be operating normally, the power generation efficiency consistently fell short of the theoretical optimal value. This solution establishes a dynamic, ideal physical behavior reference for each PV module. Taking the "PV-X100" module as an example, its maximum power output under standard test conditions provided by the manufacturer is 330W, with a power temperature coefficient of −0.4% / ℃. When the real-time monitored ambient irradiance is 800W / m²... 2 When the component backplane temperature is 45℃, the system calculates the ideal output power: If the actual output power of the component remains around 241.8W, approximately 0.45% lower than the ideal value, and this slight deviation is more pronounced during periods of high temperature and high radiation, the system identifies it as a component degradation mode caused by the new encapsulation material. Simultaneously, the system detected a systematic underreporting of approximately 2°C by some temperature sensors after long-term operation, correcting the energy storage charging limit threshold. Through adaptive adjustment of control parameters and effect verification, the overall power generation efficiency of the power station improved by approximately 0.5% within one month after the correction, and the response timing of the energy storage charging and discharging strategy became more accurate.
[0069] In another example, after a firmware upgrade, the inverter in a distributed photovoltaic power station experienced periodic small power fluctuations under specific grid load variations. Spectral analysis of the deviation sequence revealed a distinct periodic peak at 0.2 Hz, consistent with the control logic switching frequency introduced by the inverter firmware upgrade. The system decomposed this periodic fluctuation into independent components from the total deviation. After adjusting the voltage compliance threshold for the inverter's location, the threshold was reduced from ±5% to ±5.1%, preventing unnecessary protective shutdowns caused by misinterpreting internal fluctuations as grid anomalies. Subsequent analysis of operational data for the inverter's location revealed that protective shutdown events decreased by approximately 80%.
[0070] In one possible design, step S45, the step of performing two-layer verification on the adjusted control parameters, includes: S451, Record the sensor calibration factor and its changing trend over time; In this step, the system continuously records the calibration factor for each corrected sensor. and its changing trend over time Establish a time-series database of sensor calibration factors.
[0071] S452 records the long-term trend of the overall system performance; In this step, the system records key indicators reflecting overall operational performance, including the long-term trends of average daily power generation, power generation efficiency (PR) value, energy storage system charging and discharging efficiency, and grid-connected power factor.
[0072] S453, compare the changing trend of the sensor calibration factor with the long-term trend of the overall system performance, and quantify the difference between the two; In this step, the trend curve of the sensor calibration factor is aligned and compared with the trend curve of the overall system performance. Theoretically, if the correction of the sensor calibration factor accurately reflects the actual sensor drift, the system performance should stabilize or improve after the correction. If there is a significant discrepancy between the two trends, for example, the calibration factor is continuously corrected upwards but the power generation efficiency continues to decline, then the magnitude and direction of this difference are quantified.
[0073] S454, Identify potential factors that may cause the differences; In this step, potential factors causing inconsistencies in trends are systematically investigated. Potential factors may include: systematic errors in the calibration factor correction method itself; other unidentified sources of deviation besides sensor drift, such as undetected component microcracks; and environmental events, such as temporary impacts of local power grid failures or extreme weather on system performance.
[0074] S455, generate an inconsistency report.
[0075] In this step, the system generates a discrepancy report based on the above analysis results. The report includes: the correction history of the sensor calibration factor, the trend of the overall system performance, the quantitative value of the difference between the two, the identified potential factors and their probability ranking, and the suggested further investigation or corrective measures. The report is submitted to maintenance personnel or the upper-level decision-making system for reference.
[0076] In the technical solution of the above embodiments, a macro-level verification mechanism is established by comparing and quantifying the differences between the sensor-level calibration factor trend and the system-level operating performance trend in two dimensions, which ensures the long-term reliability of control parameter adjustment and the stability of system operation.
[0077] In one possible design, step S454, the step of identifying potential factors leading to the difference, includes: S4541 performs pattern analysis on differential data to identify periodic, trend, or transient patterns; In this step, the trend difference data obtained from S453 will be used as input, and a time-frequency analysis similar to that in S252 will be performed to identify periodic components in the difference, such as daily cycles and weekly cycles, trend components, such as continuously expanding differences, and instantaneous components, such as sudden deviation peaks.
[0078] S4542, based on the physical characteristics, interaction relationships and current environmental conditions of each component inside the distributed photovoltaic power station, construct the physical connection structure; In this step, a physical relationship diagram of the system's internal structure is constructed. This structure describes the causal chains between components: for example, "component degradation - decreased output power - reduced power generation," "sensor low alarm - low temperature data - energy storage charging limitation delay - battery overheating risk," and "inverter fluctuation - unstable grid-connected power - power factor deviation," etc. Each causal relationship in the chain is assigned weights and directions.
[0079] S4543, Based on the physical association structure, extract the independent influence components of different potential factors from the difference data; In this step, the physical correlation structure is used as prior knowledge to decompose the independent influence components of each potential factor from the mixed difference data. For example, the theoretical contribution of environmental factors to power difference is calculated based on the component temperature coefficient and irradiance data, and then separated from the total difference; the remaining difference is further decomposed into component attenuation contribution, sensor drift contribution, etc.
[0080] S4544, quantify the contribution of each independent influencing component to the difference; In this step, the proportion of each potential factor's independent influence in the total variance is calculated, i.e., the percentage contribution, and then sorted by contribution from highest to lowest. For example: component attenuation contributes 45%, sensor drift contributes 30%, environmental fluctuations contribute 20%, and other factors contribute 5%.
[0081] S4545, when nonlinear coupling is identified among multiple potential factors, the interaction parameters in the physical association structure are adjusted, and the type, contribution, and coupling relationship of each potential factor are output.
[0082] In this step, when the analysis reveals nonlinear coupling between two or more potential factors—for example, when both component degradation and temperature sensor drift coexist, the impact on power generation efficiency is greater than the simple superposition of their independent effects—the interaction parameters in the physical interconnection structure are adjusted to characterize this coupling effect. The final output is a complete diagnostic report containing the type, contribution, and coupling relationship of each potential factor.
[0083] In the technical solution of the above embodiments, a five-step in-depth analysis—pattern analysis, physical correlation structure construction, independent influence component stripping, contribution quantification, and coupling relationship identification—achieves a precise diagnosis of the root causes of trend differences, providing an accurate basis for targeted problem solving and system optimization.
[0084] In one possible design, step S453, the step of quantifying the difference between the changing trend of the sensor calibration factor and the long-term trend of the overall system performance, includes: S4531, Obtain environmental event records of distributed photovoltaic power stations within a specific time period, the environmental event records including local power grid failures and extreme weather events; In this step, the system retrieves environmental event records for the target time period from the event log database. Environmental events include, but are not limited to, local power grid failures, such as voltage drops, frequency shifts, and extreme weather events, such as thunderstorms, hail, and physical obstruction or damage to components caused by strong winds.
[0085] S4532, Perform a preliminary analysis on the sensor calibration factor change trend data and the long-term trend data of the overall system operation performance to identify whether there are any instantaneous, non-periodic data fluctuations that coincide with the environmental event records in time; In this step, the sensor calibration factor trend curve and the system operation performance trend curve are cross-compared with the environmental event timeline, respectively. If the trend data at a certain moment shows abnormal transient fluctuations, such as a sudden drop in power generation efficiency of 20% within an hour followed by recovery, and there happens to be an extreme weather event recorded at that moment, it is marked as a potential external event impact point.
[0086] S4533, when the transient, non-periodic data fluctuation is identified, determine whether the data fluctuation is caused by an external sudden event based on the type and occurrence time of the environmental event record; In this step, the system determines the impact mechanism of environmental events on the operation of the photovoltaic system based on the type of event. For example, a thunderstorm event corresponds to a sudden drop in irradiance due to cloud cover, rather than an internal system fault; a voltage drop event corresponds to a grid-side problem, rather than a problem with the photovoltaic system itself. Causal inference is used to determine whether data fluctuations can be attributed to external unforeseen events.
[0087] S4534, When the data fluctuation is caused by an external sudden event, the data fluctuation is isolated or corrected before differential quantification; In this step, data fluctuations confirmed to be caused by external emergencies are isolated, that is, the data for that period is excluded from the analysis or corrected. This is done by interpolating and filling in the gaps using normal data before and after the event or by replacing the data with data from adjacent normal days, to ensure that subsequent difference quantification is not affected by external emergencies.
[0088] S4535 compares the trend data of sensor calibration factor changes after isolation or correction with the long-term trend data of the overall system performance to quantify the difference between the two.
[0089] In this step, the difference indicators between the sensor calibration factor trend and the system operation performance trend are calculated using the cleaned trend data, such as the slope difference of the trend line, the offset in time series, the root mean square difference, etc., to obtain the quantitative difference that reflects the true internal state of the system after removing external event interference.
[0090] In the technical solution of the above embodiments, by identifying the interference of external environmental events on trend data and isolating or correcting them, it is ensured that the quantitative difference between the sensor calibration factor trend and the system operation performance trend reflects the real state change inside the system rather than the short-term impact of external sudden events, thereby providing an accurate data basis for subsequent root cause analysis.
[0091] In summary, the distributed photovoltaic data analysis method based on big data provided in Embodiment 1 of this invention accurately identifies and quantifies systematic deviations that are misjudged as normal by traditional data cleaning methods by establishing a dynamically changing operational performance reference. These deviations include component nonlinear degradation, inverter periodic fluctuations, and sensor systematic drift. Control parameters are adaptively adjusted based on the type, magnitude, and scope of the deviation. A closed-loop mechanism—including control pulse testing during low-impact periods, sensor response curve comparison, calibration factor correction, and dual-layer verification—ensures the reliability of control parameter adjustments. Refined diagnosis of deviations is achieved through deviation pattern analysis, decomposition, and dynamic updating of the feature library. The method in this embodiment forms a complete closed loop from data perception to deviation identification, parameter adjustment, and effect verification, overcoming the lag in control strategy response in "pseudo-normal data" scenarios encountered by traditional fixed-threshold schemes.
[0092] Example 2 Embodiment 2 of the present invention provides a distributed photovoltaic data analysis system based on big data. Figure 7 This is a block diagram illustrating a distributed photovoltaic data analysis system based on big data, according to another exemplary embodiment. (e.g.) Figure 7 As shown, the system includes: Data acquisition module 01 is used to acquire the operating data of the distributed photovoltaic system.
[0093] The deviation identification module 02 is used to establish an operational performance reference that varies with operating conditions based on the operating data and the physical characteristics of the distributed photovoltaic system, and to identify operational deviations that deviate from the operational performance reference in a specific pattern.
[0094] The parameter adjustment module 03 is used to adjust the control parameters in the distributed photovoltaic system according to the type, magnitude and range of influence of the operating deviation.
[0095] The effect verification and optimization module 04 is used to verify the effect of the adjusted control parameters and optimize the control parameters based on the results of the effect verification.
[0096] The four modules described above work together in series according to the data flow direction: the data acquisition module 01 transmits the acquired operating data to the deviation identification module 02; after the deviation identification module 02 identifies and quantifies the deviation, it transmits the deviation type, magnitude, and scope of influence information to the parameter adjustment module 03; after the parameter adjustment module 03 adjusts the control parameters, it transmits the adjusted parameters to the effect verification and optimization module 04; the effect verification and optimization module 04 verifies the adjustment effect and feeds back the results to the parameter adjustment module 03 for optimization iteration, forming a closed loop.
[0097] In summary, the distributed photovoltaic data analysis system based on big data provided in Embodiment 2 of the present invention, through the modular design and collaborative work of the data acquisition module 01, deviation identification module 02, parameter adjustment module 03, and effect verification and optimization module 04, achieves fully automated processing from the acquisition of distributed photovoltaic system operation data to the accurate identification of hidden deviations, adaptive adjustment of control parameters, and closed-loop verification and optimization of effects. This overcomes the technical problem that traditional data analysis tools are unable to process multiple parameter relationships in real time, resulting in a lag in the response of control strategies.
[0098] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0099] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. A distributed photovoltaic data analysis method based on big data, characterized in that, Includes the following steps: Obtain operational data from distributed photovoltaic systems; Based on the operational data and the physical characteristics of the distributed photovoltaic system, an operational performance reference is established that varies with operating conditions, and operational deviations that deviate from the operational performance reference in a specific pattern are identified. The operational deviations include type, magnitude, and scope of impact. Based on the type, magnitude, and scope of the operational deviation, adjust the control parameters in the distributed photovoltaic system; The effects of the adjusted control parameters are verified, and the control parameters are optimized based on the results of the verification.
2. The method according to claim 1, characterized in that, The step of establishing an operational performance reference that varies with operating conditions based on the operational data and the physical characteristics of the distributed photovoltaic system includes: Acquire environmental irradiance data and environmental temperature data; Based on the ambient temperature data, the ambient irradiance data is corrected to obtain the corrected irradiance data; Based on the corrected irradiance data and the physical characteristics of the distributed photovoltaic system, an operational performance reference is established that varies with operating conditions.
3. The method according to claim 1, characterized in that, The steps of verifying the effect of the adjusted control parameters and optimizing the control parameters based on the results of the verification include: Identify the low-impact periods of the distributed photovoltaic power station, generate and send control pulses during the low-impact periods; During the application of the control pulse, temperature sensor readings in the area affected by the control pulse are acquired at high frequency to obtain the actual response curve of the sensor. Based on the current environmental conditions, the change in charging and discharging power of the energy storage system, and the known batch characteristics and aging trend of the target sensor, the expected response curve of the sensor under ideal conditions is generated, and the actual response curve of the sensor is compared with the expected response curve to quantify the deviation between the two. The calibration factor of the sensor under the current operating conditions is corrected, the corrected sensor data is fed back to the upper control system, and the control parameters of the upper control system are adaptively adjusted according to the corrected sensor data. The adjusted control parameters are verified in two layers. The two-layer verification includes continuously monitoring the response of the sensor in subsequent control pulse detection tests to confirm the continued effectiveness of the calibration factor, and comparing the overall operating performance of the adjusted system with the expected target.
4. The method according to claim 1, characterized in that, The step of identifying operational deviations that deviate from the operational performance reference in a specific pattern includes: Collect the deviation sequence between actual operating data and the operating performance reference; The evolution characteristics of the deviation sequence are extracted by combining environmental conditions, system operation mode, and equipment operating time; Update the preset deviation pattern feature library based on the evolution characteristics of the deviation sequence; The deviation sequence is matched with the updated deviation pattern feature library to identify the pattern of the operational deviation, determine the type of the operational deviation, quantify the magnitude of the operational deviation, and assess the scope of the impact of the operational deviation.
5. The method according to claim 1, characterized in that, The step of identifying operational deviations that deviate from the operational performance reference in a specific pattern includes: Collect the deviation sequence between actual operating data and the operating performance reference; Pattern analysis is performed on the deviation sequence to identify periodic, trend, or transient characteristics; Based on the physical characteristics and interaction relationships of the components within the distributed photovoltaic system, as well as the current environmental conditions, the deviation is decomposed into multiple independent sub-deviations; Quantify the magnitude of the independent sub-biases; The influence range of the independent sub-bias is assessed based on the type and magnitude of the independent sub-bias.
6. The method according to claim 1, characterized in that, The step of identifying operational deviations that deviate from the operational performance reference in a specific pattern includes: Collect the deviation sequence between actual operating data and the operating performance reference, and detect the presence of patterns in the deviation sequence that do not match the current pattern library; From the historical data of the long-term operation of distributed photovoltaic systems, extract historical deviation sequences and corresponding historical deviation characteristics that are similar to the mismatch patterns in terms of time, environmental conditions or equipment type; The historical deviation features are compared with the features of the mismatch patterns to identify candidate deviation patterns; Add the key parameters of the candidate deviation patterns to the deviation pattern feature library; The historical data was re-analyzed to verify the accuracy and coverage of the new integrated pattern.
7. The method according to claim 3, characterized in that, The step of performing two-layer verification on the adjusted control parameters includes: Record the sensor calibration factor and its changing trend over time; Record the long-term trend of the overall system performance; By comparing the changing trend of the sensor calibration factor with the long-term trend of the overall system performance, the difference between the two is quantified. Identify the potential factors that cause the discrepancies and generate an inconsistency report.
8. The method according to claim 7, characterized in that, The step of identifying potential factors that could cause the difference includes: Perform pattern analysis on the differential data to identify periodic, trend, or transient patterns; Based on the physical characteristics, interaction relationships, and current environmental conditions of the components within the distributed photovoltaic system, a physical interconnection structure is constructed. Based on the physical correlation structure, the independent influence components of different potential factors are extracted from the differential data; Quantify the contribution of each independent influencing component to the difference; When nonlinear coupling is identified among multiple potential factors, the interaction parameters in the physical association structure are adjusted, and the type, contribution, and coupling relationship of each potential factor are output.
9. The method according to claim 7, characterized in that, The step of quantifying the difference between the trend of change of the sensor calibration factor and the long-term trend of the overall system performance includes: The system acquires environmental event records for a distributed photovoltaic system within a specific time period, including local power grid failures and extreme weather events. A preliminary analysis is conducted on the trend data of the sensor calibration factor change and the long-term trend data of the overall system performance to identify whether there are any instantaneous, non-periodic data fluctuations that coincide with the environmental event records in time. When the transient, non-periodic data fluctuations are identified, it is determined whether the data fluctuations are caused by an external sudden event, based on the type and occurrence time of the environmental event record. When the data fluctuations are caused by external unforeseen events, the data fluctuations should be isolated or corrected before differential quantification. The differences between the isolated or corrected sensor calibration factor change trend data and the long-term trend data of the overall system performance are quantified by comparing the two.
10. A distributed photovoltaic data analysis system based on big data, characterized in that, The system includes: The data acquisition module is used to acquire the operating data of the distributed photovoltaic system; The deviation identification module is used to establish an operational performance reference that varies with operating conditions based on the operational data and the physical characteristics of the distributed photovoltaic system, and to identify operational deviations that deviate from the operational performance reference in a specific pattern. The operational deviations include type, magnitude, and scope of influence. The parameter adjustment module is used to adjust the control parameters in the distributed photovoltaic system according to the type, magnitude, and scope of influence of the operating deviation. The effect verification and optimization module is used to verify the effect of the adjusted control parameters and optimize the control parameters based on the results of the effect verification.