Multi-dimensional model and data driven photovoltaic power station aggregation output attenuation evaluation method

By using multidimensional models and data-driven methods, a three-dimensional feature tensor and state-space model is constructed. Combined with a dual verification mechanism, the problems of single evaluation dimension and low reliability in photovoltaic power plant assessment are solved. This enables dynamic tracking and reliability assessment of photovoltaic power plant performance degradation, supporting accurate operation and maintenance decisions.

CN121504244APending Publication Date: 2026-02-10ECONOMIC TECH RES INST STATE GRID QIANGHAI ELECTRIC POWER +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511571162.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-30
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing technologies for evaluating the performance of photovoltaic power plants suffer from problems such as a single evaluation dimension, insufficient dynamic tracking capability, and low reliability of evaluation results. They are difficult to fully integrate multidimensional heterogeneous data, cannot reflect the nonlinear time-varying characteristics of the degradation process, and lack a verification mechanism for the reliability of model output.

Method used

Employing a multidimensional model and data-driven approach, a nonlinear dynamic model is established by constructing a three-dimensional feature tensor of time-space-attribute and combining it with state-space theory. A dual data-physical verification mechanism is introduced to identify abnormal attenuation periods, output verified attenuation state sequences and reliability identification parameters, and construct a component-level attenuation rate weighted aggregation algorithm to ensure the credibility of the evaluation results.

Benefits of technology

It achieves the systematic integration of all-time, all-space, and multi-attribute information of photovoltaic power plants, dynamically tracks the performance degradation process that cannot be directly observed, identifies and eliminates abnormal data interference, outputs a degradation state sequence with reliability indicators, ensures the credibility of the evaluation results, and supports accurate operation and maintenance decisions and risk warnings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121504244A_ABST
    Figure CN121504244A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-dimensional model and data driven photovoltaic power station aggregation output attenuation evaluation method. Comprising the following steps: constructing a multi-dimensional spatial-temporal feature matrix; establishing an output attenuation nonlinear dynamic model; outputting a reliability state sequence through a data-physics dual check mechanism; and carrying out attenuation rate aggregation and trend prediction based on a verification result. According to the method, the problems of incomplete evaluation, poor dynamic nature and low reliability in the prior art are solved, and accurate, dynamic and reliable evaluation of the output attenuation characteristics of the photovoltaic power station is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of photovoltaic power plant performance evaluation technology, and more specifically, to a multi-dimensional model and data-driven method for evaluating the aggregate output degradation of photovoltaic power plants. Background Technology

[0002] The power output performance of photovoltaic (PV) power plants inevitably declines over time due to factors such as material aging, dust and snow accumulation, hot spots, and potential-induced degradation. Accurately assessing the overall degradation status and trends of a power plant is crucial for asset valuation, operation and maintenance decisions, and power generation forecasting. Existing technologies typically employ periodic offline monitoring, statistical analysis based on a single data source, or static physical models for performance evaluation.

[0003] However, existing technologies have significant drawbacks: First, assessment methods often rely on a single data source or static model, making it difficult to fully integrate multi-dimensional heterogeneous data such as historical power output, high-precision meteorological data, component operating status, and geospatial information, resulting in incomplete descriptions of attenuation characteristics. Second, they lack the ability to dynamically track and explicitly quantify the internal health status of power plants (such as permanent performance loss, temporary pollution, and potential failure risks), failing to reflect the nonlinear time-varying characteristics of the attenuation process. Finally, existing methods generally lack a mechanism to verify the reliability of model outputs, making them susceptible to interference from abnormal data, resulting in insufficient credibility of assessment results and difficulty in supporting accurate operation and maintenance decisions and risk warnings.

[0004] To address the aforementioned shortcomings, this invention aims to solve the technical problems of single evaluation dimensions, insufficient dynamic tracking capabilities, and low reliability of evaluation results in existing technologies, and provides a comprehensive evaluation scheme for attenuation characteristics that can achieve multi-source data fusion, dynamic evolution modeling, and dual verification. Summary of the Invention

[0005] In view of the shortcomings of the prior art, the purpose of this invention is to provide a multi-dimensional model and data-driven method for evaluating the aggregate output degradation of photovoltaic power plants, so as to solve the problems of incomplete evaluation, poor dynamics and low reliability in the prior art.

[0006] To achieve the above objectives, the present invention adopts the following technical solution;

[0007] A multidimensional model and data-driven method for evaluating the aggregated output degradation of photovoltaic power plants includes the following steps:

[0008] S100. Construction of Multidimensional Spatiotemporal Feature Matrix: By integrating historical output data of photovoltaic power plants, high-precision meteorological observation data, component operating status parameters and geospatial information, a feature matrix with time-space-attribute dimensions is constructed.

[0009] S200. Dynamic evolution modeling of power output decay: Based on the multidimensional feature matrix constructed by S100, a nonlinear dynamic model of photovoltaic power output decay is established using state space theory.

[0010] The S300 data-physical dual verification mechanism, based on the output of the dynamic evolution model of S200, establishes a dual prediction architecture of data-driven and physical models. It identifies abnormal decay periods through dynamic residual analysis and outputs the verified decay state sequence and reliability identification parameters.

[0011] S400. Attenuation Quantification and Trend Assessment: Based on the verified attenuation state sequence and reliability identification parameters output by S300, a component-level attenuation rate weighted aggregation algorithm is constructed to establish an attenuation acceleration analysis model. In conjunction with the reliability score, trend prediction is performed to ensure the credibility of the assessment results.

[0012] As a further aspect of the present invention, S100 includes:

[0013] S101. Multi-source data acquisition and preprocessing: Acquire multi-source heterogeneous operation data throughout the entire life cycle of the photovoltaic power station. The multi-source heterogeneous operation data includes historical power output sequences, high-precision meteorological observation data, component operation status parameters, and geospatial information.

[0014] Subsequently, the raw data undergoes preprocessing, including spatiotemporal alignment, missing value imputation, and outlier removal, to form a clean data pool with uniform spatiotemporal granularity.

[0015] S102 Feature Engineering and Standardization: Based on the clean data pool, derived features with stronger physical meaning and discriminative ability are constructed. Since the variables in this feature set have huge differences in scale and numerical range, directly inputting them into the model will cause the features with large numerical ranges to dominate the training process. In order to eliminate the influence of scale and ensure that the model can learn the information of all features equally, the Z-Score standardization method is then used to normalize the original numerical features and derived features.

[0016] Among them, derived features include performance ratio and temperature efficiency coefficient, etc.

[0017] The original features are numerical variables obtained directly from the clean data pool of S101, such as irradiance, ambient temperature, wind speed, back panel temperature, and open circuit voltage.

[0018] In the Z-Score standardization method, for the 1st Features Its standardized value The calculation formula is:

[0019] ;

[0020] in, For the first One eigenvalue; For the first The mean of the eigenvalues; For the first The standard deviation of each eigenvalue;

[0021] S103. Construct a three-dimensional feature tensor. The data processed in S102 is used to construct a three-dimensional feature tensor based on three dimensions: timestamp, spatial location, and feature variables. This tensor fully encapsulates the spatiotemporal and multi-attribute information of the power plant's operation; the formula is expressed as:

[0022] ;

[0023] For time step, The number of spatial points. This represents the total number of feature variables.

[0024] As a further aspect of the present invention, the step of establishing a nonlinear dynamic model of photovoltaic power output decay using state-space theory based on the multidimensional feature matrix constructed based on S100 includes:

[0025] S210. Define state variables and construct a state vector to describe the internal health state of the power plant. The state vector contains three core implicit state components that cannot be directly measured but determine the power output performance. The implicit state components include:

[0026] Health : Characterizes the degree of permanent performance loss caused by material aging, etc., with a value range of . 1 represents an ideal state of health;

[0027] filth : Characterizes the degree of temporary performance loss caused by dust, snow, etc., with a value range of . 1 represents complete cleanliness;

[0028] Potential failure risks : Characterizes the probability of component failures such as hot spots and PID; it is a non-negative scalar, and the larger the value, the higher the risk.

[0029] S220. Establish the state-space equations and, using state-space theory, construct a nonlinear dynamic system model driven externally by the three-dimensional feature tensor output by S100. The nonlinear dynamic system model includes:

[0030] State equations are used to describe the internal states. How it evolves under the influence of the previous state and the current external input can be expressed by the following formula:

[0031] ;

[0032] Observation equations describe how the system's internal state and external inputs together manifest as observable values ​​that can be actually measured. The formula is expressed as:

[0033] ;

[0034] The state transition function defines the system state transition from... Time's up The dynamic evolution pattern at any given moment; The observation function defines the system's state vector. and input vector How can we collectively generate observable measurements?

[0035] For discrete-time indexing; It is a state vector; The input vector; Indicates in A set of variables that constantly describe the internal state of the system; For observation vectors; This is process noise; The observation noise is used to characterize model uncertainty and measurement error, respectively, and follows a Gaussian distribution with a mean of zero.

[0036] Mapping relationship between model variables and S100 feature tensors:

[0037] External input : for the feature tensor constructed from S100 In the middle, at the time step The set of feature vectors of all spatial points obtained by slicing; it serves as the external known condition driving the state evolution, specifically including standardized irradiance, temperature, derived features, etc.

[0038] System observation : Corresponding to the actual measured values ​​in the S100 data pool, mainly the output sequence and key electrical parameters, used to compare with the model prediction values ​​to update the state estimate;

[0039] S230. State Estimation and Prediction: A nonlinear filtering algorithm is used to solve the state-space model, with the state transition function as the process. and observation function As the kernel, with system observation and external input Given the condition, the prediction-update step is recursively executed to achieve the Bayesian optimal estimate of the most likely state vector at each time step; finally, a smooth, denoised healthy state sequence is output, thereby making the performance degradation process, which is not directly observable, explicit and quantifiable.

[0040] As a further aspect of this invention, based on the output of the dynamic evolution model of S2, a dual prediction architecture of data-driven and physical models is established. Abnormal decay periods are identified through dynamic residual analysis, and a verified decay state sequence and reliability identification parameters are output, specifically including:

[0041] S310. Establish a dual prediction architecture and construct two parallel prediction models to provide a benchmark for subsequent residual analysis. The parallel prediction models include a physical mechanism prediction model and a data-driven prediction model.

[0042] Among them, the physical mechanism prediction model is a photovoltaic module performance model based on physical mechanisms, which is based on semiconductor physics and circuit theory.

[0043] Data-driven prediction models are based on feature tensors constructed using S100. Train a time-series deep learning model as a prediction model. The data-driven prediction model learns the complex nonlinear mapping relationship between input features and output power from massive historical data.

[0044] The core of the physical mechanism prediction model can be expressed as:

[0045] ;

[0046] Photocurrent; This is the diode saturation current; For electron charge; This refers to the output voltage. Output current; It is a series resistor; This is the diode ideality factor; Boltzmann's constant; Absolute temperature; These are parallel resistors;

[0047] S320. Dynamic residual analysis and anomaly identification: By comparing the prediction results of physical mechanism prediction models and data-driven prediction models with actual measured values, high-confidence anomalous time periods are identified; specific steps include:

[0048] Calculate the dual residuals: separately calculate the residuals between the predicted output and the actual output of the physical mechanism prediction model and the data-driven prediction model.

[0049] ;

[0050] ;

[0051] Indicates at time The actual measured output value; Physical mechanism model at time The predicted output value; Data-driven model at time The predicted output value; Predict residuals for physical models; Data-driven model prediction residuals;

[0052] To avoid the inadequacy of fixed thresholds under different weather conditions, dynamic thresholds are set. The thresholds can be calculated based on recent historical residuals.

[0053] S330. Output the verified status and reliability parameters. Based on the anomaly identification results, correct and label the original health status sequence output by S200.

[0054] For state sequence verification, the health estimate based on S200 is directly adopted for non-abnormal periods. For abnormal time periods, the data for that specific moment... Labeling or interpolating using health values ​​from nearby normal periods can prevent anomalous data from contaminating the final decay trend analysis.

[0055] Calculate reliability parameters, including a residual confidence level for each time step, as a quantitative indicator of the reliability of the state estimate at that point. The calculation formula is:

[0056] ;

[0057] For a moment The reliability of the state estimation result is a quantitative indicator with a value range of (0,1]. The larger the value, the higher the reliability of the state estimation result at that moment.

[0058] The final output is a verified state sequence with reliability indicators. The state sequence not only includes the health status of the power plant, but also the reliability of each state point, providing a crucial weighting basis for trend prediction in S400.

[0059] As a further aspect of the present invention, based on the verified attenuation state sequence and reliability identification parameters output by S3, a component-level attenuation rate weighted aggregation algorithm is constructed, an attenuation acceleration analysis model is established, and trend prediction is performed in conjunction with the reliability score to ensure the credibility of the evaluation results; specifically, it includes: S410. Component-level attenuation rate weighted aggregation, performing linear fitting on the verified health sequence of each component to obtain its individual annual attenuation rate; subsequently, weights are assigned according to the importance of the components to calculate the power plant-level aggregated attenuation rate. :

[0060] ;

[0061] This represents the total number of photovoltaic modules participating in the aggregation calculation; For component indexing; No. The annual degradation rate of an individual photovoltaic module;

[0062] S420. Establish a decay acceleration analysis model, smooth the average health sequence at the power plant level, and calculate its first-order and second-order differences:

[0063] ;

[0064] ;

[0065] For a moment The rate of decay; The time interval between adjacent time steps; For a moment The average health of power plants; For a moment The average health of power plants; A persistently negative value is an early warning sign of accelerated decay. A persistently negative value is an early warning sign of accelerated decay;

[0066] S430. Combining reliability trend prediction, based on the verified state sequence with reliability identifier output by S300, a time series prediction model integrating reliability parameters is constructed to probabilistically predict the future health status of the power plant.

[0067] As a further aspect of the present invention, the method for constructing a time series prediction model that integrates reliability parameters based on the verified state sequence with reliability identifiers output by S300, and for probabilistically predicting the future health status of the power plant, includes:

[0068] Construct a reliability-weighted training dataset by extracting historical health data and its corresponding reliability parameters from the verified state sequence output by S300; when constructing model training samples, assign a sample weight coefficient to each sample, which is determined by the average value of the reliability parameters in its corresponding time period;

[0069] A prediction model integrating reliability information is established, and a Long Short-Term Memory (LSTM) network with an attention mechanism is used as the core architecture of the prediction model. The input features of the model include both the historical health sequence and the corresponding reliability parameter sequence.

[0070] A weighted training strategy is implemented. During the model training process, a reliability-based weighted loss function is adopted. This loss function is based on the reliability quantification index and sample weights of the S300 output, and uses the weighted mean square error as the loss function.

[0071] The formula for calculating the weighted loss function is:

[0072] ;

[0073] This is the total number of training samples; and They are the first The actual health status of each sample and the model prediction; It is the first The weights of each sample;

[0074] The model generates probabilistic prediction results, and its output uses a quantile regression architecture, simultaneously outputting multiple quantile values ​​for future health. Specifically, three nodes are set up in parallel at the model's output layer, each corresponding to one of the three specific quantiles of health. Their relationship is expressed as follows:

[0075] ;

[0076] in, This represents the baseline prediction (median) of future health. and These three output values ​​collectively define an 80% confidence interval; using these three output values, a predictive interval for health status can be intuitively constructed; the width of this interval... This naturally reflects the degree of uncertainty in the prediction results, providing complete reference information for power plant operation and maintenance decisions, from the most likely scenario to the risk boundary.

[0077] Compared with existing technologies, the technical solution provided by this invention has produced significant beneficial effects by systematically integrating multidimensional modeling, dynamic estimation and dual verification.

[0078] First, this invention systematically integrates multi-source heterogeneous data from photovoltaic power plants by constructing a three-dimensional feature tensor of time, space, and attributes, thus solving the problems of single data dimension and insufficient information utilization in existing technologies. This method not only achieves structured encapsulation of full-time, multi-attribute information about power plant operation, but also eliminates the influence of dimensions through feature engineering and standardization, providing a high-quality and standardized data foundation for subsequent accurate modeling.

[0079] Secondly, by introducing state-space theory to establish a nonlinear dynamic model and defining core state variables such as health, pollution, and potential failure risk, this invention achieves dynamic tracking and explicit quantification of the performance degradation process within a power plant, which is not directly observable. This overcomes the shortcomings of existing static models that cannot reflect the time-varying characteristics of degradation, and can more accurately describe the dynamic evolution of degradation.

[0080] Finally, by establishing a dual prediction architecture of data-driven and physical models, and combining dynamic residual analysis and reliability parameter calculation, this invention constructs a complete verification and credibility assessment mechanism. This mechanism effectively identifies and eliminates interference from abnormal data, outputs a decay state sequence with reliability indicators, and ensures the credibility of the final decay rate aggregation and trend prediction results, fundamentally solving the pain point of insufficient reliability of existing assessment methods. Attached Figure Description

[0081] Figure 1 A flowchart for a multidimensional model and data-driven method for evaluating the aggregated output degradation of photovoltaic power plants;

[0082] Figure 2 This is a flowchart of the method for implementing step S100;

[0083] Figure 3 The flowchart shows the steps for establishing a nonlinear dynamic model of photovoltaic power output decay using state-space theory based on the multidimensional feature matrix constructed based on S1.

[0084] Figure 4 The flowchart for the implementation method of step S300;

[0085] Figure 5 The flowchart illustrates the method for constructing a component-level attenuation rate weighted aggregation algorithm and establishing an attenuation acceleration analysis model. Detailed Implementation

[0086] To make the objectives, technical solutions, and advantages of the invention clearer, the invention will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.

[0087] Please see Figure 1This illustrates a multi-dimensional model and data-driven method for evaluating the aggregate output degradation of photovoltaic power plants, provided by an embodiment of the present invention, including the following steps:

[0088] S100. Construction of Multidimensional Spatiotemporal Feature Matrix: By integrating historical output data of photovoltaic power plants, high-precision meteorological observation data, component operating status parameters and geospatial information, a feature matrix with time-space-attribute dimensions is constructed.

[0089] S200. Dynamic evolution modeling of power output decay: Based on the multidimensional feature matrix constructed by S1, a nonlinear dynamic model of photovoltaic power output decay is established using state space theory.

[0090] The S300 data-physical dual verification mechanism, based on the output of the dynamic evolution model of S2, establishes a dual prediction architecture of data-driven and physical models. It identifies abnormal decay periods through dynamic residual analysis and outputs the verified decay state sequence and reliability identification parameters.

[0091] S400. Attenuation Quantification and Trend Assessment: Based on the verified attenuation state sequence and reliability identification parameters output by S3, a component-level attenuation rate weighted aggregation algorithm is constructed to establish an attenuation acceleration analysis model. In conjunction with the reliability score, trend prediction is performed to ensure the credibility of the assessment results.

[0092] Please refer to Figure 2 It illustrates an exemplary multidimensional model and data-driven photovoltaic power plant aggregate output degradation evaluation method of this application, a flowchart of S100, the contents of which include:

[0093] S101. Multi-source data acquisition and preprocessing: Acquire multi-source heterogeneous operation data throughout the entire life cycle of the photovoltaic power station. The multi-source heterogeneous operation data includes historical power output sequences, high-precision meteorological observation data (including irradiance, ambient temperature, and wind speed), module operation status parameters (module backsheet temperature, open-circuit voltage, short-circuit current, maximum power point current and voltage), and geospatial information (altitude, slope, and aspect).

[0094] Subsequently, the raw data undergoes preprocessing, including spatiotemporal alignment, missing value imputation, and outlier removal, to form a clean data pool with uniform spatiotemporal granularity.

[0095] S102 Feature Engineering and Standardization: Based on the clean data pool, derived features with stronger physical meaning and discriminative ability are constructed. Due to the huge differences in the dimensions and numerical ranges of the variables in this feature set (e.g., voltage is tens of volts, irradiance is hundreds of watts / square meter), directly inputting them into the model would cause features with large numerical ranges to dominate the training process. To eliminate the influence of dimensions and ensure that the model can learn the information of all features equally, the Z-Score standardization method is then used to normalize the original numerical features and derived features.

[0096] Among them, derived features include performance ratio and temperature efficiency coefficient, etc.

[0097] The original features are numerical variables obtained directly from the clean data pool of S101, such as irradiance, ambient temperature, wind speed, back panel temperature, and open circuit voltage.

[0098] In the AA-Score standardization method, for the _____ Features Its standardized value The calculation formula is:

[0099] ;

[0100] in, For the first One eigenvalue; For the first The mean of the eigenvalues; For the first The standard deviation of each eigenvalue;

[0101] S103. Construct a three-dimensional feature tensor. The data processed in S102 is used to construct a three-dimensional feature tensor based on three dimensions: timestamp, spatial location (component or string), and feature variables. This tensor fully encapsulates the full-time, multi-attribute information of the power plant's operation; the formula is expressed as:

[0102] ;

[0103] For time step, The number of spatial points. This represents the total number of feature variables.

[0104] Step S100 is the data foundation of the entire method. Its core logic lies in solving the problem of information silos and building a unified analysis framework. The performance degradation of photovoltaic power plants is a complex process affected by multiple factors such as time, space, and equipment attributes. Through multi-source data fusion and preprocessing in S101, the integrity and consistency of the data are ensured. Feature engineering and standardization in S102 aim to eliminate the influence of dimensions and construct derived indicators that can more directly reflect performance changes, providing high-quality input for subsequent models. Finally, the three-dimensional feature tensor constructed in S103 systematically organizes scattered data points into a structured data cube, which fully encapsulates the spatiotemporal and multi-attribute information of power plant operation, providing a unique and standardized data source for subsequent dynamic modeling and quantitative analysis.

[0105] Please refer to Figure 3 This application illustrates an exemplary multidimensional model and data-driven method for evaluating the aggregated output degradation of photovoltaic power plants. The steps of establishing a nonlinear dynamic model of photovoltaic output degradation based on the multidimensional feature matrix constructed using S1 and state-space theory include:

[0106] S210. Define state variables and construct a state vector to describe the internal health state of the power plant. The state vector contains three core implicit state components that cannot be directly measured but determine the power output performance. The implicit state components include:

[0107] Health : Characterizes the degree of permanent performance loss caused by material aging, etc., with a value range of . 1 represents an ideal state of health;

[0108] filth : Characterizes the degree of temporary performance loss caused by dust, snow, etc., with a value range of . 1 represents complete cleanliness;

[0109] Potential failure risks : Characterizes the probability of component failures such as hot spots and PID; it is a non-negative scalar, and the larger the value, the higher the risk.

[0110] S220. Establish the state-space equations and, using state-space theory, construct a nonlinear dynamic system model driven externally by the three-dimensional feature tensor output by S100. The nonlinear dynamic system model includes:

[0111] State equations (which describe the intrinsic evolution of states over time) are used to describe the internal states. How it evolves under the influence of the previous state and the current external input can be expressed by the following formula:

[0112] ;

[0113] The observation equation (which describes how the state is represented by external observations) describes how the system's internal state and external inputs together are represented by observable values ​​that can be actually measured. The formula is expressed as:

[0114] ;

[0115] The state transition function is a (usually nonlinear) function that defines the transition from the system state to the state transition. Time's up The dynamic evolution pattern at any given moment; The observation function is a (usually nonlinear) function that defines the system's state vector. and input vector How can we collectively generate observable measurements?

[0116] For discrete-time indexing, representing the first... Each observation time; Let be the state vector, representing the state vector at which the state vector is in the state vector. A set of variables that constantly describe the internal (non-observable) state of the system; The input vector represents External inputs that are known at any given time and affect the state of the system; Indicates in A set of variables that constantly describe the internal (non-observable) state of the system; Let be the observation vector, representing the observation vector at . The system output variables that can actually be measured at any given moment; This is process noise; The observation noise is used to characterize model uncertainty and measurement error, respectively, and follows a Gaussian distribution with a mean of zero.

[0117] Mapping relationship between model variables and S100 feature tensors:

[0118] External input : for the feature tensor constructed from S100 In the middle, at the time step The set of feature vectors of all spatial points obtained by slicing; it serves as the external known condition driving the state evolution, specifically including standardized irradiance, temperature, derived features, etc.

[0119] System observation The actual measured values ​​in the S100 data pool mainly consist of the output sequence and key electrical parameters (measured short-circuit current, open-circuit voltage, etc.), which are used to compare with the model predictions to update the state estimate.

[0120] S230. State estimation and prediction: A nonlinear filtering algorithm (Particle Filtering, PF) is used to solve the state-space model, with the state transition function as the process. and observation function As the kernel, with system observation and external input Given the condition, the prediction-update step is recursively executed to achieve the Bayesian optimal estimate of the most likely state vector at each time step; finally, a smooth, denoised healthy state sequence is output, thereby making the performance degradation process, which is not directly observable, explicit and quantifiable.

[0121] A brief description of the model's working logic: The core idea of ​​this state-space model is to connect the internally invisible dynamic process of photovoltaic power plant performance degradation through a noisy observation system (i.e., actual measured power and electrical parameters). Using a nonlinear filtering algorithm, and leveraging a series of noisy observation data and known external inputs, the most probable internal state sequence is estimated in reverse, thereby achieving dynamic and quantitative tracking of performance degradation characteristics.

[0122] Please refer to Figure 4 This application illustrates an exemplary multidimensional model and data-driven method for evaluating the aggregated output degradation of photovoltaic power plants. Based on the output of the S2 dynamic evolution model, a dual prediction architecture combining data-driven and physical models is established. Abnormal degradation periods are identified through dynamic residual analysis, and a validated degradation state sequence and reliability identification parameters are output. Specifically, this includes:

[0123] S310. Establish a dual prediction architecture and construct two parallel prediction models to provide a benchmark for subsequent residual analysis. The parallel prediction models include a physical mechanism prediction model and a data-driven prediction model.

[0124] Among them, the physical mechanism prediction model is a photovoltaic module performance model based on physical mechanisms, which is based on semiconductor physics and circuit theory.

[0125] Data-driven prediction models are based on feature tensors constructed using S100. Train a time-series deep learning model (such as LSTM or Transformer) as a prediction model. The data-driven prediction model learns the complex nonlinear mapping relationship between input features and output power from massive historical data.

[0126] The core of the physical mechanism prediction model can be expressed as:

[0127] ;

[0128] Photocurrent; This is the diode saturation current; For electron charge; This refers to the output voltage. Output current; It is a series resistor; This is the diode ideality factor; Boltzmann's constant; Absolute temperature; These are parallel resistors;

[0129] S320. Dynamic residual analysis and anomaly identification: By comparing the prediction results of physical mechanism prediction models and data-driven prediction models with actual measured values, high-confidence anomalous time periods are identified; specific steps include:

[0130] Calculate the dual residuals: separately calculate the residuals between the predicted output and the actual output of the physical mechanism prediction model and the data-driven prediction model.

[0131] ;

[0132] ;

[0133] Indicates at time The actual measured output value; Physical mechanism model at time The predicted output value; Data-driven model at time The predicted output value; Predict residuals for physical models; Data-driven model prediction residuals;

[0134] To avoid the inadequacy of fixed thresholds under different weather conditions, a dynamic threshold is set. This threshold can be calculated based on recent historical residuals. For example, the mean of the residual sequence within the window plus twice the standard deviation is expressed as:

[0135] ;

[0136] time The dynamic threshold; From time arrive The mean of the residual sequence; From time arrive The standard deviation of the residual sequence;

[0137] S330. Output the verified status and reliability parameters. Based on the anomaly identification results, correct and label the original health status sequence output by S200.

[0138] For state sequence verification, the health estimate based on S200 is directly adopted for non-abnormal periods. For abnormal time periods, the data for that specific moment... Labeling or interpolating using health values ​​from nearby normal periods can prevent anomalous data from contaminating the final decay trend analysis.

[0139] Calculate reliability parameters, including a residual confidence level for each time step, as a quantitative indicator of the reliability of the state estimate at that point. The calculation formula is:

[0140] ;

[0141] For a moment The reliability of the state estimation result is a quantitative indicator with a value range of (0,1]. The larger the value, the higher the reliability of the state estimation result at that moment.

[0142] The final output is a verified state sequence with reliability indicators. The state sequence not only includes the health status of the power plant, but also the reliability of each state point, providing a crucial weighting basis for trend prediction in S400.

[0143] Please refer to Figure 5 This paper illustrates an exemplary multidimensional model and data-driven method for evaluating the aggregated output degradation of photovoltaic power plants. Based on the verified degradation state sequence and reliability identifier parameters output in S3, a component-level degradation rate weighted aggregation algorithm is constructed, a degradation acceleration analysis model is established, and trend prediction is performed in conjunction with reliability scores to ensure the credibility of the evaluation results. Specifically, it includes: S410. Component-level degradation rate weighted aggregation: linear fitting is performed on the verified health sequence of each component to obtain its individual annual degradation rate; subsequently, weights are assigned according to the importance of the component (such as installation years, number of historical failures, and type) to calculate the power plant-level aggregated degradation rate. :

[0144] ;

[0145] This represents the total number of photovoltaic modules participating in the aggregation calculation; For component indexing; No. The annual degradation rate of an individual photovoltaic module;

[0146] The S420 establishes a decay acceleration analysis model. After smoothing the average health sequence at the power plant level, it calculates the first-order difference (decay rate) and the second-order difference (decay acceleration).

[0147] ;

[0148] ;

[0149] For a moment The rate of decay; The time interval between adjacent time steps; For a moment The average health of power plants; For a moment The average health of power plants; A persistently negative value is an early warning sign of accelerated decay. A persistently negative value is an early warning sign of accelerated decay;

[0150] S430. Combining reliability trend prediction, based on the verified state sequence with reliability identifier output by S300, a time series prediction model integrating reliability parameters is constructed to probabilistically predict the future health status of the power plant.

[0151] A reliability-weighted training dataset is constructed by extracting historical health data and corresponding reliability parameters from the verified state sequence output by S300. When constructing model training samples, each sample is assigned a sample weight coefficient, which is determined by the average value of the reliability parameters in its corresponding time period. This weighting mechanism ensures that data from high-reliability periods have a greater influence during model training, enabling the prediction model to prioritize learning reliable historical patterns.

[0152] A prediction model integrating reliability information is established, and a Long Short-Term Memory (LSTM) network with an attention mechanism is used as the core architecture of the prediction model. The input features of the model include both historical health status sequences and corresponding reliability parameter sequences. Through the attention mechanism inside the model, the model automatically identifies and focuses on historical data points with high reliability, dynamically adjusts the contribution of historical data at different times to the prediction results, and achieves effective integration of reliability parameters and health status data.

[0153] A weighted training strategy is implemented. During the model training process, a reliability-based weighted loss function is adopted. This loss function is based on the reliability quantification index and sample weights of the S300 output, and uses the weighted mean square error as the loss function.

[0154] The formula for calculating the weighted loss function is:

[0155] ;

[0156] This is the total number of training samples; and They are the first The actual health status of each sample and the model prediction; It is the first The weight of each sample, which was calculated in the first step of S430 based on the reliability parameter sequence output by S300;

[0157] The model generates probabilistic prediction results, and its output uses a quantile regression architecture, simultaneously outputting multiple quantile values ​​for future health. Specifically, three nodes are set up in parallel at the model's output layer, each corresponding to one of the three specific quantiles of health (e.g., the 10th, 50th, and 90th quantiles). Their relationship is expressed as follows:

[0158] ;

[0159] in, This represents the baseline prediction (median) of future health. and These three output values ​​collectively define an 80% confidence interval; using these three output values, a predictive interval for health status can be intuitively constructed; the width of this interval... This naturally reflects the degree of uncertainty in the prediction results, providing complete reference information for power plant operation and maintenance decisions, from the most likely scenario (baseline prediction) to the risk boundary (confidence upper and lower limits).

[0160] These quantiles can be used to construct a prediction range for the health status, visually demonstrating the possible fluctuation range of the future health status; the width of the prediction range reflects the degree of uncertainty of the prediction results, providing complete reference information from baseline prediction to risk boundary for power plant operation and maintenance decisions.

[0161] The above description is a preferred embodiment of the invention and is not intended to limit the scope of the invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the invention should be included within the scope of protection of the invention.

Claims

1. A multidimensional model and data-driven method for evaluating the aggregated output degradation of photovoltaic power plants, characterized in that, Includes the following steps: S100. Construction of Multidimensional Spatiotemporal Feature Matrix: By integrating historical output data of photovoltaic power plants, high-precision meteorological observation data, component operating status parameters and geospatial information, a feature matrix with time-space-attribute dimensions is constructed. S200. Dynamic evolution modeling of power output decay: Based on the multidimensional feature matrix constructed by S100, a nonlinear dynamic model of photovoltaic power output decay is established using state space theory. The S300 data-physical dual verification mechanism, based on the output of the dynamic evolution model of S200, establishes a dual prediction architecture of data-driven and physical models. It identifies abnormal decay periods through dynamic residual analysis and outputs the verified decay state sequence and reliability identification parameters. S400. Attenuation Quantification and Trend Assessment: Based on the verified attenuation state sequence and reliability identification parameters output by S300, a component-level attenuation rate weighted aggregation algorithm is constructed to establish an attenuation acceleration analysis model. In conjunction with the reliability score, trend prediction is performed to ensure the credibility of the assessment results.

2. The multidimensional model and data-driven method for evaluating the aggregated output degradation of photovoltaic power plants according to claim 1, characterized in that, S100 includes: S110. Multi-source data acquisition and preprocessing: Acquire multi-source heterogeneous operation data throughout the entire life cycle of the photovoltaic power station. The multi-source heterogeneous operation data includes historical power output sequences, high-precision meteorological observation data, component operation status parameters, and geospatial information. Subsequently, the raw data undergoes preprocessing, including spatiotemporal alignment, missing value imputation, and outlier removal, to form a clean data pool with uniform spatiotemporal granularity. S120. Feature Engineering and Standardization: Based on the clean data pool, derived features with stronger physical meaning and discriminative ability are constructed. The variables in the derived feature set have huge differences in scale and numerical range. Directly inputting them into the model will cause the features with large numerical range to dominate the training process. In order to eliminate the influence of scale and ensure that the model can learn the information of all features equally, the Z-Score standardization method is then used to normalize the original numerical features and derived features. Among these, derived features include performance ratio and temperature efficiency coefficient; The original features are numerical variables obtained directly from the clean data pool of S101, including irradiance, ambient temperature, wind speed, backplane temperature, and open-circuit voltage. In the Z-Score standardization method, for the 1st Features Its standardized value The calculation formula is: ; in, For the first One eigenvalue; For the first The mean of the eigenvalues; For the first The standard deviation of each eigenvalue; S130. Construct a three-dimensional feature tensor. The data processed in S102 is used to construct a three-dimensional feature tensor based on three dimensions: timestamp, spatial location, and feature variables. This tensor fully encapsulates the full-time, multi-attribute information of the power plant's operation; the formula is expressed as: ; For time step, The number of spatial points. This represents the total number of feature variables.

3. The multidimensional model and data-driven method for evaluating the aggregated output degradation of photovoltaic power plants according to claim 1, characterized in that, The steps for establishing a nonlinear dynamic model of photovoltaic power output decay using the state-space theory based on the multidimensional feature matrix constructed based on S100 include: S210 defines state variables and constructs a state vector to describe the internal health state of the power plant. This state vector contains three core implicit state components that cannot be directly measured but determine power output performance. These implicit state components include: health level. filth level and potential failure risks : S220 establishes the state-space equations and, using state-space theory, constructs a nonlinear dynamic system model driven externally by the three-dimensional feature tensor output by S100. The nonlinear dynamic system model includes: State equations are used to describe the internal states. How it evolves under the influence of the previous state and the current external input can be expressed by the following formula: ; Observation equations describe how the system's internal state and external inputs together manifest as observable values ​​that can be actually measured. The formula is expressed as: ; The state transition function defines the system state transition from... Time's up The dynamic evolution pattern at any given moment; The observation function defines the system's state vector. and input vector How can we collectively generate observable measurements? For discrete-time indexing; It is a state vector; The input vector; Indicates in A set of variables that constantly describe the internal state of the system; For observation vectors; This is process noise; The observation noise is used to characterize model uncertainty and measurement error, respectively, and follows a Gaussian distribution with a mean of zero. S230, State Estimation and Prediction, employs a nonlinear filtering algorithm to solve the state-space model, using the state transition function as a reference. and observation function As the kernel, with system observation and external input Given the condition, the prediction-update step is recursively executed to achieve the Bayesian optimal estimate of the most likely state vector at each time step; finally, a smooth, denoised healthy state sequence is output, thereby making the performance degradation process, which is not directly observable, explicit and quantifiable.

4. The multidimensional model and data-driven method for evaluating the aggregated output degradation of photovoltaic power plants according to claim 1, characterized in that, Based on the output of the S2 dynamic evolution model, a dual prediction architecture of data-driven and physical models is established. Abnormal decay periods are identified through dynamic residual analysis, and validated decay state sequences and reliability indicator parameters are output, specifically including: S310 establishes a dual prediction architecture, constructing two parallel prediction models to provide a benchmark for subsequent residual analysis. The parallel prediction models include a physical mechanism prediction model and a data-driven prediction model. S320 dynamic residual analysis and anomaly identification identifies anomalous periods with high confidence by comparing the prediction results of physical mechanism prediction models and data-driven prediction models with actual measured values. The S330 outputs verified status and reliability parameters, and based on the anomaly identification results, corrects and labels the original health status sequence output by the S200.

5. The multidimensional model and data-driven method for evaluating the aggregated output degradation of photovoltaic power plants according to claim 4, characterized in that, The physical mechanism prediction model is a photovoltaic module performance model based on physical mechanisms, which is based on semiconductor physics and circuit theory. Data-driven prediction models are based on feature tensors constructed using S100. Train a time-series deep learning model as a prediction model. The data-driven prediction model learns the complex nonlinear mapping relationship between input features and output power from massive historical data. The core of the physical mechanism prediction model can be expressed as: ; Photocurrent; This is the diode saturation current; For electron charge; This refers to the output voltage. Output current; It is a series resistor; This is the diode ideality factor; Boltzmann's constant; Absolute temperature; It is a parallel resistor.

6. The multidimensional model and data-driven method for evaluating the aggregated output degradation of photovoltaic power plants according to claim 4, characterized in that, The method for identifying high-confidence abnormal time periods by comparing the prediction results of physical mechanism prediction models and data-driven prediction models with actual measured values ​​includes: Calculate the dual residuals: calculate the residuals between the predicted output and the actual output of the physical mechanism prediction model and the data-driven prediction model, respectively. ; ; Indicates at time The actual measured output value; Physical mechanism model at time The predicted output value; Data-driven models at time The predicted output value; Predict residuals for physical models; Data-driven model prediction residuals; To avoid the incompatibility of fixed thresholds under different weather conditions, dynamic thresholds are set. The thresholds can be calculated based on recent historical residuals.

7. The multidimensional model and data-driven method for evaluating the aggregate output degradation of photovoltaic power plants according to claim 4, characterized in that, The method for correcting and crediting the original health status sequence output by S200 based on anomaly identification results includes: For state sequence verification, the health estimate based on S200 is directly adopted for non-abnormal periods. For abnormal periods, for abnormal periods Labeling or interpolating using health values ​​from nearby normal periods can prevent anomalous data from contaminating the final decay trend analysis. Calculate reliability parameters, including a residual confidence level for each time step, as a quantitative indicator of the reliability of the state estimate at that point. The calculation formula is: ; For a moment The reliability of the state estimation result is a quantitative indicator with a value range of (0,1]. The larger the value, the higher the reliability of the state estimation result at that moment. The final output is a verified state sequence with reliability indicators. The state sequence not only includes the health status of the power plant, but also the reliability of each state point, providing a crucial weighting basis for trend prediction in S400.

8. The multidimensional model and data-driven method for evaluating the aggregated output degradation of photovoltaic power plants according to claim 1, characterized in that, Based on the verified attenuation state sequence and reliability identification parameters output by S3, a component-level attenuation rate weighted aggregation algorithm is constructed to establish an attenuation acceleration analysis model. This is combined with reliability scores for trend prediction to ensure the reliability of the evaluation results. Specifically, this includes: S410 component-level attenuation rate weighted aggregation, which linearly fits the verified health sequence of each component to obtain its individual annual attenuation rate; subsequently, weights are assigned according to the importance of the components to calculate the power plant-level aggregated attenuation rate. : ; This represents the total number of photovoltaic modules participating in the aggregation calculation; For component indexing; No. The annual degradation rate of an individual photovoltaic module; The S420 establishes a decay acceleration analysis model, and after smoothing the average health sequence at the power plant level, calculates its first-order and second-order differences: ; ; For a moment The rate of decay; The time interval between adjacent time steps; For a moment The average health of power plants; For a moment The average health of power plants; A persistently negative value is an early warning sign of accelerated decay. A persistently negative value is an early warning sign of accelerated decay; S430 combines reliability trend prediction with the verified state sequence with reliability identifiers output by S300 to construct a time series prediction model that integrates reliability parameters, and makes probabilistic predictions about the future health status of the power plant.

9. The multidimensional model and data-driven method for evaluating the aggregated output degradation of photovoltaic power plants according to claim 8, characterized in that, The method for constructing a time series prediction model that integrates reliability parameters based on the verified state sequence with reliability identifiers output by S300, and for probabilistically predicting the future health status of the power plant, includes: Construct a reliability-weighted training dataset by extracting historical health data and its corresponding reliability parameters from the verified state sequence output by S300; when constructing model training samples, assign a sample weight coefficient to each sample, which is determined by the average value of the reliability parameters in its corresponding time period; A prediction model integrating reliability information is established, and a Long Short-Term Memory (LSTM) network with an attention mechanism is used as the core architecture of the prediction model. The input features of the model include both the historical health sequence and the corresponding reliability parameter sequence. A weighted training strategy is implemented. During the model training process, a reliability-based weighted loss function is adopted. The weighted loss function is based on the reliability quantification index and sample weights of the S300 output, and uses the weighted mean square error as the loss function. The formula for calculating the weighted loss function is: ; This is the total number of training samples; and They are the first The actual health status of each sample and the model prediction; It is the first The weights of each sample; The model generates probabilistic prediction results, and its output uses a quantile regression architecture, simultaneously outputting multiple quantile values ​​for future health. Specifically, three nodes are set up in parallel at the model's output layer, each corresponding to one of the three specific quantiles of health. Their relationship is expressed as follows: ; in, This represents a baseline prediction of future health. and These three output values ​​collectively define an 80% confidence interval; using these three output values, a predictive interval for health status can be intuitively constructed; the width of this interval... This naturally reflects the degree of uncertainty in the prediction results, providing complete reference information for power plant operation and maintenance decisions, from the most likely scenario to the risk boundary.