Photovoltaic module service life prediction system and method based on artificial intelligence
By using artificial intelligence technology in the life prediction system of photovoltaic modules, data preprocessing, uncertainty analysis, hierarchical cascade processing and life prediction evaluation, the problem of parameter volatility in the existing system has been solved, and more accurate and stable life prediction has been achieved.
Patent Information
- Application Number
- CN202510217875.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-10
- Estimated Expiration
- Not applicable · inactive patent
Smart Images

Figure CN120124010A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine learning, and particularly to a photovoltaic module life prediction system and method based on artificial intelligence. Background Art
[0002] The photovoltaic module life prediction system is based on machine learning technology. By collecting, preprocessing and modeling the operation data of photovoltaic modules, it realizes the estimation of the remaining life of the modules and trend analysis.
[0003] However, there are certain limitations in the existing technology during the life prediction process. The feature extraction method fails to quantify the volatility of parameters, resulting in the model being difficult to accurately reflect the dynamic changes of the equipment operation state. Relying only on static data for analysis makes the prediction results lack adaptability. There is a lack of hierarchical feature modeling method, the data structure is single, and the complex correlation between variables is not fully utilized, which affects the identification ability of the prediction model for key influencing factors and reduces the prediction accuracy. Therefore, improvements are needed. Summary of the Invention
[0004] The purpose of the present invention is to solve the deficiencies existing in the prior art, and to propose a photovoltaic module life prediction system and method based on artificial intelligence.
[0005] To achieve the above purpose, the present invention adopts the following technical solutions: The photovoltaic module life prediction system based on artificial intelligence includes:
[0006] A data acquisition and preprocessing module, which obtains voltage values, current values and power values from the output end of the photovoltaic module, normalizes the voltage values, current values and power values, batch-reorganizes and segments the normalized results according to the sampling time sequence, and generates a preprocessed data sequence;
[0007] An uncertainty analysis module, which respectively performs deviation calculation and fluctuation analysis on the voltage values, current values and power values in each time period of the preprocessed data sequence to obtain parameter fluctuation characteristics; calculates the probability density based on the parameter fluctuation characteristics, and generates uncertainty distribution characteristics;
[0008] A task decomposition and optimization module, which hierarchically cascades the probability density and deviation mean in the uncertainty distribution characteristics to obtain a hierarchical feature set; assigns weight coefficients and performs weight coefficient matrix operations on the values of each level in the hierarchical feature set to generate a decomposition and optimization matrix;
[0009] A life prediction and evaluation module, which performs linear regression on each matrix element in the decomposition and optimization matrix and the operation duration value of the photovoltaic module to obtain a life correlation factor; calculates the probability and confidence interval analysis of the remaining duration of the photovoltaic module based on the life correlation factor, and generates a life prediction result.
[0010] Preferably, the step of obtaining the preprocessed data sequence is as follows:
[0011] Collect the voltage value, current value and power value from the output end of the photovoltaic module, normalize each value to obtain the normalized data;
[0012] Based on the normalized data, calculate the comprehensive index of each sampling point, and the formula is:
[0013]
[0014] where, GQ k is the normalized voltage value, GI k is the normalized current value, GP k is the normalized power value, m is the number of sampling points, and F s is the comprehensive index of the sampling point;
[0015] According to the comprehensive index of the sampling point, perform time series analysis and batch recombination on the normalized data, divide the data segments according to the set time window, and generate the preprocessed data sequence.
[0016] Preferably, the step of obtaining the parameter fluctuation characteristics is as follows:
[0017] Perform statistical analysis on the voltage, current and power in each time period of the preprocessed data sequence to obtain the mean value of each time period and generate the basic statistical data;
[0018] Based on the basic statistical data, calculate the comprehensive deviation value of the voltage, current and power in each time period, and the formula is:
[0019]
[0020] where, V i , I i , P i are the observed values of voltage, current and power respectively, are the corresponding time period mean values respectively, n is the number of observation points, and F dev is the comprehensive deviation value;
[0021] According to the comprehensive deviation value, analyze the fluctuation characteristics of voltage, current and power during the monitoring period to obtain the parameter fluctuation characteristics.
[0022] Preferably, the step of obtaining the uncertainty distribution characteristics is as follows:
[0023] Based on the parameter fluctuation characteristics, calculate the probability density of each parameter, and the formula is:
[0024]
[0025] where x represents the measured value of voltage, current or power in each time period, μ x represents the corresponding average deviation, and σ x represents the standard deviation, and P(x) represents the probability density;
[0026] By integrating the probability densities of all time periods, an uncertainty distribution feature is generated.
[0027] Preferably, the steps for obtaining the hierarchical feature set are as follows:
[0028] Based on the uncertainty distribution feature, levels are divided according to the distribution of probability density data, and then refined levels are set according to the range of deviation mean data to obtain a preliminary hierarchical feature structure;
[0029] Based on the preliminary hierarchical feature structure, the probability density data and deviation mean data of adjacent levels are cascaded, and the transfer relationship of data between levels is sorted out to form a hierarchical feature set.
[0030] Preferably, the steps for obtaining the decomposition and optimization matrix are as follows:
[0031] Based on the hierarchical feature set, the values of each level are screened to eliminate outliers and duplicates, and standardized hierarchical feature data is generated;
[0032] Based on the standardized hierarchical feature data, an initial weight coefficient is set according to the change range of the level values, and then it is adjusted in combination with the distribution of the values within the level to form a weight coefficient allocation result;
[0033] Based on the weight coefficient allocation result, a weight coefficient matrix is constructed, the weight influence relationship between each level is analyzed, the optimization adjustment parameters between levels are determined, and a decomposition and optimization matrix is generated.
[0034] Preferably, the steps for obtaining the life correlation factor are as follows:
[0035] Based on the decomposition and optimization matrix, a linear regression model is established, and the formula is:
[0036] L rf =β 0 +β 1 X 1 +β 2 X 2 +…+β k X k
[0037] where L rf represents the predicted life correlation factor, and X 1 , X 2 ,…, X krepresent the matrix elements in the decomposition optimization matrix, β 0 , β 1 , …, β k are regression coefficients;
[0038] Calculate the life correlation factor of the photovoltaic module based on the linear regression model.
[0039] Preferably, the steps for obtaining the life prediction result are as follows:
[0040] Extract the life correlation factor, call the operation data of the photovoltaic module, screen the life samples that meet the same operating conditions, and generate a life probability calculation data set;
[0041] Based on the life probability calculation data set, calculate the remaining duration of the photovoltaic module. The calculation formula is:
[0042]
[0043] where, T rem represents the calculated remaining duration, c represents the life sample value of the photovoltaic module, F lf represents the calculated life correlation factor, θ represents the scale parameter of the life data, k represents the shape parameter, and Γ(k + 1) is the Gamma function;
[0044] Based on the remaining duration, use the confidence interval calculation to determine the confidence interval of the remaining life of the photovoltaic module and generate the life prediction result.
[0045] The present invention provides an artificial intelligence-based method for predicting the life of a photovoltaic module, including the following steps:
[0046] Obtain voltage values, current values, and power values from the output end of the photovoltaic module, normalize the values, and then batch-recombine and segment the normalized data in the order of sampling time to generate a preprocessed data sequence;
[0047] Perform deviation calculation and fluctuation analysis on the voltage values, current values, and power values within each time period in the preprocessed data sequence. Based on the obtained fluctuation characteristics, calculate the probability density and construct and obtain the uncertainty distribution characteristics;
[0048] Perform hierarchical cascading on the probability density and deviation mean in the uncertainty distribution characteristics, perform weight coefficient allocation, and perform weight coefficient matrix operations to generate a decomposition optimization matrix;
[0049] Perform linear regression analysis on the elements in the decomposition optimization matrix and the operation duration value of the photovoltaic module to obtain a life correlation factor associated with the module life;
[0050] Based on the lifetime correlation factor, perform probability calculation and confidence interval analysis on the remaining operating duration of the photovoltaic module to generate and obtain the lifetime prediction result.
[0051] Compared with the prior art, the advantages and positive effects of the present invention are as follows:
[0052] In the present invention, deviation calculation and fluctuation analysis extract the parameter fluctuation characteristics from multiple time periods, enabling the prediction to cover the dynamic changes during the operation of the module and no longer relying solely on static data. The probability density calculation combines with the parameter fluctuation characteristics to form an uncertainty distribution characteristic, which can quantify the change trend of the data during the prediction process and reduce the influence of abnormal data. The hierarchical cascading processing method enables the expression of the data correlation at different levels, enhancing the analysis ability of the prediction model for the lifetime influencing factors. The weight coefficient allocation combines matrix operations, making the contribution degree of different parameters to the lifetime prediction more hierarchical and reducing the interference of local data anomalies on the overall evaluation. Based on the linear regression analysis of the matrix elements and the operating duration, the lifetime correlation factor can describe the operating state of the photovoltaic module, and combined with probability calculation and confidence interval analysis, the lifetime prediction can not only give numerical results but also provide credibility evaluation, reduce the error range, and improve the prediction stability. Description of the Drawings
[0053] Figure 1 It is the system flowchart of the present invention. Detailed Embodiments
[0054] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0055] Please refer to Figure 1 , the technical solution provided by the present invention is as follows: The photovoltaic module lifetime prediction system based on artificial intelligence includes:
[0056] A data acquisition and preprocessing module, which obtains voltage values, current values and power values from the output end of the photovoltaic module, normalizes the voltage values, current values and power values, and batch-reorganizes and segments the normalized results according to the sampling time sequence to generate a preprocessed data sequence;
[0057] An uncertainty analysis module, which respectively performs deviation calculation and fluctuation analysis on the voltage values, current values and power values in each time period of the preprocessed data sequence to obtain parameter fluctuation characteristics; calculates the probability density based on the parameter fluctuation characteristics to generate an uncertainty distribution characteristic;
[0058] The task decomposition and optimization module hierarchically cascades the probability density and deviation mean in the uncertainty distribution characteristics to obtain a hierarchical feature set; assigns weight coefficients to the values at each level in the hierarchical feature set and performs weight coefficient matrix operations to generate a decomposition and optimization matrix.
[0059] The lifespan prediction and evaluation module performs linear regression on each matrix element in the decomposition and optimization matrix and the photovoltaic module operation duration value to obtain a lifespan correlation factor; calculates the probability and confidence interval analysis of the remaining duration of the photovoltaic module based on the lifespan correlation factor to generate a lifespan prediction result.
[0060] The steps for obtaining the preprocessed data sequence are as follows:
[0061] Collect voltage values, current values, and power values from the output terminal of the photovoltaic module, and normalize each value to obtain normalized data.
[0062] Based on the normalized data, calculate the sampling point comprehensive index for each sampling point. The formula is:
[0063]
[0064] where GQ k is the normalized voltage value, GI k is the normalized current value, GP k is the normalized power value, m is the number of sampling points, and F s is the sampling point comprehensive index.
[0065] According to the sampling point comprehensive index, perform time series analysis and batch reorganization on the normalized data, divide the data into segments according to the set time window, and generate a preprocessed data sequence.
[0066] Specifically, the voltage value, current value and power value collected from the output end of the photovoltaic module, here referring to the photovoltaic module with a rated voltage range of 0V to 40V, a current range of 0A to 10A, and a power range of 0W to 400W, are continuously recorded at a sampling frequency of 1Hz, and a one-minute period recording data set is formed by integrating the data collected every second, in which the voltage value, current value and power value are normalized one by one. The specific method of normalization is to compare each collected record with a preset value range, scale the voltage data between 0V and 40V to obtain a normalized result between 0 and 1, and scale the current data The normalized result between 0 and 1 is obtained by scaling between 0A and 10A, and the power data is scaled between 0W and 400W to obtain a normalized result between 0 and 1. After the normalization operation is completed in the time sequence of the records, the normalized results here are checked and compared as a whole, and all records whose normalized values exceed the interval of 0 to 1 are marked, and the normalized value of the exceeding record is replaced by the average value of the adjacent time or the previous and next sampling records to ensure that all sampling records are within the range of 0 to 1. Subsequently, the normalized voltage value, current value and power value are sequentially stored in the same data structure. In this way, a record sequence of normalized data is generated.
[0067] The benefit of the formula is that through the combination of the square term and the parameters in the denominator, the joint contribution of voltage, current and power to the comprehensive index at the same time can be refined and quantified, so as to more accurately evaluate the working status of the photovoltaic module in subsequent data analysis or life prediction;
[0068] GQ k The parameter acquisition steps are as follows: normalize the original voltage data in the range of 0V to 40V to obtain a normalized voltage value between 0 and 1, and continuously record the data at k moments under the condition of a sampling frequency of 1Hz to obtain a set of voltage data sequences, and normalize the sequences one by one to form {GQ 1 ,GQ 2 ,…,GQ m} and store them. For example, in a certain test, each data in the voltage record is divided by 40V to obtain a batch of normalized values in the range of 0 to 1. These normalized values are then summarized for subsequent formula calls. For example, in a specific monitoring process, the original voltage record for k=1 is 34V, and after normalization, it is obtained
[0069] And so on to get all GQ k ;
[0070] GI kThe steps for obtaining the parameter are as follows: Normalize the original current data within the range of 0 A to 10 A to obtain a normalized current value between 0 and 1. Similarly, under the condition that the sampling frequency is 1 Hz, form {GI 1 , GI 2 , …, GI m} in chronological order. During actual acquisition, the scaling can be performed by measuring multiple groups of current data in parallel and comparing the intervals of 0 A and 10 A item by item. For example, during a certain detection, when k = 1, the original current record is 8 A, and by dividing it by 10 A, we get Thus, a complete normalized current sequence is formed;
[0071] GP k The steps for obtaining the parameter are as follows: Normalize the original power data within the range of 0 W to 400 W to obtain a normalized power value between 0 and 1. The power data here is obtained by detecting with an actual test instrument. The sampling frequency is also 1 Hz, and the power change of the photovoltaic module is continuously monitored within a certain period of time, so as to obtain {GP 1 , GP 2 , …, GP m}. During a certain monitoring, if k = 1 and the original power record is 316 W, then by dividing it by 400 W, we obtain All power normalization values are obtained in this way;
[0072] The steps for obtaining the m parameter are as follows: During the real-time monitoring of voltage, current, and power, multiple sampling records will be accumulated, forming the total number of sampling times within a period of time, which is the value of m. Specifically, it can be determined according to the monitoring duration of the photovoltaic module and the sampling frequency. For example, when monitoring for 5 seconds at a frequency of 1 Hz, 5 records are generated, and at this time m = 5. In actual monitoring, a longer monitoring duration will be set according to needs to obtain more data;
[0073] Calculation process:
[0074] First step, when k = 1, the numerator is 0.85 2 + 0.80 2 + 0.79 2
[0075] = 0.7225 + 0.64 + 0.6241 = 1.9866, and the denominator is
[0076] 0.85 × 0.80 + 0.79 = 0.68 + 0.79 = 1.47. Then the result at this moment is 1.9866 / 1.47 ≈ 1.3518;
[0077] Second step, when k = 2, the numerator is 0.87 2 + 0.81 2 + 0.822
[0078] = 0.7569 + 0.6561 + 0.6724 = 2.0854, and the denominator is
[0079] 0.87 × 0.81 + 0.82 = 0.7047 + 0.82 = 1.5247, so the result at this moment is
[0080] 2.0854 / 1.5247 ≈ 1.3676;
[0081] In the third step, when k = 3, the numerator is 0.88 2 + 0.83 2 + 0.80 2
[0082] = 0.7744 + 0.6889 + 0.64 = 2.1033, and the denominator is
[0083] 0.88 × 0.83 + 0.80 = 0.7304 + 0.80 = 1.5304, so the result at this moment is
[0084] 2.1033 / 1.5304 ≈ 1.3741;
[0085] In the fourth step, when k = 4, the numerator is 0.90 2 + 0.85 2 + 0.84 2
[0086] = 0.81 + 0.7225 + 0.7056 = 2.2381, and the denominator is
[0087] 0.90 × 0.85 + 0.84 = 0.765 + 0.84 = 1.605, so the result at this moment is
[0088] 2.2381 / 1.605 ≈ 1.3942;
[0089] In the fifth step, when k = 5, the numerator is 0.93 2 + 0.86 2 + 0.88 2
[0090] = 0.8649 + 0.7396 + 0.7744 = 2.3789, and the denominator is
[0091] 0.93 × 0.86 + 0.88 = 0.7998 + 0.88 = 1.6798, so the result at this moment is
[0092] 2.3789 / 1.6798 ≈ 1.4165;
[0093] Add the results of the above five moments to obtain F s
[0094] ≈1.3518 + 1.3676 + 1.3741 + 1.3942 + 1.4165 = 6.9042;
[0095] This result indicates that when F s The numerical value is greater than the threshold, the normalized states of voltage, current and power are relatively concentrated and are at a relatively high level as a whole. Set the threshold to 5.0 in this monitoring scenario. The actually calculated 6.9042 has exceeded 5.0. It can be inferred from this that the operating state of the photovoltaic module is relatively stable and the output power level is relatively high during this time period. When this index falls in a lower range, it is necessary to further check the abnormal ranges of each normalized data.
[0096] According to the comprehensive index of sampling points and the normalized data obtained in the previous step, it is necessary to perform time series analysis and batch recombination on the two. Here, first determine a fixed time window according to the monitoring duration. For example, select 60 seconds as an interval, package all the comprehensive indexes of sampling points and the normalized data in 60 consecutive seconds in chronological order, and aggregate them synchronously in combination with the timestamp information. Check whether the data volume in each 60 - second interval meets the sample requirements to ensure that there is sufficient data density for subsequent analysis. If the data volume is insufficient or there are excessive over - limit records in a certain window, mark the data in this window additionally in the unified record structure and store it separately. Next, these packaged window intervals will be sorted together according to continuity to form a set with multiple segments of sequences. Each segment of sequence contains voltage, current, power and the corresponding comprehensive index of sampling points. Subsequently, these multiple segments of sequences will be used for further construction of time series models. The main method is to read the sequences in each time period and compare their change rules when crossing different windows. Abnormal jitters or significant fluctuations can be identified based on the high - low distribution of the records at each moment, so as to complete the feature division of the photovoltaic module under different time windows, and finally obtain the pre - processed data sequence required for subsequent life prediction.
[0097] The steps to obtain the parameter fluctuation characteristics are as follows:
[0098] Perform statistical analysis on the voltage, current and power in each time period of the pre - processed data sequence to obtain the mean value of each time period and generate basic statistical data;
[0099] Based on the basic statistical data, calculate the comprehensive deviation values of voltage, current and power in each time period. The formula is:
[0100]
[0101] Among them, V i ,Ii , P i are the observed values of voltage, current, and power respectively, are the corresponding time - period means respectively, n is the number of observation points, F dev is the comprehensive deviation value;
[0102] According to the comprehensive deviation value, analyze the fluctuation characteristics of voltage, current, and power during the monitoring period to obtain the parameter fluctuation characteristics.
[0103] Specifically, based on the pre - processed data sequence obtained previously, perform continuous statistics on the voltage data, current data, and power data within each time period. In the statistical process, first read all the records within the time period, extract the voltage value, current value, and power value of each record, and compare them with the pre - set effective range. For example, the voltage is between 0V and 40V, the current is between 0A and 10A, and the power is between 0W and 400W. Screen and mark each record with an outlier one by one, and then summarize the observed values within the effective range for mean calculation. The mean calculation method is to sum up all the qualified observed values of the same type and then divide by the number of valid records. The voltage mean can be recorded as the sum of all valid voltage values within a certain period divided by the total number of voltage records within that period. The same applies to current and power. After the calculation, record the voltage mean, current mean, and power mean of each period as basic statistical data for performing more in - depth analysis in the subsequent process by combining these basic statistical data. To prevent insufficient record numbers caused by intermittent acquisition or equipment failures, when the number of valid records within a time period is less than a certain threshold value, it is considered that the data of this period cannot represent the normal operating state, and it is necessary to check the observation instrument again or supplement the monitoring duration, and then perform the same mean calculation and archiving based on the repeatedly observed data. Through this continuous cyclic and parallel calculation and inspection, obtain multiple voltage means, current means, and power means in multiple specified time periods, arrange them in sequence for each period, and finally generate the basic statistical data.
[0104] The advantage of the formula is that by summing up the absolute deviations of voltage, current, and power respectively and adopting a quantization method under the mean dimension, parameters with different dimensions are converted into a set of comprehensive deviation values, which can uniformly compare and highlight the relative weights of deviation magnitudes at different times.
[0105] V i The steps to obtain the parameter are as follows: During the process of monitoring the photovoltaic module, record multiple voltage data at a sampling frequency of once per second, use the pre - processed data sequence obtained previously to intercept all voltage observed values within a certain time period and save them as {V 1 , V 2 , …, V n}, where V iDenote the voltage observation value of the \(i\)-th item, whose value ranges from 0V to 40V. The specific value is obtained by actually measuring the terminal voltage of the photovoltaic module with a measuring instrument. For example, within a certain period of time, there are \(n = 6\) voltage records, and the detected effective voltage values are 32.3V, 33.1V, 31.8V, 34.0V, 32.7V, and 34.2V in sequence. Then, integrating these records in chronological order can form \(\{V 1 , V 2 , V 3 , V 4 , V 5 , V 6 \}\) for subsequent calculations.
[0106] The steps to obtain the \(I i parameter are as follows: Similarly, within the specified time period, collect the current observation values at the terminals of the photovoltaic module, and record all the measured current data to form \(\{I 1 , I 2 , …, I n \}\). Its value range is usually between 0A and 10A, and the detection is completed by real-time measuring an ammeter or a sensor. For example, within the same time period, the recorded current values corresponding to the \(i\)-th sampling are 7.5A, 7.2A, 7.8A, 8.1A, 7.6A, and 7.9A in sequence. Then, store this sequence for subsequent calculation of the comprehensive deviation value. All acquisition links can be reused again. Just maintain the same sampling frequency and recording rules under different actual measurement scenarios to obtain \(I i \).
[0107] The steps to obtain the \(P i parameter are as follows: Obtain the power observation value of the photovoltaic module within this time period by multiplying the voltage observation value by the current observation value, or by real-time measurement with a power meter, etc., and record it as \(\{P 1 , P 2 , …, P n \}\). The power range can set a reference interval based on the rated power of the photovoltaic module. For example, between 0W and 400W. If the measured power values in actual monitoring are 255W, 268W, 260W, 275W, 271W, and 279W respectively, then store these observation results as a power sequence for subsequent statistical analysis. Each power record can be matched with the corresponding voltage and current data within the same monitoring period to ensure the consistency of the three parameter sources.
[0108] The steps to obtain the parameter are as follows: Based on all \(V\) within this time period iThe records are added up and then divided by the number of records \(n\) to obtain the average voltage, and the calculation method remains stable when the acquisition frequency and duration are fixed. For example, based on the previous 6 voltage records, the result of \(32.3V + 33.1V + 31.8V + 34.0V + 32.7V + 34.2V\) is divided by 6, resulting in approximately \(33.0V\), and this value can also be corrected by extending the observation duration or repeating the measurement.
[0109] The steps to obtain the parameter are as follows: Sum up all the I i current observation values within this time period and divide by \(n\) to obtain the average current. Assuming that in one continuous acquisition, the sum of \(7.5A + 7.2A + 7.8A + 8.1A + 7.6A + 7.9A\) is \(45.1A\), dividing by 6 gives approximately \(7.52A\). By calculating and recording according to the principle of continuous operation of the monitoring instrument and fixed frequency in multiple time periods, multiple groups can be continuously accumulated for subsequent analysis or prediction processes.
[0110] The steps to obtain the parameter are as follows: Similarly, add up all the P i power values within this time period and divide by \(n\) to obtain the average power. For example, the sum of \(255W + 268W + 260W + 275W + 271W + 279W\) is \(1608W\), and dividing \(1608W\) by 6 gives approximately \(268W\). Subsequently, this average value can also be stored as a reference value representing the power level of this time period for convenient comparison.
[0111] The steps to obtain the \(n\) parameter are as follows: Determine the value of \(n\) by checking the total number of sampling records within this time period, which specifically depends on the monitoring duration and sampling frequency. For example, if sampling at a frequency of 1Hz for 10 seconds, then \(n = 10\). In the previous examples, \(n = 6\) because only the operations of 6 records were demonstrated. Each time, ensure that the selection of \(n\) corresponds to the number of valid records within the same time period, and avoid directly changing \(n\) due to missing values or outliers, which may affect the correctness of the formula.
[0112] Calculation process:
[0113] First step, set the number of observation points \(n = 6\), substitute the observation values and average values of the previous examples, and calculate the corresponding deviation terms for each \(i\):
[0114]
[0115] \(\vert32.3 - 33.0\vert+\vert7.5 - 7.52\vert+\vert255 - 268\vert\)
[0116] The value is approximately equal to 0.7+0.02+13=13.72. In the same way, the values from i=2 to i=6 are added up and recorded as the sum of the numerators. The second step is to divide the denominator. according to Calculate The third step is to divide the deviation items corresponding to each record by 270.18 and sum them to get the final F dev For example, the total of all is about 57.60, and then 57.60 is divided by 270.18 to get about 0.2133, which is the comprehensive deviation value for this time period.
[0117] The results show that when F dev The higher it is, the greater the deviation of the voltage, current and power during this time period relative to their mean values. If this value is lower than 0.1, it indicates that the deviation is small. If it exceeds 0.3, it indicates that the fluctuation is more obvious.
[0118] After obtaining the comprehensive deviation value, it is necessary to process the change status of voltage data, current data and power data within the monitoring period based on the result. The specific method is to compare the comprehensive deviation value of each time period section by section, compare the time period with higher deviation value with the adjacent time period in the same category, record whether there is a significant increase or decrease, and then check the voltage value distribution, current value distribution and power value distribution in different time periods one by one, identify the significant increase or decrease that may exist, and take out the maximum and minimum values in each time period to define the range. In order to better grasp the fluctuation characteristics, each collected record can also be sorted by the absolute difference value, and the top-ranked records can be taken out separately. It can be found and continuously tracked to see whether there are large fluctuations in subsequent time periods. If the same large deviation is shown for multiple consecutive periods of time, it means that the corresponding monitoring record is continuously fluctuating in a certain direction. If the next comparison shows that the value has fallen back to a relatively stable level, the decline interval is recorded and marked. Each marked interval and record will be summarized and stored in the same data structure, with an annotation of the degree of deviation from the previous mean. This method of sequence comparison and sub-item inspection can be used to screen out fluctuations between different monitoring cycles, and then the fluctuation characteristics of voltage, current, and power can be mapped to a visual distribution sequence. Through multiple cross-time period comparisons, the parameter fluctuation characteristics can be finally obtained.
[0119] The steps to obtain the uncertainty distribution characteristics are:
[0120] Based on the parameter fluctuation characteristics, the probability density of each parameter is calculated using the formula:
[0121]
[0122] Among them, x represents the measured value of voltage, current, or power in each time period, μ x represents the corresponding average deviation, and σ x represents the standard deviation, and P(x) represents the probability density;
[0123] By integrating the probability density of all time periods, an uncertainty distribution characteristic is generated.
[0124] Specifically, the advantage of the formula is that by introducing the deviation expressions of the cubic and quartic powers in the exponential term, it can more flexibly reflect the asymmetry or heavy-tailed characteristics presented by the output parameters of photovoltaic modules in different intervals. Compared with the general normal distribution based on the quadratic power, it can describe the outlier situation or peak phenomenon more precisely.
[0125] The steps for obtaining the x parameter are as follows:
[0126] During the monitoring of the voltage, current, and power of photovoltaic modules, the single observation value in each time period is uniformly referred to as x, including voltage observation values, current observation values, or power observation values, etc. If the voltage distribution is to be analyzed in practical applications, all voltage measurement records in this time period are marked as {x 1 , x 2 , …, x m}, and sufficient numbers of records are collected from the monitoring instrument according to the sampling frequency. The terminal voltage of the photovoltaic module is measured by a sensor with a national or industry-recognized accuracy level. The numerical range is mostly between 0V and 40V. If the current is measured, the range is commonly between 0A and 10A, and the power may be around 0W to 400W. It is extended according to the specific component power level. For example, during a continuous monitoring period, several voltage values are successively collected, and these values may be 28.6V, 29.2V, 31.1V, 32.4V, etc. After all records are incorporated into the same set, they can be corresponding to x. As long as the sensor reading frequency and time window are stable, a series of x values in time series can be obtained. During a 10-minute monitoring, if the acquisition frequency is set to 1Hz, 600 voltage records can be obtained. After comparing each record with the valid range and removing the abnormal points, the remaining valid measurement values are saved as {x 1 , x 2 , …, x n} for subsequent processing and calculation of this formula.
[0127] The steps for obtaining the μ x parameter are as follows:
[0128] μ xrepresents the average deviation of the corresponding measured value during this time period. Different from the direct arithmetic mean, it is usually necessary to first calculate the error between each measured value and the ideal level or reference state, and then combine relevant statistical indicators to obtain the average deviation. For example, in the analysis of the voltage of photovoltaic modules, a reference expected value can be obtained through a large number of historical monitoring samples first, and then the difference between each actually measured x and this expected value is calculated. Furthermore, all the difference values are averaged over a relatively long statistical period to obtain μ x To illustrate in a more specific way: During the continuous operation of 300 hours of a photovoltaic module, the voltage value per hour is subtracted from a certain reference value (such as 30V) item by item to form a sufficiently large sequence of differences. Then this sequence is accumulated strictly according to time segments. After summarizing each segment of differences and dividing by the number of difference items in that segment, repeating multiple times can obtain the average deviation value of this segment. Finally, after combining all segment results and removing abnormal or extreme situations, a more stable average deviation quantity, namely μ, is comprehensively obtained x In practice, the expected value determined by industry specifications or a large-scale measured database can also be used as a reference benchmark. In a complete example, assume that the sequence of differences between the voltage recorded during a certain continuous monitoring and the reference 30V is {-1.2V, -0.8V, +1.1V, +2.4V,...}. The average value of all differences is obtained by superimposing and dividing by the number of items to get a fixed value, such as 1.5V. This 1.5V represents μ under this segment of observations x for formula substitution
[0129] σ x The steps to obtain the parameter are as follows
[0130] σ x is the standard deviation corresponding to μ x and is usually obtained by calculating the difference between each measured value and μ item by item, then summing the squares, dividing by the number of items, and taking the square root. However, here it may be necessary to make further corrections to the average deviation quantity during the actual measurement of photovoltaic modules, such as removing noise data points or extreme fault points, to ensure that σ x can truly reflect the dispersion degree of most records. For example, σ can be obtained according to the following steps x : First, subtract the corresponding μ from all x observed values to obtain a number of differences; second, square and sum these differences; third, divide by the number of observations or the number of observations minus 1 (when considering unbiased estimation), and finally take the square root of the result. For example, in a series of voltage differences monitored for 300 hours, some inaccurate measurement points during the early instrument calibration period are removed, and finally the subsequent qualified N records are calculated. If the accumulated sum of squares is 2100 (V x ) and divided by N to get 70 (V x ), and then take the square root to obtain σ 2 ) and divided by N to get 70 (V 2 ), and then take the square root to obtain σ x≈8.37 V.
[0131] Calculation process:
[0132] First step, determine the values of each parameter. For example, in a monitoring interval, select x = 28.0 V, μ x = 30.0 V and σ x = 5.0 V, and substitute them into the denominator term :
[0133]
[0134] Second step, calculate the cubic deviation part in the exponential term:
[0135] (x - μ x ) 3 = (28.0 - 30.0) 3 = (-2.0) 3 = -8.0
[0136]
[0137] Thus, the exponential part exp(0.02133) ≈ 1.02155 is obtained. Third step, calculate the quartic deviation part in the exponential term:
[0138] (x - μ x ) 4 = (28.0 - 30.0) 4 = (-2.0) 4 = 16.0
[0139]
[0140] Thus, another exponential part exp(-0.0064) ≈ 0.99362 is obtained. Fourth step, multiply the exponential parts by respectively and accumulate to obtain the final P(x):
[0141] P 1 = 0.0798 × 1.02155 ≈ 0.08156, P 2 = 0.0798 × 0.99362 ≈ 0.07929
[0142] P(x) = P 1 + P 2 = 0.08156 + 0.07929 = 0.16085
[0143] This result shows that when x = 28.0 V, the corresponding probability density is approximately 0.16085. If the same μ x and σ xWhen applied to other different voltage values, such as 25.0V or 35.0V, the entire distribution curve can also be calculated point by point according to this formula. When the probability density value is at a relatively high level, it indicates that the corresponding observed value appears in a more common area of the overall distribution. On the contrary, when the value is significantly small, it means that its occurrence probability is not very high. It can be compared with other time periods. Subsequently, if it is necessary to evaluate whether the voltage deviates from the normal state during a certain period in a statistical sense, it is possible to check whether the calculated P(x) is lower than a certain reference density threshold, such as 0.05, etc., so as to judge its degree of abnormality.
[0144] According to the probability density results of voltage, current, and power at different time periods obtained previously on multiple measurement values, first select the P(x) sequences that have been calculated for each time period, observe the more concentrated or more discrete regions point by point, and summarize all time periods. Then, compare them in the same coordinate system, record the differences in the x distribution curves corresponding to each monitoring interval, and then conduct a more refined analysis of these curves. For example, calculate the union or intersection distribution of their overlapping regions to determine in which intervals the voltage value distribution is relatively stable or the current and power fluctuation intervals are more biased towards the high value range. During the process, an index table grouped by time period can be constructed according to the previously obtained mean deviation and standard deviation information. Then, select some distribution segments that exceed the specified range or are near a specific threshold demarcation point. For example, delimit the interval where the voltage probability density is greater than 0.15 and compare it interactively with the interval where the current probability density is greater than 0.10. If there is a large amount of overlap, record the corresponding time period, and then compare the change moments of the power probability density. If there are many probability density concentration phenomena for all three in the same time period, mark this period as a high-frequency distribution area. On the contrary, if the voltage probability density in some time periods is lower than 0.02 and the power probability density is also lower than 0.02, it is regarded as a low occurrence probability segment. The special reasons for the data records in it can be further checked, and these low probability segments can be stored separately from other time periods and integrated with subsequent cross-segment data statistics. Finally, integrate the probability density information of all time periods into a unified distribution feature file, and each time period will be attached with its probability density curve and numerical interval division content. After completion, the uncertainty distribution feature is obtained.
[0145] The steps to obtain the hierarchical feature set are as follows:
[0146] Based on the uncertainty distribution feature, divide the levels according to the distribution of the probability density data, and then set the refined levels according to the range of the mean deviation data to obtain the preliminary hierarchical feature structure;
[0147] Based on the preliminary hierarchical feature structure, cascade the probability density data and mean deviation data of adjacent levels, and sort out the transfer relationship of the data between levels to form the hierarchical feature set.
[0148] Specifically, based on the probability density data and related deviation mean data contained in the uncertainty distribution characteristics, the probability density distribution corresponding to each time period recorded previously is read first, and these data lists are grouped according to the three categories of voltage, current and power, and a numerical index table is established for them. Then, according to the statistical results of this index table, several reference intervals are selected as reference thresholds for preliminary division. If there are more records in the probability density data concentrated in the interval with a value greater than 0.20, the interval is identified as the first level. If there are relatively scattered records concentrated between 0.10 and 0.20, it is defined as the second level, and the interval with a probability density below 0.10 is defined as the second level. It is classified into three levels. At the same time, combined with the range of the deviation mean data, the records with deviation mean greater than a certain empirical threshold (for example, 2.0) are distinguished with additional marks. This empirical threshold can be determined by reading similar photovoltaic module test documents or analyzing statistical charts of the past hundreds of hours of measurement. Then, it is refined on the basis of the above three levels. For example, the deviation mean in the range of 2.0 to 3.0 is further classified into a sub-level, and the deviation mean exceeding 3.0 is classified into another sub-level. After recording all the stratification results, they are superimposed on a hierarchical index chart according to the logical order of the intervals, and finally a mapping that matches the numerical interval with the corresponding probability density range is formed, thereby obtaining a preliminary stratified feature structure.
[0149] According to the preliminary hierarchical feature structure generated earlier, first read the probability density interval and deviation mean interval corresponding to each level one by one, and compare the numerical ranges between adjacent levels. During the comparison process, first check whether these levels have overlapping or similar ranges in probability density values, and further check the upper and lower limits of the deviation mean. If the probability density intervals of adjacent levels are close to each other or there is an intersection, the records of these two levels are merged for cascade operation, and the deviation mean range of the merged data is also arranged together. Then all the merged contents are compared in sequence, and the records at the junction of different levels are marked in a unified manner. If it is found that the adjacent levels of the deviation mean span a large numerical range, continue to subdivide or merge the existing level entries, and write the sorted adjacent level mapping relationship into the corresponding mapping table. After completion, review the connection of all levels and the order of mutual transmission, record the level number and its corresponding probability density and deviation mean information one by one, and a hierarchical feature set can be formed.
[0150] The steps to obtain the decomposition optimization matrix are:
[0151] Based on the hierarchical feature set, the values of each level are screened, outliers and duplicates are removed, and standardized hierarchical feature data are generated;
[0152] Based on the standardized hierarchical feature data, set the initial weight coefficient according to the variation range of the hierarchical values, and then adjust it in combination with the distribution of the internal values of the layer to form the weight coefficient allocation result;
[0153] Based on the weight coefficient allocation result, construct a weight coefficient matrix, analyze the weight influence relationship between each layer, determine the optimization adjustment parameters between layers, and generate a decomposition optimization matrix.
[0154] Specifically, in the process of screening the values of each layer based on the hierarchical feature set, first extract the voltage level value, current level value, and power level value from all the hierarchical values recorded previously, and compare all the stratified values with the corresponding interval ranges. For example, compare the voltage level value with the range of 0V to 40V, compare the current level value with the range of 0A to 10A, and compare the power level value with the range of 0W to 400W. Check item by item whether there are abnormal records exceeding the rated range of the device or lower than the reasonable lower limit. If it is found that some values deviate significantly from the previously obtained stratified interval or there are exactly the same data entries appearing multiple times in the same interval, then remove these records or mark them as duplicates. Then, refer to the sub-intervals previously divided according to the deviation mean to regroup the remaining valid data again, and integrate the records with the same stratification and extremely small numerical differences into a unified set. For intervals greater than a certain threshold or less than a certain threshold, separate segmentation processing will also be carried out when necessary. For example, the records with voltage level values exceeding 38V are classified into the high-voltage sub-interval. The threshold of 38V here is set according to the common rated voltage of 40V of the photovoltaic module. After multiple statistics, it is found that the ratio of jitter increases when the voltage exceeds 38V. Therefore, 38V is selected as an empirical threshold to distinguish the normal range from the range close to the limit. A similar method will also be applied to the current and power level values to obtain refined intervals. Integrate and re-number all the data after screening and removing duplicates and abnormal items according to their respective layers to form a numerical table containing multiple segmented hierarchical entries. Subsequently, conduct a consistency verification on the values of each layer in this table. For example, check whether there are records with both the voltage deviation mean and the current deviation mean being too low under the same layer, and check whether the power deviation mean meets the interval standard previously formulated. Remove all records that do not meet the conditions or appear repeatedly through these steps, and then format the remaining integrated data and mark the layer name and the corresponding numerical interval. Finally, name this result as the standardized hierarchical feature data.
[0155] When setting the initial weight coefficients based on the standardized hierarchical feature data according to the variation range of the hierarchical values, first select all voltage level entries from the standardized hierarchical feature data, compare the span of the voltage values item by item with the upper and lower bounds of the previously defined interval, and set an initial weight coefficient by calculating the proportion of the length of the interval where the voltage level value is located to the overall range. For example, when the voltage level value interval is between 30V and 32V, its weight coefficient can be initially set to 0.7. If the interval is between 35V and 38V, then according to the magnitude of the deviation mean and the level of fluctuation probability in the previous statistics, the initial weight coefficient is set to 0.9. These values are not randomly given, but are obtained by sorting out the historical statistical data at the same level and observing the fluctuation frequencies. After that, the same process is performed on the current and power level values, and an initial weight coefficient is calculated for each level. After unifying and recording all these coefficients, they are adjusted according to the distribution within each level. If it is found that in a certain voltage level, although the values in the interval of 30V to 32V mostly concentrate around 31V, there are also a small number of records close to the edges of 30V or 32V, it may mean that there are more obvious internal differences within this level. At this time, the weight coefficient will be re-corrected according to the degree of dispersion of the internal distribution. For example, the original 0.7 is refined to 0.75 or 0.65, and at the same time, it is checked whether there are similar internal dispersions in the current and power at the corresponding level. If so, their initial weight coefficients are corrected at the same time. After all adjustments are completed, the final version of the weight coefficients are respectively attached to the data entries of the corresponding levels and the adjustment basis is indicated. For example, if the adjustment is due to a high degree of internal dispersion, the specific deviation mean distribution and the comparison result with the previous fluctuation probability will be pointed out in the description. After the weight coefficient corrections for all levels are completed, they are recorded in a detailed list and uniformly named as the weight coefficient distribution result.
[0156] When constructing the weight coefficient matrix based on the weight coefficient allocation results, first read the corresponding weight values from the weight coefficient allocation results according to the hierarchical serial numbers, arrange these weight values in the rows and columns of the matrix to obtain a square matrix structure, and then determine the arrangement order of the positions of each row and column according to the logical order between the levels, pair and compare these weights column by column or row by row to identify possible similar or approximate coefficient values between different levels. Then, judge the magnitude of the influence relationship between levels through the difference values obtained by comparison. If the difference between the weight coefficients of two levels is lower than a certain threshold (for example, 0.05), it means that these two levels are relatively similar in terms of deviation mean and fluctuation characteristics, etc., and they can be marked at the corresponding positions in the matrix. If the difference reaches or exceeds this threshold, it will be separately identified as a significantly different part in the matrix. Here, the threshold can be determined by statistical analysis of the previous data, or can be set with reference to the amplitude of the rated output of the photovoltaic module and the laboratory test results. According to the positions of these significantly different or significantly similar cells in the matrix, split or merge the weight relationships between adjacent or associated levels to form the optimized adjustment parameters between levels. Compare these adjustment parameters with the weights of each level in the matrix and record them in the same data table, and then recombine the optimized matrix entries into a clearer decomposition and optimization matrix, and give the numerical positions corresponding to each level and the finally adjusted coefficient values to complete this decomposition and optimization matrix.
[0157] The steps for obtaining the life correlation factor are as follows:
[0158] Based on the decomposition and optimization matrix, establish a linear regression model, and the formula is:
[0159] L rf =β 0 +β 1 X 1 +β 2 X 2 +…+β k X k
[0160] Among them, L rf represents the predicted life correlation factor, X 1 , X 2 ,…, X k represent the matrix elements in the decomposition and optimization matrix, and β 0 , β 1 ,…, β k are the regression coefficients;
[0161] Calculate the life correlation factor of the photovoltaic module based on the linear regression model.
[0162] Specifically, the benefit of the formula is that through the linear combination of multi-dimensional matrix elements and regression coefficients, the multi-level characteristics exhibited by voltage observation values, current observation values, and power observation values in the previous decomposition and optimization matrix can be quantitatively mapped to a comprehensive value, thereby more directly measuring the life correlation factor of photovoltaic modules.
[0163] Parameter X 1 Obtaining steps of: X 1 is the first matrix element in the decomposition and optimization matrix, usually used to characterize the characteristic quantity corresponding to the voltage level of the photovoltaic module within a certain period, including the fluctuation statistical results of the voltage in different time windows, the cumulative deviation value of the fluctuation relative to the reference voltage, etc. To obtain X 1 , it is necessary to comprehensively process the voltage deviation, probability density interval, and normalized statistical results at the same level in the previous steps, and then allocate them to the corresponding row and column positions in the final decomposition and optimization matrix to form a numerical entry. In the example, frequency statistics can be performed on information such as the average voltage deviation within 300 hours. If it is found that a certain type of deviation accounts for 20% in the high voltage range (for example, greater than 36V and less than 40V) and the probability density concentration is relatively high, it is written into the decomposition and optimization matrix as a quantitative value. For example, this value is determined to be 1.20 after multiple aggregations and screenings. After completing the entire extraction process, 1.20 can be assigned to X 1 for subsequent regression operations.
[0164] Parameter X 2 Obtaining steps of: X 2 is the second matrix element in the decomposition and optimization matrix, mainly corresponding to the indicators of the current level of the photovoltaic module. To obtain this indicator, it is necessary to analyze the previous observation records of different current intervals and their distribution under multi-level subdivision, especially pay attention to the high-probability section and low-probability section where the current appears in the range of 0A to 10A, and then calculate the final single value in combination with the average deviation and uncertainty distribution characteristics in this section. For example, during a period of detection, the average current deviation is concentrated when it is between 6.5A and 7.5A, and the coupling degree between this concentrated section and the power change is also relatively high. A quantity reflecting the current level characteristics can be obtained through counting and weighted integration. If the integrated value is 0.95, it is recorded as X 2 and placed in the decomposition and optimization matrix to participate in the regression model. The finally obtained value must have a clear source, which can be obtained by statistically analyzing the comprehensive performance after the average current interval between 6.5A and 7.5A accounts for about 40% within all time periods based on a fixed sampling frequency of 1Hz and a monitoring duration of 120 hours.
[0165] Parameter X 3Obtaining steps: If we take k = 3 as an example, then X 3 is the third matrix element of the decomposition optimization matrix, mainly reflecting the main eigenvalue of the power level in the previous multi-layer decomposition. When obtaining it, we can first screen out some high-frequency occurrence sections from the power history monitoring data according to the previous steps, just like making a long-term observation of the power dimension in the actual field. Suppose within the range of 0W to 400W, after multiple segmentations, it is found that the frequency of occurrence and the average deviation of a certain interval (for example, 290W to 310W) are both moderately high, and the corresponding probability density is also relatively concentrated in the previous statistics. Then we can uniformly weight the proportion of this interval and its fluctuation range to obtain a quantitative value representing the power stratification situation. For the convenience of example, after summarizing 220 consecutive hours of monitoring, it is found that the concentration proportion of this power section is 0.88, and the reasonable interval is confirmed by combining the reference threshold values of previous periods. Take this 0.88 as X 3 and input it into the subsequent regression model.
[0166] Parameter β 0 Obtaining steps: β 0 represents the intercept term in the linear regression model and can be obtained by performing least squares fitting or multiple linear regression training on the historical operation samples of the photovoltaic modules. The specific approach is to establish a regression relationship between a large number of existing decomposition optimization matrices and the corresponding life observation data, and then find the most suitable intercept value in the fitting equation. To improve accuracy, it is necessary to collect samples as much as possible under different component models, different time periods, and different geographical environments to avoid over-reliance on a single condition. For example, based on the hundreds of thousands of data accumulated previously, we can compare {X 1 , X 2 , X 3} of each sample with the true life decay index, and continuously update the intercept term in a regression coefficient training table using the least squares method or the gradient descent method. Finally, if it converges to β 0 = 1.5, it indicates that when there is no other characteristic influence, the basic value of the life correlation factor is approximately 1.5. In the example, various states including early decay and stable output period can be incorporated into the training set together to ensure the 0 scientificity and universality of β.
[0167] Parameter β 1 Obtaining steps: β 1 is the regression coefficient that cooperates with X 1 and is used to measure the influence weight of the voltage level characteristics on the life correlation factor. To obtain β 1, it is also necessary to compare and train each voltage level index with the actual life attenuation data based on a large number of previous samples, fit the voltage fluctuation and deviation mean information of thousands of segments or more included in the training set, and find out the influence degree of the voltage level characteristics on the final life attenuation. If β is obtained after training 1 = 0.30, it means that when the voltage level value X 1 increases by one unit, the corresponding life correlation factor L rf will increase by 0.30. In a specific example, voltage deviation stratified data collected under different sunlight irradiation levels can be selected, and the actual performance attenuation degree of the components can be compared, and β can be converged through multiple rounds of linear regression or cross-validation 1 .
[0168] The acquisition steps of parameter β 2 : β 2 cooperates with X 2 to represent the contribution degree of the current level index of the photovoltaic module to the life correlation factor. This coefficient needs to find out the true life changes corresponding to multiple current intervals in the large-scale samples prepared in the early stage, and then obtain the coefficient value through regression fitting. For example, compare the component data monitored for up to 500 hours, count the life maintenance situation of the component in different intervals of current fluctuation, and perform regression processing on the corresponding current characteristic value X 2 and the final life result, so as to calculate β 2 . If the result shows that β 2 = 0.45 after large-scale training, it means that when X 2 increases by one unit, the life correlation factor increases by 0.45. If this coefficient is too large, it means that the current level has a more significant impact on the component life attenuation; if it is too small, it means that the current only makes a limited contribution.
[0169] The acquisition steps of parameter β 3 : β 3 cooperates with X 3 to reflect the influence of the power level characteristics of the photovoltaic module on the life correlation factor. Power is directly related to the output efficiency, so this coefficient often needs to be observed both in the actual high-power output stage and the relatively low-power operation stage. When obtaining it, the power data of thousands of hours can be associated with the component attenuation rate detected later, and multiple typical intervals can be selected for fitting training. If β 3 = 0.27 is obtained after iteration, it means that when the power level index increases by one unit, the life correlation factor will increase by 0.27 in this model. It should be noted that for different types of photovoltaic modules, as well as different temperature and humidity conditions, this coefficient may change. Therefore, before establishing a regression model, it is necessary to ensure that the sampling and training data have sufficient diversity so that β 3 can more accurately reflect the general situation.
[0170] Calculation process:
[0171] In the first step, prepare the observation data and determine the sample parameter values, taking k = 3. Let β 0 = 1.50, β 1 = 0.30, β 2 = 0.45, β 3 = 0.27, and let X 1 = 1.20, X 2 = 0.95, X 3 = 0.88;
[0172] In the second step, substitute these values into the formula:
[0173] L rf = 1.50 + 0.30×1.20 + 0.45×0.95 + 0.27×0.88
[0174] In the third step, calculate each product respectively:
[0175] 0.30×1.20 = 0.36, 0.45×0.95 = 0.4275, 0.27×0.88 = 0.2376
[0176] In the fourth step, add all the results together:
[0177] L rf = 1.50 + 0.36 + 0.4275 + 0.2376 = 2.5251
[0178] Finally, we get L rf ≈ 2.5251, and the accuracy of this value can be further corrected or compared according to the long-term operation attenuation records of the same component in subsequent actual detections.
[0179] This result indicates that the life correlation factor at this time is approximately 2.5251. If a certain interval is preset in the statistical analysis (for example, when it is lower than 2.0, it is considered that the life attenuation degree is low; when it is between 2.0 and 3.0, it is considered that the life enters the moderate attenuation state; when it is greater than 3.0, it indicates a more significant attenuation), then 2.5251 can be judged as the moderate attenuation stage, and more subsequent data can be combined to track whether there is a trend towards a higher attenuation level.
[0180] Calculate the life correlation factor of the photovoltaic module based on the linear regression model. First, find the corresponding parameters for the voltage level, current level, and power level in the regression coefficient records obtained previously. Then, perform a comparison and matching based on the values of each level extracted from the decomposition and optimization matrix. Analyze item by item the range and fluctuation amplitude of the values of each level within different time periods. If any record exceeds the pre-set abnormal threshold, it is marked as an observation item that cannot be directly used. The threshold can be selected by combining multiple discharge tests in the laboratory and the rated parameters of the photovoltaic module. For example, in the voltage level, consider 40V as a certain limit value. If it exceeds this value, it is confirmed that the operating scenario of the module has an extreme situation and needs to be monitored separately. Multiply the values of all levels within the normal range by the regression coefficients, sum up the products in the order of levels, read the previously calculated intercept term and add it to the sum result to obtain the life correlation factor values corresponding to each time period. By comparing these values, the evolution of the life correlation factor of the photovoltaic module in each time period after running continuously for hundreds of hours can be observed, and based on this, the numerical changes of the photovoltaic module in different external environments such as high sunlight, low sunlight, high temperature, and normal temperature can be sorted out in segments. If the life correlation factor in some time periods far exceeds the reference range of the same type of samples, record the specific output power, current deviation, and voltage fluctuation data of the photovoltaic module for subsequent more in-depth tracking of this time period using the hierarchical feature set, and check whether it is related to local faults or component aging through the change information of the matrix elements of the decomposition and optimization matrix. Finally, summarize the life correlation factor data of multiple time periods to obtain an observation curve that extends over time.
[0181] The steps to obtain the life prediction result are as follows:
[0182] Extract the life correlation factor, call the operation data of the photovoltaic module, screen the life samples that meet the same operating conditions, and generate a life probability calculation data set;
[0183] Based on the life probability calculation data set, calculate the remaining duration of the photovoltaic module. The calculation formula is:
[0184]
[0185] where, T rem represents the calculated remaining duration, c represents the life sample value of the photovoltaic module, F lf represents the calculated life correlation factor, θ represents the scale parameter of the life data, k represents the shape parameter, and Γ(k + 1) is the Gamma function;
[0186] Based on the remaining duration, use the confidence interval calculation to determine the confidence interval of the remaining life of the photovoltaic module and generate the life prediction result.
[0187] Specifically, after extracting the life correlation factors, check whether the voltage value is in the range of 0V to 40V, the current value is in the range of 0A to 10A, and the power value is in the range of 0W to 400W item by item according to the photovoltaic module operation data obtained previously. Then select life samples that meet the same operating conditions from these records, including the same irradiation intensity and external temperature range. If some records have been marked as abnormal or high-temperature overlimit, skip this record and continue to read backward. Then divide the sampling period into several sub-intervals, aggregate all sub-intervals that meet the same irradiation intensity and external temperature together, and statistically calculate the average deviation of voltage, current, and power within these sub-intervals. Compare the average deviation with the pre-defined fluctuation threshold, which can be set by referring to the value table generated in a large number of previous tests. For example, if it is found that the sample life remains stable when the voltage is between 30V and 32V, then take 32V as a reference threshold. If a period with an average deviation close to or exceeding 32V is encountered during the test, mark it as another category. After statistically analyzing these selected intervals, combine them to obtain a preliminary life sample. By cross-mapping item by item with the life correlation factors, a life probability calculation data set can be generated. Attach timestamp information and the corresponding voltage, current, and power values to each sample in the order of record, and at the same time write the characteristic segments of these life samples in the same summary table and mark the relationship with the operating time of the photovoltaic module at that time. If the operating time has reached a certain high value at this time, include the life samples under this segment for analyzing the remaining life. In this way, the actually detected data is split and aggregated to obtain a life probability calculation data set.
[0188] The advantage of the formula is that it can more flexibly reflect the distribution pattern of different components during the actual attenuation process by combining and integrating the life sample values of photovoltaic modules through the shape parameter, scale parameter, and life correlation factor, and obtain a continuous remaining duration estimation result in the sense of probability.
[0189] The steps to obtain the c parameter are as follows:
[0190] c represents the life sample value of a photovoltaic module, which is a comprehensive index statistically obtained from a large number of historical failure cases or attenuation rate tests. To obtain c, it is necessary to track modules of the same type over a long period, record the time from the start of their use until the output power decays to a certain critical point (such as 80% of the initial power), and record the operating time corresponding to each module. After screening these time data and excluding abnormal situations in extreme environments and excessive damage, a batch of life data points can be obtained. Subsequently, through distribution analysis, these life data points can be processed to find a value that can represent the average attenuation condition or medium attenuation level. In a test environment, after uniformly monitoring several batches of modules of the same type for more than three years, the attenuation times of more than 200 modules were extracted and statistically analyzed. If a summary average value of approximately 6000 hours is finally obtained, this value can be defined as c. When other modules are in a similar working environment or have the same rated parameters, this c can be used as an important reference value for subsequent calculations to evaluate the remaining life of new modules.
[0191] F lf The steps to obtain the parameter are as follows:
[0192] F lf F represents the calculated life correlation factor, which has been obtained in the previous linear regression or multiple data analysis. It is the result calculated by decomposing and optimizing through multi-batch test records and corresponding indicators such as voltage, current, and power and then substituting them into the regression model. To obtain this factor, after analyzing the fluctuation characteristics, uncertainty distribution, and multi-layer decomposition characteristics of the photovoltaic module, using the previously constructed regression model, the matrix elements {X 1 , X 2 , …} corresponding to each time period are combined with the regression coefficients {β 0 , β 1 , …} for calculation, and a value obtained is F lf . For example, after on-site monitoring of the voltage and current deviation for about 300 hours and obtaining the life correlation factor of 2.20 through regression fitting, 2.20 can be substituted into this formula to indicate that the module has entered the relatively medium attenuation stage.
[0193] The steps to obtain the θ parameter are as follows:
[0194] θ is the scale parameter of the lifetime data, commonly found in the exponential distribution, Gamma distribution, or Weibull distribution, and is used to determine the stretching or shrinking degree of the overall lifetime distribution curve on the abscissa. To obtain θ, it is necessary to perform distribution fitting on a large amount of historical lifetime data, and then select the distribution model that best represents the overall decay situation, such as the Gamma distribution or Weibull distribution. Then, the scale parameter θ corresponding to this distribution is obtained through maximum likelihood estimation or the least squares method. If θ = 500 hours is obtained in large-scale component testing, it indicates that there will be a relatively significant peak in the failure density of the components around 500 hours in a statistical sense. In actual operation, the decay duration data of multiple batches of components can be collected, first determine which distribution is closer to the actual situation, and then perform parameter estimation under this distribution to finally obtain the specific value of θ.
[0195] The steps to obtain the k parameter are as follows:
[0196] k represents the shape parameter, which is used to determine the shape characteristics of the lifetime distribution and is also an essential parameter in models such as the Gamma distribution or Weibull distribution. If k is relatively small, the distribution shape deviates more from symmetry; if k increases, the distribution is closer to normal or more concentrated. When specifically obtaining it, the same method can be used to first collect enough decay duration data, substitute it into the distribution fitting, and gradually adjust the shape parameter to observe the difference between the data and the fitting curve. When the difference is the smallest, the corresponding k is the optimal solution. For example, in a certain batch of components, after fitting the Gamma distribution through thousands of failure records, if it is calculated that k = 2.2 and it has the highest degree of coincidence with the actual data, then 2.2 can be set as the shape parameter and applied to this formula.
[0197] The steps to obtain the Γ(k + 1) parameter are as follows:
[0198] Γ(k + 1) is the value of the Gamma function at k + 1, which is used to assist in the probability calculation or integral analysis of the lifetime distribution. In the calculation, it can be completed by means of numerical integration methods or ready-made Gamma function tables. To obtain Γ(k + 1), first substitute the determined k value into the Gamma function operation formula. For example, when k = 2.2, Γ(3.2) ≈ 2.423965 can be obtained by looking up the table or using numerical algorithms. It is also possible to write a numerical solution process in the engineering, accumulate the integral results from small step sizes to large step sizes to obtain an approximate value of the Gamma function. If the sample size is large, usually several typical k values are pre-looked up in the table, and then the corresponding Γ(k + 1) is recorded for subsequent calculations.
[0199] Calculation process:
[0200] The first step is to determine the values of each parameter. For example, let c = 6000 hours, F lf= 2.20, θ = 500 hours, k = 2.2, and from the Gamma function table, Γ(k + 1) = Γ(3.2) ≈ 2.423965;
[0201] Step 2, write out the integral expression:
[0202]
[0203] It should be noted here that if the symbol x is just an integration variable, in practice, the integration method or segmented interval often needs to be further specified. Here, since it is an example, the general formula is maintained;
[0204] Step 3, decompose the exponential term and the coefficient:
[0205] 6000 2.2 ≈ 6000 2 × 6000 0.2 ≈ 3.6 × 10 7 × 3.30 ≈ 1.188 × 10 8
[0206]
[0207] 500 3.2 = 500 3 × 500 0.2 ≈ 1.25 × 10 8 × 2.24 ≈ 2.8 × 10 8
[0208]
[0209] Step 4, combine the constant terms:
[0210]
[0211] 0.175 × e -26.4 ≈ 0.175 × 3.53 × 10 -12 ≈ 6.18 × 10 -13
[0212] 2.20 × 6.18 × 10 -13 ≈ 1.36 × 10 -12
[0213] Step 5, the final integral result needs to accumulate x from 0 to ∞. Since the example only shows the operation of the constant part, the complete integration process often needs to consider the influence mechanism in x and the segmented calculation method. If simply measured by this constant order of magnitude, it can be determined that the remaining time part will fall within the accumulation range of the small probability region.
[0214] The results show that when θ = 500 hours, k = 2.2, and the life correlation factor F lf = 2.20, etc., under the condition that the component has a life sample value at the 6000-hour level, the corresponding integral result of the remaining duration is very small, which implies that the component has entered the relatively middle and late decay stage. For further accurate evaluation, it is necessary to discretize the integral or use numerical integration methods in the actual algorithm to check the obtained T rem in the specific numerical range between dozens and hundreds of hours. If it is greater than 300 hours, it means that it can still maintain a certain duration. If it is less than 100 hours, it means that it will face rapid decay soon.
[0215] Based on the remaining duration, first check in the recorded time period index whether the operating environment of the photovoltaic module during this period belongs to relatively stable operating conditions such as a light intensity of about 1000 W / ㎡ and an ambient temperature in the range of 25°C to 30°C. Then select these time periods suitable for comparison with the previous life samples, number them according to the previously statistically obtained remaining duration distribution range. When comparing the observed data of a certain component with the known distribution curve, it is necessary to check each moment with a large deviation one by one, such as the moment when the pressure value is higher than 10 MPa or the moment when the current value is higher than 10 A. If it is found that the proportion of such records exceeds a certain preset threshold, make a special mark after the data of this period and use interpolation for comparative analysis to judge whether this period belongs to the abnormal peak region in the life curve. After counting these marks, the predicted value of the remaining duration within the normal range can be obtained. Compare it with the confidence interval reference value in the previous similar environment, and then delimit the upper and lower boundaries of the remaining life based on the data quantile within the confidence interval. For example, according to the previously trained high confidence interval, an interval can be given within the range of the remaining duration ±50 hours. Finally, in the same list, correspond the confidence interval result with the current monitoring time and the life correlation factor to obtain the life prediction result.
Claims
1. Photovoltaic module life prediction system based on artificial intelligence, characterized by: The system comprises: The data acquisition preprocessing module obtains the voltage value, current value and power value from the output end of the photovoltaic module, normalizes the voltage value, current value and power value, and reorganizes and segments the normalized results in batches according to the sampling sequence to generate a preprocessing data sequence; The uncertainty analysis module performs deviation calculation and fluctuation analysis on the voltage value, current value and power value in each time period in the preprocessed data sequence to obtain parameter fluctuation characteristics; calculates the probability density based on the parameter fluctuation characteristics to generate uncertainty distribution characteristics; The task decomposition optimization module performs hierarchical cascading on the probability density and the deviation mean in the uncertainty distribution feature to obtain a hierarchical feature set; performs weight coefficient assignment and weight coefficient matrix operation on the values of each level in the hierarchical feature set to generate a decomposition optimization matrix; The life prediction and evaluation module performs linear regression on the matrix elements in the decomposition optimization matrix and the operating time value of the photovoltaic component to obtain a life correlation factor; performs probability calculation and confidence interval analysis on the remaining time of the photovoltaic component based on the life correlation factor to generate a life prediction result.
2. The photovoltaic module life prediction system based on artificial intelligence according to claim 1 is characterized in that: The steps of obtaining the preprocessed data sequence are: Collect voltage values, current values, and power values from the output end of the photovoltaic module, and normalize each value to obtain normalized data; Based on the normalized data, the comprehensive index of each sampling point is calculated, and the formula is: Among them, GQ k is the normalized voltage value, GI k is the normalized current value, GP k is the normalized power value, m is the number of sampling points, F s It is a comprehensive index of sampling points; According to the comprehensive index of the sampling points, the normalized data is subjected to time series analysis and batch reorganization, and the data segments are divided according to the set time windows to generate a preprocessed data sequence.
3. The photovoltaic module life prediction system based on artificial intelligence according to claim 1 is characterized in that: The steps for obtaining the parameter fluctuation characteristics are: Performing statistical analysis on the voltage, current and power in each time period in the preprocessed data sequence to obtain the mean value of each time period and generate basic statistical data; Based on the basic statistical data, the comprehensive deviation values of voltage, current and power in each time period are calculated using the formula: Among them, V i ,I i ,P i are the observed values of voltage, current and power respectively, are the means of the corresponding time periods, n is the number of observation points, F dev is the comprehensive deviation value; According to the comprehensive deviation value, the fluctuation characteristics of voltage, current and power within the monitoring period are analyzed to obtain parameter fluctuation characteristics.
4. The photovoltaic module life prediction system based on artificial intelligence according to claim 1 is characterized in that: The steps for obtaining the uncertainty distribution characteristics are: Based on the parameter fluctuation characteristics, the probability density of each parameter is calculated, and the formula is: Where x represents the measured value of voltage, current or power in each time period, μ x represents the corresponding mean deviation, σ x represents the standard deviation, and P(x) represents the probability density; By integrating the probability densities over all time periods, an uncertainty distribution signature is generated.
5. The photovoltaic module life prediction system based on artificial intelligence according to claim 1 is characterized in that: The steps for obtaining the hierarchical feature set are: Based on the uncertainty distribution characteristics, the levels are divided according to the distribution of probability density data, and the levels are further refined according to the range of deviation mean data to obtain a preliminary hierarchical feature structure; Based on the preliminary hierarchical feature structure, the probability density data and deviation mean data of adjacent levels are cascaded, and the transmission relationship of the data between the levels is sorted out to form a hierarchical feature set.
6. The photovoltaic module life prediction system based on artificial intelligence according to claim 1, characterized in that: The steps of obtaining the decomposition optimization matrix are: Based on the stratified feature set, the values of each level are screened, abnormal values and duplicates are removed, and standardized stratified feature data are generated; Based on the standardized hierarchical feature data, an initial weight coefficient is set according to the range of variation of the hierarchical value, and then adjusted in combination with the distribution of the internal values of the hierarchical level to form a weight coefficient allocation result; Based on the weight coefficient allocation result, a weight coefficient matrix is constructed, the weight influence relationship between the levels is analyzed, the optimization adjustment parameters between the levels are determined, and a decomposition optimization matrix is generated.
7. The photovoltaic module life prediction system based on artificial intelligence according to claim 1, characterized in that: The steps for obtaining the lifespan correlation factor are as follows: Based on the decomposition optimization matrix, a linear regression model is established, and the formula is: L rf =β0+β1X1+β2X2+…+β k X k Among them, L rf Represents the predicted life-related factors, X1, X2, …, X k represents the matrix elements in the decomposition optimization matrix, β0,β1,…,β k is the regression coefficient; The life-span correlation factor of the photovoltaic module is calculated based on the linear regression model.
8. The photovoltaic module life prediction system based on artificial intelligence according to claim 1, characterized in that: The steps for obtaining the life prediction result are: Extracting the life-related factors, calling the operation data of the photovoltaic modules, screening the life samples meeting the same operation conditions, and generating a life probability calculation data set; Based on the life probability calculation data set, the remaining time of the photovoltaic module is calculated using the following formula: Among them, T rem represents the calculated remaining time, c represents the sample value of the life of the photovoltaic module, and F lf represents the calculated life correlation factor, θ represents the scale parameter of the life data, k represents the shape parameter, and Γ(k+1) is the Gamma function; Based on the remaining time, a credible interval of the remaining life of the photovoltaic component is determined using confidence interval calculation to generate a life prediction result.
9. The method for predicting the life of a photovoltaic component based on artificial intelligence according to the system for predicting the life of a photovoltaic component based on artificial intelligence according to any one of claims 1 to 8, characterized in that: The following steps are involved: Obtain voltage values, current values, and power values from the output end of the photovoltaic module, normalize the values, and then batch reorganize and segment the normalized data according to the sampling time sequence to generate a preprocessed data sequence; Perform deviation calculation and fluctuation analysis on the voltage value, current value and power value in each time period in the preprocessed data sequence, calculate the probability density based on the obtained fluctuation characteristics, and construct and obtain the uncertainty distribution characteristics; The probability density and the deviation mean in the uncertainty distribution characteristics are hierarchically cascaded, weight coefficients are assigned, and weight coefficient matrix operations are performed to generate a decomposition optimization matrix; Linear regression analysis is performed between the elements in the decomposition optimization matrix and the operating time values of the photovoltaic modules to obtain the life correlation factor associated with the module life; Based on the life correlation factors, probability calculation and confidence interval analysis are performed on the remaining operating time of the photovoltaic modules to generate and obtain the life prediction results.
Citation Information
Cited By
Valve health state monitoring system and method
CN120578914A
A valve health status monitoring system and method
CN120578914B
Multi-specification wire clamp precision and conductivity integrated aging test system
CN121933861A
Photovoltaic device life prediction method based on artificial intelligence
CN122241320A