Shadow photovoltaic user identification method based on photovoltaic output and user load deviation

By constructing a multi-factor photovoltaic output model and a dynamic load reference model, combined with time-frequency domain consistency analysis, identifying shadow photovoltaic users and estimating installed capacity, the problem of insufficient identification accuracy in the existing technology is solved, and higher identification accuracy and accuracy of grid management are achieved.

CN120448997APending Publication Date: 2025-08-08ECONOMIC & TECH RES INST OF HUBEI ELECTRIC POWER COMPANY SGCC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510310190.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The existing shadow photovoltaic user identification method fails to fully utilize the physical characteristics of photovoltaic output and dynamic correlation modeling of user behavior, resulting in insufficient identification accuracy, affecting the accuracy of grid load prediction and scheduling, increasing the risk of transformer and line operation, and being unable to effectively supervise shadow photovoltaic users.

Method used

By establishing a photovoltaic output model that takes into account irradiation intensity, ambient temperature, photovoltaic panel angle, humidity and dust pollution degree, a dynamic load reference model is constructed using a self-organized mapping network, load deviation is calculated and matching degree is evaluated, and shadow photovoltaic users are identified and installed capacity is estimated in combination with time-frequency domain consistency and random forest classification model.

Benefits of technology

It significantly improves the identification accuracy of shadow photovoltaic users, reduces installed capacity estimation errors, improves the accuracy and security of power grid management, reduces regulatory loopholes, and promotes the healthy development of distributed energy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120448997A_ABST
    Figure CN120448997A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of distributed energy management of a power system, and particularly relates to a shadow photovoltaic user identification method based on photovoltaic output and user load deviation, which comprises the following steps of: establishing a photovoltaic output model by considering irradiation intensity, environment temperature, photovoltaic panel angle, humidity and dust pollution degree; constructing a dynamic load reference model by using a self-organizing mapping network, calculating a load deviation of a user load to be identified relative to the dynamic reference model, and evaluating a matching degree of a user load mode and a photovoltaic power generation mode based on the load deviation to obtain a comprehensive matching degree; and calculating the time-frequency domain consistency of the load curve and the irradiation intensity curve of the to-be-identified user, and inputting the comprehensive matching degree and the time-frequency domain consistency into a random forest classification model to predict the category of the to-be-identified user to obtain whether the category of the to-be-identified user is a shadow photovoltaic user or a non-photovoltaic user. According to the method, the shadow photovoltaic user identification precision can be remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of distributed energy management in power systems, and in particular relates to a method for identifying shadow photovoltaic users based on the deviation between photovoltaic output and user load. Background Art

[0002] With the rapid growth of distributed photovoltaic installations, some users have not reported their photovoltaic installations to the grid company. These shadow photovoltaic users have brought the following problems to grid operation and dispatch accuracy:

[0003] 1. The grid load data is inconsistent with the actual electricity demand, which affects the accuracy of load forecasting and scheduling plans.

[0004] 2. Shadow photovoltaic installations may cause reverse power transmission, increasing the operating risks of transformers and lines, especially in areas with high photovoltaic penetration rates.

[0005] 3. Grid companies are unable to effectively supervise shadow photovoltaic users, resulting in economic losses and grid management loopholes.

[0006] Photovoltaic output has significant physical properties, such as its nonlinear relationship with solar irradiance and temperature, and its diurnal fluctuations. Existing methods for identifying shadow photovoltaic users typically rely on simple statistical analysis of user loads. These methods fail to fully exploit the physical characteristics of photovoltaic output and lack the ability to model the dynamic correlation between user behavior and regional environmental conditions, resulting in insufficient identification accuracy. Summary of the Invention

[0007] The purpose of the present invention is to address the above-mentioned problems existing in the prior art and to provide a shadow photovoltaic user identification method based on photovoltaic output and user load deviation, which combines the physical characteristics of photovoltaic output with user load deviation to improve the identification accuracy of undeclared photovoltaic installed users.

[0008] To achieve the above objectives, the technical solutions of the present invention are as follows:

[0009] The present invention provides a method for identifying shadow photovoltaic users based on the deviation between photovoltaic output and user load, the method comprising:

[0010] Considering the radiation intensity, ambient temperature, photovoltaic panel angle, humidity and dust pollution level, a photovoltaic output model is established;

[0011] A dynamic load reference model is constructed using a self-organizing map network, and the load deviation of the user to be identified relative to the dynamic reference model is calculated.

[0012] Based on the load deviation of the user to be identified relative to the dynamic reference model, the matching degree between the user load pattern and the photovoltaic power generation pattern is evaluated to obtain the comprehensive matching degree;

[0013] Calculate the time-frequency domain consistency of the load curve and radiation intensity curve of the user to be identified;

[0014] The comprehensive matching degree and time-frequency domain consistency are input into the random forest classification model to predict the category of the user to be identified, and the category of the user to be identified is obtained as a shadow photovoltaic user or a non-photovoltaic user.

[0015] The photovoltaic output model is:

[0016] P pv (t) = η·A·G(t)·f T (T(t))·f α (α(t))·f D (D(t))·f H (H(t));

[0017] f T (T(t))=1-γ T ·(T(T)-T ref ) 2 ;

[0018] f α (α(t))=cos(α(t)-α opt );

[0019]

[0020] f H (H(t))=1-γ H |H(t)-H opt |;

[0021] In the above formula, P pv (t) is the photovoltaic output, that is, the output of the photovoltaic output model; η is the photovoltaic conversion efficiency of the photovoltaic panel to convert light energy into electrical energy; A is the area of the photovoltaic panel; G(t) is the irradiance; f T (T(t)) is the temperature influence function; T ref is the reference temperature; f α (α(t)) is the photovoltaic panel angle influence function; α opt is the optimal photovoltaic panel installation angle; f D (D(t)) is the dust pollution influence function; f H (H(t)) is the humidity influence function; H opt is the optimal humidity value; γ T is the temperature influence coefficient; γ H is the humidity influence coefficient; Dis the dust attenuation coefficient; α(t) is the photovoltaic panel angle; T(t) is the ambient temperature; H(t) is the humidity; and D(t) is the dust pollution index.

[0022] Based on the photovoltaic output data of the declared photovoltaic users, the parameters of the photovoltaic output model are optimized using an intelligent optimization algorithm; the intelligent optimization algorithm is a particle swarm optimization algorithm, and the parameters of the photovoltaic output model include η, γ T , γ H ,λ D .

[0023] The intelligent optimization algorithm optimizes the parameters of the photovoltaic output model based on the following objective function:

[0024]

[0025] In the above formula, N known is the number of declared PV users; is the photovoltaic output of the i-th declared photovoltaic user predicted by the photovoltaic output model; is the actual photovoltaic output of the i-th declared photovoltaic user.

[0026] The dynamic load reference model is:

[0027]

[0028] In the above formula, Y ref (t) is the reference load curve of non-PV users obtained by dynamic clustering of the self-organizing map network; SOM represents the self-organizing map network operation; These are the normalized load curves of the 1st, 2nd, and Nth non-PV users in the area, respectively; N is the total number of non-PV users.

[0029] The calculation formula of the load deviation of the user load to be identified relative to the dynamic reference model is:

[0030]

[0031] In the above formula, ΔY(t) is the load deviation of the user to be identified relative to the dynamic reference model; Y user (t) is the user load to be identified; are first-order and second-order difference operators respectively.

[0032] Based on the load deviation of the user load to be identified relative to the dynamic reference model, the matching degree between the user load pattern and the photovoltaic power generation pattern is evaluated to obtain the comprehensive matching degree:

[0033] First, construct a multidimensional correlation matrix R of the load deviation of the user to be identified relative to the dynamic reference model and the output of the photovoltaic output modelcorr :

[0034]

[0035] In the above formula, r ij represents the correlation coefficient between the i-th component of the load deviation and the j-th characteristic component of the output of the photovoltaic output model; f i 、g j Feature extraction functions corresponding to the output of load deviation and photovoltaic output model respectively; are the characteristic means of the output of the load deviation and photovoltaic output model, respectively; T represents the total number of time periods; n and m are the number of characteristics of the load deviation characteristics and the output of the photovoltaic output model, respectively;

[0036] Then calculate the comprehensive matching degree according to the following formula:

[0037]

[0038] In the above formula, R match is the comprehensive matching degree; w ij is the feature weight coefficient.

[0039] The time-frequency domain consistency of the load curve and the radiation intensity curve of the user to be identified in the area is calculated according to the following steps:

[0040] First, the time-frequency domain representation of the load curve and radiation intensity curve of the user to be identified in the area is obtained through continuous wavelet transform:

[0041]

[0042] In the above formula, W Y (a, b), W G (a, b) represent the wavelet coefficients of the load curve and irradiation intensity curve of the user to be identified respectively; Y user (t) represents the load of the user to be identified at time t; G(t) represents the radiation intensity at time t; ψ is the wavelet basis function; a is the scale parameter; b is the translation parameter;

[0043] Then the time-frequency domain consistency is calculated according to the following formula:

[0044]

[0045] Φ diff (a,b)=arg(W Y (a,b))-arg(W G (a,b))

[0046]

[0047] In the above formula, C TF(a, b) is the time-frequency domain coherence; Φ diff (a, b) is the phase difference; Φ sync It is the consistency in time-frequency domain.

[0048] The shadow photovoltaic user identification method further includes: after obtaining that the category of the user to be identified is a shadow photovoltaic user, estimating the installed capacity of the shadow photovoltaic user.

[0049] The installed capacity of shadow photovoltaic users is estimated according to the following formula:

[0050] P shadow (t) = f inv (ΔY(t), R match , Φ sync );

[0051]

[0052] In the above formula, C shadow is the installed capacity of shadow photovoltaic users; P shadow (t) is the photovoltaic output of the shadow photovoltaic user; ΔY(t) is the load deviation of the user to be identified relative to the dynamic reference model; Φ sync is the consistency in time-frequency domain; f inv is a reverse derivation function based on a neural network, which is used to map the load deviation to the possible photovoltaic output; G STC is the irradiation intensity under standard conditions; Ω clear Gather for sunny days.

[0053] Compared with the prior art, the present invention has the following beneficial effects:

[0054] 1. The shadow photovoltaic user identification method based on the deviation between photovoltaic output and user load described in the present invention considers the radiation intensity, ambient temperature, photovoltaic panel angle, humidity and dust pollution level, establishes a photovoltaic output model, uses the self-organizing map network to build a dynamic load reference model, calculates the load deviation of the user load to be identified relative to the dynamic reference model, and evaluates the matching degree between the user load pattern and the photovoltaic power generation pattern based on the load deviation of the user load to be identified relative to the dynamic reference model to obtain a comprehensive matching degree; calculates the time-frequency domain consistency of the load curve of the user to be identified and the radiation intensity curve, and inputs the comprehensive matching degree and time-frequency domain consistency into the random forest classification model for treatment. Identify the user category and make a prediction to determine whether the user to be identified is a shadow photovoltaic user or a non-photovoltaic user. The above design first considers the influence of multiple factors to construct a more accurate photovoltaic output model. Secondly, based on a self-organizing map network and time-series fluctuation feature analysis, it accurately extracts load deviations and constructs a matching degree assessment mechanism between the user load pattern and the photovoltaic power generation pattern. Feature spectrum analysis is also introduced, and a time-frequency domain feature consistency detection mechanism is constructed between the photovoltaic output and the user load through wavelet transform. The matching degree assessment mechanism and the time-frequency domain feature consistency detection mechanism are used to identify shadow photovoltaic users, which can significantly improve the identification accuracy of shadow photovoltaic users. Therefore, the present invention can improve the identification accuracy of shadow photovoltaic users. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 This is a flow chart of the shadow photovoltaic user identification method described in the present invention.

[0056] Figure 2 It is the recognition performance of the shadow photovoltaic user recognition method described in the present invention.

[0057] Figure 3 The load curve and estimated photovoltaic output curve of user A357.

[0058] Figure 4 The load time-frequency domain characteristic spectrum and load curve of user B129.

[0059] Figure 5 These are the sensitivity analysis results of the comprehensive matching threshold and the time-frequency domain consistency threshold. DETAILED DESCRIPTION

[0060] The present invention will be further described in detail below with reference to specific embodiments and the accompanying drawings.

[0061] Example:

[0062] See also Figure 1 , a shadow photovoltaic user identification method based on the deviation between photovoltaic output and user load is carried out in the following steps:

[0063] Step 1: Data collection and preprocessing;

[0064] User load data: including the real-time load curve Y of the user to be identified user (t);

[0065] Regional environmental data: including solar radiation intensity G(t), ambient temperature T(t), humidity H(t), photovoltaic panel angle α(t), and dust pollution index D(t);

[0066] Regional photovoltaic output data: including the real-time output curve P of the photovoltaic sites in the region pv (t);

[0067] Data preprocessing: including missing value filling and data normalization; missing value filling refers to the use of spatiotemporal joint prediction methods to process missing values of real-time load data and regional environmental data to ensure the continuity and spatial correlation of time series; data normalization adopts z-score standardization, and its calculation formula is:

[0068]

[0069] In the above formula, Y norm (t) is the normalized load data; μ Y is the mean value of user load; σ Y is the standard deviation of user load;

[0070] Step 2: Considering the radiation intensity, ambient temperature, photovoltaic panel angle, humidity, and dust pollution level, establish the photovoltaic output model shown below:

[0071] P pv (t) = η·A·G(t)·f T (T(t))·f α (α(t))·f D (D(t))·f H (H(t));

[0072] f T (T(t))=1-γ T ·(T(t)-T ref ) 2 ;

[0073] f α (α(t))=cos(α(t)-α opt );

[0074]

[0075] f H (H(t))=1-γ H |H(t)-Hopt |;

[0076] In the above formula, P pv (t) is the photovoltaic output, that is, the output of the photovoltaic output model; η is the photovoltaic conversion efficiency of the photovoltaic panel to convert light energy into electrical energy; A is the area of the photovoltaic panel, the size of the light-receiving area; G(t) is the solar irradiation intensity, that is, the solar radiation power per unit area; f T (T(t)) is the temperature effect function, which is used to describe the efficiency attenuation when the temperature deviates from the optimal temperature. Using a quadratic function instead of the traditional linear model can more accurately reflect the nonlinear effect of temperature change on efficiency; T ref is the reference temperature; f α (α(t)) is the photovoltaic panel angle influence function, which is used to describe the effect of the sunlight incident angle on the power generation efficiency; α opt is the optimal photovoltaic panel installation angle; f D (D(t)) is the dust pollution effect function, which is used to describe the exponential attenuation of light transmittance due to dust accumulation; f H (H(t)) is the humidity influence function, which is used to describe the effect of air humidity on photovoltaic efficiency; H opt is the optimal humidity value; γ T is the temperature influence coefficient; γ H is the humidity influence coefficient; D is the dust attenuation coefficient;

[0077] Specifically, based on the photovoltaic output data of the declared photovoltaic users, the parameters of the photovoltaic output model are optimized using an intelligent optimization algorithm; the intelligent optimization algorithm optimizes the parameters of the photovoltaic output model based on the following objective function:

[0078]

[0079] In the above formula, N known is the number of declared PV users; is the photovoltaic output of the i-th declared photovoltaic user predicted by the photovoltaic output model; is the actual photovoltaic output of the i-th declared photovoltaic user;

[0080] The intelligent optimization algorithm described is the particle swarm optimization algorithm. By simulating swarm intelligence, the particle swarm optimization algorithm searches for the optimal solution in the parameter space. Compared with traditional gradient descent methods, it is less likely to fall into local optimality and is less sensitive to initial values, making it more suitable for the complex nonlinear optimization problems in this scenario. In the particle swarm optimization algorithm, each particle represents a candidate solution in the problem space, and the problem space is searched by following the currently found optimal solution. The mathematical description of the particle swarm optimization algorithm is as follows:

[0081] Assume that in the D-dimensional search space, there is a particle swarm consisting of N particles, and the position of particle i is represented by X i =(x i1 , x i2 ,...,x iD ), speed is expressed as V i =(v i1 , v i2 ,...,v iD ). The historical optimal position of particle i (individual optimal) is recorded as P i =(p i1 , p i2 ,...,p iD ), the historical optimal position of the entire particle swarm (global optimal) is recorded as G = (g1, g2, ..., g D ); In each iteration, the particle updates its velocity and position using the following formula:

[0082]

[0083] In the above formula, ω is the inertia weight, which controls the influence of the previous speed on the current speed; c1 and c2 are acceleration constants, which respectively control the influence of individual optimum and global optimum on particle speed; r1 and r2 are random numbers between [0, 1], which increase the randomness of the search; t represents the number of iterations;

[0084] The steps for optimizing the photovoltaic output model parameters using the particle swarm algorithm are as follows:

[0085] Initialization: Set the particle swarm size N = 50, the maximum number of iterations T = 200; randomly initialize the particle position and velocity;

[0086] Fitness function definition: For each particle, the sum of squares of its prediction errors is calculated as the fitness value:

[0087]

[0088] In the above formula, N known is the number of declared PV users; is the photovoltaic output of the i-th declared photovoltaic user predicted by the photovoltaic output model; is the actual photovoltaic output of the i-th declared photovoltaic user; T data is the current iteration number;

[0089] Iterative optimization: For each particle, calculate its fitness value; update the individual optimal position P i and the global optimal position G; update the speed and position of all particles according to the speed and position update formula;

[0090] Convergence judgment: When the maximum number of iterations is reached or the change in the global optimal fitness after 30 consecutive iterations is less than the preset threshold, the iterative calculation is stopped; the global optimal position G is output as the final optimization parameter;

[0091] Step 3: First, use the self-organizing map network to build a dynamic load reference model as shown below:

[0092]

[0093] In the above formula, Y ref (t) is the reference load curve of non-PV users obtained by dynamic clustering of the self-organizing map network; SOM represents the self-organizing map network operation; are the normalized load curves of the 1st, 2nd, and Nth non-PV users in the region, respectively; N is the total number of non-PV users;

[0094] A self-organizing map (SOM) network is an unsupervised learning neural network that can map high-dimensional input data to a low-dimensional space while preserving the topological structure of the original data. The SOM network consists of an array of competitive layer neurons, each of which has a weight vector with the same dimension as the input space.

[0095] The core learning process of the self-organizing map network includes competitive learning and collaborative learning. Competitive learning: For each input sample, its distance from all neuron weight vectors is calculated, and the neuron with the smallest distance becomes the winner: Where x is the input vector, w i is the weight vector of the i-th neuron; collaborative learning: update the winner neuron and update the surrounding neurons weighted by distance: w i (t+1)=w i (t)+α(t)·h ci (t)·[xw i (t)]; where α(t) is the time-varying learning rate, which gradually decreases during the training process; h ci (t) is the neighborhood function, which represents the influence of the winner neuron c on neuron i;

[0096] A dynamic reference model is constructed using a self-organizing map network (SOM network) to automatically discover and cluster the load patterns of different types of user groups. The specific steps are as follows:

[0097] Data preparation: The user's daily load curve (96 points, 15-minute resolution) is used as input features. Each user's load curve is normalized and a self-organizing map model is built for weekdays and non-workdays.

[0098] Network structure: Input layer: 96 nodes, corresponding to load values at 96 time points per day; Competition layer: 10×10 two-dimensional neuron array, forming 100 prototype vectors; Topology: Hexagonal grid, convenient for representing complex neighborhood relationships;

[0099] Training process: Initialization: Use PCA to reduce the dimensionality of the input data and initialize the weight vector accordingly; Coarse tuning phase (first 1000 iterations): Large initial learning rate (0.8) and neighborhood radius (5.0); Fine tuning phase (last 4000 iterations): Smaller learning rate (linearly decreasing from 0.05 to 0.01) and neighborhood radius (decreasing from 2.0 to 0.5); Batch update: In each round of iteration, all samples are processed before uniformly updating the weights;

[0100] Reference model generation: Map all users to the self-organizing map network and determine the best matching unit (BMU) for each user; count the number of users corresponding to each BMU, screen out representative BMUs (the number of users exceeds the threshold), and use the weight vectors of these representative BMUs as load pattern prototypes for different types of users; for the user to be identified, calculate its real-time load Y user (t) Input the som network and find the load pattern prototype closest to it as the reference model Y ref (t);

[0101] The dynamic reference model constructed through the self-organizing map network can automatically capture the load pattern characteristics of different user groups and more accurately reflect the differences in user behavior than traditional simple average or fixed classification methods;

[0102] Then, the load deviation of the user load to be identified relative to the dynamic reference model is calculated according to the following formula:

[0103]

[0104] In the above formula, ΔY(t) is the load deviation of the real-time load of the user to be identified relative to the dynamic reference model; Y user (t) is the real-time load of the user to be identified; are first-order and second-order difference operators respectively;

[0105] The zero-order difference is used to reflect the absolute power consumption deviation, the first-order difference is used to extract the deviation of the load curve change trend, and the second-order difference is used to extract the deviation of the load curve fluctuation mode, thereby better capturing the two typical characteristics of photovoltaic power generation;

[0106] Step 4: Based on the load deviation of the user to be identified relative to the dynamic reference model, the matching degree between the user load pattern and the photovoltaic power generation pattern is evaluated to obtain the comprehensive matching degree; the specific steps are as follows:

[0107] First, construct a multidimensional correlation matrix R of the load deviation of the user to be identified relative to the dynamic reference model and the output of the photovoltaic output model corr , through the multidimensional correlation matrix R corr Describe the correlation between different characteristics of load deviation and different characteristics of PV output to comprehensively evaluate the association between load deviation and PV output:

[0108]

[0109] In the above formula, r ij represents the correlation coefficient between the i-th component of the load deviation and the j-th characteristic component of the output of the photovoltaic output model; f i 、f j Feature extraction functions corresponding to the output of load deviation and photovoltaic output model respectively; are the characteristic means of the output of the load deviation and photovoltaic output model, respectively; T represents the total number of time periods; n and m are the number of characteristics of the load deviation characteristics and the output of the photovoltaic output model, respectively;

[0110] Then calculate the comprehensive matching degree according to the following formula:

[0111]

[0112] In the above formula, R match is the comprehensive matching degree; w ij is the feature weight coefficient, which can be optimized by genetic algorithm to maximize the classification performance of PV users and non-PV users;

[0113] A genetic algorithm is an optimization algorithm that simulates the biological evolution process. It optimizes the search for solution space through genetic operations such as natural selection, crossover, and mutation. The main operations of a genetic algorithm include: selection operation: selecting excellent individuals according to a certain probability based on their fitness to enter the next generation (such as roulette selection, tournament selection, etc.); crossover operation: two parent individuals exchange some genes to generate new offspring individuals; mutation operation: randomly changing certain genes of an individual to increase population diversity. This paper uses a genetic algorithm to optimize the weight coefficient w in the multidimensional feature correlation matrix. ij , determine the optimal contribution ratio of different features in the comprehensive matching calculation; the specific steps are as follows:

[0114] Problem encoding: Flatten the n×m dimensional weight matrix W into a one-dimensional vector as the individual's genotype; the value range of each gene is [0,1], representing the weight of the corresponding feature;

[0115] Initialization: Set the population size N = 100 and the maximum evolutionary number G = 300; randomly initialize the population and normalize the weight vector of each individual to ensure Σ i,j wij =1;

[0116] Fitness function definition: Evaluate the classification performance of different weight combinations based on cross-validation:

[0117]

[0118] In the above formula, K is the cross-validation fold (in this embodiment, K is 5), F1 s core k is the F1 score of the k-th fold validation. The F1 score takes into account both precision and recall and is a comprehensive measure of classification performance.

[0119] Genetic operations: Selection: A tournament selection strategy is used, randomly selecting 5 individuals from the population each time, and the individual with the highest fitness is selected; Crossover: Arithmetic crossover is used, and the crossover probability is set to 0.8; Mutation: Adaptive Gaussian mutation is used, and the basic mutation probability is set to 0.05, and the step length decreases linearly with the increase of generations; Elite retention: The 5% individuals with the highest fitness in each generation are retained and directly enter the next generation without crossover and mutation;

[0120] Termination condition: reaching the maximum number of evolutionary generations or the improvement of the optimal fitness for 50 consecutive generations is less than 0.0001, and the optimal weight matrix obtained by the final evolution is output;

[0121] Step 5: Calculate the time-frequency domain consistency between the load curve of the user to be identified and the radiation intensity curve. The specific calculation steps are as follows:

[0122] First, the time-frequency domain representation of the load curve and radiation intensity curve of the user to be identified in the area is obtained through continuous wavelet transform:

[0123]

[0124] In the above formula, W Y (a, b), W G (a, b) represent the wavelet coefficients of the load curve and irradiation intensity curve of the user to be identified respectively; Y user (t) represents the load of the user to be identified at time t; G(t) represents the irradiation intensity at time t; ψ is the wavelet basis function, and Morlet wavelet can generally be used. This is a complex-valued wavelet basis function commonly used in time-frequency analysis. It has good detection capabilities for periodic fluctuations and is suitable for identifying the sinusoidal waveform characteristics of photovoltaic output under sunny conditions. It has the best joint resolution in the time-frequency domain, can accurately locate the time-frequency characteristics, and can effectively capture the energy changes and phase characteristics on a 1-2 hour scale; a is the scale parameter; b is the translation parameter;

[0125] The wavelet transform decomposes the time-domain signal into fluctuation components at different scales and time positions to reveal the time-frequency characteristics of the signal. Compared with the traditional Fourier transform, the wavelet transform has multi-resolution analysis capabilities and can simultaneously capture the short-term and long-term variation characteristics of the signal. It is suitable for analyzing the multi-scale fluctuation characteristics of photovoltaic output.

[0126] Then the time-frequency domain consistency is calculated according to the following formula:

[0127]

[0128] Φ diff (a,b)=arg(W Y (a,b))-arg(W G (a,b));

[0129]

[0130] In the above formula, C TF (a, b) is the time-frequency domain coherence, which is used to quantify the similarity of the fluctuations of two signals at a specific time and scale; Φ diff (a, b) is the phase difference, which is used to reflect the time delay or advance of the two signal fluctuations; Φ sync The time-frequency domain consistency is calculated by the coherence weighting method. The coherence weighting method can highlight the contribution of strongly correlated areas. A high value of the time-frequency domain consistency indicates that the user load curve and the irradiation intensity curve are highly synchronized in the time-frequency domain.

[0131] Step 6: Input the comprehensive matching degree and time-frequency domain consistency into the random forest classification model to predict the category of the user to be identified, and obtain whether the category of the user to be identified is a shadow photovoltaic user or a non-photovoltaic user. To further improve the recognition accuracy, auxiliary features can be used together with the comprehensive matching degree and time-frequency domain consistency to input the random forest classification model for category prediction, specifically:

[0132] L user = RandomForest(R match , Φ sync , F aux );

[0133] In the above formula, L user is the user category label; F aux It is an auxiliary feature, which can be an environmental condition feature;

[0134] The random forest classification model can integrate multiple features to make the final identification decision. By integrating the results of multiple decision trees, it balances recognition accuracy and robustness. When the following conditions are met at the same time, it is identified as a shadow photovoltaic user:

[0135] Condition 1: The comprehensive matching degree is greater than the corresponding threshold: R match >R thresh , R thresh The threshold is set to 0.75;

[0136] Condition 2: Time-frequency domain consistency is greater than its corresponding threshold: Φ sync >Φ thresh , Φ thresh The threshold is set to 0.68;

[0137] Condition 3: The output probability of the random forest classification model is greater than its corresponding threshold P(L user =1)>P thresh , P thresh The threshold is set to 0.8;

[0138] Step 7: After the user category to be identified is a shadow photovoltaic user, the installed capacity of the shadow photovoltaic user is estimated according to the following formula;

[0139] P shadow (t) = f inv (ΔY(t), R match , Φs ync );

[0140]

[0141] In the above formula, C shadow is the installed capacity of shadow photovoltaic users; P shadow (t) is the photovoltaic output of the shadow photovoltaic user; ΔY(t) is the deviation of the load of the user to be identified relative to the dynamic reference model; Φ sync is the consistency in time-frequency domain; f inv is a reverse derivation function based on a neural network, which is used to map the load deviation to the possible photovoltaic output; G STC The irradiation intensity under standard conditions is 1000W / m 2 ;Ω clear Gather for sunny days;

[0142] When estimating installed capacity, a neural network is used to model the complex load deviation-PV output relationship, so that the PV output can be more accurately inferred from the load deviation. Only the best data points under sunny conditions are selected for capacity estimation to reduce environmental interference. Calculation is based on the standardized irradiation intensity ratio to eliminate the impact of irradiation intensity changes.

[0143] The beneficial effects of the present invention are:

[0144] 1. Multi-factor photovoltaic output model: Breaking through the limitations of traditional single- or dual-factor models, this model comprehensively considers the impact of multiple factors on photovoltaic power generation, including radiation intensity, ambient temperature, photovoltaic panel angle, dust pollution level, and humidity. By introducing a nonlinear temperature influence function and an exponential dust attenuation function, it more accurately models the physical characteristics of photovoltaics in real environments, ultimately improving the prediction accuracy of the photovoltaic output model by approximately 35%.

[0145] 2. Load dynamic deviation feature extraction based on self-organizing map networks: A dynamic reference model is constructed using a self-organizing map network to discover the load patterns of different user groups. By extracting multi-scale load deviation features, the trend changes and fluctuation characteristics of the load curve are captured, solving the problem of identification blind spots caused by differences in user behavior.

[0146] 3. Multi-dimensional feature correlation matrix and adaptive weight optimization: This method breaks through the limitations of traditional single correlation coefficient calculation and constructs a complete feature correlation matrix to comprehensively evaluate the relationship between load and photovoltaics. It also optimizes feature weights through genetic algorithms to achieve adaptive matching for different user types, improving final recognition accuracy by approximately 15%.

[0147] 4. Time-frequency domain characteristic spectrum analysis based on wavelet transform: Wavelet transform and time-frequency domain coherence analysis technology are introduced to construct photovoltaic fluctuation characteristic spectrum, revealing the correlation pattern between load and irradiation on different time scales; through the coherence-weighted phase synchronization index, load fluctuations caused by photovoltaic power generation are effectively distinguished from fluctuations caused by other factors, significantly improving the identification accuracy of borderline cases.

[0148] 5. Capacity estimation method based on reverse deduction of neural network: Based on neural network modeling of the complex load deviation-PV output relationship, accurate reverse deduction from load deviation to PV installed capacity is achieved, reducing the installed capacity estimation error by approximately 50%, providing more accurate technical support for grid planning and user supervision.

[0149] 6. Collaborative optimization framework for multi-algorithm fusion: In the first stage, the particle swarm optimization algorithm is used to determine the optimal parameters of the photovoltaic output model. In the second stage, a dynamic reference model of user load is constructed based on the self-organizing map network. In the third stage, the genetic algorithm is used to optimize the weight coefficients of the multi-dimensional feature association matrix. This not only fully utilizes the advantages of each algorithm, the particle swarm optimization algorithm has high efficiency in searching in continuous parameter space, the genetic algorithm performs well in combinatorial optimization problems, and the self-organizing map network has unique capabilities in unsupervised pattern discovery, but also uses the optimization results of each algorithm as the input for the next stage, forming a collaborative optimization framework for multi-algorithm fusion. The algorithms complement and synergize with each other, jointly improving the performance and stability of the overall method.

[0150] The method described in this invention can promptly identify undeclared distributed photovoltaic installations, prevent the risk of reverse power transmission, and ensure the safe and stable operation of distribution networks, especially in areas with high photovoltaic penetration. By understanding the actual photovoltaic installation situation, it facilitates standardized management, reduces regulatory loopholes, provides accurate data support for energy planning and policy formulation, improves the accuracy of scheduling plans, optimizes power resource allocation, promotes the rational layout and efficient use of renewable energy, and promotes the healthy and orderly development of distributed energy.

[0151] Performance Verification:

[0152] This study used 1,000 low-voltage users within the jurisdiction of a provincial power grid company, collecting 15-minute resolution load data from June 1, 2023, to July 31, 2023. Meteorological data (including irradiation intensity, temperature, and humidity) and output data from 200 registered photovoltaic users were also collected. The calculation process was divided into two phases: offline training and online recognition. The offline training phase mainly includes:

[0153] Photovoltaic output model parameter optimization: Based on the output data of 200 declared photovoltaic users, the particle swarm optimization algorithm is used to determine the optimal parameters of the photovoltaic output model: η is 0.175, γ T is 0.0012, γ H is 0.0008, λ D is 0.15; T ref is 25℃; opt 32°; H opt is 45;

[0154] Self-organizing map network training: network size: 10×10; learning rate: initial value 0.8, decay rate 0.01; topology: hexagonal grid; neighborhood function: Gaussian function;

[0155] Feature weight optimization: Determine the optimal weight of multi-dimensional features through genetic algorithms;

[0156] Random forest classification model training: training a classification model based on a labeled dataset; parameter settings: number of decision trees: 200; maximum depth: 15; feature sampling ratio: 0.7; sample sampling ratio: 0.8;

[0157] In the test environment (Intel Xeon E5-2680 v4 CPU, 64GB memory), the offline training phase took about 3.5 hours in total, mainly focusing on PV output model parameter optimization (1.2 hours) and feature weight optimization (1.5 hours).

[0158] The calculation process for a single user in the online recognition phase includes:

[0159] Data preprocessing: about 0.2 seconds;

[0160] Load dynamic deviation feature extraction: takes about 0.5 seconds;

[0161] Feature correlation matching: takes about 0.3 seconds;

[0162] Photovoltaic fluctuation characteristic spectrum analysis: takes about 1.2 seconds;

[0163] Shadow PV user identification and capacity estimation: takes about 0.1 seconds;

[0164] The complete identification process for a single user takes an average of 2.3 seconds. For a regional network containing 10,000 users, the entire area scan can be completed within 1 hour through parallel computing, which can meet the real-time requirements of daily power grid supervision.

[0165] (1) The recognition performance of the shadow photovoltaic user identification method of the present invention is evaluated using a test set (including 800 users, of which 160 are shadow photovoltaic users). The shadow photovoltaic user identification method of the present invention is compared with two other identification methods. One is a traditional correlation analysis method, which uses the traditional Pearson correlation analysis method to analyze the correlation between the load of the user to be identified and the load of the user containing photovoltaic output. The other is a simple load statistical feature method, which uses statistical features such as load peak-to-valley difference rate, load rate, and midday climbing rate to classify them with the statistical features of the user containing photovoltaic output. The results are as follows: Figure 2 As shown. Figure 2 As can be seen, the shadow photovoltaic user identification method described in this invention can more accurately identify shadow photovoltaic users. The installed capacity of shadow photovoltaic users was estimated, and the average relative error of the estimated results was calculated to be 6.3%, the maximum relative error was 12.5%, the minimum relative error was 1.8%, and the standard deviation was 3.2%.

[0166] (2) User A357 in the test set is selected as the analysis object. Its load fluctuation characteristics are that the load during the daytime is significantly lower than that at night and is negatively correlated with the radiation intensity. The calculated comprehensive matching degree R match 0.86, time-frequency domain consistency Φ sync is 0.83, and the random forest classification model output probability P(L user =1) is 0.97, so it is identified as a shadow photovoltaic user. The estimated installed capacity is 5.8kW, and the actual installed capacity is later obtained to be 6.2kW, with a relative error of 6.5%. The load curve of user A357 and the estimated photovoltaic output are as follows Figure 3As shown, the load curve exhibits a significant dip between 09:00 and 15:00, coinciding with the peak period of photovoltaic power generation. The load fluctuation pattern exhibits a high anti-correlation with irradiance intensity variations, suggesting that photovoltaic power generation is being consumed. During short-term fluctuations in irradiance (such as cloud cover), the load curve also exhibits corresponding inverse fluctuations. These characteristics are typical of shadow photovoltaic users, demonstrating that this method can successfully distinguish them from ordinary load fluctuations. Furthermore, the multidimensional feature correlation matrix shows that the first-order difference of user A357's load correlates with photovoltaic output at a rate of -0.92, the second-order difference correlates with photovoltaic output at a rate of -0.87, and the raw load correlates with photovoltaic output at a rate of -0.83, demonstrating high correlation across multiple scales. Wavelet time-frequency domain analysis reveals that on a 1-4 hour scale (corresponding to the short-term PV fluctuation period), the phase difference between load deviation and irradiance variation is close to 180°, with a coherence as high as 0.89, further confirming the possibility of photovoltaic power generation.

[0167] (3) User B129 in the test set is selected as the analysis object. The load time-frequency domain characteristic spectrum and load curve of user B129 are as follows: Figure 4 As shown, the calculated comprehensive matching degree R match 0.73, time-frequency domain consistency Φ sync is 0.70, the random forest classification model output probability P(L user =1) is 0.83, so it is identified as a shadow photovoltaic user. Since user B129 is located near the discrimination boundary, traditional identification methods are prone to missed judgments. The present invention deeply analyzes the time-frequency domain feature map and finds that there is an obvious radiation-load deviation anti-correlation feature during the noon period, especially on the 1-2 hour scale. The phase difference between the load deviation and the radiation change is close to 180°, and the coherence reaches 0.82. This subtle but definite time-frequency domain feature, combined with the comprehensive evaluation of the multi-dimensional feature correlation matrix, finally identifies the user as a shadow photovoltaic user. Later, it was confirmed that the user did install a photovoltaic system, and the actual installed capacity was 3.8kW. The installed capacity estimated by this method is 3.5kW, with a relative error of 7.9%. The above results show that this method still has a strong recognition ability for users at the discrimination boundary.

[0168] (4) To verify the influence of key parameters on recognition performance, the comprehensive matching threshold R thresh and the time-frequency domain consistency threshold Φ thresh Sensitivity analysis was performed. Figure 5 As shown, with R thresh As R increases, the precision increases and the recall decreases; thresh When increasing from 0.65 to 0.85, the precision improves from 91.2% to 99.3%, but the recall decreases from 99.3% to 92.1%.thresh As Φ increases, the precision increases and the recall decreases; when Φ thresh When the R value increases from 0.60 to 0.80, the precision increases from 93.5% to 99.1%, but the recall decreases from 98.7% to 92.5%. thresh =0.75 and Φ thresh When ∑ = 0.68, the F1 score reaches a maximum of 97.4%, achieving the optimal balance between precision and recall. This sensitivity analysis has important guiding significance for practical applications. For example, in scenarios where high accuracy is sought (such as precision monitoring), the threshold can be increased, while in scenarios where high coverage is sought (such as primary screening), the threshold can be appropriately lowered.

[0169] When the number of decision trees in the random forest classification model increased from 50 to 200, the accuracy improved from 95.3% to 97.9%. However, when the number was further increased to 300, the accuracy only increased to 98.1%, a less significant improvement. This indicates that 200 decision trees is an ideal balance point, ensuring performance while controlling computational complexity.

Claims

1. A method for identifying shadow photovoltaic users based on the deviation between photovoltaic output and user load, characterized by: The shadow photovoltaic user identification method includes: Considering the radiation intensity, ambient temperature, photovoltaic panel angle, humidity and dust pollution level, a photovoltaic output model is established; A dynamic load reference model is constructed using a self-organizing map network, and the load deviation of the user to be identified relative to the dynamic reference model is calculated. Based on the load deviation of the user to be identified relative to the dynamic reference model, the matching degree between the user load pattern and the photovoltaic power generation pattern is evaluated to obtain the comprehensive matching degree; Calculate the time-frequency domain consistency of the load curve and radiation intensity curve of the user to be identified; The comprehensive matching degree and time-frequency domain consistency are input into the random forest classification model to predict the category of the user to be identified, and the category of the user to be identified is obtained as a shadow photovoltaic user or a non-photovoltaic user.

2. The method for identifying shadow photovoltaic users based on the deviation between photovoltaic output and user load according to claim 1, characterized in that: The photovoltaic output model is: P pv (t)=η·A·G(t)·f T (T(t))·f α (α(t))·f D (D(t))·f H (H(t)); f T (T(t))=1-γ T ·(T(t)-T ref ) 2 ; f α (α(t))=cos(α(t)-α opt ); f H (H(t))=1-γ H ·|H(t)-H opt |; In the above formula, P pv (t) is the photovoltaic output, that is, the output of the photovoltaic output model; η is the photovoltaic conversion efficiency of the photovoltaic panel to convert light energy into electrical energy; A is the area of the photovoltaic panel; G(t) is the irradiance; f T (T(t)) is the temperature influence function; T ref is the reference temperature; f α (α(t)) is the photovoltaic panel angle influence function; α opt is the optimal photovoltaic panel installation angle; f D (D(t)) is the dust pollution influence function; f H (H(t)) is the humidity influence function; H opt is the optimal humidity value; γ T is the temperature influence coefficient; γ H is the humidity influence coefficient; D is the dust attenuation coefficient; α(t) is the photovoltaic panel angle; T(t) is the ambient temperature; H(t) is the humidity; and D(t) is the dust pollution index.

3. The method for identifying shadow photovoltaic users based on the deviation between photovoltaic output and user load according to claim 2, characterized in that: Based on the photovoltaic output data of the declared photovoltaic users, the parameters of the photovoltaic output model are optimized using an intelligent optimization algorithm; the intelligent optimization algorithm is a particle swarm optimization algorithm, and the parameters of the photovoltaic output model include η, γ T , γ H ,λ D .

4. The method for identifying shadow photovoltaic users based on the deviation between photovoltaic output and user load according to claim 3 is characterized by: The intelligent optimization algorithm optimizes the parameters of the photovoltaic output model based on the following objective function: In the above formula, N known is the number of declared PV users; is the photovoltaic output of the i-th declared photovoltaic user predicted by the photovoltaic output model; is the actual photovoltaic output of the i-th declared photovoltaic user.

5. The method for identifying shadow photovoltaic users based on the deviation between photovoltaic output and user load according to any one of claims 1 to 4, characterized in that: The dynamic load reference model is: In the above formula, Y ref (t) is the reference load curve of non-PV users obtained by dynamic clustering of the self-organizing map network; SOM represents the self-organizing map network operation; These are the normalized load curves of the 1st, 2nd, and Nth non-PV users in the region; n is the total number of non-PV users.

6. The method for identifying shadow photovoltaic users based on the deviation between photovoltaic output and user load according to claim 5, characterized in that: The calculation formula of the load deviation of the user load to be identified relative to the dynamic reference model is: In the above formula, ΔY(t) is the load deviation of the user to be identified relative to the dynamic reference model; Y user (t) is the user load to be identified; are first-order and second-order difference operators respectively.

7. The method for identifying shadow photovoltaic users based on the deviation between photovoltaic output and user load according to claim 5, characterized in that: Based on the load deviation of the user load to be identified relative to the dynamic reference model, the matching degree between the user load pattern and the photovoltaic power generation pattern is evaluated to obtain the comprehensive matching degree: First, construct a multidimensional correlation matrix R of the load deviation of the user to be identified relative to the dynamic reference model and the output of the photovoltaic output model corr : In the above formula, r ij represents the correlation coefficient between the i-th component of the load deviation and the j-th characteristic component of the output of the photovoltaic output model; f i 、g j are the feature extraction functions of load deviation and photovoltaic output model output respectively; are the mean values of the characteristics of the load deviation and the photovoltaic output model output, respectively; T represents the total number of time periods; n and m represent the number of characteristics of the load deviation characteristics and the photovoltaic output model output, respectively; Then calculate the comprehensive matching degree according to the following formula: In the above formula, R match is the comprehensive matching degree; w ij is the feature weight coefficient.

8. The method for identifying shadow photovoltaic users based on the deviation between photovoltaic output and user load according to any one of claims 1 to 3, characterized in that: The time-frequency domain consistency of the load curve and the radiation intensity curve of the user to be identified in the area is calculated according to the following steps: First, the time-frequency domain representation of the load curve and radiation intensity curve of the user to be identified in the area is obtained through continuous wavelet transform: In the above formula, W Y (a, b), W G (a, b) represent the wavelet coefficients of the load curve and irradiation intensity curve of the user to be identified respectively; Y user (t) represents the load of the user to be identified at time t; G(t) represents the radiation intensity at time t; ψ is the wavelet basis function; a is the scale parameter; b is the translation parameter; Then the time-frequency domain consistency is calculated according to the following formula: Φ diff (a,b)6(W Y (a,b))-arg(W). G (a,b)) In the above formula, C TF (a, b) is the time-frequency domain coherence; Φ diff (a, b) is the phase difference; Φ sync It is the time-frequency domain consistency.

9. The method for identifying shadow photovoltaic users based on the deviation between photovoltaic output and user load according to claim 1, characterized in that: The shadow photovoltaic user identification method further includes: after obtaining that the category of the user to be identified is a shadow photovoltaic user, estimating the installed capacity of the shadow photovoltaic user.

10. The method for identifying shadow photovoltaic users based on the deviation between photovoltaic output and user load according to claim 9, characterized in that: The installed capacity of shadow photovoltaic users is estimated according to the following formula: P shadow (t)=f inv (ΔY(t),R match ,Φ sync ) In the above formula, C shadow is the installed capacity of shadow photovoltaic users; P shadow (t) is the photovoltaic output of the shadow photovoltaic user; ΔY(t) is the load deviation of the user to be identified relative to the dynamic reference model; Φ sync is the consistency in time-frequency domain; f inv is a reverse derivation function based on a neural network, which is used to map the load deviation to the possible photovoltaic output; G STC is the irradiation intensity under standard conditions; Ω clear Gather for sunny days.