Distributed photovoltaic district user data anomaly detection method and system

By employing K-means clustering and Spearman correlation coefficient methods, the problems of low efficiency and accuracy in anomaly detection in distributed photovoltaic low-voltage distribution areas were solved, achieving efficient and accurate anomaly identification and repair.

CN117312896BActive Publication Date: 2026-01-02HEFEI UNIV OF TECH +2
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202311301983.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-08
Publication Date
2026-01-02
Estimated Expiration
2043-10-08

AI Technical Summary

Technical Problem

In existing technologies, the detection of abnormal data from different types of users in low-voltage distribution areas containing distributed photovoltaic power generation suffers from problems such as low data measurement efficiency, low prediction accuracy, and low identification accuracy.

Method used

The method employs K-means clustering and Spearman correlation coefficient to cluster historical measurement data, forming cluster center curves and feasible region matrices. The Spearman correlation coefficient is then used to determine whether there are anomalies in the data of the day to be measured, thereby eliminating interference and improving detection accuracy.

Benefits of technology

It effectively improves the accuracy and efficiency of abnormal data detection, can accurately identify anomalies at multiple measurement points, repair abnormal data, and detect line loss anomalies in the transformer area.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117312896B_ABST
    Figure CN117312896B_ABST
Patent Text Reader

Abstract

The application provides a distributed photovoltaic area user data anomaly detection method and system, which comprises the following steps: taking ordinary users and photovoltaic users in a distributed photovoltaic area as research objects, collecting historical measurement data in a monitoring and data acquisition system; based on the historical measurement data of the users, performing K-means clustering on the user historical load, selecting the number of clusters according to the DBI index, and forming the cluster center curve and the feasible region matrix of each type of power consumption behavior; according to the Spearman correlation coefficient, dividing the to-be-tested daily data into a category with high correlation, comparing the to-be-tested data collection point with its feasible region, and judging whether the user data is abnormal. The application solves the technical problems of low data measurement efficiency, low prediction accuracy and low recognition accuracy in the anomaly data detection of different types of users in the distributed photovoltaic low-voltage distribution area.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of circuit system intelligent diagnosis and optimization, and particularly relates to a user data anomaly detection method and system containing a distributed photovoltaic substation. BACKGROUND

[0002] User measurement data, as one of the most important basic data of the power system, has a direct impact on the reliability of the results of state estimation, load forecasting, optimal scheduling, etc., and has a direct impact on the decision of the power system. Due to objective reasons in the actual transmission process and the measurement device itself and subjective reasons, abnormal data may exist in the measurement data, and continuous abnormal data is an important reason for non-technical loss of the distribution network. In addition, under the background of "carbon peak" and "carbon neutral", the development momentum of distributed photovoltaic access to the distribution network will become stronger, and the human abnormal photovoltaic data caused by photovoltaic subsidies will increase, making the form of user data anomaly more complex and diverse, and more difficult to identify. Therefore, it is necessary to study the user abnormal data detection of the distributed photovoltaic substation.

[0003] At present, many methods for identifying abnormal data have been proposed by domestic and foreign experts and scholars. The existing user data anomaly detection is mainly based on two categories of traditional statistical methods and artificial intelligence technology methods. The method of detecting user abnormal measurement data based on statistics is to distinguish normal data from abnormal data by analyzing the specific statistical indicators of measurement data. This method assumes the probability distribution model of measurement data, and through inconsistency test, the measurement data that deviates from the probability distribution curve is regarded as abnormal data. However, this method has limitations such as low implementation efficiency when there are many measurement data, poor diagnosis effect on abnormal measurement data in some cases, etc. The method of detecting abnormal data based on artificial intelligence technology can extract the key features of measurement data (unknown system) through learning or training, and then make correct judgment on the effectiveness of measurement data. Among them, the machine learning abnormal power consumption mode detection model is suitable for the case where the power user data set lacks training samples. For example, the existing patent application for invention with publication number CN114662584A, "A method for detecting user electricity stealing and electricity leakage in a large range based on a time convolution network", the existing method includes: data set collection and preprocessing; collecting the electricity consumption data of electricity consumption users through intelligent electric meters, and dividing the electricity consumption users into normal users and abnormal users; normalizing the electricity consumption data of electricity consumption users through a time convolution neural network, and normalizing the data set between (-1, 1); dividing the electricity consumption data into training set and test set according to the ratio of 8:2. 2. Build a time convolution network model for user electricity consumption anomaly based on the preprocessed data set; evaluate the time convolution network model. And the existing patent application for invention with publication number CN111817299A, "Intelligent identification method for abnormal causes of distribution area line loss rate based on fuzzy reasoning", the existing method includes: obtaining the line loss rate data of the distribution area; obtaining the abnormal period of the line loss rate of the distribution area; establishing the membership function in the fuzzy expert library according to the historical data of the line loss rate of the distribution area; using the membership function to judge the line loss rate of the abnormal period; analyzing the judgment result: if the abnormal line loss rate is negative line loss rate, go to the next step; if the abnormal line loss rate is high line loss rate, go to the next step; if the line loss rate is normal, output the normal line loss rate and report an error; classify the negative line loss rate and judge the abnormal reason, go to the next step; determine the abnormal reason by the line loss rate estimation formula and the Pearson coefficient, go to the next step; analyze and output the obtained reasons. However, there are many factors affecting the output of photovoltaic power in the foregoing prior art, and the prediction model is difficult to consider comprehensively, resulting in low prediction accuracy and low recognition accuracy.

[0004] In summary, in the existing technology of low-voltage distribution area with distributed photovoltaic power, there are technical problems of low data measurement efficiency, low prediction accuracy and low recognition accuracy in abnormal data detection of different types of users. SUMMARY

[0005] The technical problems to be solved by the present application are: how to solve the technical problems of low data measurement efficiency, low prediction accuracy and low recognition accuracy in abnormal data detection of different types of users in the existing technology containing distributed photovoltaic low-voltage distribution area.

[0006] The present application solves the above technical problems by adopting the following technical solutions: a user data abnormality detection method for a distributed photovoltaic area includes:

[0007] S1, for the ordinary users and photovoltaic users in the distributed photovoltaic area, using the pre-set monitoring and data acquisition system, collecting historical measurement data;

[0008] S2, clustering the user historical load in the historical measurement data by K-means, selecting the number of clusters according to the pre-set DBI index, and forming the cluster center curve and the feasible region matrix of each type of power consumption behavior;

[0009] S3, according to the to-be-measured day data and the cluster center curve, calculating the Spearman correlation coefficient, dividing the to-be-measured day data into a strong correlation category, and comparing the collection point corresponding to the to-be-measured day data with the feasible region according to the feasible region matrix to determine whether the to-be-measured day data of the user is abnormal.

[0010] The present application selects the Spearman correlation coefficient to exclude the interference of abnormal data, performs correlation analysis on the data of the to-be-measured day, divides the to-be-measured day data into the actual belonging feasible region, and finally completes the abnormality detection of the users in the distributed photovoltaic area. Since the calculation of each measurement point at each time is independent of each other, the abnormality of multiple measurement points can be identified.

[0011] The present application has high abnormal data detection capability for ordinary users and photovoltaic users, which helps to further repair abnormal measurement data and detect line loss anomalies in the area. It can effectively solve the problem of abnormal data detection of different types of users in the distributed photovoltaic low-voltage distribution area.

[0012] In a more specific technical solution, in step S1, not less than 2 user measurement points and sampling points are selected, and the historical measurement data of the pre-set continuous time interval is recorded.

[0013] In a more specific technical solution, the historical measurement data is expressed by the following logic:

[0014]

[0015] In the formula, the row vector {x i1 ,x i2 ,…,…x im}(1≤i≤n) represents the measurement data of a day, and the vertical vector {x 1j ,x 2j,…,…x nj}(1≤j≤m) represents the sampling data of n days at the same time.

[0016] In a more specific technical solution, step S2 comprises:

[0017] S21, input the number of clustering clusters k, take k data sample points from the historical measurement data as initial clustering centers, and calculate the distance D of the remaining data sample points to the clustering centers using the following logic: ij , compare the distance of the measurement data to each clustering center D ij The smaller the distance, the closer the measurement data is to the clustering center, and the corresponding sample point is assigned to the adjacent clustering center.

[0018]

[0019] In the formula, x i represents the i-th remaining measurement data (1≤i≤n), x j represents the j-th clustering center (1≤j≤k).

[0020] S22, recalculate the average value of the distance of all data sample points in each cluster to the clustering center as the new clustering center, and redivide the data sample points to the new adjacent clustering center in the new clustering center;

[0021] S23, determine whether the adjacent distance center has changed compared with the new adjacent distance center;

[0022] S24, if yes, store and output the clustering result, and execute the subsequent step S26, otherwise, replace the initial clustering center with the new average value, and jump to step S22;

[0023] S25, (14) calculate the DBI index to measure the ratio between the intra-cluster distance and the inter-cluster distance, select the number of data sample points with the smallest ratio as the applicable clustering number, and obtain the user historical data clustering;

[0024] S26, take the sampling data at a specific time from each user historical data clustering for analysis to obtain the normal sampling data threshold interval at the specific time:

[0025] S27, execute step S26 for all sampling times to obtain the feasible region matrix, according to which the abnormal measurement data is identified, and the clustering center curve is drawn according to the user historical data clustering.

[0026] In a more specific technical solution, in step S25, the DBI index is calculated using the following logic:

[0027]

[0028] wherein c i is the center of the i-th cluster, σ i is the average distance of all points in the i-th cluster to the center, d(c i ,c j ) is the distance between the cluster center c i and c j .

[0029] The present application is based on the measurement data of the data acquisition and monitoring control system, and different factors are selected for k-means clustering for photovoltaic users and ordinary users, and the best cluster number is selected according to the DBI index to divide different power consumption behaviors of the same user more accurately and improve the accuracy of the feasible region.

[0030] In a more specific technical solution, in step S26, the normal sampling data threshold interval at the jth moment is obtained by using the following logic:

[0031] z j =[z jmax ,z jmin ] (4)

[0032] wherein z jmax is the maximum value at the jth moment; and z jmin is the minimum value at the jth moment.

[0033] In a more specific technical solution, in step S27, the feasible region matrix is obtained by processing by using the following logic:

[0034]

[0035] In a more specific technical solution, step S3 includes:

[0036] S31, the to-be-measured daily data, the clustering center curve are taken as a first variable X and a second variable Y, and sorting operation is performed according to the size of the first variable X and the second variable Y, and the average value is ranked;

[0037] S32, the Pearson correlation coefficient between the ranks of the first variable X and the second variable Y is calculated by using formula (17), so as to analyze the correlation between each clustering center and the to-be-measured daily data, and a to-be-measured data correlation coefficient value table is obtained.

[0038] S33, according to the Pearson correlation coefficient, the to-be-measured curve is substituted into the corresponding feasible region curve, so as to judge the part exceeding the boundary of the feasible region as abnormal data.

[0039] In a more specific technical solution, in step S32, the Pearson correlation coefficient between the ranks of the first variable X and the second variable Y is calculated by using the following logic:

[0040]

[0041] In the formula, d represents the ranking difference of X and Y, and n represents the sample quantity.

[0042] The present application utilizes the Spearman correlation coefficient when calculating the correlation of the to-be-tested data and the cluster center, and therefore does not need to care about how the data varies and what distribution it conforms to, but only needs to care about the position of the corresponding numerical value of each variable. According to this method, the to-be-tested daily data is divided into the corresponding categories, and it is judged whether the data of the user is abnormal, thereby effectively improving the abnormal detection precision.

[0043] In a more specific technical solution, the user data anomaly detection system for a distributed photovoltaic area comprises:

[0044] A historical data acquisition module is used to acquire historical measurement data by using a preset monitoring and data acquisition system for ordinary users and photovoltaic users in the distributed photovoltaic area.

[0045] A cluster center and feasible region acquisition module is used to perform K-means clustering on the historical load of the users in the historical measurement data, select the number of clusters according to the preset DBI index, and form the cluster center curve and the feasible region matrix of the power consumption behavior of each category.

[0046] An anomaly determination module is used to calculate the Spearman correlation coefficient according to the to-be-tested daily data and the cluster center curve, divide the to-be-tested daily data into the strong correlation category, and compare the collection point corresponding to the to-be-tested daily data with the feasible region according to the feasible region matrix, so as to judge whether the to-be-tested daily data of the user is abnormal.

[0047] Compared with the prior art, the present application has the following advantages: the present application selects the Spearman correlation coefficient, excludes the interference of abnormal data, performs correlation analysis on the to-be-tested daily data, divides the to-be-tested daily data into the actual feasible region, finally completes the anomaly detection of the users in the distributed photovoltaic area, and can complete the anomaly identification of multiple measurement points since the calculation of each measurement point at each time is independent.

[0048] The present application has high abnormal data detection capability for ordinary users and photovoltaic users, is helpful to further repair the abnormal measurement data, and detects the line loss anomaly of the area.

[0049] The application is based on the measurement data of the data acquisition and monitoring control system, and different factors are selected for k-means clustering for photovoltaic users and ordinary users, and the best cluster number is selected according to the DBI index to divide different power consumption behaviors of the same user more accurately and improve the accuracy of the feasible region.

[0050] The application uses the Spearman correlation coefficient when calculating the correlation of the to-be-measured data and the cluster center, so it does not need to care about how the data changes and what distribution it conforms to, and only needs to care about the position of the corresponding value of each variable. According to this method, the to-be-measured daily data is divided into the corresponding category, whether the user data is abnormal is judged, and the abnormal detection accuracy is effectively improved.

[0051] The application solves the technical problems of low data measurement efficiency, low prediction accuracy and low recognition accuracy of abnormal data detection of different types of users in the existing low-voltage distribution area containing distributed photovoltaic. BRIEF DESCRIPTION OF DRAWINGS

[0052] Figure 1 It is a data flow processing schematic diagram of the user data abnormality detection method of the distributed photovoltaic area in embodiment 1 of the application;

[0053] Figure 2 It is a basic step schematic diagram of the user data abnormality detection method of the distributed photovoltaic area in embodiment 1 of the application;

[0054] Figure 3 It is a specific step schematic diagram of obtaining the distance center curve and the feasible region matrix in embodiment 1 of the application;

[0055] Figure 4a It is a user data clustering result graph of the ordinary user in embodiment 1 of the application;

[0056] Figure 4b It is a user data clustering result graph of the photovoltaic user in embodiment 1 of the application;

[0057] Figure 5a It is a feasible region and cluster center curve graph of the ordinary user in embodiment 1 of the application;

[0058] Figure 5b It is a feasible region and cluster center curve graph of the photovoltaic user in embodiment 1 of the application;

[0059] Figure 6 It is a specific step schematic diagram of judging the user abnormal data in embodiment 1 of the application;

[0060] Figure 7a It is an abnormality judgment graph of the to-be-measured data of the ordinary user in embodiment 1 of the application;

[0061] Figure 7bThis is a graph for judging abnormal photovoltaic user test data in Embodiment 1 of the present invention. Detailed Implementation

[0062] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0063] Example 1

[0064] like Figure 1 and Figure 2 As shown, the method for detecting data anomalies in user areas with distributed photovoltaic power distribution zones provided by the present invention, in this embodiment, selects users in a certain actual distributed photovoltaic power distribution zone in Anhui Province for data anomaly detection, and includes the following basic steps:

[0065] Step S1: Taking the selected ordinary users and photovoltaic users with distributed photovoltaic substations as the research objects, collect the historical measurement data of the users in the previous month in the monitoring and data acquisition system; in this embodiment, the historical measurement data includes, but is not limited to: temperature and historical photovoltaic data for photovoltaic users.

[0066] In step S1 of this embodiment, which involves collecting historical measurement data, there are 130 users in the area, including 15 photovoltaic users. The data acquisition meter collects data every 15 minutes, totaling 96 data points per day. One user's measurement point is selected, and 96 measurement data points are collected daily for 31 consecutive days.

[0067]

[0068] Here, the row vector represents the measurement data for one day: {x i1 ,x i2 ,…,…x im}(1≤i≤n), m is the number of sampling points; the longitudinal quantity is represented by the sampling data of n days at the same time: {x 1j ,x 2j ,…,…x nj (1≤j≤m);

[0069] Step S2: Based on the user's historical measurement data, perform K-means clustering on the user's historical load, select the number of clusters according to the DBI index, and form the cluster center curves and feasible region matrix of various electricity consumption behaviors.

[0070] like Figure 3As shown, in the embodiment, the step S2 of obtaining the distance center curve and the feasible region matrix further includes the following specific steps:

[0071] Step S21, input the clustering cluster number k=2-10, randomly take k data from the historical data as the initial clustering center, calculate the distance from the remaining sample points to the clustering center, and distribute the corresponding sample points to the nearest clustering center;

[0072] Step S22, recalculate the average value of the distance from all points in each cluster to the clustering center, take it as the new clustering center, and redivide the sample points to the nearest clustering center;

[0073] Step S23, judge whether the clustering center changes before and after, if not, store and output the clustering result, and execute step S25;

[0074] Step S24, otherwise, replace the original clustering center with the new average value, and jump to execute the aforementioned step S22;

[0075] Step S25, calculate the Davies-Bouldin Index (DBI) by using the following formula (14) to measure the ratio between the intra-cluster distance and the inter-cluster distance;

[0076] As shown in FIGS. Figure 4a and Figure 4b In the embodiment, the calculation result shows that the ratio is the smallest when the number of clusters is 2, that is, the clustering effect is the best, and the clustering of the user historical data is further completed;

[0077]

[0078] Wherein, c i is the center of the i-th clustering cluster; σ i is the average distance from all points in the i-th clustering cluster to the center;

[0079] d(c i ,c j ) is the distance between the clustering centers c i and c j .

[0080] Step S26, by respectively analyzing the sampling data of each j time of the day in the aforementioned 2 classes, the threshold interval of the normal sampling data of the j time is obtained by using the following formula (15):

[0081] z j =[z jmax ,z jmin ] (9)

[0082] Wherein, z jmaxis the maximum value at the jth moment; z jmin is the minimum value at the jth moment.

[0083] As shown in Figure 5a and Figure 5b In the present embodiment, step S27, the above-mentioned processing is performed on all 96 sampling moments, the feasible region matrix of the abnormal amount of measurement data is obtained using the following formula (16), and the clustering center curve is plotted. The feasible region and the clustering center curve of the ordinary user, and the feasible region and the clustering center curve of the photovoltaic user are obtained;

[0084]

[0085] Step S3, according to the Spearman correlation coefficient, the to-be-tested daily data is divided into a category with high correlation, and the to-be-tested data collection point is compared with its feasible region to determine whether the user's data is abnormal;

[0086] As shown in Figure 6 In the present embodiment, step S3 of judging the abnormal data of the user further includes the following specific steps:

[0087] Step S31, taking the to-be-tested daily data and the clustering center curve as two given variables X and Y, first sorting the values of the foregoing variables according to the size, replacing each value with its corresponding rank, and if there are multiple values that are the same, taking the rank of the foregoing variable as the average of these values;

[0088] Step S32, the Pearson correlation coefficient between the ranks of the two variables is calculated using the following formula (17), the correlation between each clustering center and the to-be-tested daily data is analyzed, and the correlation coefficient values of the to-be-tested data of the ordinary user are as shown in Table 1, and the correlation coefficient values of the to-be-tested data of the photovoltaic user are as shown in Table 2.

[0089]

[0090] Wherein, d represents the rank difference of X and Y, and n represents the sample quantity.

[0091] Table 1 Correlation coefficient value table of to-be-tested data of ordinary user

[0092] Cluster center curve 1 Cluster center curve 2 To-be-measured curve correlation coefficient 0.48611 0.77788

[0093] Table 2 Correlation coefficient value table of to-be-tested data of photovoltaic user

[0094] Cluster center curve 1 Cluster center curve 2 To-be-measured curve correlation coefficient 0.94359 0.83451

[0095] As shown in Figure 7a and Figure 7bAs shown, in the embodiment, step S33, according to the correlation coefficient of the to-be-tested curve and the clustering center curve, the to-be-tested curve of the ordinary user is substituted into the feasible region curve corresponding to the clustering center curve 2 with a larger correlation coefficient, in the embodiment, please refer to Figure 5a , the abnormal power consumption behavior occurs at the time of 31-34 and 49-54, i.e., 7:45-8:30 in the morning and 12:15-1:30 in the afternoon, which exceeds the boundary of the feasible region; the to-be-tested curve of the photovoltaic user is substituted into the feasible region curve corresponding to the clustering center curve 1 with a larger correlation coefficient, please refer to Figure 4b , the abnormal power consumption behavior occurs at the time of 33-68, i.e., 8:15 in the morning to 5:00 in the afternoon.

[0096] In summary, the present application selects the Spearman correlation coefficient, excludes the interference of abnormal data, performs correlation analysis on the data of the to-be-tested day, divides the data of the to-be-tested day into the actual feasible region, and finally completes the abnormal detection of the user in the distributed photovoltaic area, and since the calculation of each measurement point at each time is independent of each other, the abnormal identification of multiple measurement points can be completed.

[0097] The present application has high abnormal data detection capability for ordinary users and photovoltaic users, is helpful for further repairing the abnormal measurement data and detecting the line loss anomaly of the area, and can effectively solve the abnormal data detection problem of different types of users in the distributed photovoltaic low-voltage distribution area.

[0098] The present application is based on the measurement data of the data acquisition and monitoring control system, selects different factors for k-means clustering for photovoltaic users and ordinary users, selects the best clustering number according to the DBI index, and divides the different power consumption behaviors of the same user more accurately, thereby improving the accuracy of the feasible region.

[0099] The present application utilizes the Spearman correlation coefficient when calculating the correlation of the to-be-tested data and the clustering center, and therefore does not need to care about how the data changes and what distribution it conforms to, but only needs to care about the position of the corresponding value of each variable. According to this method, the to-be-tested day data is divided into the corresponding category, whether the data of the user is abnormal is judged, and the abnormal detection precision is effectively improved.

[0100] The present application solves the technical problems of low data measurement efficiency, low prediction accuracy and low recognition accuracy in the abnormal data detection of different types of users in the distributed photovoltaic low-voltage distribution area in the prior art.

[0101] The above examples are only used to illustrate the technical solutions of the present application, and are not intended to limit the present application; although the present application has been described in detail with reference to the foregoing examples, those skilled in the art should understand that the technical solutions recorded in the foregoing examples can be modified, or some technical features can be replaced by equivalent features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for detecting anomalies in user data from distributed photovoltaic power stations, characterized in that, The method comprises: S1, for general users and photovoltaic users in a distributed photovoltaic area, collecting historical measurement data by using a preset monitoring and data acquisition system; S2, performing K-means clustering on user historical loads in the historical measurement data, selecting a cluster number according to a preset DBI index, and forming a cluster center curve and a feasible region matrix of each type of electricity consumption behavior; S2 comprises: S21, input the number of clustering clusters k From the historical measurement data, take k data sample points as initial clustering centers, and calculate the distances of the remaining data sample points to the clustering centers by using the following logic Compare the distances of the measurement data to the clustering centers of a certain class, The smaller the distance, the closer the measurement data is to the clustering center. The corresponding sample point is assigned to the adjacent clustering center; In the formula, represents the i-th remaining measurement data , represents the j-th cluster center ; S22, recalculating the average value of the distance of all data sample points in each cluster to the cluster center as a new cluster center, and re-dividing the data sample points to a new adjacent cluster center in the new cluster center; S23, judging whether the adjacent cluster center changes compared with the new adjacent cluster center; S24, if yes, storing and outputting the clustering result, and executing subsequent step S26, otherwise, replacing the initial cluster center with the new average value, and jumping to step S22; S25, (13) calculating a DBI index to measure the ratio between intra-cluster distance and inter-cluster distance, selecting the number of data sample points with the smallest ratio as the applicable cluster number, and obtaining user historical data clustering; S26, taking specific time sampling data in each user historical data cluster for analysis to obtain a normal sampling data threshold interval at the specific time; S27, executing step S26 for all sampling times to obtain the feasible region matrix, identifying abnormal measurement data, and drawing the cluster center curve according to the user historical data clustering; S3, calculating a Spearman correlation coefficient according to the cluster center curve and the to-be-measured day data, dividing the to-be-measured day data into a strong correlation category, comparing the collection point corresponding to the to-be-measured day data with the feasible region according to the feasible region matrix, and judging whether the to-be-measured day data of the user is abnormal; S3 comprises: S31, taking the to-be-measured day data and the cluster center curve as a first variable X and a second variable Y, sorting the first variable X and the second variable Y according to their sizes, and ranking the average value; S32, calculating the Pearson correlation coefficient between the ranks of the first variable X and the second variable Y, analyzing the correlation between each cluster center and the to-be-measured day data, and obtaining a to-be-measured data correlation coefficient value table; S33, according to the Pearson correlation coefficient, substituting the to-be-measured curve into the corresponding feasible region curve, and identifying the part exceeding the boundary of the feasible region as abnormal data.

2. The method of claim 1, wherein, In step S1, no less than two user measurement points and sampling points are selected, and the historical measurement data of a preset continuous time interval is recorded. 3.The method of claim 2, wherein, The historical measurement data is expressed by the following logic: where the row vector represents the measurement data of a day, and the column vector represents the sampling data of a day at the same time. n day.

4. The method of claim 1, wherein, In step S25, the DBI index is calculated by using the following logic: wherein is the center of the i th cluster, is the average distance of all points in the i th cluster to the center, is the distance between the cluster center and the center of the kth cluster.

5. The method of claim 1, wherein, In the step S26, the normal sampling data threshold interval at the time point is calculated using the following logic: j t = (t - 1) + (t - 1) / 2 In the formula, is the maximum value at the time j is the minimum value at the time is the maximum value at the time j is the minimum value at the time 6. The method of claim 1, wherein, In step S27, the feasible region matrix is obtained by using the following logic; 。 7. The method of claim 1, wherein, In step S32, the Pearson correlation coefficient between the ranks of the first variable X and the second variable Y is calculated by using the following logic: wherein d represents the rank difference of X and Y, n represents the number of samples.

8. The system for detecting abnormal data of distributed photovoltaic users in a certain area, which is used for executing the method for detecting abnormal data of distributed photovoltaic users in a certain area according to any one of the preceding claims 1 to 7, characterized in that, The system comprises: A historical data collection module is configured to collect historical measurement data of common users and photovoltaic users in a distributed photovoltaic area by using a preset monitoring and data collection system; A cluster center and feasible region acquisition module is configured to perform K-means clustering on historical loads of users in the historical measurement data, select a number of clusters according to a preset DBI index, and form cluster center curves and feasible region matrices of power consumption behaviors of different categories. The cluster center and feasible region acquisition module is connected with the historical data collection module. An anomaly determination module is configured to calculate a Spearman correlation coefficient according to to-be-measured daily data and the cluster center curves, divide the to-be-measured daily data into a strong correlation category, compare a collection point corresponding to the to-be-measured daily data with a feasible region according to the feasible region matrix, and determine whether the to-be-measured daily data of the user is abnormal.

Citation Information

Patent Citations

  • Fuzzy reasoning-based intelligent identification method for abnormal causes of line loss rate of power distribution area

    CN111817299A

  • Method for detecting electricity stealing and leakage of user in large range based on time convolutional network

    CN114662584A

  • Photovoltaic output power prediction method and system based on curve characteristic index clustering

    CN114792156A

  • Power distribution network line fault prediction method and device

    CN115270965A