A distributed photovoltaic area metering data online self-verification method and system

By calculating the mixed correlation coefficient and DTW distance of voltage subsequences and combining the Copula model to construct a user node adjacency matrix, abnormal users can be identified. This solves the voltage distortion and clock deviation problems in the metering data verification of distributed photovoltaic stations, realizes efficient online self-verification, and improves the accuracy of metering data and the intelligence level of the smart distribution network.

CN120296291BActive Publication Date: 2025-09-05STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510787390.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-09-05
Estimated Expiration
2045-06-13

AI Technical Summary

Technical Problem

Traditional methods for verifying metering data in distributed photovoltaic areas lead to inaccurate metering data analysis and judgment due to voltage distortion, clock deviation and excessive power supply radius, resulting in reduced accuracy and reliability.

Method used

By calculating the mixed correlation coefficient, DTW distance and Copula model of voltage subsequences, the user node adjacency matrix is ​​constructed, and the chain criterion is used to identify abnormal users and realize online self-verification.

Benefits of technology

It improves the accuracy and efficiency of metering data verification, detects anomalies in a timely manner, ensures the stable operation and power supply quality of the power system, and improves the intelligence level of the smart distribution network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296291B_ABST
    Figure CN120296291B_ABST
Patent Text Reader

Abstract

The present invention relates to an online self-verification method and system for distributed photovoltaic area metering data, belonging to the field of smart distribution networks. The method comprises obtaining distributed voltage metering data in the photovoltaic area and preprocessing it to obtain voltage subsequence data for each user; treating the distribution transformer as a user; calculating the mixed correlation coefficient of the voltage subsequences corresponding to the same time period for each user; calculating the distortion measurement weight corresponding to each user based on the mixed correlation coefficient; calculating the DTW distance of the voltage subsequences corresponding to the same time period for each user, and calculating the weighted DTW distance; and calculating the time series voltage correlation of the voltage subsequences corresponding to each user based on the weighted DTW distance; constructing a user node adjacency matrix in the area based on the time series voltage correlation; traversing the elements in the matrix to obtain multiple chain structures; the chain structure including the distribution transformer is the main chain, and independent sub-chain users not connected to the main chain are abnormal users. The method realizes online self-verification of photovoltaic area metering data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of smart distribution networks, and in particular to an online self-checking method and system for metering data of distributed photovoltaic areas. Background Art

[0002] In the field of smart distribution networks, metering data verification in distributed photovoltaic substations faces numerous technical challenges. Traditional methods often suffer from inaccurate correlation analysis and judgment due to voltage distortion, clock deviation, and excessive power supply radius caused by photovoltaic access, ultimately leading to abnormal metering data self-verification results. This is manifested in the following ways:

[0003] Voltage distortion: The randomness and uncertainty of PV output leads to voltage fluctuations at distribution transformers and users, disrupting the consistency of the voltage sequence.

[0004] Clock deviation: The clocks of different metering devices are not synchronized, resulting in deviations in the voltage data on the time axis, affecting correlation analysis;

[0005] The power supply radius is too long: Long-distance power supply leads to large voltage changes, further complicating correlation analysis.

[0006] These problems make it difficult for traditional methods to adapt to the complex characteristics of distributed photovoltaic areas, resulting in inaccurate analysis and judgment of distributed photovoltaic area metering data, and reduced accuracy and reliability of verification results. Summary of the Invention

[0007] In view of the above analysis, an embodiment of the present invention aims to provide an online self-verification method for distributed photovoltaic area metering data, so as to solve the technical problem of inaccurate online verification of metering data in existing distributed photovoltaic areas due to voltage distortion, clock deviation and excessive power supply radius in the same area.

[0008] The purpose of the present invention is mainly achieved through the following technical solutions:

[0009] The present invention provides an online self-checking method for distributed photovoltaic area metering data, comprising the following steps:

[0010] Obtain distributed voltage metering data in the photovoltaic area and preprocess it to obtain voltage subsequence data for each user; the distribution transformer in the photovoltaic area is regarded as one user;

[0011] Calculating the mixed correlation coefficient of the voltage subsequences corresponding to the same time period of each two users; and calculating the distortion measurement weight of the voltage subsequences corresponding to each two users based on the mixed correlation coefficient;

[0012] Calculating the DTW distance of the voltage subsequences corresponding to the same time period of each pair of users, calculating the weighted DTW distance based on the distortion metric weight and the DTW distance, and obtaining the time series voltage correlation of the voltage subsequences corresponding to each pair of users based on the weighted DTW distance calculation;

[0013] A user node adjacency matrix within the substation is constructed based on the time-series voltage correlation of the voltage subsequences corresponding to each two users; multiple chain structures are obtained by traversing the elements in the user node adjacency matrix; wherein the chain structure including the distribution transformer is the main chain, and independent sub-chain users that are not connected to the main chain are abnormal users.

[0014] Furthermore, two users and No. Time period corresponds to voltage subsequence and The mixed correlation coefficient is as follows:

[0015] ;

[0016] in, 、 、 and The corresponding voltage subsequences are and The mixed correlation coefficient, Pearson correlation coefficient, lower tail correlation coefficient and upper tail correlation coefficient of 、 、 are the corresponding weights of the Pearson correlation coefficient, lower tail correlation coefficient, and upper tail correlation coefficient of the voltage subsequence, respectively.

[0017] Furthermore, we calculate the pairwise user and No. Time period voltage subsequence and The ranking of each element in the corresponding voltage subsequence is calculated, and then the ranking is converted into a uniform distribution value;

[0018] User-based and The Clayton Copula model is constructed using the uniform distribution value of , and the parameters of the Clayton Copula model are estimated using the lower tail log likelihood function. , maximize the lower tail log likelihood function and obtain the best parameter estimate ,based on Calculate the lower tail correlation coefficient ;

[0019] User-based and The Gumbel Copula model is constructed using the uniform distribution value of the upper tail log likelihood function. , maximize the upper tail log-likelihood function and obtain the best parameter estimate ,based on Calculate the lower tail correlation coefficient .

[0020] Furthermore, based on user and The Clayton Copula model is constructed based on the uniform distribution value of , including:

[0021] User-based and The uniformly distributed value of constructs the cumulative distribution function of the Clayton Copula model;

[0022] Constructing a probability density function of a Clayton Copula model based on the cumulative distribution function;

[0023] Constructing a lower tail log-likelihood function based on the probability density function;

[0024] The Clayton Copula model parameters are estimated by iteratively maximizing the lower tail log-likelihood function using the gradient descent method. .

[0025] Furthermore, the lower tail log-likelihood function ,as follows:

[0026] ;

[0027] in, is the probability density function of the Clayton Copula model.

[0028] Furthermore, the best parameter estimates based on the Clayton Copula model , calculate the lower tail correlation coefficient ,as follows:

[0029] .

[0030] Furthermore, based on user and The uniformly distributed values ​​of construct the Gumbel Copula model, including:

[0031] User-based and The cumulative distribution function of the Gumbel Copula model is constructed using the uniformly distributed value of

[0032] Constructing a probability density function of a Gumbel Copula model based on the cumulative distribution function;

[0033] Constructing an upper tail log-likelihood function based on the probability density function;

[0034] The upper tail log-likelihood function is solved using Newton's method to obtain the Gumbel Copula model parameters. .

[0035] Furthermore, the lower tail log-likelihood function ,as follows:

[0036] ;

[0037] in, is the probability density function of the Gumbel Copula model.

[0038] Furthermore, the optimal parameter estimation based on Gumbel Copula , calculate the lower tail correlation coefficient ,as follows:

[0039] .

[0040] Furthermore, the ranking is converted into a uniformly distributed value as follows:

[0041] ;

[0042] in, 、 They are 、 The corresponding uniform distribution value; 、 For users 、 No. The first time period voltage subsequence Voltage data 、 In the voltage subsequence 、 Ranking in is the voltage subsequence length.

[0043] Furthermore, based on the mixed correlation coefficient of the voltage subsequence Calculate the distortion metric weight corresponding to the voltage subsequence:

[0044] ;

[0045] in, Represents a user and In the The distortion metric weight of the voltage subsequence in the time period, is the total number of voltage subsequences.

[0046] Furthermore, based on the distortion metric weight and the DTW distance, a weighted DTW distance is calculated as follows:

[0047] ;

[0048] in, For user voltage subsequence and DTW distance; is the corresponding weighted DTW distance.

[0049] Furthermore, based on the weighted DTW distance, the user and No. Voltage subsequence corresponding to the time period and Timing voltage dependence ,as follows:

[0050] .

[0051] Furthermore, based on the pairwise user and The corresponding voltage subsequence temporal voltage correlation Construct the user node adjacency matrix within the station area, including:

[0052] Preset correlation threshold ;

[0053] If the timing voltage correlation , then the user and Strong correlation, users and Connect into chains to obtain the user node adjacency matrix ,as follows:

[0054] ;

[0055] in, for the number of users, is a matrix element.

[0056] Furthermore, the adjacency matrix is ​​traversed using the breadth-first search algorithm. Elements include:

[0057] Start searching from any user node, mark the searched users as visited, and for the current user , check all its unvisited neighbor user nodes;

[0058] like , then the neighbor user is added to the queue and marked as visited;

[0059] Connect each newly visited user to its neighboring users that its predecessor has visited; ultimately, multiple chain structures are obtained.

[0060] Furthermore, the users in the main chain are normal users;

[0061] Users who obtain independent sub-chains that are not connected to the main chain are abnormal users.

[0062] Furthermore, distributed voltage metering data of the photovoltaic area is collected through a preset sampling frequency;

[0063] The distributed voltage metering data includes user voltage data of multiple distributed users and voltage data on the low-voltage side of the distribution transformer; the voltage data on the low-voltage side of the distribution transformer is regarded as one user voltage data.

[0064] Furthermore, the obtained distributed voltage metering data of the photovoltaic area is preprocessed, including:

[0065] After performing data deduplication, data correction, continuity verification, outlier processing, and missing value filling, first voltage time series data is obtained;

[0066] The first voltage time series data is divided into multiple segments using a sliding window function to obtain voltage subsequence data of multiple time periods for each user.

[0067] Furthermore, the correlation threshold The value is 0.9.

[0068] This application also provides an online self-checking system for distributed photovoltaic area metering data, including:

[0069] The data acquisition module is used to obtain the distributed voltage metering data of the photovoltaic area and pre-process it to obtain the voltage subsequence data of each user; the distribution transformer in the photovoltaic area is regarded as one user;

[0070] A hybrid correlation coefficient calculation module is used to calculate the hybrid correlation coefficient of the voltage subsequences corresponding to two users in the same time period; based on the hybrid correlation coefficient, the distortion measurement weight of the voltage subsequence corresponding to two users is calculated;

[0071] a weighted DTW distance calculation module, configured to calculate the DTW distance of voltage subsequences corresponding to two users in the same time period, calculate the weighted DTW distance based on the distortion metric weight and the DTW distance, and obtain the time series voltage correlation of the voltage subsequences corresponding to two users based on the weighted DTW distance calculation;

[0072] A chain structure analysis and verification module is used to construct a user node adjacency matrix within the substation based on the time-series voltage correlation of the voltage subsequences corresponding to each of the two users; traverse the elements in the user node adjacency matrix to obtain multiple chain structures; wherein the chain structure including the distribution transformer is the main chain, and independent sub-chain users that are not connected to the main chain are abnormal users.

[0073] Compared with the prior art, the present invention can achieve at least one of the following beneficial effects:

[0074] 1. The present invention calculates the mixed correlation coefficient of the voltage subsequences of each user and uses it as the distortion measurement weight of the voltage subsequence, thereby more accurately measuring the correlation between the voltage sequences of users in the distributed photovoltaic area and effectively improving the accuracy of voltage distortion detection in metering data verification;

[0075] 2. The present invention uses DTW distance and weighted DTW distance as indicators to measure the consistency of voltage subsequence fluctuations, cleverly solving the problem of deviation between voltage subsequences caused by clock asynchrony, and ensuring the accuracy of voltage sequence correlation analysis;

[0076] 3. The present invention uses a chain criterion to form a chain of highly correlated users and identifies users disconnected from the main chain as abnormal users, effectively solving a series of technical problems caused by the long power supply radius and improving the accuracy of metering data anomaly identification;

[0077] 4. The present invention realizes efficient online self-verification function, which can monitor and verify the metering data of distributed photovoltaic areas in real time, detect anomalies in time, greatly reduce manual intervention, and improve verification efficiency. At the same time, by accurately distinguishing between normal users and abnormal users, it helps to quickly deal with metering data anomalies and ensure the stable operation of the power system and the quality of power supply;

[0078] 5. This invention integrates the Clayton Copula model, the Gumbel Copula model, and the DTW algorithm to conduct in-depth analysis and mining of metering data. This significantly enhances the intelligence level and data analysis capabilities of the smart distribution network, providing strong support for efficient management and optimized operation of the power system.

[0079] In the present invention, the above-mentioned technical solutions can be combined with each other to achieve more preferred combinations. Other features and advantages of the present invention will be described in the following description, and some advantages will become apparent from the description or be learned through practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the contents particularly pointed out in the description and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0080] The accompanying drawings are only used for the purpose of illustrating specific embodiments and are not to be considered as limiting the present invention. Throughout the drawings, the same reference symbols denote the same components.

[0081] Figure 1 This is a flow chart of an online self-verification method for distributed photovoltaic area metering data in an embodiment of the present invention;

[0082] Figure 2 Schematic diagram of the path used in DTW distance calculation according to an embodiment of the present invention;

[0083] Figure 3 This is a schematic diagram of the chain formation of terminal users in a photovoltaic area in the present invention;

[0084] Figure 4 This is a schematic diagram of a distributed photovoltaic area metering data online self-verification system module. DETAILED DESCRIPTION

[0085] The preferred embodiments of the present invention will be described in detail below in conjunction with the accompanying drawings, wherein the accompanying drawings constitute a part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, and are not used to limit the scope of the present invention.

[0086] The present invention provides a distributed photovoltaic area metering data online self-verification method system. Based on the obtained distributed voltage metering data of the distribution transformers and users in the photovoltaic area, the voltage subsequence data of each user is obtained by pre-processing, and the mixed correlation coefficient of the voltage subsequences of each two users is calculated as a measure of the voltage distortion in the period; the DTW (Dynamic Time Warping) distance is used as an indicator to measure the consistency of voltage segment fluctuations, and the correlation between the voltage subsequence data is calculated based on the measure of voltage distortion; then a chain criterion is used to chain users with high correlation, and the broken chain users are output as users with abnormal metering data, thereby realizing online self-verification of distributed photovoltaic area metering data, which is effective in greatly improving the quality management of distribution network data and helping to realize the construction of a new type of intelligent power system.

[0087] Example 1:

[0088] A specific embodiment of the present invention discloses an online self-checking method for distributed photovoltaic area metering data, such as Figure 1 As shown, the following steps are included:

[0089] Step S1: Obtain distributed voltage metering data in the photovoltaic area and pre-process it to obtain voltage subsequence data for each user; the distribution transformer in the photovoltaic area is regarded as one user;

[0090] Step S2: Calculate the mixed correlation coefficient of the voltage subsequences corresponding to the same time period of each two users; and calculate the distortion measurement weight of the voltage subsequences corresponding to each two users based on the mixed correlation coefficient.

[0091] Step S3: calculating the DTW distance of the voltage subsequences corresponding to the same time period of each user, calculating the weighted DTW distance based on the distortion metric weight and the DTW distance, and obtaining the time series voltage correlation of the voltage subsequences corresponding to each user based on the weighted DTW distance calculation;

[0092] Step S4: constructing a user node adjacency matrix within the substation based on the time-series voltage correlation of the voltage subsequences corresponding to the two users; traversing the elements in the user node adjacency matrix to obtain multiple chain structures; wherein the chain structure including the distribution transformer is the main chain, and independent sub-chain users that are not connected to the main chain are abnormal users.

[0093] A distribution substation is a power supply unit consisting of a distribution transformer, the low-voltage lines within its power supply area, and the user area. A substation is the smallest functional unit in the distribution network that converts high-voltage electricity into low-voltage electricity and distributes it directly to end users. A distributed photovoltaic substation has only one distribution transformer.

[0094] Distribution transformer voltage refers to the voltage at the low-voltage side of a distribution transformer (abbreviated as distribution transformer). A distribution transformer is a device in a power system that converts high-voltage electricity into low-voltage energy suitable for user consumption. Distribution transformer voltage typically refers to the voltage output on the low-voltage side of the distribution transformer, which directly supplies power to users.

[0095] A photovoltaic power station refers to a station where distributed photovoltaic power generation systems, such as photovoltaic panels and inverters, are installed at the user end (such as residential buildings, industrial and commercial buildings, etc.). Photovoltaic power directly supplies power to local loads, close to the power load, reducing transmission losses. When power generation exceeds load demand, excess power can be fed back to the grid. When photovoltaic power generation is insufficient, power is drawn from the grid.

[0096] The distributed metering data of the photovoltaic area includes: voltage time series data at the user end and distribution transformer end (low-voltage side of the distribution transformer), current time series data at the user end and distribution transformer end, photovoltaic power generation and power consumption, photovoltaic power generation and power consumption.

[0097] Since voltage has fluctuation consistency and the voltage fluctuation trend in the same area is consistent, voltage data is selected as the data basis for online verification of the distributed photovoltaic area in the present invention; the voltage drop is small during transmission and the values ​​are close, making the voltage data more reliable and effective during the analysis and verification process.

[0098] Step S1 includes steps S11-S12.

[0099] Step S11: Obtain distributed voltage metering data in the photovoltaic area.

[0100] Collect distributed voltage metering data of photovoltaic areas through preset sampling frequency;

[0101] The distributed voltage metering data includes user voltage data of multiple distributed users and voltage data on the low-voltage side of the distribution transformer; the voltage data on the low-voltage side of the distribution transformer is regarded as one user voltage data.

[0102] The voltage obtained is the output voltage of the low-voltage side of the distribution transformer, which is consistent with the user voltage data. The distribution transformer is regarded as a user in the distributed photovoltaic area, and the voltage data on the low-voltage side of the distribution transformer is regarded as the voltage data of a user.

[0103] (1) Data collection:

[0104] User meters and distribution transformer meters are equipped with a meter carrier module to generate a high-frequency carrier signal that carries metering data (such as voltage, current, and power).

[0105] A concentrator is installed at the distribution transformer, which contains a carrier module to intercept the carrier signals from user meters and distribution transformer meters.

[0106] The concentrator converts the received carrier signal (analog signal) into a digital signal for subsequent processing.

[0107] Data is sampled every 15 minutes, generating daily voltage time series data with 96 time stamps (24 hours * 60 minutes / 15 minutes = 96), including distribution transformer and user voltage data. This high-frequency sampling ensures detailed and real-time voltage data.

[0108] (2) Data transmission:

[0109] The data is organized into structured data and transmitted to the electricity consumption information collection system through a communication network (such as optical fiber, wireless communication);

[0110] The electricity consumption information collection system synchronizes data to the wide table in the data center, which is used to store and manage large amounts of structured data.

[0111] A wide table is a data storage structure commonly used in databases or data warehouses. Each row (record) contains multiple columns (fields), enabling the storage of multi-dimensional data. Wide tables are designed to consolidate related data into a single table for easy querying and analysis. Table 1 shows examples of structured data in wide tables relevant to this invention.

[0112] Table 1: Examples of structured data related to the present invention in a wide table

[0113] ;

[0114] (3) Data extraction: Use SQL statements to extract user voltage data.

[0115] By executing SQL statements, the required voltage time series data of distribution transformers and users are extracted from the wide table of the data center database. The SQL statements specify conditions such as time range and user ID to extract the required user voltage data.

[0116] Step S12: pre-process the acquired distributed user voltage data to obtain voltage subsequence data of each user.

[0117] The obtained distributed voltage metering data of the photovoltaic area is preprocessed, including:

[0118] After performing data deduplication, data correction, continuity verification, outlier processing, and missing value filling, first voltage time series data is obtained;

[0119] The first voltage time series data is divided into multiple segments using a sliding window function to obtain voltage subsequence data of multiple time periods for each user.

[0120] Calculate the MD5 code of the obtained voltage measurement data. Voltage data with the same MD5 code are duplicate data. Keep one voltage measurement data with the same MD5 code.

[0121] Data correction, through sensor calibration linear compensation and sliding average filtering to eliminate system errors and high-frequency noise, improve data accuracy.

[0122] Continuity check, based on timestamp alignment and breakpoint detection to ensure timing continuity.

[0123] For outlier processing, the Z-score method is used to statistically detect and identify outliers, and data smoothing is achieved through interpolation or elimination.

[0124] Use interpolation methods to fill missing values, including the following two interpolation methods: linear interpolation and adjacent mean interpolation.

[0125] (1) For voltage time series data with fewer missing points, linear interpolation is used to fill in the missing values.

[0126] If the user voltage sequence data In, if is missing, and and If all values ​​are non-missing, then linear interpolation is performed using formula (1) to fill in the missing values ​​in the voltage series, as follows:

[0127] ;

[0128] in, Indicates the voltage sequence of distribution transformer or user No. data, and Respectively The voltage data of the previous moment and the next moment.

[0129] (2) If there are many missing values, use the adjacent average interpolation.

[0130] Fill the missing values ​​in the voltage series based on the average value of the adjacent data points of the missing value, as follows:

[0131] ;

[0132] in, for The number of adjacent non-empty data points in the voltage series, .

[0133] For example, The value is 3. The value can be adjusted according to specific needs.

[0134] If the amount of voltage data used is large, such as 192-point sequential voltage over 2 days, the value can be increased appropriately. values ​​to better reflect the overall trend of the data.

[0135] If the voltage data fluctuates greatly, The value can be appropriately reduced to avoid excessive noise affecting the interpolation results.

[0136] If there are many missing points, The value can be increased appropriately to ensure that there are enough adjacent data points for interpolation.

[0137] The pre-processed user and distribution transformer voltage time series data are obtained through interpolation method.

[0138] Through interpolation data preprocessing, the integrity and continuity of distribution transformer and user voltage data are ensured, providing a reliable data basis for subsequent analysis.

[0139] The preprocessed user voltage time series data is divided into multiple time segments using a sliding window function.

[0140] The parameters of the preset sliding window function include the window length and step length The window length is , if 30 minutes ( for 2 sampling points) or 60 minutes ( 4 sampling points), adjusted according to the actual data fluctuation characteristics .

[0141] In the user Voltage sequence In the example, sliding starts from the 0th moment, and the step length of each movement is (For example, 1 sampling point, i.e. 15 minutes). Time period voltage subsequence .

[0142] Repeat the above steps until all voltage sequence data are covered and the user The set of all voltage subsequences .

[0143] The specific formula is as follows:

[0144] ;

[0145] in, represents the window length of the sliding window function, Represents a user The voltage sequence, Indicates the window moving step size each time, Represents a user No. Time period voltage subsequence; Indicates the The starting position of the time period voltage subsequence, Indicates the The end position of the voltage subsequence in the time period.

[0146] For example, The value is 5. If the value is 2, the user The voltage subsequence of the first time period , the voltage subsequence of the second time period , and so on, eventually multiple voltage subsequences for each user are obtained.

[0147] Divide the long voltage sequence data of multiple users into multiple voltage subsequences to facilitate the analysis of voltage fluctuations and distortions.

[0148] In practical applications, by adjusting the window length and step length , which can adapt to different analysis needs.

[0149] The function of step S1 is to obtain and pre-process the voltage metering data of the distributed photovoltaic area and convert it into voltage subsequence data suitable for subsequent analysis.

[0150] Step S2 includes steps S21-S22.

[0151] Step S21: Calculate the mixed correlation coefficients of the voltage subsequences corresponding to the same time period of each two users.

[0152] The mixed correlation coefficient is calculated as a measure of voltage distortion during that time period. Because all users in the same PV grid are supplied by the same distribution transformer, there is a high degree of linear correlation and fluctuation consistency between each user's voltage subsequence (the voltage subsequence on the low-voltage side of the distribution transformer is also considered a user voltage subsequence). To measure this characteristic, the mixed correlation coefficient is used to represent the degree of distortion between voltage subsequences in the same time period.

[0153] The higher the degree of distortion, the lower the mixed correlation coefficient; the lower the degree of distortion, the higher the mixed correlation coefficient.

[0154] The mixed correlation coefficient is used to quantify the degree of distortion between voltage subsequences. Its core idea is that the voltages of users in the same substation area should be highly synchronized and the voltage fluctuation trends of users should be consistent under normal conditions.

[0155] If there is distortion, the correlation between voltage subsequences will decrease. Large-scale grid connection of photovoltaics will destroy the synchronization, which is manifested as a decrease in the hybrid correlation coefficient.

[0156] The mixed correlation coefficient is obtained by weighted summation of the subsequence Pearson correlation coefficient, the subsequence lower tail correlation coefficient, and the subsequence upper tail correlation coefficient.

[0157] Among them, the Pearson correlation coefficient measures the degree of linear correlation between voltage subsequences and reflects the overall trend consistency;

[0158] The lower tail correlation coefficient measures the correlation when the values ​​of voltage subsequences take smaller values ​​at the same time, that is, when the fluctuation trend is downward, such as when they drop suddenly at the same time;

[0159] The upper tail correlation coefficient measures the situation where the values ​​of the voltage subsequences take large values ​​at the same time, that is, the correlation of the simultaneous upward fluctuation trend, such as simultaneous surges.

[0160] Two users are represented as users and .

[0161] Two users and No. Time period corresponds to voltage subsequence and The mixed correlation coefficient is as follows:

[0162] ;

[0163] in, 、 、 and The corresponding voltage subsequences are and The mixed correlation coefficient, Pearson correlation coefficient, lower tail correlation coefficient and upper tail correlation coefficient of 、 、 are the corresponding weights of the Pearson correlation coefficient, lower tail correlation coefficient, and upper tail correlation coefficient of the voltage subsequence, respectively.

[0164] For example, considering that the linear correlation and fluctuation consistency of the voltage subsequence are both high, 、 、 The values ​​are all 1 / 3; 、 、 The weight value can be adjusted according to specific needs.

[0165] The Pearson correlation coefficient of the two user voltage subsequences is used to measure the linear correlation between the two voltage subsequences. and No. Time period voltage subsequence and The Pearson correlation coefficient between is calculated as follows:

[0166] ;

[0167] in, 、 For users and No. The first time in the voltage subsequence Voltage data, The length of the user voltage subsequence is equal to the window length of the sliding window , 、 For users and No. Average value of voltage subseries in a time period. Represents a user and The Pearson correlation coefficient, Pearson logo.

[0168] Under normal circumstances, the Pearson correlation coefficient , indicating a strong linear correlation, the voltage fluctuations of users in the same substation area are synchronized, reflecting the consistency of voltage fluctuations;

[0169] In the case of distortion, the Pearson correlation coefficient decreases significantly, which may be caused by equipment failure or wiring errors.

[0170] If the Pearson correlation coefficient of two user voltage subsequences is less than a predetermined threshold ,For example , it is determined that there may be a risk of distortion in the two voltage subsequences.

[0171] For example, the preset threshold The value is 0.8.

[0172] Calculate the two users and No. Time period voltage subsequence and The ranking of each element in the corresponding voltage subsequence is calculated, and then the ranking is converted into a uniform distribution value;

[0173] User-based and The Clayton Copula model is constructed using the uniform distribution value of , and the parameters of the Clayton Copula model are estimated using the lower tail log likelihood function. , maximize the lower tail log likelihood function and obtain the best parameter estimate ,based on Calculate the lower tail correlation coefficient ;

[0174] User-based and The Gumbel Copula model is constructed using the uniform distribution value of the upper tail log likelihood function. , maximize the upper tail log-likelihood function and obtain the best parameter estimate ,based on Calculate the lower tail correlation coefficient .

[0175] Calculating users and No. Time period corresponds to voltage subsequence and The lower tail correlation coefficient and the upper tail correlation coefficient between .

[0176] For users and No. Time period voltage subsequence and , calculate the rank of each voltage value in the subsequence and then convert it into a uniformly distributed value.

[0177] Convert the ranking to a uniformly distributed value as follows:

[0178] ;

[0179] in, 、 They are 、 The corresponding uniform distribution value; 、 For users and No. The first time period voltage subsequence Voltage data 、 In the voltage subsequence 、 Ranking in is the voltage subsequence length.

[0180] For example,

[0181] user No. The voltage subsequence of the time period is ;

[0182] user No. The voltage subsequence of the time period is .

[0183] Ranking starts at 1.

[0184] User No. Taking the voltage subsequence as an example, the ranking of each element in the subsequence is as follows:

[0185] , ranking ;

[0186] , ranking ;

[0187] , ranking ;

[0188] , ranking ;

[0189] , ranking .

[0190] If there are identical data in a voltage subsequence, the identical data will be ranked in the same order and ranked in parallel. Identical values ​​will receive the same ranking, and subsequent values ​​will skip the corresponding number of digits.

[0191] If the voltage subsequence is [221, 220, 223, 220, 219], the ranking is 4, 2, 5, 2, 1; if two are tied for rank 2, rank 3 is skipped.

[0192] , based on user No. Voltage subsequence For example, convert to uniform distribution value as follows:

[0193] ;

[0194] ;

[0195] ;

[0196] ;

[0197] ;

[0198] .

[0199] user No. The voltage subsequence is , converted to uniform distribution value using the same method as above .

[0200] Calculate the uniform distribution value and map all voltage values ​​in the voltage subsequence to the [0,1] interval to facilitate subsequent analysis.

[0201] User-based and The Clayton Copula model is constructed based on the uniform distribution value of , including:

[0202] User-based and The uniformly distributed value of constructs the cumulative distribution function of the Clayton Copula model;

[0203] Constructing a probability density function of a Clayton Copula model based on the cumulative distribution function;

[0204] Constructing a lower tail log-likelihood function based on the probability density function;

[0205] The Clayton Copula model parameters are estimated by iteratively maximizing the lower tail log-likelihood function using the gradient descent method. .

[0206] The lower tail correlation coefficient is used to measure the consistency of the downward fluctuations of two voltage subsequences (i.e., the voltage drops at the same time). The parameters of the Copula are calculated by constructing a Clayton Copula model using maximum likelihood estimation. , and thus calculate the lower tail correlation coefficient .

[0207] The cumulative distribution function of the Clayton Copula model is as follows:

[0208] ;

[0209] in, For users and No. A sequence of uniformly distributed values ​​of the voltage subsequences in a time period; are the Clayton Copula model parameters to be estimated, which are used to control the shape and correlation strength of the Clayton Copula The parameter is greater than 0.

[0210] Based on the transformed uniform distribution value , estimated by maximum likelihood estimation .

[0211] The probability density function of the Clayton Copula model describes the user and No. Uniformly distributed value sequence of voltage subsequences in a time period The joint probability density of .

[0212] The probability density function of the Clayton Copula model is as follows:

[0213] ;

[0214] The probability density function of the Clayton Copula model describes the joint probability density of two voltage subsequences at different values.

[0215] The lower tail log-likelihood function ,as follows:

[0216] ;

[0217] in, is the probability density function of the Clayton Copula model.

[0218] The probability density function of Clayton Copula Substituting the log-likelihood function into formula (9), we can obtain the specific log-likelihood expression.

[0219] Using the gradient descent optimization method, find the log-likelihood function The largest value;

[0220] The parameters of the Clayton Copula model are estimated by iteratively updating the maximum log-likelihood function using gradient descent. , the steps are as follows:

[0221] The first step is to initialize the parameters, including:

[0222] Initialization initial value , illustratively, ;

[0223] Setting the learning rate , illustratively, ;

[0224] Set the maximum number of iterations. For example, the maximum number of iterations is 100.

[0225] The second step is to calculate the gradient, including:

[0226] Calculate the gradient of the log-likelihood function, that is, the derivative, as follows:

[0227] ;

[0228] According to the probability density function of Clayton Copula, the gradient is expressed as follows:

[0229] ;

[0230] Step 3: Update parameters, including:

[0231] ;

[0232] Repeat steps 1 to 3 above until the maximum number of iterations is reached or the log-likelihood function value changes less than the preset threshold; obtain the maximum log-likelihood function. value.

[0233] Based on maximizing the log-likelihood function The best parameter estimate is obtained by solving ,as follows:

[0234] ;

[0235] The best parameter estimates based on the Clayton Copula model , calculate the lower tail correlation coefficient ,as follows:

[0236] ;

[0237] The lower tail correlation coefficient is used to measure the correlation between two voltage subsequences in the lower tail.

[0238] The closer the value is to 1, the stronger the lower tail correlation is;

[0239] The closer the value is to 0, the weaker the lower-tail correlation is.

[0240] Calculate the lower-tail correlation coefficient between two voltage subseries to measure their consistency in downward fluctuations. In power systems, lower-tail correlation indicates the tendency of two voltage subseries to decrease simultaneously. This is important for detecting voltage sags and assessing power system stability.

[0241] By constructing a Clayton Copula model, the correlation of user voltage subsequences under extreme conditions can be effectively captured, providing an important basis for voltage distortion analysis and anomaly detection.

[0242] The upper tail correlation coefficient is used to measure the consistency of the upward fluctuation (voltage rises at the same time) of two voltage subsequences. By constructing the Gumbel Copula model, the parameters of the Copula are calculated using maximum likelihood estimation. , and thus calculate the lower tail correlation coefficient .

[0243] User-based and The uniformly distributed values ​​of construct the Gumbel Copula model, including:

[0244] User-based and The cumulative distribution function of the Gumbel Copula model is constructed using the uniformly distributed value of

[0245] Constructing a probability density function of a Gumbel Copula model based on the cumulative distribution function;

[0246] Constructing an upper tail log-likelihood function based on the probability density function;

[0247] The upper tail log-likelihood function is solved using Newton's method to obtain the Gumbel Copula model parameters. .

[0248] The cumulative distribution function of the Gumbel Copula model is as follows:

[0249] ;

[0250] in, is the parameter to be estimated.

[0251] The probability density function of the Gumbel Copula is as follows:

[0252] ;

[0253] The probability density function of the Gumbel Copula describes the joint probability density of two voltage subsequences at different values.

[0254] The lower tail log-likelihood function ,as follows:

[0255] ;

[0256] in, is the probability density function of the Gumbel Copula model.

[0257] The probability density function of the Gumbel Copula Substitute the log-likelihood function to obtain the specific log-likelihood expression and the specific log-likelihood function expression.

[0258] The upper tail log-likelihood function is solved using Newton's method to obtain the Gumbel Copula model parameters. , the steps are as follows:

[0259] The first step is to initialize the parameters, including:

[0260] Set initial value ; Convergence threshold ; The maximum number of iterations is 100.

[0261] The second step is to calculate the gradient and Hessian matrix, including:

[0262] (1) Calculate the gradient, that is, the first-order derivative, as follows:

[0263] ;

[0264] (2) Calculate the Hessian matrix, that is, the second-order derivative, as follows:

[0265] ;

[0266] The third step is to update the parameters as follows:

[0267] ;

[0268] Repeat steps 1 to 3 above until the maximum number of iterations is reached or the log-likelihood function value changes less than the preset threshold. ; Get the maximum log-likelihood function value.

[0269] Based on maximizing the log-likelihood function The best parameter estimate is obtained by solving ,as follows:

[0270] ;

[0271] Once the optimal parameter estimates for the Gumbel Copula are obtained , calculate the lower tail correlation coefficient according to formula (22) ;

[0272] Optimal parameter estimation based on Gumbel Copula , calculate the lower tail correlation coefficient ,as follows:

[0273] ;

[0274] The lower tail correlation coefficient is used to measure the correlation between two user voltage subsequences in the upper tail. The closer the value is to 1, the stronger the upper tail correlation is; the closer the value is to 0, the weaker the upper tail correlation is.

[0275] Based on the calculated Pearson correlation coefficient of the voltage subsequence , voltage subsequence upper tail correlation coefficient The lower tail correlation coefficient of the sum voltage subsequence , weighted calculation can be used to calculate the subsequence mixed correlation coefficient ,

[0276] Hybrid correlation coefficient based on voltage subsequence Calculate the distortion metric weight corresponding to the voltage subsequence:

[0277] ;

[0278] in, Represents a user and No. The distortion metric weight of the voltage subsequence in the time period, is the total number of voltage subsequences.

[0279] Two users and No. Distortion metric weight of time period voltage subsequence , which is calculated based on the mixed correlation coefficient between voltage subsequences and is used to reflect the degree of voltage distortion.

[0280] The function of step S2 is to calculate the mixed correlation coefficients of the voltage subsequences of two users in the same time period, and derive the distortion measurement weight based on the mixed correlation coefficients to quantify the voltage distortion degree.

[0281] Step S3, specifically.

[0282] Calculate the DTW distance of the voltage subsequences corresponding to the same time period of each user, and calculate the correlation between the voltage subsequences. Figure 2 The figure shows a path diagram for DTW distance calculation. A path exists between two voltage subsequences that minimizes the distance between them, preventing inaccuracies in correlation calculations due to clock skew. The upper and lower solid lines in the figure represent two user-series voltage data, while the dashed line represents the DTW distance calculated between the user voltage subsequences. The horizontal axis represents the sampling time, with one point collected every 15 minutes.

[0283] The DTW distance is used to measure the similarity between two voltage subsequences.

[0284] To address clock skew in existing methods, we calculated the DTW distance between each user's voltage subsequences. The DTW distance uses a dynamic programming algorithm to calculate the optimal alignment path between two time series. It effectively handles time series expansion and shifts, thereby quantifying the similarity between two voltage subsequences.

[0285] DTW distance, calculated as follows:

[0286] ;

[0287] in, Represents a user No. Time period voltage subsequence ,user No. Time period voltage subsequence The minimum dynamic time warping distance between 、 Represents users and No. The first subsequence To avoid excessive time deviation, set , where 3 represents 3 sampling time points.

[0288] After obtaining the DTW distance between each voltage subsequence, based on the distortion metric weight The weighted DTW distance is calculated from the DTW distance.

[0289] By introducing weights, smaller weights can be assigned to subsequences that are more affected by voltage distortion, thereby reducing their impact on the overall similarity during the calculation process.

[0290] Based on the distortion metric weight and the DTW distance, the weighted DTW distance is calculated as follows:

[0291] ;

[0292] in, For user voltage subsequence and DTW distance; is the corresponding weighted DTW distance.

[0293] The smaller it is, the higher the correlation is.

[0294] Based on the weighted DTW distance, calculate the user and No. Voltage subsequence corresponding to the time period and Timing voltage dependence ,as follows:

[0295] ;

[0296] The function of step S3 is to consider the timing voltage correlation between voltage distortion and clock deviation, calculate the DTW distance and weighted DTW distance between two user voltage subsequences, so as to measure the similarity of voltage subsequences and solve the clock deviation problem.

[0297] Step S4, specifically.

[0298] Based on the time-series voltage correlation of the voltage subsequences corresponding to each two users, a user node adjacency matrix within the substation is constructed, with each user being treated as a node.

[0299] Based on two users and The corresponding voltage subsequence temporal voltage correlation Construct the user node adjacency matrix within the station area, including:

[0300] Preset correlation threshold γ;

[0301] If the timing voltage correlation , then the user and Strong correlation, users and Connect into chains to obtain the user node adjacency matrix ,as follows:

[0302] ;

[0303] in, for the number of users, is a matrix element.

[0304] The correlation threshold The value is 0.9.

[0305] , indicating connection; indicating user and users There is a strong correlation in voltage;

[0306] , indicating disconnection, indicating user and users There is no strong correlation in voltage.

[0307] Use breadth-first search algorithm to traverse the adjacency matrix Elements include:

[0308] Start searching from any user node, mark the searched users as visited, and for the current user , check all its unvisited neighbor user nodes;

[0309] like , then the neighbor user is added to the queue and marked as visited;

[0310] Connect each newly visited user to its neighboring users that its predecessor has visited; ultimately, multiple chain structures are obtained.

[0311] Construct an adjacency graph, where nodes are users and edges are connections based on the adjacency matrix;

[0312] The chain structure including the distribution transformer is the main chain, and the independent sub-chain users that are not connected to the main chain are abnormal users.

[0313] The users in the main chain are normal users;

[0314] Users who obtain independent sub-chains that are not connected to the main chain are abnormal users.

[0315] Output a list of abnormal users, mark them as metering abnormal users, and record the substation number to which they belong.

[0316] Normal user list, output the normal user list in the main chain.

[0317] like Figure 3 As shown in the figure, user terminals are nodes. If the voltage correlation between nodes is greater than the threshold of 0.9, they are connected by a black solid line to form a chain. Otherwise, they are connected by a dotted line. The broken chain area is the non-main chain area, and the voltage correlation with the main chain users is less than the threshold of 0.9. Users in the broken chain area are users with abnormal metering.

[0318] Step S4 verifies metering data anomalies based on a voltage-correlation chain criterion, outputting a list of users with metering anomalies in the disconnected area and a list of users with normal metering in the main chain. This chain criterion effectively identifies users with metering anomalies, ensuring the accuracy of metering data in the substation area and the stable operation of the power system.

[0319] For users with abnormal metering data identified, determine whether their time series voltage data is within the normal range.

[0320] First, according to the voltage limit standard in the power grid, the upper and lower limit voltage amplitudes are set as The standard voltage is 220V. If it exceeds the limit, it is considered that the abnormal user voltage exceeds the limit, as follows:

[0321] ;

[0322] in, Represents a user The over-limit judgment result of the time series voltage data is 0, which means the voltage is within the limit; 1, which means the voltage is over-limit. Represents a user No. The standard voltage is 220 V, and the allowable range is [198 V, 242 V].

[0323] If the user voltage exceeds the limit instantaneously (such as a short-term sag caused by a lightning strike), it will be marked as exceeding the limit even if it is normal at other times.

[0324] And judge whether the voltage is distorted based on the fluctuation detection of adjacent points. Absolute value of voltage difference between adjacent points:

[0325] ;

[0326] in, Represents a user In the time period and The absolute value of the voltage change.

[0327] Calculating users Average difference of voltage series :

[0328] ;

[0329] in, Indicates the number of time series voltage data, is the average intensity of voltage fluctuation.

[0330] Calculating standard deviation ,as follows:

[0331] ;

[0332] Standard deviation , the discrete degree of voltage difference reflects the fluctuation stability.

[0333] Set the multiple threshold , if satisfied , it is considered that there is voltage distortion. For example, The value ranges from 1.5 to 3, depending on specific needs.

[0334] The voltage distortion is determined as follows:

[0335] ;

[0336] in, Indicates the distortion judgment result. A value of 1 indicates that the user's voltage is distorted, and a value of 0 indicates normal voltage fluctuations. If the timing is within the limit and the voltage distortion is normal, the user is reported as having a suspected abnormal user-to-user relationship. Otherwise, the abnormal user is reported as having a suspected metering data abnormality.

[0337] The function of step S4 is to construct the user node adjacency matrix in the substation area based on the time-series voltage correlation of each user's voltage subsequence, identify the main chain and the independent sub-chain users disconnected from the main chain, output the list of normal and abnormal users, and ensure the accuracy of the substation area metering data.

[0338] At distributed photovoltaic substations, power supply operators replace damaged sensors and rewire to eliminate poor connections based on abnormal user output. Furthermore, they conduct a comprehensive assessment of the photovoltaic equipment to ensure all components are operating within specifications. Finally, they check the relationship between the abnormal user and the transformer to ensure that the on-site records are consistent with the system records.

[0339] Once the field work is complete and all known issues are resolved, the latest voltage time series data for the distribution transformer and users is uploaded back to the central database. The system then uses this updated information to re-run the metering data self-verification process in this invention to verify the effectiveness of the previous corrective measures.

[0340] According to the verification results, the accuracy of an online self-verification method for distributed photovoltaic area metering data has reached This means that the method can effectively screen out most abnormal cases, greatly reducing the false alarm rate and improving the timeliness of resolving distortion problems. The method of the present invention not only enhances the intelligent level of power grid management, but also provides solid technical support for achieving greener and more efficient energy utilization.

[0341] In summary, by combining online self-verification with on-site real-time verification, a closed-loop management system can be built to continuously monitor and improve the metering data quality of photovoltaic stations, thereby promoting the healthy development of the entire power system.

[0342] Example 2:

[0343] Another embodiment of the present invention discloses an online self-verification system for distributed photovoltaic area metering data, thereby implementing the online self-verification method for distributed photovoltaic area metering data in embodiment 1. The specific implementation of each module refers to the corresponding description in embodiment 1.

[0344] like Figure 4 As shown, the system includes a data acquisition module M1, a hybrid correlation coefficient calculation module M2, a weighted DTW distance calculation module M3 and a chain structure analysis and verification module M4.

[0345] The data acquisition module M1 is used to obtain the distributed voltage metering data of the photovoltaic area and pre-process it to obtain the voltage subsequence data of each user; the distribution transformer in the photovoltaic area is regarded as one user;

[0346] A mixed correlation coefficient calculation module M2 is used to calculate the mixed correlation coefficient of the voltage subsequences corresponding to two users in the same time period; based on the mixed correlation coefficient, the distortion measurement weight of the voltage subsequence corresponding to two users is calculated;

[0347] A weighted DTW distance calculation module M3 is configured to calculate the DTW distance of voltage subsequences corresponding to two users in the same time period, calculate the weighted DTW distance based on the distortion metric weight and the DTW distance, and obtain the time series voltage correlation of the voltage subsequences corresponding to two users based on the weighted DTW distance calculation;

[0348] The chain structure analysis and verification module M4 is used to construct a user node adjacency matrix within the substation based on the time-series voltage correlation of the voltage subsequences corresponding to each two users; traverse the elements in the user node adjacency matrix to obtain multiple chain structures; wherein the chain structure including the distribution transformer is the main chain, and independent sub-chain users that are not connected to the main chain are abnormal users.

[0349] In summary, the online self-verification method for distributed photovoltaic area metering data according to the embodiment of the present invention has the following beneficial effects:

[0350] 1. The present invention calculates the mixed correlation coefficient of the voltage subsequences of each user and uses it as the distortion measurement weight of the voltage subsequence, thereby more accurately measuring the correlation between the voltage sequences of users in the distributed photovoltaic area and effectively improving the accuracy of voltage distortion detection in metering data verification;

[0351] 2. The present invention uses DTW distance and weighted DTW distance as indicators to measure the consistency of voltage subsequence fluctuations, cleverly solving the problem of deviation between voltage subsequences caused by clock asynchrony, and ensuring the accuracy of voltage sequence correlation analysis;

[0352] 3. The present invention uses a chain criterion to form a chain of highly correlated users and identifies users disconnected from the main chain as abnormal users, effectively solving a series of technical problems caused by the long power supply radius and improving the accuracy of metering data anomaly identification;

[0353] 4. The present invention realizes efficient online self-verification function, which can monitor and verify the metering data of distributed photovoltaic areas in real time, detect anomalies in time, greatly reduce manual intervention, and improve verification efficiency. At the same time, by accurately distinguishing between normal users and abnormal users, it helps to quickly deal with metering data anomalies and ensure the stable operation of the power system and the quality of power supply;

[0354] 5. This invention integrates the Clayton Copula model, the Gumbel Copula model, and the DTW algorithm to conduct in-depth analysis and mining of metering data. This significantly enhances the intelligence level and data analysis capabilities of the smart distribution network, providing strong support for efficient management and optimized operation of the power system.

[0355] Those skilled in the art will appreciate that all or part of the process steps of the above-described embodiments can be implemented by instructing related hardware through a computer program, and the program can be stored in a computer-readable storage medium, such as a magnetic disk, an optical disk, a read-only memory, or a random access memory.

[0356] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by any technician familiar with this technical field within the technical scope disclosed by the present invention should be covered by the scope of protection of the present invention.

Claims

1. A distributed photovoltaic area metering data online self-verification method, characterized in that: include: Obtain distributed voltage metering data in the photovoltaic area and preprocess it to obtain voltage subsequence data for each user; the distribution transformer in the photovoltaic area is regarded as one user; Calculating the mixed correlation coefficient of the voltage subsequences corresponding to the same time period of each two users; and calculating the distortion measurement weight of the voltage subsequences corresponding to each two users based on the mixed correlation coefficient; Calculating the DTW distance of the voltage subsequences corresponding to the same time period of each pair of users, calculating the weighted DTW distance based on the distortion metric weight and the DTW distance, and obtaining the time series voltage correlation of the voltage subsequences corresponding to each pair of users based on the weighted DTW distance calculation; Constructing a user node adjacency matrix within the substation based on the time-series voltage correlation of the voltage subsequences corresponding to each of the two users; Traversing the elements in the user node adjacency matrix to obtain multiple chain structures; wherein the chain structure including the distribution transformer is the main chain, and independent sub-chain users that are not connected to the main chain are abnormal users; Two users and No. Time period corresponds to voltage subsequence and The mixed correlation coefficient is as follows: ; in, 、 、 and The corresponding voltage subsequences are and The mixed correlation coefficient, Pearson correlation coefficient, upper tail correlation coefficient and lower tail correlation coefficient of 、 、 are the corresponding weights of the Pearson correlation coefficient, upper tail correlation coefficient, and lower tail correlation coefficient of the voltage subsequence respectively; Hybrid correlation coefficient based on voltage subsequence Calculate the distortion metric weight corresponding to the voltage subsequence: ; in, Represents a user and In the The distortion metric weight of the voltage subsequence in the time period, is the total number of voltage subsequences.

2. The method according to claim 1, characterized in that Calculate the two users and No. Time period voltage subsequence and The ranking of each element in the corresponding voltage subsequence is calculated, and then the ranking is converted into a uniform distribution value; User-based and The Clayton Copula model is constructed using the uniform distribution value of , and the upper tail log likelihood function is used to estimate the parameters of the Clayton Copula model. , maximize the upper tail log-likelihood function and obtain the best parameter estimate ,based on Calculate the upper tail correlation coefficient ; User-based and The Gumbel Copula model is constructed using the uniform distribution value of the lower tail log likelihood function. , maximize the lower tail log likelihood function and obtain the best parameter estimate ,based on Calculate the lower tail correlation coefficient .

3. The method according to claim 2, characterized in that User-based and The uniformly distributed values ​​of construct the ClaytonCopula model, including: User-based and The uniformly distributed value of constructs the cumulative distribution function of the Clayton Copula model; Constructing a probability density function of a Clayton Copula model based on the cumulative distribution function; Constructing an upper tail log-likelihood function based on the probability density function; The gradient descent method is used to iteratively maximize the upper tail log-likelihood function to estimate the Clayton Copula model parameters. .

4. The method according to claim 3, characterized in that The upper tail log-likelihood function ,as follows: ; in, is the probability density function of the Clayton Copula model; 、 For users and No. The first time in the voltage subsequence Voltage data; 、 They are 、 The corresponding uniform distribution value; is the length of the user voltage subsequence.

5. The method according to claim 4, characterized in that: The best parameter estimates based on the Clayton Copula model , calculate the upper tail correlation coefficient ,as follows: 。 6. The method according to claim 2, characterized in that: User-based and The uniformly distributed values ​​of construct the GumbelCopula model, including: User-based and The cumulative distribution function of the Gumbel Copula model is constructed using the uniformly distributed value of Constructing a probability density function of a Gumbel Copula model based on the cumulative distribution function; Constructing a lower tail log-likelihood function based on the probability density function; The Newton method is used to solve the lower tail log likelihood function to obtain the Gumbel Copula model parameters. .

7. The method according to claim 6, characterized in that The lower tail log-likelihood function ,as follows: ; in, is the probability density function of the Gumbel Copula model; 、 For users and No. The first time in the voltage subsequence Voltage data; 、 They are 、 The corresponding uniform distribution value; is the length of the user voltage subsequence.

8. The method according to claim 7, characterized in that: Optimal parameter estimation based on Gumbel Copula , calculate the lower tail correlation coefficient ,as follows: 。 9. The method according to claim 2, characterized in that: Convert the ranking to a uniformly distributed value as follows: ; in, 、 They are 、 The corresponding uniform distribution value; 、 For users 、 No. The first time period voltage subsequence Voltage data 、 In the voltage subsequence 、 Ranking in is the length of the user voltage subsequence.

10. The method according to claim 1, characterized in that: Based on the distortion metric weight and the DTW distance, the weighted DTW distance is calculated as follows: ; in, For user voltage subsequence and DTW distance; is the corresponding weighted DTW distance.

11. The method according to claim 10, characterized in that: Based on the weighted DTW distance, calculate the user and No. Voltage subsequence corresponding to the time period and Timing voltage dependence ,as follows: 。 12. The method according to claim 1, characterized in that: Based on two users and The corresponding voltage subsequence temporal voltage correlation Construct the user node adjacency matrix within the station area, including: Preset correlation threshold ; If the timing voltage correlation , then the user and Strong correlation, users and Connect into chains to obtain the user node adjacency matrix ,as follows: ; in, for the number of users, is a matrix element.

13. The method according to claim 12, characterized in that: Use breadth-first search algorithm to traverse the adjacency matrix Elements include: Start searching from any user node, mark the searched users as visited, and for the current user , check all its unvisited neighbor user nodes; like , then the neighbor user is added to the queue and marked as visited; Connect each newly visited user to its neighboring users that its predecessor has visited; ultimately, multiple chain structures are obtained.

14. The method according to any one of claims 1 to 13, characterized in that The users in the main chain are normal users; Users who obtain independent sub-chains that are not connected to the main chain are abnormal users.

15. The method according to any one of claims 1 to 13, characterized in that: Collect distributed voltage metering data of photovoltaic areas through preset sampling frequency; The distributed voltage metering data includes user voltage data of multiple distributed users and voltage data on the low-voltage side of the distribution transformer; the voltage data on the low-voltage side of the distribution transformer is regarded as one user voltage data.

16. The method according to claim 15, characterized in that The obtained distributed voltage metering data of the photovoltaic area is preprocessed, including: After performing data deduplication, data correction, continuity verification, outlier processing, and missing value filling, first voltage time series data is obtained; The first voltage time series data is divided into multiple segments using a sliding window function to obtain voltage subsequence data of multiple time periods for each user.

17. The method according to claim 12, characterized in that: The correlation threshold The value is 0.

9.

18. A distributed photovoltaic area metering data online self-checking system, characterized by: include: The data acquisition module is used to obtain the distributed voltage metering data of the photovoltaic area and pre-process it to obtain the voltage subsequence data of each user; the distribution transformer in the photovoltaic area is regarded as one user; A hybrid correlation coefficient calculation module is used to calculate the hybrid correlation coefficient of the voltage subsequences corresponding to two users in the same time period; based on the hybrid correlation coefficient, the distortion measurement weight of the voltage subsequence corresponding to two users is calculated; a weighted DTW distance calculation module, configured to calculate the DTW distance of voltage subsequences corresponding to two users in the same time period, calculate the weighted DTW distance based on the distortion metric weight and the DTW distance, and obtain the time series voltage correlation of the voltage subsequences corresponding to two users based on the weighted DTW distance calculation; A chain structure analysis and verification module is used to construct a user node adjacency matrix within the substation based on the time-series voltage correlation of the voltage subsequences corresponding to each two users; Traversing the elements in the user node adjacency matrix to obtain multiple chain structures; wherein the chain structure including the distribution transformer is the main chain, and independent sub-chain users that are not connected to the main chain are abnormal users; Two users and No. Time period corresponds to voltage subsequence and The mixed correlation coefficient is as follows: ; in, 、 、 and The corresponding voltage subsequences are and The mixed correlation coefficient, Pearson correlation coefficient, upper tail correlation coefficient and lower tail correlation coefficient of 、 、 are the corresponding weights of the Pearson correlation coefficient, upper tail correlation coefficient, and lower tail correlation coefficient of the voltage subsequence respectively; Hybrid correlation coefficient based on voltage subsequence Calculate the distortion metric weight corresponding to the voltage subsequence: ; in, Represents a user and In the The distortion metric weight of the voltage subsequence in the time period, is the total number of voltage subsequences.

Citation Information

Patent Citations

  • Power distribution area user phase identification method and system based on screened voltage data

    CN112701675A

  • Power distribution network area topology identification method and system

    CN118472923A