Distributed photovoltaic district metering data online self-checking method and system
By calculating the mixed correlation coefficient and DTW distance of the voltage subsequence, combined with the Copula model and chain structure criterion, the voltage distortion and clock deviation problems in the metrological data verification of distributed photovoltaic platform areas are solved, and efficient online self-check is achieved, improving the accuracy of the metrological data and the stability of the power system.
Patent Information
- Application Number
- CN202510787390.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-13
AI Technical Summary
In the calibration of the metrology data in the distributed photovoltaic platform area, the traditional method causes the measurement data to be misjudged due to voltage distortion, clock deviation and too long power supply radius, and the accuracy and reliability are reduced.
By calculating the mixed correlation coefficient, DTW distance and Copula model of the voltage subsequence, the user node adjacency matrix is constructed, and abnormal users are identified using chain structure criteria to achieve online self-checking.
It improves the accuracy and efficiency of metrological data verification, reduces manual intervention, ensures the stable operation of the power system and the quality of power supply, and improves the intelligence level of the intelligent distribution network.
Smart Images

Figure CN120296291A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent distribution networks, and in particular, to an online self-verification method and system for metering data of distributed photovoltaic substations. Background Art
[0002] In the field of intelligent distribution networks, the verification of metering data in distributed photovoltaic substations faces many technical challenges. When traditional methods process metering data, due to problems such as voltage distortion caused by photovoltaic access, clock deviation, and too long power supply radius, the correlation judgment is often inaccurate, and finally the self-verification result of the metering data is abnormal. The specific manifestations are as follows: Voltage distortion: The randomness and uncertainty of photovoltaic output lead to voltage fluctuations of distribution transformers and users, destroying the consistency of the voltage sequence; Clock deviation: The clocks of different metering devices are not synchronized, resulting in deviations of voltage data on the time axis and affecting the correlation analysis; Too long power supply radius: The long-distance power supply causes too large voltage changes, further exacerbating the complexity of correlation judgment.
[0003] These problems make it difficult for traditional methods to adapt to the complex characteristics of distributed photovoltaic substations, resulting in inaccurate judgment of metering data in distributed photovoltaic substations and a decrease in the accuracy and reliability of verification results. Summary of the Invention
[0004] In view of the above analysis, the embodiments of the present invention aim to provide an online self-verification method for metering data of distributed photovoltaic substations to solve the technical problem of inaccurate online verification of metering data in existing distributed photovoltaic substations due to voltage distortion, clock deviation, and too large power supply radius under the same substation.
[0005] The object of the present invention is mainly achieved through the following technical solutions: The present invention provides an online self-verification method for metering data of distributed photovoltaic substations, including the following steps: Obtain the distributed voltage metering data of the photovoltaic substation and perform preprocessing to obtain the voltage subsequence data of each user; the distribution transformer in the photovoltaic substation is regarded as a user; Calculate the mixed correlation coefficient of the voltage subsequences corresponding to two users in the same time period; calculate the distortion metric weight corresponding to the voltage subsequences of two users based on the mixed correlation coefficient; Calculate the DTW distance of the voltage subsequences corresponding to two users in the same time period, calculate the weighted DTW distance based on the distortion metric weight and the DTW distance, and calculate the temporal voltage correlation corresponding to the voltage subsequences of two users based on the weighted DTW distance; Construct the adjacent matrix of user nodes in the transformer area based on the time - series voltage correlation of the voltage subsequences corresponding to each pair of users; traverse the elements in the user node adjacent matrix to obtain multiple chain structures; among them, the chain structure including the distribution transformer is the main chain, and the independent sub - chain users without connection to the main chain are abnormal users.
[0006] Further, for each pair of users and in the time period, the mixed correlation coefficient of the corresponding voltage subsequences and is as follows: ; Among them, , , and are the mixed correlation coefficient, Pearson correlation coefficient, lower - tail correlation coefficient and upper - tail correlation coefficient of the corresponding voltage subsequences and respectively, , , are the corresponding weights of the Pearson correlation coefficient, lower - tail correlation coefficient and upper - tail correlation coefficient of the voltage subsequence respectively.
[0007] Further, calculate the rankings of each element in the voltage subsequences and in the time period for each pair of users and , and then convert the rankings into uniformly distributed values; Based on the uniformly distributed values of users and , construct a Clayton Copula model, and use the lower - tail log - likelihood function to estimate the parameters of the Clayton Copula model to maximize the lower - tail log - likelihood function, and obtain the best parameter estimate value . Based on , calculate the lower - tail correlation coefficient ; Based on the uniformly distributed values of users and , construct a Gumbel Copula model, and use the upper - tail log - likelihood function to estimate the Gumbel Copula model , to maximize the upper - tail log - likelihood function, obtain the best parameter estimate value . Based on , calculate the lower - tail correlation coefficient .
[0008] Further, a Clayton Copula model is constructed based on the uniform distribution values of the user and , including: Based on the uniform distribution values of the user and , construct the cumulative distribution function of the Clayton Copula model; Construct the probability density function of the Clayton Copula model based on the cumulative distribution function; Construct the lower tail log-likelihood function based on the probability density function; Use the gradient descent method to iteratively maximize the lower tail log-likelihood function to estimate the parameters of the Clayton Copula model.
[0009] Further, the lower tail log-likelihood function is as follows: ; wherein, is the probability density function of the Clayton Copula model.
[0010] Further, based on the optimal parameter estimate value of the Clayton Copula model, calculate the lower tail correlation coefficient , as follows: .
[0011] Further, a Gumbel Copula model is constructed based on the uniform distribution values of the user and , including: Based on the uniform distribution values of the user and , construct the cumulative distribution function of the Gumbel Copula model; Construct the probability density function of the Gumbel Copula model based on the cumulative distribution function; Construct the upper tail log-likelihood function based on the probability density function; Use Newton's method to solve the upper tail log-likelihood function to obtain the parameters of the Gumbel Copula model.
[0012] Further, the lower tail log-likelihood function is as follows: ; wherein, Is the probability density function of the Gumbel Copula model.
[0013] Furthermore, based on the optimal parameter estimation of the Gumbel Copula , calculate the lower tail correlation coefficient , as follows: .
[0014] Furthermore, convert the said ranking into uniformly distributed values, as follows: ; where , are respectively , corresponding uniformly distributed values; , are respectively the , th voltage data of the voltage subsequence of the th , in the voltage subsequences , ; is the length of the voltage subsequence.
[0015] Furthermore, calculate the distortion metric weight corresponding to the voltage subsequence based on the hybrid correlation coefficient of the voltage subsequence: ; where represents the distortion metric weight of the voltage subsequence of user and in the th time period, is the total number of voltage subsequences.
[0016] Furthermore, calculate the weighted DTW distance based on the said distortion metric weight and the DTW distance, as follows: ; where is the DTW distance between the user voltage subsequences and ; is the corresponding weighted DTW distance.
[0017] Furthermore, based on the said weighted DTW distance, calculate the voltage subsequence and of user th time period corresponding to and Time - series voltage correlation is as follows: .
[0018] Furthermore, based on the time - series voltage correlation and of the corresponding voltage subsequences of pairwise users construct the adjacency matrix of user nodes in the distribution area, including: A preset correlation threshold ; If the time - series voltage correlation , then users and have strong correlation, connect users and into a chain to obtain the adjacency matrix of user nodes as follows: ; Among them, is the number of users of is the matrix element.
[0019] Furthermore, traverse the elements in the adjacency matrix using the breadth - first search algorithm, including: Start searching from any user node, mark the searched users as visited, and for the current user , check all its unvisited neighbor user nodes; If , then add the neighbor user to the queue and mark it as visited; Connect each newly visited user with its predecessor visited neighbor user; finally, obtain multiple chain - like structures.
[0020] Furthermore, obtain the users in the main chain as normal users; Obtain the users of the independent sub - chains that have no connection with the main chain as abnormal users.
[0021] Furthermore, collect the distributed voltage measurement data of the photovoltaic distribution area through a preset sampling frequency; Among them, the distributed voltage measurement data includes the user voltage data of multiple distributed users and the voltage data of the low - voltage side of the distribution transformer; the voltage data of the low - voltage side of the distribution transformer is regarded as a user voltage data.
[0022] Furthermore, pre - process the obtained distributed voltage measurement data of the photovoltaic distribution area, including: After data deduplication, data correction, continuity verification, outlier processing, and missing value filling, the first voltage time series data is obtained; The first voltage time series data is divided into multiple segments by using a sliding window function to obtain voltage subsequence data for each user in multiple time periods.
[0023] Furthermore, the correlation threshold takes a value of 0.9.
[0024] This application also provides an online self-verification system for distributed photovoltaic substation measurement data, including: A data acquisition module, which is used to acquire distributed voltage measurement data of a photovoltaic substation and perform preprocessing to obtain voltage subsequence data for each user; the distribution transformer in the photovoltaic substation is regarded as a user; A hybrid correlation coefficient calculation module, which is used to calculate the hybrid correlation coefficient of the voltage subsequences corresponding to two users in the same time period; based on the hybrid correlation coefficient, the distortion measurement weight corresponding to two users is calculated; A weighted DTW distance calculation module, which is used to calculate the DTW distance of the voltage subsequences corresponding to two users in the same time period, calculate the weighted DTW distance based on the distortion measurement weight and the DTW distance, and calculate the time series voltage correlation of the voltage subsequences corresponding to two users based on the weighted DTW distance; A chain structure analysis and verification module, which is used to construct an adjacency matrix of user nodes in the substation based on the time series voltage correlation of the voltage subsequences corresponding to two users; traverse the elements in the user node adjacency matrix to obtain multiple chain structures; among them, the chain structure including the distribution transformer is the main chain, and the independent sub-chain users not connected to the main chain are abnormal users.
[0025] Compared with the prior art, the present invention can at least achieve one of the following beneficial effects: 1. By calculating the hybrid correlation coefficient of the voltage subsequences of two users and using it as the distortion measurement weight of the voltage subsequences, the present invention can more accurately measure the correlation between the voltage sequences of users in a distributed photovoltaic substation, effectively improving the accuracy of voltage distortion detection in measurement data verification; 2. By using the DTW distance and the weighted DTW distance as indicators to measure the fluctuation consistency of voltage subsequences, the present invention cleverly solves the deviation problem caused by clock asynchronization between voltage subsequences, ensuring the accuracy of voltage sequence correlation analysis; 3. By using a chain criterion to chain users with high correlation and identifying users disconnected from the main chain as abnormal users, the present invention effectively solves a series of technical problems caused by too long power supply radius and improves the accuracy of abnormal identification of measurement data; 4. The present invention realizes an efficient online self-checking function, can monitor and check the metering data of a distributed photovoltaic substation area in real time, discover anomalies in a timely manner, greatly reduce manual intervention, and improve the checking efficiency. At the same time, by accurately distinguishing normal users from abnormal users, it helps to quickly handle abnormal metering data problems, ensuring the stable operation of the power system and power supply quality; 5. The present invention integrates various technical means such as the Clayton Copula model, the Gumbel Copula model, and the DTW algorithm to deeply analyze and mine the metering data. It significantly improves the intelligent level and data analysis ability of the intelligent distribution network, providing strong support for the efficient management and optimized operation of the power system.
[0026] In the present invention, the above technical solutions can also be combined with each other to achieve more preferred combined solutions. Other features and advantages of the present invention will be described in the subsequent specification, and some advantages can be made obvious from the specification, or understood by implementing the present invention. The objectives and other advantages of the present invention can be realized and obtained through the content specifically pointed out in the specification and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The drawings are only for the purpose of showing specific embodiments and are not considered as limitations on the present invention. Throughout the drawings, the same reference signs denote the same components; Figure 1 It is a flowchart of an online self-checking method for metering data of a distributed photovoltaic substation area in an embodiment of the present invention; Figure 2 It is a schematic diagram of the path during DTW distance calculation in an embodiment of the present invention; Figure 3 It is a schematic diagram of the chain formation situation of each terminal user in a certain photovoltaic substation area of the present invention; Figure 4 It is a schematic diagram of the modules of an online self-checking system for metering data of a distributed photovoltaic substation area. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0028] The following will specifically describe the preferred embodiments of the present invention with reference to the drawings. The drawings form a part of this application and are used together with the embodiments of the present invention to explain the principles of the present invention, not to limit the scope of the present invention.
[0029] In the present invention, an online self-checking method and system for metering data of a distributed photovoltaic substation area is based on the obtained distributed voltage metering data of the distribution transformer and users in the photovoltaic substation area. After preprocessing, the voltage subsequence data of each user is obtained, and the hybrid correlation coefficient of the voltage subsequences of two users is calculated as a measure of voltage distortion during this period. The DTW (Dynamic Time Warping) distance is used as an index to measure the consistency of voltage segment fluctuations, and the correlation between voltage subsequence data is calculated based on the measure of voltage distortion. Then, a chain criterion is used to chain users with high correlation, and the users with broken chains are output as abnormal users for metering data, thus realizing the online self-checking of metering data in the distributed photovoltaic substation area, greatly improving the quality management of distribution network data, and contributing to the construction of a new type of intelligent power system.
[0030] Embodiment 1: A specific embodiment of the present invention discloses an online self-checking method for metering data of a distributed photovoltaic substation area, as Figure 1 shown, including the following steps: Step S1: Obtain the distributed voltage metering data of the photovoltaic substation area and perform preprocessing to obtain the voltage subsequence data of each user; the distribution transformer in the photovoltaic substation area is regarded as a user; Step S2: Calculate the hybrid correlation coefficient of the voltage subsequences corresponding to two users in the same time period; calculate the distortion measure weight of the voltage subsequences corresponding to two users based on the hybrid correlation coefficient; Step S3: Calculate the DTW distance of the voltage subsequences corresponding to two users in the same time period, calculate the weighted DTW distance based on the distortion measure weight and the DTW distance, and calculate the time-series voltage correlation of the voltage subsequences corresponding to two users based on the weighted DTW distance; Step S4: Construct an adjacency matrix of user nodes in the substation area based on the time-series voltage correlation of the voltage subsequences corresponding to two users; traverse the elements in the adjacency matrix of user nodes to obtain multiple chain structures; among them, the chain structure including the distribution transformer is the main chain, and the independent sub-chain users not connected to the main chain are abnormal users.
[0031] A distribution substation area refers to a power supply unit composed of a distribution transformer and the low-voltage lines and user areas within its power supply range. The substation area is the smallest functional unit in the distribution network that converts high-voltage electricity into low-voltage electricity and directly distributes it to end users. There is only one distribution transformer in a distributed photovoltaic substation area.
[0032] The distribution transformer voltage refers to the voltage value at the low-voltage end of the distribution transformer (abbreviation: distribution transformer). The distribution transformer is a device in the power system used to convert high-voltage electrical energy into low-voltage electrical energy suitable for users. The distribution transformer voltage usually refers to the voltage output at the low-voltage side of the distribution transformer, which directly supplies power to users.
[0033] A photovoltaic substation refers to a substation where a distributed photovoltaic power generation system such as photovoltaic panels and inverters is installed at the user side (such as residential houses, industrial and commercial buildings, etc.). The photovoltaic power directly supplies power to local loads, is close to the power consumption load, reduces transmission losses, and can feed the excess power back into the power grid when the generated power exceeds the load demand. When the photovoltaic power generation is insufficient, power is taken from the power grid.
[0034] The distributed measurement data of the photovoltaic substation includes: voltage time series data at the user side and the distribution transformer side (low-voltage side of the distribution transformer), current time series data at the user side and the distribution transformer side, photovoltaic power generation power and power consumption, photovoltaic power generation amount and power consumption amount.
[0035] Since the voltage has a consistent fluctuation, the voltage fluctuation trends in the same substation are the same, so the voltage data is selected as the data basis for online verification of the distributed photovoltaic substation in the present invention; the voltage drop is small during the transmission process and the values are close, making the voltage data more reliable and effective in the analysis and verification process.
[0036] Step S1 includes steps S11 - S12.
[0037] Step S11: Obtain the distributed voltage measurement data of the photovoltaic substation.
[0038] Collect the distributed voltage measurement data of the photovoltaic substation through a preset sampling frequency; Among them, the distributed voltage measurement data includes the user voltage data of multiple distributed users and the voltage data of the low-voltage side of the distribution transformer; the voltage data of the low-voltage side of the distribution transformer is regarded as a user voltage data.
[0039] What is obtained is the voltage output from the low-voltage side of the distribution transformer, which is consistent with the user voltage data. The distribution transformer is regarded as a user of the distributed photovoltaic substation, and the voltage data of the low-voltage side of the distribution transformer is regarded as a user voltage data.
[0040] (1) Data collection: The user electricity meter and the distribution transformer electricity meter are equipped with electricity meter carrier modules for generating high-frequency carrier signals. The high-frequency carrier signals carry measurement data (such as voltage, current, power, etc.).
[0041] A concentrator is installed at the distribution transformer. The concentrator contains a carrier module for intercepting the carrier signals from the user electricity meter and the distribution transformer electricity meter.
[0042] The concentrator converts the received carrier signals (analog signals) into digital signals for subsequent processing.
[0043] The data sampling frequency is once every 15 minutes, forming daily voltage time series data with 96 timestamp points (24 hours * 60 minutes / 15 minutes = 96), including distribution transformer and user voltage data. High-frequency sampling ensures the meticulousness and real-time nature of the voltage data.
[0044] (2) Data transmission: The data is organized into structured data and transmitted to the power consumption information acquisition system through a communication network (such as optical fiber, wireless communication); The power consumption information acquisition system synchronizes the data into a wide table in the data middle platform, and the wide table is used to store and manage a large amount of structured data.
[0045] A wide table (Wide Table) is a data storage structure, usually used in databases or data warehouses. Its characteristic is that each row (record) contains multiple columns (fields) and can store multi-dimensional data. The wide table aims to integrate relevant data into one table for easy querying and analysis. An example of the structured data related to the present invention in the wide table is shown in Table 1.
[0046] Table 1: Example of structured data related to the present invention in the wide table ; (3) Data extraction: Use SQL statements to extract user voltage data.
[0047] By executing SQL statements, the required voltage time series data of the distribution transformer and users is extracted from the wide table in the data middle platform database. The SQL statements specify conditions such as the time range and user ID to extract the required user voltage data.
[0048] Step S12: Preprocess the obtained distributed user voltage data to obtain voltage subsequence data for each user.
[0049] Preprocess the obtained distributed voltage measurement data of the photovoltaic substation area, including: After performing data deduplication, data correction, continuity verification, outlier processing, and missing value filling, the first voltage time series data is obtained; Use a sliding window function to divide the first voltage time series data into multiple segments to obtain voltage subsequence data for each user in multiple time periods.
[0050] Calculate the MD5 code of the obtained voltage measurement data. Voltage data with the same MD5 code indicates duplicate data, and only one voltage measurement data with the same MD5 code is retained.
[0051] Data correction, through sensor calibration linear compensation and moving average filtering, eliminates system errors and high-frequency noise, improving data accuracy.
[0052] Continuity check, ensuring temporal continuity based on timestamp alignment and breakpoint detection.
[0053] Outlier handling, using the Z-score method to statistically detect and identify outliers, and achieving data smoothing through interpolation or elimination.
[0054] Using interpolation methods to fill missing values, including the following two interpolation methods: linear interpolation and adjacent mean interpolation.
[0055] (1) For voltage time series data with fewer missing points, linear interpolation is used to fill in the missing values.
[0056] If the user voltage sequence data in, if is missing, and and are both non-missing values, then linear interpolation is performed using formula (1) to fill in the missing values in the voltage sequence, as follows: ; Where, represents the th data of the distribution transformer or user voltage sequence , and respectively represent the voltage data at the previous and next moments of
[0057] (2) If there are more missing values, adjacent mean interpolation is used.
[0058] Based on the average of the data points adjacent to the missing value, to fill in the missing values in the voltage sequence, as follows: ; Where, is the number of adjacent non-empty data points of in the voltage sequence,
[0059] Exemplarily, takes the value 3. The value can be adjusted according to specific requirements.
[0060] If the amount of voltage data used is large, such as 192-point time series voltage for 2 days, value can be appropriately increased to better reflect the overall trend of the data.
[0061] If the voltage data fluctuates greatly, value can be appropriately decreased to avoid excessive noise affecting the interpolation result.
[0062] If there are more missing points, The value can be appropriately increased to ensure that there are enough adjacent data points for interpolation.
[0063] The preprocessed user and distribution transformer voltage time series data are obtained through interpolation.
[0064] Through interpolation data preprocessing, the integrity and continuity of the distribution transformer and user voltage data are ensured, providing a reliable data basis for subsequent analysis.
[0065] The preprocessed user voltage time series data are divided into multiple time segments using a sliding window function.
[0066] The preset parameters of the sliding window function include the window length and the step size . The window length is . If it takes 30 minutes ( is 2 sampling points) or 60 minutes ( is 4 sampling points), it is adjusted according to the actual data fluctuation characteristics. .
[0067] In the voltage sequence of the user , starting from the 0th moment, it slides each time with a step size of (for example, 1 sampling point, that is, 15 minutes). Then the voltage subsequence of the th time period is extracted.
[0068] Repeat the above steps until all voltage sequence data are covered, and all voltage subsequence sets of the user are obtained.
[0069] The specific formula is as follows: ; where represents the window length of the sliding window function, represents the voltage sequence of the user , represents the step size of the window movement each time, represents the user the th time period voltage subsequence; represents the starting position of the th time period voltage subsequence, represents the ending position of the th time period voltage subsequence.
[0070] Exemplarily, takes a value of 5, takes a value of 2, then the 1st time period voltage subsequence of the user , the voltage subsequence of the second time period , and so on, eventually multiple voltage subsequences for each user are obtained.
[0071] The long voltage sequence data of multiple users is divided into multiple voltage subsequences to facilitate the analysis of voltage fluctuation and distortion.
[0072] In practical applications, by adjusting the window length and step length , which can adapt to different analysis needs.
[0073] The function of step S1 is to obtain and pre-process the voltage metering data of the distributed photovoltaic area and convert it into voltage sub-sequence data suitable for subsequent analysis.
[0074] Step S2 includes steps S21-S22.
[0075] Step S21, calculating the mixed correlation coefficients of the voltage subsequences corresponding to the same time period of two users.
[0076] The mixed correlation coefficient is calculated as a measure of voltage distortion in this period. Since users in the same photovoltaic area are powered by the same distribution transformer, there is a high degree of linear correlation and fluctuation consistency between the voltage subsequences of each user (the voltage subsequence on the low-voltage side of the distribution transformer is also regarded as a user voltage subsequence). To measure this characteristic, the mixed correlation coefficient is used to represent the degree of distortion between voltage subsequences in the same time period.
[0077] The higher the degree of distortion, the lower the mixed correlation coefficient; the lower the degree of distortion, the higher the mixed correlation coefficient.
[0078] The mixed correlation coefficient is used to quantify the degree of distortion between voltage subsequences. The core idea is that the same distribution transformer is used to supply power to the same area, and the voltage of users in the area should be highly synchronized. Under normal conditions, the voltage fluctuation trend of users is consistent.
[0079] If there is distortion, the correlation between voltage subsequences will decrease. Large-scale grid-connected photovoltaics will destroy synchronization, which is manifested as a decrease in the mixed correlation coefficient.
[0080] The mixed correlation coefficient is obtained by weighted summing the subsequence Pearson correlation coefficient, the subsequence lower tail correlation coefficient and the subsequence upper tail correlation coefficient.
[0081] Among them, the Pearson correlation coefficient measures the linear correlation between voltage subsequences and reflects the overall trend consistency; The lower tail correlation coefficient measures the situation where the values between voltage subsequences take smaller values at the same time, that is, the correlation of the downward fluctuation trend at the same time, such as a sudden drop at the same time; The upper tail correlation coefficient measures the situation where the values in voltage subsequences are simultaneously large, that is, the correlation of the upward fluctuation trends, such as soaring simultaneously.
[0082] Two users are represented as user and .
[0083] Two users and The time period corresponds to the mixed correlation coefficients of the voltage subsequences and are as follows: ; Among them, , , and are the mixed correlation coefficient, Pearson correlation coefficient, lower tail correlation coefficient, and upper tail correlation coefficient of the corresponding voltage subsequences and , , , are the corresponding weights of the Pearson correlation coefficient, lower tail correlation coefficient, and upper tail correlation coefficient of the voltage subsequence respectively.
[0084] Exemplarily, considering that the linear correlation and fluctuation consistency of the voltage subsequence are relatively high, take , , The values are all 1 / 3; , , The weight values can be adjusted according to specific requirements.
[0085] The Pearson correlation coefficient of the voltage subsequences of two users is used to measure the linear correlation between two voltage subsequences. The Pearson correlation coefficient between the voltage subsequences and of the time period of users and is calculated as follows: ; Among them, , are the and th voltage data in the voltage subsequences of the th time period of users, , , are the average values of the voltage subsequences for the and th time period of the user respectively. represents the Pearson correlation coefficient between the and users, and is the Pearson identifier.
[0086] Under normal circumstances, the Pearson correlation coefficient , indicating a strong linear correlation, the voltage fluctuations of users in the same substation area are synchronized, reflecting the consistency of voltage fluctuations; Under distorted circumstances, the Pearson correlation coefficient drops significantly, which may be caused by equipment failure or wiring error resulting in voltage distortion.
[0087] If the Pearson correlation coefficient of the voltage subsequences of two users is less than a predetermined threshold , for example , it is determined that there may be a risk of distortion in the two voltage subsequences.
[0088] Exemplarily, the predetermined threshold takes a value of 0.8.
[0089] Calculate the rankings of each element in the and th time period voltage subsequences of two users and in their corresponding voltage subsequences, and then convert the rankings into uniformly distributed values; Based on the uniformly distributed values of the and users, construct a Clayton Copula model, and use the lower tail log-likelihood function to estimate the parameters of the Clayton Copula model to maximize the lower tail log-likelihood function, obtaining the best parameter estimate value , and calculate the lower tail correlation coefficient based on ; Based on the uniformly distributed values of the and users, construct a Gumbel Copula model, and use the upper tail log-likelihood function to estimate the Gumbel Copula model , to maximize the upper tail log-likelihood function, obtaining the best parameter estimate value , and calculate the lower tail correlation coefficient based on .
[0090] Calculate the and the corresponding voltage subsequence of the time period and the lower-tail correlation coefficient and the upper-tail correlation coefficient therebetween.
[0091] For user and the voltage subsequence of the time period and , calculate the rank of each voltage value in the subsequence, and then convert it into a uniformly distributed value.
[0092] Convert the said rank into a uniformly distributed value as follows: ; wherein, , are respectively , the corresponding uniformly distributed values; , are respectively the and th voltage data of the voltage subsequence of the , in the voltage subsequence , ; is the length of the voltage subsequence.
[0093] Exemplarily, the th voltage subsequence of user ; the th voltage subsequence of user .
[0094] The rank starts from 1.
[0095] Taking the th voltage subsequence of user as an example, rank each element in the subsequence as follows: ; ; ; ; ; ; ; , ranking .
[0096] If there are identical data in a voltage subsequence, the identical data have the same ranking, and a tied ranking method is adopted. Identical values receive the same ranking, and the rankings of subsequent values skip the corresponding number of digits.
[0097] If the voltage subsequence is [221, 220, 223, 220, 219], the rankings are 4, 2, 5, 2, 1; the two tied rankings of 2 skip ranking 3.
[0098] , taking the th voltage subsequence of the user as an example, it is converted into uniformly distributed values as follows: ; ; ; ; ; .
[0099] The th voltage subsequence of the user is , and it is converted into uniformly distributed values by the same method as above .
[0100] Calculate the uniformly distributed values, map all voltage values in the voltage subsequence to the interval [0, 1] for subsequent analysis.
[0101] Based on the uniformly distributed values of the user and , construct a Clayton Copula model, including: Based on the uniformly distributed values of the user and , construct the cumulative distribution function of the Clayton Copula model; Based on the cumulative distribution function, construct the probability density function of the Clayton Copula model; Based on the probability density function, construct the lower tail log-likelihood function; Use the gradient descent method to iteratively maximize the lower tail log-likelihood function to estimate the parameters of the Clayton Copula model.
[0102] The lower tail correlation coefficient is used by users to measure the consistency of downward fluctuations (i.e., simultaneous voltage drops) of two voltage subsequences. By constructing a Clayton Copula model, the parameters of the Copula are calculated using maximum likelihood estimation , so as to calculate the lower tail correlation coefficient .
[0103] The cumulative distribution function of the Clayton Copula model is in the following form: ; where are the sequences of uniformly distributed values of the voltage subsequences of the and th time period of the user respectively; is the parameter of the Clayton Copula model to be estimated, which is used to control the shape and correlation strength of the Clayton Copula The parameter is greater than 0.
[0104] Based on the transformed uniformly distributed values , estimate by the maximum likelihood estimation method.
[0105] The probability density function of the Clayton Copula model describes the joint probability density of the sequences of uniformly distributed values of the voltage subsequences of the and th time period of the user .
[0106] The probability density function of the Clayton Copula model is as follows: ; The probability density function of the Clayton Copula model describes the joint probability density of two voltage subsequences at different values.
[0107] The lower tail log-likelihood function , is as follows: ; where is the probability density function of the Clayton Copula model.
[0108] Substitute the probability density function of the Clayton Copula into the log-likelihood function of formula (9) to obtain a specific log-likelihood expression.
[0109] Use the gradient descent optimization method to find the value that maximizes the log-likelihood function The largest value; Estimate the parameters of the Clayton Copula model by maximizing the log-likelihood function through iterative updates using the gradient descent method , the steps are as follows: First step, initialize the parameters, including: Initialize the initial value , exemplarily, ; Set the learning rate , exemplarily, ; Set the maximum number of iterations, exemplarily, the maximum number of iterations is 100 times.
[0110] Second step, calculate the gradient, including: Calculate the gradient of the log-likelihood function, that is, take the derivative, as follows: ; According to the probability density function of the Clayton Copula, the gradient is expressed as follows: ; Third step, update the parameters, including: ; Repeat the above steps one to three until the maximum number of iterations is reached, or the change in the log-likelihood function value is less than the preset threshold; obtain the value that maximizes the log-likelihood function.
[0111] Based on the value that maximizes the log-likelihood function, solve to obtain the optimal parameter estimate value , as follows: ; Based on the optimal parameter estimate value of the Clayton Copula model, calculate the lower tail correlation coefficient , as follows: ; The lower tail correlation coefficient is used to measure the correlation between two voltage subsequences at the lower tail.
[0112] The closer the value is to 1, the stronger the lower tail correlation;
[0113] Calculate the lower tail correlation coefficient between two voltage subsequences to measure their consistency during downward fluctuations. In a power system, lower tail correlation indicates the tendency of two voltage subsequences to decline simultaneously. This is very important for detecting voltage sags and evaluating the stability of the power system.
[0114] By constructing a Clayton Copula model, the correlation of user voltage subsequences in extreme cases can be effectively captured, providing an important basis for voltage distortion analysis and anomaly detection.
[0115] The upper tail correlation coefficient, used to measure the consistency of upward fluctuations (voltages rising simultaneously) of two voltage subsequences. By constructing a Gumbel Copula model, the parameters of the Copula are calculated using maximum likelihood estimation , thereby calculating the lower tail correlation coefficient .
[0116] Based on the user and 's uniformly distributed values, construct a Gumbel Copula model, including: Based on the user and 's uniformly distributed values, construct the cumulative distribution function of the Gumbel Copula model; Based on the cumulative distribution function, construct the probability density function of the Gumbel Copula model; Based on the probability density function, construct the upper tail log-likelihood function; Use Newton's method to solve the upper tail log-likelihood function to obtain the Gumbel Copula model parameters .
[0117] The cumulative distribution function of the Gumbel Copula model is as follows: ; where is the parameter to be estimated.
[0118] The probability density function of the Gumbel Copula is as follows: ; The probability density function of the Gumbel Copula describes the joint probability density of two voltage subsequences at different values.
[0119] The lower tail log-likelihood function is as follows: ; where It is the probability density function of the Gumbel Copula model.
[0120] Substitute the probability density function of the Gumbel Copula into the log-likelihood function to obtain a specific log-likelihood expression, and get a specific log-likelihood function expression.
[0121] Use Newton's method to solve the upper-tail log-likelihood function to obtain the Gumbel Copula model parameters , the steps are as follows: The first step is to initialize the parameters, including: Set the initial value ; the convergence threshold ; the maximum number of iterations is 100 times.
[0122] The second step is to calculate the gradient and Hessian matrix, including: (1) Calculate the gradient, that is, the first derivative, as follows: ; (2) Calculate the Hessian matrix, that is, the second derivative, as follows: ; The third step is to update the parameters, as follows: ; Repeat the above steps one to three until the maximum number of iterations is reached, or the change in the log-likelihood function value is less than the preset threshold ; obtain the value that maximizes the log-likelihood function.
[0123] Based on the value that maximizes the log-likelihood function, solve to obtain the best parameter estimate value , as follows: ; Once the best parameter estimate of the Gumbel Copula is obtained , calculate the lower-tail correlation coefficient according to formula (22) ; Based on the best parameter estimate of the Gumbel Copula , calculate the lower-tail correlation coefficient , as follows: ; The lower-tail correlation coefficient is used to measure the correlation between two user voltage subsequences in the upper tail. The closer the value is to 1, the stronger the upper-tail correlation; the closer the value is to 0, the weaker the upper-tail correlation.
[0124] Based on the calculated Pearson correlation coefficient of the voltage subsequence , the upper tail correlation coefficient of the voltage subsequence and the lower tail correlation coefficient of the voltage subsequence , and the mixed correlation coefficient of the subsequence can be calculated by weighted calculation , Based on the mixed correlation coefficient of the voltage subsequence calculate the distortion metric weight corresponding to the voltage subsequence: ; Among them, represents the distortion metric weight of the voltage subsequence of user and in the time period, is the total number of voltage subsequences.
[0125] For every two users and in the time period, the distortion metric weight of the voltage subsequence , which is calculated based on the mixed correlation coefficient between voltage subsequences and is used to reflect the degree of voltage distortion.
[0126] The function of step S2 is to calculate the mixed correlation coefficient of the voltage subsequences of every two users in the same time period, and accordingly obtain the distortion metric weight to quantify the degree of voltage distortion.
[0127] Step S3, specifically.
[0128] Calculate the DTW distance of the voltage subsequences corresponding to every two users in the same time period to calculate the correlation between voltage subsequences. As Figure 2 shown in the path diagram during DTW distance calculation; there is a path between two voltage subsequences to make the distance between voltage subsequences the shortest, which can avoid the inaccuracy of correlation calculation caused by clock deviation. The upper solid line and the lower solid line in the figure represent the time series voltage data of two users, the dotted line represents the DTW distance calculated between the user voltage subsequences, the abscissa is the sampling time point, and a point is collected every 15 minutes at the sampling frequency.
[0129] The DTW distance is used to measure the similarity between two voltage subsequences.
[0130] To solve the clock deviation problem in the existing method, calculate the DTW distance between the voltage subsequences of every two users. The DTW distance is a method to calculate the optimal alignment path between two time series through the dynamic programming algorithm. It can effectively handle the stretching and offset of time series on the time axis, thereby quantifying the similarity between two voltage subsequences.
[0131] The DTW distance is calculated as follows: ; Among them, represents the th time period voltage subsequence of the user, th time period voltage subsequence of the user, and the minimum dynamic time warping distance between them. , respectively represent the and th sub - sequence th value of the user. To avoid excessive time deviation, set , where 3 represents 3 sampling time points.
[0132] After obtaining the DTW distance between each voltage subsequence, calculate the weighted DTW distance based on the distortion metric weight and the DTW distance.
[0133] By introducing weights, smaller weights can be assigned to those subsequences that are more affected by voltage distortion, thereby reducing their impact on the overall similarity during the calculation process.
[0134] Calculate the weighted DTW distance based on the distortion metric weight and the DTW distance as follows: ; Among them, is the DTW distance between the user voltage subsequences and ; is the corresponding weighted DTW distance.
[0135] The smaller it is, the higher the correlation.
[0136] Based on the weighted DTW distance, calculate the temporal voltage correlation and th time period corresponding voltage subsequences and of the user as follows: ; ; The function of step S3 is to consider the temporal voltage correlation of voltage distortion and clock deviation, calculate the DTW distance and weighted DTW distance of pairwise user voltage subsequences to measure the similarity of voltage subsequences and solve the clock deviation problem.
[0137] Step S4, specifically.
[0138] Construct the adjacent matrix of user nodes in the substation area based on the time - series voltage correlation of the voltage sub - sequences corresponding to each pair of users. Each user is regarded as one node.
[0139] Based on each pair of users and the time - series voltage correlation of the corresponding voltage sub - sequences construct the adjacent matrix of user nodes in the substation area, including: Preset the correlation threshold γ; If the time - series voltage correlation , then users and have strong correlation. Connect users and into a chain to obtain the adjacent matrix of user nodes , as follows: ; Among them, is the number of users, is the matrix element.
[0140] The correlation threshold takes the value of 0.9.
[0141] , indicating connection; it means that users and user have strong correlation in voltage; , indicating disconnection, meaning that users and user do not have strong correlation in voltage.
[0142] Use the breadth - first search algorithm to traverse the elements in the adjacent matrix , including: Start searching from any user node. The searched users are marked as visited. For the current user , check all its unvisited neighbor user nodes; If , then add the neighbor user to the queue and mark it as visited; Connect each newly visited user to its predecessor visited neighbor user; finally, obtain multiple chain - like structures.
[0143] Construct an adjacent graph, where the nodes are users; the edges are the connection relationships based on the adjacent matrix; The chain - like structure including the distribution transformer is the main chain, and the independent sub - chain users not connected to the main chain are abnormal users.
[0144] Users obtained from the main chain are normal users; Users obtained from independent sub-chains that have no connection with the main chain are abnormal users.
[0145] Output a list of abnormal users, mark them as metering abnormal users, and record the number of the distribution transformer area to which they belong.
[0146] For the list of normal users, output the list of normal users in the main chain.
[0147] As Figure 3 shown, the user terminal is a node in the figure. If the voltage correlation between nodes is greater than the threshold of 0.9, they are connected into a chain with solid black lines; otherwise, they are connected with dotted lines. The disconnected area is the non-main chain area, and the voltage correlation with the main chain users is less than the threshold of 0.9. The users in the disconnected area are the metering abnormal users.
[0148] The function of step S4 is to verify the abnormal metering data according to the voltage correlation chain criterion, and output the list of metering abnormal users in the disconnected area and the list of normal users in the main chain. Through the chain criterion, metering abnormal users can be effectively identified to ensure the accuracy of the distribution transformer area metering data and the stable operation of the power system.
[0149] For the identified users with abnormal metering data, judge whether their time-series voltage data is within the normal range.
[0150] First, according to the standard of voltage over-limit in the power grid, set the upper and lower limit voltage amplitudes as , the standard voltage is 220V. If the limit is exceeded, it is determined that the voltage of the abnormal user is over-limit, as follows: ; Among them, represents the over-limit judgment result of the time-series voltage data of user , 0 means the voltage is not over-limit, and 1 means the voltage is over-limit. represents the th voltage value of user
[0151] If the voltage of the user is instantaneously over-limit (such as a short-term sudden drop caused by lightning), even if it is normal at other times, it will still be marked as over-limit.
[0152] And judge whether the voltage is distorted according to the adjacent point fluctuation detection. Calculate the absolute value of the voltage difference between adjacent points of user : ; Among them, represents user in the time period and The absolute value of the voltage change.
[0153] Calculate the user The average difference of the voltage sequence : ; Among them, represents the number of sequential voltage data, is the average intensity of voltage fluctuations.
[0154] Calculate the standard deviation , as follows: ; The standard deviation , the degree of dispersion of voltage differences, reflects the fluctuation stability.
[0155] Set the multiple threshold , if it satisfies , it is considered that there is voltage distortion. Exemplarily, takes values from 1.5 to 3 and is determined according to specific requirements.
[0156] The voltage distortion judgment is as follows: ; Among them, represents the distortion judgment result. When it is 1, it means that the user's voltage is distorted. When it is 0, it means that the user's voltage fluctuates normally. If the time sequence is not exceeded and the voltage distortion is normal, the user is output as a suspected abnormal household transformation relationship. Otherwise, the abnormal user is output as a suspected abnormal metering data.
[0157] The function of step S4 is to construct the adjacent matrix of user nodes in the substation area based on the time sequence voltage correlation of pairwise user voltage subsequences, identify the main chain and the independent sub-chain users disconnected from the main chain, output the normal and abnormal user lists, and ensure the accuracy of the substation area metering data.
[0158] On-site in the distributed photovoltaic substation area, the power supply operation and maintenance personnel replace the damaged sensors and rewire to eliminate problems such as poor contact based on the output abnormal users. In addition, a comprehensive assessment of the photovoltaic equipment will be carried out to ensure that all components operate according to the specifications. Finally, check the household transformation relationship of the abnormal users to ensure that the on-site and system ledger relationships are consistent.
[0159] Once the on-site work is completed and all known problems are properly solved, the latest voltage time sequence data of the distribution transformer and users will be uploaded back to the central database. Then, the system will use this updated information to execute the metering data self-checking process in the present invention again to verify the effectiveness of the previous rectification measures.
[0160] According to the verification result statistics, the accuracy of a distributed photovoltaic substation area metering data online self-checking method reaches , which means that this method can effectively screen out most abnormal cases, greatly reducing the false alarm rate while improving the timeliness of solving distortion problems. The method in the present invention not only enhances the intelligence level of power grid management, but also provides solid technical support for realizing more green and efficient energy utilization.
[0161] In summary, by combining online self-verification and on-site real-time verification, a closed-loop management system can be constructed to continuously monitor and improve the quality of metering data in photovoltaic substations, thereby promoting the healthy development of the entire power system.
[0162] Embodiment 2: Another embodiment of the present invention discloses an online self-verification system for distributed photovoltaic substation metering data, so as to implement an online self-verification method for distributed photovoltaic substation metering data in Embodiment 1. The specific implementation manners of each module refer to the corresponding descriptions in Embodiment 1.
[0163] As Figure 4 shown, the system includes a data acquisition module M1, a hybrid correlation coefficient calculation module M2, a weighted DTW distance calculation module M3, and a chain structure analysis and verification module M4.
[0164] The data acquisition module M1 is used to acquire distributed voltage metering data of the photovoltaic substation and perform preprocessing to obtain voltage subsequence data of each user; the distribution transformer in the photovoltaic substation is regarded as a user; The hybrid correlation coefficient calculation module M2 is used to calculate the hybrid correlation coefficient of the voltage subsequences corresponding to two users in the same time period; based on the hybrid correlation coefficient, the distortion metric weight corresponding to the voltage subsequences of two users is calculated; The weighted DTW distance calculation module M3 is used to calculate the DTW distance of the voltage subsequences corresponding to two users in the same time period, calculate the weighted DTW distance based on the distortion metric weight and the DTW distance, and calculate the temporal voltage correlation of the voltage subsequences corresponding to two users based on the weighted DTW distance; The chain structure analysis and verification module M4 is used to construct an adjacent matrix of user nodes in the substation based on the temporal voltage correlation of the voltage subsequences corresponding to the two users; traverse the elements in the adjacent matrix of user nodes to obtain multiple chain structures; among them, the chain structure including the distribution transformer is the main chain, and the independent sub-chain users not connected to the main chain are abnormal users.
[0165] In summary, an online self-verification method for distributed photovoltaic substation metering data according to an embodiment of the present invention has the following beneficial effects: 1. The present invention calculates the mixed correlation coefficient of the voltage subsequences of two users and uses it as the distortion measurement weight of the voltage subsequence, so as to more accurately measure the correlation between the voltage sequences of users in the distributed photovoltaic area and effectively improve the accuracy of voltage distortion detection in metering data verification; 2. The present invention uses DTW distance and weighted DTW distance as indicators to measure the consistency of voltage subsequence fluctuations, which ingeniously solves the deviation problem caused by clock asynchrony between voltage subsequences and ensures the accuracy of voltage sequence correlation analysis; 3. The present invention uses chain criteria to chain highly correlated users and identifies users disconnected from the main chain as abnormal users, effectively solving a series of technical problems caused by the long power supply radius and improving the accuracy of metering data abnormality identification; 4. The present invention realizes efficient online self-checking function, can monitor and check the metering data of distributed photovoltaic areas in real time, find abnormalities in time, greatly reduce manual intervention, and improve the efficiency of verification. At the same time, by accurately distinguishing normal users from abnormal users, it helps to quickly deal with abnormal metering data problems and ensure the stable operation of the power system and the quality of power supply; 5. This invention combines the Clayton Copula model, Gumbel Copula model and DTW algorithm to conduct in-depth analysis and mining of metering data. It significantly improves the intelligence level and data analysis capabilities of the smart distribution network, and provides strong support for the efficient management and optimized operation of the power system.
[0166] Those skilled in the art will appreciate that all or part of the processes of the above-mentioned embodiments can be implemented by instructing related hardware through a computer program, and the program can be stored in a computer-readable storage medium, wherein the computer-readable storage medium is a disk, an optical disk, a read-only storage memory, or a random access memory, etc.
[0167] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by any technician familiar with the technical field within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention.
Claims
1. An online self-verification method for metering data of a distributed photovoltaic substation area, characterized in that, Including: Obtain the distributed voltage measurement data of the photovoltaic substation area and perform preprocessing to obtain the voltage subsequence data of each user; the distribution transformer in the photovoltaic substation area is regarded as a user; Calculate the hybrid correlation coefficient of the voltage subsequences corresponding to two users in the same time period; calculate the distortion metric weight of the voltage subsequences corresponding to two users based on the hybrid correlation coefficient; Calculate the DTW distance of the voltage subsequences corresponding to two users in the same time period, calculate the weighted DTW distance based on the distortion metric weight and the DTW distance, and calculate the temporal voltage correlation of the voltage subsequences corresponding to two users based on the weighted DTW distance; Construct the user node adjacency matrix in the substation area based on the temporal voltage correlation of the voltage subsequences corresponding to the two users; Traverse the elements in the user node adjacency matrix to obtain multiple chain structures; among them, the chain structure including the distribution transformer is the main chain, and the independent sub-chain users not connected to the main chain are abnormal users.
2. The method according to claim 1, characterized in that, Two users and of the voltage subsequence corresponding to the time period and The hybrid correlation coefficient is as follows: ; Among them, , , and are the hybrid correlation coefficient, Pearson correlation coefficient, lower tail correlation coefficient, and upper tail correlation coefficient corresponding to the voltage subsequences and respectively, and , , are the corresponding weights of the Pearson correlation coefficient, lower tail correlation coefficient, and upper tail correlation coefficient of the voltage subsequence respectively.
3. The method according to claim 2, characterized in that, Calculate pairwise users and in the voltage subsequence of the time period and The rank of each element in the corresponding voltage subsequence is calculated, and then the rank is converted into a uniformly distributed value; Based on users and Construct a Clayton Copula model using uniformly distributed values, and estimate the parameters of the Clayton Copula model using the lower tail log-likelihood function to maximize the lower tail log-likelihood function and obtain the best parameter estimates Based on calculate the lower tail correlation coefficient ; Based on the user and Construct a Gumbel Copula model based on the uniformly distributed values, and estimate the Gumbel Copula model using the upper tail log-likelihood function to maximize the upper tail log-likelihood function and obtain the optimal parameter estimates Based on calculate the lower tail correlation coefficient .
4. The method according to claim 3, wherein Based on the user and Construct a Clayton Copula model based on the uniformly distributed values, including: Based on the user and Construct the cumulative distribution function of the Clayton Copula model based on the uniformly distributed values; Construct the probability density function of the Clayton Copula model based on the cumulative distribution function; Construct the lower tail log-likelihood function based on the probability density function; Estimate the parameters of the Clayton Copula model by iteratively maximizing the lower tail log-likelihood function using the gradient descent method .
5. The method according to claim 4, wherein The lower tail log-likelihood function is as follows: ; Among them, is the probability density function of the Clayton Copula model; , are respectively the th voltage data in the th time period voltage subsequence of user ; , are respectively the th corresponding uniform distribution values.
6. The method according to claim 5, wherein Optimal Parameter Estimation Values Based on the Clayton Copula Model , calculate the lower tail correlation coefficient , as follows: 。 7. The method according to claim 3, wherein Based on the user and construct a Gumbel Copula model based on the uniformly distributed values, including: Based on the user and Construct the cumulative distribution function of the Gumbel Copula model based on the uniformly distributed values; Construct the probability density function of the Gumbel Copula model based on the cumulative distribution function; Construct the upper tail log-likelihood function based on the probability density function; The Gumbel Copula model parameters are obtained by solving the upper tail log-likelihood function using Newton's method .
8. The method according to claim 7, wherein The lower tail log-likelihood function , is as follows: ; Among them, is the probability density function of the Gumbel Copula model; , are respectively the and th voltage data in the th voltage subsequence of the , are respectively the , corresponding uniformly distributed values.
9. The method according to claim 8, wherein Optimal Parameter Estimation Based on Gumbel Copula , calculate the lower tail correlation coefficient , as follows: 。 10. The method according to claim 3, characterized in that, Convert the ranking to a uniform distribution value as follows: ; Among them, and are respectively and corresponding uniformly distributed values; and are respectively the and th voltage data of the th and in the voltage subsequence and ; is the length of the voltage subsequence.
11. The method according to claim 2, wherein Hybrid correlation coefficient based on voltage subsequence Calculate the distortion measure weight corresponding to the voltage subsequence: ; Among them, represents the user and at the distortion metric weight of the voltage subsequence in the time period, is the total number of voltage subsequences.
12. The method according to claim 11, wherein Based on the distortion metric weight and the DTW distance, calculate the weighted DTW distance as follows: ; Among them, is the DTW distance of the user voltage subsequence and ; is the corresponding weighted DTW distance.
13. The method according to claim 12, characterized in that, Based on the weighted DTW distance, calculate the and for the voltage subsequences corresponding to the time periods and for the temporal voltage correlation , as follows: 。 14. The method according to claim 1, characterized in that, Based on pairwise users and the timing voltage correlation of the corresponding voltage subsequences Construct the adjacency matrix of user nodes in the substation area, including: Preset correlation threshold ; If the temporal voltage correlation , then the users and have a strong correlation. Connect the users and into a chain to obtain the user node adjacency matrix , as follows: ; Among them, is the number of users of which is a matrix element.
15. The method according to claim 14, wherein Traverse the adjacency matrix using the breadth-first search algorithm elements, including: Start the search from any user node, mark the searched users as visited, and for the current user , check all its unvisited neighbor user nodes; If , add the neighbor user to the queue and mark it as visited; Connect each newly accessed user with its previously accessed neighbor users; finally, obtain multiple chain structures.
16. The method according to any one of claims 1 to 15, characterized in that Obtain the users in the main chain as normal users; Obtain the users of the independent sub-chain not connected to the main chain as abnormal users.
17. The method according to any one of claims 1 to 15, characterized in that, Collect the distributed voltage measurement data of the photovoltaic substation area through a preset sampling frequency; Among them, the distributed voltage measurement data includes the user voltage data of multiple distributed users and the voltage data of the low-voltage side of the distribution transformer; the voltage data of the low-voltage side of the distribution transformer is regarded as a user voltage data.
18. The method according to claim 17, wherein Perform preprocessing on the obtained distributed voltage measurement data of the photovoltaic substation area, including: After data deduplication, data correction, continuity verification, outlier processing and missing value filling, obtain the first voltage time series data; Use the sliding window function to divide the first voltage time series data into multiple segments to obtain the voltage subsequence data of each user in multiple time periods.
19. The method according to claim 14, wherein The correlation threshold has a value of 0.
9.
20. A distributed photovoltaic substation metering data online self-checking system, characterized in that, Including: A data acquisition module for obtaining the distributed voltage measurement data of the photovoltaic substation area and performing preprocessing to obtain the voltage subsequence data of each user; the distribution transformer in the photovoltaic substation area is regarded as a user; A hybrid correlation coefficient calculation module for calculating the hybrid correlation coefficient of the voltage subsequences corresponding to two users in the same time period; calculating the distortion metric weight of the voltage subsequences corresponding to two users based on the hybrid correlation coefficient; A weighted DTW distance calculation module for calculating the DTW distance of the voltage subsequences corresponding to two users in the same time period, calculating the weighted DTW distance based on the distortion metric weight and the DTW distance, and calculating the temporal voltage correlation of the voltage subsequences corresponding to two users based on the weighted DTW distance; The chain - structure analysis and verification module is used to construct the adjacent matrix of user nodes in the transformer area based on the temporal voltage correlation of the voltage subsequences corresponding to each pair of users. Traverse the elements in the adjacent matrix of user nodes to obtain multiple chain - structures; among them, the chain - structure including the distribution transformer is the main chain, and the independent sub - chain users without connection to the main chain are abnormal users.
Citation Information
Patent Citations
Transformer area user-transformer relation verification method and system
CN111881189A
Power distribution area user phase identification method and system based on screened voltage data
CN112701675A
Low-voltage transformer area phase household topology identification method based on DDTW distance
CN116995653A
Data-driven user side distributed photovoltaic detection and access capacity estimation method
CN117410964A
Distributed photovoltaic abnormal data score decision-making method based on sequential sequence clustering
CN118277935A