Line loss analysis monitoring method and system based on big data
By constructing a dynamic baseline loss prediction model and a multi-scale morphological perception method using big data analysis, the problem of identifying abnormal power grid line losses was solved, achieving accurate identification and rapid location, and improving the efficiency of power grid operation and maintenance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-15
- Publication Date
- 2026-03-27
AI Technical Summary
Existing technologies cannot accurately identify abnormal power grid line losses, resulting in frequent false alarms, inability to locate the source of the anomaly, and a lack of intelligent diagnostic capabilities, leading to low operation and maintenance efficiency.
By employing a big data-based line loss analysis method, a dynamic baseline line loss prediction model and a multi-scale morphological perception method are constructed. Combined with machine learning and anomaly morphology rule base, this method enables in-depth feature mining and intelligent pre-identification of line loss residuals, thereby locating the true source of line loss anomalies.
It significantly reduced the false alarm and false alarm rates of line loss monitoring, improved the sensitivity and early warning capabilities for anomaly detection, quickly located the source of anomalies, and improved operation and maintenance efficiency.
Smart Images

Figure CN121744136A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of power grid safety monitoring technology, and in particular relates to a method and system for line loss analysis and monitoring based on big data. Background Technology
[0002] In the power industry, line loss rate is a core indicator for measuring the efficiency and management level of power grid operation. Currently, power grid companies generally face the challenge of refined line loss management: on the one hand, with the large-scale integration of distributed energy resources and the increasingly complex characteristics of loads, the deviation between traditional theoretical line loss calculations based on fixed ratios or simple formulas and actual values is increasing; on the other hand, the concealment of electricity theft and the diversification of metering device failures make the accurate identification and location of abnormal line losses increasingly difficult.
[0003] Currently, the industry mainly relies on static threshold methods based on historical averages for line loss monitoring, triggering alarms when the statistical line loss rate exceeds a preset threshold. While this method is simple and easy to implement, it has significant limitations: First, it cannot distinguish between normal line loss fluctuations caused by load fluctuations and environmental changes and genuine anomalies, leading to frequent false alarms; second, alarms can only be located at the transformer substation or line level, failing to pinpoint the source of the anomaly, which greatly hinders on-site troubleshooting; third, it lacks the ability to intelligently identify and diagnose anomaly patterns, making it difficult for maintenance personnel to quickly determine the cause of the anomaly, resulting in low response efficiency.
[0004] Therefore, developing a line loss analysis technology that can adapt to the power grid operating status, accurately identify anomalies, and intelligently diagnose and locate them has become an urgent need for the industry. Summary of the Invention
[0005] The purpose of this invention is to provide a line loss analysis and monitoring method based on big data, which aims to solve the above-mentioned technical problems.
[0006] This invention is implemented as follows: a line loss analysis and monitoring method based on big data, comprising the following steps:
[0007] Historical and real-time multi-source data of the target monitoring area are collected and preprocessed to generate a fused data sequence; the multi-source data includes power grid topology data, power consumption data, load data and environmental data;
[0008] Based on the preset dynamic baseline loss prediction model, the dynamic baseline loss rate is predicted according to the fused data sequence.
[0009] Based on real-time multi-source data, the actual line loss rate is determined, and the residual between the actual line loss rate and the dynamic baseline line loss rate is calculated to generate a time series of line loss residuals.
[0010] Based on the multi-scale morphological perception method, the actual anomalies of line loss are predicted and identified according to the line loss residual time series, and line loss monitoring results are generated.
[0011] Based on the line loss monitoring results, the true source of the line loss anomaly is located according to the line loss residual time series.
[0012] Furthermore, the training method for the dynamic baseline loss prediction model includes the following steps:
[0013] Features strongly correlated with the physical principles of line loss are extracted from the fused data sequence to form feature variables;
[0014] The theoretical line loss rate or the statistical line loss rate are used as training labels to form label variables;
[0015] A machine learning model is selected, and based on a preset loss function, the model is trained according to feature variables and label variables to obtain a preset dynamic baseline loss prediction model.
[0016] Furthermore, the machine learning model is a gradient boosting decision tree.
[0017] Furthermore, the steps of determining the actual line loss rate based on real-time multi-source data, calculating the residual between the actual line loss rate and the dynamic baseline line loss rate, and generating a time series of line loss residuals specifically include:
[0018] The actual line loss rate is determined based on the real-time total power supply and total power consumption.
[0019] The line loss residual is calculated based on the actual line loss rate and the dynamic baseline line loss rate.
[0020] Based on a preset time window, a time series of line loss residuals is generated according to the line loss residuals.
[0021] Furthermore, based on the multi-scale morphological perception method, the steps of predicting and identifying true anomalies in line loss and generating line loss monitoring results according to the line loss residual time series specifically include:
[0022] Based on the wavelet transform method, the time series of line loss residuals is decomposed into components of different scales;
[0023] For the time series of the line loss residual and its components at each scale, a multidimensional morphological feature vector containing statistical features, dynamic features and waveform features is extracted respectively.
[0024] The multidimensional morphological feature vectors are simultaneously input into the isolated forest unsupervised detection model and the predefined anomaly morphology rule base for parallel analysis and matching to obtain anomaly scores and matching results.
[0025] The anomaly score output by the isolated forest model is fused with the matching results of the morphological rule base to generate a macroscopic anomaly alarm with anomaly type prediction and confidence information.
[0026] Furthermore, the method for locating the true source of abnormal line loss includes the following steps:
[0027] Calculate the correlation between the line loss residual time series and the electricity consumption behavior data of each branch line or user in the target monitoring area, and select a candidate set containing the real abnormal sources of line loss from the target monitoring area based on the correlation.
[0028] Another objective of this invention is to provide a big data-based line loss analysis and monitoring system for implementing the aforementioned line loss analysis and monitoring method, comprising:
[0029] The data acquisition and preprocessing module is used to acquire historical and real-time multi-source data of the target monitoring area, and perform preprocessing to generate a fused data sequence; the multi-source data includes power grid topology data, power generation data, load data and environmental data;
[0030] The dynamic benchmark prediction module is used to predict the dynamic benchmark loss rate based on the preset dynamic benchmark loss prediction model and the fused data sequence.
[0031] The line loss residual determination module is used to determine the actual line loss rate based on real-time multi-source data, calculate the residual between the actual line loss rate and the dynamic baseline line loss rate, and generate a line loss residual time series.
[0032] The line loss anomaly identification module is used to predict and identify real line loss anomalies based on the line loss residual time series using a multi-scale morphological perception method, and generate line loss monitoring results.
[0033] The anomaly source localization module is used to locate the true source of line loss anomalies based on the line loss monitoring results and the line loss residual time series.
[0034] Furthermore, the dynamic benchmark prediction module specifically includes:
[0035] The feature variable determination unit is used to extract features that are strongly correlated with the physical principles of line loss from the fused data sequence, and to form feature variables.
[0036] The label variable determination unit is used to construct label variables by using theoretical line loss rate or statistical line loss rate as training labels.
[0037] The model training unit is used to train the model based on a preset loss function, according to feature variables and label variables, to obtain a preset dynamic baseline loss prediction model.
[0038] Furthermore, the line loss residual determination module includes:
[0039] The actual line loss rate determination unit is used to determine the actual line loss rate based on the real-time total power supply and total power consumption.
[0040] The line loss residual calculation unit is used to calculate the line loss residual based on the actual line loss rate and the dynamic baseline line loss rate.
[0041] The time series generation unit is used to generate a time series of line loss residuals based on a preset time window and the line loss residuals.
[0042] Furthermore, the line loss anomaly identification module specifically includes:
[0043] The time series decomposition unit is used to decompose the line loss residual time series into components of different scales based on the wavelet transform method.
[0044] The multidimensional morphological feature extraction unit is used to extract multidimensional morphological feature vectors containing statistical features, dynamic features and waveform features from the time series of line loss residuals and its various scale components.
[0045] An anomaly analysis and matching unit is used to simultaneously input the multi-dimensional morphological feature vector into the isolated forest unsupervised detection model and the predefined anomaly morphological rule base for parallel analysis and matching, and obtain anomaly scores and matching results.
[0046] An anomaly alarm generation unit is used to perform fusion decision-making based on the anomaly score output by the isolated forest model and the matching result of the morphological rule base, and generate macroscopic anomaly alarms with anomaly type prediction and confidence information.
[0047] This invention provides a big data-based line loss analysis and monitoring method. By constructing a dynamic baseline line loss prediction model, it effectively filters out the inherent influence of load and environmental factors on line loss, making the criteria for judging true line loss anomalies more scientific and significantly reducing false alarms and missed alarms. Furthermore, by introducing a multi-scale morphological perception method, this invention achieves the mining of deep features of residual sequences and intelligent pre-identification of anomaly patterns, greatly improving the detection sensitivity and early warning capability for early, slowly changing anomalies. Attached Figure Description
[0048] Figure 1 This is a flowchart illustrating a big data-based line loss analysis and monitoring method provided in an embodiment of the present invention.
[0049] Figure 2 This is a flowchart illustrating the training method for the dynamic baseline loss prediction model provided in an embodiment of the present invention.
[0050] Figure 3This is a flowchart illustrating step S300 in a big data-based line loss analysis and monitoring method provided in an embodiment of the present invention.
[0051] Figure 4 This is a flowchart illustrating step S400 in a big data-based line loss analysis and monitoring method provided in an embodiment of the present invention.
[0052] Figure 5 This is a schematic diagram of a line loss analysis and monitoring system based on big data, provided as an embodiment of the present invention.
[0053] Figure 6 This is a schematic diagram of the structure of the dynamic benchmark prediction module provided in an embodiment of the present invention.
[0054] Figure 7 This is a schematic diagram of the line loss residual determination module provided in an embodiment of the present invention.
[0055] Figure 8 This is a schematic diagram of the line loss anomaly identification module provided in an embodiment of the present invention. Detailed Implementation
[0056] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0057] like Figure 1 As shown, in one embodiment of the present invention, a method for line loss analysis and monitoring based on big data is provided, including the following steps:
[0058] S100. Collect historical and real-time multi-source data of the target monitoring area, perform preprocessing, and generate a fused data sequence; the multi-source data includes power grid topology data, power consumption data, load data, and environmental data;
[0059] S200. Based on the preset dynamic baseline loss prediction model, predict the dynamic baseline loss rate according to the fused data sequence;
[0060] S300. Based on real-time multi-source data, determine the actual line loss rate, calculate the residual between the actual line loss rate and the dynamic baseline line loss rate, and generate a time series of line loss residuals.
[0061] S400. Based on the multi-scale morphological perception method, predict and identify real anomalies in line loss according to the line loss residual time series, and generate line loss monitoring results.
[0062] S500. Based on the line loss monitoring results, locate the true source of the line loss anomaly according to the line loss residual time series.
[0063] In practical applications, the target monitoring area includes, but is not limited to, a distribution transformer area or a feeder. Power grid topology data includes the connection relationships between various branch lines and users, equipment parameters, etc.; power data includes total power supply and power consumption of each user or branch line; load data includes active power, reactive power, current, voltage, etc., of each branch line and user; environmental data includes temperature, humidity, wind speed, rainfall, etc. After data acquisition, preprocessing such as data cleaning and spatiotemporal alignment is required to eliminate noise, handle missing values, correct inconsistencies, and ensure that all data are aligned within a unified timestamp and spatial range, ultimately forming a fused data sequence. Data cleaning includes linear interpolation, smoothing, and outlier removal based on z-scores; spatiotemporal alignment can be achieved through resampling or spatial interpolation (such as inverse distance weighted interpolation) to provide high-quality input for subsequent analysis.
[0064] like Figure 2 As shown, in a preferred embodiment of the present invention, the training method of the dynamic baseline loss prediction model includes the following steps:
[0065] S210. Extract features that are strongly correlated with the physical principles of line loss from the fused data sequence to form feature variables;
[0066] S220. Use theoretical line loss rate or statistical line loss rate as training labels to form label variables;
[0067] S230. Select a machine learning model, based on a preset loss function, and train the model according to the feature variables and label variables to obtain a preset dynamic baseline loss prediction model.
[0068] Specifically, the inputs and outputs of the dynamic baseline line loss prediction model include: feature variables and label variables; wherein, features strongly correlated with the physical principles of line loss (such as active power, reactive power, current, temperature, humidity, wind speed, etc.) can be extracted from the fused data sequence to constitute feature variables; label variables refer to training label variables, which can use theoretical line loss rate or statistical line loss rate after data cleaning as training labels; this embodiment of the invention uses theoretical line loss rate as training label for illustrative purposes; wherein, theoretical line loss rate can be calculated by the following simplified physical formula:
[0069] ;
[0070] In the formula, Y t (t) represents the theoretical line loss rate at time t; R and G are the equivalent resistance and conductance of the line, respectively; w1 and w2 are correction coefficients, which can be determined through model training; P c (t) represents the transformer's fixed losses, such as iron loss, at time t; P tI(t) represents the total power supply at time t; I(t) represents the current at time t; and U(t) represents the voltage at time t.
[0071] Furthermore, the dynamic baseline loss prediction model can be trained using machine learning algorithms such as linear regression and gradient boosting decision trees. In this embodiment of the invention, gradient boosting decision trees are used as an example. Gradient boosting decision trees are an ensemble model that gradually fits the residuals through an additive model and a forward stepwise algorithm. The final output prediction result is the sum of all decision trees, as shown in the following formula:
[0072] ;
[0073] in, The prediction result output by the dynamic baseline loss prediction model is the dynamic baseline loss rate at time t. f is the function for gradient boosting decision trees; X(t) is the feature variable at time t; k is the index of the decision tree; K is the total number of decision trees; f k Let F be the output of the k-th decision tree; F is the function space consisting of all decision trees.
[0074] The training objective of gradient boosting decision trees is to minimize the loss function; the expression for the loss function is as follows:
[0075] ;
[0076] In the formula, L is the loss function of the dynamic baseline loss prediction model; l is the loss function of a single sample, such as mean squared error; Ω is the regularization term of a single decision tree, used to control the complexity of a single decision tree and prevent overfitting.
[0077] In this embodiment of the invention, machine learning algorithms (such as gradient boosting decision trees) are used to learn the complex nonlinear relationship between load, environment and line loss in historical data. The dynamic baseline line loss prediction model constructed by it can adaptively adjust the dynamic baseline line loss rate according to the power grid operating conditions (such as load level, temperature changes, etc.), so that the baseline of line loss is no longer a rigid fixed threshold. This can effectively filter out the inherent influence of load and environmental factors on line loss, making the criteria for judging the true anomalies of line loss more scientific and significantly reducing false alarms and missed alarms.
[0078] like Figure 3 As shown, in a preferred embodiment of the present invention, the step of determining the actual line loss rate based on real-time multi-source data, calculating the residual between the actual line loss rate and the dynamic baseline line loss rate, and generating a time series of line loss residuals, specifically step S300, includes:
[0079] S310. Determine the actual line loss rate based on the real-time total power supply and total power consumption;
[0080] S320. Calculate the line loss residual based on the actual line loss rate and the dynamic baseline line loss rate;
[0081] S330. Based on a preset time window, generate a time series of line loss residuals according to the line loss residuals.
[0082] Specifically, firstly, the actual line loss rate can be determined based on the real-time total power supply and total power consumption, as follows:
[0083] ;
[0084] In the formula, Y a (t) represents the actual line loss rate at time t; C t (t) represents the total electricity consumption at time t;
[0085] Next, based on the actual line loss rate and the predicted dynamic baseline line loss rate mentioned above, the line loss residual can be calculated, as follows:
[0086] ;
[0087] In the formula, Let be the line loss residual at time t;
[0088] Then, within the preset time window, the time series of line loss residuals can be expressed as follows: , where t i Let i represent the i-th time point, i=1,2,…,N, where N is the total number of time points in the time series of line loss residuals.
[0089] In this embodiment of the invention, the line loss residual time series determined by the dynamic baseline line loss rate filters out the influence of inherent factors such as load and environment, so that any significant fluctuations in the line loss residual time series can more accurately reflect non-technical losses (such as electricity theft) or sudden equipment failures (such as metering failure modes), significantly improving the signal-to-noise ratio of real abnormal signals, thereby improving the accuracy of subsequent prediction of real line loss anomalies.
[0090] like Figure 4 As shown, in a preferred embodiment of the present invention, the step of predicting and identifying actual anomalies in line loss based on the multi-scale morphological perception method and generating line loss monitoring results according to the line loss residual time series, namely step S400, specifically includes:
[0091] S410. Based on the wavelet transform method, the time series of line loss residuals is decomposed into components of different scales;
[0092] The components include low-frequency approximate components and high-frequency detail components. Wavelet transform is suitable for non-stationary signal analysis and can effectively reveal the true anomalies of line loss at different resolutions. Specifically, the original line loss residual time series can be decomposed into components of different frequencies by using discrete wavelet transform to capture the anomaly features at different time scales.
[0093] S420. For the line loss residual time series and its components at each scale, extract multi-dimensional morphological feature vectors containing statistical features, dynamic features and waveform features respectively.
[0094] The statistical characteristics include the mean, variance, first difference, skewness, and kurtosis of the line loss residual time series or its components at various scales. The dynamic characteristics include the approximate entropy, trend slope, and autocorrelation function of the line loss residual time series or its components at various scales. Approximate entropy measures the complexity and regularity of the time series; time series with strong regularity (such as periodic electricity theft) have lower approximate entropy values. The trend slope can be determined through linear fitting, reflecting the continuous deterioration or improvement trend of anomalies. The autocorrelation function quantifies the similarity between the time series and itself at different time lags, thus significantly improving the ability to identify complex and persistent anomalies. Waveform characteristics include the rate of crossing the mean and the area of persistently positive residuals. The rate of crossing the mean refers to the number of times the time series crosses its mean, reflecting the frequency of fluctuations. The area of persistently positive residuals refers to the sum of periods during which the line loss residuals are continuously positive, used to quantify the energy of the anomaly.
[0095] S430. Simultaneously input the multidimensional morphological feature vector into the isolated forest unsupervised detection model and the predefined abnormal morphological rule base for parallel analysis and matching to obtain the abnormality score and matching result.
[0096] The isolated forest unsupervised detection model is as follows: multiple isolated trees are constructed, and features and segmentation values are randomly selected for each tree. For any sample, its path length is calculated, and then the anomaly score is obtained by the expected value of the path length. The larger the anomaly score, the greater the probability of an anomaly. Isolated forest unsupervised detection isolates samples by randomly dividing the feature space. Anomalies are isolated more quickly because their features are much different from normal points (the average path length is shorter).
[0097] The anomaly pattern rule base is used to match the current line loss residual time series pattern with typical patterns for anomaly matching; for example, Rule 1 (suspected electricity theft pattern): IF (low approximate entropy AND high persistent positive residual area) THEN suspected electricity theft; Rule 2 (metering fault pattern): IF (single extreme value in detail component AND excessively large first-order difference in subsequent sequence AND significantly increased skewness) THEN suspected metering fault. In addition, it can also correlate equipment events with environmental data for anomaly matching, such as line insulation aging, tree obstruction, and line dampness.
[0098] S440. Based on the anomaly score output by the isolated forest model and the matching result of the morphological rule base, a fusion decision is made to generate a macroscopic anomaly alarm with anomaly type prediction and confidence information.
[0099] Specifically, if the anomaly score output by the isolated forest model is greater than the preset threshold, or if an abnormal pattern is triggered in the matching results of the morphological rule base, a macro-anomaly alarm is generated. The confidence level of each anomaly type prediction can be obtained by first calculating the statistical probability that the anomaly type is predicted to be true (considered as the probability that the anomaly type is predicted to be true under given evidence, which can be obtained using Bayes' theorem or logistic regression), and then taking a weighted average with the anomaly score.
[0100] In this embodiment of the invention, the time series of line loss residuals is decomposed by wavelet transform, which can simultaneously capture short-term abrupt changes (such as instantaneous damage to metering equipment) and long-term slow changes (such as slow electricity theft or insulation aging), solving the problem that single-scale analysis is insensitive to weak and slow anomalies. In addition, by extracting morphological features and matching them with a rule base, this embodiment of the invention can provide preliminary anomaly predictions during line loss monitoring and issue early warnings before the actual line loss rate exceeds the baseline.
[0101] In a preferred embodiment of the present invention, the method for locating the true source of abnormal line loss includes the following steps:
[0102] Calculate the correlation between the line loss residual time series and the electricity consumption behavior data of each branch line or user in the target monitoring area, and select a candidate set containing the real abnormal sources of line loss from the target monitoring area based on the correlation.
[0103] In practical applications, if a macroscopic anomaly alarm is triggered in the line loss monitoring results, the correlation between the line loss residual time series and the electricity consumption behavior data of each branch line or user in the target monitoring area is calculated to screen out a candidate set containing the true source of the line loss anomaly. Specifically, by calculating the Pearson correlation coefficient between the line loss residual and the rate of change of electricity consumption of each branch line or user in the target monitoring area during the abnormal period in the line loss residual time series, branch lines or users whose absolute value of the Pearson correlation coefficient is greater than a preset threshold are screened out (considered as high-suspicion objects), thereby determining the candidate set containing the true source of the line loss anomaly.
[0104] In this embodiment of the invention, by calculating the correlation between line loss residual and the electricity consumption behavior of a large number of branch lines or users, a few highly suspicious objects can be quickly screened out, narrowing the scope of investigation from the entire transformer area or line to specific branch lines or users, saving a lot of manpower and time for on-site investigation.
[0105] like Figure 5As shown, in another embodiment of the present invention, a line loss analysis and monitoring system based on big data is also provided to implement the above-mentioned line loss analysis and monitoring method, specifically including:
[0106] The data acquisition and preprocessing module 10 is used to acquire historical and real-time multi-source data of the target monitoring area, and perform preprocessing to generate a fused data sequence; the multi-source data includes power grid topology data, power data, load data and environmental data;
[0107] The dynamic benchmark prediction module 20 is used to predict the dynamic benchmark loss rate based on the preset dynamic benchmark loss prediction model and the fused data sequence.
[0108] The line loss residual determination module 30 is used to determine the actual line loss rate based on real-time multi-source data, calculate the residual between the actual line loss rate and the dynamic baseline line loss rate, and generate a line loss residual time series.
[0109] The line loss anomaly identification module 40 is used to predict and identify real line loss anomalies based on the line loss residual time series using the multi-scale morphological perception method, and generate line loss monitoring results.
[0110] The anomaly source location module 50 is used to locate the true source of line loss anomalies based on the line loss monitoring results and the line loss residual time series.
[0111] like Figure 6 As shown, in a preferred embodiment of the present invention, the dynamic benchmark prediction module 20 specifically includes:
[0112] The feature variable determination unit 21 is used to extract features that are strongly related to the physical principle of line loss from the fused data sequence to form feature variables;
[0113] Label variable determination unit 22 is used to use theoretical line loss rate or statistical line loss rate as training labels to form label variables;
[0114] Model training unit 23 is used to train the model based on a preset loss function, according to feature variables and label variables, to obtain a preset dynamic baseline loss prediction model.
[0115] like Figure 7 As shown, in a preferred embodiment of the present invention, the line loss residual determination module 30 includes:
[0116] The actual line loss rate determination unit 31 is used to determine the actual line loss rate based on the real-time total power supply and total power consumption.
[0117] The line loss residual calculation unit 32 is used to calculate the line loss residual based on the actual line loss rate and the dynamic baseline line loss rate.
[0118] The time series generation unit 33 is used to generate a time series of line loss residuals based on a preset time window and the line loss residuals.
[0119] like Figure 8 As shown, in a preferred embodiment of the present invention, the line loss anomaly identification module 40 specifically includes:
[0120] The time series decomposition unit 41 is used to decompose the line loss residual time series into components of different scales based on the wavelet transform method.
[0121] The multidimensional morphological feature extraction unit 42 is used to extract multidimensional morphological feature vectors containing statistical features, dynamic features and waveform features from the line loss residual time series and its various scale components.
[0122] Anomaly analysis and matching unit 43 is used to simultaneously input the multi-dimensional morphological feature vector into the isolated forest unsupervised detection model and the predefined anomaly morphology rule base for parallel analysis and matching, and obtain anomaly score and matching result;
[0123] The anomaly alarm generation unit 44 is used to perform fusion decision based on the anomaly score output by the isolated forest model and the matching result of the morphological rule base to generate a macroscopic anomaly alarm with anomaly type prediction and confidence information.
[0124] It should be noted that the above modules and units can be implemented as a computer program, which can run on a computer device. The computer device's memory can store the computer program that makes up the modules, enabling the processor to execute the various steps of the above method.
[0125] It should be understood that although the steps in the flowcharts of the various embodiments of the present invention are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the various embodiments may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least a portion of the sub-steps or stages of other steps.
[0126] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory.
[0127] The above embodiments merely illustrate several implementation methods of the present invention, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this patent should be determined by the appended claims.
Claims
1. A method for line loss analysis and monitoring based on big data, characterized in that, Includes the following steps: Historical and real-time multi-source data of the target monitoring area are collected and preprocessed to generate a fused data sequence; the multi-source data includes power grid topology data, power consumption data, load data and environmental data; Based on the preset dynamic baseline loss prediction model, the dynamic baseline loss rate is predicted according to the fused data sequence. Based on real-time multi-source data, the actual line loss rate is determined, and the residual between the actual line loss rate and the dynamic baseline line loss rate is calculated to generate a time series of line loss residuals. Based on the multi-scale morphological perception method, the actual anomalies of line loss are predicted and identified according to the line loss residual time series, and line loss monitoring results are generated. Based on the line loss monitoring results, the true source of the line loss anomaly is located according to the line loss residual time series.
2. The line loss analysis and monitoring method based on big data according to claim 1, characterized in that, The training method for the dynamic baseline loss prediction model includes the following steps: Features strongly correlated with the physical principles of line loss are extracted from the fused data sequence to form feature variables; The theoretical line loss rate or the statistical line loss rate are used as training labels to form label variables; A machine learning model is selected, and based on a preset loss function, the model is trained according to feature variables and label variables to obtain a preset dynamic baseline loss prediction model.
3. The line loss analysis and monitoring method based on big data according to claim 2, characterized in that, The machine learning model is a gradient boosting decision tree.
4. The line loss analysis and monitoring method based on big data according to claim 1, characterized in that, The steps of determining the actual line loss rate based on real-time multi-source data, calculating the residual between the actual line loss rate and the dynamic baseline line loss rate, and generating a time series of line loss residuals specifically include: The actual line loss rate is determined based on the real-time total power supply and total power consumption. The line loss residual is calculated based on the actual line loss rate and the dynamic baseline line loss rate. Based on a preset time window, a time series of line loss residuals is generated according to the line loss residuals.
5. The line loss analysis and monitoring method based on big data according to claim 1, characterized in that, Based on the multi-scale morphological perception method, the steps of predicting and identifying actual anomalies in line loss and generating line loss monitoring results according to the line loss residual time series specifically include: Based on the wavelet transform method, the time series of line loss residuals is decomposed into components of different scales; For the time series of the line loss residual and its components at each scale, a multidimensional morphological feature vector containing statistical features, dynamic features and waveform features is extracted respectively. The multidimensional morphological feature vectors are simultaneously input into the isolated forest unsupervised detection model and the predefined anomaly morphology rule base for parallel analysis and matching to obtain anomaly scores and matching results. The anomaly score output by the isolated forest model is fused with the matching results of the morphological rule base to generate a macroscopic anomaly alarm with anomaly type prediction and confidence information.
6. The line loss analysis and monitoring method based on big data according to claim 1, characterized in that, The method for locating the true source of abnormal line loss includes the following steps: Calculate the correlation between the line loss residual time series and the electricity consumption behavior data of each branch line or user in the target monitoring area, and select a candidate set containing the real abnormal sources of line loss from the target monitoring area based on the correlation.
7. A big data-based line loss analysis and monitoring system, used to implement the line loss analysis and monitoring method according to any one of claims 1-6, characterized in that, include: The data acquisition and preprocessing module is used to acquire historical and real-time multi-source data of the target monitoring area, and perform preprocessing to generate a fused data sequence; the multi-source data includes power grid topology data, power generation data, load data and environmental data; The dynamic benchmark prediction module is used to predict the dynamic benchmark loss rate based on the preset dynamic benchmark loss prediction model and the fused data sequence. The line loss residual determination module is used to determine the actual line loss rate based on real-time multi-source data, calculate the residual between the actual line loss rate and the dynamic baseline line loss rate, and generate a line loss residual time series. The line loss anomaly identification module is used to predict and identify real line loss anomalies based on the line loss residual time series using a multi-scale morphological perception method, and generate line loss monitoring results. The anomaly source localization module is used to locate the true source of line loss anomalies based on the line loss monitoring results and the line loss residual time series.
8. The line loss analysis and monitoring system based on big data according to claim 7, characterized in that, The dynamic benchmark prediction module specifically includes: The feature variable determination unit is used to extract features that are strongly correlated with the physical principles of line loss from the fused data sequence, and to form feature variables. The label variable determination unit is used to construct label variables by using theoretical line loss rate or statistical line loss rate as training labels. The model training unit is used to train the model based on a preset loss function, according to feature variables and label variables, to obtain a preset dynamic baseline loss prediction model.
9. The line loss analysis and monitoring system based on big data according to claim 7, characterized in that, The line loss residual determination module includes: The actual line loss rate determination unit is used to determine the actual line loss rate based on the real-time total power supply and total power consumption. The line loss residual calculation unit is used to calculate the line loss residual based on the actual line loss rate and the dynamic baseline line loss rate. The time series generation unit is used to generate a time series of line loss residuals based on a preset time window and the line loss residuals.
10. The line loss analysis and monitoring system based on big data according to claim 7, characterized in that, The line loss anomaly identification module specifically includes: The time series decomposition unit is used to decompose the line loss residual time series into components of different scales based on the wavelet transform method. The multidimensional morphological feature extraction unit is used to extract multidimensional morphological feature vectors containing statistical features, dynamic features and waveform features from the time series of line loss residuals and its various scale components. An anomaly analysis and matching unit is used to simultaneously input the multi-dimensional morphological feature vector into the isolated forest unsupervised detection model and the predefined anomaly morphological rule base for parallel analysis and matching, and obtain anomaly scores and matching results. An anomaly alarm generation unit is used to perform fusion decision-making based on the anomaly score output by the isolated forest model and the matching result of the morphological rule base, and generate macroscopic anomaly alarms with anomaly type prediction and confidence information.
Citation Information
Patent Citations
Line loss detection method and system based on big data
CN117668772A
Power grid abnormal electricity utilization user detection method and system
CN118779809A
Line loss rate prediction method based on gradient boosting tree
CN119249099A
Power transmission line intelligent management and control method and system based on power distribution data feedback
CN120087566A
Electric energy meter abnormity identification method and device based on LSTM model
CN120744474A