Line loss diagnosis method and device based on multi-source data, medium and program product
By dividing the low-voltage distribution network into topological units and analyzing its abnormal characteristics, and combining this with a multi-source data-based line loss diagnosis method, the problem of low efficiency in low-voltage distribution network line loss diagnosis has been solved. This has enabled accurate detection of line loss anomalies and tracing of fault root causes, thereby improving operation and maintenance efficiency and accuracy.
Patent Information
- Application Number
- CN202511618226.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-06
- Publication Date
- 2026-02-06
AI Technical Summary
In existing technologies, the efficiency of line loss diagnosis in low-voltage distribution networks is low, and maintenance personnel need to conduct large-scale inspections of the entire distribution area, which reduces the efficiency of line loss diagnosis.
The line loss diagnosis method based on multi-source data divides the transformer area into topological calculation units, establishes a theoretical line loss-load characteristic curve model for each unit, and uses a prediction model to obtain the load curve for future periods. It then performs real-time line loss rate comparison and anomaly detection, and combines spatial clustering analysis of abnormal feature vectors and Granger causality test to identify line loss anomalies.
It enables accurate diagnosis and anomaly location of transformer area line loss, improves the efficiency and accuracy of line loss anomaly detection, can identify complex multi-point concurrent anomalies and trace the root cause of coordinated anomaly events, and improves the comprehensiveness and accuracy of fault diagnosis.
Smart Images

Figure CN121479599A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of digital operation and maintenance technology for smart distribution networks, and in particular to line loss diagnosis methods, diagnostic equipment, media, and program products based on multi-source data. Background Technology
[0002] As the "last mile" of the power system's transmission of electricity to end users, the low-voltage distribution network's economic efficiency and reliability directly impact the profitability of power grid companies and the quality of electricity supply for users. Line loss, as one of the core indicators for measuring the economic efficiency of low-voltage distribution network operation, is a crucial link in the power sector's efforts to reduce costs, increase efficiency, and improve operation and maintenance levels through accurate monitoring, analysis, and management.
[0003] In related technologies, multi-dimensional data such as historical total load and meteorological data of low-voltage distribution areas can be collected. Time series forecasting models (such as deep learning algorithms) can then be used to train and predict the total load curve of the distribution area over a future period. Subsequently, based on the predicted total load, the system calculates the expected bus loss value for the entire distribution area during the corresponding time period and uses this as a dynamic benchmark. By comparing the actual monitored bus loss of the distribution area with this predicted dynamic benchmark, high-loss anomalies can be identified and alerted.
[0004] However, alarm information can only reflect the overall abnormality. In related technologies, after receiving an alarm, maintenance personnel need to conduct a large-scale investigation of the entire transformer area, which reduces the efficiency of line loss diagnosis. Summary of the Invention
[0005] This application provides a method, diagnostic equipment, media, and program product for diagnosing line loss based on multi-source data, which can improve the efficiency of line loss diagnosis.
[0006] Firstly, this application provides a method for diagnosing line loss in transformer substations based on multi-source data fusion, applied to transformer substation unit fusion equipment. The method includes: dividing the target transformer substation into units based on preset division rules and a digital topology model, using key branch nodes, the main meter of the substation, and user smart meters as boundaries, to obtain multiple topology calculation units. The digital topology model includes the hierarchical connection relationships and physical line parameters between transformers, branch lines, nodes, and user meters within the target substation; performing regression fitting calculations on the historical line loss rate and historical active load of these multiple topology calculation units to obtain multiple theoretical line loss-load characteristic curve models corresponding to each topology calculation unit; and within a preset future time period, based on a preset transformer substation... The line loss prediction model acquires the total load curve of the target transformer area for a preset future time period and multiple unit load curves corresponding to multiple topology calculation units. It then compares the theoretical line loss rate corresponding to the total load curve with the real-time line loss rate to determine the real-time difference. If the real-time difference reaches a preset threshold, the multiple unit load curves are mapped to multiple theoretical line loss-load characteristic curve models of the unit to obtain multiple theoretical line loss rates. Within the preset future time period, the real-time line loss rate curves of the multiple topology calculation units are compared with the multiple theoretical line loss rates to determine if a preset deviation characteristic exists. If the preset deviation characteristic exists, the topology calculation unit exhibiting the preset deviation characteristic is identified as an abnormal unit.
[0007] By adopting the above technical solutions, this application achieves accurate diagnosis and anomaly location for transformer substation line losses. Topology calculation unit partitioning decomposes the complex transformer substation network into manageable basic units, making line loss analysis more refined. The unit theoretical line loss-load characteristic curve model established by regression fitting provides a scientific benchmark for anomaly detection. The theoretical line loss rate is obtained by mapping the future period load curve obtained from the prediction model and compared with the real-time line loss rate, forming an anomaly detection process from macro to micro. Deviation feature identification completes the accurate location of abnormal units. This achieves refined management from the overall transformer substation to specific units, elevating line loss management from post-event statistics to a proactive model of pre-event prediction and in-event intervention, improving the efficiency and accuracy of line loss anomaly detection.
[0008] In conjunction with some embodiments of the first aspect, in some embodiments, after the step of determining the topology computing unit with the preset deviation feature as an abnormal unit when the preset deviation feature is determined to exist, the method further includes: selecting at least one reference unit from the plurality of topology computing units that meets preset similarity conditions with the abnormal unit; obtaining the real-time line loss rate of the reference unit for the abnormal unit and the at least one reference unit in the same time period, and calculating the difference curve between the real-time line loss rate of the abnormal unit and the real-time line loss rate of the at least one reference unit; when the total duration of the difference curve exceeding the preset lateral comparison threshold exceeds the preset duration, confirming that the abnormal unit has a line loss anomaly and determining the type of the line loss anomaly based on a preset diagnostic rule base.
[0009] By adopting the above technical solution, this application further improves the accuracy and reliability of line loss anomaly diagnosis. The introduction of a horizontal comparison mechanism using reference units provides a more objective benchmark for anomaly judgment. The system selects reference units with similar characteristics to the anomaly unit from the topology calculation units; these units should exhibit similar line loss characteristics under normal conditions. The line loss rate difference curve between the anomaly unit and the reference unit is calculated, and horizontal comparison thresholds and duration requirements are set, effectively filtering out the influence of short-term fluctuations and random errors. Combined with a preset diagnostic rule base, the anomaly type is accurately determined, not only confirming the existence of the anomaly but also identifying the specific anomaly type. This multi-dimensional cross-validation method reduces the false alarm rate, improves the credibility of diagnostic conclusions, and enhances operational efficiency.
[0010] In conjunction with some embodiments of the first aspect, in some embodiments, after determining the topology calculation unit with the preset deviation feature as an abnormal unit when the preset deviation feature is determined to exist, the method further includes: when it is determined that the number of abnormal units is not unique, based on the deviation time series of the abnormal units in the preset future time period, performing spatial clustering analysis on the abnormal feature vector of the deviation time series to obtain at least one abnormal unit cluster, wherein the deviation time series refers to the difference sequence between the unit real-time line loss rate and the unit theoretical line loss rate of the abnormal unit, and multiple abnormal units located in the same abnormal unit cluster have similar abnormal feature vectors; calculating the sum of the minimum electrical distances of all abnormal units in the abnormal unit cluster in the digital topology model of the distribution area; determining whether the sum of the minimum electrical distances is greater than a preset topology proximity threshold; and, if it is determined that the sum of the minimum electrical distances is greater than the preset topology proximity threshold, diagnosing the abnormal units of the abnormal unit cluster as collaborative abnormal events.
[0011] By adopting the above technical solution, this application achieves intelligent identification and classification of complex multi-point concurrent anomalies. When the system detects multiple anomalous units, it performs spatial clustering analysis on the anomalous feature vectors of the deviation time series, grouping anomalous units with similar behavioral patterns into a cluster. This data-driven clustering method can automatically discover potential correlation patterns. By calculating the sum of the minimum electrical distances of all units within the anomalous unit cluster in the topology model and comparing it with a preset threshold, the system can distinguish between localized concentrated faults and wide-area dispersed faults. When anomalous units with highly consistent behavior but dispersed topological locations are detected, the system diagnoses them as collaborative anomalous events. This analysis method, which combines behavioral similarity and topological distribution characteristics, breaks through the limitations of traditional single-point anomaly diagnosis, providing deeper insights and more accurate root cause analysis for multi-source anomalies in complex power grid environments, and improving the comprehensiveness and accuracy of fault diagnosis.
[0012] In conjunction with some embodiments of the first aspect, in some embodiments, before performing spatial clustering analysis on the abnormal feature vector of the deviation time series, the method further includes: performing wavelet transform on the deviation time series to obtain high-frequency wavelet energy coefficients and low-frequency wavelet energy coefficients at different time scales; determining the hierarchical depth of the abnormal unit in the digital topology model of the distribution area and the number of downstream connected users as the topological importance index of the abnormal unit; calculating the non-Gaussian statistics of the deviation time series, and combining the high-frequency wavelet energy coefficients, the low-frequency wavelet energy coefficients, the non-Gaussian statistics, and the topological importance index into the abnormal feature vector.
[0013] By adopting the above technical solutions, this application significantly improves the representational ability and diagnostic accuracy of anomaly feature vectors. By performing wavelet transform on the deviation time series, the system can capture the characteristics of anomalous signals at different time scales. High-frequency wavelet energy coefficients reflect short-term fluctuation characteristics, while low-frequency wavelet energy coefficients characterize long-term trend changes. Introducing a topological importance index ensures that the anomaly feature vector not only includes the statistical characteristics of the time series but also incorporates information on the unit's location and influence range within the power grid structure. Calculating non-Gaussian statistics effectively captures the nonlinear characteristics of anomalous behavior. Combining these multidimensional features into a comprehensive anomaly feature vector provides richer and more discriminative input for spatial clustering analysis. This makes the clustering results more accurate, the anomaly pattern classification more reasonable, lays a solid data foundation for precise diagnosis, and improves the system's adaptability to complex anomaly patterns.
[0014] In conjunction with some embodiments of the first aspect, in some embodiments, the method further includes: examining the load change among anomalous units within the anomalous unit cluster using Granger causality to obtain the power transmission timing correlation; constructing a directed acyclic graph by using anomalous units within the anomalous unit cluster as nodes and determining the power transmission timing correlation between nodes as edges; identifying the power disturbance propagation path from the root node to other nodes using a maximum weight path search algorithm, and determining the starting node of the power disturbance propagation path as the root anomalous unit of the cooperative anomalous event, wherein the root node refers to a node without any incoming edges in the directed acyclic graph; and performing the step of confirming that the anomalous unit has a line loss anomaly and determining the type of the line loss anomaly based on a preset diagnostic rule base, using the root anomalous unit as the updated anomalous unit.
[0015] By adopting the above technical solutions, this application achieves accurate tracing and location of the root causes of coordinated anomalies. Granger causality testing analyzes load changes between anomaly units, enabling the system to identify the timing correlation of power transmission between units and reveal the propagation path and sequence of disturbances. Based on these timing correlations, a directed acyclic graph (DAG) is constructed, networking the complex causal relationships and making the structure of disturbance propagation clearly visible. The maximum weighted path search algorithm can identify the main disturbance propagation paths from the root node to other nodes in the DAG. This allows for accurate identification of the source and propagation mechanism of problems in complex scenarios of multi-point coordinated anomalies, improving the targeting and efficiency of fault handling and providing more precise decision support for power grid operation and maintenance personnel.
[0016] In conjunction with some embodiments of the first aspect, in some embodiments, the step of dividing the target transformer area into units based on preset partitioning rules and a digital topology model of the transformer area, using the key branch nodes, the main meter of the transformer area, and the user smart meters as boundaries, to obtain multiple topology calculation units, specifically includes: extracting the topology diagram of the target transformer area from the digital topology model of the transformer area, where the nodes of the topology diagram represent transformers, branch points, and user meters, and the edges represent connecting lines; dividing the topology diagram into candidate topology unit sets by using a spectral clustering algorithm based on the electrical distance and load correlation between nodes; calculating the cohesion and boundary complexity of each unit in the candidate topology unit set, and adjusting candidate topology units with cohesion below a first threshold or boundary complexity above a second threshold based on preset unit balance constraints, where the boundary complexity represents the number of connections between the unit and other units, and the unit balance constraints include the balance of the number of nodes within the unit, the load distribution balance between units, and the coverage integrity of key nodes; matching the adjusted candidate topology unit set with the locations of the key branch nodes, the main meter of the transformer area, and the user smart meters of the target transformer area to generate a topology calculation unit partitioning scheme, thereby obtaining multiple topology calculation units.
[0017] By adopting the above technical solution, this application achieves a scientific and reasonable division of transformer substation topology units, laying a solid foundation for subsequent accurate line loss analysis. A topology diagram is extracted from the digital topology model of the transformer substation, clearly expressing the complex power grid structure in the form of nodes and edges. Spectral clustering is performed based on the electrical distance and load correlation between nodes, and preliminary network segmentation is achieved using advanced methods of graph theory and machine learning. The cohesion and boundary complexity of candidate units are calculated and adjusted based on preset balance constraints to ensure the rationality and practicality of the division results. The adjusted candidate units are matched with the actual locations of key nodes to generate the final topology calculation unit division scheme. The comprehensive division method combining electrical characteristics, load characteristics, and physical layout not only considers the physical connection relationship of the power grid but also takes into account the actual situation of load distribution and monitoring point layout, ensuring that the divided topology calculation units have clear electrical boundaries and good observability and controllability.
[0018] In conjunction with some embodiments of the first aspect, in some embodiments, the step of obtaining a candidate topological unit set by segmenting nodes based on the electrical distance and load correlation between nodes in the topological graph using a spectral clustering algorithm specifically includes: constructing an affinity matrix between nodes based on the electrical distance and load correlation between nodes in the topological graph, wherein the matrix elements represent the electrical connection strength between nodes, which is calculated by a weighted combination of the reciprocal of the electrical distance and the load correlation coefficient; calculating the Laplace matrix based on the affinity matrix, and solving for the eigenvalues and eigenvectors of the Laplace matrix; determining the target eigenvectors corresponding to the first K smallest non-zero eigenvalues of the Laplace matrix, and constructing a feature space based on the target eigenvectors, wherein the K value is determined by the steep descent point of the eigenvalue distribution; performing K-means clustering on the nodes in the feature space to obtain initial clustering results and silhouette coefficients; when the silhouette coefficient is lower than a preset quality threshold, adjusting the K value and performing K-means clustering on the nodes in the feature space until the silhouette coefficient is higher than the preset quality threshold or the maximum number of iterations is reached; mapping the final clustering results back to the original topological graph to form a candidate topological unit set.
[0019] By adopting the above technical solutions, this application achieves high-quality topological unit clustering, ensuring the scientific validity and effectiveness of transformer substation division. Constructing an affinity matrix between nodes organically integrates the two key factors of electrical distance and load correlation, making the basis for clustering more comprehensive. Based on the affinity matrix, the Laplace matrix is calculated, and its eigenvalues and eigenvectors are solved. Utilizing the mathematical foundation of spectral clustering, the complex network structure is mapped to a low-dimensional feature space. The optimal number of clusters K is adaptively determined by the steep descent point of the eigenvalue distribution, avoiding subjectivity caused by manual setting. The silhouette coefficient is introduced as an evaluation index for clustering quality, and iterative optimization ensures that the clustering results achieve the expected quality. The adaptive partitioning method based on spectral clustering fully utilizes the global structural information of the graph, can identify natural clusters with non-convex shapes, ensuring the reliability and stability of the partitioning results, and providing data support for subsequent line loss analysis and anomaly diagnosis.
[0020] In a second aspect, this application provides a diagnostic device comprising: one or more processors and a memory; the memory being coupled to the one or more processors, the memory being used to store computer program code including computer instructions, the one or more processors invoking the computer instructions to cause the diagnostic device to perform the method described in the first aspect and any possible implementation thereof.
[0021] Thirdly, this application provides a computer program product containing instructions that, when run on a diagnostic device, cause the diagnostic device to perform the method described in the first aspect and any possible implementation thereof.
[0022] Fourthly, this application provides a computer-readable storage medium including instructions that, when executed on a diagnostic device, cause the diagnostic device to perform the method described in the first aspect and any possible implementation thereof.
[0023] One or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages: 1. By adopting the above technical solution, due to the use of a multi-level line loss analysis architecture based on topology unit division, and a theoretical line loss-load characteristic curve modeling and real-time comparison mechanism, the problem of management lag in traditional line loss management is effectively solved. This enables refined management from the overall transformer area to specific units, and elevates line loss management from post-event statistics to a proactive mode of pre-event prediction and in-event intervention, thereby improving the detection efficiency and accuracy of line loss anomalies.
[0024] 2. By adopting the above technical solution, and by using a collaborative anomaly identification method based on spatial clustering analysis of anomaly feature vectors and topological electrical distance assessment, the technical problem of distinguishing and associating multiple concurrent anomalies is effectively solved. This enables intelligent identification and classification of complex multi-point concurrent anomalies, accurately distinguishing between localized concentrated faults and wide-area distributed faults. In particular, the ability to identify systemic problems is significantly improved, breaking through the limitations of traditional single-point anomaly diagnosis. This provides a deeper insight and more accurate root cause analysis for multi-source anomalies in complex power grid environments, improving the comprehensiveness and accuracy of fault diagnosis.
[0025] 3. By adopting the above technical solution, and by using power transmission timing analysis based on Granger causality and the maximum weight path search algorithm for directed acyclic graphs, the technical problem of difficulty in tracing the root cause of coordinated abnormal events is effectively solved. This enables accurate tracing and location of the root cause of coordinated abnormal events, reveals the propagation path and sequence of disturbances, accurately identifies the source and propagation mechanism of the problem, improves the pertinence and efficiency of fault handling, provides more accurate decision support for power grid operation and maintenance personnel, and reduces the processing time and resource consumption of complex faults. Attached Figure Description
[0026] Figure 1 This is a flowchart illustrating a line loss diagnosis method based on multi-source data in an embodiment of this application. Figure 2 This is another flowchart illustrating the line loss diagnosis method based on multi-source data in the embodiments of this application; Figure 3 This is a schematic diagram of an exemplary hardware structure of a diagnostic device in an embodiment of this application. Detailed Implementation
[0027] The terminology used in the following embodiments of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application. As used in the specification and appended claims of this application, the singular expressions “a,” “an,” “the,” “the,” “the,” and “this” are intended to include the plural expressions as well, unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in this application refers to and includes any or all possible combinations of one or more of the listed items.
[0028] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as implying or suggesting relative importance or implicitly indicating the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature, and in the description of the embodiments of this application, unless otherwise stated, "multiple" means two or more.
[0029] Please see Figure 1 This is a flowchart illustrating a line loss diagnosis method based on multi-source data in an embodiment of this application.
[0030] S101. Using the key branch nodes, the main meter of the target transformer area, and the user's smart meter as boundaries, the target transformer area is divided into units based on preset division rules and the digital topology model of the transformer area, resulting in multiple topology calculation units.
[0031] In this context, the target distribution area refers to the specific power supply area targeted in this unit division operation, typically referring to the low-voltage power grid range powered by a single distribution transformer, such as a residential community or a commercial building; a critical branch node refers to a key connection point in the power grid topology that connects multiple lines or branches out into multiple user circuits, such as a cable branch box or junction box; it is the physical location where a line branches into multiple lines; the distribution area master meter is used to represent the electricity meter installed at the low-voltage side outlet of the transformer, used to measure the total electricity consumption of the entire distribution area, and is the only entry point for electricity into this distribution area; the user smart meter refers to an intelligent electricity meter installed at the end user, with data acquisition, two-way communication, and remote control functions, and is the final metering point for electricity consumption; the distribution area digital topology model refers to a virtual network model constructed in a computer in the form of data, representing the structure of the physical power grid, including transformers, lines, switches, nodes, and meters, and their connection relationships; and the topology calculation unit represents the smallest, independent calculation and analysis area divided based on the above boundaries and rules, with each unit having a clearly defined electrical entry and exit point.
[0032] After system initialization or changes to the power grid structure (such as adding users or upgrading lines), the device retrieves the digital topology model of the designated target distribution area from its storage system or the upstream system. The model depicts all electrical connection paths starting from the distribution area's main meter, passing through various levels of lines, switches, and branch nodes, ultimately reaching each user's smart meter. The device then initiates a traversal program, scanning the entire topology model and automatically identifying and marking three types of key boundary points based on built-in device type identifiers: the unique distribution area main meter (as the root node of the topology), all key branch nodes (as intermediate nodes of the topology), and all end-user smart meters (as leaf nodes of the topology). After identification, the device performs a "cutting" operation according to preset partitioning rules. For example, the rule might be defined as follows: starting from the distribution area main meter, following the line path to the first key branch node encountered, this path (including cables, connectors, etc.) is defined as the first topology calculation unit; starting from this key branch node, following each branch path until the next key branch node or a user smart meter encountered, these paths constitute new, independent topology calculation units. This process proceeds recursively until every branchless path in the entire distribution area, from the main meter to all user meters, is divided and assigned to a unique topology calculation unit. Finally, a list of topology calculation units is generated and stored. Each unit contains its starting boundary point, ending boundary point, a list of electrical devices within the unit, and their electrical parameters, thus decomposing a complex, large power grid structure into a series of sub-units that are easy to analyze and calculate independently. Understandably, the preset partitioning rules can be a series of pre-defined logical conditions and algorithms used to guide how to divide the power grid topology based on boundary points, such as "a branchless path between any two boundary points constitutes a unit."
[0033] In some embodiments, the device can load the digital topology model of the transformer substation as a directed graph, where the main transformer substation is the root node and the current direction is the edge direction. Then, a depth-first traversal is performed starting from the root node. When the traversal path encounters a critical branch node or a user smart meter, the path from the previous boundary point (initially the main transformer substation) to the current point is recorded as a topology calculation unit. Using the currently encountered critical branch node as the new starting point, a depth-first traversal continues for all its untraversed branches until all leaf nodes (user smart meters) have been visited. Optionally, the device can also employ a breadth-first traversal (BFS) algorithm based on path tracing. The main transformer substation is placed in a queue, and then processing begins in a loop. Each time a node is taken from the queue, all its downstream adjacent nodes are checked. If an adjacent node is a critical branch node or a user smart meter, the path from the current node to that adjacent node is identified as a topology calculation unit; then, if the adjacent node is a critical branch node, it is added to the queue for subsequent processing. This process is repeated until the queue is empty.
[0034] In other embodiments, the diagnostic device can also extract a topology map from the digital topology model of the distribution area, where nodes represent transformers, branch points, and user meters, and edges represent connecting lines. Then, a spectral clustering algorithm is used to intelligently segment nodes based on their electrical distance and load correlation. Specifically, the device first constructs an affinity matrix between nodes, where the element values (electrical connection strength) are weighted by the reciprocal of the electrical distance and the load correlation coefficient. This makes nodes that are electrically close and have similar electricity consumption behaviors more likely to be grouped together. Then, by calculating the Laplace matrix and its eigenvectors of this matrix, the nodes are mapped to a low-dimensional feature space, and K-means clustering is performed in this space. The optimal number of clusters K is determined by analyzing the steep descent point of the eigenvalue distribution, and the clustering quality is evaluated by the silhouette coefficient. If the clustering is below a threshold, the K value is adjusted, and clustering is re-established. The resulting initial clustering results form a candidate topology unit set.
[0035] It is understandable that other methods can be used to achieve the above unit division process, such as combining geographic information (GIS) regional segmentation algorithms, as long as it can ensure that the entire transformer network is completely, non-overlapping, and without omissions decomposed into computational units defined by specified boundary points. No limitation is made here.
[0036] In some embodiments, the device calculates the cohesion (the tightness of internal connections within the unit) and boundary complexity (the number of connections between the unit and other units) of each candidate topology unit. Based on preset unit balance constraints (such as a suitable number of nodes within a unit, relatively balanced load distribution among units, and complete coverage of critical nodes), units with low cohesion or overly complex boundaries are merged or split for adjustment. Finally, this optimized set of candidate topology units is matched and calibrated against the locations of critical branch nodes, transformer substations, and user smart meters in the physical world to generate the final topology calculation unit partitioning scheme.
[0037] S102. Perform regression fitting calculations on the historical line loss rate and historical active load of multiple topology calculation units to obtain multiple unit theoretical line loss-load characteristic curve models corresponding to the topology calculation units.
[0038] After the unit partitioning is completed and historical operational data accumulation reaches the expected level (e.g., more than one month), the diagnostic equipment will perform a modeling process for each topology computing unit partitioned in S101. First, the equipment extracts all historical data for the specified topology computing unit from its database over a given period, forming two one-to-one time-series datasets: one is the historical active power load sequence {L1, L2, ..., L...}. n The other is the historical line loss rate sequence {R1,R2,...,R}. n This means that at any given time point tᵢ, the device has a data pair (Lᵢ, Rᵢ). The diagnostic device then treats these data pairs as a scatter plot in a two-dimensional coordinate system, where the horizontal axis represents the active load L and the vertical axis represents the line loss rate R. Theoretically, the line loss of a power grid consists of two main parts: fixed losses (independent of the load, such as transformer core losses) and variable losses (related to the load, mainly I²R losses on the line). Since power P is positively correlated with current I, variable losses are approximately proportional to the square of the load. Therefore, the relationship between the line loss rate and the load typically resembles an upward-opening parabola. The diagnostic device uses a regression fitting algorithm (such as multinomial regression) to determine a function R = f(L) (e.g., a quadratic function R = aL² + bL + c) that minimizes the overall error (e.g., the sum of squared residuals) between the actual line loss rate Rᵢ at all historical data points and the theoretical line loss rate f(Lᵢ) predicted by the model. After completing the calculation, the diagnostic device obtains the specific model parameters for this topology calculation unit. This process is repeated for all topology calculation units one by one, and finally a theoretical line loss-load characteristic curve model is generated for each unit. These models are then stored and associated with the corresponding unit ID.
[0039] S103. Within a preset future time period, based on the preset transformer area line loss prediction model, obtain the total load curve of the target transformer area for the preset future time period and the load curves of multiple units corresponding to multiple topology calculation units, and compare the theoretical line loss rate corresponding to the total load curve of the transformer area with the real-time line loss rate to determine the real-time difference.
[0040] Within a monitoring cycle, the diagnostic equipment uses its built-in preset transformer area line loss prediction model. This model comprehensively considers various input information, such as historical load data of the target transformer area over the past few weeks or even years, weather forecasts for the future period (temperature, humidity, light intensity, etc.), and calendar information (such as weekdays, weekends, and holidays), to predict the load for a "preset future period" (e.g., the next 24 hours). This prediction task produces two types of results: first, it generates a macroscopic total load curve for the transformer area; second, it generates a microscopic unit load curve for each topology calculation unit divided in S101. Simultaneously with the prediction, the diagnostic equipment performs a real-time status assessment. It sends data acquisition commands via the communication network to the transformer area's main meter and all end-user smart meters to obtain their accurate active power readings at the current moment. Based on this real-time data, the equipment calculates two core values: first, it uses the reading of the main meter as the current real-time total load; second, it calculates the current real-time line loss rate using the formula (total meter power - sum of power of all user meters) / total meter power. Subsequently, the equipment inputs the real-time total load value obtained earlier into a pre-established theoretical model that specifically describes the overall line loss characteristics of the transformer area (this model reflects the line loss pattern of the transformer area under healthy conditions), and calculates the theoretical line loss rate that should have occurred at this load level. Finally, the equipment subtracts this theoretical line loss rate from the "real-time line loss rate" to obtain the "real-time difference".
[0041] S104. When the real-time difference reaches a preset threshold, the load curves of multiple units are mapped to the theoretical line loss-load characteristic curve model of multiple units to obtain the theoretical line loss rate of multiple units.
[0042] When an anomaly in line loss is detected across the entire distribution area (i.e., the real-time difference exceeds a preset threshold), the diagnostic equipment confirms that an alarm for an anomaly in line loss across the distribution area has been triggered. It then retrieves data from its memory or database: first, the unit load curves of all topology calculation units; and second, the pre-trained and stored theoretical line loss-load characteristic curve models for these units in S102. Next, the diagnostic equipment iterates through each topology calculation unit. For any given unit, the equipment retrieves its corresponding load prediction curve (e.g., a load sequence containing 96 future points, with a value every 15 minutes) and its corresponding line loss-load model. Then, the equipment performs a "mapping" operation on the predicted load value Lᵢ at each time point on the load curve, substituting it into the aforementioned function to calculate the corresponding theoretical line loss rate Rᵢ=aLᵢ²+bLᵢ+c. This calculation process proceeds point-by-point along the entire load curve, ultimately generating a completely new "unit theoretical line loss rate" curve that perfectly corresponds to the given time point for that unit. This process will be executed in parallel or serially for all topology computing units, and the final output is multiple theoretical line loss rate curves corresponding to each unit.
[0043] Understandably, the line loss-load model can be based on daily load curve data from the past year (each curve contains 96 sampling points) and the corresponding daily line loss. To differentiate between different electricity consumption scenarios, the equipment extracts three key morphological features from each daily load curve: peak-to-valley difference, average load, and relative height of the evening peak, forming a three-dimensional feature vector. Then, the equipment uses the K-means clustering algorithm to cluster these 365 feature vectors, dividing the year into K "stages" with similar electricity consumption patterns, such as "normal days in spring and autumn" and "high-temperature days with air conditioning in summer," where the optimal number of clusters K is determined by maximizing the silhouette coefficient. After completing the stage division, the equipment performs independent regression analysis within each stage: for all days belonging to the same cluster, its 96-point load curve is first converted into a daily load rate sequence, and the root mean square load rate (ρ_rms) that comprehensively reflects the shape of the daily load is calculated. Finally, using this root mean square load rate as the independent variable and the daily line loss as the dependent variable, a quadratic regression fitting is performed to obtain the line loss prediction model specific to that stage. Finally, this electricity-based model is converted into a general theoretical line loss rate curve model ΔP_k(ρ), thereby generating a theoretical line loss-load characteristic curve for each electricity consumption stage.
[0044] In some embodiments, the diagnostic device can perform preliminary statistical filtering on the raw historical data, such as using the sigma criterion or the interquartile range of box plots, to remove data points with extremely abnormal line loss rates (e.g., much greater than 30% or less than 0), resulting in a preliminary dataset. The device then uses this preliminary dataset for the first regression fitting, obtaining an initial line loss-load characteristic model. Next, the device uses this initial model to iterate through all raw historical data points (including those previously removed), calculating the model residual for each point, i.e., the difference between the actual line loss rate and the theoretical line loss rate predicted by the model. Ideally, these residuals should follow a zero-centered normal distribution. If the residual of a data point is significantly positive and exceeds a dynamically set threshold (e.g., twice the standard deviation of the mean of all residuals), the device marks that data point as highly suspected of containing non-technical line loss. After marking all highly suspected points, the device permanently removes these points from the dataset, forming a cleaner training dataset. Finally, based on this iteratively purified final dataset, the device re-executes the regression fitting calculation to obtain the unit theoretical line loss-load characteristic curve model.
[0045] S105. Within a preset future time period, compare the real-time line loss rate curves of multiple topology computing units with the theoretical line loss rates of multiple units to determine whether there are preset deviation characteristics.
[0046] Triggered after S103 confirms an anomaly in the overall line loss of the transformer area, it continues to run for a preset future period, starting an independent monitoring thread for each topology calculation unit. In each monitoring cycle (e.g., every 5 minutes), the device sends data acquisition commands via the communication network to the unit's inlet metering point (which may be a virtual metering point on the outlet or branch node of the previous unit) and all outlet metering points (downstream user meters or the inlet of the next unit). After acquiring the instantaneous active power of all boundary points, it calculates the unit's real-time line loss rate at the current moment. Then, the device compares this real-time value with the theoretical line loss rate to obtain the instantaneous deviation value, which is stored in a time series buffer. Pattern analysis is performed on the deviation time series in the time series buffer to determine whether it meets any of the "preset deviation characteristics." For example, deviation characteristics might include: "the deviation value of 3 consecutive sampling points is greater than 1%", which corresponds to persistent electricity theft or metering failure; "the deviation value jumps by more than 5% within 5 minutes", which may indicate a sudden ground fault in the line. This comparison and feature recognition process is executed in parallel for all topology computing units. When the line loss deviation behavior of a unit matches the preset deviation feature, the unit is locked as a suspect.
[0047] In some embodiments, the determination of preset deviation characteristics can be achieved in several ways: Optionally, the diagnostic device can employ a statistical process control-based method. First, the theoretical line loss rate curve of each unit is considered as the centerline of its normal operation, and based on the deviation fluctuations during historical normal operation, a statistically reasonable upper control limit (UCL) and lower control limit (LCL) are calculated to form a control chart. Then, the device plots the line loss rate calculated in real time point by point onto this control chart. Finally, deviation characteristics are identified by applying classic anomaly detection rules (such as Westinghouse rules). For example, if a point exceeds the UCL / LCL, or multiple consecutive points fall on one side of the centerline, or the points show a clear upward / downward trend, it is determined that a preset deviation characteristic exists. Optionally, the diagnostic device can also employ a template matching method based on the Dynamic Time Warping (DTW) algorithm. In this method, the system predefines various typical abnormal line loss deviation curve "templates," such as "gradually rising type," "peak type," and "periodic type." Once real-time monitoring begins, the device extracts a sequence of unit line loss deviations over a recent period (e.g., the past hour) and uses the DTW algorithm to calculate the "morphological similarity distance" between this sequence and various anomaly templates. The DTW algorithm effectively compares the similarity of two time series that may differ in length and may be stretched or shifted along the time axis. Once a calculated distance is less than a preset similarity threshold, it indicates that the unit's real-time line loss deviation behavior highly matches a known anomaly pattern, thus confirming the existence of a preset deviation characteristic, which is not specified here.
[0048] In some embodiments, after determining an anomaly in the line loss of the transformer area in step S103, the diagnostic device, in addition to independently monitoring the deviation characteristics of each unit, also calculates a real-time deviation contribution index in parallel. At any time t, for topology calculation unit i, its deviation contribution is defined as the proportion of the abnormal line loss power of that unit (i.e., the real-time line loss power of unit i minus its theoretical line loss power) in the total abnormal line loss power of the transformer area (i.e., the total real-time line loss power of the transformer area minus its theoretical line loss power). Subsequently, the device continuously tracks the deviation contribution ranking of each unit within a sliding time window, focusing on a dynamic suspect group composed of units whose contribution ranking is consistently stable at the top (e.g., ranked in the top three for 80% of the time points in the past hour). Simultaneously, the device calculates the sum of the cumulative contributions of this suspect group at each time point. When the combination not only has a stable ranking, but its cumulative contribution also consistently exceeds a preset significance ratio (e.g., 70%), the system can determine that the combination of units together constitutes a preset deviation feature, even if the independent deviation feature of any one of the units is not significant.
[0049] S106. If it is determined that there are preset deviation characteristics, the topology calculation unit with preset deviation characteristics is identified as an abnormal unit.
[0050] After step S105 successfully matches at least one preset deviation feature from the line loss deviation behavior of one or more topology computing units, a series of pre-configured automated workflows are triggered based on the matched deviation feature type and the anomalous unit. For example, the diagnostic device immediately generates a detailed alarm message that packages all relevant "evidence," including: the unique identifier of the confirmed anomalous unit, the timestamp of the anomaly confirmation, the specific deviation feature that triggered the judgment (e.g., "deviation exceeding 2% for 30 consecutive minutes"), a comparison curve of the real-time line loss rate and the theoretical line loss rate during the anomalous period, and possible preliminary diagnostic suggestions (e.g., depending on the type of deviation feature, the system may infer "suspected electricity theft" or "suspected metering equipment failure"). Subsequently, this structured alarm message is pushed to the upper-level marketing and distribution management system, the operation and maintenance work order system, or directly sent to the mobile terminal of the designated operation and maintenance personnel, thereby initiating the manual verification and processing process.
[0051] In some embodiments, after initially identifying the abnormal unit, at least one reference unit is selected for the abnormal unit based on preset similarity conditions. These similarity conditions can comprehensively consider multiple dimensions such as topology, historical load level, line physical parameters, and the composition of the types of users supplied. Next, the diagnostic device obtains the real-time line loss rate of the abnormal unit and the reference unit in the same time period and calculates the difference curve between them. Subsequently, the device determines the total duration for which the difference curve exceeds a preset horizontal comparison threshold. When the total duration exceeds a preset duration threshold, the device confirms that the abnormal unit has an abnormal line loss and further calls a preset diagnostic rule base. Based on the combination of deviation characteristics and other electrical data of the unit, it intelligently determines the specific type of line loss anomaly, such as electricity theft, metering failure, or line aging, and generates a final diagnostic report.
[0052] In this embodiment, by employing the technical feature of refining the power grid of the distribution area into topology calculation units and establishing an independent theoretical line loss-load characteristic model for each unit, and then combining it with load forecasting for dynamic comparison, the rapid and accurate location of line loss anomalies is achieved, reducing the scope of fault investigation from the entire distribution area to the smallest topology unit, thereby improving operation and maintenance efficiency.
[0053] In the above embodiments, the diagnostic device can locate line loss anomalies by fusing real-time and theoretical longitudinal comparisons. In practical applications, under certain complex power grid operation scenarios, multiple topology calculation units may be identified simultaneously as having line loss anomalies. In this case, the system outputs a discrete list of anomaly units, and the inherent correlation between these units is unknown, reducing the accuracy of the diagnosis.
[0054] Please see Figure 2 This is another flowchart illustrating the line loss diagnosis method based on multi-source data in this application embodiment.
[0055] S201. Using the key branch nodes, the main meter of the target transformer area, and the user's smart meter as boundaries, the target transformer area is divided into units based on preset partitioning rules and the digital topology model of the transformer area, resulting in multiple topology calculation units.
[0056] S202. Perform regression fitting calculations on the historical line loss rate and historical active load of multiple topology calculation units to obtain multiple unit theoretical line loss-load characteristic curve models corresponding to the topology calculation units.
[0057] S203. Within a preset future time period, based on the preset transformer area line loss prediction model, obtain the total load curve of the target transformer area for the preset future time period and the load curves of multiple units corresponding to multiple topology calculation units, and compare the theoretical line loss rate corresponding to the total load curve of the transformer area with the real-time line loss rate to determine the real-time difference.
[0058] S204. When the real-time difference reaches a preset threshold, the load curves of multiple units are mapped to the theoretical line loss-load characteristic curve model of multiple units to obtain the theoretical line loss rate of multiple units.
[0059] S205. Within a preset future time period, compare the real-time line loss rate curves of multiple topology computing units with the theoretical line loss rates of multiple units to determine whether there are preset deviation characteristics.
[0060] S206. If it is determined that there are preset deviation characteristics, the topology calculation unit with preset deviation characteristics is identified as an abnormal unit.
[0061] Steps S201~S206 and Figure 1 Steps S101 to S106 in the illustrated embodiment are similar and will not be repeated here.
[0062] S207. When the number of anomalous units is not unique, based on the deviation time series of the anomalous units in a preset future time period, perform spatial clustering analysis on the anomalous feature vector of the deviation time series to obtain at least one anomalous unit cluster.
[0063] When the diagnostic device identifies more than one anomalous unit in S206, it retrieves the deviation time series of each marked anomalous unit from the monitoring database during the period of the anomalous occurrence. Next, the device performs feature engineering on each raw time series, including: the mean, maximum, and variance of the deviation (reflecting the overall level, peak intensity, and stability of the anomalousness), the kurtosis and skewness of the deviation curve (describing the sharpness and symmetry of the curve), the time of the first threshold exceedance, and the duration of the anomalous event. After generating corresponding feature vectors for all anomalous units, the diagnostic device treats these vectors as points in a multi-dimensional space and applies spatial clustering algorithms (such as K-means or DBSCAN) to group these points. The algorithm automatically calculates the Euclidean distance or cosine similarity between the vectors and groups vectors with similar distances (i.e., units representing similar anomalous behavior patterns) into the same "cluster." Finally, this step outputs one or more anomalous unit clusters. For example, if the deviation curves of units A, B, and C all show "stable and continuous high deviation", they may be clustered into one category; while the deviation curves of units D and E show "drastic and irregular pulse-like jumps", they may be clustered into another category.
[0064] In some embodiments, this spatial clustering analysis can be implemented in several ways: Optionally, the diagnostic device can employ a clustering method based on the K-means algorithm. First, the device extracts a three-dimensional feature vector for each anomalous unit, consisting of the mean deviation, peak value, and duration. Then, to determine the optimal number of clusters K, the device can use the silhouette coefficient method, i.e., try different K values (e.g., from 2 to half the total number of anomalous units) for clustering and calculate the silhouette coefficient for each result, finally selecting the K value that maximizes the average silhouette coefficient. Finally, the device executes the standard K-means algorithm with the selected K value, iteratively updating the cluster centers and sample affiliations until convergence, obtaining the final K anomalous unit clusters. Optionally, the diagnostic device can also employ a DBSCAN (Density-Based Noisy Spatial Clustering) algorithm. Similarly, the deviation time series is converted into a feature vector. Then, the device sets two key parameters: the neighborhood radius (Eps) and the minimum number of neighborhood samples required for the core objects (MinPts). Starting with any core object, all objects reachable from it are recursively grouped into a cluster. This process does not require pre-specifying the number of clusters and can automatically identify any "outliers" (i.e., anomalous units with unique behavioral patterns) that do not belong to any cluster, which helps to discover isolated and special types of faults. Understandably, other clustering algorithms such as hierarchical clustering and spectral clustering can also be used to achieve the same purpose, and this is not limited here.
[0065] In other embodiments, the device can perform wavelet transform on the deviation time series of each anomalous unit, decomposing the signal to obtain high-frequency and low-frequency wavelet energy coefficients at different time scales. These are used to capture the transient and slow-persistent characteristics of the anomaly, respectively. Secondly, the device extracts topological importance indicators for the unit from the digital topology model of the distribution area, such as its layer depth (electrical level relative to the transformer) and the total number of users directly or indirectly connected downstream. These indicators quantify the unit's criticality in the network structure. Non-Gaussian statistics, such as kurtosis and skewness, are calculated for the deviation time series to describe the sharpness and symmetry of the anomalous signal distribution. Finally, the diagnostic device combines the extracted high / low-frequency wavelet energy coefficients, topological importance indicators, and non-Gaussian statistics to form an anomaly feature vector. After generating corresponding feature vectors for all anomalous units, the diagnostic device treats these vectors as points in a multi-dimensional space and applies spatial clustering algorithms (such as K-means or DBSCAN) to group these points.
[0066] S208. Calculate the sum of the minimum electrical distances of all abnormal units within the abnormal unit cluster in the digital topology model of the transformer area.
[0067] The diagnostic equipment identifies a target cluster of anomalous cells and retrieves a list of all anomalous cells within that cluster. Next, the equipment loads a digital topology model of the distribution area and extracts all possible cell pairings from the cluster. For each pairing (e.g., cell A and cell B), the equipment initiates a shortest path search algorithm from graph theory (such as Dijkstra's algorithm). Starting from the node corresponding to cell A and ending at the node corresponding to cell B, the minimum cumulative impedance value connecting the two points is calculated on the topology graph, where line impedance is the path "cost." This value represents the minimum electrical distance between A and B. This calculation process is repeated for all unique cell pairs within the cluster. For example, if a cluster contains three cells {A, B, C}, the equipment needs to calculate the distances dist(A, B), dist(A, C), and dist(B, C). Finally, the diagnostic equipment sums all the calculated minimum electrical distance values to obtain a final total value. This total value is the minimum electrical distance sum.
[0068] In some embodiments, this calculation can be implemented in several ways: Optionally, the diagnostic device can employ an iterative calculation method based on multiple Dijkstra's algorithm iterations. First, the digital topology model of the transformer area is loaded as a graph data structure. Then, for a cluster containing N anomalous units, the device performs N iterations. In each iteration, taking one anomalous unit as the source point, Dijkstra's algorithm is run to calculate the shortest path from the source point to all other nodes in the graph. Finally, the shortest path values between all pairs of units within the cluster are extracted from the results of these N calculations and summed. Optionally, the diagnostic device can also use the Floyd-Warshall algorithm for a one-time calculation. This method first constructs an adjacency matrix based on the topology model, where the matrix elements represent the direct line impedance between nodes (infinite if not directly connected). Then, by executing the Floyd-Warshall algorithm, which uses dynamic programming, the shortest path between all pairs of nodes in the graph is calculated in one go, and the adjacency matrix is updated. After the calculation is completed, the device only needs to directly query and sum the minimum electrical distance between all corresponding unit pairs from the final matrix according to the member list of the abnormal unit clusters; no restrictions are imposed here.
[0069] S209. Determine whether the sum of minimum electrical distances is greater than the preset topology proximity threshold.
[0070] The diagnostic device retrieves the single value of the "minimum total electrical distance" calculated by S208 for a specific anomalous unit cluster. Simultaneously, the device reads a pre-set "topology proximity threshold" from its system configuration. This threshold is typically set based on expert experience or statistical analysis of numerous historical cases; for example, it might be set as "the average of the total electrical distances between all downstream users covered by a typical cable branch box." The diagnostic device then compares the "minimum total electrical distance" with the "pre-set topology proximity threshold."
[0071] In some embodiments, the device does not directly use the original "minimum electrical distance sum" before performing the comparison; instead, it first normalizes it. The device first calculates the "topological centroid" (e.g., the geometric or electrical center of all member nodes in the topology graph) of the anomalous cluster. Then, centered on this centroid, a minimum topology subgraph containing all members of the cluster is defined, and metrics such as "average line density" or "average node degree" are calculated for this subgraph to quantify the topological compactness of the local region where the cluster is located. Finally, the device divides the original minimum electrical distance sum by this local topology density metric to obtain the normalized distance.
[0072] S210. If the sum of the minimum electrical distances is determined to be greater than a preset topological proximity threshold, the abnormal units of the abnormal unit cluster are diagnosed as collaborative abnormal events.
[0073] The diagnostic device will execute this step only if the judgment result of S209 is "yes". It is understood that the abnormal units within the abnormal unit cluster whose minimum electrical distance sum is greater than the preset topology proximity threshold exhibit highly consistent line loss deviation curves in morphology and dynamics (from the clustering results of S207), but are geographically distant from each other in the power grid topology (from the judgments of S208 and S209). This indicates that multiple units that are electricalally unrelated or geographically distant exhibit unnatural synchronicity in their abnormal behavior in time and morphology, which cannot be explained by local faults or propagation effects, and may involve external coordination or human intervention. Therefore, the diagnostic device will characterize this abnormal unit cluster as a coordinated abnormal event and label all abnormal units within the cluster with this diagnostic tag. This diagnostic conclusion, along with relevant evidence (such as clustering results, calculated electrical distance values, etc.), will be recorded in the alarm log and pushed to maintenance personnel, guiding them to investigate from a higher level (such as voltage quality problems in the upper-level power grid, clock errors in the metering system across the entire region, or even organized wide-area electricity theft).
[0074] In some embodiments, the diagnosis can be implemented in several ways: Optionally, the diagnostic device can employ a direct diagnostic method based on a rule engine. When the condition of S209 is met, the rule is triggered, and the system directly outputs a diagnostic conclusion. Optionally, the diagnostic device can also employ a Bayesian network diagnostic model based on probabilistic inference. This model treats the cooperative anomaly event as a hidden root node, and its child nodes include observable variables such as unit behavior similarity and topological dispersion. When high behavior similarity and high topological dispersion are observed, the model calculates the posterior probability of the cooperative anomaly event through Bayesian inference. Only when this probability exceeds a preset confidence threshold will the system issue a diagnostic conclusion. Optionally, after diagnosing the anomalous unit cluster as a cooperative anomaly event, the device analyzes the lead-lag relationship of the load change time series among the anomalous units within the cluster through the Granger-causality test, thereby obtaining the power transmission time series correlation between them. Then, the device connects the unit pairs with significant causal relationships using all anomalous units within the cluster as nodes, and uses the strength of the causal relationship as the edge weight to construct a directed acyclic graph. Then, the device uses a maximum weighted path search algorithm in the graph to identify the power disturbance propagation path from the root node (i.e., the node without incoming edges, representing the initial source of the disturbance) to all other nodes. The starting node of this path, i.e., the root node, is ultimately determined as the root cause anomaly unit of this coordinated anomaly event. Finally, the diagnostic device uses this accurately identified root cause anomaly unit as the updated unique anomaly unit and executes step S211, which confirms it and determines its specific line loss anomaly type based on a preset diagnostic rule base; no specific limitations are imposed here.
[0075] S211. If the number of abnormal units is determined to be unique, confirm that the abnormal unit has a line loss abnormality and determine the type of line loss abnormality based on the preset diagnostic rule base.
[0076] After identifying the anomalous unit in S206, the diagnostic device first confirms that only one topology calculation unit is currently marked as anomalous. In this case, the device does not need to execute the multi-point anomaly clustering and collaborative analysis process of S207-S210, but directly proceeds to the diagnosis of the single-point anomaly. The device retrieves the complete monitoring dataset of the anomalous unit from its database, which includes not only the time series of deviations in line loss rate, but also the historical records of multi-dimensional electrical parameters such as voltage, current, and power factor of the unit. Next, the device formally confirms the existence of line loss anomalies in the unit. This confirmation is recorded in the system log and may trigger a series of preset workflows, such as generating alarms and sending notifications. Subsequently, the device starts its built-in diagnostic engine and calls the preset diagnostic rule base. This rule base is the "intelligent core" of the system, which embodies a large amount of expert experience and historical case analysis, and contains "feature fingerprints" of various line loss anomaly types. The diagnostic engine matches the multi-dimensional data of the anomalous unit with various patterns in the rule base. For example, if the line loss deviation of a unit exhibits "persistently high values unrelated to load levels," and "more significant deviations at night," along with "abnormally increased voltage drop," this combination of features might match the "suspected electricity theft" pattern in the rule base. Similarly, if the deviation exhibits "a sharp increase with load" and "high three-phase current imbalance," this might match the "line insulation aging" pattern. The diagnostic engine calculates the matching degree or confidence level for various possible types and selects the type with the highest matching degree as the final diagnostic result. This result, along with key evidence supporting the conclusion, is integrated into a structured diagnostic report and pushed to relevant maintenance personnel.
[0077] In this embodiment, by employing a fusion technology that vectorizes the deviation time series of abnormal units and performs spatial clustering analysis, when the system faces a complex situation with multiple abnormal units simultaneously alarming, the diagnostic device no longer treats them as isolated events to be processed one by one. Instead, it automatically groups abnormal units with similar behavior patterns into one category by calculating and comparing various abnormal behaviors. This process effectively solves the technical problems in the prior art where, when faced with concurrent alarms, maintenance personnel cannot identify the inherent correlation between alarms, can only blindly investigate, and have low analysis efficiency. It achieves the technical effect of intelligent dimensionality reduction and pattern aggregation for massive, scattered concurrent fault alarms, transforming the chaotic alarm list into a structured event cluster view. Furthermore, by employing a fusion technology that jointly analyzes the behavior clustering results with the electrical distance in the digital topology model of the transformer area, the diagnostic device, after identifying event clusters with similar behaviors, further infers the root cause by calculating the topological dispersion of members within the cluster in the power grid. This technology solves the problem in existing technologies that cannot distinguish whether multiple alarms originate from a single local fault or a systemic problem, leading to incorrect diagnostic paths. It then achieves the technical effect of tracing the source and identifying the root cause of complex faults.
[0078] The exemplary diagnostic device 300 provided in the embodiments of this application is described below. Figure 3 This is an exemplary hardware structure diagram of the diagnostic device 300 provided in this application embodiment.
[0079] In some embodiments, the diagnostic device 300 is a computer device or includes a computer device. The computer device includes a processor, memory, and a network interface connected via a system bus. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database of the computer device stores data. The network interface of the computer device is used to communicate with other external terminals or servers via a network connection. In some embodiments, the network interface can be a wired network interface; in some embodiments, the network interface can also be a wireless network interface. When the computer program is executed by the processor, it implements the methods in the embodiments of this application.
[0080] Those skilled in the art will understand that Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements. The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
[0081] As used in the above embodiments, depending on the context, the term "when..." can be interpreted as meaning "if...", "after...", "in response to determining...", or "in response to detecting...". Similarly, depending on the context, the phrase "when determining..." or "if (the stated condition or event) is interpreted as meaning "if determining...", "in response to determining...", "when (the stated condition or event) is detected", or "in response to detecting (the stated condition or event)".
[0082] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive), etc.
[0083] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.
Claims
1. A method for diagnosing line loss in transformer substations based on multi-source data fusion, characterized in that, The method, applied to a transformer substation unit fusion device, includes: Using the key branch nodes, main meter, and user smart meters of the target transformer area as boundaries, the target transformer area is divided into units based on preset partitioning rules and a digital topology model of the transformer area, resulting in multiple topology calculation units. The digital topology model of the transformer area includes the hierarchical connection relationships and physical line parameters between transformers, branch lines, nodes, and user meters within the target transformer area. The historical line loss rate and historical active power load of the multiple topology computing units are regressed and fitted to obtain multiple unit theoretical line loss-load characteristic curve models corresponding to the topology computing units. Within the preset future time period, the total load curve of the target transformer area and the load curves of multiple units corresponding to multiple topology calculation units are obtained based on the preset transformer area line loss prediction model. The theoretical line loss rate corresponding to the total load curve of the transformer area is compared with the real-time line loss rate to determine the real-time difference. When the real-time difference reaches a preset threshold, the load curves of the multiple units are mapped to the theoretical line loss-load characteristic curve models of the multiple units to obtain the theoretical line loss rate of the multiple units. Within the preset future time period, the real-time line loss rate curves of the multiple topology computing units are compared with the theoretical line loss rates of the multiple units to determine whether there are preset deviation characteristics. If the preset deviation feature is determined to exist, the topology calculation unit that has the preset deviation feature is identified as an abnormal unit.
2. The method according to claim 1, characterized in that, After the step of determining the topology computing unit with the preset deviation feature as an anomalous unit when the preset deviation feature is determined to exist, the method further includes: From the plurality of topology computing units, at least one reference unit that satisfies the preset similarity conditions to the abnormal unit is selected; After obtaining the real-time line loss rate of the abnormal unit and the at least one reference unit in the same time period, the difference curve between the real-time line loss rate of the abnormal unit and the real-time line loss rate of the at least one reference unit is calculated. If the total duration of the difference curve exceeding the preset horizontal comparison threshold exceeds the preset duration, it is confirmed that the abnormal unit has a line loss anomaly and the type of the line loss anomaly is determined based on the preset diagnostic rule base.
3. The method according to claim 1, characterized in that, After the step of determining the topology computing unit with the preset deviation feature as an anomalous unit when the preset deviation feature is determined to exist, the method further includes: When it is determined that the number of abnormal units is not unique, based on the deviation time series of the abnormal units in the preset future time period, spatial clustering analysis is performed on the abnormal feature vector of the deviation time series to obtain at least one abnormal unit cluster. The deviation time series refers to the difference sequence between the real-time line loss rate and the theoretical line loss rate of the abnormal unit. Multiple abnormal units located in the same abnormal unit cluster have similar abnormal feature vectors. Calculate the sum of the minimum electrical distances of all abnormal units within the abnormal unit cluster in the digital topology model of the transformer substation area; Determine whether the sum of the minimum electrical distances is greater than a preset topology proximity threshold; If the sum of the minimum electrical distances is determined to be greater than a preset topological proximity threshold, the abnormal units of the abnormal unit cluster are diagnosed as collaborative abnormal events.
4. The method according to claim 3, characterized in that, Before the step of performing spatial clustering analysis on the abnormal feature vectors of the deviation time series, the method further includes: Wavelet transform is performed on the deviation time series to obtain high-frequency wavelet energy coefficients and low-frequency wavelet energy coefficients at different time scales; The hierarchical depth of the abnormal unit in the digital topology model of the transformer area and the number of downstream connected users are determined as the topological importance index of the abnormal unit. The non-Gaussian statistics of the deviation time series are calculated, and the high-frequency wavelet energy coefficient, the low-frequency wavelet energy coefficient, the non-Gaussian statistics, and the topological importance index are combined into the anomaly feature vector.
5. The method according to claim 3, characterized in that, The method further includes: By examining the load variation among anomalous units within the anomalous unit cluster using Granger causality tests, the power transmission timing correlation is obtained. Using the abnormal units within the abnormal unit cluster as nodes, the power transmission timing correlation between the nodes is determined as edges, and a directed acyclic graph is constructed. The power perturbation propagation path from the root node to other nodes is identified by the maximum weight path search algorithm, and the starting node of the power perturbation propagation path is determined as the root cause anomalous unit of the cooperative anomalous event. The root node refers to a node without any incoming edges in the directed acyclic graph. Using the root cause anomaly unit as the updated anomaly unit, the steps of confirming that the anomaly unit has a line loss anomaly and determining the type of the line loss anomaly based on a preset diagnostic rule base are executed.
6. The method according to claim 1, characterized in that, The step of dividing the target transformer area into multiple topology calculation units based on preset partitioning rules and a digital topology model of the transformer area, using the key branch nodes, the main meter of the transformer area, and the user smart meters as boundaries, specifically includes: The topology diagram of the target transformer area is extracted from the digital topology model of the transformer area. The nodes of the topology diagram represent transformers, branch points and user meters, and the edges represent connecting lines. Based on the electrical distance and load correlation between nodes in the topology diagram, a candidate topology unit set is obtained by segmentation using a spectral clustering algorithm; Calculate the cohesion and boundary complexity of each unit in the candidate topology unit set, and adjust the candidate topology units whose cohesion is lower than the first threshold or whose boundary complexity is higher than the second threshold based on preset unit balance constraints. The boundary complexity represents the number of connections between the unit and other units. The unit balance constraints include the balance of the number of nodes within the unit, the balance of load distribution between units, and the integrity of coverage of key nodes. The adjusted candidate topology unit set is matched with the key branch nodes, the main meter of the substation, and the location of the user's smart meter in the target substation area to generate a topology computing unit partitioning scheme, resulting in multiple topology computing units.
7. The method according to claim 6, characterized in that, The step of obtaining a candidate topological unit set by segmenting nodes based on the electrical distance and load correlation between nodes in the topology graph using a spectral clustering algorithm specifically includes: Based on the electrical distance and load correlation between nodes in the topology diagram, an affinity matrix between nodes is constructed, where the matrix elements represent the electrical connection strength between nodes. The electrical connection strength is calculated by a weighted combination of the reciprocal of the electrical distance and the load correlation coefficient. The Laplacian matrix is calculated based on the affinity matrix, and the eigenvalues and eigenvectors of the Laplacian matrix are solved. Determine the target eigenvectors corresponding to the first K smallest non-zero eigenvalues of the Laplacian matrix, and construct a feature space based on the target eigenvectors. The K values are determined by the steep descent points of the eigenvalue distribution. K-means clustering is performed on the nodes in the feature space to obtain the initial clustering results and silhouette coefficients; When the silhouette coefficient is lower than the preset quality threshold, the K value is adjusted and the K-means clustering of nodes in the feature space is performed until the silhouette coefficient is higher than the preset quality threshold or the maximum number of iterations is reached. The final clustering results are mapped back to the original topological structure diagram to form a set of candidate topological units.
8. A diagnostic device, characterized in that, The diagnostic device includes: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code including computer instructions, and the one or more processors call the computer instructions to cause the diagnostic device to perform the method as described in any one of claims 1-7.
9. A computer program product containing instructions, characterized in that, When the computer program product is run on a diagnostic device, the diagnostic device performs the method as described in any one of claims 1-7.
10. A computer-readable storage medium comprising instructions, characterized in that, When the instructions are executed on a diagnostic device, the diagnostic device performs the method as described in any one of claims 1-7.