Low-voltage transformer area management method and system based on multi-source data and active learning
By employing multi-source data fusion and active learning methods, and utilizing the isolated forest model and DS evidence theory, the problems of model rigidity and insufficient operation and maintenance decision-making in low-voltage distribution area fault diagnosis were solved. This enabled early identification and accurate diagnosis of latent faults, dynamic adaptation to power grid changes, rational allocation of maintenance resources, and improved operation and maintenance efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- STATE GRID INFORMATION & TELECOMM GRP CO LTD
- Filing Date
- 2025-11-27
- Publication Date
- 2026-04-21
AI Technical Summary
Existing technologies have rigid models for fault diagnosis in low-voltage distribution areas, which cannot adapt to new types of faults and lack support for operation and maintenance decisions. This leads to unreasonable allocation of operation and maintenance resources and makes it difficult to effectively alleviate the pressure of latent faults.
By fusing multi-source data and active learning, unsupervised anomaly analysis is performed using the isolated forest model. Combined with DS evidence theory, the comprehensive confidence level of the fault point is determined, and the model parameters are dynamically updated to achieve early identification and accurate diagnosis of latent faults.
It significantly improves the model's sensitivity to latent faults and its diagnostic generalization ability, dynamically adapts to changes in the power grid, rationally allocates maintenance resources, reduces the risk of missed detections, and improves operation and maintenance efficiency.
Smart Images

Figure CN121901915A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power distribution network operation and maintenance technology, specifically to a low-voltage distribution area management method and system based on multi-source data and active learning. Background Technology
[0002] As distribution network reliability management becomes increasingly sophisticated, the focus is shifting from medium-voltage to low-voltage distribution areas. As the "last mile" of power supply service, low-voltage distribution areas operate in complex environments with a large user base and diverse equipment types, leading to frequent latent faults such as terminal oxidation, cable insulation aging, and loose contacts. These faults typically have a long latency period before developing into permanent failures and triggering power outage complaints, during which they exhibit abnormal electrical characteristics. Therefore, achieving early and accurate diagnosis of latent low-voltage faults is crucial for improving low-voltage reliability and user satisfaction in distribution areas.
[0003] Currently, the industry primarily utilizes intelligent diagnostic systems based on big data and artificial intelligence. These systems typically build diagnostic models based on load data to identify abnormal characteristics of individual users. For example, they can locate potential fault points by analyzing high-resistance characteristics such as voltage sags and voltage-current inverse correlations.
[0004] However, existing technologies suffer from model rigidity and a lack of operational decision support when dealing with complex low-voltage distribution area environments. The rigid models are poorly adapted to novel and rare faults. Initial diagnostic models heavily rely on the limited historical samples used during the training phase. For novel faults not fully included in the training set or rare defects with unique waveform characteristics, the model's recognition rate is extremely low. Furthermore, the inability to self-optimize based on field feedback leads to stagnant model performance after deployment, making it difficult to adapt to constantly changing operating environments and emerging new fault modes.
[0005] There is a lack of intelligent operation and maintenance decision support. Most existing systems can only output a "risk value" or "fault probability," but cannot distinguish the urgency and development stage of a fault. A newly appearing terminal oxidation and a severe overheating that has already caused the insulation layer to melt may show similar risk scores in the system. This prevents the operation and maintenance department from rationally allocating limited maintenance resources according to the actual risk, which may lead to "high-risk, low-repair" or "over-maintenance" situations, and the operation and maintenance pressure in low-voltage distribution areas cannot be effectively alleviated.
[0006] In summary, there is an urgent need for a low-voltage distribution area fault management method that has the ability to learn and evolve to adapt to new types of faults and can guide operation and maintenance decisions. Summary of the Invention
[0007] To overcome the shortcomings of the existing technology, this invention proposes a low-voltage transformer area management method based on multi-source data and active learning, comprising: Based on the fault detection requirements of low-voltage distribution areas, multiple standard data tables of multi-source data are obtained from the data platform. The multiple standard data tables are fused and dynamic impedance features are extracted to obtain a feature wide table for fault detection. An unsupervised anomaly analysis was performed on the feature wide table using an isolated forest model to identify multiple high-risk users in the low-voltage distribution area; based on the power grid topology of the low-voltage distribution area and the multiple high-risk users, multiple fault points were determined. The overall confidence level of each fault point is analyzed based on environmental data of each fault point using the DS evidence theory, and the maintenance sequence and maintenance resources of each fault point are determined based on the overall confidence level of each fault point. The isolated forest model is obtained by actively learning based on new fault sample data at preset time intervals. The active learning process includes updating its own input feature dimensions and model parameters.
[0008] Optionally, the process of fusing multiple standard data tables and extracting dynamic impedance features to obtain a wide feature table for fault detection includes: Using multiple user entities in the low-voltage distribution area as indexes, static attributes and real-time measurement data from multiple standard data tables within a preset time window are integrated into a single wide table through horizontal splicing. And by vertical association, the load time series data of each user preset time window in the wide table are aligned according to the time dimension to obtain the merged wide table; Based on the voltage and current sequences in the load time series data in the fused wide table, a least squares fitting algorithm is used to calculate multiple dynamic impedance values in real time. The maximum value, standard deviation and rate of change of the dynamic impedance values constitute a dynamic impedance characteristic sequence. The dynamic impedance feature sequence is supplemented into the fused wide table to obtain a feature wide table for fault detection.
[0009] Optionally, the anomaly analysis of the feature wide table using the isolated forest model is performed to obtain multiple high-risk users in the low-voltage distribution area, including: Based on the dynamic impedance characteristic sequence, voltage sag frequency, and total harmonic distortion of current in the feature wide table, a feature matrix is constructed; The feature matrix is input into the isolated forest model for unsupervised anomaly analysis to obtain the anomaly score for each user in the low-voltage distribution area. Multiple high-risk users are selected from the low-voltage distribution area by using the abnormal score of each user in combination with the abnormal score threshold.
[0010] Optionally, after performing anomaly analysis on the feature wide table using the isolated forest model to obtain multiple high-risk users in the low-voltage distribution area, the method further includes: A load feature matrix is constructed based on the load data in the feature wide table; The load feature matrix is input into a pre-trained TrAdaBoost transfer learning classifier for supervised risk prediction to obtain the risk probability of each user in the low-voltage distribution area. Based on the risk probability of each user in the low-voltage distribution area and the risk threshold, multiple users with abnormal risks are selected and added to the list of multiple high-risk users in the low-voltage distribution area.
[0011] Optionally, the TrAdaBoost transfer learning classifier is pre-trained in the following manner: Acquire tagged sample data from other transformer substations as source domain samples, and a small amount of tagged sample data from the low-voltage transformer substation as target domain samples. Different weights are assigned to the source domain samples and the target domain samples, and multiple iterations of training are performed on the base classifier; Each iteration of the training process is as follows: the base classifier is trained based on the weights of the source domain samples and the target domain samples to obtain the trained base classifier and the prediction error of the target domain samples; based on the prediction error of the target domain samples, the weights of the source domain samples that are misclassified on the target domain samples are reduced, and the weights of the target domain samples that are misclassified on the target domain samples are increased for use in the next iteration of training. The training continues until a preset number of iterations is reached. Then, the base classifiers that have completed training are weighted and combined to obtain a TrAdaBoost transfer learning classifier adapted to the low-voltage substation area.
[0012] Optionally, the step of using DS evidence theory to analyze the comprehensive confidence level of each fault point based on the environmental data of each fault point includes: If the fault point is a common fault point, the topology confidence of the fault point is determined based on the number of high-risk users downstream of the fault point; if the fault point is a single fault point, the topology confidence is determined based on a fixed empirical value. Based on the infrared thermometry data and historical reference temperature in the environmental data of the fault point, the confidence level of temperature anomaly at the fault point is calculated; and based on the environmental humidity and historical reference humidity in the environmental data of the fault point, the confidence level of environmental anomaly at the fault point is calculated. Using the synthesis formula of DS evidence theory, the topological confidence, temperature anomaly confidence, and environmental anomaly confidence of the fault point are synthesized to obtain the comprehensive confidence.
[0013] Optionally, the active learning process of the isolated forest model includes: The isolated forest model is updated at preset time intervals: Based on the physical characteristics of newly emerging fault points during multiple maintenance processes, the feature dimensions input to the unsupervised anomaly analysis of the isolated forest model are expanded. Based on fault data samples from multiple maintenance processes, the model parameters of the isolated forest model are fine-tuned using a meta-learning algorithm to obtain an isolated forest model that has completed active learning.
[0014] Optionally, the determination of multiple fault points based on the power grid topology of the low-voltage distribution area and the multiple high-risk users includes: Based on the power grid topology of the low-voltage distribution area, obtain multiple upstream node paths from each high-risk user to the transformer, and summarize the upstream node paths of multiple high-risk users to obtain all paths; Starting from the end of the path, search upwards level by level for nodes with multiple downstream high-risk users, and denot them as common fault points; If none of the nodes in the multiple upstream node paths of a certain high-risk user are other downstream high-risk users, then the dedicated line of the high-risk user is a single point of failure.
[0015] Based on the same inventive concept, this invention proposes a low-voltage transformer area management system based on multi-source data and active learning, comprising: The data acquisition module is used to acquire multiple standard data tables from multiple sources of data from the data platform based on the fault detection requirements of the low-voltage distribution area, and to fuse the multiple standard data tables and extract dynamic impedance features to obtain a feature wide table for fault detection. The fault identification module is used to perform unsupervised anomaly analysis on the feature wide table using an isolated forest model to obtain multiple high-risk users in the low-voltage distribution area; and to determine multiple fault points based on the power grid topology of the low-voltage distribution area and the multiple high-risk users. The transformer area management module is used to analyze the comprehensive confidence level of each fault point based on the environmental data of each fault point using the DS evidence theory, and to determine the maintenance sequence and maintenance resources for each fault point based on the comprehensive confidence level of each fault point. The isolated forest model is obtained by actively learning based on new fault sample data at preset time intervals. The active learning process includes updating its own input feature dimensions and model parameters.
[0016] Optionally, the data acquisition module is specifically used for: Using multiple user entities in the low-voltage distribution area as indexes, static attributes and real-time measurement data from multiple standard data tables within a preset time window are integrated into a single wide table through horizontal splicing. And by vertical association, the load time series data of each user preset time window in the wide table are aligned according to the time dimension to obtain the merged wide table; Based on the voltage and current sequences in the load time series data in the fused wide table, a least squares fitting algorithm is used to calculate multiple dynamic impedance values in real time. The maximum value, standard deviation and rate of change of the dynamic impedance values constitute a dynamic impedance characteristic sequence. The dynamic impedance feature sequence is supplemented into the fused wide table to obtain a feature wide table for fault detection.
[0017] Optionally, the fault identification module is specifically used for: Based on the dynamic impedance characteristic sequence, voltage sag frequency, and total harmonic distortion of current in the feature wide table, a feature matrix is constructed; The feature matrix is input into the isolated forest model for unsupervised anomaly analysis to obtain the anomaly score for each user in the low-voltage distribution area. Multiple high-risk users are selected from the low-voltage distribution area by using the abnormal score of each user in combination with the abnormal score threshold.
[0018] Optionally, the fault identification module is further configured to: A load feature matrix is constructed based on the load data in the feature wide table; The load feature matrix is input into a pre-trained TrAdaBoost transfer learning classifier for supervised risk prediction to obtain the risk probability of each user in the low-voltage distribution area. Based on the risk probability of each user in the low-voltage distribution area and the risk threshold, multiple users with abnormal risks are selected and added to the list of multiple high-risk users in the low-voltage distribution area.
[0019] Optionally, the system further includes a classifier training module for: Acquire tagged sample data from other transformer substations as source domain samples, and a small amount of tagged sample data from the low-voltage transformer substation as target domain samples. Different weights are assigned to the source domain samples and the target domain samples, and multiple iterations of training are performed on the base classifier; Each iteration of the training process is as follows: the base classifier is trained based on the weights of the source domain samples and the target domain samples to obtain the trained base classifier and the prediction error of the target domain samples; based on the prediction error of the target domain samples, the weights of the source domain samples that are misclassified on the target domain samples are reduced, and the weights of the target domain samples that are misclassified on the target domain samples are increased for use in the next iteration of training. The training continues until a preset number of iterations is reached. Then, the base classifiers that have completed training are weighted and combined to obtain a TrAdaBoost transfer learning classifier adapted to the low-voltage substation area.
[0020] Optionally, the transformer area management module is specifically used for: If the fault point is a common fault point, the topology confidence of the fault point is determined based on the number of high-risk users downstream of the fault point; if the fault point is a single fault point, the topology confidence is determined based on a fixed empirical value. Based on the infrared thermometry data and historical reference temperature in the environmental data of the fault point, the confidence level of temperature anomaly at the fault point is calculated; and based on the environmental humidity and historical reference humidity in the environmental data of the fault point, the confidence level of environmental anomaly at the fault point is calculated. Using the synthesis formula of DS evidence theory, the topological confidence, temperature anomaly confidence, and environmental anomaly confidence of the fault point are synthesized to obtain the comprehensive confidence.
[0021] Optionally, the system further includes a model update module for: The isolated forest model is updated at preset time intervals: Based on the physical characteristics of newly emerging fault points during multiple maintenance processes, the feature dimensions input to the unsupervised anomaly analysis of the isolated forest model are expanded. Based on fault data samples from multiple maintenance processes, the model parameters of the isolated forest model are fine-tuned using a meta-learning algorithm to obtain an isolated forest model that has completed active learning.
[0022] Optionally, the fault identification module is specifically used for: Based on the power grid topology of the low-voltage distribution area, obtain multiple upstream node paths from each high-risk user to the transformer, and summarize the upstream node paths of multiple high-risk users to obtain all paths; Starting from the end of the path, search upwards level by level for nodes with multiple downstream high-risk users, and denot them as common fault points; If none of the nodes in the multiple upstream node paths of a certain high-risk user are other downstream high-risk users, then the dedicated line of the high-risk user is a single point of failure.
[0023] In another aspect, this application also provides an electronic device, comprising: at least one processor and a memory; the memory and the processor are connected via a bus; The memory is used to store one or more programs; When the one or more programs are executed by the at least one processor, the low-voltage distribution area management method based on multi-source data and active learning as described above is implemented.
[0024] In another aspect, this application also provides a computer-readable storage medium having an executable program stored thereon, which, when executed, implements the low-voltage distribution area management method based on multi-source data and active learning as described above.
[0025] Compared with the closest existing technology, the present invention has the following beneficial effects: The present invention provides a low-voltage distribution area management method and system based on multi-source data and active learning, comprising: acquiring multiple standard data tables from a data platform based on the fault detection needs of the low-voltage distribution area; fusing the multiple standard data tables and extracting dynamic impedance features to obtain a feature wide table for fault detection; performing unsupervised anomaly analysis on the feature wide table using an isolated forest model to obtain multiple high-risk users of the low-voltage distribution area; determining multiple fault points based on the power grid topology of the low-voltage distribution area and the multiple high-risk users; using DS evidence theory to analyze the comprehensive confidence of each fault point based on environmental data, and determining the maintenance sequence and maintenance resources for each fault point based on the comprehensive confidence of each fault point; the isolated forest model is actively learned at preset time intervals based on new fault sample data, and the active learning process includes updating its own input feature dimensions and model parameters. This solution utilizes the periodic active learning of the Isolation Forest model to dynamically expand and optimize the input feature dimensions and model parameters. This enables the model to continuously learn and adapt to new fault modes, effectively overcoming the missed detection problem caused by the fixed features of traditional models. Ultimately, it significantly improves the model's sensitivity to latent faults and its diagnostic generalization ability, dynamically adapting to power grid changes and effectively identifying new or unknown types of latent faults, thus reducing the risk of missed detection. The solution also introduces DS evidence theory to fuse multi-source environmental data to determine the confidence level of fault points. This transforms fault detection results from single fault judgments to a detection decision basis (urgency) with multiple layers of credible environmental evidence, providing a clear quantitative basis for maintenance sequence and resource allocation. This allows limited maintenance resources to prioritize high-confidence, high-risk faults, enabling reasonable allocation of maintenance resources and effectively alleviating maintenance pressure. Simultaneously, accurate and standardized multi-source data in the data platform ensures the reliability of the subsequent fault analysis foundation. Furthermore, dynamic impedance feature sequences are extracted as key features of the Isolation Forest model, enabling the algorithm to distinguish between normal users and potential fault users with abnormal impedance fluctuation patterns, significantly improving the accuracy of initial screening. Attached Figure Description
[0026] Figure 1 A flowchart illustrating the low-voltage distribution area management method based on multi-source data and active learning provided by this invention; Figure 2 This is a schematic diagram of the low-voltage distribution area management system based on multi-source data and active learning provided by the present invention.
[0027] Figure 3 This is a schematic diagram of the structure of an electronic device provided by the present invention. Detailed Implementation
[0028] The specific embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.
[0029] Example 1 The low-voltage transformer area management method based on multi-source data and active learning provided by this invention, such as... Figure 1 As shown, it includes: S1. Based on the fault detection requirements of the low-voltage distribution area, multiple standard data tables of multi-source data are obtained from the data platform. The multiple standard data tables are fused and dynamic impedance features are extracted to obtain a feature wide table for fault detection. S2. Using the isolated forest model, perform unsupervised anomaly analysis on the feature wide table to obtain multiple high-risk users in the low-voltage distribution area; based on the power grid topology of the low-voltage distribution area and the multiple high-risk users, determine multiple fault points; S3. Using the DS evidence theory, analyze the comprehensive confidence level of each fault point based on the environmental data of each fault point, and determine the maintenance sequence and maintenance resources for each fault point based on the comprehensive confidence level of each fault point.
[0030] In step S1, S1-1, the data platform includes multiple standardized data tables from different business systems (e.g., user profile tables, electricity meter operation data tables, environmental data tables, etc., to avoid redundant development and ensure data consistency). To complete the specific task of "low-voltage latent fault analysis in transformer substations," the project team needed specific data materials. Therefore, through multi-source data tracing, they accurately selected the five most relevant and highest-quality tables from numerous standard data tables. These tables primarily originated from the data platform's electricity information collection summary table, user voltage, current, and power outage data, serving as the foundation for this project. Important fields are shown in the table below:
[0031] Among them, the User Acquisition 2.0 system is the only source that can directly provide users with real-time voltage and current measurement data, which is the core of identifying the "high resistance characteristic" fault fingerprint; the Marketing 2.0 system provides user files, transformer area relationships and meter asset information, which is the necessary foundation for accurately locating the fault location and clarifying the scope of responsibility (before / after the meter).
[0032] S1-2. Using the user entities of multiple meters in the transformer area as indexes, the static attributes and real-time measurement data of multiple standard data tables (such as user profiles and transformer area relationships) in the preset time window are integrated into the same wide table through horizontal splicing; at the same time, through vertical association, the load time series data (such as voltage and current sequences) of each user's preset time window in the wide table are aligned according to the time dimension to obtain the fused wide table after splicing.
[0033] Each row in the fused wide table corresponds to comprehensive information about a specific user at a specific point in time. That is, each row represents a snapshot of a user's data at a given time. Each column represents a feature, including the user's voltage and current, and the associated information of their transformer substation (such as substation number, total number of users in the same substation, etc.). This feature wide table contains data for all users and links them together through substation information to facilitate individual and group analysis.
[0034] Multi-source data is scattered across different systems, such as data acquisition systems, Power Management Systems (PMS), and marketing systems, each with different structures, formats, and granularities. The role of standard data tables is to break down these silos, integrating and aligning this heterogeneous data according to a unified theme (such as "user profiles" or "voltage events") to form a logically consistent data set. Raw data fields often contain a large amount of technical code or internal identifiers, with unclear business meaning. The core function of standard data tables is data governance, transforming raw data into information with clear business semantics that can be directly understood and used by business personnel and technical models through association with code tables and standardized field naming and definitions.
[0035] The reason for using a data platform to trace data in this solution is to find the most accurate and authoritative source of original data, ensuring that the analysis is based on highly reliable data. When questioning the model results, the team can clearly demonstrate the source and processing of the data, using the visual chain of data to prove the reliability of the conclusions.
[0036] S1-3, extract dynamic impedance features from the fused wide table, and supplement the dynamic impedance features into the fused wide table to obtain a feature wide table for fault detection.
[0037] Specifically, the multiple standard data tables obtained in this solution are data sets from multiple moments within a preset time window. Based on the voltage and current sequences in the load time series data in the fused wide table, a least squares fitting algorithm is used to calculate multiple dynamic impedance values in real time. The maximum value, standard deviation and rate of change of the dynamic impedance values constitute a dynamic impedance characteristic sequence. The dynamic impedance feature sequence is supplemented into the fused wide table to obtain a feature wide table for fault detection.
[0038] Traditional single resistance values only reflect the contact state at a specific instant, while dynamic impedance sequences can characterize the changing trends and instabilities of contact resistance under load fluctuations. This is a core characteristic of latent faults such as oxidation and loosening during their development phase. By analyzing the rate of change and transient response of impedance, intermittent, early faults that are difficult to detect with steady-state data can be identified more sensitively, enabling earlier warnings. This dynamic impedance feature sequence provides key temporal features for subsequent unsupervised anomaly detection in isolated forests, enabling the algorithm to distinguish between normal users and potential faulty users with abnormal impedance fluctuation patterns, significantly improving the coverage and accuracy of the initial screening.
[0039] In step S2, S2-1, an unsupervised anomaly analysis is performed on the feature wide table using an isolated forest model to obtain multiple high-risk users in the low-voltage substation area.
[0040] The reason for choosing the Isolation Forest model here is that the algorithm assumes outlier data points are "few and distinct." Data points are "isolated" by randomly selecting features and split values. Outliers, because they differ significantly from the mainstream data distribution, are usually easier to isolate (requiring fewer splits).
[0041] Based on the dynamic impedance characteristic sequence, voltage sag frequency, and total harmonic distortion rate of current in the aforementioned feature wide table, a feature matrix is constructed; The feature matrix contains the feature sequence of each meter user in the transformer area. If there are 1000 users in the transformer area, and each user's feature sequence contains 3 features (dynamic impedance feature sequence, voltage sag frequency, and total harmonic distortion of current), then the dimension of this feature matrix is 1000 rows × 3 columns. The total harmonic distortion of current refers to the ratio of the sum of the squares of the effective values of all harmonic components except the fundamental wave to the effective value of the fundamental wave by performing a fast Fourier transform on the current waveform.
[0042] The feature matrix is input into an isolated forest model for unsupervised anomaly analysis to obtain the anomaly score of each user in the transformer substation. The anomaly score of each user is combined with an anomaly score threshold to filter out multiple high-risk users from the low-voltage transformer substation.
[0043] The process involves inputting the feature matrix of the entire distribution area into the isolated forest, resulting in a list of anomaly scores corresponding to the input order. This list of anomaly scores is then sorted, and users with scores closest to 1 (with a set threshold) are selected. All these users with scores close to 1 are then used to create an unsupervised list of anomaly users. This method effectively and accurately identifies users exhibiting abnormal load behavior, even if their failure modes have never appeared in historical records.
[0044] S2-2. After performing anomaly analysis on the feature wide table using the isolated forest model to obtain multiple high-risk users in the low-voltage transformer area, the method further includes: using the TrAdaBoost transfer learning classifier to further assess the risk of users in the low-voltage transformer area to obtain the final high-risk users.
[0045] A load feature matrix is constructed based on the load data in the feature wide table; To simplify the process, the load characteristic matrix here follows the characteristic matrix from step 2-1, including the dynamic impedance characteristic sequence, voltage sag frequency, and total harmonic distortion rate of the current in the characteristic wide table.
[0046] The load feature matrix is input into a pre-trained TrAdaBoost transfer learning classifier for supervised risk prediction to obtain the risk probability of each user in the transformer area. Based on the risk probability of each user in the transformer area and the risk threshold, multiple users with abnormal risks are selected and added to the list of multiple high-risk users in the low-voltage transformer area.
[0047] The TrAdaBoost transfer learning classifier is pre-trained in the following manner: Acquire tagged sample data from other transformer substations as source domain samples, and a small amount of tagged sample data from the low-voltage transformer substation as target domain samples. Different weights are assigned to the source domain samples and the target domain samples, and multiple iterations of training are performed on the base classifier; Each iteration of the training process is as follows: the base classifier is trained based on the weights of the source domain samples and the target domain samples to obtain the trained base classifier and the prediction error of the target domain samples; based on the prediction error of the target domain samples, the weights of the source domain samples that are misclassified on the target domain samples are reduced, and the weights of the target domain samples that are misclassified on the target domain samples are increased for use in the next iteration of training. The training continues until a preset number of iterations is reached. Then, the base classifiers that have completed training are weighted and combined to obtain a TrAdaBoost transfer learning classifier adapted to the low-voltage substation area.
[0048] The TrAdaBoost algorithm gradually reduces its reliance on incompatible source domain knowledge by assigning higher initial weights to target domain samples and continuously decreasing the weights of misclassified source domain samples while increasing the weights of misclassified target domain samples during iterations. Specifically, decreasing the weights of misclassified source domain samples in the target domain weakens their influence in subsequent iterations, while increasing the weights of misclassified target domain samples focuses on difficult examples. After multiple iterations, the algorithm integrates multiple base classifiers into a single strong classifier (TrAdaBoost transfer learning classifier), which effectively incorporates general knowledge from the source domain and accurately adapts to the characteristics of the target domain.
[0049] In the above steps, the TrAdaBoost transfer learning classifier is applied as an effective supplement to the unsupervised detection of the Isolation Forest, constructing a dual guarantee mechanism of "broad unsupervised screening + precise supervised identification". While the Isolation Forest can detect anomalies in unknown patterns, it may produce false positives; the transfer learning classifier, on the other hand, utilizes existing fault knowledge of existing transformer substations to perform supervised and accurate identification of users in new substations. The combination of these two methods ensures both the sensitivity to detect novel faults and improves the accuracy and reliability of judgments through historical experience, achieving a more comprehensive and accurate initial risk screening.
[0050] This solution also includes transforming the Python scripts and SQL queries written based on the isolated forest model and TrAdaBoost transfer learning classifier from the above steps into a production-grade model suitable for stable and efficient operation on the data platform.
[0051] Specifically, the Python scripts and SQL queries initially written by data analysts are usually a validation model that can be manually executed periodically in a personal data analysis environment to process small-scale sample data. However, the data platform in this solution requires an industrial product with a production-grade model for fully automated and efficient processing of all data in a stable server cluster.
[0052] The aforementioned modifications could be: for the code itself: add error handling, logging, and parameterized configuration, and optimize SQL and Python logic to improve execution performance under large data volumes; for the execution method: encapsulate it into a task that can be scheduled by a scheduling system (such as Airflow), and configure the automatic triggering cycle, data dependencies, and output destination.
[0053] SQL queries act as filters for the model, precisely extracting data relevant to fault analysis from the vast database of the data platform. Python scripts form the core of the model's decision-making process, reprocessing the raw data extracted by SQL to create features for judgment (e.g., calculating "voltage sag" and "dynamic impedance characteristic sequence"), and containing the core judgment logic. They are responsible for executing core steps such as event determination and risk quantification, outputting a preliminary list of high-risk events.
[0054] S2-3. Based on the power grid topology of the low-voltage distribution area and the multiple high-risk users, identify multiple fault points.
[0055] Based on the power grid topology of the low-voltage distribution area, obtain multiple upstream node paths from each high-risk user to the transformer, and summarize the upstream node paths of multiple high-risk users to obtain all paths; Starting from the end of the path, search upwards level by level for nodes with multiple downstream high-risk users, and denot them as common fault points; If none of the nodes in the multiple upstream node paths of a certain high-risk user are other downstream high-risk users, then the dedicated line of the high-risk user is a single point of failure.
[0056] Specifically, based on the tree-like power grid topology of the low-voltage distribution area, the complete upstream device node path from each high-risk user to the transformer is obtained, and all these paths are aggregated to form a path set.
[0057] Starting from the end of all paths (user side), traverse all equipment nodes upstream (transformer side). For each node, check the number of high-risk users directly or indirectly connected downstream. If a node has more than one high-risk user downstream, mark that node as a potential common point of failure.
[0058] Iterate through all high-risk users. For a given user, if none of the device nodes on its upstream path are marked as a common point of failure (i.e., these nodes only have one high-risk user downstream), then the user is determined to have encountered a single point of failure. The fault location is at the user's dedicated access point, such as a terminal in its meter box, a fuse, or its dedicated user connection line.
[0059] In step S3, the DS evidence theory is used to analyze the comprehensive confidence level of each fault point based on the environmental data of each fault point, and the maintenance sequence of each fault point is determined based on the comprehensive confidence level of each fault point.
[0060] Based on each fault point: If the fault point is a common fault point, the topology confidence of the fault point is determined based on the number of high-risk users downstream of the fault point; if the fault point is a single fault point, the topology confidence is determined based on a fixed empirical value. Based on the infrared thermometry data and historical reference temperature in the environmental data of the fault point, the confidence level of temperature anomaly at the fault point is calculated; and based on the environmental humidity and historical reference humidity in the environmental data of the fault point, the confidence level of environmental anomaly at the fault point is calculated. Using the synthesis formula of the Dempster-Shafer Evidence Theory, the topological confidence, temperature anomaly confidence, and environmental anomaly confidence of the fault point are synthesized to obtain the comprehensive confidence.
[0061] Specifically, using the synthesis formula of DS evidence theory, the three functions (topological confidence, temperature anomaly confidence, and environmental anomaly confidence) are combined into a unified confidence assignment function.
[0062] The synthesized result will provide two key values: reliability and likelihood. Reliability represents the lowest overall reliability supporting the proposition "the fault point is a true fault," while likelihood represents the highest overall reliability that does not deny the proposition "the fault point is a true fault." The smaller the range of the confidence interval [reliability, likelihood] and the higher the value, the more reliable the fault point hypothesis. This confidence interval is denoted as the composite confidence level.
[0063] Maintenance personnel can prioritize allocating more maintenance resources to handle fault points with high confidence values (e.g., greater than 0.8) and narrow confidence intervals, in order of increasing confidence. For fault points with low confidence or wide confidence intervals, further inspections can be arranged in order of decreasing confidence, such as key inspections and increased temperature measurement frequency. Once a fault is discovered, maintenance can be scheduled.
[0064] This plan also includes optimizing and improving the model by conducting pilot projects in multiple power supply stations and combining positive and negative examples from the pilot units and grassroots units, including increasing feature dimensions and optimizing model parameters.
[0065] That is, the isolated forest model is actively learned based on new fault sample data at preset time intervals. The active learning process includes updating its own input feature dimensions and model parameters for the next unsupervised anomaly analysis.
[0066] The isolated forest model is deployed on a data platform, forming an automated fault detection system. This system is used in multiple pilot applications. The model acquires real-time data through the data platform, analyzes it using a series of algorithms, and outputs a list of high-risk users. Next, operations personnel conduct on-site verification based on this list and feed the verification results (confirmed risks are "positive examples," while false alarms are "negative examples"—these are extremely valuable "labeled data") back to the system. Finally, using this real feedback data from the front lines of business, the model is continuously optimized and iterated, ultimately resulting in a practically proven, accurate, and reliable fault detection model.
[0067] Specifically, the active learning process of the isolated forest model at each preset time interval includes: The isolated forest model is updated at preset time intervals: Based on the physical characteristics of newly emerging fault points during multiple maintenance processes, the feature dimensions input to the unsupervised anomaly analysis of the isolated forest model are expanded. Based on fault data samples from multiple maintenance processes, the model parameters of the isolated forest model are fine-tuned using a meta-learning algorithm to obtain an isolated forest model that has completed active learning.
[0068] In this solution, the repair sequence and repair resources for each fault point are determined based on the comprehensive confidence level. After the maintenance personnel are dispatched to perform repairs and troubleshooting based on the repair sequence and repair resources, the solution also includes: Obtain the user voltage and current waveforms after maintenance personnel have rectified the issues on-site. The dynamic time warping algorithm is used to calculate the similarity between the defect-removed waveform and the standard normal waveform library. If the similarity is lower than the preset threshold, it is determined that the inspection was not thorough and a re-inspection work order will be automatically generated.
[0069] This solution also includes a visualization interface connected to the fault detection module. This interface displays the risk value of the fault point, allowing staff to prioritize preventative maintenance based on the risk value. It also allows viewing the user's voltage and current status and historical risk changes. After maintenance records are filled in, the system automatically determines whether the maintenance was thorough. Staff obtain the risk value of the fault point through the visualization interface and then perform maintenance, including: First, inspect the inside of the meter box. The focus is on the incoming terminals, incoming switches, and the joints between incoming and outgoing lines. Pay special attention to oxidized or discolored terminal screws, meter terminals, and concealed locations such as live and neutral wire connections within multi-meter boxes. Second, inspect the cable and household connection joints. The focus is on the cable's appearance for any abnormalities, such as blistering or bulging of the cable sheath. Check for overheating and burning of the cable near the wall. For the joints between the incoming cable and the household connection line, check the clamps, insulation tape, and sheath for any signs of overheating, deformation, or oxidation. Third, inspect the upstream household connection lines and their joints. The focus is on the insulation of the household connection lines for any abnormal blistering or other problems, especially at jumper connections for oxidation. Pay particular attention to connections with multiple joints. Fourth, inspect the low-voltage outgoing lines and cable junction boxes within the distribution area. The focus of the inspection is on the cable joints. If a large number of users on the same low-voltage line in the same area are detected by the latent hazard tool and multiple users are alerted to sudden voltage drops, check whether there is abnormal oxidation of the joints in the low-voltage outgoing line or cable junction box of the area.
[0070] The method provided in this solution can be verified through practical applications to improve the fault detection effect and continuously collect sample data to optimize and improve the accuracy of the algorithm model. At the same time, this solution can also use large models to intelligently classify safety hazards and integrate "power grid map" and mobile applications to make the whole system more intuitive and convenient.
[0071] Furthermore, this solution can also be used to expand the application scope through the linkage of large and small models. That is, based on the already practical grassroots data applications (meter replacement anomalies, power quality, meter overcapacity, charging pile anomalies, etc.), combined with data from PMS, power supply service command and other systems, a knowledge base can be built, which, together with the model of this solution, can provide reasonable transformer substation renovation plans, realizing the solution of multiple problems with one maintenance.
[0072] This solution is applied to grassroots units across various regions. Through digitalization and business integration, it enables cross-disciplinary data platforms to meet grassroots needs. Staff members enjoy the improved work efficiency and work patterns brought about by data empowerment, develop small tools to solve major problems at the grassroots level, and strengthen the recognition of digital expertise at the grassroots level.
[0073] This solution offers several advantages: First, it accurately identifies potential hazards. It can detect not only obvious latent faults such as terminal corrosion and cable bulging, but also hidden latent faults like loose screws inside meter box seals and overheating of seemingly intact wires. Second, it improves the quality of power supply service. Preventative maintenance before hazards escalate into faults addresses problems at their inception, enhancing power supply reliability and user satisfaction. Third, it effectively alleviates operational and maintenance pressure. It shifts from "reactive repair" to "prevention," particularly reducing the pressure of emergency repairs during peak summer seasons. Fourth, it offers significant economic and social benefits. It reduces power loss, equipment damage, and repair costs, while also preventing losses for users due to power outages.
[0074] Example 2 Based on the same inventive concept, this invention also provides a low-voltage distribution area management system based on multi-source data and active learning, such as... Figure 2 As shown, it includes: The data acquisition module is used to acquire multiple standard data tables from multiple sources of data from the data platform based on the fault detection requirements of the low-voltage distribution area, and to fuse the multiple standard data tables and extract dynamic impedance features to obtain a feature wide table for fault detection. The fault identification module is used to perform unsupervised anomaly analysis on the feature wide table using an isolated forest model to obtain multiple high-risk users in the low-voltage distribution area; and to determine multiple fault points based on the power grid topology of the low-voltage distribution area and the multiple high-risk users. The transformer area management module is used to analyze the comprehensive confidence level of each fault point based on the environmental data of each fault point using the DS evidence theory, and to determine the maintenance sequence and maintenance resources for each fault point based on the comprehensive confidence level of each fault point. The isolated forest model is obtained by actively learning based on new fault sample data at preset time intervals. The active learning process includes updating its own input feature dimensions and model parameters.
[0075] In one possible implementation, the data acquisition module described above is specifically used for: Using multiple user entities in the low-voltage distribution area as indexes, static attributes and real-time measurement data from multiple standard data tables within a preset time window are integrated into a single wide table through horizontal splicing. And by vertical association, the load time series data of each user preset time window in the wide table are aligned according to the time dimension to obtain the merged wide table; Based on the voltage and current sequences in the load time series data in the fused wide table, a least squares fitting algorithm is used to calculate multiple dynamic impedance values in real time. The maximum value, standard deviation and rate of change of the dynamic impedance values constitute a dynamic impedance characteristic sequence. The dynamic impedance feature sequence is supplemented into the fused wide table to obtain a feature wide table for fault detection.
[0076] In one possible implementation, the fault identification module described above is specifically used for: Based on the dynamic impedance characteristic sequence, voltage sag frequency, and total harmonic distortion of current in the feature wide table, a feature matrix is constructed; The feature matrix is input into the isolated forest model for unsupervised anomaly analysis to obtain the anomaly score for each user in the low-voltage distribution area. Multiple high-risk users are selected from the low-voltage distribution area by using the abnormal score of each user in combination with the abnormal score threshold.
[0077] In one possible implementation, the fault identification module described above is further used for: A load feature matrix is constructed based on the load data in the feature wide table; The load feature matrix is input into a pre-trained TrAdaBoost transfer learning classifier for supervised risk prediction to obtain the risk probability of each user in the low-voltage distribution area. Based on the risk probability of each user in the low-voltage distribution area and the risk threshold, multiple users with abnormal risks are selected and added to the list of multiple high-risk users in the low-voltage distribution area.
[0078] In one possible implementation, the system further includes a classifier training module for: Acquire tagged sample data from other transformer substations as source domain samples, and a small amount of tagged sample data from the low-voltage transformer substation as target domain samples. Different weights are assigned to the source domain samples and the target domain samples, and multiple iterations of training are performed on the base classifier; Each iteration of the training process is as follows: the base classifier is trained based on the weights of the source domain samples and the target domain samples to obtain the trained base classifier and the prediction error of the target domain samples; based on the prediction error of the target domain samples, the weights of the source domain samples that are misclassified on the target domain samples are reduced, and the weights of the target domain samples that are misclassified on the target domain samples are increased for use in the next iteration of training. The training continues until a preset number of iterations is reached. Then, the base classifiers that have completed training are weighted and combined to obtain a TrAdaBoost transfer learning classifier adapted to the low-voltage substation area.
[0079] In one possible implementation, the aforementioned transformer area management module is specifically used for: If the fault point is a common fault point, the topology confidence of the fault point is determined based on the number of high-risk users downstream of the fault point; if the fault point is a single fault point, the topology confidence is determined based on a fixed empirical value. Based on the infrared thermometry data and historical reference temperature in the environmental data of the fault point, the confidence level of temperature anomaly at the fault point is calculated; and based on the environmental humidity and historical reference humidity in the environmental data of the fault point, the confidence level of environmental anomaly at the fault point is calculated. Using the synthesis formula of DS evidence theory, the topological confidence, temperature anomaly confidence, and environmental anomaly confidence of the fault point are synthesized to obtain the comprehensive confidence.
[0080] In one possible implementation, the system further includes a model update module for: The isolated forest model is updated at preset time intervals: Based on the physical characteristics of newly emerging fault points during multiple maintenance processes, the feature dimensions input to the unsupervised anomaly analysis of the isolated forest model are expanded. Based on fault data samples from multiple maintenance processes, the model parameters of the isolated forest model are fine-tuned using a meta-learning algorithm to obtain an isolated forest model that has completed active learning.
[0081] In one possible implementation, the fault identification module described above is specifically used for: Based on the power grid topology of the low-voltage distribution area, obtain multiple upstream node paths from each high-risk user to the transformer, and summarize the upstream node paths of multiple high-risk users to obtain all paths; Starting from the end of the path, search upwards level by level for nodes with multiple downstream high-risk users, and denot them as common fault points; If none of the nodes in the multiple upstream node paths of a certain high-risk user are other downstream high-risk users, then the dedicated line of the high-risk user is a single point of failure.
[0082] Example 3 like Figure 3 As shown, the present invention also provides an electronic device, which may be a computer device, a microcontroller device, a smart mobile device, etc. The electronic device in this embodiment may include a processor, a memory, a transceiver component, etc. The memory, processor, and transceiver component are connected via a bus; the memory can be used to store executable programs, and an exemplary executable program may include instructions; the processor is used to execute the instructions stored in the memory. The memory can also be used to store data, which can be accessed and / or modified when instructions are executed.
[0083] The processor may be a Central Processing Unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing and control core of the terminal, and it is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions in the storage medium to realize the corresponding method flow or corresponding function, so as to realize the steps of the low-voltage distribution area management method based on multi-source data and active learning in the above embodiments.
[0084] Example 4 Based on the same inventive concept, this invention also provides a readable storage medium, specifically an electronic device readable storage medium (Memory). This readable storage medium is a memory device within an electronic device used to store programs and data. It is understood that the storage medium here can include both built-in storage media within the electronic device and extended storage media supported by the electronic device. The storage medium provides storage space, which stores the terminal's operating system. Furthermore, this storage space also stores one or more instructions suitable for loading and execution by a processor. These instructions can be one or more executable programs (including program code). It should be noted that the storage medium here can be high-speed RAM or non-volatile memory, such as at least one disk storage device. Loading and executing one or more instructions stored in the storage medium by the processor can implement the steps of the low-voltage distribution area management method based on multi-source data and active learning in the above embodiments.
[0085] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0086] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0087] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0088] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0089] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit its scope of protection. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that after reading the present invention, they can still make various changes, modifications or equivalent substitutions to the specific implementation methods of the application, but these changes, modifications or equivalent substitutions are all within the scope of protection of the claims pending approval.
Claims
1. A low-voltage distribution area management method based on multi-source data and active learning, characterized in that, include: Based on the fault detection requirements of low-voltage distribution areas, multiple standard data tables of multi-source data are obtained from the data platform. The multiple standard data tables are fused and dynamic impedance features are extracted to obtain a feature wide table for fault detection. An unsupervised anomaly analysis was performed on the feature wide table using the isolated forest model to identify multiple high-risk users in the low-voltage transformer area. Based on the power grid topology of the low-voltage distribution area and the multiple high-risk users, multiple fault points were identified. The overall confidence level of each fault point is analyzed based on environmental data of each fault point using the DS evidence theory, and the maintenance sequence and maintenance resources of each fault point are determined based on the overall confidence level of each fault point. The isolated forest model is obtained by actively learning based on new fault sample data at preset time intervals. The active learning process includes updating its own input feature dimensions and model parameters.
2. The method as described in claim 1, characterized in that, The process of fusing multiple standard data tables and extracting dynamic impedance features yields a wide feature table for fault detection, including: Using multiple user entities in the low-voltage distribution area as indexes, static attributes and real-time measurement data from multiple standard data tables within a preset time window are integrated into a single wide table through horizontal splicing. And by vertical association, the load time series data of each user preset time window in the wide table are aligned according to the time dimension to obtain the merged wide table; Based on the voltage and current sequences in the load time series data in the fused wide table, a least squares fitting algorithm is used to calculate multiple dynamic impedance values in real time. The maximum value, standard deviation and rate of change of the dynamic impedance values constitute a dynamic impedance characteristic sequence. The dynamic impedance feature sequence is supplemented into the fused wide table to obtain a feature wide table for fault detection.
3. The method as described in claim 1, characterized in that, The unsupervised anomaly analysis of the feature wide table using the isolated forest model yields multiple high-risk users in the low-voltage distribution area, including: Based on the dynamic impedance characteristic sequence, voltage sag frequency, and total harmonic distortion of current in the feature wide table, a feature matrix is constructed; The feature matrix is input into the isolated forest model for unsupervised anomaly analysis to obtain the anomaly score for each user in the low-voltage distribution area. Multiple high-risk users are selected from the low-voltage distribution area by using the abnormal score of each user in combination with the abnormal score threshold.
4. The method as described in claim 1, characterized in that, After using the isolated forest model to perform anomaly analysis on the feature wide table to identify multiple high-risk users in the low-voltage distribution area, the method further includes: A load feature matrix is constructed based on the load data in the feature wide table; The load feature matrix is input into a pre-trained TrAdaBoost transfer learning classifier for supervised risk prediction to obtain the risk probability of each user in the low-voltage distribution area. Based on the risk probability of each user in the low-voltage distribution area and the risk threshold, multiple users with abnormal risks are selected and added to the list of multiple high-risk users in the low-voltage distribution area.
5. The method as described in claim 4, characterized in that, The TrAdaBoost transfer learning classifier is pre-trained in the following manner: Acquire tagged sample data from other transformer substations as source domain samples, and a small amount of tagged sample data from the low-voltage transformer substation as target domain samples. Different weights are assigned to the source domain samples and the target domain samples, and multiple iterations of training are performed on the base classifier; Each iteration of the training process is as follows: the base classifier is trained based on the weights of the source domain samples and the target domain samples to obtain the trained base classifier and the prediction error of the target domain samples; Based on the prediction error of the target domain samples, the weight of source domain samples that are misclassified on the target domain samples is reduced, and the weight of target domain samples that are misclassified on the target domain samples is increased for use in the next iteration of training. The training continues until a preset number of iterations is reached. Then, the base classifiers that have completed training are weighted and combined to obtain a TrAdaBoost transfer learning classifier adapted to the low-voltage substation area.
6. The method as described in claim 1, characterized in that, The method of using DS evidence theory to analyze the comprehensive confidence level of each fault point based on environmental data includes: If the fault point is a common fault point, the topology confidence of the fault point is determined based on the number of high-risk users downstream of the fault point; if the fault point is a single fault point, the topology confidence is determined based on a fixed empirical value. Based on the infrared thermometry data and historical reference temperature in the environmental data of the fault point, the confidence level of temperature anomaly at the fault point is calculated; and based on the environmental humidity and historical reference humidity in the environmental data of the fault point, the confidence level of environmental anomaly at the fault point is calculated. Using the synthesis formula of DS evidence theory, the topological confidence, temperature anomaly confidence, and environmental anomaly confidence of the fault point are synthesized to obtain the comprehensive confidence.
7. The method as described in claim 1, characterized in that, The active learning process of the isolated forest model includes: The isolated forest model is updated at preset time intervals: Based on the physical characteristics of newly emerging fault points during multiple maintenance processes, the feature dimensions input to the unsupervised anomaly analysis of the isolated forest model are expanded. Based on fault data samples from multiple maintenance processes, the model parameters of the isolated forest model are fine-tuned using a meta-learning algorithm to obtain an isolated forest model that has completed active learning.
8. The method as described in claim 1, characterized in that, Based on the power grid topology of the low-voltage distribution area and the multiple high-risk users, multiple fault points are identified, including: Based on the power grid topology of the low-voltage distribution area, obtain multiple upstream node paths from each high-risk user to the transformer, and summarize the upstream node paths of multiple high-risk users to obtain all paths; Starting from the end of the path, search upwards level by level for nodes with multiple downstream high-risk users, and denot them as common fault points; If none of the nodes in the multiple upstream node paths of a certain high-risk user are other downstream high-risk users, then the dedicated line of the high-risk user is a single point of failure.
9. A low-voltage distribution area management system based on multi-source data and active learning, characterized in that: include: The data acquisition module is used to acquire multiple standard data tables from multiple sources of data from the data platform based on the fault detection requirements of the low-voltage distribution area, and to fuse the multiple standard data tables and extract dynamic impedance features to obtain a feature wide table for fault detection. The fault identification module is used to perform unsupervised anomaly analysis on the feature wide table using an isolated forest model to identify multiple high-risk users in the low-voltage distribution area. Based on the power grid topology of the low-voltage distribution area and the multiple high-risk users, multiple fault points were identified. The transformer area management module is used to analyze the comprehensive confidence level of each fault point based on the environmental data of each fault point using the DS evidence theory, and to determine the maintenance sequence and maintenance resources for each fault point based on the comprehensive confidence level of each fault point. The isolated forest model is obtained by actively learning based on new fault sample data at preset time intervals. The active learning process includes updating its own input feature dimensions and model parameters.
10. The system as described in claim 9, characterized in that, The data acquisition module is specifically used for: Using multiple user entities in the low-voltage distribution area as indexes, static attributes and real-time measurement data from multiple standard data tables within a preset time window are integrated into a single wide table through horizontal splicing. And by vertical association, the load time series data of each user preset time window in the wide table are aligned according to the time dimension to obtain the merged wide table; Based on the voltage and current sequences in the load time series data in the fused wide table, a least squares fitting algorithm is used to calculate multiple dynamic impedance values in real time. The maximum value, standard deviation and rate of change of the dynamic impedance values constitute a dynamic impedance characteristic sequence. The dynamic impedance feature sequence is supplemented into the fused wide table to obtain a feature wide table for fault detection.