A transformer area line loss simulation method based on multi-source data

By analyzing multi-source data, the electricity consumption behavior characteristics and stability index of users in the transformer substation are extracted, clustering and abnormal event correlation are performed, and a comprehensive anomaly score is generated. This solves the problems of accuracy and efficiency in transformer substation line loss analysis, and realizes refined line loss management and efficient location of abnormal users.

CN121524980BActive Publication Date: 2026-05-19STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO
Filing Date
2026-01-15
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing methods for analyzing line losses in distribution areas rely on accurate grid topology, which leads to inaccurate analysis results when the topology of low-voltage distribution areas changes frequently. This makes it difficult to accurately locate abnormal users, resulting in low efficiency and high costs.

Method used

The transformer substation line loss simulation method based on multi-source data collects user power data, extracts power consumption behavior feature vectors and stability indices, performs cluster analysis, identifies abnormal line loss events, calculates the correlation strength and dynamic confidence of users, generates a comprehensive anomaly score, and filters out abnormal users.

Benefits of technology

It significantly improved the accuracy and reliability of identifying users with abnormal line losses in the transformer area, realized refined and intelligent line loss management, improved the pertinence and efficiency of on-site investigation, and reduced power loss.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121524980B_ABST
    Figure CN121524980B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on multi-source data's transformer area line loss simulation method, it is related to electric power system data analysis technical field, the method includes: collecting each user electric energy data and transformer area total table electric energy data in transformer area, and extract the power consumption behavior characteristic vector of each user and first behavior stability index, obtain user behavior characteristic set to cluster, obtain multiple homogeneous user groups and the group deviation of each user;Based on transformer area total table electric energy data, identify transformer area line loss abnormal event, and record abnormal event time period, calculate the correlation strength of the power consumption behavior of each user with transformer area line loss abnormal event as event correlation degree, and calculate the dynamic abnormal confidence degree in combination with the first behavior stability index of user;Based on the group deviation of each user and dynamic abnormal confidence degree, calculate the comprehensive abnormal score of each user, and screen out abnormal user.The application solves the technical problems of low efficiency of transformer area line loss analysis in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power system data analysis technology, specifically to a method for simulating line loss in transformer substations based on multi-source data. Background Technology

[0002] Existing methods for analyzing line losses in distribution areas heavily rely on accurate physical simulations of the power grid topology. However, distribution networks, especially low-voltage distribution areas, experience frequent topological changes, and archival data often differs from actual conditions, leading to inherent biases in simulation-based theoretical line loss calculations. When discrepancies arise between theoretical and actual statistical line losses, current technologies struggle to accurately pinpoint the specific causes and identifying abnormal users, necessitating manual on-site investigations—a process that is inefficient and costly. Summary of the Invention

[0003] This application provides a method for simulating transformer area line loss based on multi-source data, which is used to address the technical problems of low efficiency and inaccurate results in existing transformer area line loss analysis methods.

[0004] In view of the above problems, this application provides a method for simulating line loss in transformer substations based on multi-source data, the method comprising:

[0005] Collect electricity data of each user in the transformer area and electricity data of the transformer area's main meter. The electricity data includes at least time series data of voltage, current, and active power. Based on the electricity data of each user, extract the electricity consumption behavior feature vector and the first behavior stability index of each user to obtain a user behavior feature set. Based on a preset clustering algorithm, cluster the user behavior feature set to obtain multiple homogeneous user groups and the group deviation of each user.

[0006] Based on the total electricity data of the transformer area, abnormal line loss events in the transformer area are identified and the time period of the abnormal event is recorded. During the time period of the abnormal event, the correlation strength between each user's electricity consumption behavior and the abnormal line loss event in the transformer area is calculated as the event correlation degree. The dynamic anomaly confidence degree is calculated by combining the user's first behavior stability index.

[0007] Based on each user's group deviation and dynamic anomaly confidence, a comprehensive anomaly score is calculated for each user. Based on a preset comprehensive anomaly score threshold, abnormal users are selected.

[0008] One or more technical solutions provided in this application have at least the following technical effects or advantages:

[0009] This application proposes a multi-source data-based method for simulating transformer substation line losses. By comprehensively utilizing time-series data such as voltage, current, and active power from various users and the main meter within the substation area, a complete technical solution is constructed, encompassing feature extraction, behavioral clustering, and abnormal event correlation analysis. This significantly improves the accuracy and reliability of identifying abnormal users in transformer substation line loss. Compared to traditional methods, the technical solution provided in this application significantly overcomes the limitations of relying solely on simple data comparison and static threshold judgment. It achieves in-depth mining and dynamic evaluation of user electricity consumption behavior, resulting in a comprehensive improvement in the refinement and intelligence of transformer substation line loss management. It can provide maintenance personnel with clear and specific abnormal user location reports and maintenance work order guidance, thereby greatly improving the pertinence and efficiency of on-site investigations and providing strong technical support for ensuring stable power grid operation and reducing power loss. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 This is a flowchart illustrating a method for simulating line loss in transformer substations based on multi-source data, provided in an embodiment of this application.

[0012] Figure 2 This is a flowchart illustrating the process of obtaining a user behavior feature set in the method provided in the embodiments of this application. Detailed Implementation

[0013] This application provides a method for simulating transformer area line loss based on multi-source data, which addresses the technical problems of low efficiency and inaccurate results in existing transformer area line loss analysis methods.

[0014] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0015] It should be noted that the terms "comprising" and "having" are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or modules that are not explicitly listed or that are inherent to these processes, methods, products, or devices.

[0016] Examples, such as Figure 1 As shown, this application provides a method for simulating line loss in transformer substations based on multi-source data, wherein the method includes:

[0017] S10: Collect the electricity data of each user in the transformer area and the electricity data of the transformer area's main meter. The electricity data includes at least time series data of voltage, current, and active power. Based on the electricity data of each user, extract the electricity consumption behavior feature vector and the first behavior stability index of each user to obtain a user behavior feature set. Based on a preset clustering algorithm, cluster the user behavior feature set to obtain multiple homogeneous user groups and the group deviation of each user.

[0018] In transformer substation line loss analysis, traditional methods typically treat users as isolated individuals or perform simple classifications, lacking in-depth analysis of the similarities and differences in electricity consumption behavior within user groups. Because user electricity consumption behavior is complex and variable, relying on only a single or a few electricity consumption indicators is insufficient to comprehensively and accurately depict users' true electricity consumption patterns, leading to unreliable baselines in subsequent anomaly analysis.

[0019] Step S10 in the method provided in this application embodiment includes:

[0020] The electricity information collection system acquires time-series data of voltage, current, and active power uploaded by smart meters of each user in the distribution area.

[0021] The distribution automation system acquires the time-series data of voltage, current, and active power of the main meter in the distribution area for the same time period.

[0022] The collected time-series data is cleaned and time-series aligned to form a multi-source data set;

[0023] Among these, obtaining user behavior feature sets, such as Figure 2 As shown, it includes:

[0024] Extract load pattern characteristics, including calculating the dynamic time-normalized distance vector of the user's daily load curve, and statistical characteristics including maximum value, minimum value, average value, standard deviation, and peak-to-valley difference;

[0025] Extract voltage-current coupling characteristics, including calculating the average and standard deviation of the power factor of users at different power levels, and calculating the windowed cross-correlation coefficient between the voltage series and the current series;

[0026] Extracting electricity consumption entropy features, including calculating the sample entropy of the user's active power sequence;

[0027] Based on cosine similarity, obtain the first behavior stability index for each user;

[0028] The stability index of each user's first behavior is obtained, including:

[0029] Based on the user's current period's electricity consumption behavior feature vector, a cosine similarity is calculated with the average of its historical feature vectors over the past M periods, and the value of the similarity is used as the first behavior stability index.

[0030] The load form characteristics, voltage-current coupling relationship characteristics, and electricity entropy characteristics are combined to form the electricity consumption behavior feature vector of each user, and together with the first behavior stability index, they constitute the user behavior feature set.

[0031] The K-means++ clustering algorithm is used to cluster the electricity consumption behavior feature vectors of all users to obtain K homogeneous user groups.

[0032] For each user, calculate the Mahalanobis distance from their electricity consumption behavior feature vector to their cluster center, which is used as the initial group deviation for the corresponding user.

[0033] The initial group deviation is corrected using the user's first behavior stability index to obtain the corresponding user's group deviation.

[0034] The initial group deviation is corrected using the user's first behavior stability index. The correction method is as follows:

[0035] Group deviation = Initial group deviation / (First behavior stability index + α), where α is a preset smoothing factor.

[0036] In this embodiment of the application, the power consumption data of each user in the transformer area and the power consumption data of the transformer area's main meter are collected. The power consumption data includes at least time series data of voltage, current, and active power. Based on the power consumption data of each user, the power consumption behavior feature vector and the first behavior stability index of each user are extracted to obtain the user behavior feature set. Based on the preset clustering algorithm, the user behavior feature set is clustered to obtain multiple homogeneous user groups and the group deviation of each user.

[0037] Specifically, firstly, the system acquires time-series data of voltage, current, and active power uploaded by smart meters of each user within the distribution area through an electricity consumption information collection system. For example, it collects time-series data of voltage, current, and active power recorded by each user's smart meter at fixed intervals of 15 minutes over the past month.

[0038] Furthermore, through the distribution automation system, the voltage, current, and active power time-series data of the transformer area master meter for the same time period are obtained. For example, the voltage, current, and active power time-series data of the past month are also collected at fixed 15-minute intervals.

[0039] Furthermore, the collected time-series data undergoes data cleaning and time-series alignment to form a multi-source data set. Specifically, data cleaning mainly involves removing null values ​​or obviously erroneous outliers, while time-series alignment ensures that user data and master table data correspond perfectly in timestamps, ultimately forming a complete and consistent multi-source data set.

[0040] Furthermore, load pattern characteristics are extracted, including calculating the dynamic time warping distance vector of the user's daily load curve, and statistical characteristics such as maximum, minimum, average, standard deviation, and peak-to-valley difference. For example, a user's maximum load might be 2.1 kW, minimum 0.3 kW, average 0.8 kW, standard deviation 0.5 kW, and peak-to-valley difference 1.8 kW. The dynamic time warping distance vector of the user's daily load curve is a vector characterizing the local differences in shape between the two curves. For example, a typical workday is selected as the baseline daily load curve, and different daily load curves of the user are obtained. The dynamic time warping algorithm is used to calculate the optimal matching path between the user's load curve and the baseline curve. The sequence of distance values ​​at each point on the matching path is the dynamic time warping distance vector.

[0041] Furthermore, voltage-current coupling characteristics are extracted, including calculating the average and standard deviation of the power factor for users at different power levels, and calculating the windowed cross-correlation coefficient between the voltage and current sequences. Specifically, the average and standard deviation of the power factor are calculated by analyzing the power factor of users at different power consumption levels, and the cross-correlation coefficient between the voltage and current change sequences within the sliding time window, such as the Pearson coefficient, is obtained.

[0042] Furthermore, the electricity consumption entropy features are extracted. For example, the sample entropy of the user's active power sequence is calculated. The higher the sample entropy value, the more complex and unpredictable the sequence. For instance, the electricity consumption sequence of a user with a very regular work-rest schedule may have a low sample entropy, such as 0.3; while the electricity consumption sequence of a user with an irregular work-rest schedule and frequent start-stop of appliances may have a high sample entropy, such as 0.8.

[0043] Furthermore, a first behavioral stability index is obtained for each user based on cosine similarity. Specifically, the cosine similarity is calculated between the user's current period's electricity consumption behavior feature vector and the average of its historical feature vectors over the past M periods. This similarity value is used as the first behavioral stability index. For example, the user's current period's electricity consumption behavior feature vector (e.g., this week) is obtained, and its cosine similarity is calculated with the average of its historical feature vectors over the past M periods (e.g., the previous four weeks). This cosine similarity value ranges from 0 to 1; the closer the value is to 1, the more similar the user's current behavior is to past habits, and therefore the more stable the behavior.

[0044] Furthermore, the load form characteristics, voltage-current coupling relationship characteristics, and electricity entropy characteristics are combined to form the electricity consumption behavior feature vector of each user, which together with the first behavior stability index constitutes the user behavior feature set.

[0045] Furthermore, the K-means++ clustering algorithm is used to cluster the electricity consumption behavior feature vectors of all users, resulting in K homogeneous user groups. For example, users can be divided into groups such as "low-energy-consumption stable users" and "high-energy-consumption fluctuating users," with each homogeneous user group having similar electricity consumption characteristics.

[0046] Furthermore, for each user, the Mahalanobis distance from their electricity consumption behavior feature vector to their cluster center is calculated as the initial group deviation for that user. Mahalanobis distance is used to represent the distance between a point and a distribution, and is used to initially measure the difference between a user's behavior and the average behavior of their homogeneous user group.

[0047] Furthermore, the initial group deviation is corrected using the user's first behavior stability index to obtain the corresponding user's group deviation.

[0048] Specifically, the group deviation is calculated as: Initial Group Deviation / (First Behavior Stability Index + α), where α is a preset smoothing factor to prevent calculation failures due to a denominator of 0. For example, α can be set to α = 0.1. For instance, user A's initial deviation is 2.5, stability index is 0.95, and final group deviation is 2.5 / (0.95 + 0.1) = 2.38. The corrected group deviation reflects both the static difference between the user and the group and the dynamic instability of their own behavior.

[0049] By extracting multi-dimensional features encompassing load patterns, voltage-current coupling, and electricity entropy to form an electricity consumption behavior feature vector, and calculating the first behavioral stability index reflecting the temporal stability of user behavior, a user behavior feature set capable of comprehensively and deeply characterizing user electricity consumption patterns is constructed. Based on this, a clustering algorithm is used to group users with similar electricity consumption behaviors into homogeneous user groups, establishing a dynamic and accurate reference benchmark for each user. Furthermore, by calculating the group deviation of each user relative to the center of their group, the degree of difference between user behavior and the group's normal state is effectively quantified, realizing a shift from extensive individual analysis to refined group intelligence analysis, laying a scientifically reliable benchmark for subsequent anomaly detection.

[0050] S20: Based on the total electricity data of the transformer area, identify abnormal line loss events in the transformer area and record the time period of the abnormal event. During the time period of the abnormal event, calculate the correlation strength between each user's electricity consumption behavior and the abnormal line loss event in the transformer area as the event correlation degree, and calculate the dynamic anomaly confidence degree by combining the user's first behavior stability index.

[0051] After identifying anomalies in line loss across an entire transformer substation, the main challenge of existing technologies is how to accurately correlate macroscopic substation anomalies with microscopic, specific user behaviors over time. Traditional methods often rely on simple comparisons of user electricity consumption during the abnormal period, failing to reveal the intrinsic causal or strong correlation between changes in user electricity consumption behavior and the abnormal line loss event, thus hindering the effective identification of the root cause of the anomaly. Furthermore, user behavior itself is inherently unstable; some users' electricity consumption changes during abnormal periods may be due to their inherent random behavioral habits, rather than a direct cause of the line loss anomaly. Ignoring these differences in behavioral stability and relying solely on the strength of the correlation can easily lead to misjudging users with habitually irregular electricity consumption as the source of the anomaly, resulting in false alarms.

[0052] Step S20 in the method provided in this application embodiment includes:

[0053] Calculate the daily line loss rate of the transformer area and establish a line loss rate baseline based on historical line loss rate data;

[0054] When the daily line loss rate continuously exceeds the preset threshold or suddenly increases relative to the baseline, it is marked as a line loss abnormal event.

[0055] Record the start time, end time, and severity of the abnormal event;

[0056] Extract the current sequence for each user and the total abnormal power sequence for the distribution area during the period in which the abnormal event occurred;

[0057] For each user, the Granger causality between their current sequence and the total abnormal power sequence of the transformer area is calculated, and the F-test statistic is used as the degree of event association.

[0058] Based on the user's event correlation degree and first behavior stability index, the dynamic anomaly confidence degree of the corresponding user is calculated, where dynamic anomaly confidence degree = event correlation degree × (1 - first behavior stability index).

[0059] In this embodiment of the application, based on the total electricity data of the transformer area, abnormal line loss events in the transformer area are identified and the time period of the abnormal event is recorded. During the time period of the abnormal event, the correlation strength between each user's electricity consumption behavior and the abnormal line loss event in the transformer area is calculated as the event correlation degree. The dynamic anomaly confidence degree is calculated by combining the user's first behavior stability index.

[0060] Specifically, first, calculate the daily line loss rate for the transformer substation area, establishing a baseline based on historical line loss rate data. Calculate the total daily active power of the transformer substation's main meter and the total daily active power of all users. Further, use the formula: Daily line loss rate = (Total daily active power of the transformer substation's main meter - Total daily active power of all users) / Total daily active power of the transformer substation's main meter * 100%. For example, if the total daily active power of a transformer substation's main meter is 1000 kWh and the total daily active power of all users is 950 kWh, then the daily line loss rate is (1000-950) / 1000×100% = 5%. Obtain historical daily line loss rate data for the past 30 days and calculate its arithmetic mean as the baseline; for example, the obtained baseline value is 4.5%.

[0061] Furthermore, when the daily line loss rate continuously exceeds a preset threshold or experiences a sudden increase relative to the baseline, it is marked as an abnormal line loss event. For example, a preset threshold of 7% and a relative increase threshold can be set. For instance, the relative increase threshold is a 50% increase compared to the baseline. Continuously monitor the daily line loss rate: if the line loss rate on a certain day exceeds 7%, it is marked as abnormal. If the line loss rate on a certain day does not exceed 7%, but the increase relative to the baseline of 4.5% exceeds 50% (i.e., line loss rate > 4.5% * (1 + 50%) = 6.75%), it is also marked as abnormal.

[0062] Furthermore, the start time, end time, and intensity of the abnormal event are recorded. For example, starting from October 5th, the line loss rates for three consecutive days were 6.8%, 7.5%, and 8.0%, respectively. 6.8% exceeded the surge threshold, and the subsequent days all exceeded the absolute threshold (7%), thus marking an abnormal event that started on October 5th and ended on October 7th. The intensity of this abnormal event is exemplarily represented by the difference between the average line loss rate during the event period and the baseline: (6.8+7.5+8.0) / 3-4.5=2.43%.

[0063] Furthermore, the current series and total abnormal power series of the transformer area for each user are extracted during the period of the abnormal event. For example, for the identified abnormal event from October 5th to October 7th, the time-series current data and total abnormal power series data of the transformer area for users A and B are extracted for these three days. Granger causality tests are then performed on the time-series current data of users A and B and the time-series total abnormal power series data of the transformer area. The Granger causality test is a statistical hypothesis test used to determine whether the past value of one variable helps predict the current value of another variable. This can be exemplarily performed using the `grangercausalitytests` function from Python's `statsmodels` library. The F-statistic is then obtained, which outputs an F-statistic. The larger the F-value, the greater the likelihood that the user's current series is a Granger cause of the total abnormal power series of the transformer area. For example, the F-value for user A might be 9.5, while the F-value for user B might be 1.2. This indicates that the correlation between user A's electricity consumption behavior and this abnormal line loss event is much stronger than that between user B and user A.

[0064] Furthermore, based on the user's event correlation and first behavior stability index, the dynamic anomaly confidence level for the corresponding user is calculated, where dynamic anomaly confidence level = event correlation × (1 - first behavior stability index). For example, user A has an event correlation of 9.5 and a first behavior stability index of 0.65. Therefore, the dynamic anomaly confidence level = 9.5 × (1 - 0.65) = 9.5 × 0.35 = 3.325. For a user whose behavior is inherently unstable (i.e., has a low first behavior stability index), if their behavior is highly correlated with the abnormal event, then the confidence level that this user caused the anomaly, i.e., the dynamic anomaly confidence level, is higher.

[0065] By first accurately identifying abnormal line loss events in the transformer substation and their occurrence times, the analysis focuses on key time windows. Based on this, by calculating the correlation between each user's electricity consumption behavior and the abnormal line loss events, effective tracing from macroscopic anomalies to microscopic user behavior is achieved, enabling the keen detection of users whose behavior exhibits suspicious changes simultaneously with the occurrence of anomalies. More importantly, this application incorporates a first behavioral stability index into the analysis, and by calculating the dynamic anomaly confidence level, the credibility of the event correlation is corrected. This mechanism effectively reduces the risk of misjudging users who, due to unstable electricity consumption habits, occasionally exhibit high correlation during abnormal periods, thus making the final dynamic anomaly confidence level more accurately reflect the potential causal possibility between user behavior and line loss anomalies.

[0066] S30: Calculate the comprehensive anomaly score for each user based on the group deviation and dynamic anomaly confidence level, and filter out abnormal users based on the preset comprehensive anomaly score threshold.

[0067] After calculating the group deviation degree and dynamic anomaly confidence degree from two dimensions—static deviation of user behavior relative to the group and dynamic correlation between user behavior and specific abnormal events—the key to accurately identifying abnormal users lies in how to scientifically and reasonably integrate these two different but complementary indicators.

[0068] Step S30 in the method provided in this application embodiment includes:

[0069] The group deviation and dynamic anomaly confidence scores were normalized respectively.

[0070] Using preset weighting coefficients, the normalized group deviation and dynamic anomaly confidence are weighted and summed to obtain a comprehensive anomaly score;

[0071] Set a threshold for the overall anomaly score;

[0072] Users whose overall anomaly score exceeds the threshold will be marked as abnormal users;

[0073] Generate an abnormal user location report, including a list of abnormal users, a comprehensive abnormal score, and a description of the main abnormal features;

[0074] Based on the abnormal user location report, output maintenance work orders to guide on-site troubleshooting.

[0075] In this embodiment of the application, a comprehensive anomaly score is calculated for each user based on the group deviation and dynamic anomaly confidence level, and abnormal users are selected based on a preset comprehensive anomaly score threshold.

[0076] Specifically, firstly, the group deviation and dynamic anomaly confidence scores are normalized. For example, a min-max normalization method is used, with the formula: Normalized value = (Original value for a user - Minimum value among all users) / (Maximum value among all users - Minimum value among all users). The group deviation and dynamic anomaly confidence scores for all users are then calculated separately to obtain the normalized group deviation and dynamic anomaly confidence scores.

[0077] Furthermore, using preset weighting coefficients, the normalized group deviation and dynamic anomaly confidence are weighted and summed to obtain a comprehensive anomaly score. For example, the comprehensive anomaly score = (weighting coefficient W1 × normalized group deviation) + (weighting coefficient W2 × normalized dynamic anomaly confidence). The weighting coefficients W1 and W2 are preset values, and W1 + W2 = 1. For example, setting W1 = 0.4 and W2 = 0.6 indicates a greater emphasis on the user's dynamic performance in specific anomaly events. For example, user A's normalized group deviation is 0.625, and their normalized dynamic anomaly confidence is 0.82. Therefore, their comprehensive anomaly score = (0.4 × 0.625) + (0.6 × 0.82) = 0.25 + 0.492 = 0.742. The higher the score, the greater the likelihood that the user caused the line loss anomaly.

[0078] Furthermore, a comprehensive anomaly score threshold is set, for example, based on historical anomaly analysis experience, the comprehensive anomaly score threshold is set to 0.7.

[0079] Furthermore, users whose overall abnormal scores exceed the threshold are marked as abnormal users.

[0080] Furthermore, an abnormal user location report is generated, including a list of abnormal users, a comprehensive abnormal score, and a description of the main abnormal characteristics. For example, for user A, the comprehensive abnormal score is 0.742, and the abnormal characteristic description is "This user's behavior pattern deviates significantly from its group and shows a high correlation with recent abnormal line loss events, indicating low stability in its own electricity consumption behavior." Location reports for all abnormal users are generated.

[0081] Furthermore, based on the abnormal user location report, maintenance work orders are generated to guide on-site troubleshooting. The maintenance work orders clearly identify the target users requiring on-site investigation and their core abnormal characteristics, allowing maintenance personnel to prioritize and target their on-site inspections, significantly improving the efficiency and accuracy of troubleshooting.

[0082] By constructing a comprehensive anomaly score, information from two dimensions—group deviation and dynamic anomaly confidence—is effectively integrated. Through normalization of these two indicators and weighted summation with preset weights, a multi-faceted and quantitative comprehensive evaluation of user anomaly probability is achieved, avoiding the arbitrariness of relying on a single criterion and making the final judgment more comprehensive and objective. Screening anomalous users based on a preset comprehensive anomaly score threshold can more accurately identify users who deviate from the norm in behavioral patterns and show a high degree of correlation in specific anomalous events, thereby greatly improving the accuracy and reliability of the location results.

[0083] In summary, the embodiments of this application have at least the following technical effects:

[0084] This application proposes a transformer substation line loss simulation method based on multi-source data. By comprehensively utilizing time-series data such as voltage, current, and active power from various users and the main meter within the transformer substation, a complete technical solution is constructed, encompassing feature extraction, behavioral clustering, and abnormal event correlation analysis. This significantly improves the accuracy and reliability of identifying abnormal users in transformer substation line loss. Specifically, this application extracts multi-dimensional features such as load characteristics, voltage-current coupling relationships, and electricity entropy to generate a user behavior feature set, which can more comprehensively represent the actual electricity consumption patterns of users. Furthermore, it uses clustering algorithms to divide homogeneous user groups and calculate group deviation, effectively identifying abnormal individuals whose behavior patterns deviate from the normal group. Simultaneously, by performing time-series correlation analysis between abnormal line loss events and user behavior, and calculating event correlation degree and dynamic anomaly confidence, it can keenly capture users with suspicious behavior during specific abnormal periods. Finally, it generates a comprehensive anomaly score by combining the user group deviation and dynamic anomaly confidence, providing a scientific and quantitative basis for accurately screening abnormal users. The synergistic effect of these technologies enables this application to effectively distinguish between normal load fluctuations and potential management line losses such as electricity theft and metering faults, significantly reducing false alarms and missed alarms common in traditional methods and improving the reliability of the identification results. Compared with traditional methods, the technical solution provided in this application significantly overcomes the limitations of relying solely on simple data comparison and static threshold judgment, achieving in-depth mining and dynamic evaluation of user electricity consumption behavior. This results in a comprehensive improvement in the refinement and intelligence of distribution area line loss management, providing maintenance personnel with clear and specific abnormal user location reports and maintenance work order guidance, thereby greatly improving the pertinence and efficiency of on-site investigations and providing strong technical support for ensuring the stable operation of the power grid and reducing power loss.

[0085] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this specification. Additionally, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are possible or may be advantageous.

[0086] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

[0087] This specification and accompanying drawings are merely illustrative examples of this application and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from its scope. Therefore, if such modifications and modifications fall within the scope of this application and its equivalents, this application intends to include such modifications and modifications.

Claims

1. A method for simulating line loss in transformer substations based on multi-source data, characterized in that, The method includes: The system collects electricity data from each user within the transformer substation area and the substation's main meter. The electricity data includes at least time-series data of voltage, current, and active power. Based on this user electricity data, it extracts a user's electricity consumption behavior feature vector and a first behavior stability index to obtain a user behavior feature set. Using a preset clustering algorithm, it clusters the user behavior feature set to obtain multiple homogeneous user groups and the group deviation for each user. The extracted first behavior stability index for each user includes: Based on the user's current period's electricity consumption behavior feature vector, a cosine similarity is calculated with the average of its historical feature vectors over the past M periods, and the value of the similarity is used as the first behavior stability index. Based on the total electricity data of the transformer substation, abnormal line loss events are identified, and the time periods of these events are recorded. Within these periods, the correlation strength between each user's electricity consumption behavior and the abnormal line loss events is calculated as the event correlation degree. This correlation is then combined with the user's first behavior stability index to calculate the dynamic anomaly confidence level, including: Extract the current sequence for each user and the total abnormal power sequence for the distribution area during the period in which the abnormal event occurred; For each user, the Granger causality between their current sequence and the total abnormal power sequence of the transformer area is calculated, and the F-test statistic is used as the degree of event association. Based on the user's event correlation degree and first behavior stability index, the dynamic anomaly confidence degree of the corresponding user is calculated, where dynamic anomaly confidence degree = event correlation degree × (1 - first behavior stability index); Based on each user's group deviation and dynamic anomaly confidence, a comprehensive anomaly score is calculated for each user. Based on a preset comprehensive anomaly score threshold, abnormal users are selected.

2. The method for simulating transformer line loss based on multi-source data according to claim 1, characterized in that, Collect electricity data from each user within the distribution area and the total electricity data from the distribution area's main meter. The electricity data includes at least time-series data of voltage, current, and active power, including: The electricity information collection system acquires time-series data of voltage, current, and active power uploaded by smart meters of each user in the distribution area. The distribution automation system acquires the time-series data of voltage, current, and active power of the main meter in the distribution area for the same time period. The collected time-series data is cleaned and time-series aligned to form a multi-source data set.

3. The method for simulating transformer line loss based on multi-source data according to claim 1, characterized in that, Based on the electricity data of each user, the electricity consumption behavior feature vector and the first behavior stability index of each user are extracted to obtain the user behavior feature set, including: Extract load pattern characteristics, including calculating the dynamic time-normalized distance vector of the user's daily load curve, and statistical characteristics including maximum value, minimum value, average value, standard deviation, and peak-to-valley difference; Extract voltage-current coupling characteristics, including calculating the average and standard deviation of the power factor of users at different power levels, and calculating the windowed cross-correlation coefficient between the voltage series and the current series; Extracting electricity consumption entropy features, including calculating the sample entropy of the user's active power sequence; Based on cosine similarity, obtain the first behavior stability index for each user; The load form characteristics, voltage-current coupling relationship characteristics, and electricity entropy characteristics are combined to form the electricity consumption behavior feature vector of each user, which together with the first behavior stability index constitutes the user behavior feature set.

4. The method for simulating transformer line loss based on multi-source data according to claim 1, characterized in that, Based on a pre-defined clustering algorithm, user behavior feature sets are clustered to obtain multiple homogeneous user groups and the group deviation of each user, including: The K-means++ clustering algorithm is used to cluster the electricity consumption behavior feature vectors of all users to obtain K homogeneous user groups. For each user, calculate the Mahalanobis distance from their electricity consumption behavior feature vector to their cluster center, which is used as the initial group deviation for the corresponding user. The initial group deviation is corrected using the user's first behavior stability index to obtain the corresponding user's group deviation.

5. The method for simulating transformer line loss based on multi-source data according to claim 4, characterized in that, The initial group deviation is corrected using the user's first behavior stability index. The correction method is as follows: Group deviation = Initial group deviation / (First behavior stability index + α), where α is a preset smoothing factor.

6. The method for simulating transformer line loss based on multi-source data according to claim 1, characterized in that, Based on the total electricity data of the transformer substations, identify abnormal line loss events in the substations and record the time periods of these abnormal events, including: Calculate the daily line loss rate of the transformer area and establish a line loss rate baseline based on historical line loss rate data; When the daily line loss rate continuously exceeds the preset threshold or suddenly increases relative to the baseline, it is marked as a line loss abnormal event. Record the start time, end time, and severity of the abnormal event.

7. The method for simulating transformer line loss based on multi-source data according to claim 1, characterized in that, Based on each user's group deviation and dynamic anomaly confidence, a comprehensive anomaly score is calculated for each user, including: The group deviation and dynamic anomaly confidence scores were normalized respectively. Using preset weighting coefficients, the normalized group deviation and dynamic anomaly confidence are weighted and summed to obtain a comprehensive anomaly score.

8. The method for simulating transformer line loss based on multi-source data according to claim 1, characterized in that, Based on a preset comprehensive anomaly score threshold, abnormal users are filtered out, including: Set a threshold for the overall anomaly score; Users whose overall anomaly score exceeds the threshold will be marked as abnormal users; Generate an abnormal user location report, including a list of abnormal users, a comprehensive abnormal score, and a description of the main abnormal features; Based on the abnormal user location report, output maintenance work orders to guide on-site troubleshooting.