Multi-source data diagnosis method and system oriented to theoretical line loss calculation of main and distribution networks

By monitoring the convergence state of the calculation and analyzing the gradient using a differentiable digital twin model, a repair data package that conforms to the characteristics of normal data distribution is generated. This solves the problem of multi-source data quality, ensures the successful convergence and accuracy of theoretical line loss calculation, and improves the efficiency and accuracy of data governance in power grid management.

CN121881231APending Publication Date: 2026-04-17ZHANGJIAKOU POWER SUPPLY COMPANY OF STATE GRID JINBEI ELECTRIC POWER COMPANY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-27
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

In existing technologies, the quality problems of multi-source data lead to the distortion of theoretical line loss calculation results, especially the failure of power flow calculation to converge successfully or the distortion of results, which affects the economic assessment of the power grid and network optimization.

Method used

By monitoring the convergence state of the computation, analyzing the gradient using a differentiable digital twin model, accurately locating key sensitive datasets, generating repair data packages that conform to the characteristics of normal data distribution, and selecting the optimal repair scheme through multi-objective optimization decision-making, the computation is ensured to converge successfully and the results are accurate.

Benefits of technology

It enables accurate diagnosis and repair of multi-source data, improves the accuracy and reliability of theoretical line loss calculation, and provides a solid data foundation for refined power grid management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121881231A_ABST
    Figure CN121881231A_ABST
Patent Text Reader

Abstract

The invention provides a multi-source data diagnosis method and system oriented to main and distribution network theoretical line loss calculation, and relates to the technical field of power grids. According to the method, the diagnosis process is intelligently triggered by monitoring and calculating the convergence state, and the key sensitive data set causing non-convergence is accurately positioned by using gradient analysis, so that accurate positioning of diagnosis is realized, and the efficiency and pertinence of data management are improved. And then generating a plurality of repair data packets based on historical normal data distribution, performing multi-objective optimization decision in multiple groups of successfully converged candidate schemes to ensure that the repair scheme achieves optimal balance among convergence stability, data authenticity and result robustness, and ensuring that theoretical line loss calculation is successfully converged and meanwhile, improving the repair efficiency of the repair scheme. The accuracy and reliability of calculation results are improved, the problem of theoretical line loss calculation result distortion caused by data quality problems is solved, diagnosis and repair of multi-source data in the theoretical line loss calculation process of the main distribution network are achieved, and a solid data foundation is provided for power grid refined line loss management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power grid technology, and in particular to a multi-source data diagnostic method and system for calculating theoretical line losses in main and distribution networks. Background Technology

[0002] Theoretical line loss calculation is a core analytical task in power system operation, planning, and management. Its accuracy directly affects the economic assessment of the power grid, network optimization, and loss reduction decisions. Modern theoretical line loss calculation, especially power flow calculation based on numerical iterative methods such as the Newton-Raphson method, relies heavily on the completeness, accuracy, and consistency of the input data for successful convergence and accurate results.

[0003] Currently, the data supporting theoretical line loss calculations for the main grid, distribution network, and even transformer substations originates from multiple heterogeneous business systems, including Geographic Information Systems (GIS), Supervisory Control and Data Acquisition (SCADA) systems, and electricity consumption information collection systems. This multi-source nature presents significant challenges to data governance. Specifically, archival data suffers from missing or incorrectly entered equipment nameplate parameters (such as line resistance and reactance). Topology data exhibits discrepancies between electrical connections and actual field conditions, including isolated data islands or faulty loops. Operational data suffers from timing mismatches, abnormal dead data, or noise from communication interference. These data quality issues directly lead to singular Jacobian matrices, iterative oscillations, or non-convergence in power flow calculations, rendering theoretical line loss calculations impossible or producing severely distorted results. Consequently, subsequent line loss analysis and management lack a reliable basis. Summary of the Invention

[0004] This invention provides a multi-source data diagnostic method and system for calculating theoretical line loss in main distribution networks, which solves the problem of distortion in theoretical line loss calculation results caused by data quality issues, and realizes the diagnosis and repair of multi-source data in the process of calculating theoretical line loss in main distribution networks.

[0005] Firstly, this invention provides a multi-source data diagnostic method for calculating theoretical line loss in a main distribution network. The method includes: monitoring the convergence state of the theoretical line loss calculation in the main distribution network; if the convergence state is an abnormal state of non-convergence, extracting the multi-source data that triggered the abnormal state, analyzing the gradient of the abnormal state relative to various types of data in the multi-source data, and locating key sensitive datasets; based on the key sensitive datasets, and the historical normal data distribution and reasonable value range of various types of data, generating multiple repair data packets that conform to the characteristics of normal data distribution through sampling; based on the multiple repair data packets and a preset differentiable digital twin model, performing theoretical line loss calculation trials, and selecting candidate data packets that have successfully converged; based on the candidate data packets, using convergence stability, data authenticity, and result robustness as optimization objectives, performing multi-objective optimization decisions, and selecting multi-source data repair schemes.

[0006] In one possible implementation, multi-source data triggering abnormal states are extracted, and the gradients of the abnormal states relative to various types of data in the multi-source data are analyzed to locate key sensitive datasets. This includes: quantifying the abnormal states where theoretical line loss calculations fail to converge in a pre-defined differentiable digital twin model to obtain a loss function; based on the multi-source data triggering abnormal states, using the adjoint sensitivity analysis method and the inverse automatic differentiation function of the differentiable digital twin model, calculating the gradient vectors of the loss function with respect to various types of data in the multi-source data; calculating the absolute values ​​of the gradients corresponding to various types of data in the gradient vectors; and selecting data whose absolute gradient values ​​are greater than a pre-defined threshold to constitute the key sensitive dataset.

[0007] In one possible implementation, based on a key sensitive dataset and the historical normal data distribution and reasonable value range of various types of data, multiple repair data packages conforming to the characteristics of normal data distribution are generated through sampling. This includes: using the key sensitive dataset as the generation condition, a pre-trained generative model is used to perform random sampling to generate multiple preliminary candidate data; the generative model learns the joint probability distribution of multi-source data of the power grid through historical normal data; the multiple preliminary candidate data are fused with the original complete dataset that has not triggered anomalies to form multiple complete data packages for data repair; based on the reasonable value range of various types of data, the feasibility of the multiple complete data packages is verified according to the physical constraints of the power grid, and invalid data packages that do not conform to the reasonable value range are eliminated to obtain multiple repair data packages.

[0008] In one possible implementation, before generating multiple repair data packages that conform to the characteristics of normal data distribution by sampling, based on key sensitive datasets and historical normal data distributions and reasonable value ranges of various types of data, the following steps are included: collecting a health status dataset from the historical operation of the power grid, where the health status data satisfies the convergence of theoretical line loss calculations and the results conform to physical laws; preprocessing the health status dataset to construct a sample set for model training, the samples containing archive parameters, topological connectivity relationships, and operational measurement data; training the generative model based on the sample set using a conditional variational autoencoder as the basic architecture; wherein, the encoder of the generative model learns the probability distribution of the health status data and maps it to the latent space, while the decoder samples from the latent space and reconstructs data that conforms to the probability distribution according to given conditional variables; introducing power grid physical constraints as regularization terms in model training to guide the data generated by the model to satisfy the power grid physical constraints, which include Kirchhoff's current law and the basic physical laws of the allowable range of voltage amplitude; and outputting the generative model after completing the model training.

[0009] In one possible implementation, based on multiple repair data packets and a pre-defined differentiable digital twin model, theoretical line loss calculation trials are performed to screen candidate data packets that have successfully converged. This includes: performing calculations in the initial stage of the theoretical line loss calculation trial process based on multiple repair data packets and a pre-defined differentiable digital twin model to obtain initial stage calculation results; the initial stage calculation results include the initial residual norm and Jacobian matrix condition number of the power flow equation corresponding to each repair data packet; based on the initial stage calculation results and the initial stage dynamic screening strategy, preliminary screening is performed to obtain multiple first data packets; the initial stage dynamic screening strategy is to eliminate repair data packets whose initial residual norm is greater than a first threshold or whose Jacobian matrix condition number is greater than a second threshold; based on the multiple first data packets and the pre-defined differentiable digital twin model, theoretical line loss calculation trials are performed to obtain candidate data packets that have successfully converged. A differentiable digital twin model is used to perform calculations during the iterative phase of the theoretical line loss calculation trial process, obtaining the calculation results of the iterative phase. The calculation results of the iterative phase include the residual decrease rate of each first data packet calculation process in the first few iterations. Based on the calculation results of the iterative phase and the dynamic screening strategy of the iterative phase, a secondary screening is performed to obtain multiple second data packets. The dynamic screening strategy of the iterative phase is to eliminate the first data packets corresponding to calculation processes with residual decrease rates lower than a third threshold. Based on the multiple second data packets and the preset differentiable digital twin model, the theoretical line loss calculation trial is completed to determine the convergence state of the multiple second data packets and the number of iterations to reach convergence. Based on the convergence state of the multiple second data packets and the number of iterations to reach convergence, a final screening is performed to obtain multiple candidate data packets.

[0010] In one possible implementation, based on candidate data packets, multi-objective optimization decisions are made with convergence stability, data authenticity, and result robustness as optimization objectives to select multi-source data repair schemes. This includes: performing reciprocal operations on the total number of iterations and residuals recorded during the theoretical line loss calculation trial for each candidate data packet to obtain a convergence stability index for each candidate data packet; analyzing data differences between the repaired data and the original measurement data and typical equipment parameter library in each candidate data packet, and calculating a weighted sum of squared residuals to obtain a data authenticity index for each candidate data packet; and determining the numerical values ​​of each candidate data packet... Monte Carlo sampling is performed within the neighborhood to generate a small perturbation dataset. Based on the small perturbation dataset, theoretical line loss is calculated to obtain multiple sets of line loss results. Based on the multiple sets of line loss results, standard deviation is calculated to obtain the robustness index of each candidate data packet. Based on the convergence stability index, data authenticity index, and robustness index of each candidate data packet, a multi-objective optimization algorithm is used to solve for the Pareto front to obtain the Pareto optimal solution set. Based on the preset engineering decision rules, the Pareto optimal solution set is used to select a scheme to determine the final multi-source data repair scheme. The engineering decision rules include the priority order of each optimization objective.

[0011] In one possible implementation, based on preset engineering decision rules, a scheme selection process is performed on the Pareto optimal solution set to determine the final multi-source data repair scheme. This includes: sorting and filtering all solutions in the Pareto optimal solution set based on their convergence stability indices to obtain a first candidate subset; the first candidate subset consists of the top N solutions in terms of convergence stability indices, or solutions whose convergence stability indices are higher than a preset stability threshold; within the first candidate subset, a second sorting and filtering process is performed based on the data authenticity indices of each solution to obtain a second candidate subset; the second candidate subset consists of the solutions with the highest data authenticity indices, or solutions with the top M data authenticity indices; within the second candidate subset, a final decision is made based on the robustness indices of each solution to determine the final multi-source data repair scheme; the final decision is to select the solution with the highest robustness indices as the final scheme.

[0012] In one possible implementation, based on candidate data packets, and with convergence stability, data authenticity, and result robustness as optimization objectives, a multi-objective optimization decision is made to select multi-source data repair schemes. The process then includes: applying the multi-source data repair schemes to theoretical line loss calculations to obtain theoretical line loss calculation results; if the theoretical line loss calculation results are successfully converged, evaluating the rationality and physical consistency of the theoretical line loss calculation results to obtain verification results; if the verification results are successful, generating new knowledge samples based on the key sensitive datasets in the diagnostic and repair process corresponding to the multi-source data repair schemes, successful repair data packets, and multi-objective evaluation indicators; generating an incremental sample set based on the new knowledge samples and the sample set of the generative model; and incrementally learning and optimizing the generative model based on the incremental sample set to obtain an updated generative model.

[0013] In one possible implementation, the method further includes: identifying weak links in the power grid based on the theoretical line loss calculation results that have successfully converged according to the multi-source data repair scheme; extracting key factors that lead to high line loss based on the weak links, including unreasonable network structure, heavily loaded lines, low-voltage operating ranges, or renewable energy access points; and generating theoretical line loss mitigation strategy recommendations based on the key factors, including network reconfiguration schemes, reactive power compensation configuration schemes, or power source optimization layout schemes.

[0014] Secondly, embodiments of the present invention provide a multi-source data diagnostic device for calculating theoretical line loss in a main distribution network. The device includes a communication module and a processing module. The communication module monitors the convergence state of the theoretical line loss calculation in the main distribution network. The processing module, if the convergence state is an abnormal state of non-convergence, extracts the multi-source data that triggers the abnormal state, analyzes the gradient of the abnormal state relative to various types of data in the multi-source data, and locates the key sensitive dataset. Based on the key sensitive dataset, and the historical normal data distribution and reasonable value range of various types of data, multiple repair data packets conforming to the characteristics of normal data distribution are generated through sampling. Based on the multiple repair data packets and a preset differentiable digital twin model, theoretical line loss calculation trials are performed, and candidate data packets that have successfully converged are selected. Based on the candidate data packets, multi-objective optimization decisions are made with convergence stability, data authenticity, and result robustness as optimization objectives, and multi-source data repair schemes are selected.

[0015] Thirdly, embodiments of the present invention provide a multi-source data diagnostic system for calculating theoretical line losses in main distribution networks. The system includes an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor is used to call and run the computer program stored in the memory to perform the steps of the method as described in the first aspect and any possible implementation thereof.

[0016] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing a computer program, characterized in that, when executed by a processor, the computer program implements the steps of the method as described in the first aspect and any possible implementation thereof.

[0017] This invention provides a multi-source data diagnostic method and system for calculating theoretical line losses in main and distribution networks. The invention intelligently triggers the diagnostic process by monitoring the convergence status of the calculation and utilizes gradient analysis to accurately locate key sensitive datasets causing convergence failures, achieving precise diagnostic localization and improving the efficiency and targeting of data governance. Subsequently, multiple repair data packages are generated based on historical normal data distribution. Through multi-objective optimization decisions among multiple successfully converged candidate schemes, the final selected repair scheme achieves an optimal balance among convergence stability, data authenticity, and result robustness. This ensures successful convergence of theoretical line loss calculations while improving the accuracy and reliability of the calculation results, providing a solid data foundation for refined line loss management in the power grid. It solves the problem of distorted theoretical line loss calculation results caused by data quality issues, realizing the diagnosis and repair of multi-source data during the calculation of theoretical line losses in main and distribution networks. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a flowchart illustrating a multi-source data diagnostic method for calculating theoretical line loss in a main distribution network, provided by an embodiment of the present invention. Figure 2 This is a schematic diagram of the structure of a multi-source data diagnostic device for calculating theoretical line loss in a main distribution network, provided in an embodiment of the present invention. Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0020] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of the invention. However, those skilled in the art will understand that the invention can be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods are omitted so as not to obscure the description of the invention with unnecessary detail.

[0021] In the description of this invention, unless otherwise stated, " / " means "or". For example, A / B can mean A or B. The term "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. Furthermore, "at least one" and "more than one" refer to two or more. The terms "first," "second," etc., do not limit the quantity or order of execution, and "first," "second," etc., do not necessarily imply differences.

[0022] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner to facilitate understanding.

[0023] Furthermore, the terms "comprising" and "having," and any variations thereof, used in the description of this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or modules is not limited to the steps or modules listed, but may optionally include other steps or modules not listed, or may optionally include other steps or modules inherent to such process, method, product, or device.

[0024] To make the objectives, technical solutions, and advantages of the present invention clearer, the following description will be provided in conjunction with the accompanying drawings and specific embodiments.

[0025] like Figure 1 As shown, this embodiment of the invention provides a multi-source data diagnostic method for calculating theoretical line losses in main distribution networks. The method includes steps S101-S104.

[0026] S101. Monitor the convergence status of the theoretical line loss calculation of the main distribution network.

[0027] In some embodiments, the convergence status is determined based on whether the power flow calculation residual decreases to an acceptable range within a preset number of iterations. If yes, the calculation converges; otherwise, the calculation does not converge.

[0028] For example, embodiments of the present invention can monitor the convergence status of the calculation in real time during the theoretical line loss calculation process. Specifically, convergence is determined by checking indicators such as the residual norm, the condition number of the Jacobian matrix, and the number of iterations during the calculation iteration process. When the residual norm cannot be reduced to within the preset tolerance, the Jacobian matrix becomes ill-conditioned, or the number of iterations exceeds the maximum allowable value, it is determined to be an abnormal state of non-convergence.

[0029] S102. If the convergence state is an abnormal state that does not converge, then extract the multi-source data that triggered the abnormal state, analyze the gradient of the abnormal state relative to various types of data in the multi-source data, and locate the key sensitive dataset.

[0030] In some embodiments, when an abnormal state is detected, multi-source data used in the current calculation is extracted, including power grid equipment parameters, topology, load data, generator output, etc. Subsequently, a differentiable digital twin model is used to calculate the gradient of the abnormal state (e.g., high residual) relative to various data types in the multi-source data. Gradient analysis is performed through sensitivity calculations; data with large gradient values ​​indicate a significant impact on computational convergence. Based on a comparison of the absolute value of the gradient with a preset threshold, key sensitive datasets are selected; these datasets are typically the main root causes of computational non-convergence.

[0031] As one possible implementation, step S102 can be specifically implemented as steps S1021-S1024.

[0032] S1021. In the preset differentiable digital twin model, the abnormal state where the theoretical line loss calculation does not converge is quantified to obtain the loss function.

[0033] For example, in a differentiable digital twin model, anomalies where theoretical line loss calculations fail to converge can be quantified into a loss function. This loss function is typically designed based on the computational residuals; for instance, the norm of the power flow equation residuals can be used as the loss value, with larger residuals indicating a higher degree of anomaly.

[0034] S1022. Based on multi-source data that triggers abnormal states, the adjoint sensitivity analysis method is adopted, and the gradient vector of the loss function with respect to various types of data in the multi-source data is calculated through the inverse automatic differentiation function of the differentiable digital twin model.

[0035] For example, embodiments of the present invention can employ the adjoint sensitivity analysis method, utilizing the backpropagation automatic differentiation function of a differentiable digital twin model to calculate the gradient vector of the loss function with respect to various types of data from multiple sources. Specifically, firstly, the theoretical line loss is calculated forward in the digital twin model to obtain the loss function; then, through the backpropagation algorithm, the partial derivatives of the loss function with respect to the input data (such as node power, line parameters, etc.) are calculated to form the gradient vector. Each component in the gradient vector represents the sensitivity of the corresponding data to the loss function.

[0036] S1023. Calculate the absolute value of the gradient corresponding to each type of data in the gradient vector.

[0037] S1024. Select data whose absolute gradient value is greater than a preset threshold to form a key sensitive dataset.

[0038] For example, embodiments of the present invention can calculate the absolute value of the gradient corresponding to various types of data in the gradient vector. The larger the absolute value of the gradient, the more significant the impact of that data on the non-convergence of the calculation. A preset threshold is set based on historical anomaly cases or statistical experience, such as using the median or a specific quantile of the absolute gradient value as the threshold. Data with absolute gradient values ​​exceeding the threshold are selected to form a key sensitive dataset. This process automatically locates the core points of data anomalies, providing precise targets for subsequent repair.

[0039] S103. Based on the key sensitive dataset, as well as the historical normal data distribution and reasonable value range of various types of data, multiple repair data packages that conform to the characteristics of normal data distribution are generated through sampling.

[0040] In some embodiments, for key sensitive datasets, multiple repair data packages are generated through sampling, combining historical normal data distribution and reasonable value ranges. The historical normal data distribution is obtained by statistically analyzing data characteristics under power grid health conditions, and the reasonable value range is set based on equipment specifications and physical laws. The sampling process utilizes a generative model (such as a conditional variational autoencoder) to ensure that the generated data conforms to normal statistical characteristics and satisfies power grid physical constraints (such as voltage amplitude range, power balance, etc.). The generated repair data packages are then fused with the original non-abnormal data to form a complete candidate dataset.

[0041] As one possible implementation, step S103 can be specifically implemented as steps S1031-S1033.

[0042] S1031. Using key sensitive datasets as generation conditions, a pre-trained generative model is used to perform random sampling to generate multiple preliminary candidate data.

[0043] In some embodiments, the generative model learns the joint probability distribution of multi-source data from the power grid by learning from historical normal data.

[0044] For example, in embodiments of the present invention, the identified key sensitive datasets can be used as inputs to a pre-trained generative model. This generative model is a deep probabilistic generative network, whose core value lies in its mastery of the complex correlations and joint probability distribution characteristics among multiple sources of power grid data through learning from a large amount of historical normal data.

[0045] Upon receiving critical and sensitive data, the generative model randomly samples within its learned normal data distribution space. This process is not a simple random replacement, but rather an intelligent generation based on probability distribution, capable of producing multiple preliminary candidate data that both meet the given conditions (i.e., are related to the critical and sensitive data) and follow historical normal patterns. These generated data exhibit a high degree of consistency with real health data in terms of statistical characteristics.

[0046] S1032. Merge multiple preliminary candidate data with the original complete dataset that has not triggered anomalies to form multiple complete data packages for data repair.

[0047] For example, embodiments of the present invention can fuse and replace the generated preliminary candidate data with the original complete dataset that has not triggered anomalies, forming multiple complete data repair packages that can be used for theoretical line loss calculation. This process ensures that only the identified problematic data is repaired, while other normal data that is not marked as sensitive remains unchanged, thus preserving the original appearance of the data to the greatest extent possible.

[0048] S1033. Based on the reasonable value range of various types of data, perform feasibility verification on multiple complete data packets according to the physical constraints of the power grid, eliminate invalid data packets that do not conform to the reasonable value range, and obtain multiple repair data packets.

[0049] For example, in embodiments of the present invention, after data fusion is completed, a feasibility verification must be performed on each complete data packet. This step filters the generated data based on the physical constraints and reasonable value ranges of the actual operation of the power grid. The physical constraints of the power grid on which the verification is based include equipment capacity limits, line transmission limits, voltage operating ranges, power factor requirements, etc. If any data packet contains data points that violate these hard constraints (for example, generating power values ​​exceeding the rated capacity of the transformer), the data packet will be considered an invalid data packet and discarded. After this round of rigorous verification and screening, the data packets that are finally retained constitute multiple repair data packets that can be used for subsequent trial calculations. They are not only statistically reasonable but also physically feasible.

[0050] S104. Based on multiple repair data packets and a preset differentiable digital twin model, perform theoretical line loss calculations and select candidate data packets that have successfully converged in the calculation.

[0051] In some embodiments, the present invention inputs each repair data packet into a differentiable digital twin model for theoretical line loss calculation trials. The trial process includes an initial stage and an iterative stage, where data packets with poor computational performance are eliminated through a dynamic screening strategy. The initial stage evaluates the initial residuals of the power flow equations and the condition number of the Jacobian matrix, eliminating data packets with poor conditions; the iterative stage monitors the residual decay rate, eliminating data packets with slow convergence. Finally, only data packets that have successfully converged and have undergone fewer iterations are retained as candidate data packets.

[0052] As one possible implementation, step S104 can be specifically implemented as steps S1041-S1046.

[0053] S1041. Based on multiple repair data packets and a preset differentiable digital twin model, calculations are performed in the initial stage of the theoretical line loss calculation trial process to obtain the calculation results of the initial stage.

[0054] In some embodiments, the calculation results in the initial stage include the initial residual norm and Jacobian matrix condition number of the power flow equation corresponding to each repair data packet.

[0055] S1042. Based on the calculation results of the initial stage and the dynamic filtering strategy of the initial stage, preliminary filtering is performed to obtain multiple first data packets.

[0056] In some embodiments, the dynamic filtering strategy in the initial stage is to eliminate repair data packets whose initial residual norm is greater than a first threshold or whose Jacobian matrix condition number is greater than a second threshold.

[0057] For example, in this embodiment of the invention, each repair data packet can be input into a preset differentiable digital twin model at this stage to initiate the theoretical line loss calculation trial process. The calculation is not performed to completion but is interrupted at the initial stage to quickly capture early key indicators that can predict the final convergence performance. The initial stage results of the calculation output mainly include two core indicators: Initial residual norm: This indicator reflects the degree of deviation of the power flow equation at the initial iteration point. An excessively large initial residual usually means that the operating point defined by the data packet is far from the physical equilibrium point, making convergence extremely difficult. Jacobian matrix condition number: This indicator measures the ill-conditioned nature of the mathematical model used to solve the power flow equation. An excessively large condition number indicates that the model is extremely sensitive to small changes in data, has poor numerical stability, and is prone to iteration divergence.

[0058] Based on the above metrics, a dynamic filtering strategy is implemented in the initial stage. This strategy sets two thresholds: the first threshold is for the initial residual norm, and the second threshold is for the Jacobian matrix condition number. Any repair data packet whose initial residual norm exceeds the first threshold or whose Jacobian matrix condition number exceeds the second threshold is considered inherently deficient and is immediately eliminated. This stage can quickly filter out a large number of obviously unqualified data packets, significantly improving the efficiency of subsequent calculations.

[0059] S1043. Based on multiple first data packets and a preset differentiable digital twin model, calculations are performed in the iterative stage of the theoretical line loss calculation trial process to obtain the calculation results of the iterative stage.

[0060] In some embodiments, the calculation results of the iteration phase include the residual decrease rate of the first data packet calculation process in the previous several iterations.

[0061] S1044. Based on the calculation results of the iteration stage and the dynamic filtering strategy of the iteration stage, a second filtering is performed to obtain multiple second data packets.

[0062] In some embodiments, the dynamic filtering strategy in the iteration phase is to eliminate the first data packet corresponding to the computation process whose residual decrease rate is lower than a third threshold.

[0063] For example, embodiments of the present invention can use multiple first data packets retained from the initial screening to enter a deeper iterative stage for computation. In this stage, the computation process is allowed to undergo several iterations, and its convergence dynamics are closely monitored. The core monitoring metric is the residual descent rate, which quantifies how quickly the equation residual decreases after each iteration. A healthy, convergent computation process should exhibit a stable and rapid decreasing trend in its residuals. Based on the dynamic screening strategy of the iterative stage, a third threshold is set as the minimum acceptable residual descent rate. For computation processes whose residual descent rate is below this threshold, exhibiting convergence stagnation or excessively slow speed, their corresponding first data packets will be judged as having poor convergence performance and thus eliminated in the secondary screening. This stage aims to identify and retain data packets with strong and stable convergence momentum.

[0064] S1045. Based on multiple second data packets and a preset differentiable digital twin model, complete the theoretical line loss calculation trial, determine the convergence state of multiple second data packets, and the number of iterations required to achieve convergence.

[0065] S1046. Based on the convergence status of multiple second data packets and the number of iterations required to reach convergence, perform final screening to obtain multiple candidate data packets.

[0066] For example, in this embodiment of the invention, multiple second data packets remaining after the first two rounds of screening will be allowed to complete a full theoretical line loss calculation trial until the convergence criterion or the maximum number of iterations is reached. Finally, the convergence status (successful convergence or failure) of each second data packet is confirmed, and the number of iterations required to reach convergence is recorded. In the final screening, all data packets that failed to converge are first eliminated. Then, among those data packets that successfully converged, the number of iterations required to reach convergence becomes an important selection criterion. Generally, the fewer the number of iterations, the more stable and efficient the computational model corresponding to the data packet. Therefore, this invention selects data packets with fewer iterations to form the final multiple candidate data packets for the next stage of multi-objective optimization decision-making.

[0067] S105. Based on candidate data packets, with convergence stability, data authenticity and result robustness as optimization objectives, multi-objective optimization decision-making is carried out to select multi-source data repair schemes.

[0068] In some embodiments, the present invention can perform multi-objective optimization on candidate data packets from three dimensions: convergence stability, data authenticity, and result robustness. Convergence stability is achieved by calculating the number of iterations and the inverse of the residual; data authenticity is achieved by weighted calculation of the differences between the repaired data and the original measurement data and a typical parameter library; result robustness is assessed by introducing small perturbations and calculating the standard deviation of the line loss results. A multi-objective optimization algorithm (such as Pareto optimization) is used to solve for the optimal solution set, and the final repair scheme is selected by combining engineering decision rules (such as priority ranking).

[0069] As one possible implementation, step S105 can be specifically implemented as steps S1051-S1057.

[0070] S1051. Based on the total number of iterations and residuals recorded during the theoretical line loss calculation trial for each candidate data packet, perform reciprocal operations to obtain the convergence stability index of each candidate data packet.

[0071] S1052. Based on the repaired data and original measurement data in each candidate data packet, as well as the typical parameter library of the equipment, analyze the data differences, perform weighted residual sum of squares calculation, and obtain the data authenticity index of each candidate data packet.

[0072] S1053. Perform Monte Carlo sampling within the numerical neighborhood of each candidate data packet to generate a small perturbation dataset.

[0073] S1054. Based on the small disturbance dataset, perform theoretical line loss calculations to obtain multiple sets of line loss results.

[0074] S1055. Based on multiple sets of line loss results, standard deviation is calculated to obtain the robustness index of each candidate data packet.

[0075] For example, the convergence stability metric is as follows: For each candidate data packet, key performance data recorded throughout the theoretical line loss calculation trial is extracted, mainly including the total number of iterations required to achieve convergence and the residual size at the final (or key iteration step). By performing a reciprocal operation on these basic data, the original values ​​are transformed into positive metrics, meaning that the larger the value, the more stable the convergence performance and the higher the computational efficiency.

[0076] For example, a data authenticity metric assesses the degree of deviation between the restored data and the original field measurement data, as well as the typical parameter library of standard equipment. This is specifically achieved by calculating the weighted sum of squared residuals between the restored data and these benchmark data. The weighting strategy can be adjusted according to the importance of the data and the measurement accuracy; the smaller the metric value, the closer the restored data is to reality, and the less tampering with the original data.

[0077] For example, the robustness metric for the results is as follows: To test the anti-interference capability of the line loss calculation results, multiple datasets containing small perturbations are generated within the numerical neighborhood of each candidate data packet using the Monte Carlo sampling method. Theoretical line loss calculations are then performed using these perturbation datasets to obtain multiple sets of line loss results. Subsequently, the standard deviation of these results is calculated. The smaller the standard deviation, the more stable the output results of the candidate data packet are when faced with small fluctuations in the input data, i.e., the better the robustness of the results.

[0078] S1056. Based on the convergence stability index, data authenticity index, and result robustness index of each candidate data packet, a multi-objective optimization algorithm is used to solve the Pareto front to obtain the Pareto optimal solution set.

[0079] S1057. Based on preset engineering decision rules, select solutions from the Pareto optimal solution set to determine the final multi-source data repair solution; the engineering decision rules include the priority order of each optimization objective.

[0080] For example, after obtaining the above three metrics, a multi-objective optimization algorithm (such as a non-dominated sorting genetic algorithm) is used to solve the problem. This algorithm can effectively find a balance point in dealing with multiple potentially conflicting objectives and output a Pareto optimal solution set. In this solution set, any improvement in the performance of one objective in any solution will inevitably lead to a decrease in the performance of at least one other objective.

[0081] Finally, based on pre-defined engineering decision-making rules, the final repair solution is selected from the Pareto optimal solution set. These rules clarify the priority order of each optimization objective; for example, in engineering practice, priority may be given to ensuring computational convergence, followed by data accuracy, and finally the robustness of the results. Through this systematic quantitative evaluation and decision-making process, it is ensured that the selected solution achieves comprehensive optimality across multiple key dimensions.

[0082] For example, step S1057 can be specifically implemented as steps A1-A3.

[0083] A1. Based on the convergence stability index of all solutions in the Pareto optimal solution set, sort and filter them to obtain the first candidate subset.

[0084] In some embodiments, the first candidate subset consists of the top N solutions in terms of convergence stability index, or consists of solutions in terms of convergence stability index that are higher than a preset stability threshold.

[0085] For example, the primary task is to ensure that the repaired data can reliably support the convergence of theoretical line loss calculations. Therefore, all candidate solutions in the Pareto optimal solution set are first sorted according to their convergence stability index. The group with the best performance is then selected to form the first candidate subset.

[0086] In this embodiment of the invention, a fixed number (top N) of solutions with the highest convergence stability index can be selected; alternatively, the invention can set a minimum performance threshold, i.e., a preset stability threshold, where all solutions with indices exceeding this threshold can be selected. This step ensures that subsequent decisions are based on computational stability and reliability.

[0087] A2. Within the first candidate subset, a second sorting and filtering is performed based on the data authenticity index of each solution to obtain the second candidate subset.

[0088] In some embodiments, the second candidate subset consists of the solution with the highest data authenticity index, or the solution with the top M data authenticity index.

[0089] For example, within the successfully converged solution set, the next step is to maximize the authenticity of the repaired data and minimize unnecessary modifications to the original data. Within the first candidate subset, a second sorting and filtering is performed based on data authenticity metrics. Similarly, a strategy can be adopted to select the single solution with the highest authenticity metric, or to select multiple solutions with the highest ranking (the top M), thus forming the second candidate subset. This step embodies the principle of remaining as faithful to the original data as possible while ensuring computational feasibility.

[0090] A3. Within the second candidate subset, based on the robustness index of each solution, make the final decision and determine the final multi-source data repair scheme.

[0091] In some embodiments, the final decision is to select the solution with the highest robustness index as the final solution.

[0092] For example, the final decision is made from the high-quality solutions remaining after the first two rounds of screening (i.e., the second candidate subset). At this point, the robustness index becomes the key deciding factor. The final decision rule is straightforward: within the second candidate subset, the solution with the highest robustness index is selected as the definitive final multi-source data repair solution. This ensures that the selected solution can provide the most stable and reliable theoretical line loss calculation results when faced with the unavoidable minor fluctuations in the actual data.

[0093] This invention provides a multi-source data diagnostic method for calculating theoretical line losses in main and distribution networks. It intelligently triggers the diagnostic process by monitoring the convergence status of the calculation and uses gradient analysis to accurately locate key sensitive datasets causing convergence failures, achieving precise diagnostic localization and improving the efficiency and targeting of data governance. Subsequently, multiple repair data packages are generated based on historical normal data distribution. Through multi-objective optimization decisions among multiple successfully converged candidate schemes, the final selected repair scheme achieves an optimal balance between convergence stability, data authenticity, and result robustness. This ensures successful convergence of theoretical line loss calculations while improving the accuracy and reliability of the calculation results, providing a solid data foundation for refined line loss management in the power grid. It solves the problem of distorted theoretical line loss calculation results caused by data quality issues, realizing the diagnosis and repair of multi-source data during the calculation of theoretical line losses in main and distribution networks.

[0094] Optionally, the multi-source data diagnostic method for calculating theoretical line loss in the main distribution network provided in this embodiment of the invention further includes steps S201-S204 before step S103.

[0095] S201. Collect the health status dataset from the historical operation of the power grid.

[0096] In some embodiments, the health status data satisfies the theoretical convergence of line loss calculation and the results conform to physical laws.

[0097] S202. Preprocess the health status dataset to construct a sample set for model training.

[0098] In some embodiments, the sample includes archive parameters, topology connections, and operational measurement data.

[0099] For example, embodiments of the present invention require collecting datasets from massive amounts of historical power grid operation data that are in a healthy state. The selection criteria for these data are: theoretical line loss calculations based on this dataset not only converge successfully, but the calculation results are also verified and conform to the physical laws of power grid operation (such as energy conservation, voltage stability, etc.). Subsequently, the collected raw data undergoes preprocessing, including data cleaning (removing obvious noise and errors), normalization (eliminating the influence of dimensions), and feature integration, thereby constructing a high-quality sample set suitable for model training. Each sample contains the necessary archival parameters for line loss calculation (such as resistance, reactance), topological connectivity relationships (such as node-branch models), and operational measurement data (such as power injection, voltage amplitude).

[0100] S203. Based on the sample set, the conditional variational autoencoder is used as the basic architecture for training the generative model.

[0101] In this generative model, the encoder learns the probability distribution of health status data and maps it to the latent space, while the decoder samples from the latent space and reconstructs data that conforms to the probability distribution based on given condition variables.

[0102] S204. In the model training, introduce power grid physical constraints as regularization terms to guide the data generated by the model to meet the power grid physical constraints. The power grid physical constraints include Kirchhoff's current law and the basic physical laws of the allowable range of voltage amplitude.

[0103] For example, embodiments of the present invention require the use of a conditional variational autoencoder as the basic architecture of the generative model. This architecture consists of two parts: an encoder and a decoder. The encoder is responsible for learning the probability distribution of health status data, mapping high-dimensional input data to a low-dimensional, statistically significant latent space, and expressing the core features of the data in the form of a probability distribution (usually assumed to be a normal distribution) within this space.

[0104] The decoder is responsible for sampling from the latent space based on the given condition variables (which are the key sensitive datasets in this invention) and reconstructing the sampled points into new data that conforms to the probability distribution of the original data.

[0105] In model training, a novel approach is to introduce power grid physical constraints as a regularization term. This means that, in addition to the standard model loss function (such as reconstruction error), an extra penalty term is added to measure the degree to which the generated data violates physical laws. For example, the generated data can be verified to ensure it satisfies Kirchhoff's Current Law (KCL) or that the voltage amplitude is within acceptable limits. Through this mechanism, the model is guided to intrinsically understand and adhere to the fundamental physical laws of the power grid while learning the data distribution, thereby generating data that is both realistic and feasible.

[0106] S205. After completing the model training, output the generative model.

[0107] For example, after completing the above training, the resulting generative model possesses powerful data generation and repair capabilities. It is no longer a simple data copying tool, but an intelligent agent that deeply understands the inherent laws and physical constraints of power grid data, and can provide high-quality, highly reliable repair candidate data for subsequent data diagnostic processes.

[0108] Thus, by introducing a conditional variational autoencoder with physical constraint regularization terms, this invention enables generative models to not only learn the probability distribution of historical health data, but also internalize the basic physical laws of the power grid. This ensures that all generated repair data possesses both statistical rationality and physical feasibility, fundamentally improving the quality and success rate of data repair.

[0109] Optionally, the multi-source data diagnostic method for calculating theoretical line loss in the main distribution network provided in this embodiment of the invention further includes steps S301-S305 after step S105.

[0110] S301. Apply the multi-source data repair scheme to the theoretical line loss calculation to obtain the theoretical line loss calculation results.

[0111] S302. If the theoretical line loss calculation result is successfully converged, then evaluate the rationality and physical consistency of the theoretical line loss calculation result to obtain the verification result.

[0112] For example, embodiments of the present invention can apply the determined multi-source data repair scheme to the actual theoretical line loss calculation environment of the main distribution network, perform calculations, and obtain actual line loss calculation results. First, it is determined whether the calculation has successfully converged. If it has successfully converged, the calculation results are further evaluated in depth to verify their rationality and physical consistency. This evaluation includes, but is not limited to: analyzing whether the line loss distribution conforms to the network topology characteristics, and verifying whether the calculation results satisfy the basic law of energy conservation. Only the result that passes this comprehensive verification is considered the final valid verification result.

[0113] S303. If the verification result is successful, then a new knowledge sample is generated based on the key sensitive datasets, successful repair data packages, and multi-objective evaluation indicators in the diagnosis and repair process corresponding to the multi-source data repair scheme.

[0114] S304. Based on the new knowledge samples and the sample set of the generative model, generate an incremental sample set.

[0115] S305. Based on the incremental sample set, perform incremental learning and optimization on the generative model to obtain the updated generative model.

[0116] For example, once the repair solution is verified to be effective, this invention automatically integrates the key elements of the diagnostic and repair process, including the initially identified critical sensitive dataset, the ultimately successful repair data package, and various indicators from the multi-objective evaluation, to generate a new, high-quality knowledge sample. These new knowledge samples, together with the original training sample set of the generative model, constitute an expanded incremental sample set. Using this incremental sample set, the pre-trained generative model is incrementally learned and optimized. This process enables the model to absorb and learn from the experience of the latest successful cases without forgetting existing knowledge, thereby dynamically updating its internal parameters. The final output is an evolved generative model with stronger data repair capabilities and higher adaptability.

[0117] Thus, by transforming successful repair cases into knowledge samples and incrementally learning the generative model, this invention establishes a self-optimizing closed loop, enabling it to continuously accumulate diagnostic experience and constantly improve its long-term performance in repairing future data anomalies and adaptability.

[0118] Optionally, the multi-source data diagnostic method for calculating theoretical line loss in the main distribution network provided in this embodiment of the invention further includes steps S401-S403.

[0119] S401. Based on the theoretical line loss calculation results that have successfully converged according to the multi-source data repair scheme, identify the weak links in the theoretical line loss of the power grid.

[0120] For example, embodiments of the present invention can perform in-depth scanning analysis based on accurate and reliable theoretical line loss results calculated from repaired data, thereby precisely identifying high-loss links in the power grid, i.e., weak links in theoretical line loss. These links typically manifest as a line loss rate of a specific line, transformer, or regional network that is significantly higher than the reasonable level.

[0121] S402. Based on the theoretical weak links in line loss, extract the key factors that lead to high line loss.

[0122] In some embodiments, key factors include an inappropriate network structure, heavy-load lines, low-voltage operating ranges, or new energy access points.

[0123] For example, embodiments of the present invention can perform source tracing analysis on each identified weak link to extract and locate the key factors leading to high line losses. These factors are multifaceted and may include: unreasonable network structure, such as excessively long power supply radius or circuitous power supply; heavily loaded lines, i.e., lines that operate for extended periods near or exceeding their thermal stability limits; low-voltage operating ranges, leading to increased current and thus increased line losses; and the impact of the fluctuations in the output of specific renewable energy access points on local network losses.

[0124] S403. Based on key factors, generate theoretical recommendations for line loss mitigation strategies.

[0125] In some embodiments, governance strategy recommendations include network reconfiguration schemes, reactive power compensation configuration schemes, or power supply optimization layout schemes.

[0126] For example, embodiments of the present invention can automatically generate targeted theoretical line loss mitigation strategy recommendations based on clearly defined key factors. These recommendations are specific and actionable, such as: Network reconfiguration schemes: proposing optimized combinations of changing grid switching states to optimize power flow distribution and reduce overall network losses. Reactive power compensation configuration schemes: suggesting the addition or adjustment of the capacity and location of reactive power compensation equipment at specific nodes to improve voltage levels and reduce losses caused by reactive power flow. Power source optimization layout schemes: proposing optimization suggestions for the access location and capacity of distributed power sources (especially new energy sources) to mitigate their negative impact on grid losses.

[0127] Thus, by deeply analyzing the theoretical line loss results after data repair, this invention accurately locates the high-loss links and their root causes in the power grid, and automatically generates specific governance strategies such as network reconfiguration and reactive power compensation. It directly transforms the data diagnosis results into decision support for reducing losses and increasing efficiency, significantly enhancing the practical value of lean operation and maintenance in the power grid.

[0128] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0129] The following are device embodiments of the present invention. For details not described in detail, please refer to the corresponding method embodiments described above.

[0130] Figure 2 This diagram illustrates the structure of a multi-source data diagnostic device for calculating theoretical line losses in a main distribution network, according to an embodiment of the present invention. The diagnostic device 500 includes a communication module 501 and a processing module 502.

[0131] The communication module 501 is used to monitor the convergence status of the theoretical line loss calculation of the main distribution network.

[0132] The processing module 502 is used to extract the multi-source data that triggered the abnormal state when the convergence state is an abnormal state that does not converge, analyze the gradient of the abnormal state relative to various types of data in the multi-source data, and locate the key sensitive dataset; based on the key sensitive dataset, as well as the historical normal data distribution and reasonable value range of various types of data, multiple repair data packages that conform to the normal data distribution characteristics are generated by sampling; based on the multiple repair data packages and the preset differentiable digital twin model, theoretical line loss calculation is performed, and candidate data packages that have successfully converged are selected; based on the candidate data packages, multi-objective optimization decision is made with convergence stability, data authenticity and result robustness as optimization objectives, and multi-source data repair schemes are selected.

[0133] Figure 3This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. The electronic device 600 includes: a processor 601, a memory 602, and a computer program 603 stored in the memory 602 and executable on the processor 601. When the processor 601 executes the computer program 603, it implements the steps in the above-described method embodiments. Alternatively, when the processor 601 executes the computer program 603, it implements the functions of each module / unit in the above-described device embodiments.

[0134] For example, the computer program 603 may be divided into one or more modules / units, which are stored in the memory 602 and executed by the processor 601 to complete the present invention. The one or more modules / units may be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of the computer program 603 in the electronic device 600.

[0135] The processor 601 may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.

[0136] The memory 602 can be an internal storage unit of the electronic device 600, such as a hard disk or memory of the electronic device 600. The memory 602 can also be an external storage device of the electronic device 600, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD) card, flash card, etc., equipped on the electronic device 600. Furthermore, the memory 602 can include both internal and external storage units of the electronic device 600. The memory 602 is used to store the computer program and other programs and data required by the terminal. The memory 602 can also be used to temporarily store data that has been output or will be output.

[0137] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A multi-source data diagnostic method for calculating theoretical line loss in main and distribution networks, characterized in that, include: Monitor the convergence status of the theoretical line loss calculation in the main distribution network; If the convergence state is an anomalous state that does not converge, then the multi-source data that triggered the anomalous state is extracted, the gradient of the anomalous state relative to various types of data in the multi-source data is analyzed, and the key sensitive dataset is located. Based on the aforementioned key sensitive dataset, as well as the historical normal data distribution and reasonable value range of various types of data, multiple repair data packages that conform to the characteristics of normal data distribution are generated through sampling. Based on the multiple repair data packets and the preset differentiable digital twin model, theoretical line loss calculations are performed to screen out candidate data packets that have successfully converged in the calculation. Based on the candidate data packets, multi-objective optimization decisions are made with convergence stability, data authenticity, and result robustness as optimization objectives to select multi-source data repair schemes.

2. The multi-source data diagnostic method for calculating theoretical line loss in main and distribution networks according to claim 1, characterized in that, The process involves extracting multi-source data that triggers abnormal states, analyzing the gradient of the abnormal states relative to various data types within the multi-source data, and identifying key sensitive datasets, including: In the pre-defined differentiable digital twin model, the abnormal state where the theoretical line loss calculation does not converge is quantified to obtain the loss function; Based on the multi-source data that triggered the abnormal state, the adjoint sensitivity analysis method is adopted, and the gradient vector of the loss function with respect to various types of data in the multi-source data is calculated through the reverse automatic differentiation function of the differentiable digital twin model. Calculate the absolute value of the gradient corresponding to each type of data in the gradient vector; Data whose absolute gradient value is greater than a preset threshold are selected to form the key sensitive dataset.

3. The multi-source data diagnostic method for calculating theoretical line loss in main and distribution networks according to claim 1, characterized in that, Based on the key sensitive dataset and the historical normal data distribution and reasonable value range of various types of data, multiple repair data packages that conform to the characteristics of normal data distribution are generated through sampling, including: Using the aforementioned key sensitive dataset as the generation condition, a pre-trained generative model is used to perform random sampling to generate multiple preliminary candidate data; the generative model learns the joint probability distribution of multi-source data of the power grid through historical normal data; The multiple preliminary candidate data are fused with the original complete dataset that has not triggered any anomalies to form multiple complete data packages for data repair. Based on the reasonable value range of various data, the feasibility of multiple complete data packets is verified according to the physical constraints of the power grid. Invalid data packets that do not conform to the reasonable value range are eliminated, and multiple repair data packets are obtained.

4. The multi-source data diagnostic method for calculating theoretical line loss in main and distribution networks according to claim 3, characterized in that, Before generating multiple repair data packages that conform to the characteristics of normal data distribution through sampling based on the key sensitive dataset and the historical normal data distribution and reasonable value range of various types of data, the process also includes: Collect a health status dataset from the historical operation of the power grid. The health status data satisfies the convergence of theoretical line loss calculation and the results conform to physical laws. The health status dataset is preprocessed to construct a sample set for model training, the sample containing archive parameters, topological connectivity relationships and operational measurement data; Based on the sample set, the conditional variational autoencoder is used as the basic architecture of the generative model for training; wherein, the encoder of the generative model learns the probability distribution of health status data and maps it to the latent space, and the decoder samples from the latent space and reconstructs data that conforms to the probability distribution according to the given conditional variables. In model training, power grid physical constraints are introduced as regularization terms to guide the data generated by the model to meet the power grid physical constraints, which include Kirchhoff's current law and the basic physical laws of the allowable range of voltage amplitude. After completing the model training, a generative model is output.

5. The multi-source data diagnostic method for calculating theoretical line loss in main and distribution networks according to claim 1, characterized in that, The theoretical line loss calculation is performed based on the multiple repair data packets and a preset differentiable digital twin model to screen candidate data packets that have successfully converged in the calculation, including: Based on the multiple repair data packets and the preset differentiable digital twin model, calculations are performed in the initial stage of the theoretical line loss calculation trial process to obtain the calculation results of the initial stage; the calculation results of the initial stage include the initial residual norm and Jacobian matrix condition number of the power flow equation corresponding to each repair data packet. Based on the calculation results of the initial stage and the dynamic filtering strategy of the initial stage, preliminary filtering is performed to obtain multiple first data packets; the dynamic filtering strategy of the initial stage is to eliminate repair data packets whose initial residual norm is greater than a first threshold or whose Jacobian matrix condition number is greater than a second threshold. Based on the multiple first data packets and the preset differentiable digital twin model, calculations are performed in the iterative stage of the theoretical line loss calculation trial process to obtain the calculation results of the iterative stage; the calculation results of the iterative stage include the residual decrease rate of the first few iterations of the calculation process of each first data packet. Based on the calculation results of the iterative stage and the dynamic filtering strategy of the iterative stage, a second filtering is performed to obtain multiple second data packets; the dynamic filtering strategy of the iterative stage is to eliminate the first data packets corresponding to the calculation process whose residual decrease rate is lower than the third threshold. Based on the multiple second data packets and the preset differentiable digital twin model, a theoretical line loss calculation trial is completed to determine the convergence state of the multiple second data packets and the number of iterations required to achieve convergence. Based on the convergence status of multiple second data packets and the number of iterations required to reach convergence, a final selection is performed to obtain the multiple candidate data packets.

6. The multi-source data diagnostic method for calculating theoretical line loss in main and distribution networks according to claim 1, characterized in that, Based on the candidate data packets, a multi-objective optimization decision is made with convergence stability, data authenticity, and result robustness as optimization objectives to select multi-source data repair solutions, including: Based on the total number of iterations and residuals recorded during the theoretical line loss calculation trial for each candidate data packet, a reciprocal operation is performed to obtain the convergence stability index for each candidate data packet. Based on the repaired data and original measurement data in each candidate data packet, as well as the typical parameter library of the equipment, the data differences are analyzed, and the weighted sum of squared residuals is calculated to obtain the data authenticity index of each candidate data packet. Monte Carlo sampling is performed within the numerical neighborhood of each candidate data packet to generate a dataset with small perturbations; Based on the aforementioned small disturbance dataset, theoretical line loss is calculated, and multiple sets of line loss results are obtained. Based on the multiple sets of line loss results, the standard deviation is calculated to obtain the robustness index of each candidate data packet. Based on the convergence stability index, data authenticity index, and result robustness index of each candidate data packet, a multi-objective optimization algorithm is used to solve the Pareto front and obtain the Pareto optimal solution set. Based on preset engineering decision rules, scheme selection is performed on the Pareto optimal solution set to determine the final multi-source data repair scheme; the engineering decision rules include the priority order of each optimization objective.

7. The multi-source data diagnostic method for calculating theoretical line loss in main and distribution networks according to claim 6, characterized in that, The method of selecting solutions from the Pareto optimal solution set based on preset engineering decision rules to determine the final multi-source data repair solution includes: Based on the convergence stability index of all solutions in the Pareto optimal solution set, sorting and filtering are performed to obtain a first candidate subset; the first candidate subset consists of the solutions with the highest convergence stability index (ranked in the top N), or solutions with convergence stability index higher than a preset stability threshold. Within the first candidate subset, a second sorting and filtering process is performed based on the data authenticity index of each solution to obtain a second candidate subset; the second candidate subset consists of the solutions with the highest data authenticity index, or the solutions with the top M data authenticity index rankings. Within the second candidate subset, a final decision is made based on the robustness index of each solution to determine the final multi-source data repair scheme; the final decision is to select the solution with the highest robustness index as the final scheme.

8. The multi-source data diagnostic method for calculating theoretical line loss in main distribution networks according to any one of claims 1 to 7, characterized in that, After selecting a multi-source data repair scheme based on the candidate data packets, with convergence stability, data authenticity, and result robustness as optimization objectives, the process further includes: The multi-source data repair scheme was applied to theoretical line loss calculation to obtain the theoretical line loss calculation results. If the theoretical line loss calculation result is successfully converged, the rationality and physical consistency of the theoretical line loss calculation result are evaluated to obtain the verification result; If the verification result is successful, then a new knowledge sample is generated based on the key sensitive datasets, successful repair data packages, and multi-objective evaluation indicators in the diagnostic and repair process corresponding to the multi-source data repair scheme. Based on the new knowledge samples and the sample set of the generative model, an incremental sample set is generated; Based on the incremental sample set, the generative model is incrementally learned and optimized to obtain an updated generative model.

9. The multi-source data diagnostic method for calculating theoretical line loss in main distribution networks according to any one of claims 1 to 7, characterized in that, The method further includes: Based on the successfully converged theoretical line loss calculation results corresponding to the multi-source data repair scheme, weak links in the theoretical line loss of the power grid are identified. Based on the theoretical weak link of line loss, the key factors leading to high line loss are extracted. These key factors include unreasonable network structure, heavy-load lines, low-voltage operating range, or new energy access point. Based on the aforementioned key factors, theoretical line loss mitigation strategy recommendations are generated, including network reconfiguration schemes, reactive power compensation configuration schemes, or power supply optimization layout schemes.

10. A multi-source data diagnostic system for calculating theoretical line losses in main and distribution networks, characterized in that, The system includes an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor is configured to invoke and run the computer program stored in the memory to perform the method as described in any one of claims 1 to 9.