A multi-source data preprocessing method for transformer area line loss analysis and a readable storage medium

By constructing a local reference clock in the intelligent fusion terminal, utilizing a high-stability crystal oscillator and timing signal, and combining linear regression analysis, a transmission delay and clock drift model was established, which solved the problem of poor timing consistency of multi-source data in the distribution area and achieved high-precision line loss analysis.

CN122364663APending Publication Date: 2026-07-10STATE GRID JIANGSU ELECTRIC POWER CO ZHENJIANG POWER SUPPLY CO
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
STATE GRID JIANGSU ELECTRIC POWER CO ZHENJIANG POWER SUPPLY CO
Filing Date
2026-03-31
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

In existing technologies, clock drift and communication delay in multi-source data from distribution transformer areas lead to poor timing consistency, affecting the accuracy and precision of line loss analysis. Existing methods cannot effectively solve the problems of clock drift and communication delay.

Method used

By constructing a local reference clock for the intelligent fusion terminal, utilizing a high-stability crystal oscillator and timing signal, and combining linear regression analysis, a transmission delay and clock drift model is established to perform timestamp compensation and alignment of data, thereby achieving high-precision timing alignment of multi-source data.

Benefits of technology

It achieves sub-second timing alignment of multi-source data, eliminates line loss calculation errors, improves the accuracy and reliability of line loss analysis, adapts to network changes, and is low-cost and requires no hardware modification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122364663A_ABST
    Figure CN122364663A_ABST
Patent Text Reader

Abstract

This invention relates to the field of power system distribution network line loss management technology, and discloses a multi-source data preprocessing method and readable storage medium for transformer area line loss analysis. A local reference clock is constructed in an intelligent fusion terminal, which receives multi-source heterogeneous data from distributed nodes. The local timestamps and arrival timestamps of each multi-source heterogeneous data are obtained to construct transmission delay models between each distributed node and the intelligent fusion terminal, as well as clock drift models for each distributed node. The local timestamps of the data to be processed are input into the transmission delay model and the clock drift model to obtain the corresponding transmission delay compensation and clock deviation compensation. These are added to the local timestamps of the data to be processed to obtain a standard timestamp aligned with the local reference clock, completing the timing alignment of the data to be processed. This fundamentally eliminates line loss calculation errors caused by timing misalignment and achieves sub-second high-precision timing alignment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power system distribution network line loss management technology, and in particular to a multi-source data preprocessing method and readable storage medium for transformer substation line loss analysis. Background Technology

[0002] With the large-scale integration of distributed energy resources and diversified loads into the power distribution network, the operating status of transformer substations has become more complex and variable, placing higher demands on the accuracy of line loss analysis. The large-scale deployment of smart integrated terminals in transformer substations provides a massive multi-source heterogeneous data foundation for line loss analysis, including multi-frequency data such as voltage, current, and power collected by the main meter of the transformer substation, branch monitoring units, and user smart meters.

[0003] However, existing line loss analysis based on fusion terminals faces a fundamental challenge at the data preprocessing level: ensuring the temporal consistency of multi-source data is difficult. Fusion terminals, user meters distributed throughout the distribution area, and branch monitoring units typically use independent clock sources. Their inherent crystal oscillator frequency deviations and environmental temperature effects cause clock drift, accumulating over time to second- or even minute-level time discrepancies. Furthermore, data is transmitted from each distributed node to the fusion terminal via different communication paths (such as HPLC and low-power wireless), resulting in varying and potentially dynamic communication delays. This coupling effect of "clock drift" and "communication delay" means that the timestamps of the data collected at the fusion terminal cannot reflect the true, simultaneous electrical quantity relationships. In analyses requiring strict time alignment, such as power balance calculations and real-time line loss rate calculations, inaccurate timestamps lead to significant calculation errors, severely interfering with the accuracy of subsequent anomaly diagnosis and root cause localization.

[0004] In existing technologies, periodic time synchronization by the main station or access to GPS / BeiDou signals by the terminal can solve the local clock reference problem of the fusion terminal, but it cannot eliminate the clock drift of a large number of downstream user meters, branch monitoring units, and other nodes, nor can it compensate for their uncertain communication transmission delays. Some solutions attempt to perform simple delay estimations, but they do not fully consider the characteristics of data acquisition and communication in the distribution area (such as the frozen data reporting mechanism), resulting in coarse models, poor adaptability, and limited synchronization accuracy.

[0005] Therefore, there is an urgent need for a high-precision time-series alignment data preprocessing method that is deeply adapted to the characteristics of multi-source data acquisition and transmission in the distribution area, so as to fundamentally solve the analysis error problem caused by time-series misalignment. Summary of the Invention

[0006] Therefore, the technical problem to be solved by the present invention is to overcome the problem in the prior art that the poor timing consistency caused by the asynchronous clocks of multiple data sources and the lack of compensation for communication delays results in timing misalignment in the data acquisition of the transformer area, which in turn leads to low accuracy and large error in the line loss analysis of the transformer area.

[0007] To address the aforementioned technical problems, this invention provides a multi-source data preprocessing method for transformer substation line loss analysis, comprising: A local reference clock for the intelligent fusion terminal is constructed based on a preset timing signal and a high-stability crystal oscillator. The intelligent fusion terminal receives multi-source heterogeneous data from distributed nodes and obtains the timestamp events of each multi-source heterogeneous data; the timestamp events include the local timestamp of data generation and the arrival timestamp of data received by the intelligent fusion terminal; Based on the difference between the local timestamp and the arrival timestamp in multiple timestamp events in each distributed node, the transmission delay model between each distributed node and the intelligent fusion terminal is obtained. Linear regression is performed on the local timestamp and arrival timestamp of multiple timestamp events in each distributed node to obtain the clock drift model of each distributed node. Input the local timestamp of the data to be processed into the transmission delay model between the corresponding distributed node and the intelligent fusion terminal, as well as the clock drift model of the corresponding distributed node, to obtain the corresponding transmission delay compensation amount and clock deviation compensation amount. The local timestamp, transmission delay compensation, and clock skew compensation of the data to be processed are added together to obtain a standard timestamp aligned with the local reference clock, thus completing the timing alignment of the data to be processed.

[0008] Preferably, a local reference clock for the intelligent converged terminal in the distribution area is constructed based on a preset timing signal and a high-stability crystal oscillator, including: Obtain the standard timing signal from the main station or the signal from the local timing module as the preset timing signal; Based on the preset timing signal, the output frequency of the high-stability crystal oscillator is synchronized with the preset timing signal to form a local reference clock. The local time synchronization module signals include GPS signals and BeiDou signals.

[0009] Preferably, the multi-source heterogeneous data includes: the substation outlet data collected by the metering module of the substation intelligent fusion terminal itself, the user meter freezing data uploaded by the HPLC concentrator, and the branch monitoring data uploaded by the branch monitoring unit.

[0010] Preferably, the user's electricity meter periodically uploads hourly freeze data, and the branch monitoring unit periodically reports branch monitoring data.

[0011] Preferably, based on the difference between the local timestamp and the arrival timestamp in multiple timestamp events in each distributed node, a transmission delay model between each distributed node and the intelligent fusion terminal is obtained, including: For each distributed node, a difference sequence is constructed based on the difference between the local timestamp and the arrival timestamp in its multiple timestamp events; Find the minimum value in the difference sequence, and use it as the minimum delay; Calculate the standard deviation of all differences in the difference sequence as the jitter delay; Based on the principle that the sum of minimum delay and jitter delay equals transmission delay, a transmission delay model for distributed nodes is constructed.

[0012] Preferably, the transmission delay model is expressed as: ; in, Indicates the amount of transmission delay compensation. Indicates minimum delay. This indicates the preset balance coefficient. This indicates jitter delay.

[0013] Preferably, a linear regression is performed on the local timestamps and arrival timestamps of all data reporting events from each distributed node to obtain the clock drift model for each distributed node, including: For each distributed node, the difference between the arrival timestamp and the local timestamp in its corresponding timestamp event is decomposed into a fixed delay component, a variable delay component, and a data source clock drift component. Using time series analysis and linear regression, with the clock drift component of the data source as the independent variable and the clock skew compensation amount as the dependent variable, a clock drift model for each distributed node is fitted and obtained.

[0014] Preferably, the clock drift model is expressed as: ; in, This indicates the amount of clock skew compensation. Indicates clock drift rate, Indicates the initial frequency deviation. Represents the local timestamp. Indicates the reference start time.

[0015] Preferably, after completing the time sequence alignment of the data to be processed, the method further includes: performing data quality cleaning on the dataset containing the data to be processed; the data quality cleaning includes missing value completion, outlier correction, and deduplication.

[0016] This embodiment provides a computer-readable storage medium storing a computer program thereon. When the computer program is executed, it implements the steps of the multi-source data preprocessing method for transformer substation line loss analysis as described above.

[0017] Compared with the prior art, the above-described technical solution of the present invention has the following advantages: The multi-source data preprocessing method for transformer substation line loss analysis described in this invention analyzes the transmission delay between each distributed node and the intelligent fusion terminal, as well as the clock drift within the intelligent fusion terminal. It models and compensates for both the communication transmission delay and the clock drift of the distributed nodes, improving the accuracy of aligning multi-source data to a unified time base to within 100 milliseconds. This is far superior to the error levels of traditional methods, which range from several seconds to several minutes. It fundamentally eliminates line loss calculation errors caused by timing misalignment, achieving sub-second high-precision timing alignment. Furthermore, the transmission delay model and clock drift model can be updated periodically following data transmission, adapting to changes in network load and clock characteristics caused by equipment aging, maintaining high synchronization accuracy over the long term. By providing highly consistent multi-source data, this invention fundamentally guarantees the accuracy and reliability of advanced applications such as real-time power balance-based line loss calculation, transient event analysis, and phase and topology identification.

[0018] Meanwhile, this invention is based entirely on existing data collection and software algorithms, requiring no hardware upgrades or time synchronization modules for the massive deployment of user meters and branch monitoring units. It requires no hardware modification, is highly adaptable, has low implementation costs, and is easy to scale up. Attached Figure Description

[0019] To make the content of this invention easier to understand, the invention will be further described in detail below with reference to specific embodiments and accompanying drawings, wherein: Figure 1 This is a flowchart of the steps of the multi-source data preprocessing method for transformer substation line loss analysis of the present invention; Figure 2 This is a schematic diagram illustrating the principle of event-driven transmission delay and clock drift estimation. Detailed Implementation

[0020] The present invention will be further described below with reference to the accompanying drawings and specific embodiments, so that those skilled in the art can better understand and implement the present invention. However, the embodiments described are not intended to limit the present invention.

[0021] Reference Figure 1 The flowchart shown is a step-by-step flowchart of the multi-source data preprocessing method for transformer area line loss analysis of the present invention, and the specific steps include S101 to S106.

[0022] S101: Based on a preset timing signal and a high-stability crystal oscillator, a local reference clock is constructed for the intelligent fusion terminal, including: Obtain the standard timing signal from the main station or the signal from the local timing module as the preset timing signal; Based on the preset timing signal, the output frequency of the high-stability crystal oscillator is synchronized with the preset timing signal to form a local reference clock.

[0023] The local time synchronization module signals include GPS signals and BeiDou signals.

[0024] In this embodiment, the fusion terminal itself receives the main station standard timing signal and / or the local high-precision timing module (such as GPS / BeiDou) signal, and combines it with a high-stability crystal oscillator to establish and maintain a high-precision, high-stability local reference clock, providing a standard for subsequent timestamp recording.

[0025] S102: The intelligent fusion terminal receives multi-source heterogeneous data from distributed nodes and obtains the timestamp events of each multi-source heterogeneous data; the timestamp events include the local timestamp of the data generation and the arrival timestamp of the data received by the intelligent fusion terminal.

[0026] In this embodiment, the smart fusion terminal for the distribution area accesses three types of data sources through its heterogeneous communication module: distribution area outlet data collected by the terminal's own metering module, user meter frozen data uploaded by the HPLC concentrator, and branch monitoring data uploaded by the branch monitoring unit. Specifically, user meters periodically upload hourly frozen data, and the branch monitoring unit periodically reports branch monitoring data.

[0027] Specifically, for distributed nodes such as user meters and branch monitoring units, the frozen data they periodically report is used as timestamp events; the local timestamp generated on the data source side for each frozen data (such as the meter's hourly freezing time T_local) and the arrival timestamp (T_arrival) recorded according to the reference clock when the data packet arrives at the smart fusion terminal are recorded; for the same data source, multiple (T_local, T_arrival) data pairs are continuously collected for multiple periods (such as 24 hours).

[0028] S103: Based on the difference between the local timestamp and the arrival timestamp in multiple timestamp events in each distributed node, obtain the transmission delay model between each distributed node and the intelligent fusion terminal, including: S103-1: For each distributed node, construct a difference sequence based on the difference between the local timestamp and the arrival timestamp in its multiple timestamp events; S103-2: Obtain the minimum value in the difference sequence as the minimum delay; S103-3: Calculate the standard deviation of all differences in the difference sequence as the jitter delay; S103-4: Based on the principle that the sum of minimum delay and jitter delay equals transmission delay, a transmission delay model for distributed nodes is constructed, expressed as: ; in, Indicates the amount of transmission delay compensation. Indicates minimum delay. This indicates the preset balance coefficient. This indicates jitter delay.

[0029] Specifically, in this embodiment, the difference between the arrival timestamp T_arrival and the local timestamp T_local of the data source can be decomposed into a fixed delay component, a variable delay component, and a data source clock drift component. Based on the fixed delay component and the variable delay component, time series analysis and regression methods are used to dynamically estimate and update the communication transmission delay characteristic parameters from each data source to the fusion terminal, and obtain the minimum delay and delay jitter range.

[0030] S104: Perform linear regression on the local timestamp and arrival timestamp of multiple timestamp events in each distributed node to obtain the clock drift model of each distributed node, including: S104-1: For each distributed node, the difference between the arrival timestamp and the local timestamp in its corresponding timestamp event is decomposed into a fixed delay component, a variable delay component, and a data source clock drift component. S104-2: Using time series analysis and linear regression, with the clock drift component of the data source as the independent variable and the clock skew compensation amount as the dependent variable, the clock drift model of each distributed node is fitted and obtained, expressed as: ; in, This indicates the amount of clock skew compensation. Indicates clock drift rate, Indicates the initial frequency deviation. Represents the local timestamp. Indicates the reference start time.

[0031] In this embodiment, based on multiple periods of timestamp data pairs, the drift characteristics of the local clock of each distributed data source relative to the reference clock of the fusion terminal are analyzed, and a clock drift model of each data source is established. This model can at least reflect the trend of its clock deviation accumulating linearly over time. Using this model, the cumulative deviation of the data arriving later relative to the reference clock at the current moment is predicted based on the local timestamp of the source end it carries, and pre-compensation is performed.

[0032] S105: Input the local timestamp of the data to be processed into the transmission delay model between its corresponding distributed node and the intelligent fusion terminal, as well as the clock drift model of its corresponding distributed node, to obtain the corresponding transmission delay compensation amount and clock deviation compensation amount.

[0033] S106: Add the local timestamp, transmission delay compensation, and clock offset compensation of the data to be processed to obtain the standard timestamp aligned with the local reference clock, thus completing the timing alignment of the data to be processed.

[0034] Specifically, for the data to be processed arriving in real time at any distributed node, the following steps are performed: read the local timestamp of the source node carried in the data packet. Based on the clock drift model of this data source, calculate the clock skew compensation amount corresponding to time T_src. Based on the transmission delay model of the data source, estimate the transmission delay of the data packet from the source to the terminal. ; Calculate the aligned standard timestamp, represented as: The data, after being timestamped and aligned, is stored together with the data collected by the fusion terminal itself, whose timestamps are already used as the reference clock, in the time series database to form a time-consistent multi-source dataset.

[0035] After completing the time-series alignment of the data to be processed, this embodiment further includes: data quality cleaning of the dataset containing the data to be processed; the data quality cleaning includes missing value completion, outlier correction, and deduplication. That is, based on the high-quality data after time-series alignment, this embodiment can further perform preprocessing steps such as data missing value repair, outlier cleaning, and feature calculation, so as to be directly used for line loss analysis.

[0036] Reference Figure 2 The diagram shows the principle of event-driven transmission delay and clock drift estimation. This invention performs high-precision timing alignment of multi-source data in a distribution area based on event-driven and dynamic delay estimation. It aims to achieve dynamic estimation and compensation of clock drift and communication delay of distributed nodes through software algorithms without changing the existing terminal and meter hardware. This ensures that all data is aligned to a unified and accurate time reference on the fusion terminal side, providing a high-quality data foundation for subsequent accurate line loss calculation and analysis.

[0037] Based on the above embodiments, this embodiment takes a typical transformer area scenario as an example to illustrate the specific implementation of the present invention, including S201 to S205.

[0038] S201: The hardware environment in this embodiment adopts a standard-compliant smart fusion terminal for distribution areas, equipped with a quad-core ARM processor, 1GB of memory, and integrated HPLC communication module and wireless communication module. The user meters are smart meters supporting the DL / T 645-2007 protocol and featuring hourly freeze functionality. The branch monitoring unit is a smart circuit breaker with data acquisition and uploading capabilities.

[0039] S202: The intelligent converged terminal of the distribution area accesses the GPS module to obtain high-precision pulse of second (PPS) and standard time information, calibrates its local high-stability crystal oscillator, establishes a local reference clock, and directly uses this reference clock for the timestamp of its own collected data.

[0040] S203: For three consecutive days, the smart converged terminal in the distribution area records the reporting events of data from each user's electricity meter at four fixed freezing points: 0:00, 6:00, 12:00, and 18:00 each day. Each event records the freezing time carried in the meter data ( ) and the time when the terminal receives the data ( For a given user's electricity meter, 12 pairs were obtained. Data. Linear regression analysis revealed that... and There exists a stable linear relationship: The slope of 1.000003 reflects the small daily clock drift (leading) of approximately 0.26 seconds, and the intercept of 2.5 seconds reflects the average transmission delay. Based on this, an initial clock drift model (0.26 seconds fast per day) and a baseline transmission delay value (2.5 seconds) for the meter are established.

[0041] S204: On day 4 at 10:00:05 (terminal reference time), a real-time current data entry is received from the meter, carrying a local timestamp of 10:00:00; real-time data timestamp alignment is performed on it: the process includes: S204-1: Calculate clock drift compensation: Approximately 3.5 days have passed since initialization was completed to the current time, and the estimated cumulative clock drift is... The source reference time after compensation is: 10:00:00 + 0.91s = 10:00:00.91. S204-2: Superimposed transmission delay: using the initial average delay ; S204-3: Calculate the standard timestamp: (Terminal reference time); S204-4: Mark the effective time of this current data as 10:00:03.41, which corresponds precisely to the total power data collected by the terminal itself at 10:00:03.41.

[0042] S205: Model Update: The delay and drift model parameters for each meter are updated weekly using new frozen data events to accommodate possible changes.

[0043] This embodiment also provides a computer-readable storage medium storing a computer program thereon, which, when executed, implements the steps of the multi-source data preprocessing method for transformer substation line loss analysis as described above.

[0044] The multi-source data preprocessing method for transformer substation line loss analysis described in this invention analyzes the transmission delay between each distributed node and the intelligent fusion terminal, as well as the clock drift within the intelligent fusion terminal. It models and compensates for both communication transmission delay and clock drift of the distributed nodes, improving the accuracy of aligning multi-source data to a unified time base to within 100 milliseconds. This is far superior to the error levels of traditional methods, which range from several seconds to several minutes. It fundamentally eliminates line loss calculation errors caused by timing misalignment, achieving sub-second high-precision timing alignment. Furthermore, the transmission delay model and clock drift model can be updated periodically following data transmission, adapting to changes in network load and clock characteristics caused by equipment aging, maintaining high synchronization accuracy over the long term. By providing highly consistent multi-source data, this invention fundamentally guarantees the accuracy and reliability of advanced applications such as real-time power balance-based line loss calculation, transient event analysis, and phase and topology identification. Simultaneously, this invention is entirely based on existing data acquisition and software algorithms, requiring no hardware upgrades or time synchronization module additions to the massive deployment of user meters and branch monitoring units. It requires no hardware modification, exhibits strong adaptability, low implementation cost, and is easy to scale.

[0045] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0046] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0047] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0048] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0049] Obviously, the above embodiments are merely illustrative examples for clear explanation and are not intended to limit the implementation. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is neither necessary nor possible to exhaustively list all possible implementations here. However, obvious variations or modifications derived therefrom are still within the scope of protection of this invention.

Claims

1. A multi-source data preprocessing method for transformer substation line loss analysis, characterized in that, include: A local reference clock for the intelligent fusion terminal is constructed based on a preset timing signal and a high-stability crystal oscillator. The intelligent fusion terminal receives multi-source heterogeneous data from distributed nodes and obtains the timestamp events of each multi-source heterogeneous data; the timestamp events include the local timestamp of data generation and the arrival timestamp of data received by the intelligent fusion terminal; Based on the difference between the local timestamp and the arrival timestamp in multiple timestamp events in each distributed node, the transmission delay model between each distributed node and the intelligent fusion terminal is obtained. Linear regression is performed on the local timestamp and arrival timestamp of multiple timestamp events in each distributed node to obtain the clock drift model of each distributed node. Input the local timestamp of the data to be processed into the transmission delay model between the corresponding distributed node and the intelligent fusion terminal, as well as the clock drift model of the corresponding distributed node, to obtain the corresponding transmission delay compensation amount and clock deviation compensation amount. The local timestamp, transmission delay compensation, and clock skew compensation of the data to be processed are added together to obtain a standard timestamp aligned with the local reference clock, thus completing the timing alignment of the data to be processed.

2. The multi-source data preprocessing method for transformer substation line loss analysis according to claim 1, characterized in that, Based on a preset timing signal and a high-stability crystal oscillator, a local reference clock is constructed for the intelligent converged terminal in the distribution area, including: Obtain the standard timing signal from the main station or the signal from the local timing module as the preset timing signal; Based on the preset timing signal, the output frequency of the high-stability crystal oscillator is synchronized with the preset timing signal to form a local reference clock. The local time synchronization module signals include GPS signals and BeiDou signals.

3. The multi-source data preprocessing method for transformer substation line loss analysis according to claim 1, characterized in that, The multi-source heterogeneous data includes: the substation outlet data collected by the metering module of the substation intelligent fusion terminal itself, the user meter freezing data uploaded by the HPLC concentrator, and the branch monitoring data uploaded by the branch monitoring unit.

4. The multi-source data preprocessing method for transformer substation line loss analysis according to claim 3, characterized in that, User electricity meters periodically upload hourly frozen data, and branch monitoring units periodically report branch monitoring data.

5. The multi-source data preprocessing method for transformer substation line loss analysis according to claim 1, characterized in that, Based on the difference between the local timestamp and the arrival timestamp in multiple timestamp events across various distributed nodes, a transmission delay model between each distributed node and the intelligent fusion terminal is obtained, including: For each distributed node, a difference sequence is constructed based on the difference between the local timestamp and the arrival timestamp in its multiple timestamp events; Find the minimum value in the difference sequence, and use it as the minimum delay; Calculate the standard deviation of all differences in the difference sequence as the jitter delay; Based on the principle that the sum of minimum delay and jitter delay equals transmission delay, a transmission delay model for distributed nodes is constructed.

6. The multi-source data preprocessing method for transformer substation line loss analysis according to claim 5, characterized in that, The transmission delay model is expressed as: ; in, Indicates the amount of transmission delay compensation. Indicates minimum delay. This indicates the preset balance coefficient. This indicates jitter delay.

7. The multi-source data preprocessing method for transformer substation line loss analysis according to claim 1, characterized in that, A linear regression is performed on the local timestamps and arrival timestamps of all data reporting events from each distributed node to obtain the clock drift model for each distributed node, including: For each distributed node, the difference between the arrival timestamp and the local timestamp in its corresponding timestamp event is decomposed into a fixed delay component, a variable delay component, and a data source clock drift component. Using time series analysis and linear regression, with the clock drift component of the data source as the independent variable and the clock skew compensation amount as the dependent variable, a clock drift model for each distributed node is fitted and obtained.

8. The multi-source data preprocessing method for transformer substation line loss analysis according to claim 7, characterized in that, The clock drift model is expressed as follows: ; in, This indicates the amount of clock skew compensation. Indicates clock drift rate, Indicates the initial frequency deviation. Represents the local timestamp. Indicates the reference start time.

9. The multi-source data preprocessing method for transformer substation line loss analysis according to claim 1, characterized in that, After completing the time sequence alignment of the data to be processed, the process also includes: data quality cleaning of the dataset containing the data to be processed; the data quality cleaning includes missing value completion, outlier correction and deduplication.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed, it implements the steps of the multi-source data preprocessing method for transformer substation line loss analysis as described in any one of claims 1 to 9.