Data preprocessing method for multi-mode dynamic coupling and topology compensation of power distribution network

Through improved data preprocessing methods, including Hampel filter and wavelet packet decomposition, the joint modeling problem of multi-source heterogeneous data is solved, and the fault diagnosis accuracy and network adaptability of the three-phase power grid are improved.

CN120497878APending Publication Date: 2025-08-15ELECTRIC POWER RES INST OF GUANGXI POWER GRID CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510497749.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The prior art lacks the ability to jointly model multi-source heterogeneous data, resulting in limited diagnostic accuracy of three-phase power grid data and inability to adapt to dynamic networks.

Method used

The improved Hampel filter is used to correct outliers, and the low sampling rate data is resampled to high sampling rate through cubic spline interpolation, and the energy proportion of each band is extracted through wavelet packet decomposition, a weighted adjacency matrix is constructed for topological compensation, and the synchronous rotation components of voltage/current and wavelet energy are fused to generate TCN input characteristics.

Benefits of technology

Improve data quality, enhance feature extraction capabilities, improve fault diagnosis accuracy and network adaptability, improve fault classification F1-score by 11.4%, reduce fault location error by 40%, and increase delay by 6.3ms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120497878A_ABST
    Figure CN120497878A_ABST
Patent Text Reader

Abstract

The invention discloses a data preprocessing method for multi-modal dynamic coupling and topology compensation of a power distribution network, relates to the technical field of data processing, and solves the problems that the diagnosis precision is limited and a dynamic network cannot be adapted due to the lack of joint modeling capability for multi-source heterogeneous data in the prior art. According to the method, abnormal values are dynamically detected and corrected by adopting a filter in combination with median absolute deviation, and low-sampling-rate data are resampled to high-sampling-rate data through cubic spline interpolation. And then, multi-scale decomposition is carried out on the voltage / current signals through wavelet packet decomposition, the energy ratio of each frequency band is extracted, and dynamic time warping is carried out to align the change trend of the node temperature and the environment temperature, and the hysteresis effect is compensated. And calculating an electrical distance between nodes based on the admittance matrix and the power sensitivity, constructing a weighted adjacent matrix, and carrying out nonlinear temperature compensation. And finally, the synchronous rotation component, the change rate, the wavelet energy and the like of the voltage / current are fused to generate TCN input characteristics.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a data preprocessing method for multi-modal dynamic coupling and topology compensation of a distribution network. Background Art

[0002] A three-phase power grid consists of three AC power sources with the same frequency, equal amplitude, and a phase shift of 120 degrees. This system is widely used in industrial and residential power supply due to its efficient power transmission and reliability. In a three-phase power grid, voltage and current changes are dynamic and are affected by a variety of factors, such as load fluctuations, ambient temperature, and line impedance. Therefore, real-time monitoring and analysis of these changes are crucial to ensuring stable grid operation. Sensor technology enables real-time acquisition of key parameters in the three-phase power grid, such as voltage amplitude, current amplitude, power factor, and frequency. This data provides an important basis for grid operational status monitoring, fault diagnosis, and optimized control. With the widespread adoption of smart grids, the type and scale of data collected by grid sensors are growing exponentially.

[0003] Typical three-phase power grid sensor data includes time series data, environmental data, and topological data. Among them, time series data includes voltage and current waveforms (sampling frequency 1-10kHz), node temperature (1Hz sampling), etc., environmental data includes ambient temperature and humidity (1Hz sampling), and topological data includes grid node connection relationships (static or dynamic topology). The effective use of these data is highly dependent on efficient data preprocessing methods. Data preprocessing is the first step in data analysis, and its quality directly affects the accuracy and effectiveness of subsequent analysis. Existing preprocessing methods do not adequately model time correlation. Traditional sliding window root mean square (RMS) calculations ignore high-frequency transient characteristics, which easily leads to the loss of TCN input information. Traditional data preprocessing methods lack the ability to jointly model multi-source heterogeneous data (electrical + temperature + topology), resulting in limited diagnostic accuracy and inability to adapt to dynamic networks.

[0004] In view of this, a data preprocessing method for multi-modal dynamic coupling and topology compensation of distribution networks is needed. Summary of the Invention

[0005] To address the problem that existing technologies lack the ability to jointly model multi-source heterogeneous data (electrical + temperature + topology), resulting in limited diagnostic accuracy and inability to adapt to dynamic networks, this paper provides a data preprocessing method for multimodal dynamic coupling and topology compensation in distribution networks, which can improve data quality and enhance feature extraction capabilities. The specific technical solution is as follows:

[0006] A data preprocessing method for multi-modal dynamic coupling and topology compensation of a distribution network comprises the following steps:

[0007] Acquire multimodal data from each sensor node, including temperature data and electrical data, and map low-sampling data to high-sampling time series data through cubic spline interpolation;

[0008] Perform wavelet packet decomposition on time series data to extract the energy proportion of each frequency band, thereby realizing joint feature extraction in time and frequency domains;

[0009] Construct a weighted adjacency matrix of sensor nodes and establish a nonlinear compensation model for node temperature;

[0010] The synchronous rotation component, rate of change, and wavelet energy of voltage / current are fused to generate TCN input features.

[0011] Preferably, the multimodal data includes the three-phase voltage V collected by each sensor node abc (t), three-phase current I abc (t), node temperature T n (t), ambient temperature T env (t), topology connection table, three-phase voltage V abc (t) and three-phase current I abc (t) have the same sampling frequency.

[0012] Preferably, the three-phase voltage V abc (t), three-phase current I abc (t) An improved Hampel filter is used to correct outliers. The improved Hampel filter is expressed as:

[0013]

[0014] Where x(t) represents the sampling value of any phase current or voltage at the current time (t), W t is a symmetric window centered on t, median() represents the median, σ() is the median absolute deviation (MAD) within the window, Indicates the corrected sampling value.

[0015] Preferably, the low-sampling data is mapped to the high-sampling time series data by cubic spline interpolation as follows:

[0016] Use cubic spline interpolation to calculate the nodal temperature T n (t) and ambient temperature T env (t) Resample to the three-phase voltage V abc (t) or three-phase current I abc (t) With the same frequency, the resampled node temperature is expressed as:

[0017]

[0018] Among them, ti Indicates the target interpolation time point, T n '(t i ) represents the target interpolation time point t i Interpolated temperature at B j (t) is the jth B-spline basis function, c j B j (t) is the weight coefficient, m represents the number of basis functions;

[0019] Determine the target interpolation time point t according to the voltage / current sampling time series i , t i Align temperature data to high-frequency timing;

[0020] Determine the number of basis functions m according to the number of sampling time points per second of the original data;

[0021] For all original data points, the weight coefficient c j The constraint function is as follows:

[0022]

[0023] Among them, B j (t k ) represents the time point t k The jth B-spline basis function, T n (t k ) Time point t k The node temperature below.

[0024] The natural cubic spline requires the second-order derivative of the endpoints to be zero, that is:

[0025]

[0026] The continuity condition is: at the internal nodes, the interpolation function must satisfy the continuity of the first-order and second-order derivatives;

[0027] The same interpolation process is used for the ambient temperature.

[0028] Preferably, the process of joint feature extraction in time-frequency domain is as follows:

[0029] The db4 wavelet basis is used to perform 5-layer wavelet packet decomposition on the voltage and current signals, and the energy proportion of each frequency band is extracted as the TCN input feature;

[0030] Define the node temperature T n (t) and ambient temperature T env The DTW distance matrix D∈R of (t) N×M ;

[0031] Determine the optimal path P based on the DTW distance *The corresponding time shift coefficient Δt n , to compensate for the ambient temperature hysteresis effect.

[0032] Preferably, the process of constructing the weighted adjacency matrix of the sensor nodes is as follows:

[0033] The element A in the i-th row and j-th column of the weighted adjacency matrix A ij Defined as:

[0034]

[0035] Among them, γ = 0.1 is the attenuation coefficient, d ij represents the electrical distance between nodes i and j and is calculated as follows:

[0036]

[0037] Among them, Y ij is the admittance matrix element, is the power angle sensitivity.

[0038] Preferably, the nonlinear compensation model of the node temperature is established as follows:

[0039]

[0040] Among them, α n represents the ambient temperature impact intensity coefficient of node n, β represents the ambient temperature nonlinear scaling factor, Δt n Indicates the lag time of the influence of ambient temperature on node temperature, sigmoid() represents the sigmoid nonlinear activation function, and the adjustment factor α n , β are fitted by the Levenberg-Marquardt algorithm.

[0041] Preferably, the TCN input features are constructed as follows:

[0042] Generate a multidimensional time series feature matrix X TCN ∈R L×D (L is the time step, feature dimension D = 13):

[0043]

[0044] Among them, V dq , I dq are the voltage component and current component in the synchronous rotating coordinate system, They represent the rate of change of voltage and current, E kj is the wavelet packet decomposition band energy, THD V and THD I Represent the voltage and current harmonic distortion rates respectively.

[0045] A computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute the data preprocessing method for three-phase power grid sensor nodes as described above.

[0046] A processor is provided, and is used for running a program, wherein the program, when running, executes the data preprocessing method for three-phase power grid sensor nodes as described above.

[0047] Compared with the prior art, the present invention has the following beneficial effects:

[0048] The present invention adopts an improved Hampel filter, combined with the median absolute deviation (MAD) to dynamically detect and correct outliers, and resamples low sampling rate data (such as temperature) to a high sampling rate (such as voltage / current) through cubic spline interpolation. Subsequently, the voltage / current signal is multi-scale decomposed by wavelet packet decomposition (WPD), the energy proportion of each frequency band is extracted, and dynamic time warping (DTW) is used to align the changing trend of node temperature and ambient temperature to compensate for the lag effect. The electrical distance between nodes is then calculated based on the admittance matrix and power sensitivity, and a weighted adjacency matrix is constructed. Nonlinear temperature compensation: The nonlinear effect of ambient temperature on node temperature is eliminated through sigmoid function and lag time correction. Finally, the synchronous rotation component, rate of change, wavelet energy, etc. of voltage / current are fused to generate TCN input features. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly describes the drawings required for the specific embodiments or the description of the prior art. Similar elements or parts are generally identified by similar reference numerals throughout the drawings. Elements or parts in the drawings are not necessarily drawn to scale.

[0050] Figure 1 It is a flow chart of a data preprocessing method for multi-modal dynamic coupling and topology compensation of a distribution network;

[0051] Figure 2 This is a detailed flowchart of step 1;

[0052] Figure 3 This is a detailed flowchart of step 2;

[0053] Figure 4 This is a detailed flowchart of step 3;

[0054] Figure 5 This is a detailed flowchart of step 4. DETAILED DESCRIPTION

[0055] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0056] It will be understood that when used in this specification and the appended claims, the terms “comprises” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.

[0057] It should also be understood that the terms used in the present specification are only for the purpose of describing particular embodiments and are not intended to limit the present invention. As used in the present specification and the appended claims, the singular forms "a", "an", and "the" are intended to include the plural forms unless the context clearly indicates otherwise.

[0058] It should be further understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0059] like Figure 1 As shown in the flowchart, the present invention designs a data preprocessing method for multi-modal dynamic coupling and topology compensation of a distribution network, which generally includes:

[0060] Step 1: Multimodal data cleaning and alignment;

[0061] Step 2: Joint feature extraction in time-frequency domain;

[0062] Step 3: Topology-environment coupling modeling;

[0063] Step 4: Feature space mapping and normalization.

[0064] 1) Step 1: Multimodal data cleaning and alignment

[0065] like Figure 2 As shown in the flowchart, this step specifically includes:

[0066] ① Obtain the three-phase voltage V collected by each sensor node abc (t), three-phase current I abc (t), node temperature T n (t), ambient temperature T env (t), topology connection table, three-phase voltage V abc (t) and three-phase current Iabc (t) have the same sampling frequency;

[0067] ② For three-phase voltage V abc (t), three-phase current I abc (t) An improved Hampel filter is used to correct outliers. The improved Hampel filter is expressed as:

[0068]

[0069] Where x(t) represents the sampling value of any phase current or voltage at the current time (t), W t is a symmetric window centered on t (length 2k+1, k=5 is recommended), median() represents the median, σ() is the median absolute deviation (MAD) within the window, Indicates the corrected sampling value.

[0070] ③Use cubic spline interpolation to calculate the node temperature T n (t) and ambient temperature T env (t) Resample to the three-phase voltage V abc (t) or three-phase current I abc (t) With the same frequency, the resampled node temperature is expressed as:

[0071]

[0072] Among them, t i Indicates the target interpolation time point, T n '(t i ) represents the target interpolation time point t i Interpolated temperature at B j (t) is the jth B-spline basis function, c j B j The weight coefficient of (t) is solved by least square method or constrained optimization. m represents the number of basis functions. j , so that the interpolation curve passes through the original data points exactly while maintaining global smoothness.

[0073] Target interpolation time point t i Determined according to the sampling time series of voltage / current, such as t corresponding to 10kHz sampling i =0,0.0001,0.0002,...,T max sec. t i Align temperature data to high-frequency timing to ensure synchronous analysis with voltage / current data.

[0074] The number of basis functions m is determined by the number of sampling time points per second of the original data. For example, if the original data has K time points per second, when using natural cubic splines, m = K + 2 basis functions are usually required.

[0075] B j The weight coefficient c of (t) j The calculation methods include:

[0076] Suggesting the weight coefficient c based on the original data points j Constraint function:

[0077] For all original data points

[0078] The natural cubic spline requires the second-order derivative of the endpoints to be zero, that is:

[0079]

[0080] The continuity condition is: at the internal nodes, the interpolation function must satisfy the continuity of the first-order and second-order derivatives.

[0081] The same interpolation process is used for the ambient temperature.

[0082] Below is an example using cubic spline interpolation.

[0083] Assume that the original sampling rate of a node temperature is 1Hz (1 point per second) and the voltage / current sampling rate is 10kHz (10,000 points per second). Our goal is to interpolate the temperature data to a 10kHz time series.

[0084] Input data:

[0085] The original temperature series is expressed as: T n =[T(0),T(1),T(2),...,T(59)](60 seconds of data)

[0086] The target time point is expressed as: t i =0,0.0001,0.0002,...,59.9999 (a total of 600,000 points).

[0087] Interpolation process:

[0088] Construct natural cubic spline with knot τ j Set to 0,1,2,...,59.

[0089] Calculate basis function B j (t i ) and coefficient c j , and get the interpolation result T n ′(t i ).

[0090] Verification results:

[0091] At the original time point t=0,1,...,59, the interpolation result is strictly equal to the original value T n (t).

[0092] At the intermediate point t=0.5, a smooth transition is achieved through the combination of basis functions.

[0093] Through the cubic spline interpolation formula, various parameters work together to accurately and smoothly map low-sampled temperature data to a high-sampled time series, providing aligned multimodal input for the subsequent TCN and GAT networks. This method effectively suppresses noise while ensuring data fidelity, making it a key preprocessing step for multi-source data fusion in smart grids.

[0094] 2) Step 2: Joint feature extraction in time-frequency domain

[0095] like Figure 3 As shown in the flowchart, step 2 specifically includes the following steps:

[0096] ①The db4 wavelet basis is used to perform 5-layer wavelet packet decomposition on the voltage and current signals, and the energy proportion of each frequency band is extracted as the TCN input feature.

[0097] ②Define the node temperature T n (t) and ambient temperature T env The DTW (Dynamic Time Warping) distance matrix D∈R of (t) N×M (N, M represent the node temperature T n (t) and ambient temperature T env (t) the number of sampling points).

[0098] The element D(i,j) in the i-th row and j-th column of D represents the temperature T of the i-th node n (i) and the jth ambient temperature T env The DTW distance between (j) is calculated as follows:

[0099]

[0100] min means taking the minimum value.

[0101] The DTW distance matrix is used to quantify the distance between two time series T n (t) and T env (t) to find the optimal alignment path P that minimizes the cumulative distance * .

[0102] ③Determine the optimal path P based on DTW distance * The corresponding time shift coefficient Δt n Used to compensate for the ambient temperature hysteresis effect.

[0103] 3) Step 3: Topology-environment coupling modeling

[0104] like Figure 4 As shown in the flowchart, step 3 specifically includes the following steps:

[0105] ① Construct the weighted adjacency matrix A of the sensor node, where the element A in the i-th row and j-th column of the weighted adjacency matrix A is ij Defined as:

[0106]

[0107] Among them, γ = 0.1 is the attenuation coefficient, d ij represents the electrical distance between nodes i and j and is calculated as follows:

[0108]

[0109] Among them, Y ij is the admittance matrix element, is the power angle sensitivity.

[0110] The admittance matrix is the core matrix in power system analysis, which describes the electrical connection characteristics between nodes in the power grid. ij Y represents the mutual admittance between node i and node j. ij Calculated by the following formula:

[0111] Y ij =G ij +jB ij

[0112] Among them, G ij Represents conductance, reflecting the resistance characteristics of the line between nodes i and j (unit: Siemens, S). ij Susceptance, which reflects the reactance of the line between nodes i and j (unit: Siemens, S). The greater the admittance (i.e., the smaller the impedance), the stronger the electrical connection of the line, and the easier it is for current to flow.

[0113] The modulus of the admittance between nodes is inversely proportional to the line impedance (|Z ij |=1 / |Y ij |).

[0114] For example, if there is an impedance Z between nodes i and j ij =0.1+j0.5Ω line connection, then:

[0115] Y ij =1 / Zij ≈0.3846-j1.9231(S)

[0116]

[0117] represents the active power P of node i i The voltage phase angle θ at node j j The partial derivative of reflects the degree of influence of the phase angle change of node j on Pi. Based on the DC power flow approximation, Approximately equal to -B ij .

[0118] The electrical distance formula accurately quantifies the electrical coupling strength between nodes by combining the admittance matrix and power sensitivity. Its parameter design takes into account the physical connection characteristics and dynamic response capabilities of the power grid, providing subsequent GAT networks with characteristic inputs that are more consistent with the actual operating laws of the power system, significantly improving the accuracy of fault location and status prediction.

[0119] ②Establish node temperature T n Nonlinear compensation model:

[0120]

[0121] Among them, α n represents the ambient temperature impact intensity coefficient of node n, β represents the ambient temperature nonlinear scaling factor, Δt n Indicates the lag time of the influence of ambient temperature on node temperature, sigmoid() represents the sigmoid nonlinear activation function. Adjustment factor α n , β are fitted by the Levenberg-Marquardt algorithm.

[0122] α n Used to quantify the contribution of ambient temperature to the temperature of node n. The larger the value, the more significant the influence of ambient temperature on the node temperature. n ≥0, usually obtained by fitting historical data. The fitting steps are:

[0123] Data collection: Under no-load or constant-load conditions, record the node temperature T n (t) and ambient temperature T env (t) time series data.

[0124] Loss function: Minimize the mean square error between the corrected temperature and the theoretical steady-state temperature:

[0125]

[0126] Where T n,theory is the Joule heating temperature rise calculated based on the current (assuming that the environmental coupling is eliminated).

[0127] Optimization algorithm: Use Levenberg-Marquardt algorithm to jointly optimize α n , β, Δt n .

[0128] β is used to control the nonlinear saturation characteristics of the model due to the influence of ambient temperature. Larger β values cause the model to reach saturation quickly within a specific temperature range, simulating critical changes in heat dissipation efficiency. β∈[0.1,1.0] needs to be adjusted according to the climatic conditions of the grid deployment area. For example, when β=0.2, the sigmoid output increases from 0.73 to 0.88 when the ambient temperature rises from 20°C to 30°C, a gradual change. When β=1.0, the sigmoid output increases from 0.88 to 0.99 under the same temperature change, showing significant saturation characteristics.

[0129] Δt n The delay time required for the ambient temperature change to be transmitted to the node is determined by the physical structure of the node (such as the thickness of the heat sink and the thermal capacitance of the packaging material). n The steps to determine are:

[0130] Calculate T n (t) and T env The cross-correlation function of (t):

[0131]

[0132] Take τ that maximizes R(τ) as Δt n .

[0133] The nonlinear compensation model for node temperature significantly improves the physical integrity of temperature data through the use of sigmoid functions, lag time correction, and node-specific parameters. Its design, tightly integrated with the dynamics of heat conduction, provides high-quality input features for subsequent fault diagnosis algorithms and represents a key technological breakthrough in smart grid state perception.

[0134] 4) Step 4: Feature Space Mapping and Normalization

[0135] like Figure 5 As shown in the flowchart, step 4 specifically includes the following steps:

[0136] ①Construct TCN input features

[0137] Generate a multidimensional time series feature matrix X TCN ∈R L×D (L is the time step, feature dimension D = 13):

[0138]

[0139] Among them, Vdq , I dq They are the voltage component and current component in the synchronous rotating coordinate system respectively. abc (t) is mapped to the synchronous rotating coordinate system through Park transformation (dq transformation) and decomposed into the direct axis component V d (t) and the quadrature axis component V q (t), that is, V dq =(V d ,V q )(2D). Similarly, we get I dq =(I d ,I q )(2D).

[0140] (1 dimension), (1D) represents the rate of change of voltage and current, respectively. The central difference method can be used to calculate the instantaneous rate of change. The voltage rate of change can capture transient events such as voltage swells / sags and flicker, while the current rate of change can identify fast dynamic processes such as short circuits and inrush currents.

[0141] E kj That is, the wavelet packet decomposition frequency band energy. In step 2, the db4 wavelet basis is selected to perform 5-layer wavelet packet decomposition (WPD) on the voltage / current signal, and each layer of decomposition generates 2 k frequency bands (k represents the level, k=5 corresponds to 32 frequency bands), and the energy proportion of each frequency band signal (the ratio of the frequency band signal to the total frequency band energy of the layer) is calculated as the wavelet packet decomposition frequency band energy of the frequency band.

[0142] Low-frequency energy (such as 0-100Hz) reflects the fundamental and harmonic components, and high-frequency energy (such as 1-5kHz) captures high-frequency noise such as switching operations and arc discharge. We select the energy of the first four main frequency bands, covering the key frequency band of 0-5kHz, and the corresponding E kj It is 4-dimensional.

[0143] THD V (1D) and THD I (1 dimension) represents the voltage and current harmonic distortion rate, THD V The time domain expression is THD V (t), THD V The calculation method of (t) is:

[0144]

[0145] Where V1(t) is the effective value of the fundamental voltage, V h (t) is the hth harmonic component, and H is the highest harmonic component.

[0146] Similarly, THD can be calculatedI (t).

[0147] Multidimensional time series feature matrix X TCN Through carefully designed 13-dimensional features, a comprehensive description of the grid's operating status is achieved. Its design concept, which integrates time-frequency domain analysis with physical model correction, significantly improves the TCN network's ability to identify complex fault modes, providing a reliable data foundation for the efficient operation and maintenance of smart grids.

[0148] ②Construct GAT input features

[0149] GAT node feature matrix X GAT ∈R N×F (N represents the number of nodes, feature dimension F = 6) is:

[0150]

[0151] in, They represent the average voltage and average current of node n, respectively, and are defined as the average voltage and current of node n within the time window L. represents the sum of the electrical distance weights between node n and all its neighbor nodes j. Represents the temperature sequence of node n after correction The maximum value in . It represents the standard deviation of the corrected temperature, which characterizes the degree of fluctuation of the corrected temperature sequence of node n. It is calculated as follows:

[0152]

[0153] in, Indicates the average value of the corrected temperature.

[0154] Degree(n) represents the number of edges directly connected to node n in the power grid topology and is calculated as follows:

[0155]

[0156] where ||(A nj >0) is the indicator function, when the adjacency matrix element A nj It takes 1 when >0, otherwise it takes 0.

[0157] The GAT input feature matrix integrates electrical properties, thermodynamic state, and topological coupling strength, providing a multi-dimensional, high-information-density input for the graph attention mechanism. Its design fully exploits the physical meaning and statistical characteristics of power grid data, significantly improving the accuracy and interpretability of fault location. It represents a core innovation in the field of smart grid diagnostics.

[0158] ③ The voltage and current are normalized by the maximum and minimum values, and the temperature is normalized by the Z-score.

[0159] The parameters that are normalized by maximum and minimum values include: V d ,V q ,I d ,I q , harmonic distortion rate THD V ,THD I , wavelet packet energy ratio E k,j , instantaneous rate of change dV / dt,dI / dt.

[0160] Parameters standardized using Z-score include: Corrected node temperature:

[0161] Through differentiated normalization and standardization processing, the dimensional differences of multi-source data are eliminated, and the inherent characteristics of different physical quantities are retained, providing balanced and information-rich input features for TCN and GAT models.

[0162] (2) Key points of the present invention

[0163] ① Multimodal data cleaning and alignment

[0164] Outlier correction: An improved Hampel filter is used in combination with the median absolute deviation (MAD) to dynamically detect and correct outliers.

[0165] Sampling rate unification: resample low sampling rate data (such as temperature) to a high sampling rate (such as voltage / current) through cubic spline interpolation.

[0166] ②Joint feature extraction in time-frequency domain

[0167] Wavelet Packet Decomposition (WPD): Performs multi-scale decomposition on voltage / current signals to extract the energy proportion of each frequency band.

[0168] Dynamic Time Warping (DTW): Aligns the changing trends of node temperature with the ambient temperature to compensate for lag effects.

[0169] ③Topology-environment coupling modeling

[0170] Electrical distance weighted topology: Calculates the electrical distance between nodes based on the admittance matrix and power sensitivity, and constructs a weighted adjacency matrix.

[0171] Nonlinear temperature compensation: Eliminate the nonlinear effect of ambient temperature on node temperature through sigmoid function and lag time correction.

[0172] ④ Feature space mapping and normalization

[0173] Multi-dimensional time series feature construction: Fusion of synchronous rotation components, change rates, wavelet energy, etc. of voltage / current to generate TCN input features.

[0174] Node feature construction: Combine average voltage / current, corrected temperature, node degree, etc. to generate GAT input features.

[0175] Mixed normalization: Maximum and minimum values are used for voltage and current, and Z-score is used for temperature.

[0176] In summary, the present invention significantly improves system performance in many aspects. In terms of data quality, the accuracy of outlier detection is increased to more than 95% through the Hampel filter, and the error is controlled within 0.1% using cubic spline interpolation to ensure the time alignment accuracy of multi-source data. The feature extraction capability has also been enhanced. Wavelet packet decomposition increases the detection sensitivity of high-frequency transient features (such as arc discharge) by 30%, while dynamic time warping reduces the ambient temperature lag compensation error to within 0.5°C. The accuracy of topological modeling is further improved. Electrical distance weighting optimizes the sparsity of the weighted adjacency matrix to 0.18. Compared with the traditional 0-1 adjacency matrix, the fault location error of the GAT network is reduced by 40%. At the same time, nonlinear temperature compensation increases the correlation coefficient between the corrected temperature and the theoretical Joule heat rise from 0.65 to 0.92. The feature space has also been optimized. The dimensionality of TCN input features has been expanded to 13, covering multiple sources of information, including time, frequency, and thermodynamics. The dimensionality of GAT input features has been expanded to 6, integrating electrical properties, topology, and temperature. Hybrid normalization has reduced the normalized error of voltage and current features to less than 1%, and the normalized error of temperature features to less than 0.5σ. As a result, model performance has been significantly improved, with a weighted F1-score for fault classification reaching 93.7%, an 11.4% improvement over traditional methods (such as CNN-LSTM). The average node distance error (ANDE) for fault localization has been reduced to 0.9, an improvement of 2.6 nodes over traditional methods (such as pure GCN). The preprocessing delay only increases by 6.3ms, with a negligible impact on real-time performance.

[0177] Those skilled in the art will appreciate that the units of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition of each example has been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0178] In the embodiments provided by the present invention, it should be understood that the division of units is merely a logical function division, and there may be other division methods in actual implementation, for example, multiple units can be combined into one unit, one unit can be split into multiple units, or some features can be ignored, etc.

[0179] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0180] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-0nly Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk, etc., various media that can store program code.

[0181] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some or all of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present invention, and they should all be included in the scope of the claims and description of the present invention.

Claims

1. A data preprocessing method for multi-modal dynamic coupling and topology compensation of distribution network, characterized in that: The following steps are involved: Acquire multimodal data from each sensor node, including temperature data and electrical data, and map low-sampling data to high-sampling time series data through cubic spline interpolation; Perform wavelet packet decomposition on time series data to extract the energy proportion of each frequency band, thereby realizing joint feature extraction in time and frequency domains; Construct a weighted adjacency matrix of sensor nodes and establish a nonlinear compensation model for node temperature; The synchronous rotation component, rate of change, and wavelet energy of voltage / current are fused to generate TCN input features.

2. The data preprocessing method for multi-modal dynamic coupling and topology compensation of a distribution network according to claim 1 is characterized in that: Multimodal data includes the three-phase voltage V collected by each sensor node abc (t), three-phase current I abc (t), node temperature T n (t), ambient temperature T env (t), topology connection table, three-phase voltage V abc (t) and three-phase current I abc (t) have the same sampling frequency.

3. The data preprocessing method for multi-modal dynamic coupling and topology compensation of a distribution network according to claim 2 is characterized in that: Also includes the three-phase voltage V abc (t), three-phase current I abc (t) An improved Hampel filter is used to correct outliers. The improved Hampel filter is expressed as: Where x(t) represents the sampling value of any phase current or voltage at the current time (t), W t is a symmetric window centered on t, median() represents the median, σ() is the median absolute deviation (MAD) within the window, Indicates the corrected sampling value.

4. The data preprocessing method for multi-modal dynamic coupling and topology compensation of a distribution network according to claim 1 is characterized in that: Through cubic spline interpolation, the low-sampling data is mapped to the high-sampling time series data as follows: Use cubic spline interpolation to calculate the nodal temperature T n (t) and ambient temperature T env (t) Resample to the three-phase voltage V abc (t) or three-phase current I abc (t) With the same frequency, the resampled node temperature is expressed as: Among them, t i Indicates the target interpolation time point, T n '(t i ) represents the target interpolation time point t i Interpolated temperature at B j (t) is the jth B-spline basis function, c j B j (t) is the weight coefficient, m represents the number of basis functions; Determine the target interpolation time point t according to the voltage / current sampling time series i , t i Align temperature data to high-frequency timing; Determine the number of basis functions m according to the number of sampling time points per second of the original data; For all original data points, the weight coefficient c j The constraint function is as follows: Among them, B j (t k ) represents the time point t k The jth B-spline basis function, T n (t k ) Time point t k The node temperature below. The natural cubic spline requires the second-order derivative of the endpoints to be zero, that is: The continuity condition is: at the internal nodes, the interpolation function must satisfy the continuity of the first-order and second-order derivatives; The same interpolation process is used for the ambient temperature.

5. The data preprocessing method for multi-modal dynamic coupling and topology compensation of a distribution network according to claim 1 is characterized in that: The process of joint feature extraction in time-frequency domain is as follows: The db4 wavelet basis is used to perform 5-layer wavelet packet decomposition on the voltage and current signals, and the energy proportion of each frequency band is extracted as the TCN input feature; Define the node temperature T n (t) and ambient temperature T env The DTW distance matrix D∈R of (t) N×M ; Determine the optimal path P based on the DTW distance * The corresponding time shift coefficient Δt n , to compensate for the ambient temperature hysteresis effect.

6. The data preprocessing method for multi-modal dynamic coupling and topology compensation of a distribution network according to claim 1, characterized in that: The process of constructing the weighted adjacency matrix of sensor nodes is as follows: The element A in the i-th row and j-th column of the weighted adjacency matrix A ij Defined as: Among them, γ = 0.1 is the attenuation coefficient, d ij represents the electrical distance between nodes i and j and is calculated as follows: Among them, Y ij is the admittance matrix element, is the power angle sensitivity.

7. The data preprocessing method for multi-modal dynamic coupling and topology compensation of a distribution network according to claim 1 is characterized in that: The nonlinear compensation model of node temperature is established as follows: Among them, α n represents the ambient temperature impact intensity coefficient of node n, β represents the ambient temperature nonlinear scaling factor, Δt n Indicates the lag time of the influence of ambient temperature on node temperature, sigmoid() represents the sigmoid nonlinear activation function, and the adjustment factor α n , β are fitted by the Levenberg-Marquardt algorithm.

8. The data preprocessing method for multi-modal dynamic coupling and topology compensation of a distribution network according to claim 7, characterized in that: The construction of TCN input features is as follows: Generate a multidimensional time series feature matrix X TCN ∈R L×D (L is the time step, feature dimension D = 13): Among them, V dq , I dq are the voltage component and current component in the synchronous rotating coordinate system, They represent the rate of change of voltage and current, E kj is the wavelet packet decomposition band energy, THD V and THD I Represent the voltage and current harmonic distortion rates respectively.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute the data preprocessing method for a three-phase power grid sensor node according to any one of claims 1 to 8.

10. A processor, characterized in that: The processor is configured to run a program, wherein the program, when running, executes the data preprocessing method for a three-phase power grid sensor node according to any one of claims 1 to 8.