Power distribution network disturbance source identification method based on multi-modal feature fusion

CN121388732BActive Publication Date: 2026-08-07KUNMING UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
KUNMING UNIV OF SCI & TECH
Filing Date
2025-09-17
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0006]本发明的目的在于提供一种基于多模态特征融合的配电网扰动源辨识方法,旨在解决现有技术存在特征表征单一、物理可解释性弱及难以形成稳定的电力指纹的技术问题

Benefits of technology

[0049]本发明的有益效果是:本发明提出了融合时频图像特征、参数化物理特征与模量能量特征的多模态协同分析框架;通过设计以物理特征为查询、图像特征为键值的多头注意力融合机制,实现了物理信息引导下的动态特征交互与增强;结合跨模态对比学习损失约束,有效提升了模型在小样本条件下的泛化能力与辨识精度。可为电力系统行波扰动辨识、智能录波分析、保护动作评价等应用提供理论基础。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121388732B_ABST
    Figure CN121388732B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of based on multimodal feature fusion distribution network disturbance source identification method, belong to the fine disturbance monitoring and diagnosis technical field of smart grid.The method includes: the disturbance class current traveling wave data extracted is carried out time-frequency conversion, generates traveling wave panorama waveform chart and two-dimensional time-frequency chart and is fused into three-channel time-frequency image, then image feature vector is extracted;The disturbance class current traveling wave data is carried out Prony modal parameter fitting, extracts modal parameter feature vector and is projected to high-dimensional feature space;The disturbance class current traveling wave data is carried out Clark transformation, calculates zero mode energy proportion feature and is projected to high-dimensional feature space;Image feature vector, projected physical feature vector and projected modulus energy feature vector are fused;Fusion feature vector is input into classifier, and the class identification result of distribution network disturbance source is output.The present application aims to solve the technical problems that the prior art has single feature representation, weak physical interpretability and difficulty in forming stable power fingerprint.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method for identifying disturbance sources in distribution networks based on multimodal feature fusion, belonging to the field of refined disturbance monitoring and diagnosis technology for smart grids. Background Technology

[0002] During operation, power systems often generate transient current traveling waves due to events such as switching operations, load switching, lightning strikes, or intermittent electric arcs. These transient traveling wave signals contain rich information about the event type, and accurate identification of their categories is fundamental to achieving safe dispatching, rapid operation and maintenance, suppression of erroneous trips, and hazard location.

[0003] Existing methods based on temporal statistical features and threshold determination perform simple classification using indicators such as amplitude abrupt changes, slope, envelope, and kurtosis. These methods struggle to effectively distinguish perturbation types with different physical mechanisms but similar temporal morphologies; they are prone to misjudgment and attribution ambiguity in concurrent or multi-event scenarios; and overall, they are only suitable for coarse-grained tasks such as perturbation detection and origin localization, lacking interpretable physical parameters to support fine-grained discrimination.

[0004] Existing time-frequency analysis-based methods first obtain intrinsic mode functions (EMFs) through empirical mode decomposition (EMD), then calculate instantaneous frequency and energy to obtain the time-frequency spectrum. Subsequently, the energy is integrated within the target frequency band, and features such as spectral centroid, bandwidth, spectral entropy, and gray-level co-occurrence matrix texture are extracted. However, this approach is prone to energy leakage and time-frequency ridge aliasing under noise and concurrent events, making it difficult to stably separate similar mechanisms. The Hilbert-Huang transform suffers from endpoint effects and mode aliasing. Short windows reduce frequency resolution, while long windows weaken time-domain localization, making it difficult to simultaneously meet the localization and resolution requirements of short-duration damped ringing after disturbances. Spectral features are sensitive to sampling rate, anti-aliasing filtering, and range saturation, requiring a large amount of labeled data for threshold calibration or model training.

[0005] Existing deep learning-based end-to-end classification methods, while avoiding the construction of manual features, are highly dependent on large-scale, high-quality, class-balanced labeled sample sets. Field data suffers from numerous problems such as class imbalance, scarcity of rare perturbation samples, and significant differences in recording devices and operating environments at different sites. These issues lead to models being prone to overfitting and exhibiting poor cross-domain generalization performance, making them unsuitable for scenarios with scarce samples and strict data privacy requirements. Summary of the Invention

[0006] The purpose of this invention is to provide a method for identifying power distribution network disturbance sources based on multimodal feature fusion, which aims to solve the technical problems of existing technologies, such as single feature representation, weak physical interpretability, and difficulty in forming stable power fingerprints.

[0007] To achieve the above objectives, the technical solution of this invention is: a method for identifying disturbance sources in power distribution networks based on multimodal feature fusion. This method solves the problems of unstable disturbance fingerprint extraction and insufficient robustness under noise in existing methods by fusing time-frequency image features, parameterized physical features, and modulus energy distribution features of traveling wave signals. Ultimately, it achieves hierarchical identification and reliable output for path initial judgment and disturbance source subdivision, and possesses comprehensive performance including physical interpretability, noise robustness, and cross-scenario transferability. The method includes the following steps:

[0008] S1: Collect three-phase current traveling wave data from the outgoing line side of the substation and perform preprocessing to extract disturbance-type current traveling wave data;

[0009] S2: Perform time-frequency transformation on the disturbance-type current traveling wave data to generate a panoramic waveform diagram and a two-dimensional time-frequency diagram, and fuse them into a three-channel time-frequency image;

[0010] S3: Extract the depth image features of the three-channel time-frequency image to obtain the image feature vector;

[0011] S4: Perform Prony mode parameter fitting on the disturbance-type current traveling wave data, extract the mode parameter feature vector, and project the parameter feature vector to a high-dimensional feature space to obtain the projected physical feature vector;

[0012] S5: Perform Clark transformation on the disturbance-type current traveling wave data to decompose the modulus, calculate the zero-mode energy proportion feature, and project the proportion feature to a high-dimensional feature space to obtain the projected modulus energy feature vector.

[0013] S6: Fuse the image feature vector, the projected physical feature vector, and the projected modulus energy feature vector to generate a fused feature vector;

[0014] S7: Input the fused feature vector into the classifier for classification decision, and output the category identification result of the power distribution network disturbance source.

[0015] Optionally, S2 specifically includes:

[0016] S2.1: Perform continuous wavelet transform on the time-domain traveling wave of the current within a preset time window, using the analytical Morse wavelet as the mother wavelet, and set the transform frequency range to... To cover the Nyquist frequency, the wavelet complex coefficient matrix is... The amplitude value is obtained Using linear amplitude as the energy intensity characterization, the intensity distribution of each phase in the frequency-time plane is obtained. A unified color label is applied to the three-phase results, defining the range. (Summary of three phases) Take the global minimum and maximum values ​​of all values ​​as the upper and lower limits of the color axis to obtain the panoramic waveform of the traveling wave.

[0017] S2.2: Call the continuous wavelet transform to calculate the analytical Morse wavelet coefficient matrix within a given frequency band. The energy matrix is ​​obtained by taking the square of the modulus of the complex coefficients, and its expression is:

[0018]

[0019] in, The energy matrix serves as a linear measure of time-frequency energy density. Time, frequency, and energy are mapped onto a two-dimensional image, and bilinear interpolation is used to uniformly scale it to 224×224 pixels to obtain a two-dimensional time-frequency image.

[0020] S2.3: Perform independent Min–Max normalization to [0,1] on the traveling wave panoramic waveform and the two-dimensional time-frequency diagram, respectively, with the expression:

[0021] ;

[0022] In the formula, For the k-th image, an independent Min–Max normalized pixel matrix is ​​generated, with pixel values ​​linearly mapped to the interval [0,1]. This matrix is ​​then used to feed into a convolutional network or as one of the RGB three channels for synthesis. This is the pixel matrix of the original k-th image. As a constant to prevent the denominator from being zero, the normalized images of the traveling wave panoramic waveform and the two-dimensional time-frequency image are mapped to the R / G / B channels respectively, and synthesized into a 224×224×3 three-channel time-frequency image.

[0023] Optionally, S3 specifically includes:

[0024] Using a standard ResNet-18 pre-trained on ImageNet, the model's built-in global average pooling (GAP) is retained, and the final fully connected classification layer is removed. For a three-channel time-frequency image with an input of 224×224×3, the following process is performed sequentially:

[0025] The initial 7×7 convolution and max pooling are used to output a 56×56×64 three-channel time-frequency image.

[0026] After passing through Stage 1, a 56×56×64 three-channel time-frequency image is output;

[0027] After passing through Stage 2, a 28×28×128 three-channel time-frequency image is output;

[0028] After passing through Stage 3, a 14×14×256 three-channel time-frequency image is output;

[0029] After passing through Stage 4, a 7×7×512 three-channel time-frequency image is output;

[0030] By applying global average pooling (GAP) to the 7×7 space, a 512-dimensional feature vector is obtained, expressed as:

[0031]

[0032] in, The output is a 512-dimensional image feature vector. There are 512 different eigenvalues. This is a transpose operation.

[0033] Optionally, S4 specifically includes:

[0034] The start and end points of the perturbation are located by the Teager energy operator, the Hankel matrix is ​​constructed and singular value decomposition is performed to determine the model order, and a 6-dimensional modal parameter vector including frequency, attenuation factor, time constant, quality factor, amplitude and component energy is extracted. The physical parameter vector is then projected onto a 128-dimensional feature space through a two-layer multilayer perceptron to obtain the projected physical feature vector.

[0035] Optionally, S5 specifically includes:

[0036] The zero-mode and line-mode current components are obtained by using the standard Clark transformation matrix. The ratio of zero-mode energy to total energy is calculated as the zero-mode energy proportion feature. The proportion feature is then projected to a 256-dimensional feature space through a linear projection layer to obtain the projected modulus energy feature vector.

[0037] Optionally, S6 specifically includes:

[0038] S6.1: For the image feature vector and the projected physical feature vector, a linear transformation is performed on them respectively through a learnable weight matrix and a bias vector to obtain the transformed image feature vector and the transformed physical feature vector.

[0039] S6.2: The projected modulus energy feature vector is subjected to dimensionality upscaling through a linear projection layer to obtain the dimensionality upscaled modulus energy feature vector;

[0040] S6.3: Employs a physical information-guided attention fusion mechanism, concatenating the transformed physical feature vector and the upgraded modulus energy feature vector as the query vector, and using the transformed image feature vector as the key vector and value vector. The similarity between the query vector and the key vector is calculated through scaled dot product attention to obtain the attention weights, and the value vectors are weighted and summed to generate a 256-dimensional fusion feature vector.

[0041] Optionally, S6.3 specifically includes:

[0042] S6.3.1: The transformed physical feature vector and the upgraded modulus energy feature vector are concatenated along the feature dimension to form a joint physical query vector. The joint physical query vector is then linearly transformed through a learnable weight matrix and a bias vector to generate the final query vector.

[0043] S6.3.2: The transformed image feature vector is linearly transformed by two different learnable weight matrices and bias vectors to generate key vectors and value vectors. The key vector is used to calculate the similarity with the final query vector, and the value vector is used to weight the image feature information.

[0044] S6.3.3: The similarity score matrix is ​​obtained by performing a dot product operation between the final query vector and the key vector. Then, it is scaled by dividing by the square root of the key vector dimension. Finally, the scaled score is normalized to a probability distribution using the Softmax function to obtain the attention weight matrix.

[0045] S6.3.4: Use the obtained attention weight matrix to perform a weighted summation of the value vectors to generate a fused feature vector;

[0046] S6.3.5: Use different parameter matrices for each attention head, and repeat S6.3.1-S6.3.4 a preset number of times. Concatenate the fused feature vectors output by each attention head in the feature dimension, and integrate the information through a final linear transformation layer to generate a 256-dimensional fused feature vector.

[0047] Optionally, S7 specifically includes:

[0048] A two-layer fully connected neural network classifier is used. The first layer maps 256-dimensional features to 128-dimensional features and applies ReLU activation and Dropout regularization. The second layer maps features to the total number of categories and outputs a probability distribution using the Softmax function. Finally, the disturbance type is determined based on the maximum probability value, thus outputting the category identification result of the disturbance source in the power distribution network.

[0049] The beneficial effects of this invention are as follows: It proposes a multimodal collaborative analysis framework that integrates time-frequency image features, parameterized physical features, and modulus-energy features; by designing a multi-head attention fusion mechanism with physical features as queries and image features as keys, it achieves dynamic feature interaction and enhancement guided by physical information; and by combining cross-modal contrastive learning loss constraints, it effectively improves the model's generalization ability and identification accuracy under small sample conditions. This provides a theoretical basis for applications such as traveling wave disturbance identification, intelligent waveform analysis, and protection action evaluation in power systems. Attached Figure Description

[0050] Figure 1 This is a flowchart of the steps of the present invention;

[0051] Figure 2 This is a panoramic waveform diagram of the traveling wave of this invention;

[0052] Figure 3 This is a low-frequency magnified view of the traveling wave panoramic waveform of this invention;

[0053] Figure 4 This is a two-dimensional time-frequency diagram of the traveling wave panorama of the present invention;

[0054] Figure 5 This is a comparison diagram of the original signal and the reconstructed signal of this invention;

[0055] Figure 6 This is a diagram of the sinusoidal signal after Prony fitting according to the present invention. Detailed Implementation

[0056] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.

[0057] Example 1: As Figure 1 As shown, a method for identifying disturbance sources in a distribution network based on multimodal feature fusion includes the following steps:

[0058] S1: Collect three-phase current traveling wave data from the outgoing line side of the substation and perform preprocessing to extract disturbance-type current traveling wave data;

[0059] Optionally, three-phase current traveling wave data of the transmission line are synchronously acquired from a waveform recording device with a sampling frequency of not less than 1MHz. Existing signal processing technology is used for event detection, and kurtosis features of signal change degree, three-phase unbalance features, and similarity coefficient features of quantized waveform distortion are extracted to remove fault traveling wave data and other useless traveling wave data, while retaining disturbance-type current traveling wave data.

[0060] S2: Perform time-frequency transformation on the disturbance-type current traveling wave data to generate a panoramic waveform diagram and a two-dimensional time-frequency diagram, and fuse them into a three-channel time-frequency image;

[0061] Optionally, integrity checks and preprocessing are performed on the current sequence of each phase in sequence. Missing values ​​are filled by linear interpolation at the endpoints, and then the sequence is mean-removed to eliminate DC components and slow drift.

[0062] S2.1: As Figure 2 The process of performing continuous wavelet transform on the current traveling wave and drawing the panoramic waveform diagram is as follows:

[0063] A continuous wavelet transform is performed on the time-domain traveling wave of the current within a preset time window. The mother wavelet used is the analytic Morse wavelet, and the transform frequency range is set to... To cover the Nyquist frequency, the wavelet complex coefficient matrix is... The amplitude value is obtained Using linear amplitude as the energy intensity characterization, the intensity distribution of each phase in the frequency-time plane is obtained. A unified color label is applied to the three-phase results, defining the range. (Summary of three phases) Take the global minimum and maximum values ​​of all values ​​as the upper and lower limits of the color axis to obtain the panoramic waveform of the traveling wave.

[0064] S2.2: As Figure 3 The process of locating the target frequency band, drawing the low-frequency magnified image, and generating a single-channel grayscale image after performing continuous wavelet transform on the current traveling wave, as shown in the low-frequency magnified image of the traveling wave panoramic waveform, is as follows:

[0065] Using a fixed colormap (jet), with time on the horizontal axis and frequency on the vertical axis, a single-channel grayscale matrix is ​​obtained and then scaled to 224×224 using bilinear interpolation. After locating the concentrated frequency band of the main energy of the disturbance segment, the target frequency band is defined, and a high-resolution CWT is re-executed. A local magnified image is generated with plotting parameters consistent with the panoramic waveform to obtain the details of the target frequency band at high, medium, and low frequencies. This is then plotted as a single-channel grayscale image and scaled to 224×224 using bilinear interpolation.

[0066] S2.3: As Figure 4 The process of performing continuous wavelet transform on the current traveling wave, measuring energy density, plotting the two-dimensional time-frequency diagram, and generating a single-channel grayscale image, as shown in the illustrated traveling wave panorama, is as follows:

[0067] Calculate the analytical Morse wavelet coefficient matrix within a given frequency band using continuous wavelet transform. The energy matrix is ​​obtained by taking the square of the modulus of the complex coefficients, and its expression is:

[0068]

[0069] in, The energy matrix serves as a linear measure of time-frequency energy density. Time, frequency, and energy are mapped onto a two-dimensional image, and bilinear interpolation is used to uniformly scale it to 224×224 pixels. The upper and lower limits of the color levels are fixed to ensure the comparability of different samples, resulting in a two-dimensional time-frequency image and a single-channel grayscale image.

[0070] S2.4: To avoid the impact of numerical differences in images on convolutional kernel learning, the traveling wave panoramic waveform and the two-dimensional time-frequency image are independently Min–Max normalized to [0,1], as expressed by:

[0071] ;

[0072] In the formula, For the k-th image, an independent Min–Max normalized pixel matrix is ​​generated, with pixel values ​​linearly mapped to the interval [0,1]. This matrix is ​​then used to feed into a convolutional network or as one of the RGB three channels for synthesis. This is the pixel matrix of the original k-th image. This is a constant used to prevent the denominator from being zero, in this embodiment.

[0073] The normalized images of the traveling wave panoramic waveform and the two-dimensional time-frequency image are mapped to the R / G / B channels respectively, and synthesized into a 224×224×3 three-channel time-frequency image.

[0074] It should be understood that this embodiment does not perform geometric enhancement or color dithering to ensure the consistency of time-frequency color semantics and geometric structure, and only performs the above-mentioned fixed bilinear scaling.

[0075] S3: Extract the depth image features of the three-channel time-frequency image to obtain the image feature vector;

[0076] Optionally, a standard ResNet-18 pre-trained on ImageNet is used, retaining the model's built-in global average pooling GAP and removing the final fully connected classification layer. For a three-channel time-frequency image with an input of 224×224×3, the following process is performed sequentially:

[0077] The initial 7×7 convolution and max pooling are used to output a 56×56×64 three-channel time-frequency image.

[0078] After passing through Stage 1, a 56×56×64 three-channel time-frequency image is output;

[0079] After passing through Stage 2, a 28×28×128 three-channel time-frequency image is output;

[0080] After passing through Stage 3, a 14×14×256 three-channel time-frequency image is output;

[0081] After passing through Stage 4, a 7×7×512 three-channel time-frequency image is output;

[0082] By applying global average pooling (GAP) to the 7×7 space, a 512-dimensional feature vector is obtained, expressed as:

[0083]

[0084] in, The output is a 512-dimensional image feature vector. There are 512 different eigenvalues. This is a transpose operation.

[0085] Optionally, this embodiment unfreezes Stage 3, Stage 4, and BatchNorm layers by default, and the learning rate is 3-10 times lower than that of the newly added classification head, in order to adapt to the statistical distribution of time-frequency data and avoid overfitting.

[0086] S4: Perform Prony mode parameter fitting on the disturbance-type current traveling wave data, extract the mode parameter feature vector, and project the parameter feature vector to a high-dimensional feature space to obtain the projected physical feature vector;

[0087] S4.1: As Figure 5 The diagram showing the comparison between the original signal and the reconstructed signal is shown. Figure 6 The Prony-fitted sinusoidal signal plot shown is used to locate the start and end points of the current traveling wave perturbation. The process of Prony mode parameter fitting and signal reconstruction is as follows:

[0088] The start and end points of the perturbation are located using the Teager energy operator. A Hankel matrix is ​​constructed and singular value decomposition is performed to determine the model order. A 6-dimensional modal parameter vector, including frequency, attenuation factor, time constant, quality factor, amplitude, and component energy, is extracted. This physical parameter vector is then projected onto a 128-dimensional feature space using a two-layer multilayer perceptron to obtain the projected physical feature vector, specifically:

[0089] For each phase current traveling wave sequence, the instantaneous energy is calculated point-by-point using the Teager energy operator, and the curve of energy change over time is obtained; the formula is:

[0090]

[0091] in, This represents the instantaneous energy estimate at the nth sampling point; This represents the current value at the nth sampling point. and These represent the current values ​​immediately before and after the given point. Adaptive thresholds are then set for the three-phase energy curves. The moment the energy first crosses the threshold is defined as the disturbance start point for that phase, and the moment the energy falls back and remains below the threshold for a certain number of samples is defined as the disturbance end point for that phase. The earliest start point is then taken as the global start point, and the latest end point as the global end point, forming a unified time window. This time window is used to extract a clean disturbance segment from the original three-phase currents. This disturbance segment serves as the input for subsequent Prony fitting.

[0092] Furthermore, the pure perturbation segment single-phase current sequence obtained through precise positioning using the Teager energy operator is first mean-removed, the window length is evenly distributed, and the number of rows is half the window length minus one (rounded down), while the number of columns is the window length minus the number of rows plus one. A Hankel matrix is ​​constructed by shifting the previous row of data one position to the right. Two adjacent sets of columns from this matrix are then used to form the first and second pencil matrices, respectively. Singular value decomposition is performed on the first set to obtain the singular value sequence. Three criteria are used comprehensively: the relative threshold compared to the largest singular value, the elbow method, and a manually set upper limit for the order to determine the effective model order. The decomposition results are then truncated according to this order, the system matrix is ​​calculated, and its eigenvalues ​​(i.e., discrete poles) are obtained. Based on the sampling period, the discrete poles are transformed into: attenuation factor, angular frequency, dominant frequency, time constant, and quality factor.

[0093] Using each pole as a basis, a Vandermonde matrix is ​​constructed. The complex coefficients of each mode are obtained using least squares, thus yielding the amplitude and phase. The time-domain components of a single mode are:

[0094]

[0095] in, For the real-valued component of the k-th mode at the n-th sampling point, These are complex coefficients (from which amplitude and phase are derived). For discrete poles, This indicates taking the real part.

[0096] Signal reconstruction specifically involves:

[0097]

[0098] in, This is the reconstructed signal at the nth sampling point. The modal energy is obtained by multiplying the square of the magnitude of the complex coefficients by the energy of the corresponding basis column vectors. The modes (dominant frequency, attenuation factor, time constant, quality factor, amplitude, and phase) are sorted in descending order of energy.

[0099] S4.2: A 6-dimensional parameter vector with definite physical meaning obtained by fitting the Prony algorithm. Through a nonlinear mapping function (implemented by an MLP), it is transformed into a higher-dimensional, more abstract, and discriminative feature space, generating feature vectors. To better integrate with subsequent attention fusion modules, specifically:

[0100] The 6-dimensional physical parameter vector from the Prony fitting module As input:

[0101]

[0102] in, , , , , , These parameters are frequency, attenuation factor, time constant, quality factor, amplitude, and component energy. Since the dimensions and numerical ranges of these parameters vary significantly, Z-score standardization is used. During the training phase, the mean μ and standard deviation σ of each feature dimension are calculated based on the entire training set, and the input vector is transformed accordingly. Each dimension of the standardized input vector approximately follows a distribution with a mean of 0 and a standard deviation of 1, laying the foundation for stable and efficient network training. The μ and σ obtained during the training phase are saved and used for the same transformation during the inference phase.

[0103] Specifically, a Multilayer Perceptron (MLP) consists of two fully connected layers (FC), an activation function in between, and an optional normalization layer. Its structure is: Input (6-dimensional) → First hidden layer (32-dimensional) → ReLU → Batch normalization layer → Second hidden layer (128-dimensional) → ReLU → Output (128-dimensional).

[0104] The first fully connected layer is specifically as follows:

[0105]

[0106] in, The input row vector is the standardized value. This is the weight matrix for the first layer, whose function is to linearly combine the 6-dimensional input features into 32 different neuron nodes. The bias vector of the first layer This is the output after the first level of linear transformation.

[0107] Nonlinear activation specifically refers to:

[0108]

[0109] in, Represents the constant 0, used in the definition of the ReLU function;

[0110] By introducing a nonlinear transformation, the ReLU function sets all negative values ​​to zero and retains positive values, enabling the model to learn the complex nonlinear interactions between input features.

[0111] The batch normalization layer is specifically as follows:

[0112]

[0113] in, This indicates a normalization operation, and BN1 represents the output after normalization.

[0114] The output A1 after the first activation function is normalized to stabilize its mean and variance around 0 and 1.

[0115] The second fully connected layer is specifically as follows:

[0116]

[0117] in, This is the weight matrix for the second layer. This is the bias vector for the second layer. This is the output after the second-level linear transformation.

[0118] Nonlinear activation specifically refers to:

[0119]

[0120] By reintroducing nonlinearity, the expressive power of the model is further enhanced.

[0121] The output A2 after the second ReLU activation is used as the final output, which is the projected physical feature vector.

[0122] S5: Perform Clark transformation on the disturbance-type current traveling wave data to decompose the modulus, calculate the zero-mode energy proportion feature, and project the proportion feature to a high-dimensional feature space to obtain the projected modulus energy feature vector.

[0123] Optionally, the zero-mode and line-mode current components are obtained using the standard Clark transformation matrix. The ratio of zero-mode energy to total energy is calculated as the zero-mode energy proportion feature. This proportion feature is then projected onto a 256-dimensional feature space through a linear projection layer to obtain the projected modulus energy feature vector, specifically:

[0124] The Clark transformation of three-phase currents is performed using the standard Clark transformation matrix, and the calculation formula is as follows:

[0125]

[0126] in, The zero-mode current component, and For line-mode current components, , , Let A, B, and C represent the phase currents in the three-phase current, respectively; calculate the zero-mode energy within the entire time window. The formula is:

[0127]

[0128] Where L is the length of the time window;

[0129] Calculate the combined linear mode energy of the two linear mode components. The formula is:

[0130]

[0131] Calculate the proportion of zero-mode energy in the total energy. This is transformed into a dimensionless scalar feature that can be used for machine learning. The larger this value, the more zero-sequence current components are present in the disturbance current, as shown in the formula:

[0132]

[0133] in, For the total energy, the obtained The scalar is updimensionalized into a high-dimensional feature vector, namely the modulus-energy feature vector, through a dedicated linear projection layer. The formula is:

[0134]

[0135] in, For a learnable weight matrix, It is a learnable bias vector.

[0136] S6: Fuse the image feature vector, the projected physical feature vector, and the projected modulus energy feature vector to generate a fused feature vector;

[0137] Optionally, each feature vector is first mapped to a shared, high-dimensional feature space of the same dimension through independent linear transformation layers to eliminate differences in dimensions and scales. Specifically:

[0138] S6.1: For the image feature vector and the projected physical feature vector, a linear transformation is performed on them respectively through a learnable weight matrix and a bias vector to obtain the transformed image feature vector and the transformed physical feature vector.

[0139] It is understandable that the transformed image feature vector retains the deep spatiotemporal information of the original features, while the transformed physical feature vector elevates the physical parameters that characterize the nature of the disturbance to a higher-dimensional feature space.

[0140] S6.2: The projected modulus energy feature vector is subjected to dimensionality upscaling through a linear projection layer to obtain an upscaled modulus energy feature vector, which enables it to interact with other modal features.

[0141] S6.3: Employs a physical information-guided attention fusion mechanism, concatenating the transformed physical feature vector and the upgraded modulus energy feature vector as the query vector, and using the transformed image feature vector as the key vector and value vector. The similarity between the query vector and the key vector is calculated through scaled dot product attention to obtain the attention weights, and the value vectors are weighted and summed to generate a 256-dimensional fusion feature vector.

[0142] S6.3.1: The transformed physical feature vector and the upgraded modulus energy feature vector are concatenated along the feature dimension to form a joint physical query vector. The joint physical query vector is then linearly transformed through a learnable weight matrix and a bias vector to generate the final query vector.

[0143] S6.3.2: The transformed image feature vector is linearly transformed by two different learnable weight matrices and bias vectors to generate key vectors and value vectors. The key vector is used to calculate the similarity with the final query vector, and the value vector is used to weight the image feature information.

[0144] S6.3.3: The similarity score matrix is ​​obtained by performing a dot product operation between the final query vector and the key vector. Then, it is scaled by dividing by the square root of the key vector dimension. Finally, the scaled score is normalized to a probability distribution using the Softmax function to obtain the attention weight matrix.

[0145] S6.3.4: The value vector is weighted and summed using the obtained attention weight matrix to generate a fusion feature vector. The fusion feature vector can be regarded as the essence summary of image features after physical information filtering and weighting, highlighting the visual evidence most relevant to the current physical parameters.

[0146] S6.3.5: To improve the model's capacity and expressive power, different parameter matrices are used for each attention head, allowing the model to focus on different types of information in different feature subspaces. Repeat S6.3.1-S6.3.4 a preset number of times, concatenate the fused feature vectors output by each attention head along the feature dimension, and integrate the information through a final linear transformation layer to generate a 256-dimensional fused feature vector.

[0147] S7: Input the fused feature vector into the classifier for classification decision, and output the category identification result of the power distribution network disturbance source.

[0148] Optionally, a two-layer fully connected neural network classifier is used. The first layer maps the 256-dimensional features to 128 dimensions and applies ReLU activation and Dropout regularization. The second layer maps to the total number of categories and outputs a probability distribution using the Softmax function. Finally, the disturbance type is determined based on the maximum probability value, thus outputting the category identification result of the distribution network disturbance source. Specifically:

[0149] The first fully connected layer transforms the 256-dimensional fused feature vector into a 128-dimensional intermediate feature vector and introduces non-linear computational capabilities through a linear rectified activation function. During the training phase, random deactivation regularization is applied to the output of this layer, randomly setting the activation values ​​of a portion of neurons to zero with a 50% probability, effectively preventing model overfitting.

[0150] The second fully connected layer maps the regularized 128-dimensional feature vector to the dimension of the total number of categories; this output is called the log-odds vector. Subsequently, the log-odds vector is converted into a probability distribution using the Softmax function, where the value of each element in the vector represents the predicted probability that the input sample belongs to the corresponding perturbation category.

[0151] The final perturbation type determination result is determined by the category number corresponding to the maximum probability value in the probability distribution, thus completing the end-to-end processing flow from multimodal feature fusion to perturbation type identification.

[0152] Optionally, the model training employs a hybrid loss function for end-to-end optimization. The primary loss is cross-entropy loss, which directly measures the difference between the predicted probability distribution and the true label, dominating the optimization of classification accuracy. Simultaneously, a contrastive learning loss term is introduced to constrain the distance between different modal features from the same sample in the projection space, bringing positive sample pairs closer together and negative sample pairs further apart, thereby improving the model's feature representation quality and generalization ability.

[0153] The total loss is a weighted sum of the two losses mentioned above. The gradient descent algorithm and its variants are used to iteratively optimize all learnable parameters in the model (including the weights and biases of each projection layer, attention mechanism, and classifier) ​​until the model converges.

[0154] Based on the specific implementation details, the effectiveness of the technical solution of the present invention will be demonstrated through experiments.

[0155] Specifically, the experimental data comes from a traveling wave acquisition platform, which has deployed 19 high-speed acquisition devices on a 110kV network and 20 high-speed acquisition devices on a 1MHz network. The sampling frequency is 1MHz, and the recorded data includes traveling wave records from August 2012 to October 2014, totaling 5976 records, which are stored in COMTRADE format.

[0156] Step 1: Use existing signal processing techniques to perform event detection, extract kurtosis features of signal abrupt change, three-phase imbalance features, and similarity coefficient features of quantized waveform distortion to remove fault traveling wave data and other useless traveling wave data, and retain disturbance-type current traveling wave data.

[0157] Instantaneous energy sequences were calculated for each of the three-phase currents, yielding curves showing energy variation over time. For each phase, an adaptive threshold was used to identify the moment when the energy first exceeded the threshold as the start of the disturbance for that phase, and the moment when the energy fell back and remained below the threshold for a certain number of samples as the end of the disturbance for that phase.

[0158] The earliest start point of the three phases is used as the global start point, and the latest end point is used as the global end point, forming a unified disturbance time window. This window is then used to extract a clean disturbance segment from the original three-phase data. This window is used for subsequent Prony modal recognition and feature extraction, effectively suppressing background noise and irrelevant steady-state components.

[0159] Step 2: Perform continuous wavelet transform on the perturbation segment and take the amplitude of the coefficients as the energy index. In a unified time-frequency coordinate system, the three-phase energy surfaces are drawn as a three-dimensional panoramic view by superimposing transparency. The color represents the energy intensity, and the range is consistent within the global color axis after the three phases are merged to ensure inter-phase comparability and visual consistency.

[0160] The main energy concentration bands are determined on the cumulative distribution of panoramic energy. While keeping the complete time axis unchanged, higher resolution wavelet analysis and plotting are performed only on this frequency band to highlight the accumulation and attenuation rhythm of low-frequency energy.

[0161] The time-frequency energy is mapped to a two-dimensional heatmap, with fixed upper and lower limits for color levels and uniform interpolation to the resolution required by the network. If both panoramic and detail images are used simultaneously, they can be encoded as a three-channel pseudo-color input to improve the network's sensitivity to multi-scale textures.

[0162] Using a ResNet-18 model pre-trained on ImageNet, the last fully connected layer is removed, and the output layer is a global average pooling layer, ultimately extracting a 512-dimensional image feature vector. During training, only the second half of the network layers and the normalization layer are fine-tuned to maintain generalization ability while conforming to the domain characteristics of time-frequency textures.

[0163] Step 3: Construct the Hankel matrix and adjacent pencil matrices for the perturbation segment. Determine the effective order based on the strong-weak separation of singular value decomposition and the elbow position, and set an upper limit on the order to prevent overfitting. This process yields the discrete poles of several damped oscillation modes.

[0164] The poles are mapped to parameters such as attenuation factor, dominant frequency, time constant, and quality factor in the continuous domain; only physically meaningful damping modes are retained. Simultaneously, the complex coefficients of each mode are obtained through least-squares fitting, thus yielding the amplitude and phase.

[0165] The time-domain signal is reconstructed by superimposing the various modes, and the signal-to-noise ratio within the window is calculated as the fitting quality index. Several dominant component curves after energy sorting and corresponding parameter tables are derived to facilitate diagnosis and source tracing.

[0166] The six physical parameters of the dominant component—frequency, attenuation factor, time constant, quality factor, amplitude, and component energy—are extracted to form a physical feature vector. .

[0167] Will The input is a two-layer MLP, which is then projectively upscaled. The network structure is as follows: Input (6-dimensional) → First hidden layer (32-dimensional) → ReLU → Batch normalization layer → Second hidden layer (128-dimensional) → ReLU → Output (128-dimensional), outputting a 128-dimensional physical feature vector. .

[0168] Step 4: Perform Clarke transform on the three-phase current to obtain the zero-mode and line-mode components. Calculate the zero-mode and line-mode energies within the disturbance window and give the proportion of zero-mode energy in the total energy. , scalar Projecting the modulus energy feature vector through a linear layer to 256 dimensions yields the modulus energy feature vector. .

[0169] Step 5: , , Each feature is projected to 256 dimensions through independent linear layers. The concatenated vector of parametric and modal domain features serves as the query, while image features are used as keys and values. A multi-head attention mechanism aligns information sources in parallel across multiple subspaces, enabling physical priors to "actively inquire" about image evidence, thereby improving cross-modal consistency and interpretability. The fused output is a fixed-length comprehensive feature. A two-layer fully connected structure with random deactivation is used for classification, and the probability distribution of the perturbation category is output at the end. The final identification result is given based on the maximum probability criterion.

[0170] In this embodiment, cross-entropy is used as the primary supervisory loss; cross-modal contrastive loss is introduced to make the different modal features of the same sample closer in the projection space and the different samples more separated, thereby enhancing the discriminativeness and generalization of the fused sample. The two losses are weighted and summed. The total loss function is... The formula is:

[0171]

[0172] in, For cross-entropy loss, To compare the losses, As weight.

[0173] Furthermore, an adaptive optimizer with a weight decay term is adopted; a group learning rate is used for the pre-trained backbone and the newly built layers, with the learning rate of the backbone being one-third to one-tenth of that of the newly built layers; overfitting is reduced through random deactivation, early stopping, and batch normalization; and data augmentation is performed on the input signal to fit the physical scene, such as amplitude clipping, mild noise perturbation, and window randomization.

[0174] Furthermore, this embodiment also acquired 6400 current traveling wave data points. Experiments were conducted based on a framework using Python 3.08, CUDA version 12.8, and PyTorch version 2.6.0. Ten disturbance categories were selected as labels. The panoramic waveforms and panoramic two-dimensional time-frequency diagrams of different types of disturbances, Prony fitting modal parameters, and zero-mode energy ratios were used as inputs. A distribution network disturbance source identification model based on multi-modal feature fusion was used to identify the ten disturbance categories, and the overall classification accuracy reached 95.1%.

[0175] In summary, this invention proposes a multimodal fusion analysis framework that integrates time-frequency image features, parametric modal features, and modal domain energy features. This method acquires three-phase current traveling wave data from the outgoing line side of a substation, preprocesses it to extract effective disturbance traveling wave signals, constructs a three-channel time-frequency image containing both panoramic and detailed information using time-frequency transformation technology, extracts depth features from the image using a deep learning network, extracts physically meaningful traveling wave modal features through parametric fitting, obtains zero-mode energy proportion features by combining modulus analysis, and innovatively designs a multi-head attention fusion mechanism guided by physical features to achieve effective alignment and deep fusion of multimodal features. Finally, a classifier is used to accurately identify the disturbance type. It maintains high identification accuracy and robustness even in high-noise environments, providing an efficient solution for traveling wave disturbance identification in power systems.

[0176] The specific embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the spirit of the present invention.

Claims

1. A method for identifying disturbance sources in a distribution network based on multimodal feature fusion, characterized in that, The method includes the following steps: S1: Collect three-phase current traveling wave data from the outgoing line side of the substation and perform preprocessing to extract disturbance-type current traveling wave data; S2: Perform time-frequency transformation on the disturbance-type current traveling wave data to generate a panoramic waveform diagram and a two-dimensional time-frequency diagram, and fuse them into a three-channel time-frequency image; S3: Extract the depth image features of the three-channel time-frequency image to obtain the image feature vector; S4: Perform Prony mode parameter fitting on the disturbance-type current traveling wave data, extract the mode parameter feature vector, and project the parameter feature vector to a high-dimensional feature space to obtain the projected physical feature vector; S5: Perform Clark transformation on the disturbance-type current traveling wave data to decompose the modulus, calculate the zero-mode energy proportion feature, and project the proportion feature to a high-dimensional feature space to obtain the projected modulus energy feature vector. S6: Fuse the image feature vector, the projected physical feature vector, and the projected modulus energy feature vector to generate a fused feature vector; S7: Input the fused feature vector into the classifier for classification decision, and output the category identification result of the power distribution network disturbance source; Specifically, S6 is: S6.1: For the image feature vector and the projected physical feature vector, a linear transformation is performed on them respectively through a learnable weight matrix and a bias vector to obtain the transformed image feature vector and the transformed physical feature vector. S6.2: The projected modulus energy feature vector is subjected to dimensionality upscaling through a linear projection layer to obtain the dimensionality upscaled modulus energy feature vector; S6.3: Employs a physical information-guided attention fusion mechanism, concatenating the transformed physical feature vector and the upgraded modulus energy feature vector as the query vector, and using the transformed image feature vector as the key vector and value vector. The similarity between the query vector and the key vector is calculated through scaled dot product attention to obtain the attention weights, and the value vectors are weighted and summed to generate a 256-dimensional fusion feature vector. Specifically, S6.3 is as follows: S6.3.1: The transformed physical feature vector and the upgraded modulus energy feature vector are concatenated along the feature dimension to form a joint physical query vector. The joint physical query vector is then linearly transformed through a learnable weight matrix and a bias vector to generate the final query vector. S6.3.2: The transformed image feature vector is linearly transformed by two different learnable weight matrices and bias vectors to generate key vectors and value vectors. The key vector is used to calculate the similarity with the final query vector, and the value vector is used to weight the image feature information. S6.3.3: The similarity score matrix is ​​obtained by performing a dot product operation between the final query vector and the key vector. Then, it is scaled by dividing by the square root of the key vector dimension. Finally, the scaled score is normalized to a probability distribution using the Softmax function to obtain the attention weight matrix. S6.3.4: Use the obtained attention weight matrix to perform a weighted summation of the value vectors to generate a fused feature vector; S6.3.5: Use different parameter matrices for each attention head, and repeat S6.3.1-S6.3.4 a preset number of times. Concatenate the fused feature vectors output by each attention head along the feature dimension, and integrate the information through a final linear transformation layer to generate a 256-dimensional fused feature vector.

2. The method for identifying power distribution network disturbance sources based on multimodal feature fusion according to claim 1, characterized in that, Specifically, S2 is: S2.1: Perform continuous wavelet transform on the time-domain traveling wave of the current within a preset time window, using the analytical Morse wavelet as the mother wavelet, and set the transform frequency range to [specified value]. To cover the Nyquist frequency, the wavelet complex coefficient matrix is... The amplitude value is obtained Using linear amplitude as the energy intensity characterization, the intensity distribution of each phase in the frequency-time plane is obtained. A unified color label is applied to the three-phase results, defining the range. (Summary of three phases) Take the global minimum and maximum values ​​of all values ​​as the upper and lower limits of the color axis to obtain the panoramic waveform of the traveling wave. S2.2: Call the continuous wavelet transform to calculate the analytical Morse wavelet coefficient matrix within a given frequency band. The energy matrix is ​​obtained by taking the square of the modulus of the complex coefficients, and its expression is: ; in, The energy matrix serves as a linear measure of time-frequency energy density. Time, frequency, and energy are mapped onto a two-dimensional image, and bilinear interpolation is used to uniformly scale it to 224×224 pixels to obtain a two-dimensional time-frequency image. S2.3: Perform independent Min–Max normalization to [0,1] on the traveling wave panoramic waveform and the two-dimensional time-frequency diagram, respectively, with the expression: ; In the formula, For the k-th image, an independent Min–Max normalized pixel matrix is ​​generated, with pixel values ​​linearly mapped to the interval [0,1]. This matrix is ​​then used to feed into a convolutional network or as one of the RGB three channels for synthesis. This is the pixel matrix of the original k-th image. As a constant to prevent the denominator from being zero, the normalized images of the traveling wave panoramic waveform and the two-dimensional time-frequency image are mapped to the R / G / B channels respectively, and synthesized into a 224×224×3 three-channel time-frequency image.

3. The method for identifying power distribution network disturbance sources based on multimodal feature fusion according to claim 2, characterized in that, Specifically, S3 is: Using a standard ResNet-18 pre-trained on ImageNet, the model's built-in global average pooling (GAP) is retained, and the final fully connected classification layer is removed. For a three-channel time-frequency image with an input of 224×224×3, the following process is performed sequentially: The initial 7×7 convolution and max pooling are used to output a 56×56×64 three-channel time-frequency image. After passing through Stage 1, a 56×56×64 three-channel time-frequency image is output; After passing through Stage 2, a 28×28×128 three-channel time-frequency image is output; After passing through Stage 3, a 14×14×256 three-channel time-frequency image is output; After passing through Stage 4, a 7×7×512 three-channel time-frequency image is output; By applying global average pooling (GAP) to the 7×7 space, a 512-dimensional feature vector is obtained, expressed as: ; in, The output is a 512-dimensional image feature vector. There are 512 different eigenvalues. This is a transpose operation.

4. The method for identifying power distribution network disturbance sources based on multimodal feature fusion according to claim 1, characterized in that, Specifically, S4 is: The start and end points of the perturbation are located by the Teager energy operator, the Hankel matrix is ​​constructed and singular value decomposition is performed to determine the model order, and a 6-dimensional modal parameter vector including frequency, attenuation factor, time constant, quality factor, amplitude and component energy is extracted. The physical parameter vector is then projected onto a 128-dimensional feature space through a two-layer multilayer perceptron to obtain the projected physical feature vector.

5. The method for identifying power distribution network disturbance sources based on multimodal feature fusion according to claim 1, characterized in that, Specifically, S5 is: The zero-mode and line-mode current components are obtained by using the standard Clark transformation matrix. The ratio of zero-mode energy to total energy is calculated as the zero-mode energy proportion feature. The proportion feature is then projected to a 256-dimensional feature space through a linear projection layer to obtain the projected modulus energy feature vector.

6. The method for identifying power distribution network disturbance sources based on multimodal feature fusion according to claim 1, characterized in that, Specifically, S7 is: A two-layer fully connected neural network classifier is used. The first layer maps 256-dimensional features to 128-dimensional features and applies ReLU activation and Dropout regularization. The second layer maps features to the total number of categories and outputs a probability distribution using the Softmax function. Finally, the disturbance type is determined based on the maximum probability value, thus outputting the category identification result of the disturbance source in the power distribution network.

Citation Information

Patent Citations

  • Power distribution network power quality disturbance analysis method based on bimodal feature fusion

    CN115659254A

  • Non-intrusive load identification method and system based on multi-modal depth feature fusion

    CN120470456A