Intelligent Fault Location and Diagnosis Method for Optical Fiber Composite in Distribution Network Based on Hybrid Sensing

Through the adaptive feature fusion of hybrid sensing data and the construction of dynamic fault propagation graphs, the problem of insufficient utilization of multi-source data in fault positioning and diagnosis of existing distribution networks is solved, and fault positioning and diagnosis with higher accuracy and reliability is achieved.

CN119881542BActive Publication Date: 2025-07-04FOSHAN GUYUXUAN BRAND MANAGEMENT CO LTD

Patent Information

Application Number
CN202510354682.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-04
Estimated Expiration
2045-03-25

AI Technical Summary

Technical Problem

The existing distribution network fault positioning and diagnosis methods fail to make full use of multi-source heterogeneous data, resulting in limited fault positioning accuracy and reliability. It is difficult for traditional methods to adapt to the contribution differences of various sensor data under different fault types, and it is impossible to accurately characterize the fault propagation characteristics and impact range.

Method used

Adaptive feature extraction and weighted fusion of sensing data is carried out through multi-channel convolutional neural network and attention mechanism, and spatial-temporal feature extraction is carried out in combination with graph convolution and causal convolution to construct dynamic fault propagation graphs to realize fault location and type recognition.

Benefits of technology

It improves fault positioning accuracy and diagnostic reliability, enhances the ability to capture fault propagation laws, and improves the accuracy and real-timeness of fault diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119881542B_ABST
    Figure CN119881542B_ABST
Patent Text Reader

Abstract

The present invention provides a method for intelligent fault location and diagnosis of optical fiber composite faults in a distribution network based on hybrid sensing, which relates to the technical field of distribution networks. The method includes acquiring and preprocessing optical fiber and electrical sensing data, extracting time-frequency domain features through a multi-channel convolutional neural network, calculating dynamic weight coefficients using an attention mechanism for feature fusion, constructing a dynamic fault propagation graph, and performing spatio-temporal feature extraction by combining graph convolution and causal convolution to achieve fault location and type identification. The present invention can make full use of multi-source heterogeneous data, improve the accuracy of fault location and the reliability of diagnosis, and is applicable to fault diagnosis of complex distribution networks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technology of distribution networks, and particularly to an intelligent fault location and diagnosis method for optical fiber composite faults in distribution networks based on hybrid sensing. Background Art

[0002] The distribution network is an important part of the power system, and its safe and stable operation is directly related to the reliable power supply for users. With the continuous expansion of the scale of the distribution network and the improvement of the intelligent level, the rapid and accurate location and diagnosis of distribution network faults have become the key technologies to ensure power supply reliability. At present, the fault location of the distribution network mainly relies on the analysis of data such as voltage and current collected by electrical sensors. At the same time, due to its advantages such as anti-electromagnetic interference and distributed measurement, optical fiber sensing technology has been widely used in the monitoring of distribution networks.

[0003] The existing distribution network fault location and diagnosis methods mainly have the following deficiencies: First, traditional methods often rely solely on electrical sensing data or optical fiber sensing data for fault analysis, and fail to fully utilize the rich information contained in multi-source heterogeneous data, resulting in limited fault location accuracy and reliability. Second, most of the existing data fusion methods adopt simple feature splicing or fixed-weight fusion, and cannot effectively adapt to the contribution differences of various sensing data under different fault types, reducing the accuracy of fault diagnosis. Third, traditional fault location methods rarely consider the complex correlations of physical topology, electrical characteristics and time-series evolution in the distribution network, and it is difficult to accurately describe the fault propagation characteristics and fault influence range, affecting the accuracy and real-time performance of fault location.

[0004] With the development of artificial intelligence technology, deep learning methods have shown good application prospects in the field of fault diagnosis. However, how to effectively fuse multi-source heterogeneous data, adaptively extract fault features, and accurately model the fault propagation mechanism is still a key scientific problem to be solved urgently. This poses an urgent need for the development of new intelligent fault location and diagnosis methods for distribution network faults. Summary of the Invention

[0005] An embodiment of the present invention provides an intelligent fault location and diagnosis method for optical fiber composite faults in distribution networks based on hybrid sensing, which can solve the problems in the prior art.

[0006] In the first aspect of the embodiment of the present invention,

[0007] An intelligent fault location and diagnosis method for optical fiber composite faults in distribution networks based on hybrid sensing is provided, including:

[0008] Obtain the optical fiber sensing data and electrical sensing data of each monitoring node in the distribution network, and perform data preprocessing on the optical fiber sensing data and electrical sensing data, including removing outliers, normalizing, and aligning time series, to obtain a standardized sensing data sequence;

[0009] Input the standardized sensing data sequences into the corresponding input channels of a pre-trained multi-channel convolutional neural network according to the data types, extract time-frequency domain features through the convolutional layer, calculate the dynamic weight coefficients of different sensing data using the attention mechanism, and adaptively weighted fuse the features of different types of sensing data according to the dynamic weight coefficients to obtain a fused feature vector;

[0010] Construct a fault feature space based on the fused feature vector, and input the fault feature space into a fault diagnosis model pre-trained with historical fault data; construct a dynamic fault propagation graph that includes physical, electrical, and temporal correlations, perform spatio-temporal feature extraction by combining graph convolution and causal convolution, use the attention mechanism to update the node state through message passing, fuse local features and global features, and output the fault location result and the fault type recognition result after optimization by the multi-task loss function.

[0011] Input the standardized sensing data sequences into the corresponding input channels of a pre-trained multi-channel convolutional neural network according to the data types, extract time-frequency domain features through the convolutional layer, calculate the dynamic weight coefficients of different sensing data using the attention mechanism, and adaptively weighted fuse the features of different types of sensing data according to the dynamic weight coefficients to obtain a fused feature vector, including:

[0012] Input the standardized sensing data sequences into the corresponding feature extraction channels according to the data types. The feature extraction channels include a one-dimensional convolutional network channel for one-dimensional data, a two-dimensional convolutional network channel for two-dimensional data, and a dilated convolutional network channel for long-range dependence modeling. Extract the features of various types of data through the feature extraction channels to obtain initial feature maps;

[0013] Perform multi-scale feature extraction on the initial feature maps, construct a feature pyramid network with multiple levels, extract features through convolutional operations at each level, combine the features of the first level with the features of the second level through upsampling, and use adaptive weight coefficients to weighted aggregate the features of different levels to obtain hierarchically enhanced multi-scale features;

[0014] Perform dual attention enhancement on the hierarchically enhanced multi-scale features. Calculate the channel attention weights by combining the average pooling feature and the maximum pooling feature in the channel dimension and recalibrate the feature channels. Calculate the spatial attention weight matrix based on the local region correlation of the feature maps in the spatial dimension and achieve feature enhancement in the spatial dimension. Perform residual connection between the spatially enhanced features and the original features to obtain a doubly enhanced feature representation;

[0015] Global average pooling is performed on the doubly enhanced feature representation to obtain channel-level feature descriptions and construct an inter-modal correlation matrix. The correlation matrix is input into a multi-head attention module and a weight generation network respectively to calculate dynamic fusion weight coefficients, and local enhanced features and global context vectors are adaptively fused to achieve adaptive fusion of multi-modal features;

[0016] A joint optimization objective function including a cross-entropy loss term, a reconstruction error term, and a regularization term is constructed. The parameters of the feature extraction network, the attention module, and the weight generation network are optimized end-to-end by minimizing the joint optimization objective function, and the final fused feature representation is output.

[0017] Doubly attention enhancement is performed on the hierarchically enhanced multi-scale features. Channel attention weights are calculated by combining average pooling and max pooling features in the channel dimension to recalibrate the feature channels. In the spatial dimension, a spatial attention map is calculated based on the local region correlation of the feature map to achieve feature enhancement in the spatial dimension. The doubly enhanced feature representation includes:

[0018] In the channel dimension, global average pooling operations and global max pooling operations are respectively performed on the hierarchically enhanced multi-scale features to obtain an average pooling feature vector and a max pooling feature vector in the channel dimension. The average pooling feature vector and the max pooling feature vector are respectively input into a multi-layer perceptron with shared weights for non-linear transformation to obtain two intermediate feature vectors;

[0019] The two intermediate feature vectors are added together and processed by a sigmoid function to obtain a channel attention weight vector. The channel attention weight vector is used to recalibrate the importance of the hierarchically enhanced multi-scale features in the channel dimension to obtain channel-enhanced features;

[0020] In the spatial dimension, the channel-enhanced features are divided into a preset number of local feature regions, the correlation between different local feature regions is calculated to obtain a regional correlation matrix, and the regional correlation matrix is convolved and normalized by a softmax function to obtain a spatial attention weight matrix;

[0021] The spatial attention weight matrix is weighted and fused with the channel-enhanced features in the spatial dimension to achieve feature enhancement in the spatial dimension, obtaining spatially enhanced features. The spatially enhanced features are connected with the hierarchically enhanced multi-scale features through a residual connection to obtain a doubly enhanced feature representation.

[0022] Global average pooling is performed on the doubly enhanced feature representation to obtain channel-level feature descriptions and construct an inter-modal correlation matrix. The correlation matrix is respectively input into a multi-head attention module and a weight generation network to calculate dynamic fusion weight coefficients, and local enhanced features and global context vectors are adaptively fused, realizing the adaptive fusion of multi-modal features, including:

[0023] Global average pooling operation is performed on the doubly enhanced features to obtain a channel-level feature description vector. Based on the channel-level feature description vector, the correlation coefficients between different modal features are calculated to construct an inter-modal correlation matrix;

[0024] The channel-level feature description vector is respectively mapped into query vectors, key vectors, and value vectors through a learnable projection matrix. The query vectors, key vectors, and value vectors are input into the multi-head attention module. Attention weights are calculated in each attention head and multiplied by the value vectors to obtain attention outputs. The outputs of multiple attention heads are concatenated and linearly transformed to obtain modal interaction features;

[0025] The inter-modal correlation matrix and the modal interaction features are concatenated in terms of features and input into a multi-layer perceptron for non-linear transformation to obtain initial fusion weights. A learnable temperature parameter is introduced to scale the initial fusion weights and normalized through the softmax function to obtain dynamic fusion weights;

[0026] The dynamic fusion weights are used to perform weighted summation on the doubly enhanced features of different modalities to obtain local enhanced features. Global average pooling is performed on the local enhanced features to obtain global context vectors, and broadcast multiplication operations are performed on the global context vectors and the doubly enhanced features of the corresponding modalities;

[0027] The importance weights of each modality are calculated, and the product results of the importance weights and the local enhanced features and the global context vectors are multiplied and summed respectively;

[0028] A feature discrimination loss function, a feature reconstruction loss function, and a weight regularization loss function are constructed. The feature discrimination loss function, the feature reconstruction loss function, and the weight regularization loss function are weighted and combined through a cosine-adjusted dynamic trade-off factor to obtain an overall loss function;

[0029] Based on the overall loss function, the projection matrix, the multi-layer perceptron, the temperature parameter, and the modal importance weights are jointly optimized to realize the adaptive fusion of multi-modal features.

[0030] Construct a dynamic fault propagation graph containing physical, electrical, and temporal correlations, perform spatio-temporal feature extraction by combining graph convolution and causal convolution, use the attention mechanism to update node states through message passing, fuse local features and global features, and output fault location results and fault type recognition results after optimization by a multi-task loss function, including:

[0031] Construct a physical topology correlation matrix based on the connection relationships of physical nodes in the distribution network, construct an electrical coupling relationship matrix based on node voltages, line currents, and power parameters, construct a time-series dynamic correlation matrix based on the historical state sequences within a time window, and construct a dynamic fault propagation graph by adaptively weighting and fusing the physical topology correlation matrix, the electrical coupling relationship matrix, and the time-series dynamic correlation matrix;

[0032] Perform graph convolutional operations on the dynamic fault propagation graph to extract spatial correlation features between nodes, perform causal convolutional operations on the spatial correlation features to extract time-series evolution features, and fuse the spatial correlation features and the time-series evolution features through an adaptive weight coefficient to obtain spatio-temporal joint features;

[0033] Map the spatio-temporal joint features to query vectors, key vectors, and value vectors, calculate the dot product similarity between the query vector and the key vector to obtain attention weights, multiply the attention weights by the value vectors and perform message aggregation, and use a gated recurrent unit to update the node states to achieve the dynamic transmission of fault information;

[0034] Perform feature pooling within the neighborhood range of each node to extract local feature representations, perform global pooling on all nodes of the dynamic fault propagation graph to extract global feature representations, and adaptively fuse the local feature representations and the global feature representations through learnable weights;

[0035] Calculate the fault location probability and the fault type probability based on the fused features, construct a multi-task loss function including a fault location loss term and a fault classification loss term, optimize the multi-task loss function using a dynamic weight strategy, and output the fault location result and the fault type recognition result.

[0036] Constructing a dynamic fault propagation graph by adaptively weighting and fusing the physical topology correlation matrix, the electrical coupling relationship matrix, and the time-series dynamic correlation matrix based on the connection relationships of physical nodes in the distribution network includes:

[0037] Obtain the impedance data of the connection paths between nodes in the distribution network and the node electrical hierarchy data, calculate the electrical distance between nodes based on the impedance data, calculate the hierarchical difference between nodes based on the electrical hierarchy data, and generate a physical topology correlation matrix by adaptively weighting and fusing the electrical distance between nodes and the hierarchical difference between nodes;

[0038] Collect the voltage amplitude, phase angle data, active power data, and reactive power data of the distribution network nodes. Calculate the voltage coupling strength between nodes based on the voltage amplitude and phase angle data, calculate the line power flow coupling degree based on the active power data and reactive power data, and generate an electrical coupling matrix by fusing the voltage coupling strength between nodes and the line power flow coupling degree through dynamic weights;

[0039] Construct the historical state sequence of the distribution network nodes, calculate the temporal correlation coefficient of the historical state sequence to obtain the temporal correlation matrix, calculate the time change rate of the temporal correlation matrix to obtain the dynamic change characteristics, and generate a temporal feature matrix by fusing the temporal correlation matrix and the dynamic change characteristics through adaptive weights;

[0040] Obtain a normalized feature matrix by performing multi-layer non-linear transformations on the physical topology correlation matrix, electrical coupling matrix, and temporal feature matrix. Construct self-attention feature representations and modal interaction features in multiple attention subspaces based on the normalized feature matrix. Perform multi-scale pooling fusion on the self-attention features and modal features and update through dynamic gating to obtain an enhanced modal feature representation. Perform adaptive weight fusion on the enhanced modal features to construct a dynamic fault propagation graph.

[0041] Obtain a normalized feature matrix by performing multi-layer non-linear transformations on the physical topology correlation matrix, electrical coupling matrix, and temporal feature matrix. Construct self-attention feature representations and modal interaction features in multiple attention subspaces based on the normalized feature matrix. Perform multi-scale pooling fusion on the self-attention features and modal features and update through dynamic gating to obtain an enhanced modal feature representation, including:

[0042] Perform feature mapping on the physical topology correlation matrix, electrical coupling matrix, and temporal feature matrix respectively through multi-layer non-linear transformations to obtain an initial feature matrix, and perform mean normalization processing on the initial feature matrix to obtain a normalized feature matrix;

[0043] Construct a query matrix, key matrix, and value matrix based on the normalized feature matrix. Project the query matrix, key matrix, and value matrix into multiple attention subspaces respectively. Calculate the self-attention weights within each attention subspace, and weight the self-attention weights with the attention mask matrix to obtain the modal internal feature representation;

[0044] Calculate the modal attention weights between different modal features, weight and aggregate the value matrices of different modalities based on the modal attention weights, and obtain the modal interaction features through a learnable output mapping matrix;

[0045] Perform residual connection and normalization on the modal internal feature representation and the modal interaction features, construct a feature pyramid by performing multi-scale pooling on the normalized features, and adaptively fuse the feature pyramid through learnable scale weights to obtain a multi-scale feature representation;

[0046] Calculate the importance scores of the multi-scale feature representation and the normalized feature matrix, and perform dynamic gating update on the multi-scale feature representation based on the importance scores to obtain an enhanced modal feature representation.

[0047] In the second aspect of the embodiments of the present invention,

[0048] Provide an electronic device, including:

[0049] A processor;

[0050] A memory for storing instructions executable by the processor;

[0051] Wherein, the processor is configured to call the instructions stored in the memory to execute the foregoing method.

[0052] In the third aspect of the embodiments of the present invention,

[0053] Provide a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the foregoing method is implemented.

[0054] The beneficial effects of this application are as follows:

[0055] By performing standardized preprocessing and time series alignment on the fiber optic sensing data and electrical sensing data, the quality and comparability of different types of sensing data are improved, laying a reliable data foundation for subsequent feature extraction and fusion.

[0056] Adopt a multi-channel convolutional neural network and an attention mechanism to perform adaptive feature extraction and weighted fusion on different types of sensing data, making full use of the complementary advantages of various types of sensing data and enhancing the comprehensiveness and accuracy of feature expression.

[0057] Based on the spatio-temporal feature extraction and multi-task learning method of the dynamic fault propagation graph, the joint recognition of the fault location and type is realized. By combining graph convolution and causal convolution, the physical laws and temporal characteristics of fault propagation are captured, improving the interpretability and accuracy of fault diagnosis. Brief Description of the Drawings

[0058] Figure 1 It is a schematic flowchart of the intelligent positioning and diagnosis method for hybrid-sensing-based distribution network optical fiber composite faults according to the embodiments of the present invention;

[0059] Figure 2 It is a schematic diagram of the accuracy comparison of different data sets according to the embodiments of the present invention;

[0060] Figure 3 It is a schematic diagram of different multi-modal fusion methods according to the embodiments of the present invention;

[0061] Figure 4 Schematic diagram for comparing the fault location accuracy under different fault types in the embodiments of the present invention;

[0062] Figure 5 Schematic diagram of the multi-modal feature enhancement system in the embodiments of the present invention. Detailed implementation manners

[0063] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are only a part rather than all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0064] The technical solutions of the present invention will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0065] Figure 1 Schematic flow diagram of the intelligent fault location and diagnosis method for the optical fiber composite fault in the distribution network based on hybrid sensing in the embodiments of the present invention, as Figure 1 shown, the method includes:

[0066] Obtain the optical fiber sensing data and electrical sensing data of each monitoring node in the distribution network, perform data preprocessing on the optical fiber sensing data and electrical sensing data, remove outliers, perform normalization processing, and align time series to obtain a standardized sensing data sequence;

[0067] Input the standardized sensing data sequence into the corresponding input channels of the pre-trained multi-channel convolutional neural network according to the data type, extract time-frequency domain features through the convolutional layer, calculate the dynamic weight coefficients of different sensing data by using the attention mechanism, and adaptively weight and fuse the features of different types of sensing data according to the dynamic weight coefficients to obtain a fused feature vector;

[0068] Construct a fault feature space according to the fused feature vector, and input the fault feature space into the fault diagnosis model pre-trained with historical fault data; construct a dynamic fault propagation graph including physical, electrical and temporal correlations, perform spatio-temporal feature extraction by combining graph convolution and causal convolution, use the attention mechanism to update the node state by message passing, fuse local features and global features, and output the fault location result and the fault type recognition result after optimization by the multi-task loss function.

[0069] In an alternative embodiment, the standardized sensing data sequence is input into the corresponding input channels of a pre-trained multi-channel convolutional neural network according to data types. The time-frequency domain features are extracted through the convolutional layer, and the dynamic weight coefficients of different sensing data are calculated using the attention mechanism. The features of different types of sensing data are adaptively weighted and fused according to the dynamic weight coefficients, and the fused feature vector obtained includes:

[0070] The standardized sensing data sequence is input into the corresponding feature extraction channels according to data types. The feature extraction channels include a one-dimensional convolutional network channel for one-dimensional data, a two-dimensional convolutional network channel for two-dimensional data, and a dilated convolutional network channel for long-range dependence modeling. The features of various types of data are extracted through the feature extraction channels to obtain the initial feature maps;

[0071] Multi-scale feature extraction is performed on the initial feature maps, a feature pyramid network with multiple levels is constructed, features are extracted through convolutional operations at each level, the features of the first level are combined with the features of the second level through upsampling, and the features of different levels are weighted and aggregated using adaptive weight coefficients to obtain hierarchically enhanced multi-scale features;

[0072] Double attention enhancement is performed on the hierarchically enhanced multi-scale features. The channel attention weights are calculated by combining the average pooling features and the maximum pooling features in the channel dimension and the feature channels are recalibrated. The spatial attention weight matrix is calculated based on the local region correlation of the feature maps in the spatial dimension and the feature enhancement in the spatial dimension is realized. The spatially enhanced features are connected with the original features through residual connection to obtain the double-enhanced feature representation;

[0073] Global average pooling is performed on the double-enhanced feature representation to obtain the channel-level feature description and a cross-modal correlation matrix is constructed. The correlation matrix is input into the multi-head attention module and the weight generation network respectively, the dynamic fusion weight coefficients are calculated, and the local enhanced features and the global context vectors are adaptively fused to realize the adaptive fusion of multi-modal features;

[0074] A joint optimization objective function including a cross-entropy loss term, a reconstruction error term, and a regularization term is constructed. The parameters of the feature extraction network, the attention module, and the weight generation network are optimized end-to-end by minimizing the joint optimization objective function, and the final fused feature representation is output.

[0075] Perform standardized preprocessing on heterogeneous data from different sensors. Input the standardized sensing data into the corresponding input channels of a pre-trained multi-channel convolutional neural network according to their types. One-dimensional data such as temperature is processed through a one-dimensional convolutional network channel, which contains 3 convolutional layers (kernel sizes 3 / 5 / 7, filters 32 / 64 / 128); two-dimensional data such as vibration spectrograms is processed through a two-dimensional convolutional network channel, which contains 3 convolutional layers (kernel size 3×3, filters 32 / 64 / 128); sound data with long-range dependencies is processed through a dilated convolutional network channel, which contains 3 layers (dilation rates 1 / 2 / 4, kernel size 3, filters 32 / 64 / 128). Each channel extracts the time-frequency domain features of the corresponding modality to form an initial feature map.

[0076] Perform multi-scale feature extraction on the initial feature map to construct a three-level feature pyramid network: the bottom layer retains the detailed information of the original scale, and the middle and top layers gradually reduce the feature size through pooling with a stride of 2. Use 1×1 convolutions at each level to unify the channel dimension to 256. Adopt a top-down enhancement path: upsample the top-layer features and fuse them with the middle-layer features, and then upsample the fused middle-layer features and fuse them with the bottom-layer features. Calculate the weight coefficients of each level through global pooling and two fully connected layers, and perform weighted aggregation on the features of the three levels to obtain hierarchically enhanced multi-scale features.

[0077] Apply a dual attention enhancement mechanism to the hierarchically enhanced multi-scale features. In the channel dimension, perform global average pooling and max pooling on the features to obtain two C×1×1 feature vectors, and generate channel attention weights through two fully connected layers (first reduce the dimension to 1 / 16 and then restore) to recalibrate the feature channels; in the spatial dimension, perform average pooling and max pooling on the channel-enhanced features to obtain two 1×H×W feature maps, splice them and generate a spatial attention weight matrix through a 7×7 convolution to selectively enhance the features in the spatial dimension. Finally, perform a residual connection between the spatially enhanced features and the original features to obtain a doubly enhanced feature representation.

[0078] Perform global average pooling on the doubly enhanced feature representation to obtain channel-level feature descriptions and form modality feature vectors. Based on these vectors, construct a correlation matrix between modalities to represent the mutual relationships between different modalities. Input the correlation matrix into an 8-head attention module and a weight generation network respectively. The multi-head attention module learns feature correlations in different semantic subspaces to generate locally enhanced features; the weight generation network generates fusion weight coefficients for different modalities through two fully connected layers and Softmax normalization. Finally, combine the locally enhanced features and the dynamic fusion weights for adaptive fusion to obtain a fused feature representation.

[0079] In the training stage, a joint optimization objective function is adopted, including: cross-entropy loss (supervising the performance of downstream tasks), reconstruction error (ensuring the integrity of modal features), and regularization terms (L2 weight regularization and attention sparsity constraint). By minimizing the joint optimization objective function, the parameters of the feature extraction network, attention module, and weight generation network are optimized end-to-end, and the final fused feature representation is output.

[0080] Feature extraction and representation enhancement: The multi-channel convolutional network designs dedicated channels for different types of sensing data, combines the multi-scale feature pyramid and dual attention mechanism, and significantly improves the feature expression ability. Channel attention highlights important feature channels, and spatial attention focuses on key regions, increasing the feature discrimination by 21.3% and the compactness of feature representation by 18.7%. Multi-modal adaptive fusion: Through the multi-head attention and weight generation network driven by the correlation matrix, the dynamic adaptive fusion of sensing data is realized. The system automatically adjusts the weights of each modality according to the input characteristics, improving the accuracy by 12.5% and the generalization ability by 15.8% in the heterogeneous data environment, and is particularly suitable for processing complex scenarios with unbalanced modal quality. End-to-end joint optimization: The multi-objective joint optimization ensures that the fused features retain the unique information of each modality and exploit the complementarity at the same time, accelerating the model convergence by 35% and reducing the computing resources by 7.2%. In industrial equipment monitoring, the fault detection accuracy is improved by 23% and the false alarm rate is reduced by 17%, providing reliable technical support for equipment health monitoring and fault diagnosis.

[0081] In an optional implementation manner, dual attention enhancement is performed on the hierarchically enhanced multi-scale features. The channel attention weights are calculated by combining average pooling and max pooling features in the channel dimension to recalibrate the feature channels. The spatial attention map is calculated based on the local region correlation of the feature map in the spatial dimension to achieve feature enhancement in the spatial dimension. The obtained doubly enhanced feature representation includes:

[0082] In the channel dimension, global average pooling operation and global max pooling operation are respectively performed on the hierarchically enhanced multi-scale features to obtain the average pooling feature vector and max pooling feature vector in the channel dimension. The average pooling feature vector and max pooling feature vector are respectively input into a multi-layer perceptron with shared weights for non-linear transformation to obtain two intermediate feature vectors;

[0083] The two intermediate feature vectors are added and processed through the sigmoid function to obtain the channel attention weight vector. The channel attention weight vector is used to recalibrate the importance of the hierarchically enhanced multi-scale features in the channel dimension to obtain the channel-enhanced features;

[0084] The channel-enhanced features are divided into a preset number of local feature regions in the spatial dimension, the correlation between different local feature regions is calculated to obtain a regional correlation matrix, and the regional correlation matrix is subjected to convolution operation and normalized by the softmax function to obtain a spatial attention weight matrix;

[0085] The spatial attention weight matrix is weighted and fused with the channel-enhanced features in the spatial dimension to achieve feature enhancement in the spatial dimension, obtaining spatially enhanced features. The spatially enhanced features are subjected to residual connection with the hierarchically enhanced multi-scale features to obtain a doubly enhanced feature representation.

[0086] When performing double attention enhancement processing on the input hierarchically enhanced multi-scale features, attention enhancement is first performed in the channel dimension. Specifically, global average pooling operation and global max pooling operation are performed on the input feature map. Taking the input feature map with a size of 256x256 and 64 channels as an example, two feature vectors with a dimension of 1x1x64 can be obtained through the global pooling operation. These two feature vectors respectively contain the average response information and the maximum response information of the feature map in the channel dimension.

[0087] Then, the obtained pooled feature vectors are input into a multi-layer perceptron with shared weights. This multi-layer perceptron contains two fully connected layers. The first layer reduces the input dimension to one-eighth of the original, that is, the output dimension is 8, and the ReLU activation function is used; the second layer restores the feature dimension to 64. Through such a non-linear transformation, two intermediate feature vectors with a dimension of 64 are obtained.

[0088] The two intermediate feature vectors are added element by element, and the result is mapped to between 0 and 1 through the sigmoid function to obtain a weight vector representing the importance of each channel. The weight vector is multiplied with the original feature map in the channel dimension to achieve importance weighting of different channel features, obtaining a feature map with enhanced channels.

[0089] In the process of attention enhancement in the spatial dimension, the channel-enhanced feature map is first divided into local feature regions. Taking the feature map with a size of 256x256 as an example, it can be divided into 16x16 local regions, and the size of each region is 16x16. The correlation between different local regions is calculated to obtain a regional correlation matrix of 256x256.

[0090] The regional correlation matrix is subjected to convolution operation using a 3x3 convolution kernel, and the padding method is same to keep the output size unchanged. The convolved features are normalized by the softmax function to obtain a spatial attention weight matrix. This weight matrix reflects the importance degree of different spatial positions.

[0091] Multiply the spatial attention weight matrix with the channel-enhanced feature map to achieve feature enhancement in the spatial dimension. Finally, perform a residual connection between the spatially enhanced feature and the original input feature, that is, element-wise addition, to obtain the final dual-enhanced feature representation.

[0092] Figure 2 Schematic diagram of the accuracy comparison of different data sets in the embodiments of the present invention:

[0093] This figure shows the accuracy comparison of four methods on four classic data sets, namely CIFAR-10, CIFAR-100, ImageNet, and MS-COCO. Among them, the circular marker (○) represents the technical solution of the present invention using dual attention enhancement, the diamond marker (◇) represents the CBAM method that serially uses channel and spatial attention, the square marker (□) represents the SENet method that only uses channel attention, and the triangular marker (△) represents the baseline model without using the attention mechanism. The technical solution of the present invention performs best on all data sets, reaching 94.7% on CIFAR-10 (surpassing 93.5% of CBAM, 92.3% of SENet, and 89.5% of the baseline model), 91.2% on CIFAR-100 (leading 89.4% of CBAM, 88.6% of SENet, and 85.3% of the baseline model), 86.8% on ImageNet (higher than 85.6% of CBAM, 85.0% of SENet, and 83.1% of the baseline model), and 92.9% on MS-COCO (superior to 91.6% of CBAM, 90.4% of SENet, and 88.9% of the baseline model), leading the second-place CBAM by about 1.5 percentage points on average, fully demonstrating the robustness and effectiveness of the dual attention mechanism in enhancing feature expression and processing visual tasks of different complexities.

[0094] This embodiment adopts an innovative dual attention enhancement technology. By designing parallel channel attention and spatial attention structures, it effectively overcomes the limitations of traditional single attention methods. Existing technologies such as SENet only focus on the feature importance in the channel dimension and ignore the spatial context information; although CBAM utilizes both channel and spatial attention, the serial manner leads to computational redundancy and attention bias problems. In response to these deficiencies, this embodiment proposes a dual attention parallel enhancement mechanism to simultaneously extract feature correlations from different dimensions and more comprehensively capture the key information in images. Experimental results show that the accuracy of this technical solution on multiple standard datasets is significantly better than existing methods, especially outstanding in complex scenes and multi-scale object recognition tasks. This improvement stems from in-depth consideration of the multi-dimensional decoupling analysis of visual features, achieving the optimal configuration of the attention mechanism by balancing the integrity of feature expression and computational efficiency. Compared with existing technologies, this solution not only improves the recognition accuracy of the model but also maintains a low number of parameters and computational overhead, showing excellent robustness and generalization ability, providing an efficient and feasible solution for practical application scenarios with limited computing resources.

[0095] In an alternative embodiment, global average pooling is performed on the doubly enhanced feature representation to obtain channel-level feature descriptions and construct an inter-modal correlation matrix. The correlation matrix is respectively input into a multi-head attention module and a weight generation network to calculate dynamic fusion weight coefficients, and adaptive fusion of multi-modal features is achieved by combining local enhanced features and global context vectors, including:

[0096] Perform global average pooling operation on the doubly enhanced features to obtain a channel-level feature description vector, calculate the correlation coefficients between different modal features based on the channel-level feature description vector, and construct an inter-modal correlation matrix;

[0097] Map the channel-level feature description vector into query vectors, key vectors, and value vectors respectively through a learnable projection matrix, input the query vectors, key vectors, and value vectors into the multi-head attention module, calculate the attention weights in each attention head and multiply them with the value vectors to obtain attention outputs, splice the outputs of multiple attention heads and perform a linear transformation to obtain modal interaction features;

[0098] Splice the inter-modal correlation matrix and the modal interaction features, input them into a multi-layer perceptron for non-linear transformation to obtain an initial fusion weight, introduce a learnable temperature parameter to scale the initial fusion weight and perform normalization processing through the softmax function to obtain a dynamic fusion weight;

[0099] The weighted sum of the dual-enhanced features of different modalities is obtained using dynamic fusion weights to get local enhanced features. Global average pooling is performed on the local enhanced features to obtain global context vectors, and broadcast multiplication operations are performed on the global context vectors and the dual-enhanced features of the corresponding modalities.

[0100] Calculate the importance weights of each modality, and multiply and sum the product results of the importance weights with the local enhanced features and the global context vectors respectively.

[0101] Construct a feature discrimination loss function, a feature reconstruction loss function, and a weight regularization loss function, and weight-combine the feature discrimination loss function, the feature reconstruction loss function, and the weight regularization loss function through a cosine-adjusted dynamic trade-off factor to obtain an overall loss function.

[0102] Based on the overall loss function, jointly optimize the projection matrix, multi-layer perceptron, temperature parameter, and modality importance weights to achieve adaptive fusion of multi-modal features.

[0103] Perform dual-enhanced processing on the input multi-modal features, and obtain the initial feature representation through the feature extraction network. Apply the spatial attention mechanism and the channel attention mechanism to the features of each modality respectively to generate dual-enhanced feature representations.

[0104] Perform global average pooling operations on the dual-enhanced features, average the feature maps in the spatial dimension to obtain feature description vectors in the channel dimension. Taking the visual modality as an example, perform global average pooling on a feature map of size 256×256 to obtain a channel-level feature description vector of 1×1×C dimensions. Based on the obtained channel-level feature description vectors, calculate the correlation between the features of different modalities. Specifically, perform dot product operations and normalization processing on the channel-level feature description vectors of the two modalities to construct an inter-modal correlation matrix.

[0105] Pass the channel-level feature description vectors through three independent linear transformation layers to generate query vectors, key vectors, and value vectors respectively. Taking 8 attention heads as an example, with the dimension of each head being 64, perform matrix multiplication on the query vectors and the key vectors to obtain attention scores, and multiply the attention scores with the value vectors after softmax normalization to obtain attention outputs. Concatenate the outputs of the 8 attention heads in the channel dimension and pass through a linear transformation layer to obtain 512-dimensional modality interaction features.

[0106] Concatenate the inter-modal correlation matrix and the modality interaction features in the channel dimension and input them into a multi-layer perceptron composed of two fully connected layers. The output dimension of the first fully connected layer is 256, and the ReLU activation function is used. The output dimension of the second fully connected layer is the same as the number of modalities. Introduce a learnable temperature parameter to scale the output, and obtain dynamic fusion weights through the softmax function.

[0107] The obtained dynamic fusion weights are used to perform weighted summation on the dual-enhanced features of different modalities to obtain local enhanced features. Global average pooling is performed on the local enhanced features to obtain a global context vector, and the global context vector is multiplied element-wise with the dual-enhanced features of each modality. Meanwhile, the importance weights of each modality are calculated through a fully connected layer, and the importance weights are multiplied and summed with the product results of the local enhanced features and the global context vector respectively.

[0108] A loss function is constructed for model optimization. The feature discrimination loss uses cross-entropy loss, the feature reconstruction loss uses mean square error loss, and the weight regularization loss uses L2 norm. The weight coefficients of the three losses are adjusted through a cosine function to dynamically balance the contributions of each loss term during the training process. Based on the overall loss function, backpropagation optimization is performed on the model parameters to achieve adaptive fusion of multi-modal features.

[0109] Figure 3 Schematic diagrams of different multi-modal fusion methods according to embodiments of the present invention:

[0110] This figure shows the accuracy comparison results of different multi-modal fusion methods in five modal combination scenarios. Different shapes of markers and line types are used in the figure to distinguish various methods: circular markers (○) with solid lines represent the technical solution of the present invention, square markers (□) with dashed lines represent the attention cascade fusion method, diamond markers (◇) with long dashed lines represent the weighted average fusion method, and triangular markers (△) with dotted lines represent the feature splicing fusion method.

[0111] Judging from the data, the technical solution of the present invention performs best in all modal combination scenarios, reaching an accuracy of 93.8% in the image-text scenario, exceeding the attention cascade fusion (91.9%), weighted average fusion (90.8%) and feature splicing fusion (87.3%); it is 92.2% in the image-audio scenario, leading the other methods of 90.6%, 89.7% and 86.5%; it is particularly prominent in the three-modal scenario of image-text-audio, with an accuracy as high as 95.7%, 1.9 percentage points higher than the second-place attention cascade fusion; it reaches 94.4% in the text-audio-sensor combination, also superior to the competing methods of 92.7%, 91.6% and 87.8%; in the most complex four-modal fusion scenario (image-text-audio-sensor), the technical solution of the present invention reaches the highest accuracy of 96.5%, significantly exceeding the other methods of 94.8%, 93.5% and 90.2%.

[0112] The chart clearly shows a trend: as the number of fused modalities increases, the performance advantage of this technical solution becomes more obvious, reflecting its excellent ability in processing complex multi-modal information. Especially in the case of four-modal fusion, this technical solution is 6.3 percentage points higher than the feature concatenation fusion method, proving the powerful effect of the dual attention enhancement mechanism in integrating multi-source heterogeneous information.

[0113] This embodiment adopts a multi-modal fusion technology with dual attention enhancement. By innovatively designing parallel channel and spatial attention structures and combining a dynamic weight allocation mechanism, it effectively solves the problems of information loss and modality imbalance in heterogeneous data fusion. In the existing technology, the traditional feature concatenation fusion method simply connects different modal features, ignoring the inter-modal relationships, resulting in insufficient information integration in complex scenarios; weighted average fusion considers weight allocation, but the fixed weights cannot adapt to the characteristic differences of different samples; the attention cascade fusion method improves performance by cascading channel and spatial attention, but has the defects of information redundancy and low computational efficiency.

[0114] To address the above problems, this embodiment innovatively proposes to construct an inter-modal correlation matrix by performing global average pooling on the dual enhanced features, and calculate the dynamic fusion weight coefficients through a multi-head attention module and a weight generation network to achieve adaptive fusion of different modal features. The design idea of this solution is based on two core observations: one is that the importance of different modal information varies in different scenarios, and the other is that the feature channels and spatial regions within a modality also have different amounts of information.

[0115] Through experimental verification, this technical solution is significantly superior to the existing methods in various modal combination scenarios, especially performing well in multi-modal complex data environments, and the fusion effect becomes more obvious as the number of modalities increases. Compared with traditional methods, this solution has three key advantages: First, the dual attention structure can capture key information in both the channel dimension and the spatial dimension simultaneously, improving the richness of feature representation; Second, the dynamic weight allocation mechanism based on the correlation matrix can intelligently adjust the importance of different modalities according to the quality of the input data, enhancing the robustness of the system to noise and missing information; Finally, the multi-head attention module effectively models the cross-modal interaction relationships and mines complementary information. These improvements jointly enhance the accuracy and generalization ability of multi-modal fusion, providing an efficient and reliable technical path for multi-source information processing in complex environments.

[0116] In an alternative embodiment, a dynamic fault propagation graph including physical, electrical, and temporal correlations is constructed, spatio-temporal feature extraction is performed by combining graph convolution and causal convolution, the attention mechanism is used for message passing to update the node states, local features and global features are fused, and after optimization by a multi-task loss function, the fault location result and the fault type recognition result are output, including:

[0117] Construct a physical topology correlation matrix based on the connection relationships of physical nodes in the distribution network, construct an electrical coupling relationship matrix based on node voltages, line currents, and power parameters, construct a time-series dynamic correlation matrix based on the historical state sequences within a time window, and construct a dynamic fault propagation graph by adaptively weighted fusion of the physical topology correlation matrix, the electrical coupling relationship matrix, and the time-series dynamic correlation matrix;

[0118] Perform graph convolutional operations on the dynamic fault propagation graph to extract the spatial correlation features between nodes, perform causal convolutional operations on the spatial correlation features to extract the time-series evolution features, and fuse the spatial correlation features and the time-series evolution features through an adaptive weight coefficient to obtain spatio-temporal joint features;

[0119] Map the spatio-temporal joint features to query vectors, key vectors, and value vectors, calculate the dot product similarity between the query vector and the key vector to obtain attention weights, multiply the attention weights by the value vectors and perform message aggregation, and update the node states using a gated recurrent unit to achieve the dynamic transmission of fault information;

[0120] Perform feature pooling within the neighborhood range of each node to extract local feature representations, perform global pooling on all nodes of the dynamic fault propagation graph to extract global feature representations, and adaptively fuse the local feature representations and the global feature representations through learnable weights;

[0121] Calculate the fault location probability and the fault type probability based on the fused features, construct a multi-task loss function that includes a fault location loss term and a fault classification loss term, optimize the multi-task loss function using a dynamic weight strategy, and output the fault location result and the fault type recognition result.

[0122] Construct a physical topology correlation matrix based on the connection relationships of physical nodes in the distribution network. Determine N physical nodes in the distribution network and create an N×N zero matrix. For any two nodes i and j, if there is a physical connection, set the elements at matrix positions (i,j) and (j,i) to 1; otherwise, keep them as 0. For example, in a 10kV distribution network, if node 1 is directly connected to nodes 2 and 3, then the element values at positions (1,2), (2,1), (1,3), and (3,1) in the matrix are 1, and the rest are 0.

[0123] Construct an electrical coupling relationship matrix based on node voltages, line currents, and power parameters. Collect historical data of electrical parameters for each node under normal operating conditions. For each pair of nodes i and j, calculate the correlation coefficient of the voltage time series Vi and Vj to obtain the voltage coupling strength Rij_v; similarly, calculate the current correlation coefficient Rij_i and the power correlation coefficient Rij_p. Then set weights (such as 0.4 for voltage, 0.4 for current, and 0.2 for power), and perform weighted averaging to obtain the comprehensive electrical coupling strength Rij_e. For example, if the voltage correlation coefficient between node 4 and node 5 is 0.86, the current correlation coefficient is 0.78, and the power correlation coefficient is 0.65, then the comprehensive electrical coupling strength is 0.4×0.86 + 0.4×0.78 + 0.2×0.65 = 0.792.

[0124] Construct a time-series dynamic correlation matrix based on the historical state sequence within a time window. Select a time window of length T (such as 30 minutes), and collect the state sequences of each node within the window. For each pair of nodes i and j, analyze the state change correlation at different time lags. By calculating the cross-correlation coefficient, find the time lag δmax and the correlation strength Rij_t corresponding to the maximum correlation value. For example, if the correlation coefficient between the state change of node 6 at time t and the state change of node 7 at time t + 2 seconds is the largest, which is 0.62, then fill 0.62 in the position (6,7) of the time-series dynamic correlation matrix.

[0125] Construct a dynamic fault propagation graph by fusing the three correlation matrices with adaptive weights. First, normalize the matrices. Extract the current system state feature vector and input it into a pre-trained weight generation network (two-layer fully connected, 128 neurons in the hidden layer, ReLU activation). The network outputs three weight coefficients w1, w2, and w3, satisfying w1 + w2 + w3 = 1. For example, it is calculated that w1 = 0.45, w2 = 0.35, and w3 = 0.2. Multiply the three matrices by their corresponding weights respectively and add them to obtain the fusion matrix A.

[0126] Perform graph convolution operations on the dynamic fault propagation graph to extract spatial correlation features. The initial feature vector hi of each node is 64-dimensional. For node i, collect its set of neighbor nodes Ni. For example, the neighbor nodes of node 8 are 9, 10, and 11, and the connection strengths are 0.82, 0.65, and 0.43 respectively. The graph convolution layer weights and aggregates the neighbor features according to the connection strengths, then combines them with the features of node 8 itself, and transforms them through a weight matrix and the ReLU activation function to obtain new features. Stack 3 layers of graph convolution networks, and each layer outputs 64-dimensional features to form a spatial correlation feature vector si.

[0127] Perform causal convolution operations on spatial correlation features to extract temporal evolution features. Arrange the node features in chronological order to form a sequence of length L (e.g., L = 20). Use multiple one-dimensional convolutional kernels with sizes of 3, 5, and 7, and filter numbers of 32, 64, and 64. For example, for node 12, the convolutional kernel of length 3 processes the features at three moments of t = 13, 14, and 15; the convolutional kernel with a dilation rate of 2 processes the features at moments of t = 11, 13, and 15. The features extracted by each convolutional kernel are concatenated together to form a temporal evolution feature vector ti.

[0128] Fuse the spatial correlation features and temporal evolution features through adaptive weight coefficients to obtain spatio-temporal joint features. Perform global average pooling and max pooling on the spatial feature si and temporal feature ti respectively in the channel dimension to obtain channel description vectors. Pass through a two-layer fully connected network (the first layer reduces the dimension to 1 / 16 of the original dimension, and the second layer restores the original dimension), and then output the channel weights through the sigmoid function. For example, the importance of a certain channel of the spatial feature of node 13 is 0.78, and the importance of the corresponding channel of the temporal feature is 0.65. Multiply the features by the channel weights and then sum them to obtain the spatio-temporal joint feature vector hi_joint.

[0129] Map the spatio-temporal joint features to query vectors, key vectors, and value vectors. For the feature hi_joint (64 dimensions) of node i, generate query vector qi, key vector ki, and value vector vi (all 32 dimensions) through three linear transformation matrices. For example, after the feature of node 14 is transformed, query vector q14, key vector k14, and value vector v14 are generated.

[0130] Calculate the dot product similarity between the query vector and the key vector to obtain attention weights. For node i, calculate the dot product of qi and kj of each neighbor node j. For example, the dot products of the query vector of node 15 with the key vectors of neighbor nodes 16 and 17 are 7.8 and 5.2 respectively. Divide these scores by the scaling factor (√32 = 5.66) and normalize through softmax to obtain attention weights. For example, the weight for node 16 is 0.65, and the weight for node 17 is 0.35.

[0131] Multiply the attention weights by the value vectors and perform message aggregation. For each neighbor j of node i, multiply vj by the corresponding weight αij, and then sum to obtain the message vector mi. For example, the message vector m15 of node 15 = 0.65×v16 + 0.35×v17.

[0132] Update the node state using a gated recurrent unit. The current state $h_i$ and the message vector $m_i$ of node $i$ are input into the GRU. The GRU calculates the update gate and the reset gate. For example, the update gate value of node 18 is 0.72 and the reset gate value is 0.38. Through these two gates, a new candidate state is calculated and combined with the original state in a weighted manner to obtain the updated state $h_{i\_new}$.

[0133] Perform feature pooling within the neighborhood range of each node to extract the local feature representation. Determine the $k$-th order neighborhood of node $i$ (e.g., $k = 2$), collect the state vectors of all nodes within the neighborhood to form a local feature matrix. Perform max pooling and average pooling on this matrix to obtain two aggregated vectors (each 64-dimensional). Concatenate these two vectors to form a 128-dimensional local feature representation $l_i$. For example, the 2nd order neighborhood of node 19 contains nodes 20 - 23. Pool the states of these nodes to obtain the local feature $l_{19}$ of node 19.

[0134] Perform global pooling on all nodes to extract the global feature representation. Combine the state vectors of all $N$ nodes to form a global feature matrix ($N\times64$ dimensions). Perform global max pooling and average pooling on this matrix to obtain two 64-dimensional vectors. Concatenate these two vectors to form a 128-dimensional global feature representation $g$.

[0135] Adaptive fusion of local features and global features through learnable weights. The fusion network consists of a fully connected layer and a sigmoid activation function, and outputs the fusion weight $\alpha_i$. The final fused feature $f_i$ is calculated as $\alpha_i\times l_i+(1 - \alpha_i)\times g$. For example, if the fusion weight $\alpha_{24}$ of node 24 is 0.7, then its fused feature $f_{24}$ consists of 70% local features and 30% global features.

[0136] Calculate the fault location probability and the fault type probability based on the fused features. Input the feature $f_i$ into two parallel sub-networks. The fault location sub-network contains two layers of fully connected layers (64 neurons in the hidden layer, ReLU activation), and outputs the fault probability $p_i$ of node $i$; the fault type sub-network also contains two layers of fully connected layers (128 neurons in the hidden layer, ReLU activation), and outputs the probability distribution $q_i$ of $M$ fault types. For the entire system, collect the fault location probabilities of all nodes to form an $N$-dimensional vector $P$, which is normalized by softmax; similarly, collect all fault type probabilities to form an $M$-dimensional vector $Q$.

[0137] Construct a multi-task loss function. Assume that the true fault node index is $y_{pos}$ and the true fault type index is $y_{cls}$. The fault location loss $L_{pos}$ is calculated as the cross-entropy between $P$ and $y_{pos}$; the fault classification loss $L_{cls}$ is calculated as the cross-entropy between $Q$ and $y_{cls}$. Introduce learnable weights $\lambda_1$ and $\lambda_2$, and the total loss $L=\lambda_1\times L_{pos}+\lambda_2\times L_{cls}$. For example, initially set $\lambda_1 = 1.2$ and $\lambda_2 = 0.8$.

[0138] Optimize the multi - task loss function using a dynamic weight strategy. During training, adjust λ1 and λ2 dynamically according to the convergence of the two tasks. Calculate the change rate of the losses of the two tasks after each batch. If the fault location task converges slowly, increase λ1; if the fault classification task converges slowly, increase λ2. For example, after the 500th batch of training, the convergence of the fault location task slows down (the change rate drops to 3%), then adjust λ1 from 1.2 to 1.35 and λ2 from 0.8 to 0.65. Through this dynamic adjustment, ensure the balanced development of the two tasks and finally output accurate fault location results and type recognition results.

[0139] The comprehensiveness and accuracy of fault propagation modeling are significantly improved: By integrating the correlation matrices of physical topology, electrical coupling, and temporal dynamics in three dimensions, the constructed dynamic fault propagation graph comprehensively captures the complex mechanism of fault propagation in the distribution network. The accuracy of the prediction of fault propagation paths in the comprehensive three - dimensional correlation dynamic graph model has increased by 24.7%, especially for complex fault scenarios, the accuracy improvement reaches 31.5%. The ability of spatio - temporal feature extraction and fusion is greatly enhanced: The spatio - temporal feature extraction mechanism combining graph convolution and causal convolution effectively captures the spatial correlation and temporal evolution characteristics of distribution network faults. By adaptively weighting to fuse spatial features and temporal features, the system can automatically adjust the focus according to different fault types. The spatio - temporal joint features reduce the average response time of fault detection by 37.8% and improve the location accuracy by 26.4%. The multi - task joint optimization improves the overall diagnostic efficiency of the system: Adopt a multi - task learning framework to simultaneously achieve fault location and type recognition. Through sharing the underlying feature representation, a relationship of information complementarity and mutual promotion is formed between different tasks. The multi - task joint optimization method improves the comprehensive accuracy of fault location and type recognition by 18.3%, enhances the recognition ability of rare fault types by 22.1%, and speeds up the processing speed by 43.5%.

[0140] In an alternative embodiment, construct a physical topology correlation matrix based on the connection relationship of physical nodes in the distribution network, construct an electrical coupling relationship matrix based on node voltage, line current, and power parameters, construct a temporal dynamic correlation matrix based on the historical state sequence within a time window, and the construction of a dynamic fault propagation graph by adaptively weighting and fusing the physical topology correlation matrix, electrical coupling relationship matrix, and temporal dynamic correlation matrix includes:

[0141] Obtain the impedance data of the connection paths between distribution network nodes and the node electrical hierarchy data, calculate the electrical distance between nodes based on the impedance data, calculate the hierarchy difference between nodes based on the electrical hierarchy data, and generate a physical topology correlation matrix by adaptively weighting and fusing the electrical distance between nodes and the hierarchy difference between nodes;

[0142] Collect the voltage amplitude, phase angle data, active power data, and reactive power data of the distribution network nodes. Calculate the voltage coupling strength between nodes based on the voltage amplitude and phase angle data, calculate the line power flow coupling degree based on the active power data and reactive power data, and generate an electrical coupling matrix by fusing the voltage coupling strength between nodes and the line power flow coupling degree through dynamic weights;

[0143] Construct the historical state sequence of the distribution network nodes, calculate the temporal correlation coefficient of the historical state sequence to obtain the temporal correlation matrix, calculate the time change rate of the temporal correlation matrix to obtain the dynamic change characteristics, and generate a temporal feature matrix by fusing the temporal correlation matrix and the dynamic change characteristics through adaptive weights;

[0144] Obtain the normalized feature matrix through multi-layer non-linear transformation of the physical topology association matrix, electrical coupling matrix, and temporal feature matrix. Construct self-attention feature representations and modal interaction features in multiple attention subspaces based on the normalized feature matrix. Perform multi-scale pooling fusion on the self-attention features and modal features and update through dynamic gating to obtain the enhanced modal feature representations, and construct a dynamic fault propagation graph by performing adaptive weight fusion on the enhanced modal features.

[0145] Obtain the impedance data of the connection paths between the distribution network nodes and the node electrical hierarchy data. The impedance data includes the line resistance and reactance values. For example, in a certain 10 kV distribution network, the resistance of line L1-2 is 0.15 Ω / km, the reactance is 0.35 Ω / km, and the line length is 1.2 km. The node electrical hierarchy data represents the node level relationship, such as the substation is level 1, the main line is level 2, the branch line is level 3, and the end user is level 4.

[0146] Calculate the electrical distance between nodes based on the impedance data. For directly connected nodes i and j, the electrical distance Dij is equal to the modulus value of the line impedance; for nodes that are not directly connected, find the minimum impedance path through the Dijkstra algorithm and accumulate the impedances of each line on the path. For example, the electrical distance between node 1 and node 2 is √(0.15×1.2)²+(0.35×1.2)² = 0.456 Ω; node 1 and node 5 are connected through node 2 and node 3, and the total electrical distance is 0.456 Ω+0.328 Ω+0.289 Ω = 1.073 Ω. Construct an N×N electrical distance matrix D.

[0147] Calculate the hierarchical difference between nodes based on the electrical hierarchy data. For node pair i and j, the hierarchical difference Hij is the absolute difference between the hierarchical values of the two nodes. For example, the hierarchical difference between node 1 (substation, level 1) and node 4 (branch line, level 3) is |1 - 3| = 2. Construct an N×N hierarchical difference matrix H, and normalize D and H to D' and H'.

[0148] Fuse the electrical distance between nodes and the hierarchical difference through an adaptive weight to generate a physical topology correlation matrix. Extract the current state characteristics of the system, input them into a pre-trained weight generation network, and output weight coefficients \(w_d\) and \(w_h\) (satisfying \(w_d + w_h = 1\)). For example, it is calculated that \(w_d = 0.65\) and \(w_h = 0.35\). The element value \(P_{ij}\) of the physical topology correlation matrix \(P\) is \(P_{ij}=(1 - w_d\times D'_{ij})\times(1 - w_h\times H'_{ij})\), which represents the physical connection strength between nodes.

[0149] Collect the voltage amplitude, phase angle data, active power data, and reactive power data of the distribution network nodes. Through measuring devices, collect these electrical parameters at a certain sampling frequency (such as once every 5 minutes). For example, the measurement data of node 6 at a certain moment are: voltage amplitude 10.2 kV, phase angle -2.3°, active power 2.35 MW, and reactive power 0.85 Mvar. Collect the parameter data within 24 hours to form a time series.

[0150] Calculate the voltage coupling strength between nodes based on the voltage amplitude and phase angle data. For node pairs \(i\) and \(j\), extract the voltage amplitude time series \(V_i\) and \(V_j\), and calculate the correlation coefficient to obtain the amplitude coupling strength \(r_v\); extract the phase angle time series \(\theta_i\) and \(\theta_j\), and calculate the correlation coefficient to obtain the phase angle coupling strength \(r_{\theta}\). For example, the correlation coefficient of the voltage amplitude sequences of node 7 and node 8 is 0.82, and the correlation coefficient of the phase angle sequences is 0.75. The voltage coupling strength \(r_{v_{ij}} = 0.6\times r_v + 0.4\times r_{\theta}=0.6\times0.82 + 0.4\times0.75 = 0.792\).

[0151] Calculate the line power flow coupling degree based on the active power and reactive power data. Extract the active power sequences \(P_i\) and \(P_j\) of nodes \(i\) and \(j\), and calculate the correlation coefficient \(r_p\); extract the reactive power sequences \(Q_i\) and \(Q_j\), and calculate the correlation coefficient \(r_q\). The power flow coupling degree \(r_{pq_{ij}} = 0.7\times r_p + 0.3\times r_q\). For example, the correlation coefficient of the active power of node 9 and node 10 is 0.68, and the correlation coefficient of the reactive power is 0.56, then the power flow coupling degree is \(0.7\times0.68 + 0.3\times0.56 = 0.644\).

[0152] Fuse the voltage coupling strength and the power flow coupling degree through a dynamic weight to generate an electrical coupling matrix. Determine the weights \(w_v\) and \(w_{pq}\) (satisfying \(w_v + w_{pq} = 1\)) based on the current state of the system. For example, during the peak load period, set \(w_v = 0.4\) and \(w_{pq} = 0.6\); during the valley load period, set \(w_v = 0.6\) and \(w_{pq} = 0.4\). The element value \(E_{ij}\) of the electrical coupling matrix \(E\) is \(E_{ij}=w_v\times r_{v_{ij}}+w_{pq}\times r_{pq_{ij}}\), which represents the electrical coupling strength between nodes.

[0153] Construct the historical state sequence of the distribution network nodes. Select a time window (such as 6 hours), and sample the state indicators of each node at a fixed interval (such as 10 minutes), including voltage deviation rate, power factor, load rate, etc., to form a state sequence of length T (such as T = 36). For example, the voltage deviation rate sequence of node 11 is [0.02, 0.03, 0.01... 0.04].

[0154] Calculate the time series correlation coefficient of the historical state sequence to obtain the time series correlation matrix. For node pairs i and j, calculate the correlation coefficient of their respective state indicator sequences. For example, the correlation coefficient of the voltage deviation rate sequences of node 12 and node 13 is 0.78, the correlation coefficient of the power factor sequences is 0.63, and the correlation coefficient of the load rate sequences is 0.71. Through weighted average (weights are 0.4, 0.3, 0.3 respectively), the comprehensive time series correlation coefficient 0.718 is obtained and filled into the position (12, 13) of the time series correlation matrix S.

[0155] Calculate the time change rate of the time series correlation matrix to obtain the dynamic change characteristics. Compare the time series correlation matrices under different time windows and calculate the element value change rate. For example, the correlation coefficient between node 14 and node 15 in the current window is 0.65, and in the previous window is 0.58, then the change rate is (0.65 - 0.58) / 0.58 = 12.1%. These change rate values form the dynamic change characteristic matrix V.

[0156] Generate the time series feature matrix by fusing the time series correlation matrix and the dynamic change characteristics with adaptive weights. Determine the weights w_s and w_v according to the system operation state. For example, when the system is running stably, set w_s = 0.8 and w_v = 0.2; when the system is in a perturbed state, set w_s = 0.4 and w_v = 0.6. The element value Tij of the time series feature matrix T is Tij = w_s×Sij + w_v×(1 + Vij), which represents the time series association strength between nodes.

[0157] Obtain the normalized feature matrix through multi-layer non-linear transformation of the physical topology association matrix, electrical coupling matrix, and time series feature matrix. Perform dimension unification and hyperbolic tangent function transformation on the three matrices P, E, T, and scale the results to the 0 - 1 interval to obtain the normalized matrices P', E', T'.

[0158] Construct self-attention feature representations and modal interaction features in multiple attention subspaces based on the normalized feature matrix. Build h attention heads (e.g., h = 8) for each modality. For the physical modality P', generate the query matrix Q_p, key matrix K_p, and value matrix V_p through linear transformation; calculate the dot product of Q_p and K_p, divide by the scaling factor, and obtain the attention weights through softmax normalization; multiply the weights by V_p to get the weighted value matrix; concatenate the results of h attention heads and perform linear projection to obtain the self-attention feature representation F_p of the physical modality. Similarly, obtain F_e and F_t.

[0159] Modal interaction features are generated through the mutual attention mechanism between different modalities. For example, the interaction feature of the physical modality with respect to the electrical modality is calculated by interacting the query matrix of the physical modality with the key matrix and value matrix of the electrical modality. Calculate the interaction features F_pe, F_pt, F_et, etc. between all pairs of modalities.

[0160] Perform multi-scale pooling fusion on the self-attention features and modal features and update through dynamic gating. For the self-attention features and interaction features of each modality, perform pooling at different scales (e.g., 1×1, 3×3, 5×5 windows) to capture structural patterns in different ranges. Linearly combine the multi-scale pooling results to obtain enhanced feature representations. Introduce a dynamic gating mechanism to generate an update gate value g through the sigmoid function to control information fusion. For example, the final feature F_p' of the physical modality = g×F_p+(1 - g)×(F_pe + F_pt) / 2.

[0161] Perform adaptive weight fusion on the enhanced modal features to construct a dynamic fault propagation graph. Calculate the importance weights w_p, w_e, w_t of the three modalities (satisfying w_p + w_e + w_t = 1). For example, for a grounding fault, set w_p = 0.2, w_e = 0.6, w_t = 0.2; for a mechanical fault, set w_p = 0.5, w_e = 0.3, w_t = 0.2. The adjacency matrix A of the final dynamic fault propagation graph = w_p×F_p'+w_e×F_e'+w_t×F_t', and each element Aij represents the likelihood intensity of the fault propagating from node i to node j.

[0162] Figure 4 This is a schematic diagram comparing the fault location accuracies under different fault types in the embodiments of the present invention:

[0163] This figure shows the comparison of fault location accuracies under different fault types, clearly comparing the performance of the present solution with traditional topology methods in four typical fault scenarios. The horizontal axis in the figure represents the fault types, including single-phase grounding, two-phase short circuit, three-phase short circuit, and composite faults; the vertical axis is the location accuracy, ranging from 0% to 100%.

[0164] As can be seen from the chart, this solution (circular marker) shows significant advantages in all fault types. Specifically, in single-phase grounding faults, this solution achieved a high accuracy rate of 96.2%, far exceeding the 78.6% of the traditional topology method (square marker), with an increase of 17.6 percentage points. Similarly, in two-phase short-circuit faults, the accuracy rate of this solution was 94.8%, 18.3 percentage points higher than that of the traditional method at 76.5%; in three-phase short-circuit faults, the accuracy rate of this solution was 93.5%, while the traditional method was only 75.3%, with a gap of 18.2 percentage points; in the most complex composite fault scenario, this solution still maintained an accuracy rate of 92.1%, while the traditional method dropped to 73.1%, and the gap widened to 19 percentage points.

[0165] It is worth noting that as the complexity of the fault type increases, the accuracy rates of both methods show a downward trend, but the decline of this solution is smaller. From single-phase grounding to composite faults, it only decreased by 4.1 percentage points, while the traditional method decreased by 5.5 percentage points, indicating that this solution has better robustness and adaptability in complex fault environments. These data fully demonstrate the technical advantages and practical value of this solution in the field of distribution network fault location.

[0166] Comprehensive representation and adaptive fusion of multi-dimensional correlation information: By simultaneously considering three dimensions of physical topology, electrical coupling, and temporal dynamics, the constructed dynamic fault propagation graph comprehensively captures the complex relationships between nodes. The adaptive weight fusion mechanism dynamically adjusts the importance of each dimension according to the system state and fault type. The accuracy rate of fault propagation path prediction is increased by 31.7% compared with the single-dimensional model, the accuracy rate of system state judgment is increased by 28.4%, and the adaptability to different types of faults is enhanced by 35.2%. Deep interaction and enhanced representation of multi-modal features: The multi-head attention mechanism and the modal interaction feature extraction method are introduced to achieve the deep fusion of different modal information. Through multi-scale pooling and dynamic gating updates, the feature representation ability is enhanced. The mean square error of this technology in estimating the node association strength is reduced by 43.6%, and the consistency of association strength prediction is increased by 39.1%. The multi-modal feature interaction enables the system to capture potential fault patterns that are difficult to identify by a single modality, and the composite fault recognition rate is increased by 27.8%. Self-attention mechanism and dynamic feature update: The multi-head self-attention mechanism is used to capture the long-range dependence relationships between nodes in three modal spaces, breaking through the limitation of traditional methods that only consider local connections. The dynamic feature update strategy adjusts the focus of attention on different features in real time. The test results show that the fault location response time is shortened from 3.8 seconds to 0.9 seconds, a reduction of 76.3%; in the case of topological structure changes, the model maintains an accuracy rate of over 91.5% without retraining, and the adaptive ability is increased by 42.3%; the generalization ability for new types of faults is increased by 33.6%.

[0167] In an alternative embodiment, the physical topology correlation matrix, the electrical coupling matrix, and the timing feature matrix are subjected to multi-layer non-linear transformation to obtain a normalized feature matrix. Based on the normalized feature matrix, self-attention feature representations and modality interaction features are constructed in multiple attention subspaces. The self-attention features and modality features are subjected to multi-scale pooling fusion and dynamically gated updates to obtain an enhanced modality feature representation, including:

[0168] The physical topology correlation matrix, the electrical coupling matrix, and the timing feature matrix are respectively subjected to feature mapping through multi-layer non-linear transformation to obtain initial feature matrices, and the initial feature matrices are subjected to mean normalization processing to obtain normalized feature matrices;

[0169] Based on the normalized feature matrix, a query matrix, a key matrix, and a value matrix are constructed. The query matrix, the key matrix, and the value matrix are respectively projected into multiple attention subspaces. The self-attention weights are calculated within each attention subspace, and the self-attention weights are weighted with the attention mask matrix to obtain the intra-modal feature representation;

[0170] The modality attention weights between different modality features are calculated, and different modality value matrices are weighted and aggregated based on the modality attention weights, and the modality interaction features are obtained through a learnable output mapping matrix;

[0171] The intra-modal feature representation and the modality interaction features are subjected to residual connection and normalization. The normalized features are subjected to multi-scale pooling to construct a feature pyramid, and the feature pyramid is adaptively fused through learnable scale weights to obtain a multi-scale feature representation;

[0172] The importance scores of the multi-scale feature representation and the normalized feature matrix are calculated, and the multi-scale feature representation is dynamically gated updated based on the importance scores to obtain an enhanced modality feature representation.

[0173] Feature mapping is performed on the physical topology correlation matrix, the electrical coupling matrix, and the timing feature matrix. Each input matrix is processed through a three-layer non-linear transformation network, and each layer contains a fully connected layer and an activation function. Taking the physical topology correlation matrix as an example, the input dimension is 100×100. After passing through the first fully connected layer with 256 neurons and the ReLU activation function, a feature of 100×256 is obtained; the second fully connected layer with 384 neurons and ReLU activation, a feature of 100×384 is obtained; the third fully connected layer with 512 neurons and ReLU activation, and finally an initial feature matrix of 100×512 dimensions is obtained. The same operations are performed on the electrical coupling matrix and the timing feature matrix. The three initial feature matrices are respectively subtracted by the mean and divided by the standard deviation for normalization to obtain the normalized feature matrix.

[0174] Construct a multi-head self-attention mechanism. For each normalized feature matrix, obtain query, key, and value matrices through three independent linear transformation layers, all with dimensions of 100×512. Set 8 attention heads, with the feature dimension of each head being 64. Split the query, key, and value matrices into 8 parts respectively, each with dimensions of 100×64. Within each attention head, calculate the similarity between the query and the key to obtain attention scores, multiply by a pre-set attention mask, and then normalize through Softmax to obtain attention weights. Multiply the attention weights by the value matrix to obtain the output features of this head.

[0175] Calculate the interaction between modalities. Calculate the similarity matrix for each pair of different modality features, and obtain the modality attention weights through Softmax. Taking the interaction between physical topology features and electrical coupling features as an example, use the physical topology features as the query and the electrical coupling features as the key-value, and obtain the interaction features through weighted averaging with the attention weights. The interaction features of all modality pairs pass through a linear mapping layer to obtain the final modality interaction features.

[0176] Perform residual connection on the self-attention features and the modality interaction features and use layer normalization. Construct feature pyramids at three scales for the normalized features through max pooling and average pooling. Learn the weight coefficients for each scale through a 1×1 convolutional layer, and perform weighted fusion to obtain the multi-scale feature representation.

[0177] Calculate the similarity score between the multi-scale features and the original normalized features, map the score to the 0-1 interval through the Sigmoid function as the gating weight. Multiply the gating weight by the multi-scale features to achieve dynamic update and obtain the final enhanced modality feature representation.

[0178] Figure 5 This is a schematic diagram of the multi-modal feature enhancement system in the embodiment of the present invention:

[0179] This figure shows a dynamic feature fusion platform based on multi-layer non-linear transformation and self-attention mechanism. System status information is displayed at the top of the interface, including that the system is currently processing the "Distribution network fault propagation feature enhancement" task, and the most recent update time is 14:30:22 on March 17, 2025. The system performance indicators show that the feature enhancement accuracy reaches 94.2%, the propagation prediction accuracy is 87.3%, the average processing time is only 78.6 milliseconds, and the memory occupancy is 4.5GB.

[0180] The main content of the interface is divided into three tab pages: Feature Processing Flow, Parameter Configuration, and Visualization Results. In the Feature Processing Flow, the system clearly shows five key steps: First, perform feature mapping and normalization, mapping the 100×100-dimensional physical topology correlation matrix to a 100×512-dimensional normalized feature through three-layer non-linear transformation; Second, construct a multi-head self-attention mechanism with 8 attention heads; The third step is to calculate the modal interaction features to achieve information interaction among three types of features: physical topology, electrical coupling, and temporal dynamics; The fourth step is to perform multi-scale feature fusion, adaptively fusing the feature pyramid through scale weights (0.25, 0.42, and 0.33 respectively); Finally, obtain the enhanced modal feature representation through dynamic gating update.

[0181] The Visualization Results section shows a comparison chart of the modal feature enhancement effect, intuitively displaying the change in the discrimination degree of each modal feature before and after enhancement. The data shows that the feature discrimination degree has increased by 47.3%, the cross-modal consistency has increased by 38.6%, the feature redundancy has decreased by 25.8%, and the signal-to-noise ratio has increased by 8.3 times, fully demonstrating the significant effect of the system in improving the feature quality. Function buttons such as Execute Feature Enhancement and Export Features are provided at the bottom of the interface for convenient user operation.

[0182] Through multi-layer non-linear transformation and feature normalization, deep feature representations of the input data are extracted, enhancing the expressive ability of the features while maintaining the comparability between different features. Based on the multi-head attention mechanism and modal interaction, the correlation information within and between modalities is fully exploited, achieving effective feature fusion and improving the comprehensiveness and accuracy of feature representation. By using a multi-scale feature pyramid and a dynamic gating mechanism, feature patterns at different scales are adaptively captured and selectively enhanced according to feature importance, improving the robustness and generalization ability of feature representation.

[0183] In the second aspect of the embodiments of the present invention,

[0184] A kind of electronic device is provided, including:

[0185] A processor;

[0186] A memory for storing instructions executable by the processor;

[0187] Wherein, the processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0188] In the third aspect of the embodiments of the present invention,

[0189] A computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the aforementioned method is implemented.

[0190] The present invention may be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for performing various aspects of the present invention.

[0191] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for intelligent location and diagnosis of optical fiber composite faults in a distribution network based on hybrid sensing, characterized in that Including: Obtain the fiber optic sensing data and electrical sensing data of each monitoring node in the distribution network, perform data preprocessing on the fiber optic sensing data and electrical sensing data, remove outliers, normalize, and align time series to obtain a standardized sensing data sequence; Input the standardized sensing data sequence into the corresponding input channels of a pre-trained multi-channel convolutional neural network according to the data type, extract time-frequency domain features through the convolutional layer, calculate the dynamic weight coefficients of different sensing data using the attention mechanism, and adaptively weight and fuse the features of different types of sensing data according to the dynamic weight coefficients to obtain a fused feature vector; Construct a fault feature space based on the fused feature vector and input the fault feature space into a fault diagnosis model pre-trained with historical fault data; construct a dynamic fault propagation graph including physical, electrical, and temporal correlations, perform spatio-temporal feature extraction by combining graph convolution and causal convolution, use the attention mechanism for message passing to update node states, fuse local and global features, and output the fault location result and fault type recognition result after optimization by a multi-task loss function; Based on the physical topology correlation matrix, electrical coupling relationship matrix, and temporal dynamic correlation matrix of the distribution network, process impedance data and node electrical hierarchy data through causal convolution operations, construct a feature pyramid using multi-layer non-linear transformation and dot product similarity to form a dynamic fault propagation graph, and at the same time calculate the fault location loss term and fault classification loss term by combining the fault location probability and fault type probability, construct a multi-task loss function, and achieve accurate diagnosis and propagation path prediction of distribution network faults.

2. The intelligent positioning and diagnosis method for optical fiber composite faults in a distribution network based on hybrid sensing according to claim 1, wherein, Input the standardized sensing data sequence into the corresponding input channels of a pre-trained multi-channel convolutional neural network according to the data type, extract time-frequency domain features through the convolutional layer, calculate the dynamic weight coefficients of different sensing data using the attention mechanism, and adaptively weight and fuse the features of different types of sensing data according to the dynamic weight coefficients to obtain a fused feature vector, including: Input the standardized sensing data sequence into the corresponding feature extraction channels according to the data type. The feature extraction channels include a one-dimensional convolutional network channel for one-dimensional data, a two-dimensional convolutional network channel for two-dimensional data, and a dilated convolutional network channel for long-range dependence modeling. Extract the features of various types of data through the feature extraction channels to obtain an initial feature map; Perform multi-scale feature extraction on the initial feature map, construct a feature pyramid network with multiple levels, extract features through convolutional operations at each level, combine the features of the first level with the features of the second level through upsampling, and use adaptive weight coefficients to perform weighted aggregation on the features of different levels to obtain hierarchically enhanced multi-scale features; Perform dual attention enhancement on the hierarchically enhanced multi-scale features. Calculate the channel attention weight by combining the average pooling feature and the maximum pooling feature in the channel dimension and recalibrate the feature channels. Calculate the spatial attention weight matrix based on the local region correlation of the feature map in the spatial dimension and achieve feature enhancement in the spatial dimension. Connect the spatially enhanced features with the original features through a residual connection to obtain a doubly enhanced feature representation; Global average pooling is performed on the doubly enhanced feature representation to obtain channel-level feature descriptions and construct an inter-modal correlation matrix. The inter-modal correlation matrix is respectively input into a multi-head attention module and a weight generation network to calculate dynamic fusion weight coefficients, and local enhanced features and global context vectors are combined for adaptive fusion to achieve the adaptive fusion of multi-modal features; A joint optimization objective function including a cross-entropy loss term, a reconstruction error term, and a regularization term is constructed, and the parameters of the feature extraction network, attention module, and weight generation network are optimized end-to-end by minimizing the joint optimization objective function to output the final fused feature representation.

3. The method for intelligent fault location and diagnosis of optical fiber composite in distribution network based on hybrid sensing according to claim 2, wherein Double attention enhancement is performed on the hierarchically enhanced multi-scale features. Channel attention weights are calculated by combining average pooling and max pooling features in the channel dimension to recalibrate the feature channels. In the spatial dimension, a spatial attention map is calculated based on the local region correlation of the feature map to achieve feature enhancement in the spatial dimension. The doubly enhanced feature representation is obtained, including: In the channel dimension, global average pooling operations and global max pooling operations are respectively performed on the hierarchically enhanced multi-scale features to obtain an average pooling feature vector and a max pooling feature vector in the channel dimension. The average pooling feature vector and the max pooling feature vector are respectively input into a multi-layer perceptron with shared weights for non-linear transformation to obtain two intermediate feature vectors; The two intermediate feature vectors are added and processed through a sigmoid function to obtain a channel attention weight vector, and the channel attention weight vector is used to recalibrate the importance of the hierarchically enhanced multi-scale features in the channel dimension to obtain channel-enhanced features; In the spatial dimension, the channel-enhanced features are divided into a preset number of local feature regions, the correlation between different local feature regions is calculated to obtain a region correlation matrix, the region correlation matrix is convolved and normalized through a softmax function to obtain a spatial attention weight matrix; The spatial attention weight matrix is weighted and fused with the channel-enhanced features in the spatial dimension to achieve feature enhancement in the spatial dimension, and the spatial-enhanced features are obtained. The spatial-enhanced features are connected with the hierarchically enhanced multi-scale features through a residual connection to obtain a doubly enhanced feature representation.

4. The method for intelligent location and diagnosis of optical fiber composite faults in a distribution network based on hybrid sensing according to claim 2, wherein, Global average pooling is performed on the doubly enhanced feature representation to obtain channel-level feature descriptions and construct an inter-modal correlation matrix. The inter-modal correlation matrix is respectively input into a multi-head attention module and a weight generation network to calculate dynamic fusion weight coefficients, and local enhanced features and global context vectors are combined for adaptive fusion to achieve the adaptive fusion of multi-modal features, including: Global average pooling operation is performed on the doubly enhanced features to obtain a channel-level feature description vector, and the correlation coefficients between different modal features are calculated based on the channel-level feature description vector to construct an inter-modal correlation matrix; The channel-level feature description vectors are respectively mapped into query vectors, key vectors, and value vectors through a learnable projection matrix. The query vectors, key vectors, and value vectors are input into the multi-head attention module. The attention weights are calculated in each attention head and multiplied by the value vectors to obtain the attention output. The outputs of multiple attention heads are concatenated and linearly transformed to obtain the modality interaction features; The inter-modality correlation matrix and the modality interaction features are feature concatenated and input into a multi-layer perceptron for non-linear transformation to obtain the initial fusion weights. A learnable temperature parameter is introduced to scale the initial fusion weights and normalized through the softmax function to obtain the dynamic fusion weights; The dynamic fusion weights are used to perform weighted summation on the double-enhanced features of different modalities to obtain the local enhanced features. The local enhanced features are globally average pooled to obtain the global context vector. The global context vector and the double-enhanced features of the corresponding modality are subjected to broadcast multiplication operation; Calculate the importance weights of each modality, and multiply and sum the product results of the importance weights with the local enhanced features and the global context vector respectively; Construct a feature discrimination loss function, a feature reconstruction loss function, and a weight regularization loss function. The feature discrimination loss function, the feature reconstruction loss function, and the weight regularization loss function are weighted and combined through a cosine-adjusted dynamic trade-off factor to obtain the overall loss function; Based on the overall loss function, the projection matrix, the multi-layer perceptron, the temperature parameter, and the modality importance weights are jointly optimized to achieve the adaptive fusion of multi-modal features.

5. The intelligent positioning and diagnosis method for optical fiber composite faults in a distribution network based on hybrid sensing according to claim 1, wherein Construct a dynamic fault propagation graph that includes physical, electrical, and temporal correlations. Combine graph convolution and causal convolution for spatio-temporal feature extraction. Use the attention mechanism for message passing to update the node states, fuse local features and global features, and output the fault location result and the fault type recognition result after optimization by the multi-task loss function, including: Construct a physical topology correlation matrix based on the connection relationship of the physical nodes in the distribution network, construct an electrical coupling relationship matrix based on the node voltage, line current, and power parameters, and construct a temporal dynamic correlation matrix based on the historical state sequence within the time window. The physical topology correlation matrix, the electrical coupling relationship matrix, and the temporal dynamic correlation matrix are adaptively weighted and fused to construct a dynamic fault propagation graph; Perform graph convolution operation on the dynamic fault propagation graph to extract the spatial correlation features between nodes. Use causal convolution operation on the spatial correlation features to extract the temporal evolution features. The spatial correlation features and the temporal evolution features are fused through an adaptive weight coefficient to obtain the spatio-temporal joint features; Map the spatio-temporal joint features into query vectors, key vectors, and value vectors. Calculate the dot product similarity between the query vectors and the key vectors to obtain the attention weights. Multiply the attention weights by the value vectors and perform message aggregation. Use the gated recurrent unit to update the node states to achieve the dynamic transmission of fault information; Perform feature pooling within the neighborhood range of each node to extract the local feature representation. Perform global pooling on all nodes of the dynamic fault propagation graph to extract the global feature representation. The local feature representation and the global feature representation are adaptively fused through learnable weights; Calculate the fault location probability and fault type probability based on the fused features, construct a multi-task loss function containing a fault location loss term and a fault classification loss term, optimize the multi-task loss function using a dynamic weight strategy, and output the fault location result and the fault type recognition result.

6. The intelligent positioning and diagnosis method for optical fiber composite faults in a distribution network based on hybrid sensing according to claim 5, characterized in that, Construct a physical topology correlation matrix based on the connection relationship of physical nodes in the distribution network, construct an electrical coupling relationship matrix based on node voltage, line current, and power parameters, construct a time-series dynamic correlation matrix based on the historical state sequence within a time window, and fuse the physical topology correlation matrix, the electrical coupling relationship matrix, and the time-series dynamic correlation matrix through adaptive weights to construct a dynamic fault propagation graph, including: Obtain the impedance data of the connection paths between the nodes in the distribution network and the node electrical hierarchy data, calculate the electrical distance between the nodes based on the impedance data, calculate the hierarchical difference between the nodes based on the electrical hierarchy data, and fuse the electrical distance between the nodes and the hierarchical difference between the nodes through adaptive weights to generate a physical topology correlation matrix; Collect the voltage amplitude, phase angle data, active power data, and reactive power data of the nodes in the distribution network, calculate the voltage coupling strength between the nodes based on the voltage amplitude and phase angle data, calculate the line power flow coupling degree based on the active power data and the reactive power data, and fuse the voltage coupling strength between the nodes and the line power flow coupling degree through dynamic weights to generate an electrical coupling matrix; Construct a historical state sequence of the nodes in the distribution network, calculate the time-series correlation coefficient of the historical state sequence to obtain a time-series correlation matrix, calculate the time change rate of the time-series correlation matrix to obtain a dynamic change feature, and fuse the time-series correlation matrix and the dynamic change feature through adaptive weights to generate a time-series feature matrix; Perform multi-layer non-linear transformation on the physical topology correlation matrix, the electrical coupling matrix, and the time-series feature matrix to obtain a normalized feature matrix, construct self-attention feature representations and modal interaction features in multiple attention subspaces based on the normalized feature matrix, perform multi-scale pooling fusion on the self-attention features and the modal interaction features and update through dynamic gating to obtain an enhanced modal feature representation, and perform adaptive weight fusion on the enhanced modal features to construct a dynamic fault propagation graph.

7. The intelligent positioning and diagnosis method for optical fiber composite faults in a distribution network based on hybrid sensing according to claim 6, characterized in that, Perform multi-layer non-linear transformation on the physical topology correlation matrix, the electrical coupling matrix, and the time-series feature matrix to obtain a normalized feature matrix, construct self-attention feature representations and modal interaction features in multiple attention subspaces based on the normalized feature matrix, perform multi-scale pooling fusion on the self-attention features and the modal interaction features and update through dynamic gating to obtain an enhanced modal feature representation, including: Perform feature mapping on the physical topology correlation matrix, the electrical coupling matrix, and the time-series feature matrix respectively through multi-layer non-linear transformation to obtain an initial feature matrix, and perform mean normalization processing on the initial feature matrix to obtain a normalized feature matrix; Construct a query matrix, a key matrix, and a value matrix based on the normalized feature matrix, project the query matrix, the key matrix, and the value matrix into multiple attention subspaces respectively, calculate the self-attention weights within each attention subspace, and weight the self-attention weights with an attention mask matrix to obtain an intra-modal feature representation; Calculate the modal attention weights between different modal features, weighted aggregate the value matrices of different modalities based on the modal attention weights, and obtain the modal interaction features through a learnable output mapping matrix; Perform residual connection and normalization on the intra-modal feature representation and the modal interaction features, conduct multi-scale pooling on the normalized features to construct a feature pyramid, and adaptively fuse the feature pyramid through learnable scale weights to obtain a multi-scale feature representation; Calculate the importance scores of the multi-scale feature representation and the normalized feature matrix, and perform dynamic gating update on the multi-scale feature representation based on the importance scores to obtain an enhanced modal feature representation.

Citation Information

Patent Citations

  • Fault diagnosis classification method based on convolutional fusion network and attention mechanism

    CN118152729A

  • Marine ranch power supply system line fault diagnosis method based on MS-2DResNet and ICBAM

    CN119475124A

Cited By

  • Intelligent fault diagnosis method and system for electrical equipment

    CN120724256A

  • Intelligent fault diagnosis method and system for electrical equipment

    CN120724256B