An online diagnostic device and method for multispectral fusion faults in power equipment
By using a dynamic kernel generation network and a dual-path attention mechanism, combined with wavelet packet decomposition and Hilbert spectral analysis, the problem of insufficient cross-modal nonlinear correlation modeling in multispectral image fusion is solved, enabling accurate diagnosis and early warning of power equipment faults.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGDONG SENXU GENERAL EQUIP TECH CO LTD
- Filing Date
- 2025-12-31
- Publication Date
- 2026-06-02
AI Technical Summary
Existing multispectral image fusion methods are insufficient in modeling cross-modal nonlinear correlations, making it difficult to uncover the nonlinear complementary relationships between visible light, thermal infrared, and ultraviolet images. Furthermore, they lack the ability to track the dynamic evolution of faults, resulting in low detection accuracy and failing to meet the needs of smart grids for early warning and precise location.
By using a dynamic kernel generation network and a dual-path attention mechanism, a multimodal feature map is generated. Wavelet packet multi-scale decomposition and Hilbert spectral analysis are then performed to construct a multi-scale spatiotemporal feature sequence. Combined with a multilayer perceptron classification network, the probability distribution of fault types and diagnostic reports are generated.
It achieves deep interactive fusion of multispectral image features, which can accurately characterize the evolution law of the entire life cycle of faults, improve the early warning capability of early faults and latent defects and the reliability of diagnostic results.
Smart Images

Figure CN122134613A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power equipment fault diagnosis technology, and in particular to an online diagnostic device and method for multispectral fusion faults in power equipment. Background Technology
[0002] With the deepening of smart grid construction, power equipment condition monitoring and fault diagnosis technologies based on multispectral image detection are gradually evolving from traditional periodic maintenance to intelligent online diagnosis. Currently, multispectral image detection methods have become an important research direction in this field. By simultaneously acquiring visible light, thermal infrared, and ultraviolet spectral images, and utilizing the differences in the responses of different spectra to abnormal equipment states, they enable the visual detection of typical faults such as corona discharge, localized overheating, and insulation degradation. Existing technologies mainly revolve around two levels: feature-level fusion and decision-level fusion. In terms of image feature extraction, convolutional neural networks are used to automatically learn features from each spectral image, effectively improving the ability to represent fault areas. In terms of multispectral image fusion, information from visible light, thermal infrared, and ultraviolet images is integrated through attention mechanisms, weighted averaging, or tensor decomposition to overcome the detection blind spots of single-spectral images under complex backgrounds, lighting changes, or partial occlusion. In recent years, temporal modeling methods have been introduced into the multispectral image detection process, extending from single-frame static image detection to multi-frame temporal evolution perception through dynamic analysis of continuous frame image sequences.
[0003] Despite some progress in existing technologies, two key challenges remain in achieving high-precision and robust online image detection: First, existing multispectral image fusion methods have limited ability to model cross-modal feature interactions. Most employ channel stitching or simple weighting strategies, failing to fully exploit the nonlinear complementary relationships between the texture details of visible light images, the temperature distribution of thermal infrared images, and the discharge intensity of ultraviolet spectral images. This results in insufficient discriminative power of the fused image features under complex conditions such as strong interference, low signal-to-noise ratio, or target occlusion, affecting the accurate detection of fault areas. Second, existing image detection models are generally based on single or multiple static images for analysis, lacking effective modeling of the temporal evolution of faults. They cannot continuously track the dynamic process from early weak anomalies to obvious faults, leading to low sensitivity in image detection of latent defects and initial faults, making it difficult to meet the actual needs of smart grids for early warning and precise location. Summary of the Invention
[0004] In view of the aforementioned existing problems, the present invention is proposed.
[0005] Therefore, this invention provides an online diagnostic method for multispectral fusion faults in power equipment, which solves the problems of low detection accuracy caused by insufficient modeling of cross-modal nonlinear correlations and lack of fault dynamic evolution tracking capabilities in existing multispectral image fusion.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: In a first aspect, the present invention provides an online diagnostic method for multispectral fusion faults in power equipment, comprising: collecting visible light images, thermal infrared images, and ultraviolet spectral images of the power equipment and performing registration processing to generate a multispectral image data cube; inputting the multispectral image data cube into a dynamic kernel generation network to output a multimodal feature map set, and performing a dual-path attention mechanism on the multimodal feature map set to generate a fusion feature map focusing on key regions; performing wavelet packet multi-scale decomposition on the cross-modal fusion feature map, and constructing a second-level transient analysis sequence and a minute-level steady-state analysis sequence to form a multi-scale spatiotemporal feature sequence characterizing the fault evolution law; performing noise-assisted decomposition and Hilbert spectral analysis on the multi-scale spatiotemporal feature sequence, and generating a comprehensive spatiotemporal feature vector through a tensor chain decomposition algorithm; inputting the comprehensive spatiotemporal feature vector into a multilayer perceptron classification network to output the probability distribution of fault types, calculating the prediction confidence index, and generating a fault type diagnostic report.
[0007] As a preferred embodiment of the online diagnostic method for multispectral fusion faults in power equipment according to the present invention, the specific steps for generating the multispectral image data cube are as follows: Visible light images, thermal infrared images, and ultraviolet spectral images of power equipment are acquired, and non-uniformity correction and point spread function compensation processing are performed to obtain enhanced multispectral images. The enhanced multispectral image is segmented to generate a set of superpixel blocks; The matching cost of superpixel blocks between different spectral images in the superpixel block set is calculated by a multi-dimensional feature weighted fusion algorithm, and a superpixel-level correspondence is established. Control point pairs are determined based on superpixel-level correspondence, and pixel coordinate transformation relationships from thermal infrared images and ultraviolet spectral images to visible light images are established based on the control point pairs. By utilizing pixel coordinate transformation relationships, thermal infrared and ultraviolet spectral images are resampled and spatially aligned with visible light images to generate a multispectral image data cube.
[0008] As a preferred embodiment of the online diagnostic method for multi-spectral fusion faults in power equipment according to the present invention, the specific steps for outputting the multi-modal feature map set are as follows: Spectral dimension decoupling and edge enhancement processing are performed on the multispectral image data cube to generate an enhanced multispectral feature map; Based on the enhanced multispectral feature map, cross-modal association weights are obtained through a dynamic kernel generation network, and feature extraction is performed on the enhanced multispectral feature map to obtain the convolutionally enhanced feature map; Interactive enhancement of the feature maps extracted by convolution is performed through a cross-modal attention mechanism to generate attention-weighted multimodal feature maps; Attention-weighted feature maps are concatenated along the channel dimension and normalized by layers to generate a multimodal feature map set.
[0009] As a preferred embodiment of the online diagnostic method for multispectral fusion faults in power equipment according to the present invention, the specific steps for generating the fusion feature map focusing on key regions are as follows: Tensor transformation and dimensionality reduction are performed on the multimodal feature map to obtain the core feature tensor; Based on the core feature tensor, a fusion weight matrix is obtained through a dual-path attention mechanism; The core feature tensors are filtered for importance and weighted by a fusion weight matrix to obtain the enhanced fusion features; The enhanced fusion features are refined and gating at multiple scales to generate a fusion feature map that focuses on key regions.
[0010] As a preferred embodiment of the online diagnostic method for multispectral fusion faults in power equipment according to the present invention, the specific steps for forming a multi-scale spatiotemporal feature sequence characterizing the fault evolution law are as follows: Wavelet packet multi-scale decomposition is performed on the cross-modal fusion feature map to obtain multi-scale time-frequency feature sub-bands; Dynamic modal analysis was performed on multi-scale time-frequency characteristic subbands, and the dominant modal features of fault evolution were extracted. Construct second-level transient analysis sequences and minute-level steady-state analysis sequences based on dominant mode characteristics; Multi-scale fusion of transient and steady-state analysis sequences generates a multi-scale spatiotemporal feature sequence that characterizes the evolution of faults.
[0011] As a preferred embodiment of the online diagnostic method for multispectral fusion faults in power equipment according to the present invention, the specific steps for performing noise-assisted decomposition and Hilbert spectral analysis on the multi-scale spatiotemporal feature sequence are as follows: The intrinsic mode function components are obtained by performing noise-assisted decomposition on multi-scale spatiotemporal feature sequences using the ensemble empirical mode decomposition method. Hilbert spectral analysis was performed on the intrinsic mode function components, and instantaneous amplitude-frequency characteristics were extracted.
[0012] As a preferred embodiment of the online diagnostic method for multispectral fusion faults in power equipment according to the present invention, the specific steps for generating the comprehensive spatiotemporal feature vector are as follows: Based on instantaneous amplitude-frequency characteristics, feature fusion is performed using the tensor chain decomposition algorithm to obtain fused feature tensors; Multi-resolution optimization and dimensionality reduction are performed on the fused feature tensor to generate a comprehensive spatiotemporal feature vector.
[0013] As a preferred embodiment of the online diagnostic method for multispectral fusion faults in power equipment according to the present invention, the specific steps for obtaining the probability distribution of fault types are as follows: The integrated spatiotemporal feature vector is input into the multilayer perceptron classification network for forward propagation calculation, and the original classification score is output. Based on the Platt scaling algorithm, logistic regression fitting and calibration are performed on the original classification scores to obtain the probability distribution of fault types.
[0014] As a preferred embodiment of the online diagnostic method for multispectral fusion faults in power equipment according to the present invention, the specific steps for generating the fault type diagnostic report are as follows: The prediction confidence index is obtained based on the probability distribution of fault types, and the confidence score is calculated. The confidence score and diagnostic confidence threshold are used to perform decision analysis on the probability distribution of fault types, and a fault type diagnostic report is generated.
[0015] Secondly, this invention provides an online diagnostic device for multispectral fusion faults in power equipment, comprising a data acquisition module, a feature fusion module, a sequence construction module, a feature generation module, and an intelligent diagnostic module. The data acquisition module acquires visible light images, thermal infrared images, and ultraviolet spectral images of the power equipment, performs registration processing, and generates a multispectral image data cube. The feature fusion module inputs the multispectral image data cube into a dynamic kernel generation network, outputs a multimodal feature map set, and performs a dual-path attention mechanism on the multimodal feature map set to generate a fusion feature map focusing on key regions. The sequence construction module performs wavelet packet multi-scale decomposition on the cross-modal fusion feature map and constructs a second-level transient analysis sequence and a minute-level steady-state analysis sequence to form a multi-scale spatiotemporal feature sequence characterizing the fault evolution law. The feature generation module performs noise-assisted decomposition and Hilbert spectral analysis on the multi-scale spatiotemporal feature sequence, and generates a comprehensive spatiotemporal feature vector through a tensor chain decomposition algorithm. The intelligent diagnostic module inputs the comprehensive spatiotemporal feature vector into a multilayer perceptron classification network, outputs the probability distribution of fault types, calculates the prediction confidence index, and generates a fault type diagnostic report.
[0016] The beneficial effects of this invention are as follows: By using a dynamic kernel generation network and a dual-path attention mechanism, the transformation of multispectral image features from simple superposition to deep interactive fusion is realized, solving the problems of shallow mining of complex nonlinear correlations between visible light, thermal infrared, and ultraviolet images and poor robustness of fused features under complex working conditions such as occlusion and interference in existing image detection methods; by using wavelet packet multi-scale decomposition and constructing second-level transient and minute-level steady-state analysis sequences, the breakthrough of fault signal detection in multispectral images from single-frame static detection to cross-time-series dynamic full-process tracking is realized, which can accurately characterize the evolution law of the entire life cycle of faults, improve the early warning capability of image detection for faults and latent defects, and enhance the reliability of diagnostic results. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a flowchart of an online diagnostic method for multispectral fusion faults in power equipment.
[0019] Figure 2 A flowchart for generating a multispectral image data cube.
[0020] Figure 3 A flowchart for generating a fused feature map that focuses on key regions.
[0021] Figure 4 A flowchart for generating multi-scale spatiotemporal feature sequences. Detailed Implementation
[0022] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0023] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0024] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.
[0025] Reference Figures 1-4 As one embodiment of the present invention, this embodiment provides an online diagnostic method for multispectral fusion faults in power equipment, comprising the following steps: S1. Acquire visible light, thermal infrared, and ultraviolet spectral images of power equipment, perform registration processing, and generate a multispectral image data cube.
[0026] Visible light images, thermal infrared images, and ultraviolet spectral images of power equipment are acquired, and non-uniformity correction and point spread function compensation processing are performed to obtain enhanced multispectral images.
[0027] The specific process involves simultaneously deploying visible light imaging, thermal infrared imaging, and ultraviolet spectral imaging equipment to acquire images of the same power equipment under the same observation angle and time window. The visible light imaging equipment captures the visible light reflection information on the surface of the power equipment, the thermal infrared imaging equipment records the thermal radiation distribution generated during the operation of the power equipment, and the ultraviolet spectral imaging equipment detects the ultraviolet light signal excited by the power equipment during corona discharge or partial discharge. These three types of imaging equipment are calibrated and aligned in space and synchronized at the frame level in time to ensure that the acquired visible light image, thermal infrared image, and ultraviolet spectral image correspond to the same working state and geometric orientation of the power equipment. Non-uniformity correction is performed on these three types of images to eliminate image brightness or temperature distribution deviations caused by inconsistent sensor responses. Then, point spread function compensation processing is performed to restore the detail information lost due to optical blurring effects, ultimately obtaining an enhanced multispectral image.
[0028] It should be noted that point spread function (PSF) compensation is an image restoration technique used to correct blur caused by optical imperfections. During imaging, a point light source diffuses into a light spot, and its diffusion characteristics are described by the point spread function. PSF compensation estimates the point spread function and uses methods such as deconvolution to perform inverse filtering on the image, restoring details and edges, thereby obtaining an enhanced image.
[0029] The enhanced multispectral image is segmented to generate a set of superpixel blocks.
[0030] The specific process includes using a superpixel segmentation method to divide the enhanced multispectral image into several local regions with similar spectral and spatial characteristics. The pixels within each local region exhibit similar multispectral response characteristics in visible light images, thermal infrared images, and ultraviolet spectral images, thereby forming a set of compact and semantically consistent superpixel blocks, and finally generating a superpixel block set.
[0031] It should be noted that superpixel segmentation is a method of dividing an image into several local regions composed of pixels with similar colors, textures, and spatial locations. Each region is called a superpixel. It is used to reduce computational complexity and preserve structural information. Common methods include SLIC, Felzenszwalb algorithm, etc.
[0032] Multispectral response characteristics refer to the reflection, radiation, or excitation response features of the same physical location in visible light images, thermal infrared images, and ultraviolet spectral images to different electromagnetic bands. Specifically, in the visible light band, it mainly reflects the reflectivity of an object's surface to visible light; in the thermal infrared band, it reflects the intensity of thermal radiation generated by temperature differences in the object; and in the ultraviolet band, it captures the ultraviolet photon signals excited during corona discharge or partial discharge. These responses correspond to the same region spatially, but they have different characteristics in terms of numerical values and distribution patterns, collectively constituting the multispectral response characteristics of that location.
[0033] A multi-dimensional feature weighted fusion algorithm is used to measure the cross-modal similarity of superpixel block sets, obtain the matching cost of superpixel blocks between different spectral images, and establish superpixel-level correspondence.
[0034] The specific process includes extracting features of each superpixel block in multiple dimensions such as spectral response, texture distribution, and spatial coordinates based on the superpixel block sets corresponding to the visible light image, thermal infrared image, and ultraviolet spectral image. A multi-dimensional feature weighted fusion algorithm is used to assign corresponding weights to these features and fuse them to obtain the comprehensive difference degree between any two superpixel blocks from different spectral images. The comprehensive difference degree is used as the matching cost. The optimal pairing between superpixel blocks in different spectral images is determined according to the principle of minimizing the matching cost, thus establishing the superpixel-level correspondence between the visible light image, thermal infrared image, and ultraviolet spectral image.
[0035] It should be noted that the multi-dimensional feature weighted fusion algorithm is used to combine the values of superpixel blocks in visible light images, thermal infrared images and ultraviolet spectral images in multiple feature dimensions such as spectral response, texture and spatial location according to their respective weights to form a comprehensive feature representation.
[0036] Control point pairs are determined based on superpixel-level correspondence, and pixel coordinate transformation relationships from thermal infrared images and ultraviolet spectral images to visible light images are established based on the control point pairs.
[0037] The specific process includes selecting several pairs of spatially precisely corresponding superpixel center points as control point pairs from the established superpixel-level correspondence between visible light images, thermal infrared images, and ultraviolet spectral images. Using the control point pairs, the pixel coordinate mapping parameters of thermal infrared images and ultraviolet spectral images relative to visible light images are obtained through geometric transformation methods, thereby constructing pixel coordinate transformation relationships from thermal infrared images to visible light images and from ultraviolet spectral images to visible light images.
[0038] By utilizing pixel coordinate transformation relationships, thermal infrared and ultraviolet spectral images are resampled and spatially aligned with visible light images to generate a multispectral image data cube.
[0039] The specific process includes mapping and interpolating the pixel positions of the thermal infrared image and the ultraviolet spectral image according to the pixel coordinate transformation relationship from the thermal infrared image to the visible light image and from the ultraviolet spectral image to the visible light image, respectively, so that the spatial coordinates of the thermal infrared image and the ultraviolet spectral image are consistent with the pixel grid of the visible light image. Then, the resampled thermal infrared image, ultraviolet spectral image and visible light image are aligned pixel by pixel at the same spatial position and stacked along the spectral dimension to form a three-dimensional structure containing the visible light image, thermal infrared image and ultraviolet spectral image, generating a multispectral image data cube.
[0040] S2. Input the multispectral image data cube into the dynamic kernel generation network, output a multimodal feature map set, and perform a dual-path attention mechanism on the multimodal feature map set to generate a fused feature map focusing on key regions.
[0041] The spectral dimension is decoupled and edge enhancement is performed on the multispectral image data cube to generate an enhanced multispectral feature map.
[0042] The specific process includes performing spectral dimension decoupling on the multispectral image data. By analyzing the correlation between images of different bands in the multispectral image data cube, independent spectral components are separated, reducing redundant information between different bands. An edge enhancement algorithm is then applied to highlight important structural features in the image, such as equipment outlines or fault area boundaries. This process enhances the representation of features related to fault diagnosis in the image while maintaining the spectral characteristics of the image. Finally, an enhanced multispectral feature map is generated for subsequent fault detection and analysis steps.
[0043] It should be noted that edge enhancement algorithms are image processing methods that highlight areas of abrupt grayscale changes (such as contours and details) in an image through filtering. Common techniques include Sobel, Laplacian, and unsharpened masks, which are used to strengthen object boundaries to improve image clarity or analytical results.
[0044] Based on the enhanced multispectral feature map, cross-modal association weights are obtained through a dynamic kernel generation network, and feature extraction is performed on the enhanced multispectral feature map to obtain the convolutionally enhanced feature map.
[0045] The specific process includes: based on the enhanced multispectral feature map, a dynamic kernel generation network is used to process the mutual information matrix, identify and quantify the correlation strength between modes, and thus obtain the cross-modal correlation weights. These cross-modal correlation weights are then used to weight the enhanced multispectral feature map, emphasizing features that are more critical in fault diagnosis. Deep convolution operations are then used to further process the weighted feature map, extracting more abstract and high-level feature representations, ultimately obtaining the convolutionally enhanced feature map, which provides richer information support for subsequent steps. This process ensures that the information extracted from the multispectral image can effectively reflect the equipment's status and potential faults.
[0046] Furthermore, the training process of the dynamic kernel generation network is carried out on a labeled multispectral image dataset. The enhanced multispectral feature maps are input into the dynamic kernel generation network, which adaptively generates weight parameters for convolution operations based on the mutual information matrix between the enhanced multispectral feature maps. During the training phase, the parameters of the dynamic kernel generation network are iteratively optimized according to the loss function of the fault classification task through the backpropagation algorithm, so that the dynamic kernel generation network can learn effective association patterns between different spectral modes and generate cross-modal association weights that can highlight key fault features. The entire training process relies on an end-to-end supervised learning mechanism, with the goal of making the final output convolutionally enhanced feature maps have stronger discriminative ability in subsequent fault classification tasks.
[0047] By using a cross-modal attention mechanism, the feature maps extracted from convolution are interactively enhanced to generate attention-weighted multimodal feature maps.
[0048] The specific process includes interactive enhancement of the feature maps extracted by convolution through a cross-modal attention mechanism. This involves using the feature maps extracted by convolution from visible light images, thermal infrared images, and ultraviolet spectral images respectively to obtain the interdependencies between the feature maps in spatial location and channel dimension. Attention weights are generated based on these interdependencies, and then these attention weights are applied to the feature maps extracted by convolution from each modality to strengthen the semantically consistent or complementary regional responses, suppress irrelevant or conflicting information, and generate attention-weighted multimodal feature maps.
[0049] Attention-weighted feature maps are concatenated along the channel dimension and normalized by layers to generate a multimodal feature map set.
[0050] The specific process involves concatenating attention-weighted feature maps extracted from different spectral images according to their respective channel dimensions to form a new feature map set containing information from all input feature maps. To ensure the consistency of data distribution within the new feature map set and to reduce instability factors during training (such as vanishing or exploding gradients, overfitting, etc.), layer normalization is performed on the concatenated feature maps. Layer normalization adjusts the values within each feature map so that their mean is 0 and their variance is 1, resulting in a more optimized feature representation. The result after channel-dimensional concatenation and layer normalization is the multimodal feature map set, which integrates key information from multiple spectral sources.
[0051] The multimodal feature map is transformed into tensor form and dimensionality reduced to obtain the core feature tensor.
[0052] The specific process includes concatenating attention-weighted feature maps from visible light images, thermal infrared images, and ultraviolet spectral images along the channel dimension and performing layer normalization to form a multimodal feature map set. The dimensional structure is then rearranged to convert the two-dimensional feature map sequence into a unified high-dimensional tensor. Principal component analysis or linear projection and other dimensionality reduction methods are applied to the high-dimensional tensor to compress the number of channels and data size while retaining the main discriminative information, thereby obtaining the core feature tensor.
[0053] Based on the core feature tensor, a fusion weight matrix is obtained through a dual-path attention mechanism.
[0054] The specific process involves analyzing the core feature tensor and identifying the key information regions it contains. This is achieved through a dual-path attention mechanism: one path obtains the response statistics (such as activation values after global average pooling) of each channel of the core feature tensor in the channel dimension, and generates channel attention weights after nonlinear transformation to evaluate the contribution of different channels to fault detection; the other path aggregates the core feature tensor along the channel direction in the spatial dimension (such as summing or maximizing) to form a two-dimensional spatial map, and then generates spatial attention weights through convolution operations to locate fault-related spatial regions in the image. The two paths output channel attention maps and spatial attention maps respectively, and then multiply and fuse them element-wise to form a comprehensive attention representation. This comprehensive attention representation is normalized in the channel and spatial dimensions using the Softmax or Sigmoid function to obtain the weight coefficients corresponding to each position and channel, ultimately forming a fused weight matrix.
[0055] The importance of the core feature tensors is screened and weighted using a fusion weight matrix to obtain the enhanced fusion features.
[0056] The specific process involves multiplying each weight coefficient in the fusion weight matrix element-wise with the feature values of the corresponding positions and channels in the core feature tensor, so that the feature responses of high-weight regions are preserved or enhanced, while the feature responses of low-weight regions are suppressed or weakened, thereby highlighting information that is discriminative for fault diagnosis (e.g., features of specific types of fault modes, abnormal behaviors, or damaged components), and finally obtaining the enhanced fusion features.
[0057] The enhanced fusion features are refined and gating at multiple scales to generate a fusion feature map that focuses on key regions.
[0058] The specific process includes extracting contextual information of the enhanced fusion features at multiple scales using convolutional kernels of different scales, generating spatial or channel weights based on feature responses using a gating mechanism, weighting and filtering features at each scale, suppressing non-critical regions and retaining fault-related regions, and finally generating a fusion feature map focusing on key regions.
[0059] S3. Perform wavelet packet multi-scale decomposition on the cross-modal fusion feature map, and construct a second-level transient analysis sequence and a minute-level steady-state analysis sequence to form a multi-scale spatiotemporal feature sequence that characterizes the fault evolution law.
[0060] Wavelet packet multi-scale decomposition is performed on the cross-modal fusion feature map to obtain multi-scale time-frequency feature sub-bands.
[0061] The specific process includes taking the cross-modal fusion feature map as the input signal, and using wavelet packet transform to recursively decompose the cross-modal fusion feature map into multiple sub-bands in both time and frequency dimensions. Each decomposition layer further subdivides the low-frequency and high-frequency parts to obtain multi-scale time-frequency feature sub-bands covering different frequency bands and spatial scales.
[0062] Dynamic modal analysis was performed on multi-scale time-frequency characteristic subbands, and the dominant modal features of fault evolution were extracted.
[0063] The specific process includes taking multi-scale time-frequency characteristic sub-bands as input, using dynamic modal analysis to perform modal separation on the time series signals in each sub-band, identifying the inherent oscillation modes that characterize the fault evolution process of power equipment (e.g., periodic pulses caused by partial discharge, low-frequency drift caused by insulation degradation, or high-frequency resonance generated by corona discharge), and selecting modal components with concentrated energy and clear physical meaning as dominant modal features.
[0064] It should be noted that dynamic modal analysis is a signal analysis method that extracts dominant oscillation modes from time-series data. By expanding multi-scale time-frequency characteristic subbands into snapshot sequences over time, constructing linear propagation operators, and performing singular value decomposition and eigenvalue decomposition, intrinsic modal components with specific frequencies, attenuation rates, and spatial structures are obtained, which are used to characterize key evolutionary patterns in dynamic behavior.
[0065] Based on the dominant mode characteristics, second-level transient analysis sequences and minute-level steady-state analysis sequences are constructed.
[0066] The specific process includes dividing the dominant modal features according to the time dimension, sampling at second-level time intervals to form a second-level transient analysis sequence that reflects the rapid changes in the fault, and simultaneously aggregating or smoothing the dominant modal features at minute-level time intervals to form a minute-level steady-state analysis sequence that characterizes the long-term stable evolution trend of the fault.
[0067] Multi-scale fusion of transient and steady-state analysis sequences generates a multi-scale spatiotemporal feature sequence that characterizes the evolution of faults.
[0068] The specific process includes aligning and integrating the second-level transient analysis sequence and the minute-level steady-state analysis sequence in terms of time and feature dimensions, and using a multi-scale fusion method to organically combine short-term high-frequency transient change information with long-term low-frequency steady-state evolution trends to form a unified multi-scale spatiotemporal feature sequence that contains fault dynamic characteristics at different time granularities and characterizes the fault evolution law.
[0069] It should be noted that fault dynamic characteristics refer to the non-stationary and nonlinear behavior of electrical, thermal, or optical physical quantities of power equipment as they evolve over time during a fault. Fault dynamic characteristics are reflected in the changes in the amplitude, frequency, phase, energy distribution, or spatial morphology of signals during the initial, development, and deterioration stages of a fault. Examples include an increased repetition rate of partial discharge pulses, a continuous rise in hotspot temperature, and a sudden increase in ultraviolet radiation intensity. Fault dynamic characteristics reflect the fault type, severity, and evolution trend, and are a key basis for achieving early warning and accurate diagnosis.
[0070] S4. Perform noise-assisted decomposition and Hilbert spectral analysis on the multi-scale spatiotemporal feature sequences, and generate a comprehensive spatiotemporal feature vector through the tensor chain decomposition algorithm.
[0071] The intrinsic mode function components are obtained by performing noise-assisted decomposition on multi-scale spatiotemporal feature sequences using the ensemble empirical mode decomposition method.
[0072] The specific process includes adding white noise with different realizations multiple times to the multi-scale spatiotemporal feature sequence, performing empirical mode decomposition on the sequence after each noise addition to obtain several sets of intrinsic mode function components, and then averaging the intrinsic mode function components of the corresponding order in all sets to suppress the mode aliasing effect and improve the decomposition stability, and finally obtaining the intrinsic mode function components.
[0073] It should be noted that the ensemble empirical mode decomposition method is an improved empirical mode decomposition method. By adding different white noise multiple times to the multi-scale spatiotemporal feature sequence, performing empirical mode decomposition on each sequence, and averaging the intrinsic mode function components of the same order, mode aliasing is suppressed, and more stable intrinsic mode function components are finally obtained.
[0074] Hilbert spectral analysis was performed on the intrinsic mode function components, and instantaneous amplitude-frequency characteristics were extracted.
[0075] The specific process includes applying the Hilbert transform to generate corresponding analytic signals for the intrinsic mode function components. The analytic signals are composed of the original intrinsic mode function components as the real parts and the Hilbert transform result as the imaginary parts. The instantaneous phase is extracted from the analytic signals, and the instantaneous frequency is obtained by differentiating the instantaneous phase with respect to time. At the same time, the instantaneous amplitude is obtained from the magnitude of the analytic signals. The instantaneous amplitude value and instantaneous frequency corresponding to each moment are paired to form an instantaneous amplitude-frequency feature that can characterize the local oscillation characteristics of the signal, which is used to characterize the dynamic behavior of the fault in the time-frequency domain.
[0076] Based on instantaneous amplitude-frequency characteristics, feature fusion is performed using the tensor chain decomposition algorithm to obtain fused feature tensors.
[0077] The specific process includes organizing the instantaneous amplitude-frequency features from different intrinsic mode function components into a high-dimensional tensor form, using the tensor chain decomposition algorithm to perform low-rank approximate decomposition on the high-dimensional tensor form, compressing the high-dimensional features along multiple dimensions while retaining the intrinsic correlation structure, thereby reducing redundancy while achieving effective fusion of cross-modal and cross-scale features, and finally obtaining the fused feature tensor.
[0078] It should be noted that the tensor chain decomposition algorithm is a low-rank approximation method for high-dimensional tensors. It decomposes a high-order tensor into a series of interconnected third-order core tensors (chain segments). By controlling the connection dimension of each chain segment, it can compress data while retaining the main structural information. It is suitable for high-dimensional feature fusion and dimensionality reduction.
[0079] Multi-resolution optimization and dimensionality reduction are performed on the fused feature tensor to generate a comprehensive spatiotemporal feature vector.
[0080] The specific process includes enhancing the details and smoothing the structure of the fused feature tensor at different resolution levels to retain key spatiotemporal dynamic information and suppress redundant fluctuations, and using existing dimensionality reduction methods such as principal component analysis or linear projection to compress the dimension, thereby reducing the number of features while maintaining discriminative ability and generating a comprehensive spatiotemporal feature vector.
[0081] It should be noted that key spatiotemporal dynamic information refers to discriminative information that simultaneously encompasses the temporal variation patterns and spatial distribution characteristics during the evolution of power equipment faults. In the temporal dimension, it manifests as the trend, periodicity, suddenness, or attenuation characteristics of fault-related signals (such as discharge pulses, temperature rises, and enhanced ultraviolet radiation) over time. In the spatial dimension, it reflects the specific location, extent, and form of abnormal phenomena on the equipment surface or within its internal structure, such as hotspot areas, corona occurrence points, or insulation damage sites. Key spatiotemporal dynamic information integrates the dual clues of "when it happened" and "where it happened," accurately depicting the initiation, development, and deterioration of faults, and serving as the core basis for achieving high-precision fault diagnosis and location.
[0082] S5. Input the comprehensive spatiotemporal feature vector into the multilayer perceptron classification network, output the probability distribution of fault types, calculate the prediction confidence index, and generate a fault type diagnosis report.
[0083] The comprehensive spatiotemporal feature vector is input into a multilayer perceptron classification network for forward propagation calculation, and the original classification score is output, expressed as: ; in, Indicates the first The original classification score for each fault type An index variable representing the fault type. No. Trainable weight column vectors for each fault type Represents the transpose of a vector. Represents the comprehensive spatiotemporal feature vector. Indicates the first Trainable bias scalars for each fault type Indicates an adjustable scaling factor. This represents the hyperbolic tangent activation function. This represents the index variable used to traverse each radial basis function unit. This represents the total number of radial basis function elements. Indicates the first The mixed weighting coefficients of the radial basis function units, This represents the natural exponential function. Indicates the first The center point vector of the radial basis functions Indicates the first The width parameter scalar of each radial basis function.
[0084] It should be noted that before inputting the comprehensive spatiotemporal feature vector into the multilayer perceptron classification network for forward propagation calculation, the comprehensive spatiotemporal feature vector was dimensionally unified through layer normalization. Layer normalization calculates the mean and variance of all feature dimensions for each sample in the comprehensive spatiotemporal feature vector, and uses the mean and variance to standardize the feature values of each dimension, so that all features have zero mean and unit variance, thereby eliminating the influence of differences in physical dimensions or numerical scales of different features and achieving dimension unification.
[0085] It should also be noted that the integrated spatiotemporal feature vector is a unified feature representation that integrates temporal evolution and spatial distribution information. It includes the instantaneous amplitude-frequency features, the energy of the dominant mode, frequency trends and other temporal characteristics extracted from multi-scale spatiotemporal feature sequences, as well as spatial characteristics such as spatial location and structural consistency in visible light images, thermal infrared images and ultraviolet spectral images. After dimensionality reduction, it forms a vector of fixed length.
[0086] The adjustable scaling factor is a parameter learned along with the offset term in the Platt scaling algorithm by fitting a sigmoid function on a labeled validation set. It is used to control the steepness of the mapping from the original classification score to the probability.
[0087] The mixed weighting coefficient is a preset parameter used to weight the information entropy term, the maximum probability term, and the average deviation term when calculating the confidence score. It is determined by using a grid search or optimization algorithm on a labeled validation set with the goal of optimizing the consistency between confidence and diagnostic accuracy.
[0088] The specific process includes feeding the integrated spatiotemporal feature vector as the initial input into the first fully connected layer of the multilayer perceptron classification network. The first fully connected layer performs a linear weighted summation of the input features and combines it with a bias term. The output of the first layer is generated through a nonlinear activation function (such as ReLU). This output continues to be used as the input of the next fully connected layer, and the linear transformation and nonlinear activation are repeated layer by layer until the output layer is reached. The number of neurons in the output layer is consistent with the number of fault types. Each neuron corresponds to one fault type, and the output value is the original classification score of the fault type. The original classification score has not yet been normalized and only reflects the preliminary discrimination response of the multilayer perceptron classification network to various fault types.
[0089] Furthermore, the training process of the multilayer perceptron classification network is carried out on a labeled training dataset, which consists of a comprehensive spatiotemporal feature vector and corresponding fault type labels. At the beginning of training, the comprehensive spatiotemporal feature vector is input into the multilayer perceptron classification network, and the original classification score is obtained through forward propagation. Then, it is calibrated to the probability distribution of fault types by the Platt scaling algorithm. The error is calculated based on the cross-entropy loss between the probability distribution of fault types and the true fault type labels, and the gradient is calculated layer by layer using the backpropagation algorithm. The weights and bias parameters of each fully connected layer in the multilayer perceptron classification network are updated by the optimizer. This process is iterated repeatedly on the training dataset until the loss function converges or reaches the preset number of training rounds, and finally a multilayer perceptron classification network that can accurately map the comprehensive spatiotemporal feature vector to the fault type probability distribution is obtained.
[0090] Based on the Platt scaling algorithm, logistic regression fitting and calibration are performed on the original classification scores to obtain the probability distribution of fault types.
[0091] The specific process includes taking the original classification score output by the multilayer perceptron classification network as input, using the Platt scaling algorithm to perform nonlinear mapping on the original classification score through separately trained logistic regression, with the logistic regression using the sigmoid function as the link function to compress the original classification score to between zero and one, and then normalizing the output values corresponding to all fault types to make their sum equal to one, thereby obtaining the probability distribution of fault types that conforms to the probabilistic interpretation.
[0092] It should be noted that the Platt scaling algorithm is a method for calibrating classifier output scores to obtain probability estimates. By fitting a sigmoid function to the validation data, the original classification scores are mapped to a range of 0 to 1 and then normalized to form a probability distribution of fault types that conforms to probabilistic interpretation.
[0093] The prediction confidence index is obtained based on the probability distribution of fault types, and the confidence score is calculated using the following expression: ; in, This represents the confidence score. The weighting coefficients of the information entropy term. This indicates the total number of fault types. Indicates the first The probability value of each fault type This represents the weighting coefficient of the term with the highest probability. This represents the weighting coefficient of the average deviation term.
[0094] It should be noted that the weighting coefficients of the information entropy term are preset parameters determined by using a grid search or optimization algorithm on a labeled validation set with the goal of ensuring consistency between the predicted confidence and the actual diagnostic results.
[0095] The weighting coefficient of the highest probability term is a preset parameter determined by using a grid search or optimization algorithm on a labeled validation set with the goal of improving the consistency between the confidence score and the actual diagnostic accuracy.
[0096] The weighting coefficient of the average deviation term is a preset parameter determined by using a grid search or optimization algorithm on a labeled validation set with the goal of optimizing the ability of the confidence score to judge the reliability of the fault diagnosis results.
[0097] Indicates the first The probability value for each fault type is obtained by normalizing the original classification score output by the multilayer perceptron classification network after applying the Platt scaling algorithm.
[0098] The specific process includes: using the probability values of each category in the probability distribution of fault types to calculate the information of the probability distribution of fault types to measure the uncertainty of the prediction results; extracting the maximum probability value in the probability distribution of fault types to reflect the confidence level of the dominant fault category; calculating the average absolute deviation between each probability value in the probability distribution of fault types and the uniform distribution to assess the overall discriminative power; and weighting and summing the information entropy, the maximum probability value, and the average deviation with the corresponding weight coefficients to obtain a scalar confidence score.
[0099] The confidence score and diagnostic confidence threshold are used to perform decision analysis on the probability distribution of fault types, and a fault type diagnostic report is generated.
[0100] The specific process includes comparing the confidence score with a pre-set diagnostic confidence threshold. If the confidence score is greater than or equal to the diagnostic confidence threshold, the category with the highest probability in the probability distribution of the fault type is adopted as the final diagnostic result, and the fault type, its corresponding probability, and confidence score are output in the fault type diagnostic report. If the confidence score is less than the diagnostic confidence threshold, the current diagnostic result is determined to be unreliable, and the diagnostic confidence is marked as insufficient in the fault type diagnostic report, with further testing or verification recommended.
[0101] It should be noted that the diagnostic confidence threshold is preset based on historical data of power equipment fault diagnosis and actual operation and maintenance needs. It is determined by analyzing the relationship curve between confidence score and diagnostic accuracy on a labeled validation set and weighing the false alarm rate and the false negative rate. The specific value range is usually set between 0.3 and 0.95 to ensure that in power scenarios with high reliability requirements, only high confidence prediction results are clearly diagnosed.
[0102] This embodiment also provides an online diagnostic device for multispectral fusion faults in power equipment, comprising: a data acquisition module, a feature fusion module, a sequence construction module, a feature generation module, and an intelligent diagnostic module. The data acquisition module acquires visible light images, thermal infrared images, and ultraviolet spectral images of the power equipment, performs registration processing, and generates a multispectral image data cube. The feature fusion module inputs the multispectral image data cube into a dynamic kernel generation network, outputs a multimodal feature map set, and performs a dual-path attention mechanism on the multimodal feature map set to generate a fusion feature map focusing on key regions. The sequence construction module performs wavelet packet multi-scale decomposition on the cross-modal fusion feature map and constructs a second-level transient analysis sequence and a minute-level steady-state analysis sequence to form a multi-scale spatiotemporal feature sequence characterizing the fault evolution law. The feature generation module performs noise-assisted decomposition and Hilbert spectral analysis on the multi-scale spatiotemporal feature sequence, and generates a comprehensive spatiotemporal feature vector through a tensor chain decomposition algorithm. The intelligent diagnostic module inputs the comprehensive spatiotemporal feature vector into a multilayer perceptron classification network, outputs the probability distribution of fault types, calculates the prediction confidence index, and generates a fault type diagnostic report.
[0103] In summary, this invention achieves a transformation from simple superposition to deep interactive fusion of multispectral image features through a dynamic kernel generation network and a dual-path attention mechanism. This solves the problems of existing image detection methods, such as shallow mining of complex nonlinear correlations between visible light, thermal infrared, and ultraviolet images, and poor robustness of fused features under complex conditions such as occlusion and interference. By decomposing wavelet packets at multiple scales and constructing second-level transient and minute-level steady-state analysis sequences, this invention achieves a leap from single-frame static detection of fault signals in multispectral images to dynamic full-process tracking across time series. This can accurately characterize the evolution law of the entire life cycle of faults, improve the early warning capability of image-based detection for faults and latent defects, and enhance the reliability of diagnostic results.
[0104] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. An online diagnostic method for multispectral fusion faults in power equipment, characterized in that: include, The system acquires visible light, thermal infrared, and ultraviolet spectral images of power equipment, performs registration processing, and generates a multispectral image data cube. A cube of multispectral image data is input into a dynamic kernel generation network, which outputs a multimodal feature map set. A dual-path attention mechanism is then applied to the multimodal feature map set to generate a fused feature map that focuses on key regions. Wavelet packet multi-scale decomposition is performed on the cross-modal fusion feature map, and second-level transient analysis sequence and minute-level steady-state analysis sequence are constructed to form a multi-scale spatiotemporal feature sequence characterizing the fault evolution law; Noise-assisted decomposition and Hilbert spectral analysis are performed on multi-scale spatiotemporal feature sequences, and a comprehensive spatiotemporal feature vector is generated by tensor chain decomposition algorithm. The comprehensive spatiotemporal feature vector is input into the multilayer perceptron classification network, the probability distribution of fault types is output, the prediction confidence index is calculated, and a fault type diagnosis report is generated.
2. The online diagnostic method for multispectral fusion faults in power equipment as described in claim 1, characterized in that: The specific steps for generating the multispectral image data cube are as follows. Visible light images, thermal infrared images, and ultraviolet spectral images of power equipment are acquired, and non-uniformity correction and point spread function compensation processing are performed to obtain enhanced multispectral images. The enhanced multispectral image is segmented to generate a set of superpixel blocks; The matching cost of superpixel blocks between different spectral images in the superpixel block set is calculated by a multi-dimensional feature weighted fusion algorithm, and a superpixel-level correspondence is established. Control point pairs are determined based on superpixel-level correspondence, and pixel coordinate transformation relationships from thermal infrared images and ultraviolet spectral images to visible light images are established based on the control point pairs. By utilizing pixel coordinate transformation relationships, thermal infrared and ultraviolet spectral images are resampled and spatially aligned with visible light images to generate a multispectral image data cube.
3. The online diagnostic method for multispectral fusion faults in power equipment as described in claim 2, characterized in that: The specific steps for outputting the multimodal feature map set are as follows. Spectral dimension decoupling and edge enhancement processing are performed on the multispectral image data cube to generate an enhanced multispectral feature map; Based on the enhanced multispectral feature map, cross-modal association weights are obtained through a dynamic kernel generation network, and feature extraction is performed on the enhanced multispectral feature map to obtain the convolutionally enhanced feature map; Interactive enhancement of the feature maps extracted by convolution is performed through a cross-modal attention mechanism to generate attention-weighted multimodal feature maps; Attention-weighted feature maps are concatenated along the channel dimension and normalized by layers to generate a multimodal feature map set.
4. The online diagnostic method for multispectral fusion faults in power equipment as described in claim 3, characterized in that: The specific steps for generating the fused feature map focusing on the key regions are as follows. Tensor transformation and dimensionality reduction are performed on the multimodal feature map to obtain the core feature tensor; Based on the core feature tensor, a fusion weight matrix is obtained through a dual-path attention mechanism; The core feature tensors are filtered for importance and weighted by a fusion weight matrix to obtain the enhanced fusion features; The enhanced fusion features are refined and gating at multiple scales to generate a fusion feature map that focuses on key regions.
5. The online diagnostic method for multispectral fusion faults in power equipment as described in claim 4, characterized in that: The specific steps for forming a multi-scale spatiotemporal feature sequence characterizing the evolution of faults are as follows: Wavelet packet multi-scale decomposition is performed on the cross-modal fusion feature map to obtain multi-scale time-frequency feature sub-bands; Dynamic modal analysis was performed on multi-scale time-frequency characteristic subbands, and the dominant modal features of fault evolution were extracted. Construct second-level transient analysis sequences and minute-level steady-state analysis sequences based on dominant mode characteristics; Multi-scale fusion of transient and steady-state analysis sequences generates a multi-scale spatiotemporal feature sequence that characterizes the evolution of faults.
6. The online diagnostic method for multispectral fusion faults in power equipment as described in claim 5, characterized in that: The specific steps for performing noise-assisted decomposition and Hilbert spectral analysis on the multi-scale spatiotemporal feature sequences are as follows. The intrinsic mode function components are obtained by performing noise-assisted decomposition on multi-scale spatiotemporal feature sequences using the ensemble empirical mode decomposition method. Hilbert spectral analysis was performed on the intrinsic mode function components, and instantaneous amplitude-frequency characteristics were extracted.
7. The online diagnostic method for multispectral fusion faults in power equipment as described in claim 6, characterized in that: The specific steps for generating the comprehensive spatiotemporal feature vector are as follows: Based on instantaneous amplitude-frequency characteristics, feature fusion is performed using the tensor chain decomposition algorithm to obtain fused feature tensors; Multi-resolution optimization and dimensionality reduction are performed on the fused feature tensor to generate a comprehensive spatiotemporal feature vector.
8. The online diagnostic method for multispectral fusion faults in power equipment as described in claim 7, characterized in that: The probability distribution of the output fault type is determined through the following steps. The integrated spatiotemporal feature vector is input into the multilayer perceptron classification network for forward propagation calculation, and the original classification score is output. Based on the Platt scaling algorithm, logistic regression fitting and calibration are performed on the original classification scores to obtain the probability distribution of fault types.
9. The online diagnostic method for multispectral fusion faults in power equipment as described in claim 8, characterized in that: The specific steps for generating the fault type diagnostic report are as follows: The prediction confidence index is obtained based on the probability distribution of fault types, and the confidence score is calculated. The confidence score and diagnostic confidence threshold are used to perform decision analysis on the probability distribution of fault types, and a fault type diagnostic report is generated.
10. An online diagnostic device for multispectral fusion faults in power equipment, based on the online diagnostic method for multispectral fusion faults in power equipment according to any one of claims 1 to 9, characterized in that: It includes a data acquisition module, a feature fusion module, a sequence construction module, a feature generation module, and an intelligent diagnosis module. The data acquisition module is used to acquire visible light images, thermal infrared images, and ultraviolet spectral images of power equipment, and perform registration processing to generate a multispectral image data cube. The feature fusion module is used to input multispectral image data cubes into a dynamic kernel generation network, output a multimodal feature map set, and perform a dual-path attention mechanism on the multimodal feature map set to generate a fused feature map focusing on key regions. The sequence construction module is used to perform wavelet packet multi-scale decomposition on cross-modal fusion feature maps and construct second-level transient analysis sequences and minute-level steady-state analysis sequences to form multi-scale spatiotemporal feature sequences that characterize the evolution of faults. The feature generation module is used to perform noise-assisted decomposition and Hilbert spectral analysis on multi-scale spatiotemporal feature sequences, and to generate comprehensive spatiotemporal feature vectors through the tensor chain decomposition algorithm. The intelligent diagnostic module is used to input the comprehensive spatiotemporal feature vector into the multilayer perceptron classification network, output the probability distribution of fault types, calculate the prediction confidence index, and generate a fault type diagnostic report.