Ultraviolet imaging and acoustic imaging fusion fault classification method and system based on deep learning
By using a deep learning-based dual-channel feature encoding network to extract features and perform heterogeneous alignment on ultraviolet and acoustic imaging data, the problem of heterogeneous feature alignment in ultraviolet and acoustic imaging data is solved, thereby improving the accuracy and robustness of fault classification.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-23
- Publication Date
- 2026-03-31
AI Technical Summary
Existing methods fail to adequately address the issue of misalignment of heterogeneous features between ultraviolet imaging and acoustic imaging data due to differences in physical properties, thus limiting the accuracy of fault classification.
A deep learning-based dual-channel feature encoding network is adopted. By optimizing model parameters through feature extraction and heterogeneous alignment error calculation, cross-modal feature space alignment and fusion decision error constraints are achieved, thereby improving the accuracy of fault classification.
By using multimodal feature hierarchical extraction and heterogeneous feature alignment, the accuracy and robustness of fault classification are effectively improved, making it suitable for fault detection in complex scenarios.
Smart Images

Figure CN121010801B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and more specifically, to a fault classification method and system based on deep learning that fuses ultraviolet imaging and acoustic imaging. Background Technology
[0002] In the field of fault detection for critical systems such as industrial equipment and energy facilities, ultraviolet imaging technology can capture abnormal ultraviolet signals such as partial discharge, while acoustic imaging technology can acquire acoustic wave features such as abnormal vibrations. However, existing methods often directly stitch together or simply fuse shallow features of the two types of data, failing to fully address the alignment problem of heterogeneous features caused by the differences in physical properties between ultraviolet and acoustic imaging data. This results in insufficient utilization of complementary information and limited accuracy in fault classification. Summary of the Invention
[0003] The purpose of this invention is to provide a fault classification method and system based on deep learning that fuses ultraviolet imaging and acoustic imaging.
[0004] In a first aspect, embodiments of the present invention provide a fault classification method based on the fusion of ultraviolet imaging and acoustic imaging using deep learning, comprising:
[0005] Acquire an imaging data set consisting of device ultraviolet imaging data and device acoustic imaging data;
[0006] A dual-channel feature encoding network is used to perform feature extraction operations on the device ultraviolet imaging data and device acoustic imaging data in the imaging data group to obtain the basic ultraviolet imaging features and enhanced ultraviolet imaging features corresponding to the device ultraviolet imaging data, as well as the basic acoustic imaging features and enhanced acoustic imaging features corresponding to the device acoustic imaging data.
[0007] Using the basic ultraviolet imaging feature as the first heterogeneous alignment reference point, a first heterogeneous alignment error value is determined based on the first heterogeneous alignment reference point and the enhanced acoustic imaging feature; and using the basic acoustic imaging feature as the second heterogeneous alignment reference point, a second heterogeneous alignment error value is determined based on the second heterogeneous alignment reference point and the enhanced ultraviolet imaging feature.
[0008] The fusion decision error value is determined based on the basic ultraviolet imaging features, the enhanced ultraviolet imaging features, the basic acoustic imaging features, and the enhanced acoustic imaging features;
[0009] The model parameters of the dual-channel feature coding network are optimized based on the fusion decision error value, the first heterogeneous alignment error value, and the second heterogeneous alignment error value to obtain the target dual-channel feature coding network.
[0010] The target dual-channel feature coding network is used to determine the fault in the equipment monitoring data; the equipment monitoring data includes the target equipment's ultraviolet imaging data and the target equipment's acoustic imaging data.
[0011] In a second aspect, embodiments of the present invention provide a server system, including a server, the server being used to execute the method described in the first aspect.
[0012] Compared to existing technologies, the beneficial effects of this invention include: It discloses a deep learning-based method and system for fault classification by fusing ultraviolet (UV) imaging and acoustic imaging. This involves acquiring imaging data sets composed of UV and acoustic imaging data from the equipment; extracting basic and enhanced features from the two types of data using a dual-channel feature coding network; calculating the heterogeneous alignment error value using the basic UV and acoustic imaging features as reference points to achieve cross-modal feature space alignment; determining the fusion decision error value by combining the basic and enhanced features; optimizing network parameters to obtain the target dual-channel feature coding network; and finally, using the target network to determine faults in the equipment monitoring data. This invention effectively improves the accuracy and robustness of fault classification through multimodal feature layered extraction, heterogeneous feature alignment, and fusion decision error constraints, making it suitable for fault detection in complex scenarios such as power equipment. Attached Figure Description
[0013] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of the present invention and should not be considered as limiting the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0014] Figure 1 A flowchart illustrating the steps of the deep learning-based ultraviolet imaging and acoustic imaging fusion fault classification method provided in this embodiment of the invention;
[0015] Figure 2 A schematic block diagram of the structure of a computer device provided in an embodiment of the present invention. Detailed Implementation
[0016] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0017] The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0018] In order to solve the technical problems mentioned in the background art Figure 1 This is a flowchart illustrating the deep learning-based ultraviolet imaging and acoustic imaging fusion fault classification method provided in this embodiment. The following is a detailed description of the deep learning-based ultraviolet imaging and acoustic imaging fusion fault classification method.
[0019] Step S201: Obtain an imaging data set consisting of device ultraviolet imaging data and device acoustic imaging data;
[0020] Step S202: Perform feature extraction operation on the device ultraviolet imaging data and device acoustic imaging data in the imaging data group through a dual-channel feature coding network to obtain the basic ultraviolet imaging features and enhanced ultraviolet imaging features corresponding to the device ultraviolet imaging data, as well as the basic acoustic imaging features and enhanced acoustic imaging features corresponding to the device acoustic imaging data.
[0021] Step S203: Using the basic ultraviolet imaging feature as the first heterogeneous alignment reference point, determine the first heterogeneous alignment error value based on the first heterogeneous alignment reference point and the enhanced acoustic imaging feature; and using the basic acoustic imaging feature as the second heterogeneous alignment reference point, determine the second heterogeneous alignment error value based on the second heterogeneous alignment reference point and the enhanced ultraviolet imaging feature.
[0022] Step S204: Determine the fusion decision error value based on the basic ultraviolet imaging features, the enhanced ultraviolet imaging features, the basic acoustic imaging features, and the enhanced acoustic imaging features;
[0023] Step S205: Optimize the model parameters of the dual-channel feature coding network based on the fusion decision error value, the first heterogeneous alignment error value, and the second heterogeneous alignment error value to obtain the target dual-channel feature coding network;
[0024] Step S206: Fault judgment is performed on the device monitoring data through the target dual-channel feature coding network; the device monitoring data includes the target device ultraviolet imaging data and the target device acoustic imaging data.
[0025] In this embodiment of the invention, exemplarily, the following describes in detail each step of the method of the present invention in the context of a partial discharge fault detection scenario in power equipment (taking a 110kV insulator in a substation as an example). The server, as the executing entity, integrates an ultraviolet imager, an acoustic imager, and a deep learning computing module to achieve accurate classification of insulator faults.
[0026] The server first receives equipment monitoring data synchronously collected by the ultraviolet imager and the acoustic imager via a data interface. In the specific scenario, the detection objects are 10 sets of 110kV porcelain insulators (numbered #1-#10) within the substation, with each set corresponding to one complete detection cycle. The ultraviolet imager uses a solar-blind ultraviolet camera to collect ultraviolet image data with a resolution of 640×512 pixels and a sampling frame rate of 25fps, recording radiation characteristics such as ultraviolet photon count, spot area, and spot center coordinates. The acoustic imager uses a 32-channel microphone array to collect acoustic signals with a sampling rate of 48kHz and a frequency band of 20Hz-40kHz, recording the spatial distribution of sound wave intensity (acoustic image) and time-domain waveform data. To ensure data correlation, the server uses the timestamp (accuracy ±1ms) of the imager's built-in GPS module to align the ultraviolet imaging data (each frame corresponding to 1 / 25s) with the acoustic imaging data (generating one acoustic image frame every 10ms) along the time axis, forming 10 sets of "ultraviolet-acoustic imaging data sets" (denoted as G1 to G...). 10 Each group contains synchronized ultraviolet and acoustic sequence data. For example, data group G1 corresponds to the ultraviolet imaging data (500 frames, including photon number matrix and spot coordinates) and acoustic imaging data (1000 frames of time-frequency graph, including 256×256 acoustic intensity distribution) of insulator #1.
[0027] Subsequently, the server invokes a pre-built dual-channel feature encoding network to perform feature extraction on the ultraviolet and acoustic imaging data in the imaging data set, outputting basic ultraviolet imaging features, enhanced ultraviolet imaging features, basic acoustic imaging features, and enhanced acoustic imaging features. For the ultraviolet imaging data, the server first performs radiation feature enhancement: the basic radiation characterization data is filtered by Gaussian to suppress background noise and retain the photon number distribution reflecting the corona discharge intensity on the insulator surface; the enhanced radiation characterization data is enhanced by a multi-scale radiation enhancement algorithm to highlight the gradient at the edge of the light spot (the boundary of the discharge region), and a cumulative photon number histogram feature is superimposed (statistically analyzing the time-series changes in the area of the light spot over 500 frames). Next, the ultraviolet spatial feature extraction component (composed of a 3-layer convolutional neural network) extracts the spatial texture features of the local discharge region (such as the shape and density distribution of the light spot) from the basic radiation characterization, and outputs the basic ultraviolet imaging features through global average pooling; the ultraviolet enhancement feature extraction component (using a residual network structure) captures multi-scale radiation features (such as the spatiotemporal variation of discharge intensity and the dynamic expansion trend of the light spot) from the enhanced radiation characterization, and outputs the enhanced ultraviolet imaging features.
[0028] For acoustic imaging data, the server first segments the raw acoustic signal based on a 50ms sliding time window (25ms step size), and converts it into a time-frequency frame sequence (each frame is a time-spectrum graph containing energy distribution in frequency and time dimensions) through a short-time Fourier transform. The acoustic time-frequency feature extraction component (composed of a convolutional time-frequency feature extractor, a temporal compression layer, and a global aggregation layer) extracts local time-frequency features (such as characteristic frequency segments of discharge acoustic waves) from the time-frequency frame sequence. The temporal compression layer (using bidirectional LSTM) aggregates the time-period features, and then the global aggregation layer (attention mechanism) assigns weights to the features of different time periods (with higher weights for active discharge periods), outputting basic acoustic imaging features. The acoustic enhancement feature extraction component (using a Transformer encoder) performs long-term temporal dependency modeling on the time-frequency frame sequence to capture the non-stationary characteristics of discharge acoustic waves (such as distinguishing between burst pulses and continuous noise), outputting enhanced acoustic imaging features.
[0029] Next, the server calculates the heterogeneous alignment error and fusion decision error to optimize the network. In the heterogeneous alignment error calculation, the first heterogeneous alignment error uses the basic UV imaging features as the reference point: for each data set, the basic UV imaging features of that set are used as the reference, the enhanced acoustic imaging features within the same data set are used as "homogeneous related samples" (should be similar to the reference point), and the enhanced acoustic imaging features of other data sets are used as "heterogeneous related samples" (should differ significantly from the reference point). By comparing the similarity between these samples and the reference point, the first heterogeneous alignment error value is calculated (the goal is to make the similarity between homogeneous samples and the reference point higher than that of heterogeneous samples). Similarly, the second heterogeneous alignment error uses the basic acoustic imaging features as the reference point, the enhanced UV imaging features within the same data set as homogeneous related samples, and the enhanced UV imaging features of other data sets as heterogeneous related samples, to calculate the second heterogeneous alignment error value.
[0030] In the calculation of fusion decision error, the server first fuses enhanced ultraviolet imaging features and enhanced acoustic imaging features into multi-source joint features (achieved through element-weighted multiplication, with weights dynamically adjusted based on feature importance). This dynamic weight adjustment is implemented through a modal attention mechanism: the server assigns attention weights (ω) to the ultraviolet and acoustic modes respectively. U ω A ω U +ω A =1), and the weight values are obtained through training. For example, for corona discharge faults, the ultraviolet spatial features and acoustic temporal features have a high correlation, and ω is calculated during training. U =0.55、ω A =0.45 (highlighting ultraviolet spatial distribution); for surface leakage faults, the acoustic continuous noise characteristics are more critical, ω U =0.3、ω A=0.7 (highlighting acoustic frequency band characteristics). The first fusion decision error uses the basic ultraviolet imaging features as "single-source feature anchors": For each data set, the basic ultraviolet imaging features of that set are used as anchors, the multi-source joint features of the same data set are used as "intrinsic joint representation samples" (which should be semantically consistent with the anchors), and the multi-source joint features of other data sets are used as "external joint representation samples" (which should have significantly different semantics from the anchors). The first fusion decision error value is calculated by comparing the semantic consistency between the samples and the anchors. The second fusion decision error is similarly calculated by using the basic acoustic imaging features as anchors, comparing the multi-source joint features of the same data set with those of other data sets, and finally, the fusion decision error value is the sum of the two.
[0031] The server uses a weighted sum of the decision error, the first heterogeneous alignment error, and the second heterogeneous alignment error as the total loss function. It then optimizes the model parameters of the dual-channel feature encoding network (such as convolutional kernel weights, LSTM hidden layer parameters, and attention weights) through backpropagation. During training, the server proportionally divides the data into training and validation sets, iteratively updating the network parameters until the validation set loss converges. The converged network parameters are then saved as the target dual-channel feature encoding network.
[0032] Finally, the server deploys a target dual-channel feature encoding network to determine faults based on real-time equipment monitoring data. For the insulator fault classification task, the server selects a "fault classifier" as the task adapter head module (composed of a fully connected layer and a Softmax activation function, outputting fault category probabilities). After receiving the ultraviolet imaging data and acoustic imaging data from the target equipment, the encoder of the target network extracts four types of features and fuses them into multi-source joint features. These features are then input into the fault classifier to obtain the probability distribution of various fault types (such as corona discharge, surface leakage, and normal state). The category with the highest probability is output as the fault determination result and pushed to the substation monitoring system.
[0033] This embodiment achieves high-precision classification of power equipment faults through multimodal feature fusion and cross-modal alignment. In practical applications, it can effectively improve the identification accuracy of faults such as partial discharge and mechanical loosening of insulators.
[0034] In this embodiment of the invention, the step of performing feature extraction operations on the device ultraviolet imaging data and device acoustic imaging data in the imaging data group through a dual-channel feature coding network to obtain the basic ultraviolet imaging features and enhanced ultraviolet imaging features corresponding to the device ultraviolet imaging data, as well as the basic acoustic imaging features and enhanced acoustic imaging features corresponding to the device acoustic imaging data, can be implemented through the following examples.
[0035] The ultraviolet spatial feature extraction component and ultraviolet enhancement feature extraction component of the dual-channel feature coding network are used to perform feature extraction operations on the device ultraviolet imaging data in the imaging data group, respectively, to obtain the basic ultraviolet imaging features and enhanced ultraviolet imaging features corresponding to the device ultraviolet imaging data.
[0036] The acoustic time-frequency feature extraction component and the acoustic enhancement feature extraction component of the dual-channel feature coding network are used to perform feature extraction operations on the device acoustic imaging data in the imaging data group, respectively, to obtain the basic acoustic imaging features and enhanced acoustic imaging features corresponding to the device acoustic imaging data.
[0037] In this embodiment of the invention, taking the ultraviolet imaging data of data group G1 (corresponding to insulator #1) as an example, the server first performs radiation feature enhancement on it to obtain basic radiation characterization data (a 640×512 photon density matrix after Gaussian filtering, which retains the photon distribution of the corona discharge region after suppressing background noise) and enhanced radiation characterization data (a light spot edge gradient map enhanced by the Retinex algorithm, superimposed with a time cumulative histogram of the light spot area of 500 frames, with dimensions of 640×512×3). Subsequently, the server calls the ultraviolet branch component of the dual-channel feature encoding network: the ultraviolet spatial feature extraction component: this component consists of 3 layers of convolutional neural networks (the first layer is a 3×3 convolutional kernel with 64 channels to extract local light spot texture, the second layer is a 5×5 convolutional kernel with 128 channels to aggregate mesoscale features, and the third layer is global average pooling). The server inputs basic radiometric characterization data into this component, which captures the spatial distribution features of the discharge region in the UV image of insulator #1 (such as the circularity of the light spot and density gradient) through convolution operations. This data is then compressed into a 256-dimensional vector using global average pooling, representing the basic UV imaging features of insulator #1 (containing basic spatial texture information of the discharge region, such as "the light spot is concentrated at the edge of the insulator skirt, with a density of 0.35 photons / pixel"). The UV enhancement feature extraction component employs a ResNet-18 residual network structure (containing 4 residual blocks, each containing 2 layers of 3×3 convolutions). The server inputs enhanced radiation characterization data into this component, avoids gradient vanishing through residual connections, and captures multi-scale radiation features: the low-level network extracts subtle gradients at the edge of the light spot (clarity of the discharge region boundary), and the high-level network aggregates the trend of light spot area change over time (such as the dynamic process of the light spot area increasing from 120 pixels² to 210 pixels² in 500 frames), finally outputting 512-dimensional enhanced ultraviolet imaging features (containing spatiotemporal dynamic information of discharge intensity, such as "the skirt discharge region expands over time, with an edge gradient change rate of 0.12 / frame"). For the acoustic imaging data of data group G1, the server first segments the original acoustic signal (time series with a sampling rate of 48kHz) based on a 50ms sliding time window (step size 25ms), and converts it into a 200-frame time-frequency frame sequence (each frame is a 256×256 time-frequency spectrum, with the frequency axis from 0-24kHz, the time axis corresponding to the 50ms window, and the pixel value being the acoustic energy density). Subsequently, the server invokes the acoustic branch component of the dual-channel feature encoding network: the acoustic time-frequency feature extraction component, which consists of a convolutional time-frequency feature extractor (2 layers of 3D convolution, 3×3×3 kernels, 64 / 128 channels), a temporal compression layer (bidirectional LSTM, 64 hidden units), and a global aggregation layer (attention mechanism).The server inputs the time-frequency frame sequence into this component. The convolutional layer extracts local time-frequency features (such as the energy peak in the 2-10kHz frequency band, corresponding to the characteristic frequency of insulator corona discharge); the temporal compression layer models the inter-frame dependencies using LSTM to capture the intermittency of the discharge sound waves (such as a pulse sound wave appearing every 1.2s); the global aggregation layer assigns attention weights to the 200 frame time-segment features (0.8 weight for active discharge periods and 0.2 weight for background noise periods), finally aggregating them into 256-dimensional basic acoustic imaging features (containing the basic time-frequency characteristics of the sound waves, such as "energy proportion of 2-10kHz frequency band 0.6, pulse interval 1.2s"). Acoustic enhancement feature extraction component: uses a 4-layer Transformer encoder (8-head self-attention). The server inputs the time-frequency frame sequence into the component, captures long-term time-series dependencies (such as the amplitude attenuation pattern of three consecutive pulse sound waves: the first pulse amplitude is 0.8mV, the second pulse is 0.6mV, and the last pulse is 0.4mV) through a self-attention mechanism, and distinguishes between discharge pulses and environmental noise (such as excluding interference from 50Hz power frequency noise) through multi-head attention, outputting 512-dimensional enhanced acoustic imaging features (containing fine-grained time-frequency dynamic information, such as "pulse sequence amplitude attenuation rate 0.2mV / pulse, noise suppression ratio 30dB"). Through the above processing, the server outputs the basic ultraviolet imaging features, enhanced ultraviolet imaging features, basic acoustic imaging features, and enhanced acoustic imaging features of insulator #1 for data group G1, and similarly completes the feature extraction for the remaining data groups.
[0038] In this embodiment of the invention, the following implementation methods may also be provided.
[0039] Radiation characteristic enhancement is performed on the ultraviolet imaging data of the device to obtain basic radiation characterization data and enhanced radiation characterization data corresponding to the ultraviolet imaging data of the device;
[0040] The ultraviolet spatial feature extraction component and ultraviolet enhancement feature extraction component of the dual-channel feature coding network perform feature extraction operations on the device ultraviolet imaging data in the imaging data group to obtain the basic ultraviolet imaging features and enhanced ultraviolet imaging features corresponding to the device ultraviolet imaging data, including:
[0041] The ultraviolet spatial feature extraction component of the dual-channel feature coding network performs feature extraction operation on the basic radiometric characterization data to obtain the basic ultraviolet imaging features corresponding to the device's ultraviolet imaging data.
[0042] The enhanced ultraviolet imaging features corresponding to the device's ultraviolet imaging data are obtained by performing feature extraction operations on the enhanced radiation characterization data through the ultraviolet enhancement feature extraction component of the dual-channel feature coding network.
[0043] In an embodiment of the present invention, taking the ultraviolet imaging data of substation #1 insulator (500 frames, each frame containing a 640×512 pixel photon number matrix and the center coordinates of the light spot) as an example, the server executes the following process: The server first performs radiation feature enhancement on the original ultraviolet imaging data of #1 insulator, and outputs basic radiation characterization data and enhanced radiation characterization data: Basic radiation characterization data: The server reads the photon number matrix of the 500 frames of original ultraviolet images (each pixel value represents the number of ultraviolet photons per unit area, ranging from 0 to 255), and suppresses background noise (such as scattered ultraviolet light from the sky and interference from reflections of metal parts of the tower) through Gaussian filtering (σ=1.5), while retaining the core radiation information of the corona discharge area on the surface of the insulator. For example, in the original image of #1 insulator frame 200, the average number of photons in the background region is 12, and the average number of photons in the discharge region (skirt on the upper surface of the insulator) is 85. After filtering, the number of photons in the background region is reduced to below 5, and the number of photons in the discharge region is retained to be 82±3, forming a basic radiation characterization data of 640×512×1 (single-channel matrix, reflecting the spatial distribution of discharge intensity). Enhanced radiative characterization data: The server performs multi-dimensional radiative feature enhancement on the original ultraviolet image: Spatial detail enhancement: The Retinex algorithm is used to separate the illumination component and reflection component of the image, highlighting the edge gradient of the discharge area (e.g., the grayscale difference between the discharge spot and the non-discharge area is increased from the original 20 to 45); Temporal dynamic enhancement: The temporal series change of the discharge spot area in 500 frames is statistically analyzed (e.g., the spot area is stable at 150±20 pixels² in frames 1-100, and increases to 280±30 pixels² in frames 101-500), and this time-cumulative feature (normalized to 0-1) is used as the third channel; Finally, 640×512×3 enhanced radiative characterization data is formed (three-channel matrix: channel 1 is the filtered photon number distribution, channel 2 is the edge gradient map, and channel 3 is the time-cumulative spot area). The server calls the ultraviolet branch component of the dual-channel feature encoding network to perform feature extraction on the basic radiometric characterization data and the enhanced radiometric characterization data respectively: Extraction of basic ultraviolet imaging features: The server inputs the basic radiometric characterization data (640×512×1) into the ultraviolet spatial feature extraction component. This component consists of three layers of convolutional neural networks: The first layer (3×3 convolutional kernel, 64 channels, ReLU activation): extracts local photon density texture (such as the circularity and edge smoothness of the discharge spot), and outputs a 640×512×64 feature map; The second layer (5×5 convolutional kernel, 128 channels, BatchNorm): aggregates mesoscale spatial features (such as the relative position of the spot and the insulator outline, the discharge spot of #1 insulator is concentrated at 3 / 4 of the upper surface skirt), and outputs a 320×256×128 feature map; The third layer (global average pooling): compresses the spatial dimension and outputs a 256-dimensional basic ultraviolet imaging feature vector (containing basic spatial distribution information of the discharge area of #1 insulator, such as "spot spatial concentration 0.72, density gradient 0.35 / pixel").Enhanced UV imaging feature extraction: The server inputs enhanced radiation characterization data (640×512×3) into the UV enhancement feature extraction component. This component adopts a ResNet-18 residual network structure (containing 4 residual blocks, each block containing 2 layers of 3×3 convolutions and skip connections): Low-level network (first 2 residual blocks): extracts multi-channel detail features (in the edge gradient map of channel 2, the edge sharpness of the discharge area is 0.85, and that of the non-discharge area is 0.21); High-level network (last 2 residual blocks): aggregates spatiotemporal dynamic features (the time-cumulative spot area display in channel 3 shows that the discharge intensity of insulator #1 increases linearly over time, with a growth rate of 0.25 pixels² / frame); Finally, a 512-dimensional enhanced UV imaging feature vector is output through global max pooling (containing spatiotemporal dynamic information of discharge intensity, such as "the skirt discharge area expands over time, with an edge gradient change rate of 0.18 / frame"). Through the above processing, the server extracts 256-dimensional basic ultraviolet imaging features (spatial static features) and 512-dimensional enhanced ultraviolet imaging features (spatiotemporal dynamic features) from the ultraviolet imaging data of insulator #1, providing a foundation for subsequent cross-modal feature alignment and fusion.
[0044] In this embodiment of the invention, the following implementation methods are also provided.
[0045] The acoustic imaging data of the device is segmented into time-frequency frame sequences based on time windows;
[0046] The acoustic time-frequency feature extraction component and acoustic enhancement feature extraction component of the dual-channel feature coding network perform feature extraction operations on the device acoustic imaging data in the imaging data group to obtain the basic acoustic imaging features and enhanced acoustic imaging features corresponding to the device acoustic imaging data, including:
[0047] The acoustic time-frequency feature extraction component of the dual-channel feature coding network performs feature extraction operation on the time-frequency frame sequence to obtain the basic acoustic imaging features corresponding to the device acoustic imaging data.
[0048] The acoustic enhancement feature extraction component of the dual-channel feature coding network performs feature extraction on the time-frequency frame sequence to obtain the enhanced acoustic imaging features corresponding to the device acoustic imaging data.
[0049] In an embodiment of the present invention, taking the acoustic imaging data of insulator #1 of a substation (sampling rate 48kHz, total duration 10 minutes of acoustic time series, including corona discharge pulses and environmental noise) as an example, the server executes the following process: The server reads the original acoustic imaging data of insulator #1 (single-channel acoustic time series, total number of sampling points = 48000Hz × 600s = 28,800,000 points, voltage amplitude range -10mV to 10mV), and segments it based on a 50ms sliding time window (window length = 48000 × 0.05 = 2400 sampling points) and a 25ms step size (step size = 48000 × 0.025 = 1200 sampling points), resulting in a total of (28,800,000 - 2400) / 1200 + 1 = 23999 frames of segmented signals. A short-time Fourier transform (STFT, Hann window, window length 1024 points, overlap rate 50%) is performed on each signal segment to convert the time-domain signal into a 256×256 time-frequency spectrum (frequency axis 0-24kHz, energy distribution within a 50ms window on the time axis, pixel value is normalized energy density 0-1), finally forming a 23999×256×256×1 time-frequency frame sequence (each frame corresponds to the acoustic time-frequency characteristics of a 50ms time period, such as the 1000th frame time-frequency spectrum of insulator #1, where the energy density of the 2-10kHz frequency band reaches 0.7, corresponding to the corona discharge characteristic frequency). The server invokes the acoustic branch component of the dual-channel feature encoding network to perform feature extraction on the time-frequency frame sequence: Extraction of basic acoustic imaging features: The server inputs the time-frequency frame sequence into the acoustic time-frequency feature extraction component (consisting of a convolutional time-frequency feature extractor, a temporal compression layer, and a global aggregation layer): Convolutional time-frequency feature extractor: 2-layer 3D convolution (first layer 3×3×3 convolution kernel 64 channels, ReLU activation; second layer 3×3×3 convolution kernel 128 channels, BatchNorm), extracting local time-frequency features from the time-frequency frame sequence. For example, for the time-frequency frame sequence of insulator #1, the first convolutional layer captures the energy peak (discharge characteristic frequency) in the 2-10kHz frequency band, and the second convolutional layer aggregates the time-frequency correlation of three adjacent frames (such as the energy change between frames before and after the discharge pulse), outputting a frame-level feature of 23999×128×128×128; the temporal compression layer: the bidirectional LSTM (64 hidden units) performs temporal dependency modeling on the frame-level features, capturing the intermittent regularity of the discharge sound wave (such as the discharge pulse interval of insulator #1 is about 1.2s, corresponding to an energy peak occurring once every 48 frames), outputting a time-period feature of 23999×64; the global aggregation layer: the attention mechanism assigns weights to the 23999 time-period features (0.85 weight for active discharge periods and 0.15 weight for background noise periods), and after weighted aggregation, outputs a 256-dimensional basic acoustic imaging feature (including the basic time-frequency characteristics of the sound wave, such as "energy proportion of 2-10kHz frequency band 0.62, pulse interval 1.2s±0.1s").Enhanced acoustic imaging feature extraction: The server inputs the time-frequency frame sequence into the acoustic enhancement feature extraction component (4-layer Transformer encoder, 8-head self-attention): The encoder calculates the correlation weight of any two frames of time-frequency features through the self-attention mechanism, capturing long-term temporal dependencies (e.g., in frame 1000-1048 of #1 insulator (corresponding to 1.2s), the energy density of the first pulse frame (frame 1000) is 0.82, the energy density of the second pulse frame (frame 1048) is 0.75, and the energy density of the last pulse frame (frame 1096) is 0.68, forming a "decaying pulse sequence" feature); the multi-head attention focuses on the temporal changes of different frequency bands (2-5kHz, 5-10kHz) respectively, and finally outputs 512-dimensional enhanced acoustic imaging features through the linear layer (including fine-grained time-frequency dynamics, such as "pulse sequence amplitude decay rate 0.12 / pulse, the decay rate of the 5-10kHz band is faster than that of 2-5kHz"). Through the above processing, the server extracts 256-dimensional basic acoustic imaging features (reflecting core time-frequency characteristics) and 512-dimensional enhanced acoustic imaging features (reflecting long-term dynamic laws) from the acoustic imaging data of #1 insulator, providing acoustic modality discriminative information for subsequent cross-modal feature alignment and fusion.
[0050] In this embodiment of the invention, the acoustic time-frequency feature extraction component of the dual-channel feature coding network includes a convolutional time-frequency feature extractor, a temporal compression layer, and a global aggregation layer; the feature extraction operation performed on the time-frequency frame sequence by the acoustic time-frequency feature extraction component of the dual-channel feature coding network to obtain the basic acoustic imaging features corresponding to the device acoustic imaging data can be implemented through the following example.
[0051] The convolutional time-frequency feature extractor performs feature extraction on the time-frequency frame sequence to obtain frame-level features;
[0052] The time-segmentation features are obtained by performing feature aggregation operations on the frame-level features through the temporal compression layer.
[0053] The basic acoustic imaging features corresponding to the device acoustic imaging data are obtained by performing feature aggregation operations on the time period features through the global aggregation layer.
[0054] In an embodiment of the present invention, taking, for example, the time-frequency frame sequence (23999 frames, each frame a 256×256×1 time-frequency spectrum, corresponding to 10 minutes of acoustic data) of the #1 insulator of a substation, the server performs the following feature extraction process through the acoustic time-frequency feature extraction component: The server inputs the time-frequency frame sequence (23999×256×256×1) into the convolutional time-frequency feature extractor (containing two 3D convolutional layers): First convolutional layer: 3×3×3 convolutional kernel (64 channels, ReLU activated), the dimension of the 3D convolutional kernel is defined as (frequency axis convolution) The kernel size is calculated as follows: kernel size × time axis convolution kernel size × frame sequence convolution kernel size. For example, in a 3×3×3 kernel, the first two dimensions (3×3) correspond to the frequency-time dimension of the time spectrum (such as the local 3×3 energy distribution in the 256×256 time spectrum). The third dimension (3) corresponds to the sequence association of three consecutive time-frequency frames (capturing the time-frequency dynamic changes in the short time domain). Convolution is performed on the local time-frequency region (3×3 frequency × 3×3 time window) of each frame's time spectrum to extract the energy peak features of the 2-10kHz frequency band (the core frequency band of #1 insulator discharge). For example, in the time spectrum of frames 1000-1002, the energy density of the 2-5kHz frequency band increases from 0.6 to 0.8. After convolution, a 64-channel feature map is output, marking this frequency band as the "discharge active area". The second layer of convolution: 3×3×3 convolution kernel (128 channels, BatchNorm) aggregates the mesoscale time-frequency association of three adjacent frames (such as the energy change gradient of the frames before and after the discharge pulse). The final output is a 23999×128×128×128 frame-level feature (each frame corresponds to a 128-dimensional local time-frequency feature, including the frequency energy distribution within a single frame and the correlation information between adjacent frames in the time domain). The server inputs the frame-level feature (23999×128×128×128) into a temporal compression layer (bidirectional LSTM, 64 hidden units): the LSTM performs temporal dependency modeling on the frame-level feature along the time axis (23999 frames) to capture the intermittent pattern of the discharge sound wave of #1 insulator. For example, an energy peak (discharge pulse) is detected every 48 frames (1.2s), with the energy gradually increasing in the 5 frames before the pulse (from 0.3 to 0.7), and rapidly decreasing in the 5 frames after the pulse (from 0.7 to 0.2); through forward / backward propagation of the bidirectional LSTM, the 23999 frame-level features are compressed into a 23999×64 time-segment feature (each frame corresponds to a 64-dimensional temporal aggregation feature, including the correlation between the discharge state of the current frame and the frames before and after).The server inputs the time-segment features (23999×64) into the global aggregation layer (attention mechanism): The attention mechanism calculates the weight distribution of the 23999 time-segment features: the weight of the time-segment feature during the active discharge period of #1 insulator (accounting for approximately 30% of the total duration) is set to 0.8 (e.g., frames 1000-1048), and the weight of the background noise period (70%) is set to 0.2; the time-segment features are weighted and summed, compressed into a 256-dimensional basic acoustic imaging feature vector, containing the core time-frequency characteristics of the #1 insulator sound wave: "energy proportion of 2-10kHz frequency band 0.65, discharge pulse interval 1.2s±0.1s, energy fluctuation amplitude of active period 0.3±0.05". Through the above process, the server extracts 256-dimensional basic acoustic imaging features from the time-frequency frame sequence, reflecting the basic time-frequency and temporal patterns of the device's acoustic signal.
[0055] In this embodiment of the invention, the number of imaging data groups is multiple; the step of using the basic ultraviolet imaging feature as the first heterogeneous alignment reference point and determining the first heterogeneous alignment error value based on the first heterogeneous alignment reference point and the enhanced acoustic imaging feature can be implemented through the following example.
[0056] For the device ultraviolet imaging data of each imaging data group, the basic ultraviolet imaging feature corresponding to the targeted device ultraviolet imaging data is determined as the first heterogeneous alignment reference point, the enhanced acoustic imaging feature corresponding to the device acoustic imaging data belonging to the same imaging data group is determined as the first homogeneous association sample, and the enhanced acoustic imaging feature corresponding to the device acoustic imaging data belonging to different imaging data groups is determined as the first heterogeneous association sample. The first heterogeneous alignment error value is determined based on the first heterogeneous alignment reference point, the first homogeneous association sample, and the first heterogeneous association sample.
[0057] The step of using the basic acoustic imaging features as the second heterogeneous source alignment reference point, and determining the second heterogeneous source alignment error value based on the second heterogeneous source alignment reference point and the enhanced ultraviolet imaging features, includes:
[0058] For the device acoustic imaging data of each imaging data group, the basic acoustic imaging feature corresponding to the targeted device acoustic imaging data is determined as the second heterogeneous alignment reference point, the enhanced ultraviolet imaging feature corresponding to the device ultraviolet imaging data belonging to the same imaging data group is determined as the second homogeneous association sample, and the enhanced acoustic imaging feature corresponding to the device ultraviolet imaging data belonging to different imaging data groups is determined as the second heterogeneous association sample. The second heterogeneous alignment error value is determined based on the second heterogeneous alignment reference point, the second homogeneous association sample, and the second heterogeneous association sample.
[0059] In an embodiment of the present invention, for example, 10 sets of insulator imaging data groups (G1 to G) from a substation are used.10 Taking insulators #1-#10 (each group containing synchronous ultraviolet and acoustic imaging data) as an example, the server determines the first heterogeneous alignment error value and the second heterogeneous alignment error value through the following process: For each imaging data group's device ultraviolet imaging data, the server performs error calculation of "reference point - same source sample - heterogeneous sample": Taking data group G1 (#1 insulator) as an example: First heterogeneous alignment reference point: The server sets the basic ultraviolet imaging feature F of insulator #1 as the reference point. Ubase1 This is designated as the reference point. F Ubase1 This is a 256-dimensional vector containing the basic spatial characteristics of the UV discharge of insulator #1 (e.g., "the light spot is concentrated on the upper surface skirt, spatial concentration is 0.72, and photon density is 0.35 / pixel"). First source-related sample: The enhanced acoustic imaging feature F corresponding to the acoustic imaging data of the device in the same group as G1 is selected. Aenh1 F Aenh1 This is a 512-dimensional vector containing the long-term dynamic characteristics of the acoustic wave of insulator #1 (such as "2-10kHz frequency band pulse sequence, amplitude attenuation rate 0.12 / cycle, interval 1.2s"), and F Ubase1 Describe the fault state (corona discharge) of the same equipment. First heterogeneous correlation sample: Enhanced acoustic imaging features F from three randomly selected data groups (e.g., G2#2 insulator, G5#5 insulator, G8#8 insulator). Aenh2 F Aenh5 F Aenh8 As a heterogeneous sample. Among them, insulator #2 is in the normal state (F Aenh2 The characteristics are "uniform energy across the entire frequency band, no pulse sequence"; #5 insulator has surface leakage (F). Aenh5 The characteristic is "continuous noise in the 10-20kHz frequency band"), which is different from the fault type of insulator #1. The server calculates the first heterogeneous alignment error component of G1 by comparing the similarity (such as cosine similarity) between the reference point and the sample: the goal is to make F Ubase1 With F Aenh1 The similarity (e.g., 0.82) is higher than that with F. Aenh2 (0.31), F Aenh5 The similarity is (0.28), if the actual similarity does not meet expectations (e.g., F). Ubase1 With F Aenh2 If the similarity is 0.45, the error component increases. The server processes G1 to G... 10 The above operation is repeated for each data group, summing the 10 error components to obtain the first heterogeneous alignment error value (reflecting the overall alignment degree of the UV-acoustic cross-modal features). The server performs a similar process for the device acoustic imaging data of each imaging data group: taking data group G3 (#3 insulator, fault type: surface leakage) as an example: Second heterogeneous alignment reference point: The basic acoustic imaging feature F of insulator #3... Abase3This is designated as the reference point. F Abase3 The vector is 256-dimensional and contains the basic time-frequency characteristics of the acoustic wave from insulator #3 (e.g., "continuous noise in the 10-20kHz frequency band, energy percentage 0.68, no obvious pulse"). Second source-correlated sample: Enhanced ultraviolet imaging features F corresponding to the ultraviolet imaging data of devices in the same group as G3 are selected. Uenh3 F Uenh3 This is a 512-dimensional vector containing the spatiotemporal dynamic characteristics of the ultraviolet discharge of insulator #3 (such as "the light spot is dispersed across the entire surface of the insulator, the area is stable at 320 pixels², and the edge gradient is 0.15 / frame"), and F Abase3 Describe the surface leakage current state of the same device. Second heterogeneous correlation sample: Enhanced ultraviolet imaging features F of three randomly selected different data groups (e.g., G1#1 insulator, G4#4 insulator, G9#9 insulator). Uenh1 F Uenh4 F Uenh9 As a heterogeneous sample. Among them, insulator #1 is a corona discharge (F Uenh1 The characteristics are "concentrated light spots on the skirt edge, with an area growth trend of 0.25 pixels² / frame"; #4 insulator is in normal condition (F). Uenh4 The characteristics are "no obvious light spot, average photon number 5"), which are different from the fault type of insulator #3. The server calculates the second heterogeneous alignment error component of G3 by comparing similarity: the goal is to make F Abase3 With F Uenh3 The similarity (e.g., 0.79) is higher than that with F. Uenh1 (0.35), F Uenh4 The similarity of (0.22), if F Abase3 With F Uenh1 If the similarity is abnormally high (e.g., 0.50), the error component increases. The server performs similarity checks on G1 to G... 10 The above operation is repeated for each data set, and the 10 error components are summed to obtain the second heterogeneous alignment error value (reflecting the overall alignment degree of acoustic-ultraviolet cross-modal features). Through the above process, the server uses cross-modal comparison of multiple data sets to ensure that ultraviolet and acoustic features of the same type of fault are close in semantic space, while features of different types of faults are far apart, providing a basis for alignment loss for subsequent network optimization.
[0060] In this embodiment of the invention, the first heterogeneous associated sample and the second heterogeneous associated sample are dynamically updated through a rolling feature cache.
[0061] In this embodiment of the invention, for example, the server maintains rolling feature caches (capacity 20) for the first and second heterogeneous related samples respectively, and dynamically updates the sample library to reflect the latest data distribution. Taking the first heterogeneous related sample as an example: the server initializes the cache to store G1-G 10 Enhanced acoustic imaging features FAenh1 - 10 Every 5 new datasets (e.g., G) are processed 11 -G 15 ), will the new feature F Aenh11 - 15 Add to the cache and remove the oldest F. Aenh1 -5, retain the most recent 20 features. Process G 16 At that time, the server retrieves data from the updated cache (including G6-G). 20 F Aenh Three samples are randomly selected as the first heterogeneous correlation sample to avoid overfitting caused by using fixed samples for a long time. The second heterogeneous correlation sample is similarly selected. The ultraviolet imaging features are enhanced by dynamically updating the rolling buffer to ensure that the heterogeneous samples continuously reflect the changes in the distribution of device status.
[0062] In this embodiment of the invention, the fusion decision error value includes a first fusion decision error value and a second fusion decision error value; the determination of the fusion decision error value based on the basic ultraviolet imaging features, the enhanced ultraviolet imaging features, the basic acoustic imaging features, and the enhanced acoustic imaging features can be implemented through the following example.
[0063] Multi-source joint features are determined based on the enhanced ultraviolet imaging features and the enhanced acoustic imaging features;
[0064] Using the basic ultraviolet imaging features as the first single-source feature anchor point, a first fusion decision error value is determined based on the first single-source feature anchor point and the multi-source joint features;
[0065] Using the basic acoustic imaging features as the second single-source feature anchor point, the second fusion decision error value is determined based on the second single-source feature anchor point and the multi-source joint features.
[0066] In an embodiment of the present invention, for example, substation 10 is used to form an image data group (G1 to G...). 10 Taking insulators #1-#10 as an example, the server determines the fusion decision error value through the process of "multi-source joint feature construction - dual anchor point error calculation": For each data group, the server will enhance the ultraviolet imaging feature (F... Uenh ) and enhanced acoustic imaging features (F Aenh ) fused into multi-source joint features (F joint Taking data set G1 (#1 insulator, corona discharge) as an example: Enhanced feature input: F Uenh1 A 512-dimensional vector (containing spatiotemporal dynamic features of ultraviolet discharge: "skirt spot expands over time, edge gradient change rate 0.18 / frame"); F Aenh1The vector is 512-dimensional (containing acoustic wavelength temporal features: "2-10kHz pulse sequence, amplitude attenuation rate 0.12 / pulse"). The fusion method uses weighted element-wise multiplication (initially weighted at 0.5, dynamically optimized during training) to highlight complementary features (ultraviolet spatial distribution and acoustic temporal patterns), resulting in a 512-dimensional multi-source joint feature F. joint1 For example, F joint1 The correlation dimension between "spot expansion trend" and "pulse attenuation rate" is 0.75 (reflecting the spatiotemporal consistency of corona discharge). The server uses the basic ultraviolet imaging feature as the first single-source feature anchor point and compares its semantic consistency with the multi-source joint features: taking G1 as an example: the anchor point is the basic ultraviolet imaging feature F of insulator #1. Ubase1 (256 dimensions, including "spot spatial concentration 0.72, photon density 0.35 / pixel"); Intrinsic joint characterization samples: Selected F from the same group joint1 (Describe the combined corona discharge characteristics of insulator #1); External source combined characterization samples: randomly selected data from 3 different data groups (G2#2 normal insulator, G5#5 surface leakage insulator) F joint2 F joint5 As an exogenous sample (F joint2 Its characteristics are "uniform energy across the entire frequency band + no light spots", F joint5 The characteristics are "10-20kHz continuous noise + scattered light spot". The server calculates F. Ubase1 With F joint1 The cosine similarity (target ≥ 0.85, actual 0.82) and F joint2 (0.31), F joint5 The similarity is (0.29). If F Ubase1 With F joint1 If the similarity is below a threshold (e.g., 0.78), or if the similarity with the exogenous sample is too high (e.g., 0.4), the first fusion decision error component increases. For G1 to G... 10 The summation yields the first fusion decision error value. Similarly, the server uses the basic acoustic imaging features as the second single-source feature anchor point: taking G1 as an example: the anchor point is the basic acoustic imaging feature F of insulator #1. Abase1 (256 dimensions, including "energy percentage of 2-10kHz band 0.65, pulse interval 1.2s"); Intrinsic joint characterization sample: F in the same group joint1 External source combined characterization sample: Selected G3#3 normal insulator (F joint3 Features include "uniform energy across the entire frequency band + no light spots"), G7#7 mechanical loosening (F joint7 The F-type light source is characterized by "5-8kHz periodic vibration noise + no fixed light spot". joint3 F joint7 The server calculates F. Abase1 With F joint1The cosine similarity (target ≥ 0.83, actual 0.80) with F joint3 (0.27), F joint7 The similarity is (0.30). If the similarity does not meet expectations (e.g., F...), the similarity score will be lower than the expected score. Abase1 With F joint1 If the similarity is 0.75, then the second fusion decision error component increases. For G1 to G... 10 The second fusion decision error value is obtained by summing the values. Finally, the fusion decision error value is the sum of the first and second error values, reflecting the semantic consistency between the single-source basic features and the multi-source joint features.
[0067] In this embodiment of the invention, the determination of multi-source joint features based on the enhanced ultraviolet imaging features and the enhanced acoustic imaging features can be implemented through the following examples.
[0068] The enhanced ultraviolet imaging feature and the corresponding enhanced acoustic imaging feature are merged to obtain the enhanced merged feature;
[0069] An aggregation operation is performed on the enhanced merged features to obtain aggregated cluster features;
[0070] The aggregated cluster features are identified as multi-source joint features.
[0071] In an embodiment of the present invention, taking, for example, the imaging data group of insulator #1 in a substation (fault type: corona discharge), the server determines the multi-source joint features through the following steps: the server combines the enhanced ultraviolet imaging features (F...) of insulator #1... Uenh1 ) and enhanced acoustic imaging features (F Aenh1 Merging: Input features: F Uenh1 A 512-dimensional vector (containing spatiotemporal dynamic features of ultraviolet discharge: "skirt spot expands over time, edge gradient change rate 0.18 / frame"); F Aenh1 The vector is 512-dimensional (containing acoustic wavelength temporal features: "2-10kHz pulse sequence, amplitude attenuation rate 0.12 / cycle"). Merging method: Feature concatenation is used to join the two vectors end-to-end into a 1024-dimensional enhanced merged feature. For example, the first 512 dimensions retain F... Uenh1 The "spot expansion trend" dimension (value 0.68) retains F in the last 512 dimensions. Aenh1The "pulse decay rate" dimension (value 0.72) forms a joint vector that simultaneously contains ultraviolet spatial dynamics and acoustic temporal patterns. The server performs an aggregation operation on the enhanced merged features (through a fully connected layer + ReLU activation function): the fully connected layer compresses the 1024-dimensional merged features into 512 dimensions, and the weight matrix learns to highlight cross-modal complementary features (such as the correlation between ultraviolet "spot spread" and acoustic "pulse interval"). For example, the correlation dimension between "spot spread rate 0.18 / frame" and "pulse interval 1.2s" in the merged features is increased to a weight of 0.85 after aggregation, while irrelevant dimensions (such as background noise frequency band energy) are weakened. The output is a 512-dimensional aggregated cluster feature, whose dimensions reflect the synergistic relationship between ultraviolet and acoustic features (such as the "corona discharge spatiotemporal consistency" dimension value of 0.78). The server directly determines the aggregated cluster features as multi-source joint features (F joint1 F joint1 The cross-modal integrated characteristics of the corona discharge of #1 insulator are: "the spatiotemporal correlation between the skirt spot expansion and the 2-10kHz pulse sequence is 0.82, and the co-variation rate of the edge gradient and amplitude attenuation is 0.75", which provides a joint semantic representation for subsequent fusion decision error calculation.
[0072] In this embodiment of the invention, the step of using the basic ultraviolet imaging features as the first single-source feature anchor point and determining the first fusion decision error value based on the first single-source feature anchor point and the multi-source joint features can be implemented through the following example.
[0073] For each of the imaging data groups containing device ultraviolet imaging data, the basic ultraviolet imaging feature corresponding to the targeted device ultraviolet imaging data is determined as the first single-source feature anchor point, the multi-source joint feature corresponding to the imaging data group containing the targeted device ultraviolet imaging data is determined as the first intrinsic joint characterization sample, and the multi-source joint feature corresponding to the imaging data group not containing the targeted device ultraviolet imaging data is determined as the first extrinsic joint characterization sample. A first fusion decision error value is determined based on the first single-source feature anchor point, the first intrinsic joint characterization sample, and the first extrinsic joint characterization sample.
[0074] The step of using the basic acoustic imaging features as the second single-source feature anchor point, and determining the second fusion decision error value based on the second single-source feature anchor point and the multi-source joint features, includes:
[0075] For each of the imaging data groups containing device acoustic imaging data, the basic acoustic imaging feature corresponding to the targeted device acoustic imaging data is determined as the second single-source feature anchor point, the multi-source joint feature corresponding to the imaging data group containing the targeted device ultraviolet imaging data is determined as the second intrinsic joint characterization sample, and the multi-source joint feature corresponding to the imaging data group not containing the targeted device ultraviolet imaging data is determined as the second extrinsic joint characterization sample. The second fusion decision error value is determined based on the second single-source feature anchor point, the second intrinsic joint characterization sample, and the second extrinsic joint characterization sample.
[0076] In an embodiment of the present invention, for example, substation 10 is used to form an image data group (G1 to G...). 10 Taking insulators #1-#10 (including fault types such as corona discharge, surface leakage, and normal state) as an example, the server calculates the fusion decision error value through a "dual anchor point-sample comparison" process to ensure the semantic consistency between single-source basic features and multi-source joint features. For each imaging data group's equipment ultraviolet imaging data, the server performs error calculations based on "anchor point-intrinsic sample-external sample," the core of which is to verify whether single-source ultraviolet basic features and multi-source joint features describe the same equipment state. Taking data group G1 (#1 insulator, fault type: corona discharge) as an example: First single-source feature anchor point: The server compares the basic ultraviolet imaging features F of insulator #1... Ubase1 Designated as anchor point. F Ubase1 The vector is 256-dimensional and contains the basic spatial characteristics of ultraviolet discharge: "The light spot is concentrated on the skirt of the upper surface of the insulator, with a spatial concentration of 0.72 (normalized to 0-1), a photon density of 0.35 photons / pixel, and a spot circularity of 0.81" (reflecting the typical spatial distribution of corona discharge). The first intrinsic joint characterization sample: Select the multi-source joint feature F corresponding to G1, which contains the #1 ultraviolet imaging data. joint1 F joint1 It is a 512-dimensional vector, composed of the enhanced ultraviolet imaging features (F) of #1. Uenh1 "The skirt-edge light spot expands over time, with an edge gradient change rate of 0.18 / frame" and enhanced acoustic imaging features (F Aenh1 The data was obtained by fusing a pulse sequence in the 2-10kHz frequency band with an amplitude attenuation rate of 0.12 / pulse and an interval of 1.2s. It includes cross-modal integrated features: "The spatiotemporal correlation between the ultraviolet spot spread and the acoustic pulse sequence is 0.82, and the co-variation rate of the edge gradient and pulse amplitude is 0.75" (describing the ultraviolet-acoustic co-variation characteristics of the corona discharge of insulator #1). The first external source joint characterization sample: Multi-source joint features from other data groups that do not contain the ultraviolet imaging data of #1 were selected, such as the F-value of G2 (insulator #2, normal state). joint2 F of G5 (#5 insulator, surface leakage) joint5 Among them: Fjoint2 The characteristics are "uniform acoustic energy across the entire frequency band (no pulse sequence) + no light spots in the ultraviolet light (photon density 0.05 / pixel)" (normal state cross-modal characteristics); F joint5 The characteristics are "continuous noise in the 10-20kHz frequency band (without attenuation pulses) + ultraviolet light spots dispersed across the entire surface (spatial concentration 0.32)" (cross-modal characteristics of surface leakage), both of which are different from the corona discharge type of #1. Error calculation logic: The server measures the semantic consistency between the anchor point and the sample using cosine similarity (the closer the value is to 1, the more semantically consistent). The objective is F. Ubase1 With F joint1 The similarity (e.g., 0.85) was significantly higher than that with F. joint2 (0.32), F joint5 The similarity is (0.29). If in the actual calculation, F... Ubase1 With F joint1 The similarity dropped to 0.75 (below expectations), or it may be similar to F. joint2 If the similarity increases to 0.40 (semantic confusion), the first fusion decision error component of G1 increases. The server then performs a fusion decision on G1 to G... 10 Repeat the above operation for each data group (e.g., anchor point F of G2 normal insulator). Ubase2 With intrinsic sample F joint2 The similarity must be ≥0.83, and the F-value with the exogenous sample G1 must be... joint1 (Similarity must be ≤0.35). The error components of the 10 data groups are summed to obtain the first fusion decision error value. The server performs a similar process for the device acoustic imaging data of each imaging data group to verify the semantic consistency between single-source acoustic features and multi-source joint features. Taking data group G3 (#3 insulator, fault type: surface leakage) as an example: Second single-source feature anchor: The server will use the basic acoustic imaging features F of the #3 insulator... Abase3 Designated as anchor point. F Abase3 The vector is 256-dimensional and contains the basic time-frequency characteristics of the sound wave: "Energy proportion in the 10-20kHz frequency band is 0.68 (normalized to 0-1), no obvious pulse sequence (pulse interval characteristic value 0.12), and continuous noise proportion is 0.85" (reflecting typical acoustic characteristics of surface leakage). The second intrinsic joint characterization sample: The multi-source joint feature F corresponding to G3, which contains the acoustic imaging data of #3, is selected. joint3 F joint3 It is a 512-dimensional vector, composed of the enhanced acoustic imaging features (F) of #3. Aenh3 "10-20kHz continuous noise, amplitude fluctuation 0.15 / second") and enhanced ultraviolet imaging features (F Uenh3The data was obtained by fusing the following: "The light spot is dispersed across the entire surface of the insulator, with an area stable at 320 pixels², and an edge gradient of 0.15 / frame." This data includes cross-modal integrated features: "The correlation between continuous acoustic noise and the dispersed UV light spot is 0.78, and the co-stability between the noise frequency band and the light spot area is 0.81" (describing the UV-acoustic co-characteristics of leakage current on the surface of insulator #3). Second external source joint characterization sample: Multi-source joint features from other data groups that do not contain acoustic imaging data from #3, such as the F-value of G1 (#1 insulator, corona discharge). joint1 F of G7 (#7 insulator, mechanically loose) joint7 Among them: F joint1 The characteristics are "2-10kHz pulse sequence + concentrated spot on the skirt" (cross-modal characteristics of corona discharge); F joint7 The characteristics are "5-8kHz periodic vibration noise (frequency 50Hz) + no fixed UV spot" (cross-modal characteristics of mechanical loosening), both of which are different from the surface leakage type of #3. Error calculation logic: The target is F Abase3 With F joint3 The cosine similarity (e.g., 0.84) is significantly higher than that with F. joint1 (0.30), F joint7 The similarity is (0.28). If F Abase3 With F joint3 The similarity dropped to 0.76, or with F joint7 If the similarity increases to 0.38 (semantic confusion), the second fusion decision error component of G3 increases. The server compares G1 to G... 10 Repeat the above operation for each data group (e.g., G7 mechanically loosened anchor point F). Abase7 With intrinsic sample F joint7 The similarity must be ≥0.82, and the F-value with the exogenous sample G3 must be... joint3 (Similarity must be ≤0.30). The error components of the 10 data groups are summed to obtain the second fusion decision error value. The server directly adds the first fusion decision error value to the second fusion decision error value to obtain the total fusion decision error value. This value comprehensively reflects the semantic consistency between single-source basic features (UV / acoustic) and multi-source joint features. The goal is to minimize this value through network training to ensure that the fusion features retain single-source discriminativeness while integrating cross-modal complementary information.
[0077] In this embodiment of the invention, the following implementation methods are also provided.
[0078] Based on the basic ultraviolet imaging features, the enhanced ultraviolet imaging features, the basic acoustic imaging features, and the enhanced acoustic imaging features, the single-source consistency error value is determined;
[0079] The step of optimizing the model parameters of the dual-channel feature coding network based on the fusion decision error value, the first heterogeneous alignment error value, and the second heterogeneous alignment error value to obtain the target dual-channel feature coding network includes:
[0080] The model parameters of the dual-channel feature coding network are optimized based on the single-source consistency error value, the fusion decision error value, the first heterogeneous alignment error value, and the second heterogeneous alignment error value to obtain the target dual-channel feature coding network.
[0081] In an embodiment of the present invention, for example, substation 10 is used to form an image data group (G1 to G...). 10 Taking data set G1 (including insulators #1-#10) as an example, the server determines the target network through a process of "single-source feature consistency verification - multi-error joint optimization": The server verifies the semantic consistency of basic features and enhanced features for ultraviolet and acoustic modes respectively (whether the basic features are accurately inherited by the enhanced features within the same mode). The error value is the sum of the consistency errors of the two modes, and the single-source consistency error component is calculated using (1-cosine similarity). Taking data set G1 (insulator #1, corona discharge) as an example: Ultraviolet mode consistency error: compare the basic ultraviolet imaging features (FUbase1) and the enhanced ultraviolet imaging features (FUenh1). FUbase1 contains "spot spatial concentration 0.72, photon density 0.35 / pixel"; FUenh1 contains "spot expansion trend + edge gradient change rate 0.18 / frame". Both should describe the same ultraviolet discharge state, and consistency should be calculated using cosine similarity (target ≥ 0.8, actual 0.78). If the similarity is below the threshold (e.g., 0.75), the ultraviolet consistency error component increases (e.g., the ultraviolet error component of #1 is 0.05). Acoustic modal consistency error: Compare the basic acoustic imaging features (FAbase1) and the enhanced acoustic imaging features (FAenh1). FAbase1 contains "2-10kHz energy percentage 0.65, pulse interval 1.2s"; FAenh1 contains "pulse decay rate 0.12 / time + long time-series dependence". Both should describe the same acoustic discharge state, with a similarity target ≥ 0.78 (actual 0.76). If the target is not met (e.g., similarity 0.7), the acoustic consistency error component increases (e.g., the acoustic error component of #1 is 0.04). Server for G1 to G 10For each data set, the UV + acoustic consistency error component is calculated, and the summation yields the single-source consistency error value (e.g., a total error value of 1.2, reflecting the single-source feature inheritance of the 10 data sets). The server constructs a total loss function (including single-source consistency error Lsingle, fusion decision error Lfusion, first heterogeneous alignment error LAlign1, and second heterogeneous alignment error LAlign2), and optimizes the network after weighted summation: Total loss function: Ltotal = 0.3Lsingle + 1.0Lfusion + 0.5LAlign1 + 0.5LAlign2 (weights are preset according to feature importance, with a single-source consistency weight of 0.3 to ensure the stability of basic and enhanced features). Optimization process: The Adam optimizer (learning rate 1e-4) is used, iterating on the training set (G1-G7): Ltotal is calculated in each round, and the convolutional kernel weights, attention weights, and other parameters of the dual-channel feature encoding network are updated through backpropagation (e.g., increasing the association weight of UV basic-enhanced features and reducing heterogeneous sample similarity). Convergence condition: Validation set (G8-G7) 10 The network is stopped when the Ltotal value decreases by less than 1e-5 for five consecutive rounds. At this point, the single-source consistency error drops to 0.5 (similarity of each modality ≥ 0.85), the fusion decision error drops to 0.8 (similarity between anchor point and intrinsic sample ≥ 0.88), and the heterogeneous alignment error drops to 0.6 (similarity of homogeneous samples is more than 0.5 higher than that of heterogeneous samples). The server saves the converged network parameters to obtain the target network. This network can stably output semantically consistent single-source features, cross-modal aligned heterogeneous features, and highly discriminative multi-source joint features, providing reliable feature support for subsequent fault diagnosis.
[0082] In this embodiment of the invention, the following implementation methods are also provided.
[0083] The device monitoring samples for the device diagnostic task are processed by the target dual-channel feature coding network to obtain diagnostic feedback.
[0084] Based on the diagnostic feedback, the model parameters of the task adapter head module in the target dual-channel feature coding network are optimized to obtain the target dual-channel feature coding network that has been trained.
[0085] The step of using the target dual-channel feature coding network to determine faults in equipment monitoring data includes:
[0086] The trained dual-channel feature encoding network is used to determine faults in the equipment monitoring data.
[0087] In this embodiment of the invention, taking the substation insulator fault classification task (diagnostic task mode: "fault classification", task adapter module: "fault classifier") as an example, the server achieves high-precision fault judgment through the process of "task adapter fine-tuning - completing training network deployment": The server selects historical monitoring samples (insulator data with manual labels) for the equipment diagnostic task, inputs the target dual-channel feature encoding network (the encoder part has been trained, and the task adapter module is a fault classifier with initial parameters), and outputs diagnostic feedback (classification loss, accuracy). Monitoring sample and task adapter settings: Equipment monitoring samples: Select 10 sets of labeled historical data (G 11 To G 20 (Corresponding to insulators #11-#20), each group contains synchronous ultraviolet imaging data (500 frames of photon number matrix), acoustic imaging data (1000 frames of time-frequency diagram), and manually labeled fault tags (3 categories: corona discharge (C1), surface leakage (C2), normal (C3)). For example, G 11 (#11 insulator) Labeled as C1, G 15 (#15 insulator) Labeled as C2, G 20 (#20 insulator) is labeled C3. Task adapter module: The fault classifier consists of two fully connected layers (first layer 512→256 dimensions, ReLU activation; second layer 256→3 dimensions, Softmax activation). Initial parameters are randomly initialized. Input is the multi-source joint features (Fjoint) output from the encoder. Output is the probability distribution of three types of faults (e.g., [0.85, 0.12, 0.03] corresponds to the predicted probability of C1). Diagnostic feedback calculation: The server will calculate G... 11 -G 20 The training set is divided into a 7:3 ratio for fine-tuning (G). 11 -G 17 ) and fine-tuning the validation set (G 18 -G 20 Input to target network: Encoder output: Encoder of the target network (dual-channel feature encoding network) for G 11 FUbase11, FUenh11, FAbase11, and FAenh11 were extracted from the ultraviolet and acoustic imaging data and fused to obtain Fjoint. 11 (512-dimensional, including cross-modal features of class C1: "skirt-edge spot expansion + 2-10kHz pulse sequence"). Classifier prediction: Fault classifier against Fjoint. 11The output predicted probability [0.62, 0.28, 0.10] (predicted label C1) is compared with the true label C1, and the cross-entropy loss is calculated (target loss < 0.3, actual loss 0.45). Diagnostic feedback results: the accuracy of the first round of fine-tuning on the training set is 68% (7 / 10), the accuracy on the validation set is 67% (2 / 3), and the average classification loss is 0.52 (not meeting the convergence criterion). The server uses the classification loss in the diagnostic feedback as the optimization target, only updates the model parameters of the task adaptation head module (fault classifier), and freezes the encoder parameters (to avoid destroying the already trained feature extraction capabilities). Parameter optimization process: Optimizer and hyperparameters: The Adam optimizer is used (learning rate 1e-5, lower than the encoder training stage, to avoid overfitting). In each iteration, the cross-entropy loss is calculated on the fine-tuned training set, and the weights of the fully connected layers of the classifier are updated through backpropagation (such as the 512×256 weight matrix of the first layer and the 256×3 weight matrix of the second layer). Iteration and Convergence: A total of 20 iterations were performed. After each iteration, the validation set accuracy was calculated: 5th iteration: Training set loss decreased to 0.32, validation set accuracy was 83% (5 / 6); 15th iteration: Training set loss was 0.21, validation set accuracy was 92% (11 / 12); 20th iteration: Validation set accuracy remained stable at 95% for three consecutive iterations (19 / 20), and the average classification loss was 0.18 (≤0.2), at which point optimization ceased. The trained network was saved: The parameters of the fault classifier at this point were saved (e.g., the weight of the "spot expansion-pulse sequence" association dimension in the first layer weight matrix was increased to 0.88), and combined with the encoder to form the "completely trained target dual-channel feature encoding network" (denoted as Netfinal). Netfinal was deployed on the server to process new equipment monitoring data (unlabeled) uploaded in real-time from the substation and output fault judgment results. Taking the new monitoring data D (#21 insulator, unknown state) as an example: Data input: The server receives the ultraviolet imaging data (500 frames, photon number matrix) and acoustic imaging data (1000 frames time-frequency map) of D, and inputs them into Netfinal after synchronization and alignment. Feature extraction and fusion: The Netfinal encoder performs feature extraction on D: the ultraviolet branch outputs FUbase ("spots are scattered across the entire surface, spatial concentration 0.31") and FUenh ("spot area is stable, edge gradient 0.12 / frame"); the acoustic branch outputs FAbase ("energy proportion in the 10-20kHz band 0.65, continuous noise") and FAenh ("no pulse sequence, noise amplitude fluctuation 0.15 / second"); fusion yields Fjoint (512 dimensions, containing the feature "scattered spot + 10-20kHz continuous noise"). Classifier prediction: The fault classifier outputs a probability distribution [0.08, 0.87, 0.05] for Fjoint (C2 class probability 0.87 is the highest).Fault diagnosis result: The server selects the category with the highest probability, C2 (surface leakage), as the final result, with a confidence level of 0.87, and pushes it to the substation monitoring system (e.g., displaying "#21 Insulator Fault Type: Surface Leakage, Confidence Level 87%"). Through the above process, the trained target network can accurately integrate ultraviolet and acoustic features to achieve efficient classification of insulator faults.
[0088] In this embodiment of the invention, the fault judgment of the device monitoring data by the trained target dual-channel feature encoding network can be performed through the following example.
[0089] Acquire equipment monitoring data and corresponding diagnostic task modes;
[0090] In the trained target dual-channel feature encoding network, a target dual-channel feature encoding network that matches the diagnostic task mode is selected; the matched target dual-channel feature encoding network includes an encoder and a task adapter head module, and the task adapter head module includes one of a fault classifier, a region locator, a multi-source correlator, or a diagnostic inference engine.
[0091] The fault judgment result is obtained by using the matched target dual-channel feature coding network to perform fault judgment on the device monitoring data.
[0092] In this embodiment of the invention, taking a substation equipment intelligent diagnostic system as an example, the server needs to call the matching task adapter module according to different diagnostic needs (such as fault type identification, discharge location positioning, cross-modal feature association verification, and fault cause inference), and complete the fault judgment through the "task mode matching - dedicated network inference" process: the server receives real-time equipment monitoring data through the substation IoT gateway and obtains the diagnostic task mode (the specific diagnostic needs preset by the user or system) from the scheduling system. Taking the #22 insulator (newly commissioned equipment, running for 1 month) as an example: equipment monitoring data: synchronously collected ultraviolet imaging data (500 frames, 640×512 pixel photon number matrix, including suspected discharge areas) and acoustic imaging data (1000 frames time-frequency map, 256×256 sound wave intensity distribution, including abnormal frequency band energy), the data are timestamped (accuracy ±1ms) to form a group of imaging data D. 22Diagnostic task mode: The scheduling system issues the task instruction "Regional localization + diagnostic inference" (requiring the location of the discharge area coordinates and inference of the fault cause and development trend). The server, having completed training on the target dual-channel feature encoding network library, selects the corresponding task adapter module based on the diagnostic task mode and combines it with the general encoder to form a dedicated network (encoder parameters are fixed, adapter parameters are fine-tuned according to the task). For the "Regional Localization" and "Diagnostic Inference" tasks: General Encoder: All tasks share the trained dual-channel feature encoding network, responsible for extracting ultraviolet / acoustic features and fusing them into multi-source joint features (Fjoint). For example, for D... 22 The ultraviolet data was extracted using FUbase ("spatial concentration of suspected spot region 0.68") and FUenh ("spot edge gradient 0.22 / frame"); the acoustic data was extracted using FAbase ("energy proportion of 3-8kHz band 0.58") and FAenh ("pulse interval 1.5s"). These were then fused to obtain Fjoint. 22 (512-dimensional, including spatiotemporal-frequency correlation features of the discharge region). Matching task adapter module: Region locator (corresponding to the "Region Location" task): Consists of two fully connected layers (first layer 512→256 dimensions, ReLU activation; second layer 256→4 dimensions, linear activation). Input Fjoint, output the rectangular bounding box coordinates (x1, y1, x2, y2, pixel coordinates) of the discharge region in the ultraviolet image. Parameters are fine-tuned using historical location samples (ultraviolet images of manually labeled discharge regions), with a location error ≤ 5 pixels. Diagnostic inference engine (corresponding to the "Diagnostic Inference" task): Employs a graph neural network (GNN) structure. Input Fjoint and equipment metadata (insulator model, operating time, ambient humidity 65%), output the fault cause (e.g., "surface contamination causing local field strength distortion") and the development trend within 30 days (e.g., "discharge intensity increase probability 0.75, planned maintenance required"). Parameters are fine-tuned using a historical fault case library (including cause-trend labels), with an inference accuracy ≥ 85%. The server calls a dedicated matching network to process tasks D. 22 Data outputs fault diagnosis results. "Area Localization" task processing: Input: Fjoint output from the encoder. 22 (Including the associated features of "spatial-temporal distribution of suspected light spot areas + time-frequency location of 3-8kHz pulses"). Localization process: The region locator decodes Fjoint through a fully connected layer. 22 Based on the spatial coordinate characteristics of the image and the resolution of the ultraviolet image (640×512), the coordinates of the rectangular frame of the discharge region are output (x1=280, y1=150, x2=320, y2=190) (corresponding to the inner side of the third umbel skirt on the upper surface of the insulator, with a physical location "25cm from the top, 120° circumferential direction"). "Diagnostic Reasoning" Task Processing: Input: Fjoint 22(Including dynamic features such as "spot edge gradient 0.22 / frame + pulse interval 1.5s") and device metadata (model XP-100, running for 30 days, humidity 65%). Inference process: The diagnostic inference engine's GNN will use Fjoint... 22 Comparing the characteristics of the "surface contamination discharge" template in the fault case library (similarity 0.82), and considering the humidity (65% is close to the critical humidity of 70% for contamination discharge), the output is: Fault cause: "Contamination on the insulator surface (salt density 0.15mg / cm²) causes local field strength distortion, triggering corona discharge"; Development trend: "The probability of discharge intensity increasing within the next 30 days is 0.75 (current spot area is 180 pixels², expected to increase to 300 pixels²), cleaning and maintenance are recommended within 7 days." The server integrates the results of each task and generates a structured report, which is pushed to the maintenance terminal: #22 Insulator fault judgment result: Fault type: Corona discharge (confidence 0.92); Discharge area: Inner side of the 3rd shed on the upper surface (coordinates 280, 150-320, 190, physical location 25cm from the top); Fault cause: Surface contamination causes local field strength distortion; Development trend: High risk of discharge intensity increasing within 30 days (probability 0.75), cleaning and maintenance are recommended within 7 days. By combining a multi-task adapter head with a universal encoder, the server achieves full-process automation from data acquisition to multi-dimensional diagnostic result output, meeting the needs of refined operation and maintenance of substation equipment.
[0093] This invention provides a computer device 100, which includes a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device 100 executes the aforementioned deep learning-based ultraviolet imaging and acoustic imaging fusion fault classification method. Figure 2 As shown, Figure 2 This is a structural block diagram of a computer device 100 provided in an embodiment of the present invention. The computer device 100 includes a memory 111, a processor 112, and a communication unit 113. To enable data transmission or interaction, the memory 111, processor 112, and communication unit 113 are electrically connected to each other directly or indirectly. For example, these components can be electrically connected to each other through one or more communication buses or signal lines.
[0094] For illustrative purposes, the foregoing description has been made with reference to specific embodiments. However, the foregoing illustrative discussions are not intended to be exhaustive or to limit the present disclosure to the precise forms disclosed. Numerous modifications and variations are possible in accordance with the foregoing teachings. These embodiments were chosen and described in order to best illustrate the principles of the present disclosure and its practical application, thereby enabling those skilled in the art to best utilize the disclosure and to employ various embodiments with different modifications to suit a particular intended application.
Claims
1. A deep learning-based ultraviolet imaging and acoustic imaging fusion fault classification method, characterized in that, The method comprises: acquiring an imaging data set composed of device ultraviolet imaging data and device acoustic imaging data; performing feature extraction operations on the device ultraviolet imaging data in the imaging data set through an ultraviolet spatial feature extraction component and an ultraviolet enhanced feature extraction component of a dual-channel feature coding network to obtain basic ultraviolet imaging features and enhanced ultraviolet imaging features corresponding to the device ultraviolet imaging data; performing feature extraction operations on the device acoustic imaging data in the imaging data set through an acoustic time-frequency feature extraction component and an acoustic enhanced feature extraction component of the dual-channel feature coding network to obtain basic acoustic imaging features and enhanced acoustic imaging features corresponding to the device acoustic imaging data; the number of imaging data sets is multiple; for the device ultraviolet imaging data of each imaging data set, determining the basic ultraviolet imaging features corresponding to the device ultraviolet imaging data as first hetero-alignment reference points, determining the enhanced acoustic imaging features corresponding to the device acoustic imaging data belonging to the same imaging data set as first homologous correlation samples, determining the enhanced acoustic imaging features corresponding to the device acoustic imaging data belonging to different imaging data sets as first hetero-correlation samples, and determining first hetero-alignment error values according to the first hetero-alignment reference points, the first homologous correlation samples, and the first hetero-correlation samples; and for the device acoustic imaging data of each imaging data set, determining the basic acoustic imaging features corresponding to the device acoustic imaging data as second hetero-alignment reference points, determining the enhanced ultraviolet imaging features corresponding to the device ultraviolet imaging data belonging to the same imaging data set as second homologous correlation samples, determining the enhanced acoustic imaging features corresponding to the device ultraviolet imaging data belonging to different imaging data sets as second hetero-correlation samples, and determining second hetero-alignment error values according to the second hetero-alignment reference points, the second homologous correlation samples, and the second hetero-correlation samples; determining fusion decision error values according to the basic ultraviolet imaging features, the enhanced ultraviolet imaging features, the basic acoustic imaging features, and the enhanced acoustic imaging features; the fusion decision error values include first fusion decision error values and second fusion decision error values; determining multi-source joint features according to the enhanced ultraviolet imaging features and the enhanced acoustic imaging features; for the device ultraviolet imaging data in each imaging data set, determining the basic ultraviolet imaging features corresponding to the device ultraviolet imaging data as first single-source feature anchor points, determining multi-source joint features corresponding to the imaging data set containing the device ultraviolet imaging data as first intrinsic joint representation samples, determining the multi-source joint features corresponding to the imaging data set not containing the device ultraviolet imaging data as first extrinsic joint representation samples, and determining first fusion decision error values according to the first single-source feature anchor points, the first intrinsic joint representation samples, and the first extrinsic joint representation samples; For the device acoustic imaging data in each of the imaging data sets, the basic acoustic imaging feature corresponding to the device acoustic imaging data is determined as a second single-source feature anchor point, the multi-source joint feature corresponding to the imaging data set containing the device acoustic imaging data is determined as a second intrinsic joint representation sample, and the multi-source joint feature corresponding to the imaging data set not containing the device acoustic imaging data is determined as a second extrinsic joint representation sample, and a second fusion decision error value is determined according to the second single-source feature anchor point, the second intrinsic joint representation sample and the second extrinsic joint representation sample; The model parameters of the dual-channel feature coding network are optimized according to the fusion decision error value, the first heterogenous alignment error value and the second heterogenous alignment error value, to obtain a target dual-channel feature coding network; The device monitoring data is subjected to fault judgment through the target dual-channel feature coding network; the device monitoring data includes target device ultraviolet imaging data and target device acoustic imaging data.
2. The method of claim 1, wherein, The method further comprises: Radiation feature enhancement is performed on the device ultraviolet imaging data to obtain basic radiation representation data and enhanced radiation representation data corresponding to the device ultraviolet imaging data; the basic radiation representation data is a single-channel matrix obtained by performing Gaussian filtering on the device ultraviolet imaging data; and the enhanced radiation representation data is three-channel matrix data obtained by performing multi-dimensional radiation feature enhancement on the device ultraviolet imaging data: the first channel is a filtered photon number distribution, the second channel is a discharge region edge gradient graph highlighted by a Retinex algorithm, and the third channel is a normalized feature of time cumulative spot area; The device ultraviolet imaging data in the imaging data set is subjected to feature extraction operation by the ultraviolet spatial feature extraction component and the ultraviolet enhanced feature extraction component of the dual-channel feature coding network to obtain basic ultraviolet imaging features and enhanced ultraviolet imaging features corresponding to the device ultraviolet imaging data, including: The basic radiation representation data is subjected to feature extraction operation by the ultraviolet spatial feature extraction component of the dual-channel feature coding network to obtain the basic ultraviolet imaging features corresponding to the device ultraviolet imaging data; The enhanced radiation representation data is subjected to feature extraction operation by the ultraviolet enhanced feature extraction component of the dual-channel feature coding network to obtain the enhanced ultraviolet imaging features corresponding to the device ultraviolet imaging data.
3. The method of claim 1, wherein, The method further comprises: The device acoustic imaging data is segmented into a time-frequency frame sequence based on a time window; The device acoustic imaging data in the imaging data set is subjected to feature extraction operation by the acoustic time-frequency feature extraction component and the acoustic enhanced feature extraction component of the dual-channel feature coding network to obtain basic acoustic imaging features and enhanced acoustic imaging features corresponding to the device acoustic imaging data, including: The time-frequency frame sequence is subjected to feature extraction operation by the acoustic time-frequency feature extraction component of the dual-channel feature coding network to obtain the basic acoustic imaging features corresponding to the device acoustic imaging data; The acoustic enhancement feature extraction component of the dual-channel feature coding network performs a feature extraction operation on the time-frequency frame sequence to obtain enhanced acoustic imaging features corresponding to the device acoustic imaging data.
4. The method of claim 3, wherein, The acoustic time-frequency feature extraction component of the dual-channel feature coding network includes a convolutional time-frequency feature extractor, a temporal compression layer, and a global aggregation layer; the acoustic time-frequency feature extraction component of the dual-channel feature coding network performs a feature extraction operation on the time-frequency frame sequence to obtain basic acoustic imaging features corresponding to the device acoustic imaging data, including: The convolutional time-frequency feature extractor performs a feature extraction operation on the time-frequency frame sequence to obtain frame-level features; The temporal compression layer performs a feature aggregation operation on the frame-level features to obtain period features; The global aggregation layer performs a feature aggregation operation on the period features to obtain the basic acoustic imaging features corresponding to the device acoustic imaging data.
5. The method of claim 1, wherein, The method further includes: According to the basic ultraviolet imaging features, the enhanced ultraviolet imaging features, the basic acoustic imaging features, and the enhanced acoustic imaging features, a single-source consistency error value is determined; the model parameters of the dual-channel feature coding network are optimized according to the fusion decision error value, the first hetero-source alignment error value, and the second hetero-source alignment error value to obtain a target dual-channel feature coding network, including: According to the single-source consistency error value, the fusion decision error value, the first hetero-source alignment error value, and the second hetero-source alignment error value, the model parameters of the dual-channel feature coding network are optimized to obtain a target dual-channel feature coding network; According to the basic ultraviolet imaging features, the enhanced ultraviolet imaging features, the basic acoustic imaging features, and the enhanced acoustic imaging features, a single-source consistency error value is determined, including: The cosine similarity of the basic ultraviolet imaging features and the enhanced ultraviolet imaging features is calculated to determine an ultraviolet modal consistency error component; The cosine similarity of the basic acoustic imaging features and the enhanced acoustic imaging features is calculated to determine an acoustic modal consistency error component; The ultraviolet modal consistency error component and the acoustic modal consistency error component are summed to obtain the single-source consistency error value.
6. The method according to any one of claims 1 to 5, characterized in that, The method further includes: A device monitoring sample of a device diagnosis task is processed by the target dual-channel feature coding network to obtain a diagnosis feedback; According to the diagnosis feedback, the model parameters of the task adaptation head module in the target dual-channel feature coding network are optimized to obtain a target dual-channel feature coding network that has completed training; The fault judgment of the device monitoring data by the target dual-channel feature coding network includes: Obtaining device monitoring data and corresponding diagnosis task modes; In the target dual-channel feature coding network that has completed training, a target dual-channel feature coding network that matches the diagnosis task mode is selected; the matching target dual-channel feature coding network includes an encoder and a task adaptation head module, and the task adaptation head module includes one of a fault classifier, a region locator, or a diagnosis reasoner. The matched target double-channel feature coding network is used for fault judgment on the equipment monitoring data, and a fault judgment result is obtained.
7. A server system, characterized by The method comprises the steps of: a server is used to execute the method of any one of claims 1-6.
Citation Information
Patent Citations
Direct current corona discharge mode identification method and system based on YOLOv8 and CNN model
CN119516314A
Method, device and system based on acoustic-thermal multi-modal fusion imaging
CN119901818A