Intelligent fault diagnosis and positioning method based on acousto-optic thermal multi-feature fusion

By acquiring and processing multi-source data, and using a visual reconstruction generator network to extract visual reconstruction residual maps and perform dual-channel feature encoding, the problem of missing detection of "cold" faults in existing technologies is solved, and accurate classification and location of faults in complex equipment are achieved.

CN122432924APending Publication Date: 2026-07-21STATE GRID SHANDONG ELECTRIC POWER CO
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610573214.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-28
Publication Date
2026-07-21

Smart Images

  • Figure CN122432924A_ABST
    Figure CN122432924A_ABST
Patent Text Reader

Abstract

The application discloses a fault intelligent diagnosis and positioning method based on acousto-optic thermal multi-feature fusion, and relates to the field of fault intelligent diagnosis. First, a visual reconstruction generator is used to extract a visual reconstruction residual graph to capture slight appearance structure abnormalities, so that multi-source data is logically decoupled into an energy channel (acoustics / thermal imaging) representing energy radiation and a structure channel (visual residual / visible light) representing physical texture. On this basis, features are extracted through double-channel parallel coding, and a modal arbitration gating mechanism is introduced to adaptively calculate the dynamic weights of the energy and structure channels according to different fault mechanisms. When facing a'silent and cold' fault, the mechanism can automatically increase the decision weight of the structure channel, effectively utilize the visual texture information to make up for the loss of acoustic and thermal signals, and finally realize accurate classification and positioning of various faults in combination with class activation mapping.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent fault diagnosis, and more specifically, to a method for intelligent fault diagnosis and localization based on the fusion of multiple features including acoustics, light, and heat. Background Technology

[0002] As modern industrial equipment develops towards larger scale, greater complexity, and higher intelligence, higher demands are placed on the accurate monitoring and fault diagnosis of equipment operating status. Traditional single-modal diagnostic methods are insufficient to comprehensively capture the health status of equipment. For example, acoustic monitoring excels at capturing vibration and noise anomalies, infrared thermal imaging can effectively identify overheating faults, while visible light images intuitively reflect the external morphology of the equipment. Therefore, fault diagnosis technology based on the fusion of acoustic, optical, and thermal information has emerged, aiming to achieve comprehensive perception of equipment faults by integrating features from different physical dimensions, thereby improving the robustness and accuracy of diagnosis.

[0003] However, existing fault diagnosis technologies based on acoustic-optical-thermal fusion have significant limitations in practical applications, especially when facing "cold" or "silent" faults. This limitation stems from the fact that existing technologies often employ an energy-release-centric fusion paradigm, which assumes that all faults are accompanied by significant energy radiation (such as abnormal sounds or temperature rises). Under this assumption, the fusion model tends to look for the intersection of energy anomalies in the acoustic and thermal imaging modes as diagnostic criteria. However, in "cold" fault scenarios such as structural micro-deformation, surface corrosion, insulation micro-cracks, loose bolts, or oil leaks, the fault itself does not generate active heat sources or significant vibration noise. In these cases, the acoustic and thermal imaging data appear as ineffective background noise, causing the energy anomaly-based fusion mechanism to fail. Furthermore, in existing fusion systems, the visible light mode is usually only considered as an auxiliary role in providing scene context, and its feature extraction is limited to shallow contour recognition, lacking in-depth mining of visual texture and structural residuals. Because of the lack of a modal arbitration mechanism that can dynamically switch the diagnostic authority to visual structural features when acoustic and thermal signals are "silent," existing systems struggle to capture fault signs hidden in fine visual textures, making it easy to miss the detection of the aforementioned "silent and cold" faults.

[0004] Therefore, how to construct an intelligent diagnostic method that can decouple energy characteristics from structural characteristics and perform adaptive weighting and localization based on different fault mechanisms is a technical problem that urgently needs to be solved. Summary of the Invention

[0005] To address the aforementioned problems in the existing technology, according to one aspect of this application, a fault intelligent diagnosis and localization method based on the fusion of multiple acoustic, optical, and thermal features is provided, comprising: Acquire raw audio data, raw thermal imaging, and raw visible light images; The raw audio data, raw thermal imaging, and raw visible light images are standardized preprocessed to obtain standardized acoustic time-spectrum maps, standardized thermal imaging maps, and standardized visual images. Based on a pre-trained visual reconstruction generator network, visual appearance reconstruction and residual feature extraction are performed on standardized visual images to obtain visual reconstruction residual maps. Dual-channel parallel feature encoding is performed on the standardized acoustic time-spectrum map, standardized thermal image map, visual reconstruction residual map and standardized visual image to obtain the energy channel fusion feature vector and the structure channel fusion feature vector; Modal arbitration gating and decision fusion are performed on the energy channel fusion feature vector and the structure channel fusion feature vector to obtain the predicted fault category; The predicted fault category and the original visible light image are used to locate and visualize the fault to obtain the final fault location map.

[0006] Compared with existing technologies, this application provides a fault intelligent diagnosis and localization method based on the fusion of multiple acoustic, optical, and thermal features. Addressing the problem that existing technologies, due to their over-reliance on acoustic and thermal energy features, often miss "cold" faults that do not generate active heat sources or vibrations, this method first utilizes a visual reconstruction generator to extract visual reconstruction residual maps to capture subtle structural anomalies. This logically decouples multi-source data into an energy channel (acoustic / thermal imaging) representing energy radiation and a structural channel (visual residual / visible light) representing physical texture. Based on this, features are extracted through dual-channel parallel encoding, and a modal arbitration gating mechanism is introduced to adaptively calculate the dynamic weights of the energy and structural channels according to different fault mechanisms. When facing "silent and cold" faults, this mechanism automatically increases the decision weight of the structural channel, effectively utilizing visual texture information to compensate for the lack of acoustic and thermal signals. Finally, combined with class activation mapping, it achieves accurate classification and localization of various fault types. Attached Figure Description

[0007] The above and other objects, features, and advantages of this application will become more apparent from the more detailed description of the embodiments of this application in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this application and form part of the specification. They are used together with the embodiments of this application to explain this application and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.

[0008] Figure 1 This is a flowchart of a fault intelligent diagnosis and localization method based on the fusion of multiple features (acoustic, optical, and thermal) according to an embodiment of this application.

[0009] Figure 2 This is a schematic diagram of data flow for a fault intelligent diagnosis and localization method based on the fusion of multiple features (acoustic, optical, and thermal) according to an embodiment of this application.

[0010] Figure 3 This is a flowchart of step 3 in the fault intelligent diagnosis and localization method based on the fusion of multiple features of acoustic, optical and thermal according to an embodiment of this application.

[0011] Figure 4 This is a flowchart of step 5 in the fault intelligent diagnosis and localization method based on the fusion of multiple features of acoustic, optical and thermal according to an embodiment of this application.

[0012] Figure 5 This is a schematic diagram of the data flow in step 6 of the fault intelligent diagnosis and localization method based on the fusion of multiple features of acoustic, optical and thermal according to an embodiment of this application. Detailed Implementation

[0013] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0014] To address the problems mentioned in the background art, this application proposes a fault intelligent diagnosis and localization method based on the fusion of multiple features including acoustics, light, and heat. Figure 1 This is a flowchart of a fault intelligent diagnosis and localization method based on the fusion of multiple features (acoustic, optical, and thermal) according to an embodiment of this application. Figure 2 This is a schematic diagram of the data flow in the fault intelligent diagnosis and localization method based on the fusion of multiple features (acoustic, optical, and thermal) according to an embodiment of this application. Figure 1 and Figure 2 As shown in the embodiments of this application, the intelligent fault diagnosis and localization method based on acoustic-optical-thermal multi-feature fusion includes: Step 1, acquiring raw audio data, raw thermal imaging, and raw visible light images; Step 2, performing standardized preprocessing on the raw audio data, raw thermal imaging, and raw visible light images to obtain standardized acoustic time-spectrum maps, standardized thermal imaging maps, and standardized visual images; Step 3, based on a pre-trained visual reconstruction generator network, performing visual appearance reconstruction and residual feature extraction on the standardized visual images to obtain visual reconstruction residual maps; Step 4, performing dual-channel parallel feature encoding on the standardized acoustic time-spectrum maps, standardized thermal imaging maps, visual reconstruction residual maps, and standardized visual images to obtain energy channel fusion feature vectors and structure channel fusion feature vectors; Step 5, performing modal arbitration gating and decision fusion on the energy channel fusion feature vectors and structure channel fusion feature vectors to obtain predicted fault categories; Step 6, performing fault localization and visualization on the predicted fault categories and raw visible light images to obtain the final fault localization map.

[0015] In step 1, raw audio data, raw thermal imaging, and raw visible light images are acquired. It should be understood that, given the increasingly complex operating environments of industrial equipment, the physical manifestations of equipment failures exhibit diverse and unstructured characteristics. These include active failures accompanied by severe energy release, such as high-frequency noise generated by abnormal vibrations or localized temperature rises caused by electrical overloads, as well as passive "cold" failures that do not produce obvious acoustic or thermal signals, such as micro-cracks in insulating materials, fatigue deformation of metal components, or loosening and displacement of fasteners. To overcome the perception blind spots of single-modal sensors when facing such differentiated failure mechanisms, and to address the problem of missed detection of silent failures due to over-reliance on energy characteristics in existing technologies, a full-dimensional perception space capable of simultaneously covering sound field fluctuations, thermal distribution, and surface texture structure needs to be constructed at the data acquisition source. Therefore, this step aims to simultaneously acquire raw audio data, raw thermal imaging, and raw visible light images, thereby providing a complete and spatiotemporally aligned multi-source physical data foundation for subsequently decoupling failure characteristics into energy and structural channels.

[0016] In one possible implementation, the processing of step 1 is as follows: In this embodiment, the process of acquiring raw audio data, raw thermal imaging and raw visible light images is performed based on an integrated multimodal data acquisition device or a distributed sensor network, which ensures the temporal synchronization and spatial field of view consistency of different modal data.

[0017] First, for acquiring the raw audio data, a high-sensitivity microphone array is used to sample the sound field around the target device. The microphone array consists of multiple omnidirectional microphone units arranged in a specific topology (such as a uniform circular array or a linear array). This configuration not only captures the amplitude information of the sound waves but also preserves the spatial orientation information of the sound source. During implementation, the analog sound signal is converted into a digital signal through a high-precision analog-to-digital converter. To ensure effective capture of high-frequency signals generated by early mechanical faults such as bearing micro-wear, the sampling rate is preset to a high value, such as 48kHz or 96kHz, to satisfy the Nyquist sampling theorem and avoid spectral aliasing. The acquired raw audio data is represented as a time-series voltage signal of a preset duration, such as 1 to 5 seconds. This signal fully records the acoustic fingerprint of the device during operation, including background noise and possible abnormal impacts or periodic pulse components, serving as the basis for subsequent generation of a standardized acoustic time-spectrum graph.

[0018] Secondly, the acquisition of raw thermal images is achieved using an infrared thermal imager. This imager is equipped with an uncooled infrared focal plane array detector, operating in the 8-14 micrometer long-wave infrared spectrum. The device detects the infrared radiation intensity on the surface of the target device, converts it into an electrical signal, and processes it to generate a two-dimensional temperature distribution matrix. During acquisition, the imager's emissivity parameter is calibrated according to the material of the device under test (such as a metal casing or composite insulating material) to ensure accurate temperature measurement. The output raw thermal image is a single-channel grayscale or pseudo-color image with a resolution of, for example, 640×512 pixels. Each pixel value in the image directly corresponds to the absolute temperature or radiant flux at that point on the device surface. This data is primarily used to capture thermal anomalies caused by current overload, poor contact, or mechanical friction, providing a thermodynamic basis for energy channel feature extraction.

[0019] Meanwhile, the acquisition of raw visible light images is accomplished using a high-resolution industrial-grade visible light camera. This camera employs a CMOS or CCD image sensor, capable of capturing high-definition appearance information of the equipment within the visible spectrum. During implementation, the camera automatically or manually adjusts exposure time, gain, and white balance based on ambient lighting conditions to obtain clear images with high color fidelity. The acquired raw visible light images are three-channel RGB color images with a high resolution set to standards such as 1920×1080 pixels or 4K resolution to ensure that minute surface texture details can be clearly distinguished. This data is crucial for identifying "cold" faults because it records physical structural states that do not possess thermoacoustic characteristics, such as surface corrosion spots, minute cracks in insulator skirts, and angular misalignment of nuts. It serves as the core data source for subsequent visual reconstruction and residual analysis.

[0020] Strict time synchronization control is essential during the execution of the three acquisition sub-steps described above. Through hardware trigger signals or Network Time Protocol (NTP), the microphone array, infrared thermal imager, and visible light camera are triggered to start acquisition simultaneously, ensuring that each frame of image and each segment of audio corresponds to the exact same instant of device operation. Finally, the acquired raw audio data (time series vector), raw thermal image (temperature matrix), and raw visible light image (RGB pixel matrix) are packaged and stored, and then uniformly aggregated and managed through a cloud platform, serving as the raw input data stream for the next stage.

[0021] In step 2, the raw audio data, raw thermal imaging, and raw visible light images are standardized preprocessed to obtain standardized acoustic time-spectrum maps, standardized thermal imaging maps, and standardized visual images. Correspondingly, due to the significant heterogeneity in data structure, physical dimensions, and sampling frequency of the raw audio data, raw thermal imaging, and raw visible light images acquired in the previous stage, directly inputting these multi-source heterogeneous data into the subsequent deep neural network model will lead to inconsistent feature space distribution, resulting in problems such as difficulty in model convergence or unstable weight updates. Specifically, the raw audio data is a one-dimensional high-frequency time series, representing the dynamic changes in sound pressure level, while the raw thermal imaging and visible light images are two-dimensional spatial matrices, representing the radiation temperature distribution and light reflection characteristics of an object's surface, respectively. Furthermore, their resolutions and field of view are often not perfectly matched. Without unifying the spatiotemporal references and standardizing the numerical ranges of these data, the fusion model will be unable to correctly correlate the characteristic manifestations of the same fault event under different modalities; for example, it will be impossible to logically align an acoustic impact at a certain moment with a visual crack at the corresponding location. Therefore, a standardized preprocessing workflow is used to eliminate the dimensional differences of heterogeneous data, convert one-dimensional time-domain signals into two-dimensional time-frequency domain features, map images of different resolutions to a unified size space, and ensure the synchronization of multimodal data in physical events through spatiotemporal alignment.

[0022] In one possible implementation, step 2 is processed as follows: First, the raw data collected in the previous steps is received, including a one-dimensional audio voltage signal sequence with a duration of T = 1 second and a sampling rate of 48kHz, a raw thermal imaging temperature matrix with a resolution of 640×512, and a raw visible light RGB image with a resolution of 1920×1080.

[0023] The first stage is spatiotemporal alignment. Although a rough synchronization was ensured during the data acquisition stage through a hardware triggering mechanism, further fine calibration is required in the preprocessing stage. For temporal alignment, the processing unit checks the timestamp metadata of each modal data packet to ensure that the deviation between the start acquisition time of the audio data and the exposure center time of the thermal imaging and visible light images is within the allowable tolerance range, such as less than 10 milliseconds. For spatial alignment, considering the parallax between the physical installation positions of the infrared thermal imager and the visible light camera, geometric correction of the image is required using a pre-calibrated homography matrix. Let the pixel coordinates in the original visible light image be... The corresponding thermal imaging pixel coordinates are Through formula Perform coordinate mapping, where It is a 3×3 homography matrix. Based on the mapping results, the original thermal image or the original visible light image is cropped and registered so that the two have a completely overlapping field of view in subsequent processing, ensuring that the same pixel position in the image corresponds to the same physical point on the device.

[0024] The second stage involves performing acoustic feature transformation on the original audio data to generate a standardized acoustic time-spectrum. Given that fault-generated acoustic signals are often non-stationary, a simple time-domain waveform is insufficient to reveal the evolution of their frequency components over time. Therefore, a Short-Time Fourier Transform (STFT) is used to convert the one-dimensional audio signal into a two-dimensional time-spectrum. Let the original audio signal be... The default window function is Hann window was chosen to reduce spectral leakage; window length... Set to 1024 points, overlap step length The value is set to 512 points. The formula for calculating STFT is as follows:

[0025] In the formula, Indicates the index of the time frame. Indices representing frequency points. For the generated complex spectrum matrix, The imaginary unit is used. To obtain characteristics that reflect the energy distribution, the square of the amplitude of the complex spectrum needs to be calculated, i.e. Because the dynamic range of sound signals is extremely wide, in order to simulate the nonlinear perception of sound loudness by the human ear and compress the data range, a logarithmic scaling transformation is then performed on the amplitude spectrum, as shown in the formula: In the formula, For a very small positive number, such as This is used to prevent numerical errors caused by taking the logarithm of zero. The resulting... It is a two-dimensional matrix, where the number of rows corresponds to the frequency resolution and the number of columns corresponds to the number of time frames. To adapt to the input size requirements of subsequent convolutional neural networks (CNNs), the size of this two-dimensional matrix is ​​adjusted to a preset uniform size, such as 256×256, using a bicubic interpolation algorithm. Finally, the spectrum after adjustment is normalized to ensure that its values ​​are distributed within the [0,1] interval, as shown in the formula:

[0026] and These are the global minimum and maximum values ​​in the two-dimensional matrix, respectively. The values ​​are in a two-dimensional matrix, from which the following is obtained. This is a standardized acoustic time-frequency spectrum diagram, which visually displays the energy distribution of the sound from a device in terms of time and frequency.

[0027] The third stage involves image processing of the original thermal image and the original visible light image to obtain a standardized thermal image and a standardized visual image, respectively. For the original thermal image, the original data is a floating-point temperature value matrix. First, abnormal noise in the background is removed according to a preset effective temperature range, such as the ambient temperature to the device's maximum allowable temperature. Subsequently, to unify the input specifications of the feature extraction network, a bilinear interpolation algorithm is used to downsample the 640×512 temperature matrix to 256×256 pixels. To enhance the contrast of thermal features and eliminate the influence of the absolute temperature magnitude, the resampled image is Z-score normalized or linearly normalized. In this embodiment, linear normalization is used to map the temperature values ​​to the [0,1] interval, using the same formula as above. For some single-channel infrared images, to enrich the input feature dimensions, in an optional implementation, they can be copied and expanded into three-channel data, or directly processed as a single-channel tensor in subsequent steps. The final data obtained here is the standardized thermal image. For the original visible light image, irrelevant background is first removed according to the cropping area determined in the first stage. Next, if the subsequent visual reconstruction network only requires structural information, the RGB image can be converted to a grayscale image. However, in this preferred embodiment, to preserve visual features such as rust color and oil color, the RGB three-channel information is retained. A bicubic interpolation algorithm is used to adjust the image resolution from 1920×1080 to a uniform 256×256 pixels. Regarding brightness and contrast normalization, to eliminate the influence of changes in lighting conditions (such as cloudy days or direct sunlight) on image features, a Limit Contrast Adaptive Histogram Equalization (CLAHE) algorithm is used to process the image, enhancing local texture details. The processed image is also normalized, mapping pixel values ​​from the integer domain [0,255] to the floating-point domain [0,1], resulting in a normalized visual image.

[0028] In step 3, based on the pre-trained visual reconstruction generator network, visual appearance reconstruction and residual feature extraction are performed on the standardized visual image to obtain a visual reconstruction residual map. It is understandable that in traditional fault diagnosis systems, the detection of "cold" faults such as microcracks in insulators, early corrosion of metal components, or slight loosening of bolts often faces the problem of feature overload. These types of faults do not produce significant thermal radiation or induce severe acoustic vibrations; their abnormal information is mainly hidden in the subtle texture structure of the visible light image. However, when directly analyzing the original standardized visual image, complex background environments (such as trees, sky, and complex mechanical structures) often occupy the main information of the image, causing subtle fault features to be masked by high-frequency background textures. To solve this problem, this application introduces a visual reconstruction generator network based on unsupervised learning, utilizing the network's ability to memorize normal sample distributions to ideally reconstruct the input standardized visual image. By comparing the real input with the reconstruction result, background information can be effectively suppressed, and potential fault features can be transformed into highly significant residual signals, thereby generating a visual reconstruction residual map, which provides pure fault texture information with background noise removed for the feature encoding of subsequent structural channels.

[0029] Figure 3 This is a flowchart of step 3 in the fault intelligent diagnosis and localization method based on multi-feature fusion of acoustic, optical, and thermal features according to an embodiment of this application. Figure 3 As shown, in one possible implementation, step 3, based on a pre-trained visual reconstruction generator network, performs visual appearance reconstruction and residual feature extraction on the standardized visual image to obtain a visual reconstruction residual map, including: step 31, inputting the standardized visual image into the pre-trained visual reconstruction generator network to perform visual appearance reconstruction to obtain a reconstructed image; step 32, calculating the residual between the standardized visual image and the reconstructed image to obtain a visual reconstruction residual map.

[0030] In the above implementation, step 3 is processed as follows: First, step 31 is performed. The pre-trained visual reconstruction generator network adopts a deep convolutional autoencoder architecture, or more preferably, a U-Net generator architecture with skip connections, aiming to learn the latent distribution of the data through dimensionality reduction and dimensionality expansion operations. The network mainly consists of an encoder and a decoder. The encoder part contains multiple stacked convolutional layer blocks, each block consisting of a convolutional layer, a batch normalization layer, and a LeakyReLU activation function. For example, the first convolutional kernel size is 3×3 with a stride of 1, mapping the input 256×3 image to a 256×256×64 feature map; the subsequent downsampling layer gradually reduces the spatial resolution of the feature map and increases the number of channels through a convolutional operation with a stride of 2, successively passing through 128×128×128, 64×64×256, and finally compressing it to a bottleneck layer latent feature vector of 32×32×512. This process aims to extract high-level semantic features of the image and filter out redundant information. The decoder performs the inverse operation of the encoder, using transposed convolutions or upsampling layers to progressively restore the image size, ultimately outputting a reconstructed image of size 256×256×3. The pre-training process for this network is completed offline. The training dataset consists of tens of thousands of standardized visual images containing only healthy equipment (such as intact insulators and rust-free towers). The loss function during training... Defined as input image With network output image The pixel-level mean square error or L1 distance between them is calculated using the following formula:

[0031] in, , , These represent the image's height, width, and number of channels, respectively. and They are respectively and The values ​​at each position in the image are used. By continuously minimizing this reconstruction loss using the Adam or SGD optimizer, the network weights are updated, forcing the network to remember the texture, shape, and color distribution patterns of normal samples. Since the training set does not contain any fault samples, the network has never learned the feature representations of cracks, rust, or deformation. In step 31 of the online diagnostic phase, this set of pre-trained weight parameters is loaded in real time. When the normalized visual image output from step 2 is fed into the network for forward propagation, the network attempts to reconstruct the input image based on its learned normal knowledge. If the input normalized visual image is of a perfectly healthy device, the network can reconstruct the image with extremely high fidelity, with minimal difference between the two. However, if the input image contains a "cold" fault, such as a fine black crack (abnormal area) on the surface of the insulator, since the network has never seen crack features during training, its latent space feature vector cannot effectively encode this anomalous information. During decoding, the network tends to mentally fill in the crack area as a smooth insulator surface based on the surrounding normal texture context. Therefore, the reconstructed image output by the network will present an intact, crack-free appearance at the location corresponding to the crack.

[0032] Next, step 32 is performed, which essentially compares the differences between the actual observations and the ideal predictions at the pixel level. This is done if the input standardized visual image is represented as a tensor. The generated reconstructed image is represented as a tensor. Perform pixel-by-pixel subtraction and take the absolute value to quantify the magnitude of the difference. The specific calculation formula is as follows: ,in, For the original residual tensor, and These represent the values ​​at various locations in the standardized visual image and the reconstructed image, respectively. To obtain a single-channel visual reconstruction residual map for subsequent use as a structural attention feature, the residual values ​​of the three channels are aggregated, for example, by calculating the average value or Euclidean norm of the channel dimensions. In this embodiment, a channel average value strategy is used, and the calculation formula is as follows: , The values ​​at each position in the original residual tensor are the final values ​​obtained. This means a visual reconstruction residual map with dimensions of 256×256×1. Alternatively, a three-channel format of 256×256×3 can be retained to preserve color deviation information. In this technical solution, to retain more information for subsequent encoder use, it is preferable to retain the three-channel residual, i.e., use it directly. This serves as a visual reconstruction residual image. To illustrate this process more intuitively, a specific scenario is used. For example, in a power transmission line inspection scenario, a standardized visual image captures an insulator with a tiny rust spot at coordinates (100, 100) on its surface. In the standardized visual image... In the image, the pixel value at that coordinate (after normalization) might appear as dark brown, with an RGB value of approximately [0.3, 0.2, 0.1]. However, in the visual appearance reconstruction step, the generator network infers that the location should be white based on the surrounding normal white ceramic texture, thus outputting a reconstructed image. The pixel value at coordinates (100, 100) might be [0.9, 0.9, 0.9]. Moving to the residual calculation step, the absolute difference between the two channels is calculated: R channel difference: |0.3 - 0.9| = 0.6; G channel difference: |0.2 - 0.9| = 0.7; B channel difference: |0.1 - 0.9| = 0.8. This significant difference vector [0.6, 0.7, 0.8] constitutes a bright pixel in the visual reconstruction residual map. Conversely, for the normal blue sky background region at coordinates (50, 50) in the image, the input image pixel values ​​are approximately [0.1, 0.4, 0.8], and the generator can accurately reconstruct the blue sky [0.11, 0.39, 0.81], with a residual value close to the zero vector [0.01, 0.01, 0.01], which appears as a low-value black area in the residual map. Through the above processing, the resulting visual reconstruction residual map is actually a fault saliency map. In this image, all backgrounds that conform to the normal distribution (such as towers, sky, and intact insulator bodies) are effectively subtracted and suppressed to a black background, while any structural anomalies that do not conform to the normal distribution (such as cracks, foreign objects, and deformations) are preserved and highlighted.

[0033] In step 4, dual-channel parallel feature encoding is performed on the standardized acoustic time-spectrum map, standardized thermal imaging map, visual reconstruction residual map, and standardized visual image to obtain the energy channel fusion feature vector and the structure channel fusion feature vector. It should be understood that in the complex scenarios of industrial equipment fault diagnosis, the physical manifestations of faults exhibit significant duality: one type is the "thermal / noise" fault accompanied by intense energy release, such as localized temperature rise caused by electrical short circuits or high-frequency vibrations generated by mechanical wear; the other type is the "cold / quiet" fault lacking active energy radiation, such as micro-cracks in insulating materials, fatigue deformation of metal parts, or loosening and displacement of fasteners. Existing single fusion models often tend to be dominated by high-intensity acoustic and thermal energy features, leading to the suppression of key texture features as background noise when facing weak structural faults. To resolve this contradiction at the physical mechanism level, this application employs a dual-channel parallel feature encoding strategy in the feature extraction stage, constructing an energy channel and a structure channel respectively. It utilizes a deep convolutional neural network to independently mine the spatiotemporal energy distribution patterns in the acoustic-thermal mode and the geometric structural residual patterns in the visual mode, thereby generating two complementary and semantically independent feature vectors. This provides a high-dimensional feature foundation for subsequent adaptive modal arbitration and precise localization based on fault mechanisms.

[0034] In one possible implementation, step 4, performing dual-channel parallel feature encoding on the standardized acoustic time-spectrum map, standardized thermal image map, visual reconstruction residual map, and standardized visual image to obtain energy channel fusion feature vector and structure channel fusion feature vector, includes: step 41, stitching the standardized acoustic time-spectrum map and standardized thermal image map along the channel dimension to obtain the energy channel input tensor; step 42, stitching the visual reconstruction residual map and standardized visual image along the channel dimension to obtain the structure channel input tensor; step 43, inputting the energy channel input tensor into a predefined energy channel encoder network to obtain the energy channel fusion feature vector; and step 44, inputting the structure channel input tensor into a predefined structure channel encoder network to obtain the structure channel fusion feature vector.

[0035] In the above implementation, step 4 is processed as follows: First, step 41 is performed. The physical significance of this step is to construct a composite energy view that can simultaneously characterize the frequency distribution of the sound field and the spatial distribution of the thermal field of the device. The input data includes a standardized acoustic time-frequency spectrum with a size of 256×256×1, denoted as... And a standardized thermal image with a size of 256×256×1, denoted as Although the two are physically distinct—the former has one dimension of logarithmic frequency and the other of time frame, while the latter has two dimensions of spatial coordinates—after the normalization process in step 2, they have been mapped to the same pixel space grid. During implementation, the processing unit calls the tensor stitching function to merge these two single-channel matrices in the depth direction. Specifically, let... The data distribution is , The data distribution is The spliced ​​energy channel input tensor The dimensions become 256×256×2. At this time, The first channel carries the energy intensity and frequency characteristics of the audio signal, while the second channel carries the temperature gradient information of the device surface. This splicing method allows subsequent convolutional kernels to simultaneously detect changes in both the sharpness of the sound and the temperature increase during a single convolution operation, thereby capturing potential acoustic-thermal correlation features.

[0036] Next, proceed to step 42. This step aims to provide the structured channel encoder with complete visual context and anomaly cues. The input data includes the visual reconstruction residual map with dimensions of 256×256×3 generated in step 3, denoted as... The standardized visual image with a size of 256×256×3 output from step 2 provides contextual information for the structured channel, denoted as . The highlighted areas of the visual reconstruction residual map indicate where something looks wrong, i.e., the potential location of the fault, while the normalized visual image provides semantic information about what component it is, such as an insulator, bolt, or tower material. The processing unit stacks these two three-channel tensors along the channel dimension. Specifically, this involves... RGB channels and The RGB channels are arranged in order to generate a new six-channel tensor, namely the structured channel input tensor. Its dimensions are 256×256×6. This composite input structure ensures that when the network extracts features, it can focus on small cracks or deformations (guided by the residual map) and understand the component structure to which the crack is attached (provided by the original map), thus avoiding the semantic loss problem caused by simply relying on the residual map.

[0037] Subsequently, the core feature extraction stage is entered, implementing step 43. In this embodiment, to balance the depth of feature extraction with the real-time performance of edge computing devices, a lightweight variant based on ResNet-18 (residual network) is used as the backbone network. This network includes an input layer, five convolutional stages, and a global pooling layer. The specific architecture and processing flow of the energy channel encoder network are as follows: First, the input layer receives an energy channel input tensor with a dimension of 256×256×2. Because the input has only 2 channels, unlike the standard ResNet's 3-channel input, the first convolutional kernel (Conv1) of the network is modified to accept 2 channels of input. This layer contains 64 convolutional kernels of size 7×7 with a stride of 2 and padding of 3. After Conv1 convolution, batch normalization, and ReLU activation, the output feature map size becomes 128×128×64. This is followed by a 3×3 max-pooling layer with a stride of 2, further reducing the feature map size to 64×64×64. Subsequently, the data flows through four residual layer stages (Layer 1 to Layer 4). Each stage consists of two basic residual blocks concatenated. Each residual block contains two 3×3 convolutional layers and introduces skip connections, meaning the input data is directly added to the output data. This structure effectively solves the gradient vanishing problem in deep networks. In Layer 1, the feature map size remains unchanged at 64×64×64. In Layer 2, the convolution stride of the first residual block is set to 2, and downsampling is performed, resulting in an output feature map of 32×32×128. In Layer 3, downsampling is performed again, resulting in an output feature map of 16×16×256. In Layer 4, downsampling is performed for the final time, resulting in an output feature map of 8×8×512. At this point, the feature map output by Layer 4 is the final feature map of the energy channel, denoted as... Its dimensions are 8×8×512. Each spatial pixel in this deep feature map, a total of 8×8=64 points, corresponds to the acoustic and thermal energy pattern within a relatively large receptive field in the original image. For example, a point in the upper left corner might encode a joint feature of high-frequency howling accompanied by localized high temperature. This feature map... This will be retained and passed to subsequent steps for generating heatmap localization. To obtain the global feature vector used for classification decisions, [the following steps are performed]. Perform a global average pooling (GAP) operation. The GAP layer compresses the 8×8 spatial dimension into a single value, that is, it calculates the average of 64 values ​​across each channel. The calculation formula is: ,in, =8, =8, Iterate through numbers 1 to 512. The final output is... It is a one-dimensional vector of length 512, namely the energy channel fusion feature vector. It highly summarizes the global state of the input sample in the acoustic and thermodynamic dimensions. For example, higher values ​​of some elements in the vector may indicate the presence of periodic thermal pulses. The weight parameters of the network are obtained during the training phase through backpropagation algorithm, based on a large amount of labeled acoustic and thermal sample data, with the goal of minimizing the cross-entropy loss function. This allows the network to automatically learn which acoustic and thermal combinations correspond to which fault categories.

[0038] Meanwhile, step 44 is implemented in parallel. The structured channel encoder network adopts the same ResNet-18 backbone architecture as the energy channel encoder, but is completely independent in terms of input layer configuration and weight parameters. This homogeneous but heterogeneous design ensures consistency between the two channels at the feature abstraction level, facilitating subsequent fusion operations. The specific implementation details of the structured channel encoder network are as follows: the input layer receives a structured channel input tensor with a dimension of 256×256×6. Therefore, the first convolutional kernel of this network is customized to accept 6 input channels, corresponding to 3 residual channels + 3 original image channels. This layer also contains 64 7×7 convolutional kernels with a stride of 2. As the convolutional kernel slides across the input tensor, it simultaneously performs a weighted summation of white bright spots (fault points) in the residual image and texture edges (object contours) in the original image. The data also flows sequentially through a max-pooling layer and four residual layer stages. In shallow networks such as Layer 1 and Layer 2, the convolutional kernels are mainly trained to recognize low-level geometric features, such as straight cracks, circular rust spots, or irregular edge defects. At this time, since the input includes a visually reconstructed residual image, the network pays special attention to regions with high response values ​​in the residual image. For example, if a region in the residual image has a high pixel value (indicating an anomaly), the convolutional layer will activate the corresponding filter to extract the texture details (such as cracks) of that region in the original image. As the number of network layers increases (Layer 3 and Layer 4), the extracted features gradually become more abstract and semantic. The feature map output by Layer 4 is the deep feature map of the structured channel, denoted as... The dimensions are also 8×8×512. In this feature map, a channel with a high activation value may represent a complex, high-level semantic concept such as edge damage to the insulator skirt, rather than just a simple line. This feature map The spatial structure information is also preserved and will be transmitted for subsequent fault location. Finally, for A global average pooling (GAP) operation is performed, using the same formula as for the energy channel. This ultimately yields a structured channel fusion feature vector of length 512. This vector highly condenses structural health status information in the visual modality. Let's illustrate this process with a concrete numerical example: Imagine a system detecting a complex fault where internal discharge causes an insulator to burst. In the energy channel: It displays a high energy band at a 50Hz harmonic. It shows a bright spot with high temperature in the central area. After inputting the energy encoder, the network recognizes this combination of audio frequency and high temperature, and the final result... In the vector, elements 100 to 120, corresponding to discharge characteristics, will become very large, for example, [..., 5.2, 6.1, 4.8, ...]. In the structured channel: due to physical defects caused by the explosion, The damaged area appears as a bright white region. The shape of the missing edge is clearly shown. After inputting the structural encoder, the network identifies structural anomalies that disrupt the object's contour, resulting in the final... In the vector, the values ​​of elements 200 to 220 (corresponding to physical damage characteristics) will also become very large, for example, [..., 3.9, 4.5, 3.2, ...]. Conversely, if it is a simple bolt loosening fault (cold fault): the energy channel is normal, the sound is stable, and the temperature is normal. Feature extraction results This will manifest as a low-value distribution close to background noise, such as [..., 0.1, 0.05, 0.1, ...]. In the structured channel, due to bolt angle offset, A bright residue will appear at the nut position. After encoding, the network identifies patterns of abnormal bolt angles. The dimensions in the vector corresponding to the loosening feature will show high response values, such as [..., 7.8, 8.2, ...].

[0039] In step 5, modal arbitration gating and decision fusion are performed on the fused feature vectors of the energy channel and the fused feature vectors of the structure channel to obtain the predicted fault category. Correspondingly, in existing multimodal fault diagnosis systems, due to significant differences in the physical mechanisms of fault generation, the contribution of data from different modes to specific fault types is often unbalanced. For example, electrical faults are usually accompanied by strong acoustic and thermal energy release, in which case acoustic and infrared data contain the main diagnostic information; while mechanical structural faults are often silent, only showing subtle texture changes in visible light images, in which case acoustic and thermal data degenerate into interference noise. If traditional feature splicing or fixed-weight fusion strategies are used, noise information in invalid modes will dilute or even mask key features in effective modes, leading to a decrease in the robustness of the model when facing "cold" faults. To address this problem of dynamically changing feature contributions, this application constructs a modal arbitration gating module, aiming to enable the diagnostic model to have attention switching capabilities similar to human experts, that is, focusing on the energy channel when a significant energy signal is detected, and automatically shifting the decision weight to the structure channel when the energy signal is missing.

[0040] Figure 4 This is a flowchart of step 5 in the fault intelligent diagnosis and localization method based on acoustic-optical-thermal multi-feature fusion according to an embodiment of this application. Figure 4As shown, in one possible implementation, step 5, performing modal arbitration gating and decision fusion on the energy channel fusion feature vector and the structure channel fusion feature vector to obtain the predicted fault category, includes: step 51, performing modal arbitration gating on the energy channel fusion feature vector and the structure channel fusion feature vector to obtain the energy channel dynamic weight and the structure channel dynamic weight; step 52, based on the energy channel dynamic weight and the structure channel dynamic weight, performing weighted feature fusion on the energy channel fusion feature vector and the structure channel fusion feature vector to obtain the multimodal fusion feature vector; step 53, performing classification decision on the multimodal fusion feature vector to obtain the predicted fault category.

[0041] In the above implementation, step 5 is processed as follows: First, step 51 is implemented. The core of this step is to use a lightweight fully connected network, i.e., a gated network, to learn the complementary relationship between the two modes. In one possible implementation, step 51, modal arbitration gating is performed on the energy channel fusion feature vector and the structure channel fusion feature vector to obtain the dynamic weights of the energy channel and the structure channel, including: modal arbitration gating is performed on the energy channel fusion feature vector and the structure channel fusion feature vector using the following formula, wherein the formula is:

[0042]

[0043]

[0044] in, This is the energy channel fusion feature vector. For structured channel fusion feature vectors, For characteristic cascade functions, For joint feature vectors, and These are the learnable gate matrix and the learnable gate bias vector, respectively. for Activation function and These are the dynamic weights for the energy channel and the dynamic weights for the structure channel, respectively. The processing unit first concatenates the two input feature vectors—the fused feature vector for the energy channel and the fused feature vector for the structure channel—along the channel dimension. After this concatenation operation, a joint feature vector is obtained. ,at this time, It is a long vector with dimensions 1024, which fully preserves all compressed features of the four dimensions of sound, heat, light, and residual. Then, The input is fed into a learnable linear mapping layer, which aims to map the high-dimensional feature space to a two-dimensional weight decision space, i.e. In this formula, This is a gating matrix with a shape of 2×1024. The weight parameters of this matrix are automatically learned during the model training phase using the backpropagation algorithm. Each row represents a filter used to identify whether a modality is trustworthy from the joint features. For example, The first row of weights may impart a positive gain to high activation values ​​in energy features, while imposing a negative penalty on low-variance noise. It is a 2×1 bias vector used to adjust the reference activation level. These are the raw scores for the energy channel and the raw score for the structure channel, respectively. These two real values ​​represent the model's initial assessment of the importance of the energy channel and the structure channel; a larger value indicates a more important channel. To convert these two unbounded real values ​​into a probability distribution that sums to 1 for use as weighting coefficients, the Softmax activation function is then applied: The specific calculation of the Softmax function is as follows: , The final result and That is, the dynamic weights of the energy channel and the dynamic weights of the structure channel, and satisfying the following conditions: To understand this process more intuitively, let's take a specific fault scenario as an example. The input sample is a typical "cold" fault involving a micro-crack in the insulator skirt. In step 4, since the energy channel does not detect discharge sound or temperature rise, its output feature vector... The values ​​are generally small and exhibit a random distribution (similar to noise), for example, their feature norms are low. In contrast, the structured channel, due to the strong response of the visual reconstruction residual map at high-frequency textures, outputs feature vectors... It contains significant high-value activation patterns. When these two vectors are concatenated... And after inputting into the gating network, It can identify patterns where the first half (energy) is noise and the second half (structure) has a strong signal. After matrix multiplication, the calculated raw score may be... =0.5, =4.5. After Softmax calculation: ≈0.018, The value is approximately 0.982, resulting in an extremely high structural weight of 0.982 and an extremely low energy weight of 0.018. This means the model intelligently decides to ignore acoustic and thermal data and rely primarily on visual structural data for judgment. Conversely, if an internal discharge fault (thermal / acoustic fault) occurs... It will contain strong features, and It may perform poorly due to the lack of significant changes in appearance. In this case, the gating network will output something similar. =0.85, With a weight allocation of 0.15, the dominant role is returned to the energy channel.

[0045] Next, step 52 is implemented. This step is a process of adaptive enhancement and suppression of the original features. In one possible implementation, step 52, based on the dynamic weights of the energy channel and the dynamic weights of the structure channel, performs weighted feature fusion on the energy channel fusion feature vector and the structure channel fusion feature vector to obtain a multimodal fusion feature vector, including: performing weighted feature fusion on the energy channel fusion feature vector and the structure channel fusion feature vector using the following formula, where the formula is:

[0046] in, This is a multimodal fusion feature vector. As a scalar coefficient, it will be broadcast and multiplied. Similarly, process it in every dimension. Then, the two weighted vectors are added element-wise. Continuing with the numerical example of microcrack failure mentioned above: due to... ≈0.018, original The noise component in the image was reduced by nearly 50 times, and its interference with the final result was almost eliminated; while Multiplying by 0.982, the crack texture features it carries are almost completely preserved. The final result is... It is a cleaned and purified 512-dimensional multimodal fusion feature vector that not only integrates information, but more importantly, it shields the interference of invalid modalities.

[0047] Finally, step 53 is performed. This step is completed by a classifier sub-network. The classifier consists of one or more fully connected layers. The input is... The number of output nodes equals the preset total number of fault categories N, which is determined by the actual application scenario. For example, N=5, corresponding to: normal, crack, corrosion, loosening, and discharge. The processing procedure is as follows: First, it goes through a linear layer. ,in The weight matrix is ​​5×512. This is the bias vector. The resulting vector... It contains 5 components, each representing the unnormalized confidence score for each fault type. Then, the Softmax function is applied again to... Convert to probability distribution For example, regarding the aforementioned crack failure, The structural features are highly preserved. The classifier identifies the feature pattern as corresponding to the crack category, and the final output probability distribution may be: P(normal) = 0.01; P(discharge) = 0.02; P(crack) = 0.95; P(corrosion) = 0.01; P(loosening) = 0.01. Based on the maximum probability principle (ArgMax), the category with the highest probability value is selected as the final predicted fault category, i.e., determined to be a crack. It should be noted that the above gating matrix... Classifier matrix Bias The encoder parameters, along with those in the preceding steps, are determined within a unified end-to-end training process. During training, the cross-entropy loss between the predicted class distribution and the true label is calculated, and the error gradient is backpropagated. This mechanism forces the gating network to learn that increasing the structural weights minimizes the loss when the label is "crack" and the acoustic / thermal input is empty, and vice versa. In this way, the model automatically learns a dynamic modal arbitration strategy based on the fault mechanism.

[0048] In step 6, the predicted fault category and the original visible light image are used for fault location and visualization to obtain the final fault location map. That is, after completing the intelligent identification of equipment fault types, the actual operation and maintenance needs in industrial sites do not stop there. For frontline maintenance personnel, simply knowing the abstract category label of insulator damage or joint overheating is far from sufficient. When facing towering transmission towers or complex substation equipment groups, the key to determining repair efficiency lies in how quickly and accurately pinpointing the specific coordinates of the fault in physical space. Furthermore, because this invention employs a dynamic weighting mechanism based on modal arbitration, the decision logic of the diagnostic model dynamically switches between energy-dominated and structure-dominated modes. If the internal reasoning process of this black-box model cannot be intuitively presented to the user, it will reduce the human-machine trust of the system. Therefore, this application ultimately transforms abstract mathematical features into intuitive fault location heatmaps by generating class activation graphs that match the fault mechanism and using modal arbitration weights for adaptive fusion. This helps maintenance personnel quickly locate “cold” faults that are difficult to detect with the naked eye, such as micro-cracks, or accurately locate the specific physical components corresponding to infrared overheating points, thus achieving a complete closed loop from fault discovery to fault location.

[0049] Figure 5 This is a schematic diagram of the data flow in step 6 of the fault intelligent diagnosis and localization method based on multi-feature fusion of acoustic, optical, and thermal features according to an embodiment of this application. Figure 5As shown, in one possible implementation, step 6, performing fault localization and visualization on the predicted fault category and the original visible light image to obtain the final fault localization map, includes: step 61, performing dual-channel parallel class activation on the final feature map of the energy channel and the visual reconstruction residual map based on the predicted fault category to obtain an energy channel class activation map and a structure channel class activation map; step 62, performing adaptive fusion of the energy channel class activation map and the structure channel class activation map based on arbitration weights based on the dynamic weights of the energy channel and the dynamic weights of the structure channel to obtain a fault localization heatmap; step 63, performing image fusion of the fault localization heatmap and the original visible light image to obtain the final fault localization map.

[0050] In the above implementation, step 6 is processed as follows: First, step 61 is performed. The purpose of this step is to find evidence regions supporting the final classification decision in two independent channels. For the energy channel, gradient-weighted class activation mapping (Grad-CAM) is used. Since the energy channel encoder is a deep convolutional neural network, its output... The 512 feature channels, each measuring 8×8×512, contain highly abstract acoustic and thermal semantic features. To determine which channels among these 512 feature channels contribute most to the predicted category C, the processing unit first performs a backpropagation calculation. Let... For the classifier output layer, calculate the raw score corresponding to class C before the Softmax operation, and then calculate this score relative to the final feature map of the energy channel. Each element The partial derivative (gradient), i.e. .here, Indicates channel index 1~512, This represents spatial coordinates 1 to 8. Next, global average pooling is performed on the gradient map of each channel to obtain the global importance weight for that channel. The calculation formula is: ,in, =8×8=64. The physical meaning of this weight is: if the first... Increasing the value of each channel can significantly improve the score of category C. The values ​​are positive and relatively large. For example, in a transformer overheating fault, the feature channel detecting the high-temperature center will receive extremely high weights. After obtaining the weights, the 512 feature channels are linearly weighted and combined, and the ReLU activation function is applied to filter out negatively correlated information (i.e., suppress those regions that inhibit the prediction results), generating an energy channel class activation map. : From this, we obtain It is an 8×8 two-dimensional matrix, and its highlighted areas correspond to the spatial locations of acoustic and thermal signal anomalies. For the structured channel, considering the visual reconstruction residual map... The image, measuring 256×256×3, has already been precisely isolated at the pixel level and structural anomalies highlighted through the reconstruction mechanism of the generative adversarial network in step 3. It is itself the highest resolution attention map. For "cold" faults, the bright pixels in the residual map directly indicate the physical boundaries of cracks or deformations, and its localization accuracy is far higher than that of deep feature maps that have undergone multiple downsampling. Therefore, this embodiment directly uses the visual reconstruction residual map as the localization benchmark. First, the calculation... The average value of the channel dimension is converted into a single-channel grayscale image, and then the maximum and minimum values ​​are normalized to ensure that the values ​​are distributed in the interval [0,1], thus obtaining the structured channel class activation map. The image is 256×256 pixels and retains extremely fine texture positioning information.

[0051] It is understandable that in the process of obtaining the energy channel class activation map based on the predicted fault category, the final predicted fault category score is not determined by a single mode, but rather by the dynamic weighted fusion of the energy channel and structural channel feature vectors through modal arbitration gating. This leads to the complexity of the gradient backpropagation path. If the gradient of the final classification score relative to the energy channel feature map is directly calculated, the feature strength of the structural channel and the weight allocation of the gating unit will inevitably couple with the gradient of the energy channel. This coupling may lead to gradient contamination, i.e., strong signals of the structural channel, such as obvious visual cracks, may artificially amplify or suppress the gradient of the energy channel, causing the generated energy channel heatmap to reflect not the features of the acoustic and thermal data itself, but the shadow of visual features. Therefore, to ensure the interpretability of fault diagnosis and achieve complete decoupling of energy features and structural features at the gradient flow level, thereby accurately evaluating the independent contribution of acoustic and thermal modes to the final decision, an isolated gradient-weighted class activation mapping is adopted. This method blocks interference from structured channel information by constructing virtual independent contribution scores, ensuring that the generated activation map purely reflects the supporting evidence identified by the model from the standardized acoustic time-spectrum map and the standardized thermal imaging map.

[0052] Based on this, in a possible preferred implementation, step 61, based on the predicted fault category, performs dual-channel parallel class activation on the final feature map of the energy channel and the visual reconstruction residual map to obtain an energy channel class activation map and a structure channel class activation map, including: Calculate the contribution score of the energy channel. This separates the portion belonging solely to the energy channel from the final mixed decision result. Mathematically, the linear transformation of the final classifier can be decomposed to clarify the origin of each component. In practice, based on the fully connected classifier formula obtained from the preceding steps:

[0053] Multimodal feature vector fusion Substitute the components and expand:

[0054] Further breakdown into:

[0055] in, This is the output vector of the final classifier. The classifier weight matrix is... Let be the bias vector. Through this decomposition, the first term in the equation... This is defined as the contribution score of the energy channel. The effect of this step is to construct a virtual classification score space driven solely by the energy channel, defining an independent range for subsequent gradient calculations, thus making subsequent calculations independent of the energy channel's contribution score. The impact of the item.

[0056] Calculate the contribution score of the energy channel to category C. This step aims to focus on a specific prediction objective and quantify the degree to which the energy channel supports that specific fault category C. In practice, this is done by analyzing the classifier weight matrix. Extract the row vector corresponding to category C. And calculate its inner product with the weighted energy eigenvector: In the formula, This represents the contribution score of the energy channel to class C, which is a scalar value. This maps the high-dimensional feature contribution to a single confidence score, which intuitively reflects the degree to which the energy channel considers the current sample to belong to class C under the current gating weights. If... Very small (e.g., in a cold failure), even Fluctuations exist. It will also be suppressed to a minimum value, which is in line with logical expectations.

[0057] The contribution score is calculated relative to the gradient of each channel in the final feature map of the energy channel to obtain the set of importance weights. This step is performed to identify the final feature map of the energy channel. Which feature channels in the matrix played a key role in improving the contribution score? In practice, the chain rule is used to calculate... Compared to The kth channel The partial derivatives of the partial derivatives are then used to perform global average pooling over the spatial dimensions: ,in, The importance weight of the k-th channel. This represents the total number of spatial pixels (width × height) of the feature map. For spatial coordinate indexing. During this process, due to... The definition does not include a structured channel term; therefore, when calculating partial derivatives during backpropagation, the characteristics of the structured channel are not considered. and weight Treated as a constant, the gradient flow does not leak into the structured channel or gating unit, thus accurately capturing the characteristic response patterns within the energy channel, and generating weights. It can objectively reflect the true importance of acoustic and thermal feature maps for classification.

[0058] The importance weight set is weighted along the channel dimension of the final energy channel feature map, and the weighted features are then activated using the ReLU activation function to obtain the energy channel class activation map. This step aims to transform the abstract channel importance into an intuitive spatial heatmap distribution. In practice, the calculated weights are... With the corresponding feature channels Perform a linear weighted summation and filter out negative activation values:

[0059] in, For the generated energy channel class activation map, the ReLU function is used to suppress regions that have a negative impact on class C (i.e., suppress prediction), retaining only the positive contribution regions, thereby generating a visualization image that can truly reflect the focus of energy channels, greatly improving the interpretability of the diagnostic system.

[0060] To illustrate the practical significance of the above process, let's take a specific example of an insulator microcrack fault (a typical "silent and cold" fault): Imagine detecting a microcrack on the surface of an insulator. During the modal arbitration phase, the gated network identifies the acoustic and thermal data as invalid background noise, while the visual residual features are significant. Therefore, it outputs an extreme weight allocation, for example... =0.018, =0.982. If isolated calculation is not used, the total score is directly calculated... Differentiation is necessary because the structured channel contributes the majority of the score. The returned gradient may be strongly influenced by visual crack features, causing the generated energy heatmap to incorrectly highlight crack areas, creating the illusion for the user that cracks are also detected by sound. However, the isolated method proposed in this application: 1. First, the... The value is extremely small because it is strongly suppressed by the coefficient 0.018. 2. In calculating the gradient... At that time, based solely on this slight noise feature map of the input The relationship between them. 3. The final generated The result will appear as a messy, weak distribution of low values, or completely black. This truly reflects the state of the energy channel, which did not see any effective information in this diagnosis. When this weak energy activation map is fused with the clear and bright structural activation map (from the visual reconstruction residual) in subsequent steps at a ratio of 0.018:0.982, the noise in the energy channel is perfectly filtered out. The final location map accurately points to the crack, and the diagnostic logic is clear and explicit: the fault judgment is mainly based on visual structural anomalies, rather than acoustic and thermal anomalies. This pure attribution mechanism not only ensures the accuracy of the location but also provides a reliable basis for subsequent model debugging (such as determining whether it is a sensor fault or an error in algorithm weight allocation).

[0061] Next, step 62 is implemented. This step resolves conflicts arising from inconsistencies in multimodal localization results and ensures consistency between visualization results and diagnostic logic. First, spatial size alignment is performed. This is due to the energy channel activation map... The size is 8×8, and the structured channel class activation graph The size of the image is 256×256, and the two cannot be directly added. Therefore, a bilinear interpolation algorithm is used to smoothly upsample the 8×8 energy activation map to 256×256. The interpolation process uses a weighted average of the gray values ​​of four adjacent pixels to ensure the continuity of the thermal distribution. Let the upsampled energy map be denoted as... Subsequently, weighted fusion is performed. The dynamic weights calculated using the modal arbitration gating in step 5, i.e., the dynamic weights of the energy channel, are then utilized. and structured channel dynamic weights The two activation maps of the same size are summed pixel-by-pixel with weights. The calculation formula is:

[0062] in, This is the fused location map. To illustrate this process more specifically, the current detection object is a string of insulators with micro-cracks in the insulator skirts (a typical cold fault). In step 5, due to the lack of acoustic and thermal signals, extreme weighting is output after gating processing: =0.018, =0.982. In step 61, the energy channel generates [something] because the input is mainly background noise. It may manifest as random, low-value specks; while structured channel generation There are clear, bright lines (pixel values ​​close to 1) at the crack location, while the rest of the background is black (pixel values ​​close to 0). During the blending operation: for the pixel (x, y) at the crack location: ≈0.98. For pixels at the background location: ≈0. It can be seen that by introducing dynamic weights, noise interference in the energy channel is effectively suppressed (multiplied by 0.018), while the crack characteristics of the structure channel are significantly preserved. The final result is... It's a grayscale image that clearly outlines the crack shape. Conversely, in the case of an "overheating" fault, the weights are flipped, and the heatmap will appear as a blurry infrared hotspot shape. Finally, to adapt the positioning map to the original high-definition image, it is again... Upsampling is performed, adjusting the resolution from 256×256 to the same resolution as the original visible light image, such as 1920×1080, to obtain the final single-channel fault location heatmap.

[0063] Finally, step 63 is implemented. This step aims to generate augmented reality (AR) images that conform to human visual habits. First, pseudo-color mapping is performed. While a single-channel heatmap contains location information, grayscale changes are not sufficiently noticeable. The system uses a predefined color lookup table, such as JET or HOT, to map floating-point values ​​in the range [0,1] to RGB color pixels. Value 0 is mapped to transparent or blue (cool tones), value 1 to deep red (warm tones), and intermediate values ​​transition to green and yellow. This yields a color heatmap. Next, image overlay and fusion are performed. The color heatmap is semi-transparently overlaid on the original visible light image. Above. The calculation formula is: ,in The transparency factor is preset based on experience and is usually set between 0.4 and 0.6. For example, this application uses 0.5, which ensures that the details of the equipment in the background image are clearly visible, while also clearly displaying the thermal distribution. Finally, in the generated... Information is labeled on the image. Text information is rendered in the upper left corner or other blank area of ​​the image. For example, based on the previous numerical example, the label would be: "Predicted Category: Insulator Crack; Diagnostic Confidence: 95%; Modal Weights: Energy (1.8%) - Structure (98.2%)". The final output image... This is the final fault location map. Users will see a clear, high-definition photo of the insulator, with cracks highlighted in translucent dark red, along with detailed diagnostic data. This final fault location map and its diagnostic results can be pushed to the maintenance terminal via a cloud platform. This intuitive presentation method eliminates the need for maintenance personnel to possess advanced algorithmic knowledge, greatly improving the efficiency and safety of on-site operations.

[0064] In summary, the intelligent fault diagnosis and localization method based on the fusion of acoustic, optical, and thermal features, as described in this application, is elucidated. Addressing the problem in existing technologies where over-reliance on acoustic and thermal energy features leads to missed detections of "cold" faults that do not generate active heat sources or vibrations, this method first utilizes a visual reconstruction generator to extract visual reconstruction residual maps to capture subtle structural anomalies. This logically decouples multi-source data into an energy channel (acoustic / thermal imaging) representing energy radiation and a structural channel (visual residual / visible light) representing physical texture. Based on this, features are extracted through dual-channel parallel encoding, and a modal arbitration gating mechanism is introduced to adaptively calculate the dynamic weights of the energy and structural channels according to different fault mechanisms. When facing "silent and cold" faults, this mechanism automatically increases the decision weight of the structural channel, effectively utilizing visual texture information to compensate for the lack of acoustic and thermal signals. Finally, combined with class activation mapping, it achieves accurate classification and localization of various fault types.

Claims

1. A fault intelligent diagnosis and localization method based on the fusion of multiple features including acoustics, optical motion, and thermal imaging, characterized in that... include: Acquire raw audio data, raw thermal imaging, and raw visible light images; The raw audio data, raw thermal imaging, and raw visible light images are standardized preprocessed to obtain standardized acoustic time-spectrum maps, standardized thermal imaging maps, and standardized visual images. Based on a pre-trained visual reconstruction generator network, visual appearance reconstruction and residual feature extraction are performed on standardized visual images to obtain visual reconstruction residual maps. Dual-channel parallel feature encoding is performed on the standardized acoustic time-spectrum map, standardized thermal image map, visual reconstruction residual map and standardized visual image to obtain the energy channel fusion feature vector and the structure channel fusion feature vector; Modal arbitration gating and decision fusion are performed on the energy channel fusion feature vector and the structure channel fusion feature vector to obtain the predicted fault category; The predicted fault category and the original visible light image are used to locate and visualize the fault to obtain the final fault location map.

2. The fault intelligent diagnosis and localization method based on multi-feature fusion of acoustic, optical, and thermal features according to claim 1, characterized in that, Based on a pre-trained visual reconstruction generator network, visual appearance reconstruction and residual feature extraction are performed on standardized visual images to obtain visual reconstruction residual maps, including: A standardized visual image is input into a pre-trained visual reconstruction generator network to reconstruct the visual appearance and obtain a reconstructed image. The residual between the normalized visual image and the reconstructed image is calculated to obtain the visual reconstruction residual map.

3. The fault intelligent diagnosis and localization method based on multi-feature fusion of acoustic, optical, and thermal features according to claim 1, characterized in that, Dual-channel parallel feature encoding is performed on the standardized acoustic time-spectrum map, standardized thermal image map, visual reconstruction residual map, and standardized visual image to obtain energy channel fusion feature vectors and structure channel fusion feature vectors, including: The standardized acoustic time-spectrum map and the standardized thermal image map are stitched together along the channel dimension to obtain the energy channel input tensor; The visual reconstruction residual map and the normalized visual image are stitched together along the channel dimension to obtain the structured channel input tensor; The energy channel input tensor is input into a predefined energy channel encoder network to obtain the energy channel fusion feature vector; The structured channel input tensor is input into a predefined structured channel encoder network to obtain the structured channel fused feature vector.

4. The fault intelligent diagnosis and localization method based on multi-feature fusion of acoustic, optical, and thermal features according to claim 1, characterized in that, Modal arbitration gating and decision fusion are performed on the energy channel fusion feature vector and the structure channel fusion feature vector to obtain the predicted fault category, including: Modal arbitration gating is applied to the energy channel fusion feature vector and the structure channel fusion feature vector to obtain the dynamic weights of the energy channel and the structure channel. Based on the dynamic weights of the energy channel and the dynamic weights of the structure channel, a weighted feature fusion is performed on the fused feature vectors of the energy channel and the structure channel to obtain a multimodal fused feature vector. The multimodal fusion feature vector is used to make a classification decision to obtain the predicted fault category.

5. The fault intelligent diagnosis and localization method based on multi-feature fusion of acoustic, optical, and thermal features according to claim 4, characterized in that, Modal arbitration gating is applied to the energy channel fusion feature vector and the structure channel fusion feature vector to obtain the dynamic weights of the energy channel and the structure channel, including: applying modal arbitration gating to the energy channel fusion feature vector and the structure channel fusion feature vector using the following formula: ; ; ;in, This is the energy channel fusion feature vector. For structured channel fusion feature vectors, For characteristic cascade functions, and These are the learnable gate matrix and the learnable gate bias vector, respectively. for Activation function and These are the dynamic weights for the energy channel and the dynamic weights for the structure channel, respectively.

6. The fault intelligent diagnosis and localization method based on multi-feature fusion of acoustic, optical, and thermal features according to claim 4, characterized in that, Based on the dynamic weights of the energy channel and the structure channel, a weighted feature fusion is performed on the energy channel fusion feature vector and the structure channel fusion feature vector to obtain a multimodal fusion feature vector. This includes performing a weighted feature fusion on the energy channel fusion feature vector and the structure channel fusion feature vector using the following formula: ;in, This is a multimodal fusion feature vector.

7. The fault intelligent diagnosis and localization method based on multi-feature fusion of acoustic, optical, and thermal features according to claim 1, characterized in that, The predicted fault category and the original visible light image are used to locate and visualize the fault to obtain the final fault location map, including: Based on the predicted fault category, dual-channel parallel class activation is performed on the final feature map of the energy channel and the visual reconstruction residual map to obtain the energy channel class activation map and the structure channel class activation map. Based on the dynamic weights of the energy channel and the structure channel, an adaptive fusion of the activation maps of the energy channel class and the structure channel class is performed based on arbitration weights to obtain a fault location heatmap. The fault location heatmap is fused with the original visible light image to obtain the final fault location map.