Method, system and device for diagnosing defects of live equipment based on sound-heat and visible light data fusion

By simultaneously acquiring data using an acoustic array, an infrared thermal imager, and a visible light camera, multimodal data depth calibration and feature extraction are performed. Combined with adaptive beamforming and a deep learning model, the problem of lacking cross-modal fusion in existing technologies is solved, enabling efficient identification and diagnosis of internal defects in live equipment.

CN121899705BActive Publication Date: 2026-07-21SICHUAN YAAN ELECTRIC POWER (GRP) CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SICHUAN YAAN ELECTRIC POWER (GRP) CO LTD
Filing Date
2026-03-25
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

In existing defect detection technologies for live equipment, acoustic and thermal imaging data processing lacks cross-modal deep correlation and fusion algorithms, relies on human experience, lacks robustness, and is difficult to accurately identify early, weak, or internally hidden defects.

Method used

Data is collected simultaneously by an acoustic array, an infrared thermal imager, and a visible light camera. Multimodal data depth calibration and feature extraction are performed. An adaptive beamforming algorithm is used to suppress noise. A physical correlation model between sound velocity and temperature is constructed. Features are fused using a multi-channel deep learning model. Combined with historical detection records, a three-dimensional fused temperature field and defect diagnosis report are generated.

Benefits of technology

It enables the identification of internal defects in equipment under conditions of weak surface thermal imaging or visual invisibility, improving the defect detection rate and diagnostic confidence, and generating structured reports to facilitate on-site decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121899705B_ABST
    Figure CN121899705B_ABST
Patent Text Reader

Abstract

The application provides a charged equipment defect diagnosis method, system and equipment based on sound heat and visible light data fusion, acquires acoustic signals, infrared thermal image data and visible light images; based on visible light image features and infrared hot spot distribution, real-time alignment is performed in a spatial coordinate system, and a mapping relationship between a three-dimensional coordinate system and image pixel coordinates is established; an adaptive beam forming algorithm based on feature subspace decomposition is used to suppress environmental acoustic noise in the acoustic signals; according to the measured distance and environmental meteorological parameters obtained in real time, dynamic atmospheric transmittance compensation and radiation temperature inversion are performed on the infrared thermal image data, and the corrected absolute temperature field of the equipment surface is obtained; a three-dimensional fusion temperature field reflecting the internal medium state of the equipment is generated through an iterative algorithm; the preprocessed acoustic signals, three-dimensional fusion temperature field and visible light images are input into a trained multi-channel deep learning model, and thus the defect category, three-dimensional coordinates and risk index are output.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of defect detection technology for live equipment, and in particular to a method, system and device for defect diagnosis of live equipment based on the fusion of acoustic, thermal and visible light data. Background Technology

[0002] With the continuous growth of power system scale and the demand for automated operation and maintenance, the requirements for online, accurate, and early defect detection of live equipment are increasing. Power companies are gradually transforming from post-maintenance to condition-aware and predictive maintenance. To meet the needs of high-frequency inspections and uninterrupted operation on-site testing, non-destructive online sensing technologies such as acoustic and infrared thermal imaging have been widely introduced into on-site testing equipment, forming a monitoring system based on simultaneous acoustic-thermal dual-modal acquisition. This system is used to capture multi-source abnormal signals such as equipment discharge, contact overheating, and mechanical vibration, and to assist in operation and maintenance decision-making.

[0003] In related technologies, although there are detection devices that use common optical paths or coaxial structures to achieve integrated acoustic and thermal acquisition, these solutions usually only achieve "hard synchronization" in time and display space, treating acoustic and thermal images as independent data streams and lacking cross-modal deep correlation and fusion algorithms based on the physical mechanism of the fault. The diagnostic process still relies heavily on subjective judgment based on human experience and has failed to form an end-to-end automated intelligent diagnostic closed loop. At the same time, the existing systems are not robust enough to strong electromagnetic interference, background noise and changes in ambient temperature and humidity in complex sites such as substations, resulting in limited detection rate and diagnostic confidence for early weak or internal hidden defects. Summary of the Invention

[0004] In view of this, it is necessary to provide a method, system and device for defect diagnosis of live equipment based on the fusion of acoustic, thermal and visible light data, which can overcome at least one of the above defects.

[0005] In a first aspect, embodiments of this application provide a method for diagnosing defects in live electrical equipment based on the fusion of acoustic, thermal, and visible light data, including:

[0006] Acoustic signals, infrared thermal image data, and visible light images are acquired simultaneously using an acoustic array, an infrared thermal imager, and a visible light camera.

[0007] Based on visible light image features and infrared hotspot distribution, spatial coordinate system is aligned in real time, and depth information obtained by laser ranging device is used to perform depth calibration on multimodal data in order to establish the mapping relationship between the three-dimensional coordinate system and image pixel coordinates.

[0008] An adaptive beamforming algorithm based on feature subspace decomposition is used to suppress ambient acoustic noise in the acoustic signal;

[0009] A physical correlation model between sound speed and temperature is constructed, and the sound speed field distribution of the medium is initialized by applying the physical correlation model between sound speed and temperature to correct the sound wave propagation path and thus realize the sound source relocation.

[0010] The acoustic inversion temperature indication value is calculated by an iterative algorithm, and the absolute temperature field of the equipment surface is assimilated with the internal acoustic characteristics to generate a three-dimensional fused temperature field that reflects the state of the medium inside the equipment.

[0011] The preprocessed acoustic signal, the three-dimensional fused temperature field, and the visible light image are input into the trained multi-channel deep learning model.

[0012] The multi-channel deep learning model applies a cross-modal attention mechanism to fuse spatial features and combines historical detection records to perform time-series feature analysis, thereby filtering transient interference and identifying the development trend of defects, and outputting defect category, three-dimensional coordinates and risk index.

[0013] In one embodiment, the step of suppressing ambient acoustic noise in the acoustic signal using an adaptive beamforming algorithm based on feature subspace decomposition includes:

[0014] The acoustic signal is subjected to covariance matrix eigenvalue decomposition to identify and extract the signal subspace and noise subspace;

[0015] During the spatial beamforming process, a spatial null is constructed based on the direction of the interference source determined by the noise subspace, and gain compensation is performed on the direction of the target source pointed to by the signal subspace to extract the equipment discharge signal or equipment vibration signal.

[0016] In one embodiment, the method further includes:

[0017] Based on the real-time acquired measurement distance and environmental meteorological parameters, dynamic atmospheric transmittance compensation and radiation thermometry inversion are performed on the infrared thermal image data to obtain the corrected absolute temperature field of the equipment surface.

[0018] Access the meteorological monitoring interface to obtain environmental humidity, ambient temperature, and atmospheric pressure parameters;

[0019] The real-time atmospheric absorption coefficient is calculated based on the ambient humidity, ambient temperature, atmospheric pressure parameters, and measurement distance.

[0020] The dynamic atmospheric transmittance under the current path is calculated using a radiative transfer model, and the gain correction of the original infrared radiation energy is performed in combination with the preset equipment surface emissivity to eliminate temperature measurement deviations caused by environmental fluctuations and changes in detection distance.

[0021] In one embodiment, the iterative process of the physical correlation model between sound speed and temperature includes:

[0022] Based on the absolute temperature field, establish the spatial equivalent temperature distribution gradient and calculate the initial sound velocity value of the corresponding spatial grid node.

[0023] The refraction path tracing algorithm is used to correct the propagation trajectory of sound waves in non-uniform media;

[0024] By minimizing the weighted objective function of acoustic positioning residual and infrared measured temperature residual, the sound velocity distribution parameters of the medium inside the device are recursively updated, thereby obtaining the equivalent acoustic temperature inside the device.

[0025] In one embodiment, the method further includes:

[0026] The three-dimensional fused temperature field is converted into a multi-layer semi-transparent isothermal cloud map, and combined with the defect category and three-dimensional coordinate label, it is dynamically superimposed and rendered in the visible light real-time video stream according to the mapping relationship using augmented reality technology.

[0027] Based on the trend curve generated by the risk index and time series characteristics analysis, and in conjunction with the handling plans in the expert knowledge base, a diagnostic report is generated.

[0028] In one embodiment, the step of inputting the preprocessed acoustic signal, the three-dimensional fused temperature field, and the visible light image into the trained multi-channel deep learning model includes:

[0029] The apparent texture features of the visible light image and the spatial thermal distribution features of the three-dimensional fused temperature field are extracted using residual network branches, and the Mel-frequency cepstral coefficient features of the acoustic signal are extracted using a one-dimensional convolutional network.

[0030] The multi-channel deep learning model constructs a spatial cross-correlation matrix in the feature fusion layer to calculate the correlation weights of different modal feature maps at the same spatial location, thereby enhancing the ability to extract defect features.

[0031] In one embodiment, the step of combining historical detection records to perform time-series feature analysis includes:

[0032] The current frame fusion features are compared with the historical feature sequence within a preset time window, and the evolution trend of device status over time is extracted using a recurrent neural unit to obtain historical baseline values.

[0033] By calculating the offset and rate of change of the current state characteristics relative to historical benchmark values, transient disturbances caused by ambient temperature fluctuations or sudden load changes, as well as trend defects caused by insulation degradation or mechanical loosening, can be distinguished to improve the stability of diagnostic results.

[0034] In one embodiment, the synchronous acquisition and establishment of the mapping relationship between the three-dimensional coordinate system and the image pixel coordinates includes:

[0035] A field-programmable gate array is used to generate a synchronous pulse signal, which drives the acoustic sensor, infrared detector and visible light image sensor to perform equal-interval sampling;

[0036] By identifying a preset calibration reference in the target device area or utilizing the geometric edge features of the device itself, the perspective transformation matrix between multiple sensors is calculated to correct the parallax caused by the misalignment of the mounting axes of each sensor, thereby establishing the mapping relationship.

[0037] Secondly, embodiments of this application provide a fault diagnosis system for live equipment based on the fusion of acoustic, thermal, and visible light data, the system comprising:

[0038] The information acquisition and calibration module is used to simultaneously acquire acoustic signals, infrared thermal image data and visible light images from the acoustic array, infrared thermal imager and visible light camera; it performs real-time spatial coordinate system alignment based on visible light image features and infrared hotspot distribution, and applies depth information obtained by laser ranging device to perform depth calibration on multimodal data in order to establish the mapping relationship between the three-dimensional coordinate system and image pixel coordinates.

[0039] The data preprocessing module is used to suppress environmental acoustic noise in the acoustic signal using an adaptive beamforming algorithm based on feature subspace decomposition; and to perform dynamic atmospheric transmittance compensation and radiation thermometry inversion on the infrared thermal image data based on the real-time acquired measurement distance and environmental meteorological parameters to obtain the corrected absolute temperature field of the equipment surface.

[0040] The data fusion module is used to construct a physical correlation model between sound speed and temperature, and to initialize the sound speed field distribution of the medium using the sound speed and temperature physical correlation model to correct the sound wave propagation path and thus achieve sound source relocation; it calculates the acoustic inversion temperature indication value through an iterative algorithm, and assimilates the absolute temperature field of the equipment surface with the internal acoustic features to generate a three-dimensional fused temperature field reflecting the state of the medium inside the equipment; the preprocessed acoustic signal, the three-dimensional fused temperature field and the visible light image are input into a trained multi-channel deep learning model;

[0041] The defect detection module uses the multi-channel deep learning model to detect defects. The multi-channel deep learning model uses a cross-modal attention mechanism to fuse spatial features and combines historical detection records to perform time series feature analysis in order to filter transient interference and identify the development trend of defects, thereby outputting defect category, three-dimensional coordinates and risk index.

[0042] Thirdly, embodiments of this application provide an electronic device, including:

[0043] Processor; and

[0044] The memory stores computer-readable instructions for controlling the processor to execute the electrical equipment defect diagnosis method based on acoustic, thermal, and visible light data fusion as described in the second aspect.

[0045] This application provides a method, system, and device for defect diagnosis of live equipment based on the fusion of acoustic, thermal, and visible light data. Through pixel-level registration and depth calibration of acoustic, thermal, and optical modalities, combined with correction of the sound wave propagation path using a sound speed-temperature-based physical model, the system can recover the equivalent temperature field inside the equipment and achieve three-dimensional sound source relocation even when surface thermal images are weak or visually invisible. Simultaneously, dynamic atmospheric transmittance compensation is applied to infrared data to obtain the physical absolute temperature. A multi-channel deep learning model fuses modal features under spatial cross-correlation and cross-modal attention mechanisms, and combines historical sequence analysis to output defect categories, three-dimensional coordinates, confidence levels, and risk indices. Finally, the diagnostic results are presented intuitively in visible light video in the form of a semi-transparent isothermal cloud map and augmented reality overlay, generating a structured report for easy on-site decision-making and maintenance execution. Attached Figure Description

[0046] Figure 1 This is a schematic flowchart of a method for diagnosing defects in live equipment based on the fusion of acoustic, thermal, and visible light data, provided in an embodiment of this application.

[0047] Figure 2 This is a schematic diagram of the overlay of visible light video frames and diagnostic results provided in an embodiment of this application.

[0048] Figure 3 This is a schematic diagram of the overlay of visible light video frames and diagnostic results provided in another embodiment of this application.

[0049] Figure 4 This is a schematic diagram of a defect diagnosis system module for live equipment based on the fusion of acoustic, thermal, and visible light data, provided in an embodiment of this application.

[0050] Figure 5 This is a schematic diagram of the modules of an electronic device provided in an embodiment of this application.

[0051] Explanation of main component symbols

[0052] Detailed Implementation

[0053] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them.

[0054] It should be noted that, in the embodiments of this application, "at least one" refers to one or more, and "more than one" refers to two or more. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in the specification of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application.

[0055] It should be noted that in the embodiments of this application, the terms "first," "second," etc., are used only for descriptive purposes and should not be construed as indicating or implying relative importance, nor as indicating or implying order. Features specified as "first" or "second" may explicitly or implicitly include one or more of the stated features. In the description of the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design scheme described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0056] Based on the embodiments described in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0057] With the continuous growth of power system scale and the demand for automated operation and maintenance, the requirements for online, accurate, and early defect detection of live equipment are increasing. Power companies are gradually transforming from post-maintenance to condition-aware and predictive maintenance. To meet the needs of high-frequency inspections and uninterrupted operation on-site testing, non-destructive online sensing technologies such as acoustic and infrared thermal imaging have been widely introduced into on-site testing equipment, forming a monitoring system based on simultaneous acoustic-thermal dual-modal acquisition. This system is used to capture multi-source abnormal signals such as equipment discharge, contact overheating, and mechanical vibration, and to assist in operation and maintenance decision-making.

[0058] In related technologies, although there are detection devices that use common optical paths or coaxial structures to achieve integrated acoustic and thermal acquisition, these solutions usually only achieve "hard synchronization" in time and display space, treating acoustic and thermal images as independent data streams and lacking cross-modal deep correlation and fusion algorithms based on the physical mechanism of the fault. The diagnostic process still relies heavily on subjective judgment based on human experience and has failed to form an end-to-end automated intelligent diagnostic closed loop. At the same time, the existing systems are not robust enough to strong electromagnetic interference, background noise and changes in ambient temperature and humidity in complex sites such as substations, resulting in limited detection rate and diagnostic confidence for early weak or internal hidden defects.

[0059] In view of this, the method, system, and equipment for defect diagnosis of live equipment based on the fusion of acoustic, thermal, and visible light data provided in this application, through pixel-level registration and depth calibration of acoustic, thermal, and optical modes, and combined with the correction of sound wave propagation path based on a sound speed-temperature physical model, enables the system to recover the equivalent temperature field inside the equipment and achieve three-dimensional sound source relocation even when the surface thermal image is weak or visually invisible. Simultaneously, dynamic atmospheric transmittance compensation is applied to infrared data to obtain the physical absolute temperature. A multi-channel deep learning model fuses modal features under spatial cross-correlation and cross-modal attention mechanisms, and combines historical sequence analysis to output defect category, three-dimensional coordinates, confidence level, and risk index. Finally, the diagnostic results are presented intuitively in visible light video in the form of a semi-transparent isothermal cloud map and augmented reality overlay, generating a structured report to facilitate on-site decision-making and maintenance execution.

[0060] Figure 1 This is a schematic flowchart of a method for diagnosing defects in live equipment based on the fusion of acoustic, thermal, and visible light data, provided in an embodiment of this application. Figure 1 The method for defect diagnosis of live equipment based on the fusion of acoustic, thermal, and visible light data, as shown, includes at least the following steps: S100: Acoustic signals, infrared thermal image data, and visible light images are simultaneously acquired using an acoustic array, an infrared thermal imager, and a visible light camera; S200: Spatial coordinate system is aligned in real time based on visible light image features and infrared hotspot distribution, and depth information obtained from a laser ranging device is used to perform depth calibration on the multimodal data to establish a mapping relationship between the three-dimensional coordinate system and the image pixel coordinates; S300: Environmental acoustic noise in the acoustic signal is suppressed using an adaptive beamforming algorithm based on feature subspace decomposition; S400: Dynamic atmospheric transmittance compensation and radiation thermometry inversion are performed on the infrared thermal image data based on the real-time acquired measurement distance and environmental meteorological parameters to obtain the corrected equipment surface... Absolute temperature field; S500: Construct a physical correlation model between sound speed and temperature, and apply the physical correlation model to initialize the sound speed field distribution of the medium to correct the sound wave propagation path and thus achieve sound source relocation; S600: Calculate the acoustic inversion temperature indication value through an iterative algorithm, and assimilate the absolute temperature field of the equipment surface with the internal acoustic features to generate a three-dimensional fused temperature field reflecting the state of the internal medium of the equipment; S700: Input the preprocessed acoustic signal, the three-dimensional fused temperature field and the visible light image into the trained multi-channel deep learning model; S800: The multi-channel deep learning model applies a cross-modal attention mechanism to fuse spatial features and combines historical detection records to perform time series feature analysis to filter transient interference and identify the development trend of defects, thereby outputting defect category, three-dimensional coordinates and risk index.

[0061] S100: Simultaneously acquires acoustic signals, infrared thermal image data, and visible light images through an acoustic array, an infrared thermal imager, and a visible light camera.

[0062] In this embodiment of the application, the method for diagnosing defects in live equipment based on the fusion of acoustic, thermal and visible light data includes, in step S100: simultaneously acquiring acoustic signals, infrared thermal image data and visible light images through an acoustic array, an infrared thermal imager and a visible light camera.

[0063] Specifically, in step S100, an acoustic array composed of linear or area array acoustic sensors, an infrared thermal imager with known radiation characteristics and pixel calibration information, and a high-resolution visible light camera simultaneously acquire data; acoustic channel bandpass filtering and anti-aliasing sampling are completed at the sensor end; visible light and infrared images are synchronized frame by frame and the timestamps of hardware triggering or network time synchronization (such as PTP, GPS, or hardware triggering signals) are recorded; the laser ranging device outputs the distance measurement of the corresponding field of view within the same sampling period; the acquired data is buffered and written to local or cloud storage according to a unified time reference for subsequent registration and processing.

[0064] Understandably, by ensuring the temporal consistency of multimodal data at both the physical level (hard triggering or timestamp) and the sampling level, a reliable spatiotemporal correspondence can be provided for subsequent pixel-level registration, acoustic-thermal correlation, and time-series analysis, thereby reducing errors caused by asynchrony and improving diagnostic reliability.

[0065] S200: Based on visible light image features and infrared hotspot distribution, the spatial coordinate system is aligned in real time, and depth information obtained by the laser rangefinder is used to perform depth calibration on multimodal data in order to establish a mapping relationship between the three-dimensional coordinate system and the image pixel coordinates.

[0066] In this embodiment of the application, the method for diagnosing defects in live equipment based on the fusion of acoustic, thermal and visible light data includes the following steps in step S200: real-time alignment of the spatial coordinate system based on visible light image features and infrared hotspot distribution, and depth calibration of multimodal data using depth information obtained by a laser ranging device, so as to establish a mapping relationship between the three-dimensional coordinate system and the image pixel coordinates.

[0067] Specifically, firstly, structured feature points (such as corner points, edges, or semantic key points) are extracted from visible light images, and hotspot centers and temperature gradient features are extracted from infrared thermal images. An initial correspondence is established based on cross-modal matching of visible light and infrared features. Initial mapping from image pixels to 3D camera coordinates is calculated using monocular or binocular camera calibration parameters and extrinsic parameter estimation. Subsequently, sparse or dense point clouds or depth values ​​are obtained using laser ranging, and the depth information is aligned pixel by pixel with the pixel coordinates through point cloud registration (such as ICP or feature-based point cloud registration). Finally, a mapping table or projection matrix from each sensor's field of view to the device's 3D coordinate system is generated (and the internal and external calibration errors are recorded).

[0068] Understandably, this step not only solves the spatial deviation problem between thermal images and visible light, but also introduces real depth measurement for pixel-level calibration, making subsequent positioning, 3D rendering, and accurate annotation of defect coordinates on the visible light image possible, and providing spatial scale information for physical modeling.

[0069] S300: Suppresses ambient acoustic noise in acoustic signals using an adaptive beamforming algorithm based on feature subspace decomposition.

[0070] In this embodiment of the application, the method for diagnosing defects in live equipment based on the fusion of acoustic, thermal and visible light data includes, in step S300: using an adaptive beamforming algorithm based on feature subspace decomposition to suppress ambient acoustic noise in the acoustic signal.

[0071] Specifically, the array acoustic signal is first subjected to short-time Fourier transform or time-frequency analysis to obtain a spectral-time domain representation. Adaptive weights are constructed using feature subspace decomposition (e.g., separating the signal subspace and noise subspace based on SVD or PCA). The array weighting coefficients are dynamically calculated using beamforming techniques (e.g., minimum variance distortionless response MVDR or normalized delay summation) to suppress time-varying environmental noise and enhance the energy of the target sound source. Acoustic features such as time difference of arrival (TDOA), azimuth angle (DOA), instantaneous sound pressure level, and power spectrum are extracted from the processed signal, and the sound source confidence is estimated.

[0072] Understandably, adaptively separating background noise and suppressing it during beamforming using the feature subspace method can significantly improve the signal-to-noise ratio and positioning accuracy of the sound source, thereby providing a more reliable input for subsequent acoustic-based internal state inversion and assimilation with the thermal field.

[0073] S400: Based on the real-time acquired measurement distance and environmental meteorological parameters, it performs dynamic atmospheric transmittance compensation and radiation thermometry inversion on the infrared thermal image data to obtain the corrected absolute temperature field of the equipment surface.

[0074] In this embodiment of the application, the defect diagnosis method for live equipment based on the fusion of acoustic, thermal and visible light data includes the following steps in step S400: dynamic atmospheric transmittance compensation and radiation thermometry inversion are performed on the infrared thermal image data according to the real-time acquired measurement distance and environmental meteorological parameters to obtain the corrected absolute temperature field of the equipment surface.

[0075] Specifically, based on the object distance obtained by laser ranging and the environmental meteorological parameters (including ambient temperature, relative humidity, atmospheric pressure, visibility, or aerosol optical thickness) measured at the time of acquisition, an atmospheric transmittance model is used to compensate for the scene radiation of the infrared thermal image by distance and atmospheric absorption or scattering. Combining the sensor response function and the emissivity information of the target surface, the corrected radiance is mapped to the absolute temperature field of the equipment surface through a radiation thermometry inversion algorithm (and confidence level or correction strategy is given for low emissivity areas or areas with reflection interference).

[0076] Understandably, by dynamically considering distance measurement and environmental parameters for transmittance compensation and temperature inversion, we can obtain the absolute temperature value in a physical sense, eliminate temperature deviations caused by differences in distance and weather conditions, and ensure the comparability of observation results at different times and locations and the accuracy of physical modeling input.

[0077] S500: Construct a physical correlation model between sound speed and temperature, and apply the physical correlation model between sound speed and temperature to initialize the sound speed field distribution of the medium in order to correct the sound wave propagation path and thus achieve sound source relocation.

[0078] In this embodiment of the application, the defect diagnosis method for live equipment based on the fusion of acoustic, thermal and visible light data includes the following steps in step S500: constructing a physical correlation model between sound speed and temperature, and applying the physical correlation model between sound speed and temperature to initialize the sound speed field distribution of the medium in order to correct the sound wave propagation path and thereby achieve sound source relocation.

[0079] Specifically, a sound speed-temperature physical correlation model is established (e.g., based on the relationship between the physical properties of the medium, an approximation is made). Or a more precise empirical or theoretical expression. Indicates temperature as Local sound velocity at time Reference temperature The speed of sound at that time Temperature at the current spatial location, (Reference temperature), using the absolute surface temperature field obtained in step S400 as the boundary or initial value to interpolate the internal temperature distribution of the medium or to obtain the initial three-dimensional sound velocity field distribution based on physical prior inference; based on this sound velocity field, acoustic ray tracing or wave numerical simulation is used to correct the sound wave propagation path, and combined with the measurement information such as TDOA or DOA in step S300, the sound source relocation and local sound field structure estimation are achieved through time difference inversion or wave field inversion methods.

[0080] It is understandable that temperature has a direct impact on the speed of sound. Using thermal field information to initialize the speed of sound and correct the propagation model can reduce the positioning error in non-uniform media, and has a significant improvement effect on the acoustic positioning of internal or hidden defects.

[0081] S600: It calculates the acoustic inversion temperature indication value through an iterative algorithm, and assimilates the absolute temperature field of the equipment surface with the internal acoustic characteristics to generate a three-dimensional fused temperature field that reflects the state of the medium inside the equipment.

[0082] In this embodiment of the application, the defect diagnosis method for live equipment based on the fusion of acoustic, thermal and visible light data includes the following steps in step S600: calculating the acoustic inversion temperature indication value through an iterative algorithm, and assimilating the absolute temperature field of the equipment surface with the internal acoustic features to generate a three-dimensional fused temperature field that reflects the state of the internal medium of the equipment.

[0083] Specifically, an acoustic-thermal forward model is constructed to characterize the influence of the internal medium temperature distribution on the acoustic response (such as amplitude, phase, and time delay spectrum). An iterative inversion algorithm (such as regularized least squares inversion, Bayesian assimilation, or variational assimilation) is used to combine the acoustic features obtained in step S300 with the surface temperature field obtained in step S400 to solve the internal temperature indication field. During the iteration process, the surface temperature is used as a boundary constraint and spatial smoothing and physical priors are introduced to stabilize the inversion. Finally, a three-dimensional fused temperature field represented in voxel or mesh form and the corresponding uncertainty assessment are output.

[0084] It is understandable that by assimilating acoustic information and surface thermal information, the internal temperature or medium state of the equipment can be restored to a certain extent, enabling three-dimensional indication of internal local overheating or discharge points, thereby compensating for hidden defects that are difficult to detect by a single mode (sound or heat only).

[0085] S700: Input the pre-processed acoustic signal, the three-dimensional fused temperature field, and the visible light image into the trained multi-channel deep learning model.

[0086] In this embodiment of the application, the method for diagnosing defects in live equipment based on the fusion of acoustic, thermal and visible light data includes step S700: inputting the preprocessed acoustic signal, the three-dimensional fused temperature field and the visible light image into the trained multi-channel deep learning model.

[0087] Specifically, the acoustic features output in step S300 (such as TDOA or DOA, time spectrum, waveform features), the three-dimensional fused temperature field generated in step S600 (represented by a fixed voxel size or multi-resolution pyramid and projected onto the same coordinate system as visible light), and the visible light image (containing semantic or structural features) are normalized, interpolated, or resampled and data augmented, and then fed into the trained multi-channel deep learning model as multiple input channels. The model can adopt a multi-branch structure (2D-CNN, 3D-CNN, or a hybrid of Transformer). Each branch is responsible for modal feature extraction in the shallow layer, and information interaction and alignment are performed in the middle and high layers through the cross-modal fusion module. The model outputs intermediate representations for defect classification, localization, and risk assessment and calculates the corresponding loss for online or offline retraining.

[0088] Understandably, using preprocessed multi-source heterogeneous features as model input preserves physical priors (from the assimilation and localization process) while utilizing deep models to mine cross-modal high-order features, which helps improve classification accuracy, reduce false alarms, and enhance the ability to identify defect types under complex working conditions.

[0089] S800: The multi-channel deep learning model applies a cross-modal attention mechanism to fuse spatial features and combines historical detection records to perform time series feature analysis in order to filter transient interference and identify the development trend of defects, thereby outputting defect category, three-dimensional coordinates and risk index.

[0090] In this embodiment of the application, the method for diagnosing defects in live equipment based on the fusion of acoustic, thermal and visible light data includes the following steps in step S800: a multi-channel deep learning model applies a cross-modal attention mechanism to fuse spatial features and combines historical detection records to perform time series feature analysis in order to filter transient interference and identify the development trend of defects, thereby outputting defect category, three-dimensional coordinates and risk index.

[0091] Specifically, the multi-channel deep learning model integrates a cross-modal attention mechanism (for weighted integration of spatial features) and a time series module (such as Temporal Transformer, TCN, or LSTM) to access historical detection records and perform sequence modeling; it detects transient interference by comparing the changing trends of current and historical features and performs smoothing or discrimination; and it outputs the defect category, the three-dimensional coordinates of the defect in the device's three-dimensional coordinate system, the defect risk index (weighted by defect severity, location importance, and development rate), and its confidence level based on classification, regression, and confidence branches in parallel; the model can also provide defect development trend predictions and suggested maintenance levels.

[0092] Understandably, cross-modal attention ensures the correct weighting of modal information under different observation conditions, while time series analysis can filter out sudden disturbances and capture the evolution trend of defects, so that the diagnostic results can not only provide a judgment on the current state, but also support risk ranking and predictive maintenance decisions.

[0093] In this embodiment, an adaptive beamforming algorithm based on eigenspace decomposition is used to suppress ambient acoustic noise in acoustic signals. This includes: performing covariance matrix eigenvalue decomposition on the acoustic signal to identify and extract the signal subspace and noise subspace. During spatial beamforming, spatial nulls are constructed based on the interference source direction determined by the noise subspace, and gain compensation is performed on the target source direction pointed to by the signal subspace to extract equipment discharge signals or equipment vibration signals.

[0094] Specifically, the steps include: estimating the covariance matrix from the acquired array time-domain signal.

[0095]

[0096] in For array receive vector, This indicates the conjugate transpose. This indicates the expected value or the average value of a sliding window.

[0097] right Perform eigenvalue decomposition (EVD) and sort by energy to separate the signal subspace from the noise subspace:

[0098]

[0099] Will ,in The eigenvector matrix of the signal subspace. Let be the eigenvector matrix of the noise subspace.

[0100] Construct the spatial null projection matrix based on the noise subspace:

[0101]

[0102] in This is the identity matrix. The projection is used to eliminate noise subspace contributions.

[0103] Optimization using subspace constraints in beamforming (taking MVDR as an example):

[0104]

[0105] in For array weighting vectors, For the target direction The guiding vector. This can be introduced when there are subspace constraints. Or directly restrict in the optimization constraints and Orthogonality.

[0106] To adapt to time-varying noise, the covariance matrix is ​​updated using exponential weighting or a sliding window:

[0107]

[0108] The processed output includes the enhanced time-domain waveform, TDOA or DOA estimates, sound pressure level (SPL) time series, and corresponding confidence indices (which can be quantized by subspace energy ratio or estimated signal-to-noise ratio). This is the instantaneous sampling vector.

[0109] Understandably, through the above subspace projection and constraint optimization, nulls can be formed in the spatial domain under complex backgrounds and multiple interference sources, and gain compensation can be implemented for the target direction, thereby improving the distinguishability and positioning accuracy of the target sound source; the use of sliding or weighted updates improves the adaptability to environmental changes, and the confidence output can be used for uncertainty weighting during subsequent assimilation and fusion.

[0110] In this embodiment, dynamic atmospheric transmittance compensation and radiation thermometry inversion are performed on infrared thermal image data, including: retrieving environmental humidity, environmental temperature, and atmospheric pressure parameters from a meteorological monitoring interface; calculating the real-time atmospheric absorption coefficient based on the environmental humidity, environmental temperature, atmospheric pressure parameters, and measurement distance; calculating the dynamic atmospheric transmittance under the current path using a radiative transfer model; and performing gain correction on the raw infrared radiation energy in conjunction with a preset device surface emissivity to eliminate temperature measurement deviations caused by environmental fluctuations and changes in detection distance.

[0111] Specifically, the steps include: obtaining environmental and ranging parameters: ambient temperature T, relative humidity RH, air pressure p, and distance d given by the rangefinder.

[0112] Estimating the unit path absorption coefficient ( (where the center wavelength of the operating band is used), and the transmittance is calculated.

[0113]

[0114] Correcting observed radiance using radiative transfer relationships , to obtain the target radiance :

[0115]

[0116] in This represents the atmospheric self-radiation component along the path (which can be estimated by an atmospheric model).

[0117] Considering the target surface emissivity Inverting radiance into brightness temperature or absolute temperature (Based on Planck's equation or sensor calibration curve):

[0118]

[0119] in This represents the blackbody radiance (Planck function). The reflectance brightness (estimated using visible light images or multi-frame methods) is given by the Planck function:

[0120]

[0121] in Let be Planck's constant. At the speed of light, is the Boltzmann constant.

[0122] Solve based on the above relationships. And provide a confidence estimate for each pixel (considering) , , (uncertainty propagation).

[0123] Understandably, using real-time transmittance correction combined with emissivity or reflection correction can significantly reduce the deviation of temperature measurement caused by distance and environmental changes, making the obtained temperature field physically consistent and suitable for subsequent physical modeling and assimilation.

[0124] In this embodiment, the iterative process of the sound velocity and temperature physical correlation model includes: establishing a spatial equivalent temperature distribution gradient based on the absolute temperature field and calculating the initial sound velocity values ​​for the corresponding spatial grid nodes; correcting the propagation trajectory of sound waves in a non-uniform medium using a refraction path tracing algorithm; and recursively updating the sound velocity distribution parameters of the internal medium of the device by minimizing a weighted objective function of the acoustic positioning residual and the infrared measured temperature residual, thereby inverting to obtain the equivalent acoustic temperature inside the device.

[0125] Specifically, the steps include: Meshing and initialization: dividing the three-dimensional domain into voxels or finite element meshes based on boundary (surface) temperature. Initialize internal temperature According to the sound speed-temperature relationship Obtain the initial sound velocity field .

[0126] Acoustic forward modeling and residual calculation: in the k-th iteration, based on Calculate the theoretical arrival time or acoustic response k (e.g., through ray tracing or numerical fluctuation simulation) and obtain the acoustic residual.

[0127]

[0128] Temperature residual: will be determined by the current field The predicted surface temperature is compared with the measured infrared temperature to obtain the temperature residual.

[0129]

[0130] in, Indicates the current iteration number. Indicates the first The acoustic medium parameter field obtained in the next iteration This represents measured acoustic observation data. Indicates the first The theoretical acoustic response calculated using the forward model in the next iteration. This represents the difference between observed and predicted values. Indicates the first The temperature field distribution obtained in the next iteration, This represents the actual temperature data obtained from infrared measurements. Indicates the first In the next iteration, the current temperature field The predicted temperature obtained from the calculation This represents the difference between the measured infrared temperature and the model-predicted temperature.

[0131] Define the weighted objective function and solve and update it:

[0132]

[0133] in , As the acoustic and thermal residual weights, For regularization terms (e.g., smoothing terms). This is the regularization coefficient. It is obtained through gradient descent or variational methods. And through the inverse function Update the temperature field.

[0134] Convergence criterion: for example, when Alternatively, iteration may stop once the maximum number of iterations is reached.

[0135] Understandably, by simultaneously minimizing acoustic and temperature observation errors and introducing regularization constraints, this iterative assimilation process gradually approximates the equivalent acoustic temperature inside the device under physical consistency, and adjusts... , The relative influence of the two modes in the inversion can be controlled.

[0136] In this embodiment, the method further includes: converting the three-dimensional fused temperature field into a multi-layer semi-transparent isothermal cloud map, and combining it with defect categories and three-dimensional coordinate labels, dynamically overlaying and rendering it onto a real-time visible light video stream according to a mapping relationship using augmented reality technology. Based on the trend curve generated by the risk index and time series feature analysis, and in conjunction with processing plans in the expert knowledge base, a diagnostic report is generated.

[0137] Specifically, the steps include: isothermal layer and projection: the voxel temperature field is divided into isothermal layers to obtain a set of isothermal surfaces. Assign transparency to each isothermal surface Color mapping; using projection matrices 3D coordinates Projection as pixel coordinates :

[0138]

[0139] Multi-layer rendering and interaction: Composite by layers (visible light background + isothermal cloud map + defect annotation + information panel), supports transparency adjustment and interactive query; defect labels include information such as 3D coordinates, risk index R, confidence level p.

[0140] Report generation: Key fields (collection time, sensor parameters, positioning accuracy, maximum or average temperature, risk score, recommended contingency plan) are structured into report formats (such as JSON or PDF) and archived synchronously.

[0141] Understandably, through the above projection and multi-layer rendering, maintenance personnel can intuitively see the internal and surface thermal information and defect locations on the visible light real-time stream, while the structured report provides reliable data support for subsequent trend analysis and model training.

[0142] In this embodiment, the preprocessed acoustic signal, the three-dimensional fused temperature field, and the visible light image are input into a trained multi-channel deep learning model. This includes: extracting the apparent texture features of the visible light image and the spatial thermal distribution features of the three-dimensional fused temperature field using residual network branches, and extracting the Mel-frequency cepstral coefficient features of the acoustic signal using a one-dimensional convolutional network. The multi-channel deep learning model constructs a spatial cross-correlation matrix at the feature fusion layer to calculate the correlation weights of different modal feature maps at the same spatial location, thereby enhancing the ability to extract defect features.

[0143] Specifically, the implementation steps include: performing frame segmentation on the acquired acoustic waveform according to the sampling rate (e.g., 48kHz) (frame length approximately 20–30ms, frame shift 10ms), calculating the Mel spectrum for each frame, taking the logarithm, and then performing a discrete cosine transform to obtain Mel cepstral coefficients (MFCCs) to obtain a time-band feature sequence; feeding this sequence into a one-dimensional convolutional network (e.g., several layers of 1D-Conv+BatchNorm+ReLU+pooling) to extract temporal local patterns, and finally converging them into a fixed-length acoustic feature vector or acoustic feature map through global temporal pooling or learnable attention.

[0144] Visible light images are input into residual network branches (e.g., several residual blocks of a ResNet-like structure) to extract multi-scale apparent texture and structural features; the 3D fused temperature field can be projected onto a multi-layer heatmap using voxel projection, or directly input into 3D-ResNet or 3D-CNN branches to extract volumetric thermal distribution features, or a U-Net-like structure can be used to preserve multi-scale spatial information. The output of each branch is a shape-aligned feature map (or aligned to the same H×W grid after projection).

[0145] Let the visible light characteristic map be Temperature characteristic map is For each spatial location Calculate the dot product similarity and normalize it to generate association weights:

[0146]

[0147] in To prevent division by zero for small constants, This is the temperature scaling factor. Based on... Weighted fusion of features between modalities can be performed; for example, the fused features can be represented as...

[0148]

[0149] in This is a channel alignment transformation (1×1 convolution). These are learnable modal weights. The acoustic branch, after time alignment, is injected into the system via attention or gating. The channel dimension or multimodal interactions can be processed in parallel in subsequent fusion layers.

[0150] The fused features are fed into a multi-task head, which outputs (a) defect category (classification head, cross-entropy loss); (b) defect 3D location or regression offset (localization head, Smooth-L1 loss); and (c) risk index or severity regression (MSE loss). A confidence branch is also introduced. Apply uncertainty weighting (e.g., uncertainty loss by Nix & Weigend) and a time consistency regularization term. During training, a joint loss weighting strategy is used, and dynamic weight adjustment (uncertainty weighting or manual weighting) is available.

[0151] Understandably, by extracting appearance and semantic texture through residual networks, preserving the spatial structure of the temperature field through 3D or voxel branches, extracting acoustic time-frequency features through 1D-Conv, and calculating the point-by-point correlation weights between modes using cross-correlation matrices at the spatial level, the model can automatically determine the importance of each mode at the same spatial location, thereby enhancing the ability to perceive hidden or weak signal defects. Multi-task output and confidence estimation help to provide interpretable categories, precise locations, and risk measurements simultaneously in practical applications, facilitating operation and maintenance decisions.

[0152] In this embodiment, time-series feature analysis is performed in conjunction with historical detection records, including: comparing the fused features of the current frame with the historical feature sequence within a preset time window, and using a recurrent neural unit to extract the evolution trend of the device state over time to obtain a historical baseline value. By calculating the offset and rate of change of the current state features relative to the historical baseline value, transient interference caused by environmental temperature fluctuations or sudden load changes, as well as trend defects caused by insulation degradation or mechanical loosening, are distinguished to improve the stability of the diagnostic results.

[0153] Specifically, historical windowing and normalization include setting the length of the time window. (For example, 10–60 detection frames, which can be adjusted according to the inspection frequency), for historical fusion feature sequences Perform intra-batch normalization and detrending (e.g., subtract the exponential moving average EMA) to eliminate long-term drift.

[0154] Temporal modeling involves inputting the normalized sequence into a temporal network (such as a bidirectional LSTM, temporal Transformer, or Temporal Convolution Network) to extract the hidden states. And from this, the historical baseline value is obtained. (For example, the reconstructed value output by an EMA or a timing encoder).

[0155] Offset and rate determination includes calculating the difference between the current feature and the baseline. And calculate the first or second order difference (velocity and acceleration) to assess the rate of change; for Statistical tests (e.g., z-score) or threshold rules are applied to determine whether the shift is significant. If the shift consistently exceeds the threshold and remains within the window (e.g., for M consecutive frames), it is marked as a "trend anomaly"; if the shift is brief and recovers within a short time, it is considered a "transient interference".

[0156] The decision-making and output fusion process involves feeding the time-series judgment results as input to the main classification or localization branch (e.g., by adjusting the final confidence level or threshold), and adding trend weights to the risk index calculation to reflect the evolving risk. Historical records are also used for online fine-tuning or adaptive threshold updates.

[0157] Understandably, by combining historical benchmarks with speed or persistence criteria, short-term anomalies caused by environment or operation can be effectively filtered out, reducing false alarms; at the same time, it is sensitive to small deviations that are continuously deteriorating, and can elevate early trend defects to high priority, thereby improving predictive maintenance capabilities and diagnostic stability.

[0158] In this embodiment, the synchronous acquisition and establishment of the mapping relationship between the three-dimensional coordinate system and the image pixel coordinates includes: generating a synchronization pulse signal using a field-programmable gate array (FPGA) and driving the acoustic sensor, infrared detector, and visible light image sensor to perform equally spaced sampling. By identifying a preset calibration reference object in the target device area or utilizing the device's own geometric edge features, the perspective transformation matrix between the multiple sensors is calculated to correct the parallax caused by the misalignment of the mounting axes of each sensor, thereby establishing the mapping relationship.

[0159] Specifically, time synchronization is achieved by having the FPGA or real-time control unit generate hardware synchronization pulses and broadcast them to the trigger ports of each sensor, or by using PTP or GPS to assist in secondary verification of the timestamp to ensure that the sampling time deviation is within a controllable range (e.g., <±1ms); each frame of data is accompanied by a unified timestamp and written to the buffer.

[0160] The spatial calibration process includes setting up reference objects with known geometric dimensions (such as checkerboards, April tags, or reflective markers) and acquiring multi-view samples in visible light or thermal imaging during the calibration phase; and solving for camera intrinsic parameters using conventional camera calibration methods (such as Zhang calibration). Solve for extrinsic parameters using solvePnP If laser ranging or point cloud is used, the point cloud is aligned with the camera coordinate system via ICP or feature-based point cloud-image registration, and the projection matrix is ​​finally obtained.

[0161]

[0162] It is used to project three-dimensional coordinates onto the pixel plane.

[0163] Runtime calibration includes performing online extrinsic fine-tuning (e.g., online pose optimization based on vision-laser integration) periodically or triggered (when a change in the calibration board or scene is detected) to address extrinsic drift caused by vibration or slight displacement.

[0164] Understandably, by ensuring time alignment through hardware triggering and establishing robust intrinsic and extrinsic parameter mapping through calibration references or point cloud registration, the three-dimensional defect coordinates can be accurately projected into pixel coordinates on the visible light video stream, thereby achieving augmented reality-style visual annotation and spatially consistent positioning display; at the same time, the online fine-tuning mechanism ensures that the system can maintain mapping accuracy during long-term field deployment, reducing installation and maintenance costs.

[0165] In the embodiments of this application, the following are combined with Figure 2 and Figure 3 The complete workflow of the defect diagnosis method for live equipment based on the fusion of acoustic, thermal and visible light data provided in this application is illustrated by way of example. Figure 2This is a schematic diagram of the overlay of visible light video frames and diagnostic results provided in an embodiment of this application. Figure 3 This is a schematic diagram of the overlay of visible light video frames and diagnostic results provided in another embodiment of this application. Figure 2 The Chinese method identified two insulation bushing anomalies on the visible light image using marked boxes and text annotations, which were determined to be "floating discharges" with confidence levels of 0.85 and 0.66, respectively. Figure 3 An SF6 circuit breaker (SF6CB) was identified in the image as exhibiting a "point discharge" phenomenon with a confidence level of 0.86. Voxel-level temperature indications and uncertainty assessments were provided for the three suspected defect locations, and the defects were dynamically rendered in the visible light flow as semi-transparent isothermal cloud maps and defect IDs for intuitive viewing by maintenance personnel.

[0166] Specifically, the process is as follows: In step S100, the system synchronously acquires data using an acoustic array, an infrared thermal imager, and a visible light camera (time synchronization is ensured by hard triggering, PTP, or GPS); in step S200, cross-modal registration is performed based on visible light feature points and infrared hotspots, and pixel-level depth calibration is obtained using laser ranging, establishing a mapping between the three-dimensional coordinate system and the pixel plane; in step S300, covariance matrix eigenvalue decomposition is performed on the array signal, and subspace-based adaptive beamforming is applied to denoise the signal to obtain high-confidence TDOA or DOA and acoustic spectral features; in step S400, atmospheric transmittance compensation is performed by combining real-time ranging and meteorological parameters, and pixel-level absolute temperature field is obtained through radiometric inversion; in step S500, based on surface... Temperature initialization of the sound velocity field and refraction or ray tracing correction of the sound wave propagation path are used to achieve sound source relocation. In S600, an iterative assimilation framework is used to jointly invert the acoustic inversion temperature indication and the surface temperature field to generate a three-dimensional fused temperature field and estimate the uncertainty of the field. In S700, the preprocessed acoustic features (such as MFCC and TDOA or DOA), the three-dimensional fused temperature voxel mapping (aligned to an H×W grid according to projection), and the visible light image are respectively fed into the multi-channel depth model to extract modal features, and the point-by-point modal weights are calculated through the spatial cross-correlation matrix in the feature fusion layer. In S800, the model outputs the defect category, three-dimensional coordinates, confidence level, and risk index based on cross-modal attention and time series modules combined with historical detection records. Taking this example, after the fusion layer and time series analysis, the model gives the "floating discharge" label and confidence levels of 0.85 and 0.66 to the two insulating bushings, respectively, and calculates their three-dimensional coordinates (already mapped to the equipment's three-dimensional coordinate system); a confidence level of 0.86 is given for the tip discharge of SF6CB. Subsequently, the rendering module maps the 3D coordinates onto the visible light pixel plane using a projection matrix.

[0167] Understandably, through the aforementioned end-to-end acoustic-thermal-optical multimodal fusion and time-series analysis, this example can improve defect detection rate and location accuracy even when a single mode is difficult to distinguish or is masked by environmental noise. On the one hand, acoustic inversion provides depth or direction indication for internal or concealed discharge sources, thermal imaging provides surface temperature evidence, and visual images provide structural semantics and location references. On the other hand, cross-modal cross-correlation and confidence weighting, as well as trend discrimination of historical sequences, can effectively distinguish between transient interference and true trend defects, reduce false alarms, and provide actionable risk prioritization for operation and maintenance (e.g., giving high-priority treatment recommendations for insulation bushings with a confidence level of 0.85, and recommending short-term re-inspection and increased observation frequency for bushings with a confidence level of 0.66). Finally, the system integrates the detection results, 3D coordinates, isothermal cloud map, confidence level, and recommended contingency plans into a structured diagnostic report, and stores this report, along with the original acquired data and model output, in a historical database to support subsequent trend tracking and model iterative optimization.

[0168] Figure 4 This is a schematic diagram of a fault diagnosis system module for live equipment based on the fusion of acoustic, thermal, and visible light data, provided in an embodiment of this application. Figure 4 The equipment defect diagnosis system 10 based on acoustic, thermal and visible light data fusion shown includes at least the following parts: information acquisition and calibration module 11, data preprocessing module 12, data fusion module 13, and defect detection module 14.

[0169] In this embodiment, the information acquisition and calibration module 11 is used to simultaneously acquire acoustic signals, infrared thermal image data, and visible light images from the acoustic array, infrared thermal imager, and visible light camera; it performs real-time spatial coordinate system alignment based on visible light image features and infrared hotspot distribution, and applies depth information obtained from the laser ranging device to perform depth calibration on the multimodal data, thereby establishing a mapping relationship between the three-dimensional coordinate system and the image pixel coordinates. Please refer to [link / reference] for details. Figure 1 , Figure 2 and Figure 3 The details and their corresponding descriptions are not elaborated here.

[0170] In this embodiment, the data preprocessing module 12 is used to suppress environmental acoustic noise in the acoustic signal using an adaptive beamforming algorithm based on feature subspace decomposition; and to perform dynamic atmospheric transmittance compensation and radiation thermometry inversion on the infrared thermal image data according to the real-time acquired measurement distance and environmental meteorological parameters to obtain the corrected absolute temperature field of the equipment surface. Please refer to [link / reference] for details. Figure 1 , Figure 2 and Figure 3 The details and their corresponding descriptions are not elaborated here.

[0171] In this embodiment, the data fusion module 13 is used to construct a physical correlation model between sound speed and temperature, apply the sound speed and temperature physical correlation model to initialize the sound speed field distribution of the medium, correct the sound wave propagation path, and thus achieve sound source relocation; calculate the acoustic inversion temperature indication value through an iterative algorithm, and assimilate the absolute temperature field of the equipment surface with the internal acoustic features to generate a three-dimensional fused temperature field reflecting the state of the internal medium of the equipment; input the preprocessed acoustic signal, the three-dimensional fused temperature field, and the visible light image into the trained multi-channel deep learning model. Please refer to [link / reference] for details. Figure 1 , Figure 2 and Figure 3 The details and their corresponding descriptions are not elaborated here.

[0172] In this embodiment, the defect detection module 14 applies the multi-channel deep learning model for defect detection. The multi-channel deep learning model uses a cross-modal attention mechanism to fuse spatial features and combines historical detection records for time-series feature analysis to filter transient interference and identify the development trend of defects, thereby outputting the defect category, three-dimensional coordinates, and risk index. Please refer to [link to details]. Figure 1 , Figure 2 and Figure 3 The details and their corresponding descriptions are not elaborated here.

[0173] Figure 5 This is an electronic device 20 provided in one embodiment of this application. For example... Figure 5 As shown, the electronic device 20 includes at least the following components: a processor 21 and a memory 22.

[0174] In this embodiment, the memory 22 is used to store executable instructions of the processor 21, which, when configured to execute instructions, implement... Figure 1 The method for diagnosing defects in live equipment based on the fusion of acoustic, thermal, and visible light data is shown in the figure.

[0175] In one embodiment of this application, the program operating in the electronic device 20 may be a program that controls a central processing unit (CPU) or similar device to achieve the functions of the above-described embodiments of the present invention (a program that enables the computer to function). The information processed by these devices is then temporarily stored in random access memory (RAM) during processing, and subsequently stored in various ROMs such as read-only memory (Flash ROM) and hard disk drives (HDDs), and read, corrected, and written by the CPU as needed.

[0176] It should be noted that a portion of the electronic device 20 described above can also be implemented using a computer. In this case, the program for implementing the control function can be recorded on a computer-readable recording medium, and the program recorded on the recording medium can be read into the computer and executed.

[0177] It should be noted that the term "computer" as used here refers to a computer built into electronic device 20, employing hardware including an operating system and peripheral devices. Furthermore, "computer-readable recording media" refers to removable media such as floppy disks, magneto-optical disks, ROMs, and CD-ROMs, as well as storage devices such as hard drives built into the computer.

[0178] Furthermore, a "computer-readable recording medium" can include: a medium that dynamically stores a program for a short period of time, such as a communication line used when transmitting a program via a network such as the Internet or a communication line such as a telephone line; or a medium that stores a program for a fixed period of time, such as volatile memory inside a computer that serves as a server or client in this case. In addition, the aforementioned program can be a program used to implement the above-mentioned functions, or it can be a program that can implement the above-mentioned functions by combining with programs already recorded in the computer.

[0179] It is understood that the method, system, and equipment for diagnosing defects in live equipment based on the fusion of acoustic, thermal, and visible light data provided in this application, through pixel-level registration and depth calibration of acoustic, thermal, and optical modalities, and combined with the correction of sound wave propagation paths based on a sound speed-temperature physical model, enables the system to recover the equivalent temperature field inside the equipment and achieve three-dimensional sound source relocation even when surface thermal images are weak or visually invisible. Simultaneously, dynamic atmospheric transmittance compensation is applied to infrared data to obtain the physical absolute temperature. A multi-channel deep learning model fuses modal features under spatial cross-correlation and cross-modal attention mechanisms, and combines historical sequence analysis to output defect categories, three-dimensional coordinates, confidence levels, and risk indices. Finally, the diagnostic results are presented intuitively in visible light video in the form of a semi-transparent isothermal cloud map and augmented reality overlay, generating a structured report for easy on-site decision-making and maintenance execution.

[0180] Furthermore, this solution significantly reduces the false alarm rate and improves the system's robustness in complex environments (such as strong winds, noise, and varying weather conditions) through adaptive beamforming based on feature subspace decomposition to suppress environmental noise, a time-series module to distinguish between transient interference and trend-based degradation, and pixel-level confidence and uncertainty propagation mechanisms. The system supports online historical data archiving and model iteration, enabling early warning of hidden defects, risk prioritization, and predictive maintenance decisions. This reduces the burden of manual inspections, shortens fault response time, extends equipment lifespan, and enhances the safety and reliability of power system operation.

[0181] Those skilled in the art should recognize that the above embodiments are only used to illustrate this application and are not intended to limit this application. Any appropriate changes and variations made to the above embodiments within the essential spirit and scope of this application fall within the scope of protection claimed in this application.

Claims

1. A method for defect diagnosis of live equipment based on the fusion of acoustic, thermal, and visible light data, characterized in that, The method includes: Acoustic signals, infrared thermal image data, and visible light images are acquired simultaneously using an acoustic array, an infrared thermal imager, and a visible light camera. Based on visible light image features and infrared hotspot distribution, spatial coordinate system is aligned in real time, and depth information obtained by laser ranging device is used to perform depth calibration on multimodal data in order to establish the mapping relationship between the three-dimensional coordinate system and image pixel coordinates. An adaptive beamforming algorithm based on feature subspace decomposition is used to suppress ambient acoustic noise in the acoustic signal; A physical correlation model between sound speed and temperature is constructed, and the sound speed field distribution of the medium is initialized by applying the physical correlation model between sound speed and temperature to correct the sound wave propagation path and thus realize the sound source relocation. The acoustic inversion temperature indication value is calculated by an iterative algorithm, and the absolute temperature field of the equipment surface is assimilated with the internal acoustic features to generate a three-dimensional fused temperature field that reflects the state of the internal medium of the equipment. The method of obtaining the absolute temperature field of the equipment surface includes: dynamic atmospheric transmittance compensation and radiation thermometry inversion of the infrared thermal image data based on the real-time measurement distance and environmental meteorological parameters to obtain the corrected absolute temperature field of the equipment surface. The preprocessed acoustic signal, the three-dimensional fused temperature field, and the visible light image are input into the trained multi-channel deep learning model. The multi-channel deep learning model applies a cross-modal attention mechanism to fuse spatial features and combines historical detection records to perform time-series feature analysis, thereby filtering transient interference and identifying the development trend of defects, and outputting defect category, three-dimensional coordinates and risk index.

2. The method for defect diagnosis of live equipment based on the fusion of acoustic, thermal, and visible light data according to claim 1, characterized in that, The method of using an adaptive beamforming algorithm based on feature subspace decomposition to suppress ambient acoustic noise in the acoustic signal includes: The acoustic signal is subjected to covariance matrix eigenvalue decomposition to identify and extract the signal subspace and noise subspace; During the spatial beamforming process, a spatial null is constructed based on the direction of the interference source determined by the noise subspace, and gain compensation is performed on the direction of the target source pointed to by the signal subspace to extract the equipment discharge signal or equipment vibration signal.

3. The method for defect diagnosis of live equipment based on the fusion of acoustic, thermal, and visible light data according to claim 2, characterized in that, The method further includes: Access the meteorological monitoring interface to obtain environmental humidity, ambient temperature, and atmospheric pressure parameters; The real-time atmospheric absorption coefficient is calculated based on the ambient humidity, ambient temperature, atmospheric pressure parameters, and measurement distance. The dynamic atmospheric transmittance under the current path is calculated using a radiative transfer model, and the gain correction of the original infrared radiation energy is performed in combination with the preset equipment surface emissivity to eliminate temperature measurement deviations caused by environmental fluctuations and changes in detection distance.

4. The method for defect diagnosis of live equipment based on the fusion of acoustic, thermal, and visible light data according to claim 3, characterized in that, The iterative process of the physical correlation model between sound speed and temperature includes: Based on the absolute temperature field, establish the spatial equivalent temperature distribution gradient and calculate the initial sound velocity value of the corresponding spatial grid node. The refraction path tracing algorithm is used to correct the propagation trajectory of sound waves in non-uniform media; By minimizing the weighted objective function of acoustic positioning residual and infrared measured temperature residual, the sound velocity distribution parameters of the medium inside the device are recursively updated, thereby obtaining the equivalent acoustic temperature inside the device.

5. The method for defect diagnosis of live equipment based on the fusion of acoustic, thermal, and visible light data according to claim 4, characterized in that, The method further includes: The three-dimensional fused temperature field is converted into a multi-layer semi-transparent isothermal cloud map, and combined with the defect category and three-dimensional coordinate label, it is dynamically superimposed and rendered in the visible light real-time video stream according to the mapping relationship using augmented reality technology. Based on the trend curve generated by the risk index and time series characteristics analysis, and in conjunction with the handling plans in the expert knowledge base, a diagnostic report is generated.

6. The method for defect diagnosis of live equipment based on the fusion of acoustic, thermal, and visible light data according to claim 1, characterized in that, The process of inputting the preprocessed acoustic signal, the three-dimensional fused temperature field, and the visible light image into the trained multi-channel deep learning model includes: The apparent texture features of the visible light image and the spatial thermal distribution features of the three-dimensional fused temperature field are extracted using residual network branches, and the Mel-frequency cepstral coefficient features of the acoustic signal are extracted using a one-dimensional convolutional network. The multi-channel deep learning model constructs a spatial cross-correlation matrix in the feature fusion layer to calculate the correlation weights of different modal feature maps at the same spatial location, thereby enhancing the ability to extract defect features.

7. The method for defect diagnosis of live equipment based on the fusion of acoustic, thermal, and visible light data according to claim 6, characterized in that, The aforementioned time-series feature analysis, combined with historical detection records, includes: The current frame fusion features are compared with the historical feature sequence within a preset time window, and the evolution trend of device status over time is extracted using a recurrent neural unit to obtain historical baseline values. By calculating the offset and rate of change of the current state characteristics relative to historical benchmark values, transient disturbances caused by ambient temperature fluctuations or sudden load changes, as well as trend defects caused by insulation degradation or mechanical loosening, can be distinguished to improve the stability of diagnostic results.

8. The method for defect diagnosis of live equipment based on the fusion of acoustic, thermal, and visible light data according to claim 1, characterized in that, The synchronous acquisition and establishment of the mapping relationship between the three-dimensional coordinate system and the image pixel coordinates include: A field-programmable gate array is used to generate a synchronous pulse signal, which drives the acoustic sensor, infrared detector and visible light image sensor to perform equal-interval sampling; By identifying a preset calibration reference in the target device area or utilizing the geometric edge features of the device itself, the perspective transformation matrix between multiple sensors is calculated to correct the parallax caused by the misalignment of the mounting axes of each sensor, thereby establishing the mapping relationship.

9. A defect diagnosis system for live equipment based on the fusion of acoustic, thermal, and visible light data, characterized in that, The system includes: The information acquisition and calibration module is used to simultaneously acquire acoustic signals, infrared thermal image data and visible light images from the acoustic array, infrared thermal imager and visible light camera; it performs real-time spatial coordinate system alignment based on visible light image features and infrared hotspot distribution, and applies depth information obtained by laser ranging device to perform depth calibration on multimodal data in order to establish the mapping relationship between the three-dimensional coordinate system and image pixel coordinates. The data preprocessing module is used to suppress environmental acoustic noise in the acoustic signal using an adaptive beamforming algorithm based on feature subspace decomposition; and to perform dynamic atmospheric transmittance compensation and radiation thermometry inversion on the infrared thermal image data based on the real-time acquired measurement distance and environmental meteorological parameters to obtain the corrected absolute temperature field of the equipment surface. The data fusion module is used to construct a physical correlation model between sound speed and temperature, and to initialize the sound speed field distribution of the medium using the sound speed and temperature physical correlation model to correct the sound wave propagation path and thus achieve sound source relocation. It calculates the acoustic inversion temperature indication value through an iterative algorithm, and assimilates the absolute temperature field of the equipment surface with the internal acoustic features to generate a three-dimensional fused temperature field reflecting the state of the medium inside the equipment. The acquisition method of the absolute temperature field of the equipment surface includes: performing dynamic atmospheric transmittance compensation and radiation thermometry inversion on the infrared thermal image data based on the real-time acquired measurement distance and environmental meteorological parameters to obtain the corrected absolute temperature field of the equipment surface; and inputting the preprocessed acoustic signal, the three-dimensional fused temperature field, and the visible light image into a trained multi-channel deep learning model. The defect detection module uses the multi-channel deep learning model to detect defects. The multi-channel deep learning model uses a cross-modal attention mechanism to fuse spatial features and combines historical detection records to perform time series feature analysis in order to filter transient interference and identify the development trend of defects, thereby outputting defect category, three-dimensional coordinates and risk index.

10. An electronic device, characterized in that, include: processor; as well as A memory having computer-readable instructions stored thereon for controlling the processor to execute the method for diagnosing defects in live equipment based on acoustic, thermal, and visible light data fusion as described in any one of claims 1 to 8.