Power equipment defect diagnosis method and system based on multi-modal data fusion

The power equipment defect diagnosis method based on multimodal data fusion simultaneously acquires multiple signals and performs three-dimensional spatial mapping and spatiotemporal analysis, which overcomes the limitations of single-modal detection, realizes a comprehensive reflection of the power equipment status and early prediction of defects, and improves the accuracy and reliability of diagnosis.

CN121980475AInactive Publication Date: 2026-05-05SICHUAN YANYUAN HUADIAN NEW ENERGY CO LTD +3
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-07
Publication Date
2026-05-05
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing methods for diagnosing defects in power equipment mainly rely on single-mode detection, which makes it difficult to fully reflect the equipment status and lacks the ability to discover early warning signs of defects and predict future development trends, resulting in high failure risks and increased maintenance costs.

Method used

A multimodal data fusion method is adopted to simultaneously acquire ultrasonic, UHF, infrared thermal imaging and visible light video signals. Physical field component sequences are constructed through three-dimensional spatial grid mapping, spatiotemporal evolution law analysis is performed, multimodal precursor features are extracted, and a pre-trained defect evolution trend prediction model is used for diagnosis.

Benefits of technology

It enables a comprehensive reflection of the multimodal physical state of power equipment, allowing for the early detection of potential defects, improving the accuracy and reliability of defect diagnosis, and reducing failure risks and maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121980475A_ABST
    Figure CN121980475A_ABST
Patent Text Reader

Abstract

The invention provides an electrical equipment defect diagnosis method and system based on multi-modal data fusion, and relates to the technical field of electrical equipment detection.The method comprises the steps that firstly, a multi-modal original signal set and a spatial position mark set synchronously collected by electrical equipment are obtained and mapped to a three-dimensional space grid, and a physical field component sequence is constructed; analyzing a spatio-temporal evolution law, and extracting a defect early-stage multi-modal precursor feature set; inputting a pre-training model to deduce future multi-modal physical field distribution; and finally, comparing the predicted physical field component with the actual physical field component, identifying a residual error abnormal region, judging a defect type, and generating a final diagnosis result. The method can comprehensively and accurately diagnose the defects of the power equipment, discover potential problems in advance and predict the development trend.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power equipment testing technology, and more specifically, to a method and system for diagnosing power equipment defects based on multimodal data fusion. Background Technology

[0002] In the operation and maintenance of power equipment, traditional methods for diagnosing power equipment defects mainly rely on single-mode detection signals, such as obtaining equipment status information by only one of ultrasonic detection, ultra-high frequency detection, infrared thermography detection, or visible light detection.

[0003] However, single-mode detection signals have significant limitations. While ultrasonic testing can effectively detect defects such as partial discharges inside equipment, its ability to detect defects such as abnormal temperatures on the equipment surface is limited. Ultra-high frequency (UHF) testing is sensitive to partial discharge signals but cannot fully reflect the overall physical state of the equipment. Infrared thermography can visually display the temperature distribution of the equipment, but it cannot effectively identify some non-thermal defects. Visible light testing can provide an image of the equipment's appearance, but it is difficult to detect potential defects inside the equipment.

[0004] Furthermore, most existing diagnostic methods only analyze the detection data at the current moment, lacking the ability to uncover early warning signs of defects and predict their future development trends. This makes it difficult to detect and take timely measures when defects first appear, often only being noticed when the defects have developed to a more serious stage, increasing the risk of equipment failure and maintenance costs. Summary of the Invention

[0005] In view of the aforementioned problems, and in conjunction with the first aspect of the present invention, the present invention provides a method for diagnosing defects in power equipment based on multimodal data fusion, the method comprising:

[0006] Acquire a set of multimodal raw signals synchronously collected from power equipment and a set of spatial location markers corresponding to the set of multimodal raw signals. The set of multimodal raw signals includes ultrasonic raw signals, ultra-high frequency raw signals, infrared thermal image raw frames, and visible light video raw frames. The multimodal original signal set is mapped to a unified three-dimensional spatial grid based on the spatial location marker set, and corresponding physical field components are constructed on each three-dimensional grid node to obtain a physical field component sequence. The physical field component sequence includes ultrasonic energy density, ultra-high frequency field strength amplitude, infrared radiation temperature and visible light reflection intensity. Each physical field component together characterizes the multimodal physical state of the three-dimensional grid node. Spatiotemporal evolution law analysis is performed on the physical field component sequence to extract the variation trend of each modal physical quantity in the continuous time series and the coupling perturbation characteristics of the interaction between modes, and generate a set of multimodal precursor features in the early stage of defects. The multimodal precursor feature set is input into a pre-trained defect evolution trend prediction model to extrapolate the multimodal physical field distribution at future times, thereby obtaining a predicted physical field component sequence. By comparing the predicted physical field component sequence with the actual real-time physical field components, the multimodal prediction residuals on each three-dimensional grid node are calculated. Based on the spatial distribution and clustering degree of the multimodal prediction residuals, abnormal residual regions are identified and defect types are determined. Finally, a diagnostic result containing defect location coordinates and defect category identifiers is generated.

[0007] Furthermore, the present invention also provides a power equipment defect diagnosis system based on multimodal data fusion, comprising: A processor; a machine-readable storage medium for storing machine-executable instructions of the processor; wherein the processor is configured to execute the above-described method for diagnosing power equipment defects based on multimodal data fusion by executing the machine-executable instructions.

[0008] In another aspect, the present invention also provides a computer program product, the computer program product including machine-executable instructions stored in a computer-readable storage medium, wherein a processor of a power equipment defect diagnosis system based on multimodal data fusion reads the machine-executable instructions from the computer-readable storage medium, and the processor executes the machine-executable instructions, causing the power equipment defect diagnosis system based on multimodal data fusion to perform the aforementioned power equipment defect diagnosis method based on multimodal data fusion.

[0009] Based on the above, by simultaneously acquiring multimodal raw signal sets from power equipment, including ultrasonic, UHF, infrared thermal imaging, and visible light video, and marking their spatial locations, the multimodal raw signals are mapped to a unified three-dimensional spatial grid to construct a physical field component sequence. This allows physical quantities of different modes to be characterized and analyzed within the same spatial framework, achieving the organic fusion of multimodal data and more comprehensively and accurately reflecting the multimodal physical state of power equipment at various locations. Spatiotemporal evolution law analysis of the physical field component sequence extracts early-stage multimodal precursor feature sets of defects, enabling early detection of potential equipment defects and providing time for timely maintenance. By using a pre-trained defect evolution trend prediction model to extrapolate the multimodal physical field distribution at future moments, comparing the predicted physical field component sequence with the actual acquired real-time physical field components, calculating the multimodal prediction residuals, and identifying abnormal residual regions, the defect type can be accurately determined, and a final diagnostic result containing defect location coordinates and category identifiers can be generated, improving the accuracy and reliability of defect diagnosis. Attached Figure Description

[0010] Figure 1 This is a schematic diagram of the execution flow of the power equipment defect diagnosis method based on multimodal data fusion provided in the embodiments of the present invention.

[0011] Figure 2 This is a schematic diagram of exemplary hardware and software components of a power equipment defect diagnosis system based on multimodal data fusion provided in an embodiment of the present invention. Detailed Implementation

[0012] Figure 1 This is a flowchart illustrating a power equipment defect diagnosis method based on multimodal data fusion, provided in one embodiment of the present invention. A detailed description follows.

[0013] Step S110: Obtain the set of multimodal raw signals synchronously collected for power equipment and the set of spatial location markers corresponding to the set of multimodal raw signals. The set of multimodal raw signals includes ultrasonic raw signals, ultra-high frequency raw signals, infrared thermal image raw frames and visible light video raw frames.

[0014] In this embodiment, a quadrupedal inspection robot or an embodied intelligent robot is used to perform live-line testing on a 220kV oil-immersed transformer in a substation. After the robot navigates to the preset testing station, the central control unit broadcasts a synchronous acquisition trigger signal to each sensor module. This trigger signal is generated based on a power frequency phase-locked loop to ensure that the sampling clocks of all sensors are phase-locked with the 50Hz power frequency signal of the power system.

[0015] The ultrasonic sensor module includes four adhesive ultrasonic probes, which are installed at preset detection points on the transformer tank wall via a quick-change tool head at the end of a robotic arm. The analog signals acquired by each probe are pre-amplified by 40dB and converted to digital signals by 12-bit analog-to-digital conversion to generate discrete raw ultrasonic signal sequences U1(t) to U4(t). The UHF sensor module includes two built-in UHF antennas, installed at the transformer body's drain valve and the high-voltage bushing riser, respectively. The received signals are bandpass filtered with a bandwidth of 300MHz to 1500MHz and down-converted to generate raw UHF intermediate frequency signals UH1(t) and UH2(t). The infrared thermal imager module uses an uncooled focal plane detector with a resolution of 640×480 pixels and a thermal sensitivity better than 0.05 degrees Celsius. It continuously acquires raw infrared thermal image frame sequences IR_frm(m,n,t).

[0016] The visible light camera module employs a global shutter CMOS sensor with a resolution of 1920×1080 pixels, synchronously acquiring raw visible light video frame sequences VIS_frm(m, n, t). The set of spatial position markers acquired synchronously with the raw signals includes: the robot chassis's 3D spatial coordinates (x_base, y_base, z_base) and attitude angles (roll_base, pitch_base, yaw_base) in the global map; the joint angle values ​​θ1 to θ6 of the six-DOF robotic arm; the precise pose transformation matrix T_tool of the sensor quick-change tool head mounting flange in the robot's base coordinate system, obtained through laser tracker calibration; and the intrinsic parameter matrices K_ir, K_vis and distortion coefficient vectors D_ir, D_vis of the infrared thermal imager and visible light camera. All raw signals and their spatial marker data are encapsulated into timestamp-aligned data frames and transmitted to a memory buffer via the robot's onboard industrial control computer's PCIe bus.

[0017] Step S120: Map the multimodal original signal set to a unified three-dimensional spatial grid according to the spatial location marker set, and construct corresponding physical field components on each three-dimensional grid node to obtain a physical field component sequence. The physical field component sequence includes ultrasonic energy density, ultra-high frequency field strength amplitude, infrared radiation temperature and visible light reflection intensity.

[0018] Step S121: Perform spatial meshing processing based on the three-dimensional structural model of the power equipment to generate a three-dimensional spatial mesh composed of multiple hexahedral mesh units inside and on the surface of the power equipment. The vertices of each hexahedral mesh unit serve as three-dimensional mesh nodes and are assigned unique spatial coordinate identifiers.

[0019] A 3D CAD model of the transformer was retrieved from the power equipment knowledge graph, containing triangular facet geometric information of components such as the tank shell, core column, high and low voltage windings, clamps, and insulation barriers. Based on this model, a structured hexahedral mesh covering the internal oil-immersed region and the external shell surface of the transformer was constructed using an open-source mesh generation tool library. Dense regions with 5mm mesh cells were set near the high-voltage bushing equalization ball, at the ends of the winding coils, and on the core grounding lead; sparse regions with 50mm mesh cells were set in the flat areas of the tank wall; and a gradient mesh was used in the oil-paper insulation transition region. The generated 3D spatial mesh contains approximately 5 million hexahedral cells, each with 8 vertices as 3D mesh nodes. Each node is assigned a unique identifier based on its spatial coordinates. The coordinates of node N(i,j,k) are (X_min+i×dx_i, Y_min+j×dy_j, Z_min+k×dz_k), where dx_i, dy_j, and dz_k are the adaptive grid step sizes in each direction. All node coordinates are unified to a local coordinate system with the center of the bottom of the transformer tank as the origin.

[0020] Step S122: Analyze the spatial coordinates and beam pointing parameters of the ultrasonic sensor in the spatial location marker set, perform energy envelope extraction processing on each raw ultrasonic signal to obtain the instantaneous ultrasonic energy value at each sampling time, propagate the instantaneous ultrasonic energy value back to each three-dimensional grid node of the three-dimensional spatial grid according to the ultrasonic propagation attenuation model in the medium, calculate the ultrasonic energy density contribution value received by each three-dimensional grid node at the corresponding time, and superimpose and fuse the contribution values ​​of all raw ultrasonic signals to generate the ultrasonic energy density tensor component of each three-dimensional grid node in the continuous time series.

[0021] The spatial coordinates (x_us1, y_us1, z_us1) to (x_us4, y_us4, z_us4) of the four ultrasonic sensors in the local coordinate system of the transformer, as well as the beam pointing unit vectors (dx_us1, dy_us1, dz_us1) to (dx_us4, dy_us4, dz_us4), were obtained analytically. Bandpass filtering was applied to each raw ultrasonic signal U_n(t) to remove mechanical vibration noise below 50kHz and high-frequency interference above 300kHz, retaining the effective signal component with a center frequency of 150kHz. A Hilbert transform was performed on the filtered signal to construct an analytic signal, and the instantaneous energy envelope E_n(t) was obtained by taking the modulus of the analytic signal. A propagation attenuation model for ultrasonic waves in a transformer's multilayer dielectric is constructed, comprehensively considering: spherical diffusion attenuation factor A_s(r) = 1 / r², where r represents the straight-line distance traveled by the ultrasonic wave from the sound source (i.e., the 3D mesh node P_ijk) to the location of the ultrasonic sensor; dielectric absorption attenuation factor A_a(r, f) = e^(-αf²r), where f represents the center frequency of the ultrasonic signal, and α represents the dielectric absorption coefficient of the ultrasonic wave propagating in the transformer's insulating oil and solid insulating materials; interface transmission attenuation factor A_t(θ) = 4Z1Z2cosθ / (Z1cosθ+Z2)², where Z1 and Z2 are the acoustic impedances of the dielectrics on both sides. For any 3D mesh node P_ijk, the spatial distance r_ijk_n between it and the sensor n is calculated. Using the instantaneous energy value E_n(t) of sensor n at time t as input, the inverse propagation equation is solved to obtain the contribution value C_us_ijk_n(t) of node P_ijk to the received value. The contribution value C_us_ijk_n(t) = E_n(t) / [A_s(r_ijk_n) × A_a(r_ijk_n, f_c) × A_t(θ_ijk_n)], where f_c is the center frequency of 150kHz and θ_ijk_n is the angle between the line connecting the node and the sensor and the beam pointing direction. The contribution values ​​of all four sensors at time t are linearly superimposed to obtain the ultrasonic energy density value U_ijk(t) of node P_ijk at time t. This process is repeated for each sampling time to generate the ultrasonic energy density tensor component U_ijk(t) of each node in the continuous time series.

[0022] Step S123: Analyze the UHF sensor spatial coordinates and antenna pattern parameters in the spatial location marker set, perform field strength amplitude demodulation processing on each UHF raw signal to obtain the UHF instantaneous field strength amplitude at each sampling time, map the UHF instantaneous field strength amplitude to each three-dimensional grid node of the three-dimensional spatial grid according to the radiation propagation model of UHF electromagnetic waves in space, calculate the field strength amplitude contribution value of each three-dimensional grid node at the corresponding time, and perform vector synthesis on the contribution values ​​of all UHF raw signals to generate the UHF field strength amplitude tensor component of each three-dimensional grid node in the continuous time series.

[0023] The spatial coordinates (x_uhf1, y_uhf1, z_uhf1) and (x_uhf2, y_uhf2, z_uhf2) of two UHF sensors and their antenna pattern parameters are obtained through analysis. The pattern is described in spherical coordinates to represent the antenna gain G(φ, θ) in the azimuth and elevation directions. The amplitude of each original UHF intermediate frequency signal UH_n(t) is demodulated, and digital down-conversion is used to shift the intermediate frequency signal to the baseband. The in-phase component I_n(t) and quadrature component Q_n(t) of the baseband signal are extracted using a low-pass filter, and the instantaneous field strength amplitude A_n(t) is calculated. A radiation propagation model of UHF electromagnetic waves in the complex structure of a transformer is constructed. Based on the uniform diffraction theory, the attenuation in free space propagation is considered to be inversely proportional to the distance. The reflection coefficient of the metal conductor surface is related to the incident angle and the conductor conductivity, and the dielectric attenuation of the insulating oil is related to the frequency and the dielectric loss tangent. For any node P_ijk, calculate its line-of-sight path with sensor n and determine whether the path intersects with the metal components of the transformer. Based on the number of reflections, diffractions, and medium crossings along the path, calculate the comprehensive propagation coefficient T_uhf_ijk_n from node P_ijk to sensor n. This complex coefficient contains amplitude attenuation and phase delay information. By solving the inverse problem, divide the amplitude A_n(t) received by sensor n at time t by the propagation coefficient amplitude |T_uhf_ijk_n| to obtain the contribution amplitude of node P_ijk to the received value. Multiply this by the polarization unit vector p_ijk_n determined by the direction of the propagation path to obtain the contribution vector H_ijk_n(t). Perform vector synthesis on the two sensor contribution vectors to obtain H_ijk(t) = H_ijk_1(t) + H_ijk_2(t). Take the magnitude of this vector as the scalar U_uhf_ijk(t) of the node's ultra-high frequency field strength amplitude, while retaining the three orthogonal components of the vector. Repeat this process at each sampling time point to generate the UHF field strength amplitude tensor component UH_ijk(t) for each node.

[0024] Step S124: Analyze the infrared thermal imager spatial coordinates and imaging projection matrix in the spatial location marker set, perform non-uniformity correction processing on each original infrared thermal image frame to obtain the corrected infrared temperature distribution matrix, establish the mapping relationship between the infrared image pixel coordinates and the coordinates of the three-dimensional grid nodes in the three-dimensional space according to the imaging projection matrix, extract the infrared radiation temperature value corresponding to each three-dimensional grid node from the infrared temperature distribution matrix through bilinear interpolation, and generate the infrared radiation temperature tensor component of each three-dimensional grid node in the continuous time series.

[0025] The spatial coordinates (x_ir, y_ir, z_ir) of the infrared thermal imager and its imaging projection matrix P_ir are obtained through analysis. This 3×4 matrix is ​​composed of the intrinsic parameter matrix K_ir, the extrinsic parameter rotation matrix R_ir, and the translation vector t_ir, i.e., P_ir = K_ir × [R_ir|t_ir]. Non-uniformity correction is performed on each raw infrared thermal image frame IR_frm(m, n, t). Two pre-stored correction coefficients, the gain coefficient matrix G(m, n) and the offset coefficient matrix O(m, n), are called in Flash. The corrected infrared temperature distribution matrix T_ir(m, n, t) = G(m, n) × IR_frm(m, n, t) + O(m, n) is calculated. For each 3D mesh node P_ijk, its local coordinates (x_i, y_j, z_k) are converted to homogeneous coordinates (x_i, y_j, z_k, 1). Multiplying these coordinates by the projection matrix P_ir yields the image homogeneous coordinates (u_h, v_h, w_h). After normalization, the floating-point pixel coordinates u_ijk = u_h / w_h and v_ijk = v_h / w_h are obtained. It is then determined whether (u_ijk, v_ijk) are within the valid image regions [0, 639] and [0, 479]. If the temperature value is within the effective area, bilinear interpolation is used to extract the temperature value: u_ijk is rounded down to u0 and up to u1, v_ijk is rounded down to v0 and up to v1, and the horizontal weight α_u = u_ijk - u0 and the vertical weight α_v = v_ijk - v0 are calculated; the temperature values ​​T00, T01, T10, and T11 of four integer pixels are read from the T_ir matrix; first, the horizontal interpolation is T0 = T00 × (1 - α_u) + T10 × α_u and T1 = T01 × (1 - α_u) + T11 × α_u; then the vertical interpolation is T_ijk(t) = T0 × (1 - α_v) + T1 × α_v. The above operations are performed on each frame of infrared image and each node to generate the infrared radiation temperature tensor component IR_ijk(t) of each node in the continuous time series.

[0026] Step S125: Analyze the visible light camera spatial coordinates and imaging projection matrix in the spatial location marker set, perform illumination uniformity compensation processing on each raw visible light video frame to obtain the compensated visible light reflection intensity distribution matrix, establish the mapping relationship between visible light image pixel coordinates and three-dimensional grid node coordinates based on the imaging projection matrix, extract the visible light reflection intensity value corresponding to each three-dimensional grid node from the visible light reflection intensity distribution matrix using bilinear interpolation, and generate the visible light reflection intensity tensor component of each three-dimensional grid node in the continuous time series.

[0027] The visible light camera's spatial coordinates (x_vis, y_vis, z_vis) and its imaging projection matrix P_vis are obtained through analysis. This matrix consists of intrinsic parameters K_vis and extrinsic parameters R_vis and t_vis. Illumination uniformity compensation is performed on each raw visible light video frame VIS_frm(m, n, t), converting the RGB three-channel image to a grayscale image GRAY_frm(m, n, t). A pre-calibrated vignetting correction coefficient matrix V_corr(m, n) is used to describe the reciprocal of the illumination attenuation at the lens edge, and the compensated reflection intensity matrix R_vis(m, n, t) = GRAY_frm(m, n, t) × V_corr(m, n). If overexposed or underexposed areas are detected, local contrast enhancement is performed using adaptive histogram equalization based on the reference region. For each 3D mesh node P_ijk, its local coordinates (x_i, y_j, z_k) are mapped to the image coordinate system through the projection matrix P_vis, obtaining floating-point pixel coordinates (u_ijk', v_ijk'). Using the same bilinear interpolation method as in step S124, the visible light reflectance intensity value Vis_ijk(t) corresponding to the node is extracted from the compensated visible light reflectance intensity distribution matrix R_vis(m,n,t). The above operation is performed on each frame of visible light image and each node to generate the visible light reflectance intensity tensor component VIS_ijk(t) of each node in the continuous time series.

[0028] Step S126: Organize the ultrasonic energy density tensor component, ultra-high frequency field strength amplitude tensor component, infrared radiation temperature tensor component, and visible light reflection intensity tensor component of each three-dimensional mesh node at the same moment as independent data channels aligned in spatial location to form multi-channel data characterizing the node state at that moment.

[0029] For any sampling time t, for each three-dimensional grid node P_ijk, its four corresponding components U_ijk(t), UH_ijk(t), IR_ijk(t), and VIS_ijk(t) are combined into a four-dimensional feature vector F_ijk(t). These four components strictly correspond to the same physical location point in space and the same sampling time in time, together constituting the multimodal physical state representation of the node at time t.

[0030] Step S127: Reorganize the multi-channel data of all three-dimensional mesh nodes at the same time according to the spatial arrangement order of the three-dimensional mesh nodes to generate a three-dimensional spatial data volume with multiple independent channels, where each channel corresponds to a physical field distribution of a mode.

[0031] The multi-channel data of all three-dimensional mesh nodes at time t are reorganized into a four-dimensional data structure according to their spatial indices (i, j, k), i.e., a three-dimensional spatial data volume with shape (I_max, J_max, K_max, 4). The first three dimensions correspond to the three directions of the spatial mesh, and the fourth dimension, with a size of 4, corresponds to the distribution of the four modal physical fields: ultrasonic energy density, ultra-high frequency field strength amplitude, infrared radiation temperature, and visible light reflection intensity.

[0032] Step S128: Stack the physical field components at multiple consecutive moments in chronological order to generate the physical field component sequence containing spatiotemporal dimension information.

[0033] For each sampling time t1 to tT, steps S126 and S127 are repeated to generate a series of three-dimensional spatial data volumes. These data volumes are stacked strictly in chronological order along the time dimension to form a five-dimensional physical field component sequence tensor with the shape (T, I_max, J_max, K_max, 4). This sequence contains complete spatiotemporal evolution information of various physical fields inside and on the surface of the transformer from the past to the present.

[0034] Step S130: Perform spatiotemporal evolution law analysis on the physical field component sequence, extract the variation trend of each modal physical quantity in the continuous time series and the coupling perturbation characteristics of the interaction between modes, and generate a set of multimodal precursor features for the early stage of defects.

[0035] Step S131: Perform trend decomposition processing on the ultrasonic energy density time series of each three-dimensional grid node in the physical field component sequence. Use the local weighted regression scatter smoothing method to decompose the ultrasonic energy density time series into long-term trend component, periodic fluctuation component and random disturbance component. Extract the slope change rate of the long-term trend component as the ultrasonic energy accumulation rate feature.

[0036] For the ultrasonic energy density time series U_ijk(t) of each node P_ijk, a local weighted regression scatter smoothing method is applied for trend decomposition. This method selects a neighborhood window near each time point and performs weighted linear regression on the data points within the window. The weight function is w(d) = (1 - (|d| / h)³)³, where d is the distance and h is the window width. The long-term trend component T_u_ijk(t) is obtained through iterative robust fitting. The original sequence is subtracted from the long-term trend component to obtain the remaining sequence. Periodogram analysis is performed on the remaining sequence to identify the dominant periodic component, and the periodic fluctuation component P_u_ijk(t) is obtained through Fourier series fitting. Subtraction again yields the random disturbance component R_u_ijk(t). The long-term trend component T_u_ijk(t) is numerically differentiated to calculate its first derivative sequence. Then, the rate of change of this first derivative within the sliding window, i.e., the second derivative, is calculated and used as the ultrasonic energy accumulation rate feature f_us_rate_ijk.

[0037] Step S132: Perform pulse analysis processing on the time series of ultra-high frequency field strength amplitude of each three-dimensional grid node in the physical field component sequence, detect transient pulse events in the time series of ultra-high frequency field strength amplitude, count the occurrence frequency and amplitude distribution of transient pulse events in each time window, and generate ultra-high frequency pulse activity intensity characteristics and ultra-high frequency pulse phase distribution characteristics.

[0038] Pulse detection is performed on the time series UH_ijk(t) of the ultra-high frequency field strength amplitude for each node P_ijk. An adaptive threshold Th_uhf = μ_uhf + k × σ_uhf is set, where μ_uhf is the sequence mean, σ_uhf is the standard deviation, and k is a coefficient ranging from 3 to 5. Local maxima points with amplitudes exceeding Th_uhf and durations in the nanosecond to microsecond range are marked as transient pulse events. The occurrence time t_p and peak amplitude A_p of each pulse event are recorded. The time axis is divided into equal-length windows, such as 20 milliseconds per power frequency cycle, and the number of pulse events occurring within each window is counted to obtain the pulse frequency characteristic f_uhf_count_ijk. Histogram statistics are performed on all pulse amplitudes to obtain the amplitude distribution characteristic f_uhf_amp_hist_ijk. By utilizing the synchronization information of the power frequency phase-locked loop, the time of each pulse occurrence is converted into the corresponding power frequency phase angle φ_p, and the phase distribution of the pulse within the power frequency period is statistically analyzed to obtain the pulse phase distribution characteristics f_uhf_phase_hist_ijk.

[0039] Step S133: Perform thermodynamic analysis on the infrared radiation temperature time series of each three-dimensional grid node in the physical field component sequence, calculate the first-order difference sequence and the second-order difference sequence of the infrared radiation temperature time series, extract the peak feature of temperature rise rate from the first-order difference sequence, and extract the feature of the starting point of temperature acceleration change from the second-order difference sequence.

[0040] Thermodynamic analysis is performed on the infrared radiation temperature time series IR_ijk(t) for each node P_ijk. The first-order difference sequence ΔIR_ijk(t) = IR_ijk(t+Δt) - IR_ijk(t) is calculated, representing the instantaneous rate of temperature change. Local maxima of ΔIR_ijk(t) are detected, and the maximum value is taken as the peak feature of the temperature rise rate f_ir_rate_max_ijk. The second-order difference sequence ΔΔIR_ijk(t) = ΔIR_ijk(t+Δt) - ΔIR_ijk(t) is calculated, representing the acceleration of temperature change. Zero-crossing points where ΔΔIR_ijk(t) changes from negative to positive and continues to rise are detected, and the time corresponding to this point is taken as the starting point feature of the accelerated temperature change f_ir_acc_start_ijk, recording the index value of that time.

[0041] Step S134: Perform texture change analysis on the time series of visible light reflection intensity of each three-dimensional mesh node in the physical field component sequence, calculate the structural similarity index of the visible light reflection intensity tensor components at adjacent times, and extract the texture abrupt change time and abrupt change amplitude features based on the time change curve of the structural similarity index.

[0042] Texture variation analysis is performed on the visible light reflectance intensity time series VIS_ijk(t) of each node P_ijk. Note that a single node only provides intensity values; the node's neighborhood must be considered for texture analysis. A three-dimensional neighborhood block with a window size of N×N×N, centered on node P_ijk, is extracted. The VIS values ​​of all nodes within this neighborhood block constitute the local texture block B_ijk(t) at that time. For adjacent times t and t+Δt, the structural similarity index SSIM_ijk(t) of the two texture blocks B_ijk(t) and B_ijk(t+Δt) is calculated. SSIM calculation is based on the comparison of three components: brightness, contrast, and structure. Specifically, SSIM(x,y) = [l(x,y)^α] × [c(x,y)^β] × [s(x,y)^γ], where l is the brightness comparison function based on the mean, c is the contrast comparison function based on the standard deviation, and s is the structure comparison function based on the covariance. The SSIM_ijk(t) sequence is calculated for each time t. The point in the sequence that is below a set threshold and has the largest decrease is detected. The corresponding time is taken as the texture mutation time feature f_vis_change_t_ijk, and the difference between the SSIM values ​​before and after the mutation is taken as the mutation magnitude feature f_vis_change_mag_ijk.

[0043] Step S135: On each three-dimensional mesh node, retain the ultrasonic energy accumulation rate feature extracted from the ultrasonic signal, the ultra-high frequency pulse activity intensity feature and ultra-high frequency pulse phase distribution feature extracted from the ultra-high frequency signal, the temperature rise rate peak feature and temperature acceleration change start point feature extracted from the infrared signal, and the texture change time and change amplitude feature extracted from the visible light signal. These features together constitute the multimodal trend feature set of the three-dimensional mesh node.

[0044] For each node P_ijk, the extracted modal features are summarized to form the multimodal trend feature set F_trend_ijk for that node. This feature set includes: ultrasonic energy accumulation rate f_us_rate_ijk; UHF pulse frequency f_uhf_count_ijk, amplitude distribution histogram f_uhf_amp_hist_ijk, phase distribution histogram f_uhf_phase_hist_ijk; infrared temperature rise rate peak f_ir_rate_max_ijk, temperature acceleration start point f_ir_acc_start_ijk; visible light texture abrupt change time f_vis_change_t_ijk, abrupt change amplitude f_vis_change_mag_ijk. These features constitute a high-dimensional feature vector, where the histogram features exist in the form of multiple bin count values.

[0045] Step S136: Perform spatial correlation analysis on the trend characteristics of adjacent three-dimensional mesh nodes under the same mode. For the peak characteristics of ultrasonic energy accumulation rate, ultra-high frequency pulse activity intensity, and infrared temperature rise rate, calculate the spatial covariance matrix of each three-dimensional mesh node and its neighboring three-dimensional mesh nodes on each feature. Extract the spatial coupling strength characteristics reflecting the coordinated changes in the local area based on the eigenvalue decomposition results of the spatial covariance matrix.

[0046] For each node P_ijk, its three-dimensional neighborhood is defined, for example, a 26-neighborhood includes all adjacent grid nodes. For the ultrasonic energy accumulation rate feature, the values ​​of node P_ijk and all its neighboring nodes on this feature are collected to form a vector V_us_local. The covariance matrix C_us_local between this vector and the eigenvalues ​​of each node in the neighborhood is calculated, with the matrix dimension being the number of neighboring nodes × the number of neighboring nodes. Eigenvalue decomposition is performed on the covariance matrix to obtain eigenvalues ​​λ1≥λ2≥...≥λn. The ratio of the first eigenvalue λ1 to the sum of all eigenvalues ​​is taken as the ultrasonic modal spatial coupling strength feature f_us_couple_ijk, which reflects the proportion of principal components that coordinate changes in the local region on this feature. Similarly, the above neighborhood covariance calculation and eigenvalue decomposition are performed on the ultra-high frequency pulse activity intensity feature f_uhf_count and the infrared temperature rise rate peak feature f_ir_rate_max, respectively, to obtain the corresponding spatial coupling strength features f_uhf_couple_ijk and f_ir_couple_ijk.

[0047] Step S137: Calculate the time-varying cross-correlation function between ultrasonic energy density and UHF field strength amplitude at each three-dimensional grid node, and extract the maximum correlation coefficient and its corresponding delay time from the time-varying cross-correlation function as the ultrasonic-UHF coupling delay feature.

[0048] For each node P_ijk, its ultrasonic energy density time series U_ijk(τ) and ultra-high frequency field strength amplitude time series UH_ijk(τ) within the time window [tw, t] are taken, where τ is the time within the window. The time-varying cross-correlation function R_uh_ijk(τ, d) of the two series is calculated, where d is the time delay variable. The cross-correlation function is defined as R(d) = E[(U(τ) - μ_U) × (UH(τ+d) - μ_UH)] / (σ_U × σ_UH), where E is the expected value, i.e. the average value, μ_U and μ_UH are the mean values ​​of the series within the window, and σ_U and σ_UH are the standard deviations. For the center time t of each window, the cross-correlation function curve with delay d as the independent variable is obtained. The maximum value R_max_ijk(t) and its corresponding delay time d_max_ijk(t) are extracted from this curve. R_max_ijk and d_max_ijk are taken as the ultrasonic-UHF coupling delay features f_us_uhf_corr_ijk and f_us_uhf_delay_ijk of the node at time t.

[0049] Step S138: Calculate the local gradient mutual information value between infrared radiation temperature and visible light reflection intensity at each three-dimensional mesh node, and extract the thermal-optical coupling anomaly index feature based on the time evolution curve of the local gradient mutual information value.

[0050] For each node P_ijk, consider its three-dimensional neighborhood. For the infrared channel, calculate the gradients of the temperature values ​​of each node within the neighborhood in three spatial directions, obtaining gradient fields G_ir_x, G_ir_y, and G_ir_z. Similarly, calculate the gradient fields G_vis_x, G_vis_y, and G_vis_z for the visible light reflectance intensity. Discretize the two gradient fields into a finite number of bins, constructing a joint histogram p(g_ir, g_vis) and edge histograms p(g_ir) and p(g_vis). Calculate the gradient mutual information value MI_gradient = ΣΣp(g_ir, g_vis) × log[p(g_ir, g_vis) / (p(g_ir) × p(g_vis))]. For each time t, calculate MI_gradient_ijk(t) and obtain its time evolution curve. The curve's trend is detected, and its mean and standard deviation within the sliding window are calculated. If the current MI value deviates from the historical mean by more than three times the standard deviation, then this deviation is used as the thermal-optical coupling anomaly index feature f_ir_vis_mi_anom_ijk.

[0051] Step S139: The calculated spatial coupling strength characteristics reflecting spatial cooperative changes, ultrasonic-ultra-high frequency coupling delay characteristics reflecting cross-modal interactions, and thermal-optical coupling anomaly index characteristics, together with the multimodal trend feature set of the corresponding three-dimensional mesh node, are used as the enhanced multimodal evolution feature set of that node.

[0052] For each node P_ijk, the multimodal trend feature set F_trend_ijk obtained in step S135 is concatenated and combined with the spatial coupling strength features f_us_couple_ijk, f_uhf_couple_ijk, and f_ir_couple_ijk obtained in step S136, the ultrasonic-UHF coupling delay features f_us_uhf_corr_ijk and f_us_uhf_delay_ijk obtained in step S137, and the thermal-optical coupling anomaly index feature f_ir_vis_mi_anom_ijk obtained in step S138 to form the enhanced multimodal evolution feature set F_enhanced_ijk for that node. This feature set is a high-dimensional vector that contains the node's own trend information, neighborhood spatial cooperation information, and cross-modal interaction information.

[0053] Step S1310: Perform cluster analysis on the enhanced multimodal evolution feature set of all three-dimensional mesh nodes to identify three-dimensional mesh node clusters exhibiting abnormal evolution patterns. Take the area covered by the three-dimensional mesh node clusters as the defect precursor area, and take the enhanced multimodal evolution feature set of all three-dimensional mesh nodes in the defect precursor area as the early multimodal precursor feature set of the defect.

[0054] The enhanced multimodal evolutionary feature vectors F_enhanced_ijk of all 3D mesh nodes are combined into a high-dimensional feature space point set, where each point corresponds to a 3D mesh node, and the point coordinates are composed of all feature values ​​of that node. Density-based spatial clustering is performed on this point set, with a neighborhood radius parameter Eps and a minimum number of neighborhood points parameter MinPts. Core points are identified by traversing each point and calculating the number of neighboring points within its Eps radius. Mutually reachable core points and their neighborhood boundary points are connected to form clusters, generating multiple initial feature clusters. Spatial continuity is verified for each initial feature cluster by extracting the spatial coordinates of all nodes contained in the cluster, constructing a spatial adjacency graph, and checking for isolated nodes or fragmented regions. If such isolated nodes or fragmented regions exist, they are removed from the current cluster, and the spatial region corresponding to the removed cluster is recalculated to obtain spatially continuous candidate feature clusters. For each candidate feature cluster, the mean vector of the enhanced multimodal evolutionary feature vectors of all nodes within the cluster is used as the cluster center. The Euclidean distance between each node's feature vector and the cluster center is calculated, and the sum of squared Euclidean distances is used as the cluster compactness index. The mean point of the spatial coordinates of all nodes within the cluster is used as the cluster space center. The spatial distance between each node and the cluster space center is calculated, and the sum of squared spatial distances is used as the spatial cohesion index. A weighted sum of the cluster compactness index and the spatial cohesion index is used to calculate the comprehensive anomaly score for each candidate feature cluster. The weighting coefficients are pre-set as α and β based on the importance of feature similarity and spatial clustering in historical defect samples, with α+β=1. Candidate feature clusters whose comprehensive anomaly scores exceed a preset dynamic threshold are identified as anomalous evolutionary pattern clusters. This dynamic threshold is adaptively determined based on the statistical distribution of the comprehensive anomaly scores of all current candidate feature clusters, such as taking the mean plus twice the standard deviation. For each anomalous evolution mode cluster, the temporal evolution trajectory of the feature vectors of its nodes is analyzed. The variation curves of each modality's physical quantity over consecutive time periods are extracted from the physical field component sequence. The first and second derivatives of each curve in the time dimension are calculated to identify the inflection point or the starting moment of accelerated change. The earliest inflection point or acceleration is taken as the defect precursor start time for that anomalous evolution mode cluster. Based on the cluster space center coordinates and defect precursor start time of each anomalous evolution mode cluster, corresponding spatial regions and time windows are marked on a three-dimensional spatial grid. The marked spatial regions are designated as defect precursor regions. The enhanced multimodal evolution feature sets of all three-dimensional grid nodes within this defect precursor region are organized into a spatiotemporally labeled data structure, serving as the early multimodal precursor feature set F_early for the defect.

[0055] Step S1311: Perform a similarity search between the multimodal precursor feature set and the historical defect case database of power equipment, calculate the Mahalanobis distance between the current precursor feature set and each historical case precursor feature set, select the K historical cases with the smallest Mahalanobis distance and their corresponding final defect types as early reference defect type candidates for the current defect, and add the reference defect type candidates to the multimodal precursor feature set.

[0056] A historical defect case library is loaded from the power equipment knowledge graph. Each case contains a multimodal precursor feature set F_history_i collected before the defect occurred and the finally confirmed defect type Type_i. The currently obtained multimodal precursor feature set F_early is treated as a high-dimensional vector, and the Mahalanobis distance D_maha_i between it and each historical case F_history_i is calculated as follows: D_maha_i = √[(F_early - F_history_i)^T × Σ^-1 × (F_early - F_history_i)], where Σ is the covariance matrix of all precursor feature vectors in the historical case library. Mahalanobis distance considers the correlation between features and can more accurately measure similarity. The top 3 historical cases with the smallest Mahalanobis distance are selected, and their corresponding final defect types Type_1, Type_2, and Type_3 are selected as candidate reference defect types. The information of the above candidate types is appended to F_early in the form of labels to form a multimodal precursor feature set with reference labels, which is used as the prior knowledge input for subsequent prediction models.

[0057] Step S140: Input the multimodal precursor feature set into the pre-trained defect evolution trend prediction model to deduce the multimodal physical field distribution at future times and obtain the predicted physical field component sequence.

[0058] Step S141: Construct the defect evolution trend prediction model. The defect evolution trend prediction model is based on an encoder-predictor-decoder architecture, wherein the encoder is composed of a spatiotemporal convolutional long short-term memory network, the predictor is composed of multiple parallel modality prediction branches, and the decoder is composed of a three-dimensional deconvolutional network.

[0059] A defect evolution trend prediction model is constructed, adopting an encoder-predictor-decoder architecture. The encoder uses a spatiotemporal convolutional long short-term memory (ConvLSTM) network, which replaces fully connected operations with convolutional operations on top of the traditional LSTM, enabling simultaneous capture of the temporal and spatial dependencies of the spatiotemporal data. The encoder consists of three stacked ConvLSTM layers, each with a 3×3×3 kernel size, a stride of 1, and padding of 1 to maintain spatial dimensions. The number of hidden state channels is 64, 128, and 256, respectively. The predictor consists of four parallel modal prediction branches, handling the future evolution prediction of four modalities: ultrasound, ultra-high frequency, infrared, and visible light. The decoder uses a three-dimensional deconvolutional network, containing three 3D deconvolutional layers, progressively restoring the spatial resolution to the original input size. Each deconvolutional kernel is 4×4×4 with a stride of 2, and the number of output channels is 128, 64, and 4 respectively. The final output channel count of 4 corresponds to the physical field components of the four modalities.

[0060] Step S142: Reorganize the multimodal precursor feature set according to the spatial order of the three-dimensional mesh nodes in three-dimensional space to generate a precursor feature tensor with spatial structure, wherein each channel of the precursor feature tensor corresponds to a precursor feature type.

[0061] The multimodal precursor feature set F_early obtained in step S141 is reorganized. F_early was originally a set of high-dimensional feature vectors based on nodes, and its spatial structure needs to be restored. For each node P_ijk in the precursor region, its enhanced multimodal evolution features already contain multiple features. The features of all nodes are arranged according to the spatial index (i, j, k), and each feature type is treated as an independent channel to generate a five-dimensional precursor feature tensor X_input with the shape (T_input, I_zone, J_zone, K_zone, C), where T_input is the length of the historical time window, I_zone, J_zone, and K_zone are the grid sizes of the spatial sub-regions where the precursor region is located, and C is the feature channel number equal to the dimension of the enhanced multimodal evolution features.

[0062] Step S143: Input the precursor feature tensor into the encoder, and model the dependency relationship of the precursor features in the time and space dimensions through the gating mechanism of the spatiotemporal convolutional long short-term memory network to extract the hidden state tensor containing historical evolution information.

[0063] The precursor feature tensor X_input is input into the ConvLSTM layer of the encoder step by step. The core calculation process of ConvLSTM is: Input gate i_t = σ(W_xi) X_t+W_hi H_{t-1}+W_ci C_{t-1}+b_i), forget gate f_t=σ(W_xf X_t+W_hf H_{t-1}+W_cf C_{t-1}+b_f), output gate o_t=σ(W_xo) X_t+W_ho H_{t-1}+W_co C_t+b_o), cell state update C_t=f_t C_{t-1}+i_t tanh(W_xc) X_t+W_hc H_{t-1}+b_c), hidden state H_t=o_t tanh(C_t). Where This represents the convolution operation. Let σ represent the Hadamard product, σ be the sigmoid activation function, and W and b be learnable parameters. After processing through 3 layers of ConvLSTM, the hidden state H_last at the last time step is obtained as the hidden state tensor output by the encoder. Its shape is (I_zone, J_zone, K_zone, 256), which contains a compressed representation of the input historical spatiotemporal sequence.

[0064] Step S144: Input the hidden state tensor into each modal prediction branch of the predictor. Each modal prediction branch uses a recurrent neural network with a different structure to capture the unique evolution law of the corresponding mode. The ultrasonic modal prediction branch uses a bidirectional long short-term memory network to process the nonlinear growth trend of ultrasonic energy accumulation. The ultra-high frequency modal prediction branch uses a gated recurrent unit network to process the intermittent burst mode of ultra-high frequency pulse activity. The infrared modal prediction branch uses a convolutional long short-term memory network to process the spatial propagation process of thermal field diffusion. The visible light modal prediction branch uses a self-attention network to process the abrupt change characteristics of texture changes.

[0065] The hidden state tensor H_last is input into four parallel modal prediction branches. The ultrasound modal branch uses a bidirectional long short-term memory (Bi-LSTM) network, expanding H_last into a sequence along the spatial dimension. This sequence is input into the Bi-LSTM to capture forward and backward dependencies, outputting the ultrasound future evolution feature F_us_pred. The ultra-high frequency (UHF) modal branch uses a gated recurrent unit (GRU) network. GRU simplifies the LSTM structure by merging the forget gate and input gate into an update gate, making it more suitable for capturing the temporal dependencies of intermittent burst modes, outputting the UHF future evolution feature F_uhf_pred. The infrared modal branch uses a convolutional long short-term memory (ConvLSTM) network. While it also uses ConvLSTM, its structure differs from the encoder, focusing on iteratively predicting the future thermal field diffusion process from the current hidden state, outputting the infrared future evolution feature F_ir_pred. The visible light modality branch uses a self-attention network, treating H_last as a set of feature vectors. By calculating the self-attention matrix, it captures the long-distance dependencies between features, making it more suitable for handling non-local changes such as texture mutations. It outputs the visible light future evolution feature F_vis_pred.

[0066] Step S145: The intermediate prediction features output by each modality prediction branch are concatenated and fused along the feature dimension to obtain a comprehensive prediction feature tensor that integrates multimodal evolution information.

[0067] The intermediate predicted features F_us_pred, F_uhf_pred, F_ir_pred, and F_vis_pred output from the four modal branches are concatenated along the feature channel dimension to obtain the comprehensive predicted feature tensor F_fusion, which integrates multimodal evolution information. Its shape is (I_zone, J_zone, K_zone, C_us+C_uhf+C_ir+C_vis), where C is the number of feature channels output by each branch.

[0068] Step S146: Input the comprehensive prediction feature tensor into the decoder, and use the three-dimensional deconvolution network to restore the spatial resolution and reconstruct the details of the comprehensive prediction features to generate the initial predicted physical field components for the first future time.

[0069] The F_fusion input is used to construct a 3D deconvolutional network for the decoder. The first 3D deconvolutional layer upsamples the feature map size by a factor of 2 and halves the number of output channels. The second layer upsamples by a factor of 2 again. The third layer reduces the number of output channels to 4 and restores the spatial dimensions to the original input dimensions I_zone, J_zone, and K_zone. The final output tensor Y_pred_t1 has the shape (I_zone, J_zone, K_zone, 4), corresponding to the initial predicted physical field components at the first future time t1. The four channels represent the predicted values ​​of ultrasonic energy density, ultra-high frequency field strength amplitude, infrared radiation temperature, and visible light reflection intensity, respectively.

[0070] Step S147: Combine the initial predicted physical field components of the first future time with the actual physical field components of the historical time to form a new input sequence. Repeat the forward calculation process of the encoder, predictor and decoder to iteratively generate the predicted physical field components of multiple subsequent future times.

[0071] The generated Y_pred_t1 is used as the latest time-stamp data and combined with the original input history sequence X_input to form a new input sequence X_input_new = [X_input(:,:,:,:,2:end), Y_pred_t1]. X_input_new is then input into the encoder-predictor-decoder network again, and the same forward computation is performed to generate the predicted physics component Y_pred_t2 for the second future time t2. This iterative process is repeated until all predicted physics components Y_pred_t1 to Y_pred_tT for the preset future time T_future are generated.

[0072] Step S148: Perform modal consistency adjustment on the predicted physical field components generated at each future time to ensure that the physical quantities of different modes at the same time satisfy the preset physical constraint relationship in spatial distribution, and generate the adjusted predicted physical field components.

[0073] Modal consistency adjustment is performed on the predicted tensor Y_pred_t at each future time t. Preset physical constraints include: the ultrasonic energy density and the ultra-high frequency field strength amplitude satisfy local correlation constraints in space, i.e., hypersonic regions are usually accompanied by enhanced ultra-high frequency activity; the infrared radiation temperature and visible light reflectance intensity satisfy structural consistency constraints in spatial gradients. Adjustment is performed using the projection gradient method: Y_pred_t is projected onto a manifold that satisfies the constraints, i.e., the minimum adjustment amount is solved to ensure that the adjusted tensor Y_adj_t satisfies the constraints and has the minimum Euclidean distance to Y_pred_t. In the specific iterative calculation, projections to the correlation constraint set and the gradient consistency constraint set are performed alternately until convergence. The adjusted predicted physical field components Y_adj_t are obtained.

[0074] Step S149: Arrange all the adjusted predicted physical field components for future times in chronological order to form the predicted physical field component sequence.

[0075] The adjusted predicted physical field components Y_adj_t1 to Y_adj_tT at each time point are stacked in chronological order to form the predicted physical field component sequence Y_pred_seq, with the shape (T_future, I_zone, J_zone, K_zone, 4).

[0076] Step S150: Compare the predicted physical field component sequence with the actual acquired real-time physical field components, calculate the multimodal prediction residual on each three-dimensional grid node, identify abnormal residual regions and determine the defect type based on the spatial distribution clustering degree of the multimodal prediction residual, and generate a final diagnostic result containing defect location coordinates and defect category identifier.

[0077] Step S151: At each future moment, acquire the real-time physical field components that are actually collected for that future moment.

[0078] At each future time t, the robot returns to the same detection station and repeats steps S110 to S128 to collect and process the actual physical field component Y_real_t at that time, with the same shape as the predicted tensor, i.e. (I_zone, J_zone, K_zone, 4).

[0079] Step S152: For each three-dimensional mesh node, calculate the absolute difference between the ultrasonic energy density value of the three-dimensional mesh node in the real-time physical field component and the predicted ultrasonic energy density value of the three-dimensional mesh node at the corresponding time in the predicted physical field component sequence, and use it as the ultrasonic modal residual of the three-dimensional mesh node.

[0080] For each time t and each node P_ijk, calculate the ultrasonic modal residual R_us_ijk_t=|Y_real_t(i,j,k,1)-Y_pred_t(i,j,k,1)|, where index 1 represents the ultrasonic channel.

[0081] Step S153: For each three-dimensional grid node, calculate the absolute difference between the UHF field strength amplitude of the three-dimensional grid node in the real-time physical field component and the predicted value of the UHF field strength amplitude of the three-dimensional grid node at the corresponding time in the predicted physical field component sequence, and use it as the UHF modal residual of the three-dimensional grid node.

[0082] Calculate the UHF modal residual R_uhf_ijk_t=|Y_real_t(i,j,k,2)-Y_pred_t(i,j,k,2)|, where index 2 represents the UHF channel.

[0083] Step S154: For each three-dimensional mesh node, calculate the absolute difference between the infrared radiation temperature value of the three-dimensional mesh node in the real-time physical field component and the predicted infrared radiation temperature value of the three-dimensional mesh node at the corresponding time in the predicted physical field component sequence, and use it as the infrared modal residual of the three-dimensional mesh node.

[0084] Calculate the infrared modal residual R_ir_ijk_t=|Y_real_t(i,j,k,3)-Y_pred_t(i,j,k,3)|, where index 3 represents the infrared channel.

[0085] Step S155: For each three-dimensional mesh node, calculate the absolute difference between the visible light reflection intensity value of the three-dimensional mesh node in the real-time physical field component and the predicted value of the visible light reflection intensity of the three-dimensional mesh node at the corresponding time in the predicted physical field component sequence, and use it as the visible light mode residual of the three-dimensional mesh node.

[0086] Calculate the visible light mode residual R_vis_ijk_t=|Y_real_t(i,j,k,4)-Y_pred_t(i,j,k,4)|, where index 4 represents the visible light channel.

[0087] Step S156: For each mode, use the mean and standard deviation of the residual values ​​of all three-dimensional mesh nodes under historical normal operating conditions to calculate the mode's residuals. Perform Z-score normalization on the residuals of the mode to obtain normalized mode residuals. For each three-dimensional mesh node, retain its normalized ultrasonic mode residuals, ultra-high frequency mode residuals, infrared mode residuals, and visible light mode residuals to form the multi-mode residual spectrum of the three-dimensional mesh node.

[0088] For ultrasonic modes, the mean μ_us and standard deviation σ_us of all node residuals are statistically analyzed from historical normal operating condition data. The normalized ultrasonic residual Z_us_ijk_t = (R_us_ijk_t - μ_us) / σ_us is calculated. Similarly, Z_uhf_ijk_t, Z_ir_ijk_t, and Z_vis_ijk_t are obtained. For each node, the four normalized residuals are combined to form the multimodal residual spectral vector of that node, Z_ijk_t = [Z_us, Z_uhf, Z_ir, Z_vis].

[0089] Step S157: For each mode, perform spatial distribution analysis on the normalized modal residuals of all three-dimensional mesh nodes to generate a normalized residual distribution map for each mode.

[0090] For the ultrasonic mode, all the normalized residual Z_us_ijk_t are arranged according to the spatial indices (i, j, k) to obtain the three-dimensional residual distribution map Map_us_t. Similarly, Map_uhf_t, Map_ir_t, and Map_vis_t are obtained.

[0091] Step S158: Perform local maximum detection on the normalized residual distribution map. For any three-dimensional grid node, if its residual value is greater than the residual values of all its adjacent three-dimensional grid nodes, and the difference between its residual value and the average value of the residual values of all adjacent three-dimensional grid nodes exceeds a preset neighborhood residual threshold, then this three-dimensional grid node is identified as a candidate abnormal seed point.

[0092] For Map_us_t, traverse each node and compare Z_us_ijk_t with the residual values of all nodes within its 26-neighborhood. If Z_us_ijk_t is greater than the residual values of all neighborhood nodes, and Z_us_ijk_t minus the neighborhood mean μ_us_neighbor is greater than the preset threshold Th_us_seed, then this node is marked as a candidate abnormal seed point for the ultrasonic mode. Similarly, the same detection is performed for other modes, and the union of the seed points detected in all modes is used as the set of candidate abnormal seed points Seeds_t.

[0093] Step S159: Starting from each candidate abnormal seed point, respectively combine the residual distributions of each mode, and include the three-dimensional grid nodes that are spatially connected and have similar residual values to the seed point in at least one mode into the same residual anomaly region, obtaining multiple residual anomaly regions.

[0094] For each seed point P_seed, use the region growing algorithm. Define the growth criterion: for a candidate neighborhood node P_neighbor, if there exists at least one mode m such that |Z_m_neighbor - Z_m_seed| < Th_similar, and P_neighbor is spatially connected to the current region, then this node is included in the region. Th_similar is the residual similarity threshold. Iteratively grow from the seed point until no new nodes are added. The connected region obtained by growing from each seed point is used as a residual anomaly region Region_r.

[0095] Step S1510: For each residual anomaly region, calculate the average value of the normalized modal residuals of all three-dimensional grid nodes within this residual anomaly region in each mode to obtain the multi-modal anomaly intensity spectrum of this residual anomaly region, and calculate the average value of the spatial coordinates of all three-dimensional grid nodes within this residual anomaly region as the region center coordinate.

[0096] For the residual anomaly region Region_r, the mean values ​​of the normalized residuals of all nodes in the region across the four modes are calculated as μZ_us_r, μZ_uhf_r, μZ_ir_r, and μZ_vis_r, forming the multimodal anomaly intensity spectrum A_r = [μZ_us, μZ_uhf, μZ_ir, μZ_vis]. The mean values ​​of the spatial coordinates of all nodes in the region are calculated to obtain the coordinates of the region center C_r = (avg(x_i), avg(y_j), avg(z_k)).

[0097] Step S1511: Extract the ratio distribution histogram between the ultrasonic modal residual and the ultra-high frequency modal residual of the three-dimensional mesh nodes in each residual abnormal region. Match the ratio distribution histogram with the preset defect type template library according to the shape characteristics of the histogram, and take the defect type with the highest matching degree as the defect category identifier of the residual abnormal region.

[0098] Step S1511-1: For each residual abnormal region, collect the ultrasonic modal residual values ​​of all three-dimensional mesh nodes in the residual abnormal region to form an ultrasonic residual set, and collect the ultra-high frequency modal residual values ​​of all three-dimensional mesh nodes in the residual abnormal region to form an ultra-high frequency residual set.

[0099] For Region_r, collect the original residuals (unnormalized) of all nodes, R_us_ijk and R_uhf_ijk, to form sets S_us_r and S_uhf_r.

[0100] Step S1511-2: Pair the ultrasonic residual set and the ultra-high frequency residual set element by element so that each three-dimensional mesh node corresponds to a pair of ultrasonic residual values ​​and ultra-high frequency residual values.

[0101] Since the elements of the set come from the same node, they are naturally paired, with each node corresponding to a pair (R_us_ijk, R_uhf_ijk).

[0102] Step S1511-3: Calculate the ratio of ultrasonic residual value to ultra-high frequency residual value at each three-dimensional mesh node to obtain the residual ratio of the three-dimensional mesh node, forming a set of residual ratios for the residual abnormal region.

[0103] For each node, calculate the ratio Q_ijk = R_us_ijk / (R_uhf_ijk + ε), where ε is a very small positive number to prevent division by zero. This yields the set of residual ratios for that region, S_q_r = {Q_ijk}.

[0104] Step S1511-4: Perform statistical distribution analysis on the set of residual ratios, divide the range of residual ratio values ​​into multiple continuous equal-width intervals, count the number of residual ratio values ​​falling into each interval, and generate a residual ratio distribution histogram with the interval as the horizontal axis and the frequency as the vertical axis.

[0105] Divide the interval from minimum to maximum value in S_q_r into N_bin equal subintervals, for example, N_bin is 20. Count the number of Q values ​​in each subinterval, count_n, to obtain the histogram H_r_raw.

[0106] Step S1511-5: Normalize the residual ratio distribution histogram by dividing the frequency of each interval by the total number of nodes to obtain the normalized residual ratio probability density histogram.

[0107] Calculate the total number of nodes N_node_r in the region, and normalize the histogram H_r_prob=count_n / N_node_r so that the sum of probabilities of all intervals is 1.

[0108] Step S1511-6: Load multiple reference ratio distribution templates corresponding to multiple defect types from the preset defect type template library. Each reference ratio distribution template is a normalized probability density histogram pre-calculated on the sample area of ​​a known defect type using the same method.

[0109] A pre-built defect type template library contains reference histograms H_template_type corresponding to various defect types such as electrical tree discharge, floating potential discharge, oil gap breakdown, and mechanical vibration.

[0110] Step S1511-7: Calculate the Barcol distance between the residual ratio probability density histogram and each reference ratio distribution template to obtain multiple Barcol distance values. The smaller the Barcol distance, the more similar the two distributions are.

[0111] The Bach distance D_B(H_r_prob, H_template) = -ln(Σ√(H_r_prob(n) × H_template(n))). Calculate the Bach distance with each template.

[0112] Step S1511-8: Select the defect type corresponding to the reference ratio distribution template with the smallest Bartlett distance as the candidate defect type.

[0113] Take the template with the smallest D_B and the corresponding type Type_candidate.

[0114] Step S1511-9: Further calculate the correlation coefficient between the residual ratio probability density histogram and the candidate defect type reference ratio distribution template. If the correlation coefficient is greater than the preset similarity threshold, then confirm that the candidate defect type is the defect category identifier of the residual abnormal region.

[0115] Calculate the Pearson correlation coefficient ρ = cov(H_r_prob, H_template_candidate) / (σ_r × σ_template). If ρ > Th_corr (e.g., 0.8), then the defect category is confirmed as Type_candidate.

[0116] Step S1511-10: If the correlation coefficient is not greater than the similarity threshold, the reference ratio distribution template with the second smallest Bach distance is selected and the correlation coefficient calculation is repeated until a defect type that meets the correlation coefficient threshold requirement is found. If the correlation coefficients of all reference ratio distribution templates do not meet the requirements, the defect category of the residual abnormal area is marked as an unknown type.

[0117] If ρ ≤ Th_corr, then select the template with the second smallest D_B and repeat the calculation of ρ until a template that meets the threshold is found or all templates are traversed. If neither is met, mark the defect type as Unknown.

[0118] Step S1512: Take the center coordinates of the residual abnormal region in the multimodal abnormal intensity spectrum where the regional abnormal intensity exceeds the preset threshold as the defect location coordinates, and associate the defect category identifier of the residual abnormal region with the defect location coordinates to generate the final diagnostic result containing the defect location coordinates and defect category identifier.

[0119] For each residual anomaly region Region_r, calculate the square root of its comprehensive intensity S_r = (μZ_us_r^2 + μZ_uhf_r^2 + μZ_ir_r^2 + μZ_vis_r^2) of its multimodal anomaly intensity spectrum A_r. If S_r is greater than the preset comprehensive intensity threshold Th_intensity, then the region is considered a valid defect. The center coordinates C_r of the valid defect region are used as the defect location coordinates, and the defect category identifier Type_r determined in step S1511 is associated with these coordinates to generate the final diagnostic result Diag_r = {C_r, Type_r}. The diagnostic results for all valid defect regions are summarized to form a final diagnostic result list.

[0120] Step S1513: Extract the ratio distribution histogram between the ultrasonic modal residual and the ultra-high frequency modal residual of the three-dimensional mesh nodes in each residual abnormal region. Match the ratio distribution histogram with the preset defect type template library according to the shape characteristics of the histogram, and take the defect type with the highest matching degree as the defect category identifier of the residual abnormal region.

[0121] For each residual abnormal region, Region_r, collect the ultrasonic modal residual values ​​R_us_ijk and UHF modal residual values ​​R_uhf_ijk of all 3D mesh nodes in the region at the original scale. For each node, calculate the residual ratio Q_ijk = R_us_ijk divided by the sum of R_uhf_ijk and a minimum positive number ε to prevent the denominator from being zero. Construct the residual ratio set S_q_r for the region using the Q_ijk values ​​of all nodes. Divide the numerical range of S_q_r into twenty equal-width intervals, count the number of nodes falling into each interval, and obtain the original frequency histogram. Divide the frequency of each interval of this histogram by the total number of nodes in the region to obtain the normalized residual ratio probability density histogram H_r_prob. Load reference ratio distribution templates corresponding to various defect types from a preset defect type template library. Each template is a normalized probability density histogram H_template_type pre-calculated on a known defect type sample region using the same method. Calculate the Barcol distance D_B = -ln(∑√(H_r_prob(n) × H_template(n))) between H_r_prob and each H_template_type, obtaining a series of Barcol distance values. The template with the smallest Barcol distance corresponds to the defect type as the candidate defect type Type_candidate. Further calculate the Pearson correlation coefficient ρ between H_r_prob and the Type_candidate template. If ρ is greater than a preset threshold of 0.8, the candidate defect type is confirmed as the defect category identifier Type_r for that region. If ρ is not greater than 0.8, select the template with the second smallest Barcol distance and repeat the calculation of ρ until a template that meets the threshold requirement is found or all templates are traversed. If none of these conditions are met, the defect category identifier is marked as an unknown type.

[0122] Step S1514: Take the center coordinates of the residual abnormal region in the multimodal abnormal intensity spectrum where the regional abnormal intensity exceeds the preset threshold as the defect location coordinates, and associate the defect category identifier of the residual abnormal region with the defect location coordinates to generate the final diagnostic result containing the defect location coordinates and defect category identifier.

[0123] For each residual abnormal region Region_r, calculate the comprehensive intensity S_r of its multimodal abnormal intensity spectrum A_r. S_r is equal to the square root of the sum of the squares of the normalized residual means μZ_us_r, μZ_uhf_r, μZ_ir_r, and μZ_vis_r of the four modes: ultrasound, UHF, infrared, and visible light. If S_r is greater than the preset comprehensive intensity threshold Th_intensity, the region is considered a valid defect. The center coordinates C_r of the valid defect region are used as the defect location coordinates. The defect category identifier Type_r determined in step S1513 is associated with these coordinates to generate the final diagnostic result Diag_r={C_r, Type_r}. The diagnostic results of all valid defect regions are summarized to form a final diagnostic result list and output to the robot monitoring platform.

[0124] Furthermore, the method may also include: step S160: extracting the ultrasonic original signal segment and ultra-high frequency original signal segment corresponding to the spatial location from the multimodal original signal set according to the defect location coordinates in the final diagnostic result, performing deconvolution processing on the ultrasonic original signal segment to remove the multiple reflection and aliasing effects of the ultrasonic signal on the propagation path, recovering the waveform characteristics of the original acoustic emission signal at the defect source, and obtaining the ultrasonic waveform of the defect source.

[0125] Based on the defect location coordinates C_r in the final diagnostic result, the ultrasonic raw signal segment corresponding to the spatial location is extracted retrospectively from the multimodal raw signal set acquired and stored in step S110. During extraction, with C_r as the center, considering the spatial location of the ultrasonic sensor and beam pointing, the raw signal segment U_local(t) of the sensor channel receiving the strongest signal from that region is selected. This signal segment is then deconvolved to eliminate propagation path effects. First, based on the ultrasonic propagation attenuation model used in step S122, a system transfer function H_sys(ω) is constructed from the defect source location C_r to the corresponding sensor. This function includes the combined effects of medium absorption, interface transmission, and geometric diffusion. A Fourier transform is performed on U_local(t) to obtain the frequency domain representation U_local(ω). In the frequency domain, Wiener filtering and deconvolution are performed: U_source(ω) = U_local(ω)H_sys (ω) / (|H_sys(ω)|²+λ), where λ is the estimated noise-to-signal power ratio. Performing an inverse Fourier transform on U_source(ω) yields the time-domain waveform u_source(t), which is the recovered ultrasonic waveform of the defect source.

[0126] Step S161: Perform time-frequency analysis on the ultrasonic waveform of the defect source, extract the energy distribution of the ultrasonic waveform of the defect source at different frequency components, and construct an ultrasonic fingerprint feature vector characterizing the microscopic motion mode of the defect based on the number of dominant frequency peaks and the proportional relationship between the dominant frequencies in the energy distribution.

[0127] A short-time Fourier transform is performed on the recovered ultrasonic waveform u_source(t) of the defect source, decomposing the waveform into an energy distribution spectrum S_stft(f, τ) on the time-frequency plane. The average power spectral density P(f) = ∫S_stft(f, τ)dτ is obtained by integrating the spectrum along the frequency axis. Local maxima in P(f) are detected, and the three peak frequencies f_peak1, f_peak2, and f_peak3 with the largest amplitudes and their corresponding amplitudes A_peak1, A_peak2, and A_peak3 are extracted. The proportional relationships between the peak frequencies are calculated, such as A_peak2 divided by A_peak1 and A_peak3 divided by A_peak1. Simultaneously, the energy concentration index of the spectrum on the time axis, i.e., the instantaneous frequency variance, is calculated. The number of peak frequencies, the peak frequency values, the amplitude ratios, and the energy concentration index are combined to form an ultrasonic fingerprint feature vector F_us_fingerprint.

[0128] Step S162: Perform phase analysis processing on the original UHF signal segment, synchronously segment the original UHF signal segment according to the power frequency cycle, extract the amplitude of the UHF signal at equal phase intervals in each power frequency cycle, generate the phase-resolved amplitude matrix of the UHF signal in the power frequency cycle, perform singular value decomposition on the phase-resolved amplitude matrix, extract the first singular vector reflecting the stability of the discharge phase and the second singular vector reflecting the fluctuation of the discharge amplitude, and combine the first singular vector and the second singular vector to form an UHF fingerprint feature vector characterizing the defect discharge mode.

[0129] Extract the original UHF signal segment UH_local(t) corresponding to the defect location coordinates C_r. Using the synchronization signal provided by the power frequency phase-locked loop, divide UH_local(t) according to the power frequency period, with each period being 20 milliseconds. For the signal within each period, divide it into 360 equal phase intervals (one interval per degree). Extract the maximum or effective value of the signal amplitude within each phase interval to obtain the phase-resolved amplitude vector V_phase_k of that period, with a length of 360. Stack all N periods's V_phase_k row-wise to form a phase-resolved amplitude matrix M_prpd, with a size of N×360. Perform singular value decomposition on M_prpd: M_prpd=U×Σ×V^T, where U is the left singular matrix, Σ is the diagonal singular value matrix, and V is the right singular matrix. Take the first right singular vector V1 (the first column of V), which reflects the distribution stability of discharge probability at different phases; take the second right singular vector V2 (the second column of V), which reflects the fluctuation pattern of discharge amplitude. The first and last parts of V1 and V2 are concatenated to form a UHF fingerprint feature vector F_uhf_fingerprint with a length of 720.

[0130] Step S163: Extract the infrared thermal image raw frame sequence corresponding to the spatial location from the multimodal raw signal set according to the defect location coordinates, perform dynamic modeling on the temperature change curve of the defect area in the infrared thermal image raw frame sequence, calculate the zero-crossing time of the second derivative of the temperature change with time, take the zero-crossing time as the inflection point feature of the heat source energy release rate, and combine the contrast between the temperature field of the defect area and the temperature field of the surrounding normal area to construct an infrared fingerprint feature vector characterizing the thermal properties of the defect.

[0131] Based on the defect location coordinates C_r, the average temperature of a small region centered on the corresponding pixel in the infrared image projected onto C_r is extracted from the original infrared thermal image frame sequence, resulting in the temperature change curve T_region(t) for that region. After smoothing the curve to remove noise, its second derivative T''(t) is calculated. The zero-crossing points of T''(t), i.e., the points where T''(t) changes from positive to negative or from negative to positive, are detected, and the moment corresponding to the first significant zero-crossing point is selected as the inflection point feature t_inflection of the heat source energy release rate. Simultaneously, the average temperature of the defect region centered on C_r, T_defect, and the average temperature of the surrounding normal region, T_normal, are extracted, and the thermal contrast C_thermal = (T_defect - T_normal) divided by T_normal is calculated. The inflection point moment t_inflection and the thermal contrast C_thermal are combined to form the infrared fingerprint feature vector F_ir_fingerprint.

[0132] Step S164: Extract the visible light video original frame sequence corresponding to the spatial location from the multimodal original signal set according to the defect location coordinates, perform fractal dimension calculation on the texture image of the region where the defect is located in the visible light video original frame sequence, obtain the fractal dimension change curve of the texture image at different scales, extract the scale parameter and dimension parameter corresponding to the curve slope change point from the fractal dimension change curve, and construct a visible light fingerprint feature vector characterizing the evolution of the defect surface morphology.

[0133] Based on the defect location coordinates C_r, a sequence of sub-image blocks I_block(t) centered on the corresponding pixel in the visible light image projected onto the original visible light video frame sequence is extracted. For each sub-image block at each time step, its fractal dimension is calculated using the differential box counting method. Specifically, the image block size is normalized, covered with grids of different scales r, and the number of boxes N(r) required to cover the image is counted. The slope of the straight line log(N(r)) and log(1 / r) is fitted in a double logarithmic coordinate system, and this slope is the fractal dimension D_f. The fractal dimension changing with time curve D_f(t) and the fractal dimension changing with scale r curve D_f(r) are obtained. The first derivative of the D_f(r) curve is calculated, and abrupt change points are detected. The scale parameter r_break and dimension parameter D_break corresponding to the abrupt change points are extracted. r_break and D_break are combined into a visible light fingerprint feature vector F_vis_fingerprint.

[0134] Step S165: Input the ultrasonic fingerprint feature vector, the ultra-high frequency fingerprint feature vector, the infrared fingerprint feature vector, and the visible light fingerprint feature vector into a pre-constructed multimodal fingerprint fusion network. Use the autoencoder structure of the multimodal fingerprint fusion network to reduce the dimensionality and reconstruct the fingerprint features of each modality. Calculate the low-dimensional latent space vector with the smallest reconstruction error as the unified micro-fingerprint encoding of the defect.

[0135] A multimodal fingerprint fusion network is constructed using a stacked autoencoder architecture. First, the fingerprint feature vectors F_us_fingerprint, F_uhf_fingerprint, F_ir_fingerprint, and F_vis_fingerprint from the four modalities are concatenated end-to-end to form a fused feature vector F_fuse_raw. The autoencoder consists of an encoder and a decoder. The encoder comprises three fully connected layers with 512, 256, and 128 neurons respectively, using ReLU activation to progressively reduce the dimensionality of the fused features to a 128-dimensional latent space vector Z_defect. The decoder structure is symmetrical to the encoder, consisting of three fully connected layers with 256 and 512 neurons respectively. The output layer dimension is the same as F_fuse_raw, and the output layer uses linear activation while the remaining layers use ReLU activation. Z_defect is reconstructed back into the original feature space to obtain F_fuse_recon. During the training phase, a large number of historical defect samples are used to train the network parameters with the objective of minimizing the reconstruction error L_recon=||F_fuse_raw-F_fuse_recon||^2. During the inference phase, the fused features of the current defect are input into the trained autoencoder to obtain the latent space vector Z_defect, which is the unified micro-fingerprint encoding of the defect.

[0136] Step S166: Perform similarity matching between the unified micro-fingerprint code of the defect and the mechanism fingerprint codes in the pre-stored defect physical mechanism library. The defect physical mechanism library contains standard micro-fingerprint codes corresponding to different defect types at different development stages. Determine the micro-evolution mechanism type of the current defect based on the defect physical mechanism corresponding to the mechanism fingerprint code with the highest matching degree.

[0137] A pre-built defect physics mechanism library includes various micro-evolution mechanisms such as the gradual growth mechanism of electrical tree channels, the hotspot formation mechanism induced by partial discharge, and the fatigue cracking mechanism caused by mechanical vibration. Each mechanism corresponds to a standard micro-fingerprint encoding vector Z_mechanism_i, which is obtained by calculating the cluster centers of the micro-fingerprint encodings of known mechanism samples. The cosine similarity sim_i between the current defect encoding Z_defect and each mechanism encoding Z_mechanism_i is calculated as (Z_defect·Z_mechanism_i) divided by (||Z_defect||×||Z_mechanism_i||). The micro-evolution mechanism type corresponding to the mechanism encoding with the highest similarity is taken as the current defect's micro-evolution mechanism type M_type.

[0138] Step S167: Retrieve the corresponding defect development dynamic equation from the defect physical mechanism library according to the micro-evolution mechanism type. The defect development dynamic equation contains a set of differential equations describing the changes of key defect characteristics over time.

[0139] Based on the determined microscopic evolution mechanism type M_type, the corresponding defect development kinetic equation is retrieved from the defect physics mechanism library. If M_type is the gradual growth mechanism of the electrical tree channel, the electrical tree growth kinetic equation is retrieved, which includes the differential equation dL / dt=k1×(V-V_th)^n×exp(-Ea / (kT)) describing the change of the electrical tree length L over time, and the differential equation describing the change of the discharge quantity Q. If M_type is the hot spot formation mechanism induced by partial discharge, the thermodynamic equation is retrieved, which includes the differential equation dT_hot / dt=(P_discharge-P_loss)divided by (ρ×c×V) describing the change of the hot spot temperature T_hot over time, and the equation describing the expansion of the thermal damage region. If M_type is the fatigue cracking mechanism caused by mechanical vibration, the fatigue crack propagation equation Paris formula da / dN=C×(ΔK)^m is retrieved, where a is the crack length, N is the number of stress cycles, and ΔK is the stress intensity factor amplitude.

[0140] Step S168: Using some components of the unified micro-fingerprint encoding as the initial state parameters of the defect development dynamics equation, numerically integrate the differential equation system in the time domain to deduce the size expansion curve, discharge intensity evolution curve and temperature rise curve of the defect in the future time series.

[0141] Components corresponding to the initial state of the kinetic equations are extracted from the unified micro-fingerprint encoding Z_defect. For example, for the electrical tree growth mechanism, the components in Z_defect corresponding to the current electrical tree length L0 and the current discharge quantity Q0 are used as initial conditions. The fourth-order Runge-Kutta numerical integration method is used to solve the differential equation system in the time domain, with a time step set to one hour, to predict the defect size expansion curve L(t), discharge intensity evolution curve Q(t), and temperature rise curve T(t) over the next month.

[0142] Step S169: Compare the deduced defect size expansion curve, discharge intensity evolution curve, and temperature rise curve with the trend of corresponding physical quantities in the predicted physical field component sequence obtained by the defect evolution trend prediction model. Calculate the dynamic time warping distance between the multiple curves. If the dynamic time warping distance is less than a preset matching threshold, output the deduced curve as a supplementary prediction result of the defect development trajectory, and use it together with the final diagnosis result to guide subsequent operation and maintenance decisions.

[0143] From the predicted physical field component sequence Y_pred_seq obtained in step S140, extract the ultrasonic energy density prediction sequence U_pred(t), the ultra-high frequency field strength amplitude prediction sequence UH_pred(t), and the infrared radiation temperature prediction sequence IR_pred(t) corresponding to the defect location C_r. Compare the extrapolated curves with the predicted sequences: compare the trend of the size expansion curve L(t) with U_pred(t), the trend of the discharge intensity curve Q(t) with UH_pred(t), and the trend of the temperature rise curve T(t) with IR_pred(t). Calculate the dynamic time warping distances DTW_LU, DTW_QUH, and DTW_TIR, and the comprehensive distance D_dtw = (DTW_LU + DTW_QUH + DTW_TIR) / 3. If D_dtw is less than the preset matching threshold Th_dtw, the results of the two models are considered to be consistent. The inference curve is output as a supplementary prediction result. The final diagnosis result and the supplementary prediction curve are sent to the operation and maintenance decision system together for the purpose of formulating maintenance plans and risk assessments.

[0144] Furthermore, the method may also include: step S170: based on the defect location coordinates in the final diagnostic result, extracting a subset of multimodal original signals of the defect location region and its surrounding neighborhood at multiple historical moments from the multimodal original signal set, and constructing a historical signal spatiotemporal cube of the defect region.

[0145] Based on the defect location coordinates C_r in the final diagnostic results, a three-dimensional cubic spatial region with C_r as the center and a side length of L is determined. From the complete historical multimodal raw signal set acquired and stored in step S110, the raw signal subsets corresponding to multiple historical times t1 to tM of this spatial region are extracted retrospectively. For each time moment, the ultrasonic raw signal segments, UHF raw signal segments, infrared thermal image raw frame sub-blocks, and visible light video raw frame sub-blocks corresponding to all grid nodes in this spatial region are extracted. The above data from all times are organized into a five-dimensional spatiotemporal cubic data C_hist according to spatial and temporal indices, with a shape of (M, Lx, Ly, Lz, 4), where M is the number of historical times, Lx, Ly, and Lz are the grid sizes of the spatial region, and 4 represents the four modes.

[0146] Step S171: Perform a three-dimensional Fourier transform on the ultrasonic signal in the historical signal spatiotemporal cube to obtain the distribution spectrum of the ultrasonic signal in the spatial wavenumber domain and frequency domain. Extract the amplitude variation curve of the wavenumber component corresponding to the defect depth direction from the distribution spectrum as the ultrasonic characterization curve of defect depth development.

[0147] A three-dimensional Fourier transform is performed on the ultrasonic channel data C_us_hist(t, x, y, z) in the historical signal spatiotemporal cube C_hist to obtain its distribution spectrum F_us(k_x, k_y, k_z, f) in the spatial wavenumber and frequency domains. The wavenumber axis corresponding to the defect depth direction, i.e., the direction from the equipment surface to the interior, is determined, for example, the z-axis. In the frequency domain, a frequency slice corresponding to the dominant frequency of the ultrasonic signal is selected, and the wavenumber spectrum distribution along the k_z direction under this frequency slice is extracted. For each historical time t, the wavenumber spectrum in the k_z direction at the corresponding time is extracted from F_us, and the energy integral within a specific wavenumber segment related to the defect depth is calculated to obtain the depth characterization value D_us(t) at that time. The D_us(t) values ​​at each time are connected to form a curve to obtain the ultrasonic characterization curve of the defect depth development.

[0148] Step S172: Perform a three-dimensional short-time Fourier transform on the UHF signal in the historical signal spatiotemporal cube to obtain the time-frequency distribution map of the UHF signal in the spatial and temporal dimensions. Identify the frequency band peak trajectory that moves continuously with the development of the defect from the time-frequency distribution map, and extract the slope of the frequency band peak trajectory as the migration rate feature of the defect discharge channel.

[0149] A three-dimensional short-time Fourier transform (SFT) is performed on the UHF channel data C_uhf_hist(t, x, y, z) in the historical signal spatiotemporal cube. This involves performing a SFT on the time series at each spatial point and simultaneously performing a Fourier transform in the spatial dimension, resulting in the time-frequency distribution spectrum S_uhf(f, τ, k_x, k_y, k_z). ​​In the joint frequency-spatial wavenumber domain, the peak frequency of the frequency band that continuously shifts with time τ is searched. For each time window τ, the center frequency f_peak(τ) of the frequency band with the largest amplitude in S_uhf and its corresponding dominant spatial wavenumber direction are detected. The trajectory of f_peak(τ) over time is considered a representation of the defect discharge channel migration. A linear fit is performed on this trajectory to obtain the slope k_freq_rate, which serves as the feature of the defect discharge channel migration rate.

[0150] Step S173: Perform three-dimensional heat flow inversion calculation on the infrared thermal image sequence in the historical signal spatiotemporal cube, solve the inverse problem of heat conduction, reconstruct the three-dimensional spatial distribution of the heat source intensity inside the defect and its change over time, and generate a dynamic evolution image sequence of the heat source inside the defect.

[0151] A three-dimensional heat flow inversion calculation is performed on the infrared channel data C_ir_hist(t, x, y, z) in the spatiotemporal cube of historical signals. This calculation is based on the heat conduction equation. T / t=α ²T + Q / (ρc), where T is temperature, α is the thermal diffusivity, Q is the internal heat source intensity, ρ is density, and c is specific heat capacity. Given the surface temperature distribution T_surface(t, x, y), i.e., the infrared measurements of the nodes on the equipment surface, as well as the initial temperature distribution and boundary conditions, the internal heat source intensity Q(x, y, z, t) is reconstructed by solving the inverse problem. The conjugate gradient method is used for iterative optimization to minimize the error between the surface temperature obtained from the forward modeling and the actual measured value. After iterative convergence, the three-dimensional heat source intensity distribution sequence Q_recon(t, x, y, z) is obtained. Playing this sequence in chronological order yields a dynamic evolution image sequence of the heat sources inside the defect.

[0152] Step S174: Perform three-dimensional structured light reconstruction on the visible light video sequence in the historical signal spatiotemporal cube, calculate the three-dimensional point cloud data of the defect surface morphology at different times, and extract the surface morphology evolution curves of the opening area, depth and volume of the defect over time from the three-dimensional point cloud data.

[0153] For the visible light channel data C_vis_hist(t, x, y) in the historical signal spatiotemporal cube, combined with binocular vision or structured light sensor data at the same time, a 3D reconstruction algorithm is used to generate point cloud data P_cloud(t) of the defect region surface morphology at each time t. The defect region is segmented from the point cloud, and the defect opening area A_open(t), i.e., the projected area of ​​the region enclosed by the defect edge on the equipment surface, the maximum defect depth D_max(t), i.e., the distance between the lowest point in the point cloud and the reference surface, and the defect volume V_defect(t), i.e., the spatial volume enclosed by the defect region point cloud and the reference surface, are calculated. Connecting the area, depth, and volume values ​​at each time point yields the surface morphology evolution curve.

[0154] Step S175: Input the ultrasonic characterization curve, the discharge channel migration rate characteristics, the dynamic evolution image sequence, and the surface morphology evolution curve into the defect development process inversion model. Through the multi-physics coupling calculation of the defect development process inversion model, output a three-dimensional reproduction animation of the complete development process of the defect from the initial stage to the current moment.

[0155] A defect development process inversion model is constructed based on a multi-physics coupled finite element solver. The depth development curve D_us(t) obtained in step S171 is used as the acoustic boundary condition, the migration rate feature k_freq_rate obtained in step S172 is used as the electromagnetic field evolution driving parameter, the heat source dynamic evolution image sequence Q_recon(t) obtained in step S173 is used as the thermal load, and the surface morphology evolution curve obtained in step S174 is used as the geometric boundary constraint. Substituting these inputs into the finite element model, the complete physical field evolution process from the initial defect stage t0 to the current time t_current is solved in the time domain. During the solution process, the acoustic field, electromagnetic field, temperature field, and structural mechanical field are coupled within the model, iteratively updating the material properties and geometry of the defect region. Finally, a three-dimensional animation of the complete defect development process from the initial stage to the current time is output and stored in VTK format.

[0156] Step S176: Automatically label the key stages of defect development in the three-dimensional reconstruction animation. The key stages include the moment of the first sudden increase in ultrasonic energy, the moment of the first change in the ultra-high frequency discharge mode, the moment of the first formation of the infrared heat source, and the moment of the first change in the visible light surface morphology.

[0157] Key stages were detected in the 3D reconstructed animation. Traversing the entire timeline, the moment when the global energy first exceeded three standard deviations of the background noise in the ultrasonic energy density field was marked as the first surge in ultrasonic energy (t_us). The moment when the phase-resolved spectrum first showed a significant change (i.e., the correlation coefficient with the historical average spectrum was below 0.6) in the UHF field was marked as the first change in the UHF discharge mode (t_uhf). The moment when the local heat source intensity first exceeded twice that of the surrounding area and persisted in the heat source intensity field was marked as the first formation of the infrared heat source (t_ir). The moment when the defect opening area first exceeded 1 square millimeter in the surface topography point cloud sequence was marked as the first change in the visible light surface topography (t_vis). These four moments were highlighted on the 3D reconstructed animation playback progress bar and labeled with text.

[0158] Step S177: Integrate the 3D animation of the defect development process after annotation with the digital twin model of the power equipment. In the digital twin model, the defect development process is dynamically played according to the actual timeline, and the actual acquired multimodal raw signal waveforms at the corresponding time are displayed synchronously during the playback.

[0159] The annotated 3D animation is imported into the digital twin platform of the power equipment and spatially registered and fused with the precise geometric model of the transformer. A time slider is set in the digital twin model to play the defect development process from t0 to t_current according to the actual timeline. During the playback, when the time slider reaches a certain moment, the actual ultrasonic wave waveform, ultra-high frequency wave waveform, infrared thermogram, and visible light image collected at that moment are retrieved from the historical database and displayed in a floating window on the twin model interface. This achieves spatiotemporal synchronous playback of virtual evolution and real historical data.

[0160] Step S178: Calculate the average development rate of the defect based on the time intervals of different stages in the three-dimensional reconstruction animation of the defect development process, and use the average development rate to correct the defect remaining lifetime prediction value obtained by the defect evolution trend prediction model, thereby generating a corrected remaining lifetime value that combines historical development trajectory and future trend prediction.

[0161] Based on the key stage times marked in step S176, calculate the time intervals between adjacent stages: Δt1 = t_uhf - t_us, Δt2 = t_ir - t_uhf, Δt3 = t_vis - t_ir. Calculate the weighted average of the development rates of each stage as the average defect development rate v_avg, with the weights related to the magnitude of change in defect characteristic quantities at each stage. From the predicted physical field component sequence output by the defect evolution trend prediction model, extract the time required for the defect characteristic quantity to reach a preset threshold from the current time as the remaining lifetime prediction value RUL_pred. Correct RUL_pred using the average development rate v_avg, with the correction formula being RUL_corrected = α × RUL_pred + (1 - α) × (D_remain / v_avg), where D_remain is the distance from the current defect characteristic quantity to the failure threshold, and α is the confidence weight coefficient determined based on the historical accuracy of the prediction model.

[0162] Step S179: Combine the corrected remaining lifetime value, the 3D reconstruction animation of the defect development process, and the annotation information of the key stages into a defect backtracking analysis report, and store the defect backtracking analysis report in the historical case knowledge base of the power equipment as a reference template for subsequent diagnosis of similar defects.

[0163] The corrected remaining lifetime value (RUL_corrected), the path to the 3D reconstruction animation file of the defect development process, key stage annotation information (t_us, t_uhf, t_ir, t_vis), and the defect category identifier (Type_r) are combined to form a structured defect retrospective analysis report. This report is stored in JSON format and includes fields such as defect ID, device ID, discovery time, location coordinates, defect type, development rate, key stage timestamps, remaining lifetime, and animation file index. This report is uploaded to the power equipment historical case knowledge base and indexed. When diagnosing similar defects later, this report can be retrieved using feature matching as a reference template to assist in diagnosis and lifetime prediction.

[0164] For example, before the step of inputting the multimodal precursor feature set into the pre-trained defect evolution trend prediction model to deduce the multimodal physical field distribution at future times and obtain the predicted physical field component sequence, the method may further include: step S210: extracting nonlinear dynamic parameters for the enhanced multimodal evolution features of each three-dimensional grid node in the multimodal precursor feature set, calculating the Lyapunov exponent of the ultrasonic energy accumulation rate feature changing with time at each node, and determining whether the ultrasonic energy evolution at the three-dimensional grid node has entered a chaotic state based on the positive or negative sign of the Lyapunov exponent.

[0165] Phase space reconstruction is performed on the time series of the ultrasonic energy accumulation rate feature f_us_rate_ijk for each node P_ijk. Using the delayed coordinate method, the delay time τ and embedding dimension m are selected to reconstruct the one-dimensional time series into an orbit in m-dimensional phase space. The delay time τ is determined by calculating the first minimum value of the average mutual information function of the sequence, and the embedding dimension m is determined by the pseudo-nearest neighbor method. The Lyapunov exponent λ_max = (1 / N)∑ln(d_j(Δt) / d_j(0)) is calculated. If λ_max is greater than 0, the ultrasonic energy evolution of that node is determined to have entered a chaotic state; if λ_max is less than or equal to 0, it is a regular motion state.

[0166] Step S220: Perform recursive graph analysis on the intensity characteristic sequence of UHF pulse activity on each three-dimensional grid node, calculate the deterministic percentage and average diagonal length in the recursive graph, and quantify the periodicity or randomness of UHF pulse activity based on the deterministic percentage and average diagonal length.

[0167] Phase space reconstruction is performed on the time series of the UHF pulse activity intensity feature f_uhf_count_ijk for each node P_ijk, using the same method as step S210. The distance between any two points in the reconstructed phase space is calculated, and a distance matrix is ​​constructed. A threshold ε is set, and points with a distance less than ε are marked as recursive points 1, with the rest marked as 0, resulting in a recursion graph matrix R. The deterministic percentage DET is calculated from the recursion graph, which is the proportion of the number of recursive points on the diagonal structure to the total number of recursive points, reflecting the degree of determinism of the system. The average diagonal length L_avg is calculated, which is the average diagonal length, reflecting the average period of the system. If DET is close to 1 and L_avg is large, it indicates that the pulse activity exhibits strong periodicity; if DET is close to 0 and L_avg is small, it indicates that the pulse activity exhibits randomness.

[0168] Step S230: Perform detrended fluctuation analysis on the infrared radiation temperature time series of each three-dimensional grid node, calculate the scaling index, and determine the long-range correlation characteristics of temperature fluctuations based on the numerical range of the scaling index.

[0169] Detrended fluctuation analysis was performed on the infrared radiation temperature time series IR_ijk(t) for each node P_ijk. First, the series was integrated to obtain the cumulative sum series Y(k) = ∑(IR_ijk(t) - avg(IR)). Y(k) was divided into non-overlapping windows of length n. Within each window, the local trend Y_n(k) was fitted using the least squares method, and the detrended fluctuation function F(n) = √((1 / N)∑(Y(k) - Y_n(k))^2) was calculated. This calculation was repeated for different window scales n to obtain the relationship curve between logF(n) and logn. The slope of the curve is the scaling exponent α. If α is greater than 0.5 and less than 1, it indicates that the temperature fluctuation has a long-range positive correlation, i.e., persistence; if α equals 0.5, it indicates white noise with no correlation; if α is less than 0.5, it indicates an inverse correlation.

[0170] Step S240: Combine the Lyapunov exponent, determinism percentage, average diagonal length and scaling exponent calculated for each three-dimensional mesh node into the dynamic state feature vector of that three-dimensional mesh node.

[0171] For each node P_ijk, the Lyapunov exponent λ_ijk calculated in step S210, the percentage of certainty DET_ijk and the average diagonal length L_avg_ijk calculated in step S220, and the scaling exponent α_ijk calculated in step S230 are combined into a four-dimensional dynamic state feature vector F_dynamics_ijk=[λ_ijk, DET_ijk, L_avg_ijk, α_ijk].

[0172] Step S250: Perform spatial clustering on the dynamic state feature vectors of all three-dimensional mesh nodes to identify spatial regions with the same or similar dynamic states. Group the three-dimensional mesh nodes belonging to the same spatial region into one dynamic region, and select representative nodes in each dynamic region for subsequent prediction calculations.

[0173] K-means clustering is performed on the dynamic state feature vectors F_dynamics_ijk of all nodes, with the number of clusters K determined by the silhouette coefficient method. After clustering, each node obtains a dynamic state label L_dynamics. Nodes with the same label and spatial adjacency are merged into the same dynamic region Region_dyn_r. Within each dynamic region, the node whose feature vector is closest to the region center is selected as the representative node N_rep_r, and the enhanced multimodal evolution features of this node are used as input for subsequent prediction models.

[0174] Step S260: Based on the dynamic state feature vectors of different dynamic regions, select a prediction model structure for each dynamic region. The specific selection rules are as follows: if the Lyapunov exponent in the dynamic state feature vector is greater than 0, then a long short-term memory network is selected as the prediction model for the residual abnormal region; if the Lyapunov exponent is not greater than 0 and the scaling exponent is greater than the preset strong correlation threshold, then a gated recurrent unit network is selected as the prediction model for the residual abnormal region; if the Lyapunov exponent is not greater than 0 and the scaling exponent is not greater than the strong correlation threshold, then a standard recurrent neural network is selected as the prediction model for the residual abnormal region.

[0175] For each dynamic region (Region_dyn_r), a prediction model is selected based on the dynamic state feature vector of its representative node. A preset strong correlation threshold, Th_alpha, is set to 0.8. If λ_rep is greater than 0, a Long Short-Term Memory (LSTM) network is chosen as the prediction model for that region because it can capture the long-range dependencies of chaotic systems. If λ_rep is not greater than 0 and α_rep is greater than Th_alpha, a gated recurrent unit (GRU) network is chosen because it has fewer parameters and is suitable for handling strongly correlated but non-chaotic time series. If λ_rep is not greater than 0 and α_rep is not greater than Th_alpha, a standard recurrent neural network is chosen to handle weakly correlated sequences.

[0176] Step S270: Input the multimodal precursor features of the representative node in each dynamic region into the prediction model selected for the residual abnormal region, perform forward calculation independently through each prediction model, and output the predicted multimodal feature vector of the representative node in the residual abnormal region at future time.

[0177] For each dynamic region, Region_dyn_r, the multimodal precursor features (enhanced multimodal evolution features F_enhanced_rep) representing its node N_rep_r are used as input and fed into the prediction model selected for that region. Each region's prediction model performs forward computation independently. The Long Short-Term Memory (LSTM) network model includes input gates, forget gates, output gates, and cell state updates, processing the input sequence and outputting predicted features. The Gated Recurrent Unit (GRU) network model includes update and reset gates, simplifying computation. The standard recurrent neural network is a simple fully connected layer with feedback connections. Each model outputs a sequence of predicted multimodal feature vectors F_pred_rep(t) representing the node at future time steps T_future.

[0178] Step S280: For the three-dimensional mesh nodes that are not representative nodes in each dynamic region, based on their spatial positional relationship with the representative nodes and the similarity of their dynamic state feature vectors, the spatial interpolation method is used to interpolate the predicted multimodal feature vectors of these non-representative nodes from the prediction results of the representative nodes.

[0179] For each non-representative node P_ijk within a dynamic region, calculate its spatial distance d_ijk_rep with the representative node N_rep and the similarity of its dynamic state feature vectors sim_ijk_rep = cosine similarity (F_dynamics_ijk, F_dynamics_rep). Use inverse distance weighted interpolation, with interpolation weights w_ijk_rep = (sim_ijk_rep / d_ijk_rep^2) to normalize. Then, normalize the predicted features F_pred_rep(t) of the representative node. w_ijk_rep yields the predicted multimodal feature vector F_pred_ijk(t) for the non-representative node. Interpolation is repeated for each future time t.

[0180] Step S290: Reorganize the predicted multimodal feature vectors of all three-dimensional grid nodes according to their spatial locations to form the predicted physical field component sequence, and record the prediction model type identifier used by each three-dimensional grid node, which will be used to apply different residual weights to different regions in subsequent prediction residual analysis.

[0181] The predicted multimodal feature vectors F_pred_ijk(t) obtained from interpolation of all nodes are reorganized according to spatial indices (i, j, k) to form a complete predicted physical field component sequence Y_pred_seq with the shape (T_future, I_max, J_max, K_max, 4). Simultaneously, the predicted model type identifier Model_type_ijk corresponding to the dynamic region to which each node belongs is recorded. In the residual analysis of the subsequent step S150, different residual weights are applied to the outputs of different types of predicted models. For example, a higher confidence level (lower residual weight) is given to the prediction results of the Long Short-Term Memory network, while a lower confidence level (higher residual weight) is given to the prediction results of the standard Recurrent Neural Network, in order to balance the impact of differences in the accuracy of different predicted models on the final diagnosis.

[0182] In one exemplary embodiment, a power equipment defect diagnosis system based on multimodal data fusion is provided. This system can be a terminal, server, etc., and its internal structure diagram can be as follows: Figure 2 As shown, the power equipment defect diagnosis system based on multimodal data fusion includes a processor, memory, input / output interface, communication interface, display unit, and input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interface. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides the environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface is used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, near-field communication, or other technologies. When the computer program is executed by the processor, it implements a power equipment defect diagnosis method based on multimodal data fusion. The display unit is used to form a visually visible image and can be a display screen, projection device, or virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device can be a touch layer covering the display screen, or a button, trackball, or touchpad set on the casing of a power equipment defect diagnosis system based on multimodal data fusion, or an external keyboard, touchpad, or mouse, etc.

[0183] It should be noted that the above technical solution can be fully implemented by an embodied intelligent robot. Specifically, using an embodied intelligent robot as a mobile carrier, it can autonomously navigate to a pre-set inspection station in the substation. Relying on its onboard central control unit and power frequency phase-locked loop, it can simultaneously trigger multi-modal sensors such as ultrasonic, UHF, infrared thermal imaging, and visible light cameras to achieve spatiotemporal aligned acquisition of multi-source signals. The embodied intelligent robot uses a quick-change tool head at the end of its robotic arm to precisely place the adhesive ultrasonic probe at the inspection point on the transformer tank wall. Simultaneously, it uses a built-in UHF antenna, infrared thermal imager, and visible light camera to acquire the original signals and their corresponding spatial position markers in a unified spatial coordinate system. The acquired multi-modal data is transmitted to a memory buffer via the PCIe bus of the embodied intelligent robot's industrial control computer. Then, a series of complex calculation processes can be completed on the embodied intelligent robot itself or on a cloud platform, including three-dimensional spatial mesh mapping, physical field component construction, spatiotemporal evolution law analysis, precursor feature extraction, defect evolution trend prediction, multi-modal residual calculation, and defect type identification. Ultimately, the robot uploads the diagnostic results (including defect location coordinates, category identification, and remaining life prediction) to the monitoring platform in real time, and can generate defect backtracking analysis reports and 3D reconstruction animations as needed, realizing full-process automation and intelligence from data collection, fusion analysis to diagnostic decision-making.

[0184] It should be noted that, in order to simplify the description of the present invention and thus help to understand one or more embodiments of the invention, multiple features may sometimes be grouped into one embodiment, drawing or description thereof in the foregoing description of the embodiments of the present invention.

Claims

1. A method for diagnosing defects in power equipment based on multimodal data fusion, characterized in that, The method includes: Acquire a set of multimodal raw signals synchronously collected from power equipment and a set of spatial location markers corresponding to the set of multimodal raw signals. The set of multimodal raw signals includes ultrasonic raw signals, ultra-high frequency raw signals, infrared thermal image raw frames, and visible light video raw frames. The multimodal original signal set is mapped to a unified three-dimensional spatial grid based on the spatial location marker set, and corresponding physical field components are constructed on each three-dimensional grid node to obtain a physical field component sequence. The physical field component sequence includes ultrasonic energy density, ultra-high frequency field strength amplitude, infrared radiation temperature and visible light reflection intensity. Each physical field component together characterizes the multimodal physical state of the three-dimensional grid node. Spatiotemporal evolution law analysis is performed on the physical field component sequence to extract the variation trend of each modal physical quantity in the continuous time series and the coupling perturbation characteristics of the interaction between modes, and generate a set of multimodal precursor features in the early stage of defects. The multimodal precursor feature set is input into a pre-trained defect evolution trend prediction model to extrapolate the multimodal physical field distribution at future times, thereby obtaining a predicted physical field component sequence. By comparing the predicted physical field component sequence with the actual real-time physical field components, the multimodal prediction residuals on each three-dimensional grid node are calculated. Based on the spatial distribution and clustering degree of the multimodal prediction residuals, abnormal residual regions are identified and defect types are determined. Finally, a diagnostic result containing defect location coordinates and defect category identifiers is generated.

2. The power equipment defect diagnosis method based on multimodal data fusion according to claim 1, characterized in that, The step of mapping the multimodal raw signal set to a unified three-dimensional spatial grid based on the spatial location marker set, and constructing corresponding physical field components at each three-dimensional grid node, includes: Based on the three-dimensional structural model of the power equipment, a spatial mesh is generated inside and on the surface of the power equipment, consisting of multiple hexahedral mesh units. Each vertex of the hexahedral mesh unit is used as a three-dimensional mesh node and is assigned a unique spatial coordinate identifier. The spatial coordinates and beam pointing parameters of the ultrasonic sensors in the spatial location marker set are analyzed. Energy envelope extraction processing is performed on each raw ultrasonic signal to obtain the instantaneous ultrasonic energy value at each sampling time. According to the ultrasonic propagation attenuation model in the medium, the instantaneous ultrasonic energy value is propagated back to each three-dimensional grid node of the three-dimensional spatial grid. The ultrasonic energy density contribution value received by each three-dimensional grid node at the corresponding time is calculated. The contribution values ​​of all raw ultrasonic signals are superimposed and fused to generate the ultrasonic energy density tensor component of each three-dimensional grid node in the continuous time series. The spatial coordinates and antenna pattern parameters of the UHF sensor in the spatial location marker set are analyzed. The field strength amplitude demodulation processing is performed on each UHF raw signal to obtain the UHF instantaneous field strength amplitude at each sampling time. According to the radiation propagation model of UHF electromagnetic waves in space, the UHF instantaneous field strength amplitude is mapped to each three-dimensional grid node of the three-dimensional spatial grid. The field strength amplitude contribution value of each three-dimensional grid node at the corresponding time is calculated. The contribution values ​​of all UHF raw signals are vector synthesized to generate the UHF field strength amplitude tensor component of each three-dimensional grid node in the continuous time series. The spatial coordinates and imaging projection matrix of the infrared thermal imager in the spatial location marker set are analyzed. Non-uniformity correction is performed on each raw infrared thermal image frame to obtain the corrected infrared temperature distribution matrix. The mapping relationship between the pixel coordinates of the infrared image and the coordinates of the three-dimensional grid nodes in the three-dimensional space is established based on the imaging projection matrix. The infrared radiation temperature value corresponding to each three-dimensional grid node is extracted from the infrared temperature distribution matrix by the bilinear interpolation method to generate the infrared radiation temperature tensor component of each three-dimensional grid node in the continuous time series. The spatial coordinates and imaging projection matrix of the visible light camera in the spatial location marker set are analyzed. Illumination uniformity compensation is performed on each raw visible light video frame to obtain the compensated visible light reflection intensity distribution matrix. A mapping relationship between the pixel coordinates of the visible light image and the coordinates of the three-dimensional grid nodes is established based on the imaging projection matrix. The visible light reflection intensity value corresponding to each three-dimensional grid node is extracted from the visible light reflection intensity distribution matrix using the bilinear interpolation method to generate the visible light reflection intensity tensor component of each three-dimensional grid node in the continuous time series. The ultrasonic energy density tensor component, ultra-high frequency field strength amplitude tensor component, infrared radiation temperature tensor component, and visible light reflectance intensity tensor component of each three-dimensional mesh node at the same moment are organized as independent, spatially aligned data channels to form multi-channel data characterizing the node state at that moment. The multi-channel data of all three-dimensional mesh nodes at the same time are reorganized according to the spatial arrangement order of the three-dimensional mesh nodes to generate a three-dimensional spatial data volume with multiple independent channels, where each channel corresponds to a physical field distribution of a mode. The physical field components at multiple consecutive moments are stacked in chronological order to generate the physical field component sequence containing spatiotemporal dimension information.

3. The power equipment defect diagnosis method based on multimodal data fusion according to claim 1, characterized in that, The analysis of the spatiotemporal evolution of the physical field component sequence extracts the variation trends of each modal physical quantity in the continuous time series and the coupling perturbation characteristics of the interaction between modes, generating a set of multimodal precursor features for the early stage of defects, including: The ultrasonic energy density time series of each three-dimensional grid node in the physical field component sequence is subjected to trend decomposition processing. The ultrasonic energy density time series is decomposed into long-term trend component, periodic fluctuation component and random disturbance component by using the local weighted regression scatter smoothing method. The slope change rate of the long-term trend component is extracted as the ultrasonic energy accumulation rate feature. Pulse analysis processing is performed on the time series of ultra-high frequency field strength amplitude of each three-dimensional grid node in the physical field component sequence to detect transient pulse events in the time series of ultra-high frequency field strength amplitude. The occurrence frequency and amplitude distribution of transient pulse events within each time window are statistically analyzed to generate ultra-high frequency pulse activity intensity characteristics and ultra-high frequency pulse phase distribution characteristics. The infrared radiation temperature time series of each three-dimensional grid node in the physical field component sequence is subjected to thermodynamic analysis. The first-order difference sequence and the second-order difference sequence of the infrared radiation temperature time series are calculated. The peak feature of temperature rise rate is extracted from the first-order difference sequence, and the feature of the starting point of temperature acceleration change is extracted from the second-order difference sequence. Texture change analysis is performed on the time series of visible light reflection intensity of each three-dimensional grid node in the physical field component sequence. The structural similarity index of the visible light reflection intensity tensor components at adjacent time points is calculated. The texture abrupt change time and abrupt change amplitude features are extracted based on the time change curve of the structural similarity index. At each three-dimensional mesh node, the ultrasonic energy accumulation rate feature extracted from ultrasonic signals, the ultra-high frequency pulse activity intensity feature and ultra-high frequency pulse phase distribution feature extracted from ultra-high frequency signals, the temperature rise rate peak feature and temperature acceleration change start point feature extracted from infrared signals, and the texture change time and change amplitude feature extracted from visible light signals are retained. These features together constitute the multimodal trend feature set of the three-dimensional mesh node. Spatial correlation analysis is performed on the trend characteristics of adjacent three-dimensional mesh nodes under the same mode. For the peak characteristics of ultrasonic energy accumulation rate, ultra-high frequency pulse activity intensity, and infrared temperature rise rate, the spatial covariance matrix of each three-dimensional mesh node and its neighboring three-dimensional mesh nodes is calculated on each characteristic. Based on the eigenvalue decomposition results of the spatial covariance matrix, the spatial coupling strength characteristics reflecting the coordinated changes in the local area are extracted. Calculate the time-varying cross-correlation function between ultrasonic energy density and UHF field strength amplitude at each three-dimensional grid node, and extract the maximum correlation coefficient and its corresponding delay time from the time-varying cross-correlation function as ultrasonic-UHF coupling delay characteristics. Calculate the local gradient mutual information value between infrared radiation temperature and visible light reflection intensity at each three-dimensional mesh node, and extract the thermal-optical coupling anomaly index feature based on the time evolution curve of the local gradient mutual information value; The calculated spatial coupling strength characteristics reflecting spatial cooperative changes, ultrasonic-ultra-high frequency coupling delay characteristics reflecting cross-modal interactions, and thermal-optical coupling anomaly index characteristics, together with the multimodal trend feature set of the corresponding three-dimensional mesh node, are used as the enhanced multimodal evolution feature set of that node. Cluster analysis is performed on the enhanced multimodal evolution feature sets of all three-dimensional mesh nodes to identify three-dimensional mesh node clusters exhibiting abnormal evolution patterns. The area covered by the three-dimensional mesh node clusters is taken as the defect precursor region, and the enhanced multimodal evolution feature sets of all three-dimensional mesh nodes in the defect precursor region are taken as the early multimodal precursor feature sets of the defect.

4. The method for diagnosing power equipment defects based on multimodal data fusion according to claim 1, characterized in that, The step of inputting the multimodal precursor feature set into a pre-trained defect evolution trend prediction model to extrapolate the multimodal physical field distribution at future times and obtain a predicted physical field component sequence includes: The defect evolution trend prediction model is constructed based on an encoder-predictor-decoder architecture, wherein the encoder is composed of a spatiotemporal convolutional long short-term memory network, the predictor is composed of multiple parallel modality prediction branches, and the decoder is composed of a three-dimensional deconvolutional network. The multimodal precursor feature set is reorganized according to the spatial order of the three-dimensional mesh nodes in three-dimensional space to generate a precursor feature tensor with spatial structure. Each channel of the precursor feature tensor corresponds to a precursor feature type. The precursor feature tensor is input into the encoder, and the dependency relationship of the precursor features in the time and space dimensions is modeled through the gating mechanism of the spatiotemporal convolutional long short-term memory network to extract the hidden state tensor containing historical evolution information. The hidden state tensor is input into each modal prediction branch of the predictor. Each modal prediction branch uses a recurrent neural network with a different structure to capture the unique evolution law of the corresponding mode. The ultrasonic modal prediction branch uses a bidirectional long short-term memory network to process the nonlinear growth trend of ultrasonic energy accumulation. The ultra-high frequency modal prediction branch uses a gated recurrent unit network to process the intermittent burst mode of ultra-high frequency pulse activity. The infrared modal prediction branch uses a convolutional long short-term memory network to process the spatial propagation process of thermal field diffusion. The visible light modal prediction branch uses a self-attention network to process the abrupt change characteristics of texture changes. The intermediate prediction features output from each modality prediction branch are concatenated and fused along the feature dimension to obtain a comprehensive prediction feature tensor that integrates multimodal evolution information. The comprehensive prediction feature tensor is input into the decoder, and the spatial resolution of the comprehensive prediction feature is restored and details are reconstructed through the three-dimensional deconvolution network to generate the initial prediction physics components for the first future time. The initial predicted physical field components of the first future time are combined with the actual physical field components of the historical time to form a new input sequence. The forward calculation process of the encoder, predictor and decoder is repeated to iteratively generate the predicted physical field components of multiple subsequent future times. Modal consistency adjustment is performed on the predicted physical field components generated at each future time to ensure that the physical quantities of different modes at the same time satisfy the preset physical constraint relationship in spatial distribution, and the adjusted predicted physical field components are generated. Arrange all the adjusted predicted physical field components for future moments in chronological order to form the predicted physical field component sequence.

5. The power equipment defect diagnosis method based on multimodal data fusion according to claim 1, characterized in that, The process involves comparing the predicted physical field component sequence with the actual acquired real-time physical field components, calculating the multimodal prediction residual at each 3D grid node, identifying abnormal residual regions and determining defect types based on the spatial distribution and clustering of the multimodal prediction residuals, and generating a final diagnostic result containing defect location coordinates and defect category identifiers. At each future moment, acquire the real-time physical field components that are actually collected for that future moment; For each 3D mesh node, the absolute difference between the ultrasonic energy density value of the 3D mesh node in the real-time physical field component and the predicted ultrasonic energy density value of the 3D mesh node at the corresponding time in the predicted physical field component sequence is calculated as the ultrasonic modal residual of the 3D mesh node. For each 3D grid node, the absolute difference between the UHF field strength amplitude of the 3D grid node in the real-time physical field component and the predicted UHF field strength amplitude of the 3D grid node at the corresponding time in the predicted physical field component sequence is calculated as the UHF modal residual of the 3D grid node. For each 3D mesh node, the absolute difference between the infrared radiation temperature value of the 3D mesh node in the real-time physical field component and the infrared radiation temperature prediction value of the 3D mesh node at the corresponding time in the predicted physical field component sequence is calculated as the infrared modal residual of the 3D mesh node. For each 3D mesh node, the absolute difference between the visible light reflection intensity value of the 3D mesh node in the real-time physical field component and the predicted visible light reflection intensity value of the 3D mesh node at the corresponding time in the predicted physical field component sequence is calculated as the visible light mode residual of the 3D mesh node. For each mode, the mean and standard deviation of the residual values ​​of all three-dimensional mesh nodes under historical normal operating conditions are used to calculate the residual of the mode. The residual of the mode is then normalized by Z-score to obtain the normalized mode residual. For each three-dimensional mesh node, its normalized ultrasonic mode residual, ultra-high frequency mode residual, infrared mode residual and visible light mode residual are retained to form the multimodal residual spectrum of the three-dimensional mesh node. For each mode, the spatial distribution analysis of the normalized modal residuals of all three-dimensional mesh nodes is performed to generate the normalized residual distribution map for each mode; Local maximum detection is performed on the normalized residual distribution map. For any three-dimensional grid node, if its residual value is greater than the residual values ​​of all its neighboring three-dimensional grid nodes, and the difference between its residual value and the average residual value of all its neighboring three-dimensional grid nodes exceeds a preset neighborhood residual threshold, then the three-dimensional grid node is identified as a candidate anomaly seed point. Starting from each candidate anomaly seed point, and combining the residual distribution of each mode, the surrounding three-dimensional mesh nodes that have similar residual values ​​to the seed point in at least one mode and are spatially connected are included in the same residual anomaly region, resulting in multiple residual anomaly regions. For each residual anomaly region, the average value of the normalized modal residuals of all three-dimensional mesh nodes in each mode is calculated to obtain the multimodal anomaly intensity spectrum of the residual anomaly region. The average value of the spatial coordinates of all three-dimensional mesh nodes in the residual anomaly region is then calculated as the coordinates of the region center. Extract the ratio distribution histogram between the ultrasonic modal residual and the ultra-high frequency modal residual of the three-dimensional mesh nodes in each residual abnormal region. Match the ratio distribution histogram with a preset defect type template library based on the shape characteristics of the histogram, and take the defect type with the highest matching degree as the defect category identifier of the residual abnormal region. The center coordinates of residual abnormal regions in the multimodal abnormal intensity spectrum where the regional abnormal intensity exceeds a preset threshold are used as the defect location coordinates. The defect category identifier of the residual abnormal region is associated with the defect location coordinates to generate the final diagnostic result containing the defect location coordinates and defect category identifier.

6. The method for diagnosing power equipment defects based on multimodal data fusion according to claim 5, characterized in that, The step involves extracting a histogram of the ratio distribution between the ultrasonic modal residuals and the ultra-high frequency modal residuals of the three-dimensional mesh nodes within each residual abnormal region. Based on the shape characteristics of the histogram, it is matched against a preset defect type template library. The defect type with the highest matching degree is used as the defect category identifier for that residual abnormal region, including: For each residual abnormal region, the ultrasonic modal residual values ​​of all three-dimensional mesh nodes in the residual abnormal region are collected to form an ultrasonic residual set, and the ultra-high frequency modal residual values ​​of all three-dimensional mesh nodes in the residual abnormal region are collected to form an ultra-high frequency residual set. The ultrasonic residual set and the ultra-high frequency residual set are paired element by element so that each three-dimensional mesh node corresponds to a pair of ultrasonic residual values ​​and ultra-high frequency residual values; Calculate the ratio of ultrasonic residual value to ultra-high frequency residual value at each three-dimensional mesh node to obtain the residual ratio of that three-dimensional mesh node, forming a set of residual ratios for that residual abnormal region; Statistical distribution analysis is performed on the set of residual ratios, dividing the range of residual ratio values ​​into multiple continuous equal-width intervals, counting the number of residual ratio values ​​falling into each interval, and generating a histogram of residual ratio distribution with the interval as the horizontal axis and the frequency as the vertical axis. The residual ratio distribution histogram is normalized by dividing the frequency of each interval by the total number of nodes to obtain the normalized residual ratio probability density histogram. Load reference ratio distribution templates corresponding to multiple defect types from the preset defect type template library. Each reference ratio distribution template is a normalized probability density histogram pre-calculated on a sample area of ​​a known defect type using the same method. Calculate the Barcol distance between the residual ratio probability density histogram and each reference ratio distribution template to obtain multiple Barcol distance values. The smaller the Barcol distance, the more similar the two distributions are. The defect type corresponding to the reference ratio distribution template with the smallest Bartholomew distance is selected as the candidate defect type; Further calculate the correlation coefficient between the residual ratio probability density histogram and the candidate defect type reference ratio distribution template. If the correlation coefficient is greater than the preset similarity threshold, then confirm that the candidate defect type is the defect category identifier of the residual abnormal region. If the correlation coefficient is not greater than the similarity threshold, the reference ratio distribution template with the second smallest Bach distance is selected and the correlation coefficient calculation is repeated until a defect type that meets the correlation coefficient threshold requirement is found. If the correlation coefficients of all reference ratio distribution templates do not meet the requirements, the defect category of the residual abnormal area is marked as an unknown type.

7. The power equipment defect diagnosis method based on multimodal data fusion according to claim 3, characterized in that, The enhanced multimodal evolution features of all 3D mesh nodes are clustered to identify clusters of 3D mesh nodes exhibiting abnormal evolution patterns. The region covered by these clusters is designated as the defect precursor region, and the set of enhanced multimodal evolution features of all 3D mesh nodes within this defect precursor region is designated as the early-stage multimodal precursor feature set of the defect. This includes: The enhanced multimodal evolution feature vectors of all three-dimensional mesh nodes are combined into a set of points in a high-dimensional feature space. Each point corresponds to a three-dimensional mesh node. The coordinates of the point are jointly composed of the ultrasonic energy accumulation rate feature, the ultra-high frequency pulse activity intensity feature, the ultra-high frequency pulse phase distribution feature, the peak temperature rise rate feature, the temperature acceleration change start point feature, the texture change time and change amplitude feature, the spatial coupling strength feature, the ultrasonic-ultra-high frequency coupling delay feature, and the thermal-optical coupling anomaly index feature of the three-dimensional mesh node. Density-based spatial clustering is performed on the point set in the high-dimensional feature space. A neighborhood radius parameter and a minimum number of neighborhood points parameter are set. Core points are identified by traversing each point and calculating the number of neighboring points contained in its neighborhood radius. Mutually reachable core points and their neighborhood boundary points are connected into clusters to generate multiple initial feature clusters. For each initial feature cluster, spatial continuity is verified. The spatial coordinates of all three-dimensional mesh nodes contained in the initial feature cluster are extracted. A spatial adjacency graph of the initial feature cluster is constructed. The presence of isolated nodes or fragmented regions in the spatial adjacency graph is checked. If such nodes or fragmented regions exist, they are removed from the current initial feature cluster. The spatial regions corresponding to the removed initial feature clusters are recalculated to obtain spatially continuous candidate feature clusters. For each candidate feature cluster, the mean vector of the enhanced multimodal evolution feature vectors of all three-dimensional mesh nodes in the candidate feature cluster is calculated as the cluster center. The Euclidean distance between the feature vector of each three-dimensional mesh node and the cluster center is calculated. The sum of the squares of all Euclidean distances is used as the intra-cluster compactness index of the candidate feature cluster. For each candidate feature cluster, the mean point of the spatial coordinates of all three-dimensional grid nodes in the candidate feature cluster is calculated as the cluster space center. The spatial distance between each three-dimensional grid node and the cluster space center is calculated, and the sum of the squares of all spatial distances is used as the spatial cohesion index of the candidate feature cluster. The comprehensive anomaly score of each candidate feature cluster is calculated based on the weighted sum of the intra-cluster density index and the spatial cohesion index. The weighting coefficients are pre-set according to the importance of feature similarity and spatial clustering in historical defect samples. Candidate feature clusters whose comprehensive anomaly scores exceed a preset dynamic threshold are identified as anomalous evolution pattern clusters. The preset dynamic threshold is adaptively determined based on the statistical distribution of the comprehensive anomaly scores of all current candidate feature clusters. For each anomalous evolution mode cluster, analyze the temporal evolution trajectory of the feature vectors of the three-dimensional mesh nodes within the cluster, extract the change curves of each modal physical quantity of the three-dimensional mesh nodes from the physical field component sequence in the past continuous time period, calculate the first and second derivatives of each curve in the time dimension, identify the starting time of the inflection point or accelerated change of the change curve, and take the earliest inflection point or accelerated change starting time as the defect precursor starting time of the anomalous evolution mode cluster. Based on the cluster space center coordinates and defect precursor start time of each abnormal evolution mode cluster, the corresponding spatial region and time window are marked on the three-dimensional spatial grid. The marked spatial region is taken as the defect precursor region, and the enhanced multimodal evolution feature set of all three-dimensional grid nodes in the defect precursor region is organized into a data structure with spatiotemporal labels, which serves as the multimodal precursor feature set of the early stage of the defect. The multimodal precursor feature set is similar to the historical defect case database of power equipment. The Mahalanobis distance between the current precursor feature set and each historical case precursor feature set is calculated. The top K historical cases with the smallest Mahalanobis distance and their corresponding final defect types are selected as early reference defect type candidates for the current defect. The reference defect type candidates are then added to the multimodal precursor feature set.

8. The method for defect diagnosis of power equipment based on multimodal data fusion according to claim 1, characterized in that, After the steps of comparing the predicted physical field component sequence with the actual acquired real-time physical field components, calculating the multimodal prediction residual on each three-dimensional grid node, identifying abnormal residual regions and determining the defect type based on the spatial distribution clustering degree of the multimodal prediction residual, and generating a final diagnostic result containing defect location coordinates and defect category identifiers, the method further includes: Based on the defect location coordinates in the final diagnostic results, ultrasonic original signal segments and ultra-high frequency original signal segments corresponding to the spatial locations are extracted from the multimodal original signal set. The ultrasonic original signal segments are deconvolved to remove the multiple reflection and aliasing effects of the ultrasonic signal on the propagation path, and the waveform characteristics of the original acoustic emission signal at the defect source are recovered to obtain the ultrasonic waveform of the defect source. Time-frequency analysis is performed on the ultrasonic waveform of the defect source to extract the energy distribution of the ultrasonic waveform of the defect source at different frequency components. Based on the number of dominant frequency peaks and the proportional relationship between the dominant frequencies in the energy distribution, an ultrasonic fingerprint feature vector characterizing the microscopic motion mode of the defect is constructed. The original UHF signal segment is subjected to phase analysis processing. The original UHF signal segment is synchronously segmented according to the power frequency cycle. The amplitude of the UHF signal is extracted at equal phase intervals within each power frequency cycle to generate a phase-resolved amplitude matrix of the UHF signal within the power frequency cycle. Singular value decomposition is performed on the phase-resolved amplitude matrix to extract a first singular vector reflecting the stability of the discharge phase and a second singular vector reflecting the fluctuation of the discharge amplitude. The first singular vector and the second singular vector are combined to form an UHF fingerprint feature vector characterizing the defect discharge mode. Based on the defect location coordinates, the corresponding spatial location of the infrared thermal image raw frame sequence is extracted from the multimodal raw signal set. The temperature change curve of the defect area in the infrared thermal image raw frame sequence is dynamically modeled, and the zero-crossing time of the second derivative of the temperature change with time is calculated. The zero-crossing time is used as the inflection point feature of the heat source energy release rate. Combined with the contrast between the temperature field of the defect area and the temperature field of the surrounding normal area, an infrared fingerprint feature vector characterizing the thermal properties of the defect is constructed. Based on the defect location coordinates, the visible light video original frame sequence corresponding to the spatial location is extracted from the multimodal original signal set. The fractal dimension of the texture image of the region where the defect is located in the visible light video original frame sequence is calculated to obtain the fractal dimension change curve of the texture image at different scales. The scale parameter and dimension parameter corresponding to the curve slope change point are extracted from the fractal dimension change curve to construct a visible light fingerprint feature vector characterizing the evolution of the defect surface morphology. The ultrasonic fingerprint feature vector, the ultra-high frequency fingerprint feature vector, the infrared fingerprint feature vector, and the visible light fingerprint feature vector are input into a pre-constructed multimodal fingerprint fusion network. The autoencoder structure of the multimodal fingerprint fusion network is used to reduce the dimensionality and reconstruct the fingerprint features of each modality. The low-dimensional latent space vector with the smallest reconstruction error is calculated as the unified micro-fingerprint encoding of the defect. The unified micro-fingerprint code of the defect is matched with the mechanism fingerprint codes in the pre-stored defect physical mechanism library. The defect physical mechanism library contains standard micro-fingerprint codes corresponding to different defect types at different development stages. Based on the defect physical mechanism corresponding to the mechanism fingerprint code with the highest matching degree, the micro-evolution mechanism type of the current defect is determined. The micro-evolution mechanism type includes the gradual growth mechanism of electrical tree channels, the hot spot formation mechanism induced by partial discharge, or the fatigue cracking mechanism caused by mechanical vibration. According to the type of micro-evolution mechanism, the corresponding defect development dynamic equation is retrieved from the defect physical mechanism library. The defect development dynamic equation contains a set of differential equations describing the changes of key defect characteristics over time. Using some components of the unified micro-fingerprint encoding as the initial state parameters of the defect development dynamics equation, the differential equation system is numerically integrated in the time domain to deduce the size expansion curve, discharge intensity evolution curve and temperature rise curve of the defect in the future time series. The deduced defect size expansion curve, discharge intensity evolution curve, and temperature rise curve are compared with the trend of the corresponding physical quantity in the predicted physical field component sequence obtained by the defect evolution trend prediction model. The dynamic time warping distance between the multiple curves is calculated. If the dynamic time warping distance is less than the preset matching threshold, the deduced curve is output as a supplementary prediction result of the defect development trajectory, and together with the final diagnosis result, it is used to guide subsequent operation and maintenance decisions.

9. A power equipment defect diagnosis system based on multimodal data fusion, characterized in that, include: processor; A machine-readable storage medium for storing machine-executable instructions of the processor; The processor is configured to execute the power equipment defect diagnosis method based on multimodal data fusion as described in any one of claims 1 to 8 by executing the machine-executable instructions.

10. A computer program product, characterized in that, The computer program product includes machine-executable instructions stored in a computer-readable storage medium. The processor of the power equipment defect diagnosis system based on multimodal data fusion reads the machine-executable instructions from the computer-readable storage medium and executes the machine-executable instructions, causing the power equipment defect diagnosis system based on multimodal data fusion to perform the power equipment defect diagnosis method based on multimodal data fusion as described in any one of claims 1 to 8.