Intelligent monitoring and diagnosis system for partial discharge based on infrared and acoustic fusion

CN122545959APending Publication Date: 2026-08-11HUANENG LIAOCHENG THERMAL POWER CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-19
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

其一、在多模态数据处理层面,各类传感器(声学、光学、红外及环境传感器)采集的数据缺乏高精度的时空同步与对齐机制,导致后续融合分析存在基准偏差;

Benefits of technology

(1) 通过数据处理模块在完成声学数据的声压级计算与噪声滤除后,将得到的有效声学数据、视频数据、红外温度数据以及环境数据统一进行时间戳同步和对齐处理,生成时间同步的多模态目标数据,该机制为后续的多维度特征提取、空间映射和复合评估提供了绝对一致的时间主轴,彻底消除了多传感器异步采集导致的基准偏差问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122545959A_ABST
    Figure CN122545959A_ABST
Patent Text Reader

Abstract

This invention relates to the field of intelligent partial discharge detection and discloses an intelligent partial discharge monitoring and diagnostic system based on infrared and acoustic fusion. The system simultaneously acquires acoustic, video, infrared, and environmental data through a data acquisition and processing module, aligning the timestamps. A sound source spatial localization module calculates the sound source coordinates based on the array's time difference of arrival and overlays the video to generate a fused image. A composite evaluation module combines infrared temperature and acoustic data at the localized location for composite comparison, classifying and labeling risk levels. A spectrum analysis module extracts acoustic data for specific risk labels, locally reconstructs the PRPD spectrum, and outputs partial discharge characteristic analysis results using a pre-set model. A fusion diagnosis module combines the above characteristics with multimodal data for comprehensive anomaly assessment and closed-loop verification. Finally, a diagnostic report generation module calls the multimodal data to generate a structured report. This system achieves progressive purification and mutual verification of multimodal data, significantly improving the accuracy and reliability of partial discharge diagnosis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent partial discharge detection technology, and more specifically to an intelligent partial discharge monitoring and diagnostic system based on the fusion of infrared and acoustic technologies. Background Technology

[0002] Currently, in the field of power equipment condition monitoring, partial discharge detection typically employs single-dimensional sensors for independent operation. For example, ultrasonic sensors are used to collect spatial acoustic signals for time-frequency domain feature analysis, or infrared thermometers are used to obtain thermal radiation images of the equipment surface. In multimodal collaborative monitoring scenarios, acoustic signals, infrared temperature images, and video images are collected separately through independent channels. For the collected acoustic signals, existing spatial positioning algorithms typically calculate the sound source location based on the assumption of an ideal constant sound speed by performing beamforming or time difference of arrival calculations. After obtaining the acoustic positioning coordinates, they are directly superimposed onto the visual monitoring screen through geometric projection, and combined with sound pressure level data rendering to generate a heat map to visualize the spatial distribution of the sound source.

[0003] However, the existing technology still has the following drawbacks: Firstly, at the level of multimodal data processing, the data collected by various sensors (acoustic, optical, infrared and environmental sensors) lacks a high-precision spatiotemporal synchronization and alignment mechanism, which leads to a benchmark deviation in subsequent fusion analysis. Secondly, at the sound source localization algorithm level, although existing equipment has the ability to sense temperature, humidity and air pressure, its localization calculation based on the time difference of arrival of the microphone array is usually based on the assumption of an ideal constant speed of sound. It fails to combine the actual environmental parameters on site to make dynamic physical corrections to the sound wave propagation delay, resulting in limited spatial positioning accuracy in complex on-site environments. Thirdly, at the level of visualization mapping, most existing technologies simply perform geometric projection superposition and heat map rendering on the calculated sound source coordinates, lacking in-depth identification and effectiveness verification analysis of the texture energy characteristics of local areas of the local heat map, making it difficult to distinguish between real local discharge "point source" and background environment "area source" interference. Fourth, at the diagnostic logic level, existing PRPD map deep learning analysis usually executes directly on the original acoustic signal independently, lacking a mechanism for composite risk pre-assessment and hierarchical filtering by combining multimodal parameters such as infrared temperature before map reconstruction. This results in a flat diagnostic process and the failure to form a closed-loop comprehensive anomaly assessment system that progressively purifies and mutually verifies the multimodal data. Summary of the Invention

[0004] In order to overcome the above-mentioned defects of the prior art, the present invention provides a partial discharge intelligent monitoring and diagnostic system based on the fusion of infrared and acoustic technologies to solve the problems existing in the background art.

[0005] This invention provides the following technical solution: a partial discharge intelligent monitoring and diagnostic system based on the fusion of infrared and acoustic technologies, comprising: Data acquisition module: used to acquire multimodal data on partial discharge phenomena of the device under test, specifically including acquiring acoustic data generated by partial discharge, video data of the device under test, and infrared temperature data of the partial discharge area, and simultaneously acquiring current environmental data; Data processing module: used to preprocess the collected multimodal data, calculate the sound pressure level and filter out noise based on the preset decibel threshold to obtain effective acoustic data, and perform time stamp synchronization and alignment processing on the effective acoustic data, video data, infrared temperature data and environmental data to obtain time-synchronized multimodal target data; Sound source spatial localization module: used to extract effective acoustic data and video data from time-synchronized multimodal target data, calculate the sound source localization based on the arrival time difference between the signals of each channel of the microphone array in the effective acoustic data, obtain the spatial coordinate position of the sound source, and overlay the spatial coordinate position of the sound source into the video data in the form of a heat map to generate a fused image; Composite assessment module: Based on the spatial coordinates of the sound source, it locates the sound source in the multimodal target data, extracts the infrared temperature value and acoustic intensity value of the local area corresponding to the spatial coordinates of the sound source, performs composite comparison, classifies the risk level based on the comparison results, and assigns the corresponding risk level label to the data of the local area, thus obtaining multimodal target data with risk level labels; The spectrum analysis module is used to extract effective acoustic data with conventional risk labels at the spatial coordinates of the sound source from multimodal target data with risk level labels. Based on the effective acoustic data, a PRPD spectrum is generated locally and analyzed using a local preset recognition model to output the partial discharge feature analysis results. Fusion Diagnosis Module: Based on the results of partial discharge feature analysis and multimodal target data with risk level labels, it combines the spatial coordinates of the sound source with the fused image to perform a comprehensive anomaly assessment and outputs a comprehensive diagnostic result for partial discharge. Diagnostic report generation module: Based on the comprehensive diagnostic results and the category of risk level labels, it calls the corresponding fused images, PRPD maps, and time-synchronized multimodal target data to generate a structured partial discharge diagnostic report.

[0006] Preferably, the data acquisition module performs multimodal data acquisition on the partial discharge phenomenon of the device under test, specifically including: The system collects spatial acoustic wave signals generated by partial discharge through a MEMS microphone array, acquires optical video signals of the device under test through an autofocus camera, acquires thermal radiation infrared signals of the partial discharge area through an uncooled infrared sensor connected via a Type-C interface, and simultaneously acquires physical signals of temperature, humidity and air pressure of the current space through an environmental sensor. Spatial acoustic signals, optical video signals, thermal radiation infrared signals, and physical signals of temperature, humidity, and air pressure are synchronously input to the local ARM processing chip for analog-to-digital conversion, which converts them into corresponding digital acoustic data, digital video data, digital infrared temperature data, and digital environmental data, respectively.

[0007] Preferably, the data processing module preprocesses the collected multimodal data, specifically including: data cleaning, data transformation, and data normalization. Temperature, humidity, and air pressure values ​​are extracted from the preprocessed environmental data. The dynamic sound velocity correction coefficient for the current environment is calculated using a preset atmospheric sound velocity physical model. Based on the dynamic sound velocity correction coefficient, propagation delay compensation is calculated for the initial arrival timestamps of each channel signal in the preprocessed MEMS microphone array acoustic data. Simultaneously, time-frequency domain conversion calculation is performed on the preprocessed acoustic data. Sound pressure level comparison and frequency band energy integration are performed in conjunction with a preset decibel threshold. After filtering out background noise, the signal-to-noise ratio improvement index is calculated. The signal-to-noise ratio (SNR) enhancement index is compared with the preset SNR pass threshold in real time. If the SNR enhancement index is greater than or equal to the preset SNR pass threshold, the comparison is deemed successful, and the filtered acoustic data corresponding to the current moment is marked as valid and used as valid acoustic data. Finally, the effective acoustic data, video data, and infrared temperature data calculated after propagation delay compensation are resampled and aligned using the frame synchronization pulse of the video data as the main time axis, and time-synchronized multimodal target data is output.

[0008] Preferably, the sound source spatial localization module extracts effective acoustic data and video data after propagation delay compensation calculation from the time-synchronized multimodal target data. By extracting the arrival time difference between the signals of each channel of the MEMS microphone array in the effective acoustic data, and combining it with the dynamic sound speed correction coefficient, it performs beamforming spatial scanning calculation based on the arrival time difference. It also introduces the frequency band energy integral value in the preset target frequency band as the localization calculation weight to calculate the spatial coordinate position of the partial discharge sound source corresponding to the peak sound source energy. Based on the spatial coordinates of the sound source, spatial mapping and visual perspective analysis are performed in conjunction with video data. The three-dimensional spatial coordinates of the sound source are projected and transformed into the two-dimensional pixel coordinate system of the video image. The sound pressure level energy at the coordinate position is quantified and color gradient mapped according to the preset dynamic range of the sound pressure level. The quantized sound pressure level energy distribution is rendered as a heat map at the preset imaging frame rate and superimposed on the video data to generate a fused image. After generating the merged image, it is also used to perform the following determination process: Extract the preset local window region corresponding to the spatial coordinate position of the sound source in the fused image, calculate the sound pressure level energy variance of the heat map pixels in the local window region and the sound pressure level spatial gradient correlation coefficient between adjacent pixels, and perform weighted fusion of the sound pressure level energy variance and the sound pressure level spatial gradient correlation coefficient to calculate the partial discharge heat source mapping effectiveness coefficient. The partial discharge heat source mapping validity coefficient is compared with a preset heat source mapping validity threshold. If the partial discharge heat source mapping validity coefficient is greater than or equal to the preset heat source mapping validity threshold, the partial discharge heat source mapping is determined to be valid. The spatial coordinate position of the sound source is updated to the time-synchronized multimodal target data as valid positioning coordinates, generating multimodal target data with valid positioning coordinates and outputting it to the composite evaluation module. If the partial discharge heat source mapping validity coefficient is less than the preset heat source mapping validity threshold, the partial discharge heat source mapping is determined to be invalid. An abnormal interference label is assigned to the spatial coordinate position of the sound source, triggering an environmental interference warning for the local area. A video recording containing the fused image is generated and output to the terminal. At the same time, the abnormal interference label is synchronized to the time-synchronized multimodal target data to block the following diagnostic process.

[0009] Preferably, the composite assessment module receives multimodal target data with valid positioning coordinates, uses the valid positioning coordinates as a spatial index reference, extracts the highest temperature value of the local area from the infrared temperature data of the multimodal target data with valid positioning coordinates, and extracts the average sound pressure level energy of the local area from the valid acoustic data as a specific characterization parameter of the acoustic intensity value. By inputting the highest temperature value and the average sound pressure level energy into a preset two-dimensional risk mapping model for normalized weighted fusion calculation, the partial discharge composite hazard index is obtained. The partial discharge composite hazard index is matched and compared with the preset multi-level risk classification threshold range step by step. The multi-level risk classification threshold range includes, from high to low, the high-risk defect judgment threshold, the general anomaly judgment threshold, and the normal monitoring judgment threshold. Execute the corresponding logic based on the specific interval level of the partial discharge composite hazard index within the above-mentioned judgment thresholds: If the partial discharge composite hazard index is greater than or equal to the high-risk defect judgment threshold, the data is determined to be in the highest threshold range, and a high-risk defect label is assigned to the data in that local area, triggering direct system early warning intervention. If not, the system continues to determine whether the partial discharge composite hazard index is greater than or equal to the general anomaly judgment threshold. If yes, the data is determined to be in the intermediate threshold range, and a regular risk label is assigned to the data in that local area. If not, the system continues to determine whether the partial discharge composite hazard index is greater than or equal to the normal monitoring judgment threshold. If yes, the data is determined to be in the transition range between normal and abnormal, and a slight early warning label is assigned to the data in that local area. If not, i.e., the partial discharge composite hazard index is less than the normal monitoring judgment threshold, the data is determined to be in the absolutely normal range, and a normal monitoring label is assigned to the data in that local area. Finally, the local area data containing the above labels are integrated to obtain multimodal target data with risk level labels and output.

[0010] Preferably, the spectrum analysis module extracts effective acoustic data with conventional risk labels at the spatial coordinates of the sound source from the multimodal target data with risk level labels, performs phase synchronization processing on the effective acoustic data to obtain the discharge pulse sequence under the power frequency cycle, and constructs a two-dimensional scatter distribution matrix based on the discharge amplitude and power frequency phase of the discharge pulse sequence to reconstruct and generate the PRPD spectrum locally. The PRPD map is input into a local pre-built deep learning model for feature extraction and classification calculation, resulting in a feature matching confidence index output by the model. It is then determined whether the feature matching confidence value is greater than or equal to a preset confidence threshold. If so, the partial discharge type label corresponding to the highest confidence value is directly output, including corona discharge type or surface discharge type. If not, the current PRPD map features are determined to be unclear, and an unknown weak feature label is assigned to the partial discharge data. Finally, the partial discharge type label or unknown weak feature label is output as the partial discharge feature analysis result.

[0011] Preferably, the fusion diagnostic module acquires the partial discharge feature analysis results and multimodal target data with risk level labels, and combines the spatial coordinates of the sound source with the fused image to sequentially perform comprehensive anomaly assessment and evidence chain verification. In the comprehensive anomaly assessment process, the infrared temperature change rate, sound pressure level energy fluctuation value, and discharge pulse repetition rate in the partial discharge characteristic analysis results of the local area are extracted from the multimodal target data with risk level labels. The infrared temperature change rate, sound pressure level energy fluctuation value, and discharge pulse repetition rate are input into the preset comprehensive anomaly assessment model for weighted calculation to obtain the comprehensive anomaly assessment value of partial discharge. During the evidence chain verification process, the partial discharge type label and risk level label are used as joint query conditions. A table lookup and matching are performed in the preset discharge feature and risk benchmark mapping matrix to obtain the preset abnormal benchmark threshold that uniquely corresponds to the joint query condition. It is then determined whether the comprehensive abnormal assessment value of the partial discharge is greater than or equal to the preset abnormal benchmark threshold. If the comprehensive abnormal assessment value of partial discharge is greater than or equal to the preset abnormal benchmark threshold, it is determined that the multimodal data features under the current labeling system corroborate each other, the evidence chain is closed, and a comprehensive diagnostic result of multidimensional features including discharge type, risk level and spatial coordinates is directly output. If the comprehensive anomaly assessment value of partial discharge is less than the preset anomaly benchmark threshold, it is determined that there is a contradiction in the multimodal data characteristics and the evidence chain is broken. Environmental data is extracted from the time-synchronized multimodal target data. Based on the environmental data, dynamic environmental compensation and frequency band optimization calculations are performed on the effective acoustic data to obtain the optimized acoustic data after environmental compensation. The optimized acoustic data is fed back to the spectrum analysis module to trigger the spectrum analysis module to re-execute phase synchronization processing, PRPD spectrum reconstruction and deep learning model feature matching based on the optimized acoustic data, and to receive the partial discharge feature analysis results re-output by the spectrum analysis module; Extract the tag field attributes from the re-received partial discharge feature analysis results, and determine whether the tag field attributes belong to a preset set of known discharge type enumeration values, wherein the preset set of known discharge type enumeration values ​​includes corona discharge type and surface discharge type. If the attribute of the tag field belongs to the preset set of known discharge type enumeration values, the re-identification is determined to be successful. The newly received partial discharge feature analysis results are then fused and updated with the multimodal target data with risk level tags to output the final diagnostic conclusion. If the attribute of the tag field does not belong to the known discharge type enumeration value set, that is, the attribute of the tag field is an unknown weak feature tag, then the re-identification is determined to fail, and the corresponding video data, infrared temperature data, original acoustic features and environmental data are packaged to generate a manual intervention data package and output.

[0012] Preferably, the diagnostic report generation module matches the risk alarm layout format with the discharge type, spatial coordinates, and risk level label categories in the comprehensive diagnostic results, and associates and calls the fused screen, PRPD map, and local area multimodal feature fragments directionally extracted from time-synchronized multimodal target data. The module then encapsulates the above data in a structured manner to generate a partial discharge diagnostic report, and sends the partial discharge diagnostic report to the terminal for display.

[0013] The technical effects and advantages of this invention are as follows: (1) After the acoustic data sound pressure level calculation and noise filtering are completed by the data processing module, the obtained effective acoustic data, video data, infrared temperature data and environmental data are uniformly timestamped and aligned to generate time-synchronized multimodal target data. This mechanism provides an absolutely consistent time axis for subsequent multi-dimensional feature extraction, spatial mapping and composite evaluation, and completely eliminates the benchmark deviation problem caused by asynchronous acquisition of multiple sensors.

[0014] (2) When the positioning algorithm is used by the sound source spatial positioning module to calculate based on the arrival time difference of each channel of the microphone array, it deeply integrates the synchronously collected field environmental data (temperature, humidity, air pressure), calculates the dynamic sound speed correction coefficient in real time through the atmospheric sound speed physical model, and participates it in the sound wave propagation delay compensation and beamforming scanning, thereby effectively overcoming the influence of complex field environment on sound speed and significantly improving the accuracy and robustness of partial discharge spatial positioning.

[0015] (3) After generating the fused image of sound source coordinates and video superimposed by the sound source spatial localization module, it does not output directly, but extracts the sound pressure level energy variance and spatial gradient correlation coefficient of the local window area heat map pixels, and quantifies the effectiveness coefficient of partial discharge heat source mapping. It can accurately capture the difference between the bright divergent texture of the real partial discharge "point source" and the smooth and consistent texture of the background environment "area source" interference, thereby effectively identifying and eliminating environmental interference and blocking invalid positioning data from entering the subsequent process.

[0016] (4) By combining the infrared maximum temperature and acoustic intensity value through the composite evaluation module, the risk level classification and data labeling are completed in advance; then, the spectral analysis module extracts data of specific risk levels (such as conventional risk labels) based on multimodal target data with risk level labels to perform PRPD spectral reconstruction and identification; finally, the fusion diagnosis module performs comprehensive anomaly assessment and evidence chain verification of spectral features, risk labels and fused images; thus, the multimodal data is purified step by step, computing resources are accurately allocated and multidimensional features are mutually verified, which greatly improves the accuracy of qualitative and quantitative diagnosis of partial discharge and the closed-loop capability of the system. Attached Figure Description

[0017] Figure 1 This is a system structure block diagram of the present invention.

[0018] Figure 2 This is a diagram illustrating the method steps of the present invention. Detailed Implementation

[0019] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. In addition, the forms of the various structures described in the following embodiments are merely illustrative. The partial discharge intelligent monitoring and diagnosis system based on infrared and acoustic fusion involved in the present invention is not limited to the structures described in the following embodiments. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] like Figure 1 The embodiment shown provides a partial discharge intelligent monitoring and diagnostic system based on the fusion of infrared and acoustic technologies, including: Data acquisition module: used to acquire multimodal data on partial discharge phenomena of the device under test, specifically including acquiring acoustic data generated by partial discharge, video data of the device under test, and infrared temperature data of the partial discharge area, and simultaneously acquiring current environmental data.

[0021] In this embodiment, the data acquisition module performs multimodal data acquisition on the partial discharge phenomenon of the device under test, specifically including: The system collects spatial acoustic wave signals generated by partial discharge through a MEMS microphone array, acquires optical video signals of the device under test through an autofocus camera, acquires thermal radiation infrared signals of the partial discharge area through an uncooled infrared sensor connected via a Type-C interface, and simultaneously acquires physical signals of temperature, humidity and air pressure of the current space through an environmental sensor. Spatial acoustic signals, optical video signals, thermal radiation infrared signals, and physical signals of temperature, humidity, and air pressure are synchronously input to the local ARM processing chip for analog-to-digital conversion, which converts them into corresponding digital acoustic data, digital video data, digital infrared temperature data, and digital environmental data, respectively.

[0022] It should be specifically noted that the data acquisition process of the data acquisition module is essentially a hardware execution process of synchronously mapping multi-physics analog signals to the digital domain. In terms of physical spatial arrangement, the autofocus-enabled camera and the uncooled infrared sensor connected via a Type-C interface are integrated at the front of the same handheld terminal housing, and the optical and infrared lenses of both adopt a parallel optical axis or coaxial design to achieve strict overlap between the optical video field of view and the infrared thermal imaging field of view; a MEMS microphone array with no less than 100 channels is arranged around the camera and the infrared sensor, so that the central sensitivity area of ​​the acoustic array and the central area of ​​the optical / infrared field of view form a fixed spatial mapping correspondence.

[0023] During the signal acquisition phase, a MEMS microphone array is used to pick up weak spatial acoustic wave signals generated by partial discharge at high density to obtain sufficient spatial sound pressure difference information. The camera acquires optical video signals containing the macroscopic shape of the device. At the same time, the uncooled infrared core converts the thermal radiation infrared signal of the partial discharge area into a level signal (this uncooled architecture and the Type-C high-speed interface are designed to meet the constraints of low power consumption and lossless transmission of handheld devices). The environmental sensor acquires temperature, humidity and air pressure physical signals characterizing the current spatial medium state in parallel. Subsequently, the above four continuously changing analog physical signals are synchronously input to the local ARM processing chip. The multi-channel analog-to-digital converter (ADC) integrated in the chip performs discretization sampling and quantization under the trigger of the same hardware clock beat, forcibly converting the continuous sound waves, light intensity, thermal radiation and environmental parameters into digital acoustic data, digital video data, digital infrared temperature data and digital environmental data that can be directly read by subsequent algorithms.

[0024] For example, in the scenario of partial discharge detection of live equipment in power transmission and distribution systems (such as high-voltage insulators in substations), the ultrasonic waves, heat, and light signals emitted by the insulator surface during discharge are simultaneously captured by the front-end sensor in microseconds. Since a MEMS microphone array with no fewer than 100 channels provides high-density spatial sampling points, combined with real-time temperature, humidity, and air pressure data from environmental sensors, it can accurately capture subtle differences in partial discharge sound pressure. At this point, because the acoustic center and optical axis center have been pre-calibrated, when the ARM chip converts the aforementioned signal slices into four sets of digital matrix streams with a unified reference time within the same clock cycle, a pixel in the video image, a temperature point in the infrared image, and a sound source coordinate point calculated by the microphone array naturally point to the same physical spatial coordinates in the mathematical model. This provides absolutely consistent source data support for subsequent modules to achieve precise pixel-level superposition of the heat map and the video image.

[0025] Data processing module: used to preprocess the collected multimodal data, calculate the sound pressure level and filter out noise based on the preset decibel threshold to obtain effective acoustic data, and perform time stamp synchronization and alignment processing on the effective acoustic data, video data, infrared temperature data and environmental data to obtain time-synchronized multimodal target data.

[0026] In this embodiment, the data processing module preprocesses the collected multimodal data, specifically including: data cleaning, data conversion, and data normalization. Temperature, humidity, and air pressure values ​​are extracted from the preprocessed environmental data. The dynamic sound velocity correction coefficient for the current environment is calculated using a preset atmospheric sound velocity physical model. Based on the dynamic sound velocity correction coefficient, propagation delay compensation is calculated for the initial arrival timestamps of each channel signal in the preprocessed MEMS microphone array acoustic data. Simultaneously, time-frequency domain conversion calculation is performed on the preprocessed acoustic data. Sound pressure level comparison and frequency band energy integration are performed in conjunction with a preset decibel threshold. After filtering out background noise, the signal-to-noise ratio improvement index is calculated. The signal-to-noise ratio (SNR) enhancement index is compared with the preset SNR pass threshold in real time. If the SNR enhancement index is greater than or equal to the preset SNR pass threshold, the comparison is deemed successful, and the filtered acoustic data corresponding to the current moment is marked as valid and used as valid acoustic data. Finally, the effective acoustic data, video data, and infrared temperature data calculated after propagation delay compensation are resampled and aligned using the frame synchronization pulse of the video data as the main time axis, and time-synchronized multimodal target data is output.

[0027] It should be specifically noted that noise filtering is based on a dual filtering mechanism of sound pressure level comparison and frequency band energy integration calculation. It quantifies the filtering effect by calculating the signal-to-noise ratio improvement index of the target frequency band signal relative to the full frequency band signal. At the same time, the timestamp synchronization and alignment processing is a resampling and truncation mapping of the high-frequency acoustic and infrared continuous data streams based on the video frame rate. In addition, a complete fault-tolerant logic is provided as a supplement to another comparison case when performing index evaluation. That is, when the real-time comparison result shows that the signal-to-noise ratio improvement index is less than the preset qualified threshold, the system determines that the current data is contaminated by strong environmental interference, directly marks the data at that moment as invalid and discards it from the cache, and only allows the data that passes the comparison to enter the subsequent fusion process.

[0028] The specific execution process for data cleaning, data transformation, and data normalization is as follows: Data cleaning refers to performing outlier removal and missing value interpolation on the multimodal raw data stream. For video and infrared image frames, a median filtering algorithm is used to remove outlier pixels caused by transient sensor jumps. For time-continuous acoustic and environmental data sequences, a linear interpolation method based on adjacent effective sampling points is used to fill in occasional missing data packets. Data transformation refers to unifying the non-standard physical quantity formats output by different sensors into standard engineering units. Specifically, the raw voltage digital signal acquired by the MEMS microphone is converted into absolute sound pressure in Pascals using a sound pressure sensitivity coefficient. The pressure value is calculated, and the original grayscale matrix of infrared temperature measurement is converted into an absolute temperature matrix in degrees Celsius according to the factory calibration curve. Data normalization processing refers to mapping the converted heterogeneous data to a unified numerical range required for model calculation. Specifically, the global maximum value of the sound pressure absolute value sequence is extracted and a division operation is performed to force mapping to the floating-point range of [-1, 1]. The infrared temperature absolute value sequence is forced to be mapped to the dimensionless range of [0, 1] by using the maximum and minimum value scaling formula (i.e., the current value minus the global minimum value and then divided by the difference between the global maximum value and the minimum value). This completely eliminates the magnitude gap between multi-source sensor data and provides a standardized input benchmark for subsequent time-frequency domain conversion and deep learning models.

[0029] Furthermore, taking the partial discharge detection scenario of live equipment in power transmission and distribution systems (such as high-voltage insulators in substations) as an example, the specific analysis process of the data processing module is as follows: First, in the environmental compensation and propagation delay calculation stage: the system extracts the temperature value T and humidity value H from the preprocessed environmental data and substitutes them into the preset atmospheric sound speed physical model (specifically, the modified empirical formula for atmospheric sound speed). 331.4 is 0 The standard speed of sound at that time (0.6 and 0.0124 are temperature and humidity compensation constants, respectively) is used to calculate the actual speed of sound C in the current air medium; the actual speed of sound C is then divided by the standard reference speed of sound. (Usually taken as 340m / s), the dynamic sound speed correction coefficient K is calculated, and the specific calculation formula is as follows: Subsequently, based on the dynamic sound velocity correction coefficient K, the initial arrival timestamps of no less than 100 channels of signals in the preprocessed acoustic data were determined. The calculation of propagation delay compensation is performed using the following formula: This eliminates array positioning errors caused by differences in sound wave propagation speed due to changes in temperature and humidity.

[0030] Simultaneously, in the exponent calculation and noise filtering stages: the system performs a short-time Fourier transform (STFT) on the preprocessed acoustic data to obtain a time-frequency spectrum, compares the sound pressure level with a set decibel threshold, and removes background noise waveforms below the threshold. Subsequently, it performs frequency band energy integration calculation, specifically: summing the squares of the amplitudes at all discrete frequency points within the user-defined target partial discharge frequency band (within the range of 1kHz-100kHz, such as a characteristic frequency band centered at 40kHz) on the spectrum to obtain the target frequency band energy. Simultaneously, the total frequency band energy is obtained by summing the squares of the amplitudes at all discrete points across the entire frequency band. The ratio of the two ( / Take the logarithm to the base 10 and multiply by 10 to calculate the signal-to-noise ratio enhancement index in decibels.

[0031] Subsequently, in the data validity determination stage: the system compares the calculated signal-to-noise ratio (SNR) enhancement index with the preset SNR pass threshold in real time (this pass threshold is based on statistics from a typical partial discharge sample library, for example, set to 6dB, which means that the target partial discharge signal energy must reach more than 4 times the background residual energy). If the SNR enhancement index is greater than or equal to 6dB, the comparison is deemed to have passed, and the filtered acoustic data corresponding to the current moment is marked as valid and used as valid acoustic data (otherwise, the comparison is deemed to have failed and the data is discarded and blocked).

[0032] Finally, during the timestamp synchronization and alignment phase, the system will undergo the aforementioned propagation delay compensation calculation (i.e., carrying...). The effective acoustic data, video data, and infrared temperature data (timestamp) are resampled and aligned using the frame synchronization pulse of the video data as the main time axis (specifically, based on the 40-millisecond time interval of a 25fps video frame, the effective acoustic data stream and infrared temperature matrix are resampled and sliced, and precisely aligned and bound to the corresponding video image for each frame), outputting time-synchronized multimodal target data.

[0033] Sound source spatial localization module: It is used to extract effective acoustic data and video data from time-synchronized multimodal target data, perform sound source localization calculation based on the arrival time difference between the signals of each channel of the microphone array in the effective acoustic data, obtain the spatial coordinate position of the sound source, and overlay the spatial coordinate position of the sound source into the video data in the form of a heat map to generate a fused picture.

[0034] In this embodiment, the sound source spatial localization module extracts effective acoustic data and video data after propagation delay compensation calculation from the time-synchronized multimodal target data. By extracting the arrival time difference between the signals of each channel of the MEMS microphone array in the effective acoustic data, and combining it with the dynamic sound speed correction coefficient, beamforming spatial scanning calculation based on the arrival time difference is performed. The frequency band energy integral value in the preset target frequency band is introduced as the localization calculation weight to calculate the spatial coordinate position of the partial discharge sound source corresponding to the peak sound source energy. Based on the spatial coordinates of the sound source, spatial mapping and visual perspective analysis are performed in conjunction with video data. The three-dimensional spatial coordinates of the sound source are projected and transformed into the two-dimensional pixel coordinate system of the video image. The sound pressure level energy at the coordinate position is quantified and color gradient mapped according to the preset dynamic range of the sound pressure level. The quantized sound pressure level energy distribution is rendered as a heat map at the preset imaging frame rate and superimposed on the video data to generate a fused image. After generating the merged image, it is also used to perform the following determination process: Extract the preset local window region corresponding to the spatial coordinate position of the sound source in the fused image, calculate the sound pressure level energy variance of the heat map pixels in the local window region and the sound pressure level spatial gradient correlation coefficient between adjacent pixels, and perform weighted fusion of the sound pressure level energy variance and the sound pressure level spatial gradient correlation coefficient to calculate the partial discharge heat source mapping effectiveness coefficient. The partial discharge heat source mapping validity coefficient is compared with a preset heat source mapping validity threshold. If the partial discharge heat source mapping validity coefficient is greater than or equal to the preset heat source mapping validity threshold, the partial discharge heat source mapping is determined to be valid. The spatial coordinate position of the sound source is updated to the time-synchronized multimodal target data as valid positioning coordinates, generating multimodal target data with valid positioning coordinates and outputting it to the composite evaluation module. If the partial discharge heat source mapping validity coefficient is less than the preset heat source mapping validity threshold, the partial discharge heat source mapping is determined to be invalid. An abnormal interference label is assigned to the spatial coordinate position of the sound source, triggering an environmental interference warning for the local area. A video recording containing the fused image is generated and output to the terminal. At the same time, the abnormal interference label is synchronized to the time-synchronized multimodal target data to block the following diagnostic process.

[0035] It needs to be specifically explained that the complete execution logic for "combining dynamic sound velocity correction coefficients to perform beamforming spatial scanning calculations based on time difference of arrival, and introducing the frequency band energy integral value within the preset target frequency band as the positioning calculation weight" is as follows: First, the preceding data processing module performs a fast Fourier transform on the effective acoustic data to obtain a spectrum. Between the upper and lower limits of the set target partial discharge frequency band, the sound pressure amplitude corresponding to all discrete frequency points is squared and summed to calculate the frequency band energy integral value, which is then transmitted with the data stream. Subsequently, discrete grid points are established within the preset three-dimensional search space. For each grid point, the spatial geometric distance from the grid point to each element in the microphone array is used, combined with the dynamic sound velocity correction coefficient calculated by the preceding module (which incorporates the physical distance...). Dividing by the correction coefficient, the theoretical time difference of the sound wave arriving at each array element when the sound source is located at the grid point is calculated in reverse. Then, the actual arrival time difference of each channel is coherently aligned and delayed and added to the theoretical time difference to obtain the initial output power of the grid point. During this calculation, the frequency band energy integral value transmitted above is simultaneously multiplied into the initial output power as a multiplicative factor for weighted enhancement. Finally, all grid points are scanned and the energy peak is found in the output power spectrum after weighted enhancement. The grid point corresponding to the peak is the three-dimensional spatial coordinate position of the sound source. The "visual perspective analysis" is a camera intrinsic parameter model built based on the pinhole imaging principle to realize the dimensionality reduction mapping from three-dimensional physical coordinates to two-dimensional pixel coordinates.

[0036] The effectiveness coefficient of partial discharge heat source mapping is a discrimination model constructed based on the texture differences between the real partial discharge "point source" and the environmental interference "area source" on the heat map. The real partial discharge point source appears as a very small and very bright core on the screen, with extremely large differences in sound pressure values ​​between pixels (i.e., extremely large variance), and the sound pressure decays from the center outwards. The direction of sound pressure change (gradient vector) of adjacent pixels is divergent (i.e., the spatial gradient correlation coefficient is extremely low). On the other hand, the environmental area source interference appears as a large area of ​​smooth color block, with small variance and highly consistent gradient direction (extremely high correlation coefficient). To quantify this directional difference, the specific calculation process of the spatial gradient correlation coefficient is as follows: First, a preset gradient operator (such as the Sobel operator) is used to extract the sound pressure level of each pixel in the local window in the horizontal and vertical directions. The change in sound pressure is used to construct a two-dimensional gradient vector for each pixel. Then, the cosine similarity of the gradient vectors between all adjacent pixels within the window is calculated and averaged to quantify the consistency of the direction of sound pressure change in adjacent regions (the random divergence of gradient vectors in real partial discharge results in a very low mean cosine similarity, while the parallel and unidirectional gradient vectors of environmental interference result in a very high mean). Finally, by weighting and summing the "variance" representing energy concentration with the "gradient correlation coefficient deviation (1 minus the mean cosine similarity)" representing texture divergence, the effectiveness coefficient of partial discharge heat source mapping can be quantified. Simultaneously, the "preset heat source mapping effectiveness threshold" is a boundary judgment value derived statistically from a typical partial discharge point source sample library, used to define whether the thermal pattern conforms to the physical characteristics of point source discharge.

[0037] Furthermore, taking the partial discharge detection of live equipment in power transmission and distribution systems (such as high-voltage insulators in substations) as an example, and combining the output of the preceding module, the specific analysis and calculation execution process of the sound source spatial positioning module is as follows: The system first extracts 100 channels of effective acoustic data from MEMS after dynamic sound velocity correction coefficient compensation and 25fps video data. It then uses the time difference of arrival of each channel to guide beamforming algorithm for spatial grid scanning. Simultaneously, the energy integral value of the target partial discharge frequency band (e.g., 40kHz) obtained previously is added as a multiplicative weighting factor to the grid output power to achieve feature enhancement, thus calculating the three-dimensional spatial coordinates of the sound source. Subsequently, it calls the camera intrinsic parameter model to perform visual perspective projection, transforming the three-dimensional coordinates into the two-dimensional pixel coordinate system of the video image. Color gradient mapping is performed based on the dynamic range of the sound pressure level, and a heatmap is generated at 25fps and superimposed onto the video data to output a fused image. After generating the fused image, the system extracts 31 [unclear text - possibly related to sound source pixel coordinates]. A 31-pixel local window is used to calculate the statistical variance of the sound pressure level of all pixels within the window. The cosine similarity between the direction vectors of sound pressure level change of adjacent pixels is calculated using the gradient operator as the spatial gradient correlation coefficient. The normalized variance and the normalized gradient deviation are weighted and summed with a weight of 0.5 each to calculate the effectiveness coefficient of the partial discharge heat source mapping. Finally, the system compares this coefficient with the heat source mapping effectiveness threshold (e.g., 0.75) set based on the point source sample library. If it is greater than or equal to 0.75, it is determined to meet the point source characteristics (mapped effectively). The spatial coordinates of the sound source are updated as effective positioning coordinates to the time-synchronized multimodal target data, generating multimodal target data with effective positioning coordinates and outputting it to the downstream composite evaluation module. If it is less than 0.75, it is determined to be area source smoothing interference (mapped ineffective). An abnormal interference label is assigned to the spatial coordinates of the sound source, triggering an environmental interference warning, generating a video recording containing fused images and outputting it to the terminal. At the same time, the abnormal interference label is synchronized to the time-synchronized multimodal target data to block the downstream diagnostic process.

[0038] Composite assessment module: Based on the spatial coordinates of the sound source, it locates the sound source in the multimodal target data, extracts the infrared temperature value and acoustic intensity value of the local area corresponding to the spatial coordinates of the sound source, performs a composite comparison, classifies the risk level based on the comparison results, and assigns the corresponding risk level label to the data of the local area, thus obtaining multimodal target data with risk level labels.

[0039] In this embodiment, the composite assessment module receives multimodal target data with valid positioning coordinates. Using the valid positioning coordinates as a spatial index reference, it extracts the highest temperature value of the local area from the infrared temperature data of the multimodal target data with valid positioning coordinates, and extracts the average sound pressure level energy of the local area from the valid acoustic data as a specific characterization parameter of the acoustic intensity value. By inputting the highest temperature value and the average sound pressure level energy into a preset two-dimensional risk mapping model for normalized weighted fusion calculation, the partial discharge composite hazard index is obtained. The partial discharge composite hazard index is matched and compared with the preset multi-level risk classification threshold range step by step. The multi-level risk classification threshold range includes, from high to low, the high-risk defect judgment threshold, the general anomaly judgment threshold, and the normal monitoring judgment threshold. Execute the corresponding logic based on the specific interval level of the partial discharge composite hazard index within the above-mentioned judgment thresholds: If the partial discharge composite hazard index is greater than or equal to the high-risk defect judgment threshold, the data is determined to be in the highest threshold range, and a high-risk defect label is assigned to the data in that local area, triggering direct system early warning intervention. If not, the system continues to determine whether the partial discharge composite hazard index is greater than or equal to the general anomaly judgment threshold. If yes, the data is determined to be in the intermediate threshold range, and a regular risk label is assigned to the data in that local area. If not, the system continues to determine whether the partial discharge composite hazard index is greater than or equal to the normal monitoring judgment threshold. If yes, the data is determined to be in the transition range between normal and abnormal, and a slight early warning label is assigned to the data in that local area. If not, i.e., the partial discharge composite hazard index is less than the normal monitoring judgment threshold, the data is determined to be in the absolutely normal range, and a normal monitoring label is assigned to the data in that local area. Finally, the local area data containing the above labels are integrated to obtain multimodal target data with risk level labels and output.

[0040] It should be specifically noted that the two-dimensional risk mapping model essentially constructs a dimensionless two-dimensional state assessment space. Its core lies in eliminating the physical scale difference between infrared temperature (typically in °C) and acoustic intensity (typically in dB). The specific execution logic is as follows: The upper limit value Tmax of the highest infrared temperature and the upper limit value Smax of the average sound pressure level energy from the historical typical partial discharge sample library are extracted as benchmarks. The real-time acquired highest temperature value T and average sound pressure level energy S are normalized to obtain the temperature normalization index Tnorm (i.e., T / Tmax) and the acoustic normalization index Snorm (i.e., S / Smax). The upper limits Tmax and Smax are not arbitrarily set but are strictly based on this system. The hardware physical limits of the handheld terminal and the actual testing conditions are determined as follows: Specifically, the value of Smax needs to match the upper limit of the dynamic range of sound source localization preset by the system (e.g., 115dB) to ensure that the high-intensity discharge sound signal can be completely mapped to the normalized space without saturation truncation; Tmax needs to be comprehensively calibrated by combining the upper limit of the temperature measurement range of the uncooled infrared core connected to the front end, and the extreme value of the allowable temperature rise of the power transmission and distribution equipment under the maximum load condition (e.g., set to 150℃), so that the benchmark of this two-dimensional evaluation space is strictly anchored within the actual physical boundary of the equipment; subsequently, based on the contribution of "heat accumulation" and "acoustic emission" in the power transmission and distribution equipment to the deterioration of insulation defects, a fixed fusion weight is preset for the two (e.g., setting a temperature weight). The acoustic weight is 0.6. (0.4), derived from the formula for the partial discharge composite hazard index. A continuous hazard index with values ​​strictly between 0 and 1 is calculated. The multi-level risk classification threshold range is based on this 0-1 normalized index space, and the discretized boundary judgment values ​​are set according to the power equipment condition maintenance guidelines. The high-risk defect judgment threshold is set at 0.75 (its physical meaning is that the temperature or acoustic index has approached 75% of the peak value of historical serious defect samples, indicating that the insulation degradation has entered an irreversible critical state); the general anomaly judgment threshold is set at 0.45 (its physical meaning is that the energy accumulation of partial discharge has reached a significant abnormal level that needs to be included in the recent maintenance plan, indicating that the defect is accelerating its expansion); and the normal monitoring judgment threshold is set at 0.20 (its physical meaning is that the discharge intensity is only slightly higher than the baseline fluctuation margin of the background ambient noise, indicating that the equipment is in a transitional state of basic health but requires continuous monitoring). Through these three thresholds, the continuous hazard index is forcibly divided into four non-overlapping hierarchical intervals, thereby realizing the quantitative mapping from physical measurement values ​​to semantic labels of equipment health status.

[0041] The spectrum analysis module is used to extract effective acoustic data with conventional risk labels at the spatial coordinates of the sound source from multimodal target data with risk level labels. Based on this effective acoustic data, a PRPD spectrum is generated locally and analyzed using a local preset recognition model to output the partial discharge characteristic analysis results.

[0042] In this embodiment, the spectrum analysis module extracts effective acoustic data with conventional risk labels at the spatial coordinates of the sound source from the multimodal target data with risk level labels. The effective acoustic data is processed for phase synchronization to obtain the discharge pulse sequence under the power frequency cycle. A two-dimensional scatter distribution matrix is ​​constructed based on the discharge amplitude and power frequency phase of the discharge pulse sequence to reconstruct and generate the PRPD spectrum locally. The PRPD map is input into a local pre-built deep learning model for feature extraction and classification calculation, resulting in a feature matching confidence index output by the model. It is then determined whether the feature matching confidence value is greater than or equal to a preset confidence threshold. If so, the partial discharge type label corresponding to the highest confidence value is directly output, including corona discharge type or surface discharge type. If not, the current PRPD map features are determined to be unclear, and an unknown weak feature label is assigned to the partial discharge data. Finally, the partial discharge type label or unknown weak feature label is output as the partial discharge feature analysis result.

[0043] It should be specifically explained that the phase synchronization processing and the construction of the two-dimensional scatter distribution matrix essentially involve forcibly mapping the continuous time-domain acoustic pulses to the standard 50Hz power frequency phase coordinate system of the power system. The specific execution logic is as follows: the system extracts the phase window (0°-360°) of the ambient power frequency voltage in real time using a zero-crossing detection algorithm. The discharge pulse amplitude in the filtered effective acoustic data is then precisely projected onto the corresponding phase interval according to the time of occurrence, thereby generating a two-dimensional scatter plot with the phase angle on the horizontal axis and the discharge amplitude on the vertical axis. The scatter plot's clustering shape directly reflects the physical mechanism of the discharge. Simultaneously, to meet the constraint of the handheld terminal operating completely offline without a network, the locally pre-built deep learning model is specifically a convolutional neural network that has undergone lightweight pruning (e.g., using a depthwise separable convolutional structure). The definition benchmark for the model's preset discharge types (e.g., corona discharge, surface discharge) is strictly derived from data pre-collected at the power transmission and distribution equipment site using a MEMS microphone array with no fewer than 100 channels. The pure acoustic partial discharge sample library means that the model learns the texture and morphology features of the acoustic partial discharge spectrum rather than the traditional electrical pulse features, thus ensuring an absolute match between the model and the physical mechanism of acoustic acquisition in the system. Its input is the reconstructed PRPD scatter matrix image, and the output is the probability distribution array of each preset discharge type calculated by the Softmax layer. The feature matching confidence value is the maximum probability value in the probability distribution array. The confidence threshold (e.g., set to 0.65) is the boundary value for minimizing misclassification cost obtained by cross-validation statistics based on the intra-class clustering of various typical partial discharge spectra in the acoustic sample library and the inter-class dispersion of the spectrum of complex environmental interference (such as non-real discharge sound waves caused by mechanical vibration and electromagnetic noise). Its physical meaning is that the confidence of the current recognition result must reach the minimum safety redundancy to avoid misjudgment under strong interference environment. It is used to quantitatively define the confidence of the model in the current spectrum texture feature recognition, thereby avoiding the output of incorrect qualitative conclusions when features are ambiguous.

[0044] Furthermore, taking the partial discharge detection of live equipment in power transmission and distribution systems (such as high-voltage insulators in substations) as an example, the specific execution process of the spectrum analysis module is as follows: After the system extracts valid acoustic data with the 'routine risk' label at a certain insulator location, it extracts the discharge pulses within that time period through power frequency synchronization, maps their amplitudes (such as 55dB, 62dB, etc.) to the phase axis of 0°-360° with microsecond-level timestamps, and reconstructs a PRPD scatter plot with a clear 'double-peak asymmetry' morphology; the system then inputs this spectrum matrix into the main unit... The lightweight convolutional neural network pre-installed in the ARM chip extracts the cluster density and phase shift features of the scattered points and outputs a probability distribution array (e.g., corona discharge probability 0.82, surface discharge probability 0.12, unknown interference probability 0.06). The system determines that the maximum probability value of 0.82 is greater than the preset confidence threshold of 0.65, confirms that the identification is effective, and directly outputs the 'corona discharge type' label corresponding to the highest confidence as the partial discharge feature analysis result, and simultaneously transmits it to the downstream fusion diagnostic module for final determination.

[0045] Fusion Diagnosis Module: Based on the results of partial discharge feature analysis and multimodal target data with risk level labels, it combines the spatial coordinates of the sound source with the fused image to perform a comprehensive anomaly assessment and outputs a comprehensive diagnostic result for partial discharge.

[0046] In this embodiment, the fusion diagnostic module acquires the partial discharge feature analysis results and multimodal target data with risk level labels, and combines the spatial coordinates of the sound source with the fused image to perform comprehensive anomaly assessment and evidence chain verification in sequence. In the comprehensive anomaly assessment process, the infrared temperature change rate, sound pressure level energy fluctuation value, and discharge pulse repetition rate in the partial discharge characteristic analysis results of the local area are extracted from the multimodal target data with risk level labels. The infrared temperature change rate, sound pressure level energy fluctuation value, and discharge pulse repetition rate are input into the preset comprehensive anomaly assessment model for weighted calculation to obtain the comprehensive anomaly assessment value of partial discharge. During the evidence chain verification process, the partial discharge type label and risk level label are used as joint query conditions. A table lookup and matching are performed in the preset discharge feature and risk benchmark mapping matrix to obtain the preset abnormal benchmark threshold that uniquely corresponds to the joint query condition. It is then determined whether the comprehensive abnormal assessment value of the partial discharge is greater than or equal to the preset abnormal benchmark threshold. If the comprehensive abnormal assessment value of partial discharge is greater than or equal to the preset abnormal benchmark threshold, it is determined that the multimodal data features under the current labeling system corroborate each other, the evidence chain is closed, and a comprehensive diagnostic result of multidimensional features including discharge type, risk level and spatial coordinates is directly output. If the comprehensive anomaly assessment value of partial discharge is less than the preset anomaly benchmark threshold, it is determined that there is a contradiction in the multimodal data characteristics and the evidence chain is broken. Environmental data is extracted from the time-synchronized multimodal target data. Based on the environmental data, dynamic environmental compensation and frequency band optimization calculations are performed on the effective acoustic data to obtain the optimized acoustic data after environmental compensation. The optimized acoustic data is fed back to the spectrum analysis module to trigger the spectrum analysis module to re-execute phase synchronization processing, PRPD spectrum reconstruction and deep learning model feature matching based on the optimized acoustic data, and to receive the partial discharge feature analysis results re-output by the spectrum analysis module; Extract the tag field attributes from the re-received partial discharge feature analysis results, and determine whether the tag field attributes belong to a preset set of known discharge type enumeration values, wherein the preset set of known discharge type enumeration values ​​includes corona discharge type and surface discharge type. If the attribute of the tag field belongs to the preset set of known discharge type enumeration values, the re-identification is determined to be successful. The newly received partial discharge feature analysis results are then fused and updated with the multimodal target data with risk level tags to output the final diagnostic conclusion. If the attribute of the tag field does not belong to the known discharge type enumeration value set, that is, the attribute of the tag field is an unknown weak feature tag, then the re-identification is determined to fail, and the corresponding video data, infrared temperature data, original acoustic features and environmental data are packaged to generate a manual intervention data package and output.

[0047] It needs to be specifically explained that the comprehensive anomaly assessment model is essentially a joint situational awareness function of the rate of change of multiple physical quantities. Its core lies in capturing the 'dynamic acceleration' of anomaly development rather than absolute static values. To eliminate the problem of not being able to directly weight different sensing physical quantities (such as ℃ / s, dB, times / second) due to differences in the physical units (such as ℃ / s, dB, times / second), the model incorporates an absolute normalization mapping mechanism before calculation. The specific execution logic is as follows: the difference in infrared temperature within a set time window is extracted as the rate of change of infrared temperature, the root mean square difference of the sound pressure level energy sequence is extracted as the sound pressure level energy fluctuation value, and the number of discharge pulses per unit power frequency cycle is counted as the discharge pulse repetition rate. Subsequently, the physical limit peak values ​​of the three (such as the peak temperature change rate of 2.0℃ / s, the peak sound pressure fluctuation of 10dB, and the peak pulse repetition rate of 500 times / second) are extracted from the historical typical partial discharge sample library, and the three types of dynamic data acquired in real time are divided by their corresponding limit peak values. The values ​​are forcibly converted into three dimensionless rate-of-change indices (i.e., temperature dynamic index, acoustic dynamic index, and pulse dynamic index) with values ​​strictly between 0 and 1. Then, preset weights (e.g., 0.4, 0.3, 0.3) are assigned to the three and linearly weighted summation is performed to mathematically calculate a comprehensive anomaly assessment value for partial discharge with values ​​between 0 and 1, thereby quantifying the transient severity of equipment defect deterioration. The discharge characteristic and risk benchmark mapping matrix is ​​a two-dimensional association lookup table constructed based on prior knowledge of power equipment condition maintenance. Its horizontal axis is the risk level label (e.g., conventional risk), and the vertical axis is the partial discharge type label (e.g., corona discharge, surface discharge). Considering that the previous assessment values ​​have been normalized to the 0-1 interval, the preset anomaly benchmark thresholds stored in the matrix are also marked as boundary values ​​between 0 and 1 (e.g., the benchmark threshold corresponding to the above combination of 'conventional risk + corona discharge' is marked as 0).60, its physical meaning is: under the current risk level, the minimum comprehensive dynamic activity required to determine that this type of discharge is truly active and not subject to environmental interference must reach 60% of the limit state), used to define whether the current apparent characteristics and internal mechanisms are consistent; in addition, the dynamic environmental compensation and frequency band optimization calculation triggered when the evidence chain breaks is essentially an adaptive optimization closed loop that re-finds the physical characteristics of the real signal in a complex data stream. The specific calculation execution process is divided into two steps: the first step is the secondary dynamic compensation of the acoustic phase. The system extracts the latest collected ambient temperature and humidity values ​​at the current moment and substitutes them into the sound speed constructed based on atmospheric physical characteristics. The temperature and humidity dynamic correction model (defined as substituting real-time temperature and humidity parameters into an empirical sound velocity formula including a water vapor partial pressure correction term to calculate the current absolute sound velocity propagation reference value in the air medium) calculates the current actual sound velocity and the difference between this actual sound velocity and the sound velocity used in the previous positioning calculation. This difference is then used to perform a secondary proportional correction calibration on the timestamps of each channel of the valid acoustic data in the cache, eliminating drift errors caused by transient temperature and humidity changes during periods of evidence chain breakage, which could lead to delays in sound wave arrival at each array element. The second step is automatic frequency band sliding optimization, where the system recalculates the secondary-compensated acoustic data using a short-time... Fourier transform yields the full-band time spectrum. Then, within the frequency range of 1kHz-100kHz, a sliding frequency band window is set with a preset step size (e.g., 2kHz). Target frequency bands are successively intercepted, and the integral of the signal energy and background noise energy within the band is calculated. The ratio of these two values ​​is then used as the local signal-to-noise ratio (SNR). During the sliding process, the system continuously compares and updates the maximum local SNR value. Simultaneously, this maximum local SNR value is compared with a preset effective evaluation threshold (defined as the lowest reliable SNR obtained statistically based on the noise suppression limit of the system's front-end microphone array under strong electromagnetic and wind noise conditions in a typical substation). The gain boundary, for example, is set to 6dB. Its physical meaning is that only when the signal-to-noise ratio (SNR) of the truncated frequency band is 6dB or more higher than the original SNR of the full band is it considered that the optimization has extracted the true partial discharge characteristic frequency band rather than random noise fluctuations. After the sliding window traverses the entire frequency band, if the maximum local SNR value is greater than or equal to the effective optimization threshold, the center frequency and bandwidth range corresponding to the maximum value are extracted. A bandpass filter is used to accurately extract the data stream within this optimal frequency band from the original compensation data. This is used as the optimized acoustic data after environmental compensation, thereby completely eliminating the noise pollution caused by occasional environmental interference frequency bands on the spectrum reconstruction. For the secondary dynamic compensation of the acoustic phase in the first step mentioned above, the complete closed-loop calculation execution process is as follows: The system first reads the time when the evidence chain breaks (denoted as ). Ambient temperature values ​​synchronously locked in multimodal target data With ambient humidity value And substitute it into the sound speed-temperature and humidity dynamic correction model (expression: ), solve The actual speed of sound in a state of interference that could cause the chain of evidence to break. Simultaneously, the system retrieves the baseline time when the PRPD map was successfully constructed without any breaks (denoted as...). Historical ambient temperature values With humidity value The reference speed of sound was calculated using the same formula. Then, the actual speed of sound was calculated. Relative to the reference speed of sound Difference in change (Right now ), and extract the current raw timestamp sequence Traw(n) of no less than 100 microphone array channels from the valid acoustic data, and then use the time compensation formula The timestamp of each channel is scaled and corrected point-by-point to generate a phase-calibrated acoustic data sequence that eliminates time delay drift errors. For the second step of automatic frequency band sliding optimization, the complete closed-loop calculation process is as follows: the system performs a short-time Fourier transform (STFT) on the phase-calibrated acoustic data sequence to obtain the time-frequency matrix, and calculates the total signal energy of the initial full-band (1kHz-100kHz). Total noise floor energy The initial full-band signal-to-noise ratio was obtained. Subsequently, the system initializes the lower limit frequency of the sliding frequency band window. The frequency band of the sliding window increases in increments of 2kHz (i.e., the frequency band of the j-th sliding window is...). (kHz), for each truncation frequency band, the local signal energy within that band is calculated by integration. With local noise energy The local signal-to-noise ratio of the j-th frequency band is obtained. The system calculates each time... Then, it is compared with the maximum local signal-to-noise ratio of the currently recorded data. Perform a comparison, if Then Updated to And simultaneously record the center frequency corresponding to that frequency band. With bandwidth range When the lower limit frequency of the sliding window After increasing the frequency to over 98kHz and completing a full-band traversal, the system calculates the improvement in maximum local signal-to-noise ratio relative to the full-band signal-to-noise ratio. Determine the amount of increase If the value is greater than or equal to the optimization validity threshold of 6dB, and the threshold is met, then the optimization is considered valid, and the result is directly read from the value. The final center frequency and bandwidth range of the binding are used as the optimal filtering parameters. Based on this, the system instantiates an FIR bandpass filter with corresponding parameters, performs time-domain convolution filtering calculation on the phase-calibrated acoustic data sequence, and uses the time-domain sequence of the filtered output as the final optimized acoustic data after environmental compensation.

[0048] Furthermore, taking the partial discharge detection of live equipment in power transmission and distribution systems (such as high-voltage insulators in substations) as an example, the specific execution process of the fusion diagnostic module is as follows: The system receives the 'routine risk' tag and the 'corona discharge' tag at a certain insulator location, extracts the infrared temperature change rate (e.g., 0.5℃ / s), sound pressure level energy fluctuation value (e.g., 3dB), and discharge pulse repetition rate (e.g., 150 times / second) within 1 second in that area, and performs normalization calculations based on the aforementioned extreme peak values ​​to obtain the temperature dynamic index (0.5 / 2.0=0.25) and the acoustic dynamic index (3 / 10). =0.30), pulse dynamic index (150 / 500=0.30), and weighted summation with weights of 0.4, 0.3, and 0.3 (0.25×0.4+0.30×0.3+0.30×0.3) yields a comprehensive partial discharge anomaly assessment value of 0.28. The system uses 'normal risk' and 'corona discharge' as joint conditions to look up a table and obtains a preset anomaly baseline threshold of 0.60. Since 0.28 is less than 0.60, the system determines that the chain of evidence is broken (i.e., corona discharge exhibiting normal risk should not be accompanied by such a low threshold). (The activity level is contradictory). The system immediately interrupts the original conclusion, extracts the current temperature and humidity data, and performs dynamic environmental compensation and frequency band optimization on the acoustic data. Specifically, the system first calculates the time delay difference based on the latest temperature and humidity using a correction model and corrects the sound wave time delay of each channel. Then, it calculates the signal-to-noise ratio of each frequency band in the range of 1kHz-100kHz with a sliding step size of 2kHz. It finds that the local signal-to-noise ratio of the frequency band centered at 45kHz reaches 8dB, which is greater than the preset effective optimization threshold of 6dB. Based on this, the system determines that the optimization is effective and filters out the clean acoustic data of this frequency band. The data is fed back to the spectrum analysis module for secondary identification. If the label output by the secondary identification is 'surface discharge' (which belongs to the known discharge type enumeration value), the system determines that the re-identification is successful, updates the risk level and merges it with the 'surface discharge' label, and outputs the final comprehensive diagnostic conclusion. If the label output by the secondary identification is still 'unknown weak feature label', the system determines that the re-identification is unsuccessful, and packages the video image, infrared thermogram, original acoustic wave time-frequency map, and temperature and humidity records at that moment into a complete manual intervention data package, which is then output to the terminal for experts to conduct offline in-depth review.

[0049] Diagnostic report generation module: Based on the comprehensive diagnostic results and the category of risk level labels, it calls the corresponding fused images, PRPD maps, and time-synchronized multimodal target data to generate a structured partial discharge diagnostic report.

[0050] In this embodiment, the diagnostic report generation module matches the risk alarm layout format with the discharge type, spatial coordinates, and risk level label categories in the comprehensive diagnostic results, and associates and calls the fused screen, PRPD map, and local area multimodal feature fragments directionally extracted from time-synchronized multimodal target data. The module then encapsulates the above data in a structured manner to generate a partial discharge diagnostic report, and sends the partial discharge diagnostic report to the terminal for display.

[0051] It should be noted that the diagnostic report generation module performs the following structured data mapping and assembly process based on the comprehensive diagnostic results and the risk level label categories: The discharge type and spatial coordinates are extracted from the comprehensive diagnostic results as the core diagnostic fields of the report. The report template library is queried according to the risk level label category to determine the corresponding risk alarm background color and layout format. The fused screen is used as spatial visualization evidence for the report, and the PRPD spectrum corresponding to the partial discharge feature analysis results is used as discharge phase feature evidence for the report. From the time-synchronized multimodal target data, local area video frames, local area infrared temperature data, and sound pressure level time-series waveforms of effective acoustic data associated with spatial coordinate positions are selectively extracted. The local area video frames, local area infrared temperature data, sound pressure level time-series waveforms, risk alarm background colors, core diagnostic fields, spatial visualization supporting materials, and discharge phase characteristic supporting materials are then structured and encapsulated according to the aforementioned layout format to generate the partial discharge diagnostic report.

[0052] like Figure 2 The embodiment shown provides a partial discharge intelligent monitoring and diagnosis method based on the fusion of infrared and acoustic technologies, including the following steps: Step 1: Used for multimodal data acquisition of partial discharge phenomena of the device under test, specifically including acquiring acoustic data generated by partial discharge, video data of the device under test, and infrared temperature data of the partial discharge area, and simultaneously acquiring current environmental data; Step 2: This step is used to preprocess the collected multimodal data. Based on a preset decibel threshold, the acoustic data is used to calculate the sound pressure level and filter out noise to obtain effective acoustic data. The effective acoustic data, video data, infrared temperature data, and environmental data are then time-stamped and aligned to obtain time-synchronized multimodal target data. Step 3: Used to extract effective acoustic data and video data from time-synchronized multimodal target data. Based on the arrival time difference between the signals of each channel of the microphone array in the effective acoustic data, the sound source localization is calculated to obtain the spatial coordinate position of the sound source. The spatial coordinate position of the sound source is then superimposed on the video data in the form of a heat map to generate a fused image. Step 4: Locate the sound source in the multimodal target data based on the spatial coordinates of the sound source, extract the infrared temperature value of the local area corresponding to the spatial coordinates of the sound source and perform a composite comparison, classify the risk level based on the comparison results, and assign the corresponding risk level label to the data of the local area to obtain multimodal target data with risk level labels; Step 5: Extract effective acoustic data with conventional risk labels at the spatial coordinates of the sound source from the multimodal target data with risk level labels. Based on the effective acoustic data, reconstruct and generate a PRPD map locally. Analyze the PRPD map using a local preset recognition model and output the partial discharge feature analysis results. Step 6: Based on the partial discharge characteristic analysis results and multimodal target data with risk level labels, combine the spatial coordinates of the sound source with the fused image to perform a comprehensive anomaly assessment and output the comprehensive diagnostic results of partial discharge. Step 7: Based on the comprehensive diagnostic results and the risk level label category, retrieve the corresponding fused image, PRPD map, and time-synchronized multimodal target data to generate a structured partial discharge diagnostic report.

[0053] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

[0054] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. An intelligent monitoring and diagnosis system for partial discharge based on infrared and acoustic fusion, characterized in that, include: Data acquisition module: used to acquire multimodal data on partial discharge phenomena of the device under test, specifically including acquiring acoustic data generated by partial discharge, video data of the device under test, and infrared temperature data of the partial discharge area, and simultaneously acquiring current environmental data; Data processing module: used to preprocess the collected multimodal data, calculate the sound pressure level and filter out noise based on the preset decibel threshold to obtain effective acoustic data, and perform time stamp synchronization and alignment processing on the effective acoustic data, video data, infrared temperature data and environmental data to obtain time-synchronized multimodal target data; Sound source spatial localization module: used to extract effective acoustic data and video data from time-synchronized multimodal target data, calculate the sound source localization based on the arrival time difference between the signals of each channel of the microphone array in the effective acoustic data, obtain the spatial coordinate position of the sound source, and overlay the spatial coordinate position of the sound source into the video data in the form of a heat map to generate a fused image; Composite assessment module: Based on the spatial coordinates of the sound source, it locates the sound source in the multimodal target data, extracts the infrared temperature value and acoustic intensity value of the local area corresponding to the spatial coordinates of the sound source, performs composite comparison, classifies the risk level based on the comparison results, and assigns the corresponding risk level label to the data of the local area, thus obtaining multimodal target data with risk level labels; The spectrum analysis module is used to extract effective acoustic data with conventional risk labels at the spatial coordinates of the sound source from multimodal target data with risk level labels. Based on the effective acoustic data, a PRPD spectrum is generated locally and analyzed using a local preset recognition model to output the partial discharge feature analysis results. Fusion Diagnosis Module: Based on the results of partial discharge feature analysis and multimodal target data with risk level labels, it combines the spatial coordinates of the sound source with the fused image to perform a comprehensive anomaly assessment and outputs a comprehensive diagnostic result for partial discharge. Diagnostic report generation module: Based on the comprehensive diagnostic results and the category of risk level labels, it calls the corresponding fused images, PRPD maps, and time-synchronized multimodal target data to generate a structured partial discharge diagnostic report. 2.The partial discharge intelligent monitoring and diagnosis system based on infrared and acoustic fusion of claim 1, wherein, The data acquisition module performs multimodal data acquisition on the partial discharge phenomenon of the device under test, specifically including: The system collects spatial acoustic wave signals generated by partial discharge through a MEMS microphone array, acquires optical video signals of the device under test through an autofocus camera, acquires thermal radiation infrared signals of the partial discharge area through an uncooled infrared sensor connected via a Type-C interface, and simultaneously acquires physical signals of temperature, humidity and air pressure of the current space through an environmental sensor. Spatial acoustic signals, optical video signals, thermal radiation infrared signals, and physical signals of temperature, humidity, and air pressure are synchronously input to the local ARM processing chip for analog-to-digital conversion, which converts them into corresponding digital acoustic data, digital video data, digital infrared temperature data, and digital environmental data, respectively. 3.The partial discharge intelligent monitoring and diagnosis system based on infrared and acoustic fusion of claim 2, characterized in that, The data processing module preprocesses the collected multimodal data, specifically including: data cleaning, data transformation, and data normalization. Temperature, humidity, and air pressure values ​​are extracted from the preprocessed environmental data. The dynamic sound velocity correction coefficient for the current environment is calculated using a preset atmospheric sound velocity physical model. Based on the dynamic sound velocity correction coefficient, propagation delay compensation is calculated for the initial arrival timestamps of each channel signal in the preprocessed MEMS microphone array acoustic data. Simultaneously, time-frequency domain conversion calculation is performed on the preprocessed acoustic data. Sound pressure level comparison and frequency band energy integration are performed in conjunction with a preset decibel threshold. After filtering out background noise, the signal-to-noise ratio improvement index is calculated. The signal-to-noise ratio (SNR) enhancement index is compared with the preset SNR pass threshold in real time. If the SNR enhancement index is greater than or equal to the preset SNR pass threshold, the comparison is deemed successful, and the filtered acoustic data corresponding to the current moment is marked as valid and used as valid acoustic data. Finally, the effective acoustic data, video data, and infrared temperature data calculated after propagation delay compensation are resampled and aligned using the frame synchronization pulse of the video data as the main time axis, and time-synchronized multimodal target data is output.

4. The partial discharge intelligent monitoring and diagnosis system based on infrared and acoustic fusion of claim 3, characterized in that, The sound source spatial localization module extracts effective acoustic data and video data after propagation delay compensation calculation from the time-synchronized multimodal target data. By extracting the arrival time difference between the signals of each channel of the MEMS microphone array in the effective acoustic data, and combining it with the dynamic sound speed correction coefficient, it performs beamforming spatial scanning calculation based on the arrival time difference. It also introduces the frequency band energy integral value in the preset target frequency band as the localization calculation weight to calculate the spatial coordinate position of the partial discharge sound source corresponding to the peak sound source energy. Based on the spatial coordinates of the sound source, spatial mapping and visual perspective analysis are performed in conjunction with video data. The three-dimensional spatial coordinates of the sound source are projected and transformed into the two-dimensional pixel coordinate system of the video image. The sound pressure level energy at the coordinate position is quantified and color gradient mapped according to the preset dynamic range of the sound pressure level. The quantized sound pressure level energy distribution is rendered as a heat map at the preset imaging frame rate and superimposed on the video data to generate a fused image. After generating the merged image, it is also used to perform the following determination process: Extract the preset local window region corresponding to the spatial coordinate position of the sound source in the fused image, calculate the sound pressure level energy variance of the heat map pixels in the local window region and the sound pressure level spatial gradient correlation coefficient between adjacent pixels, and perform weighted fusion of the sound pressure level energy variance and the sound pressure level spatial gradient correlation coefficient to calculate the partial discharge heat source mapping effectiveness coefficient. The partial discharge heat source mapping validity coefficient is compared with a preset heat source mapping validity threshold. If the partial discharge heat source mapping validity coefficient is greater than or equal to the preset heat source mapping validity threshold, the partial discharge heat source mapping is determined to be valid. The spatial coordinate position of the sound source is updated to the time-synchronized multimodal target data as valid positioning coordinates, generating multimodal target data with valid positioning coordinates and outputting it to the composite evaluation module. If the partial discharge heat source mapping validity coefficient is less than the preset heat source mapping validity threshold, the partial discharge heat source mapping is determined to be invalid. An abnormal interference label is assigned to the spatial coordinate position of the sound source, triggering an environmental interference warning for the local area. A video recording containing the fused image is generated and output to the terminal. At the same time, the abnormal interference label is synchronized to the time-synchronized multimodal target data to block the following diagnostic process.

5. The partial discharge intelligent monitoring and diagnostic system based on the fusion of infrared and acoustic technologies according to claim 4, characterized in that, The composite assessment module receives multimodal target data with valid positioning coordinates. Using the valid positioning coordinates as a spatial index reference, it extracts the highest temperature value of the local area from the infrared temperature data of the multimodal target data with valid positioning coordinates, and extracts the average sound pressure level energy of the local area from the valid acoustic data as a specific characterization parameter of the acoustic intensity value. By inputting the highest temperature value and the average sound pressure level energy into a preset two-dimensional risk mapping model for normalized weighted fusion calculation, the partial discharge composite hazard index is obtained. The partial discharge composite hazard index is matched and compared with the preset multi-level risk classification threshold range step by step. The multi-level risk classification threshold range includes, from high to low, the high-risk defect judgment threshold, the general anomaly judgment threshold, and the normal monitoring judgment threshold. Execute the corresponding logic based on the specific interval level of the partial discharge composite hazard index within the above-mentioned judgment thresholds: If the partial discharge composite hazard index is greater than or equal to the high-risk defect judgment threshold, the data is determined to be in the highest threshold range, and a high-risk defect label is assigned to the data in that local area, triggering direct system early warning intervention. If not, the system continues to determine whether the partial discharge composite hazard index is greater than or equal to the general anomaly judgment threshold. If yes, the data is determined to be in the intermediate threshold range, and a regular risk label is assigned to the data in that local area. If not, the system continues to determine whether the partial discharge composite hazard index is greater than or equal to the normal monitoring judgment threshold. If yes, the data is determined to be in the transition range between normal and abnormal, and a slight early warning label is assigned to the data in that local area. If not, i.e., the partial discharge composite hazard index is less than the normal monitoring judgment threshold, the data is determined to be in the absolutely normal range, and a normal monitoring label is assigned to the data in that local area. Finally, the local area data containing the above labels are integrated to obtain multimodal target data with risk level labels and output.

6. The partial discharge intelligent monitoring and diagnosis system based on infrared and acoustic fusion of claim 5, wherein, The spectrum analysis module extracts effective acoustic data with conventional risk labels at the spatial coordinates of the sound source from the multimodal target data with risk level labels. It performs phase synchronization processing on the effective acoustic data to obtain the discharge pulse sequence under the power frequency cycle, and constructs a two-dimensional scatter distribution matrix based on the discharge amplitude and power frequency phase of the discharge pulse sequence to reconstruct and generate the PRPD spectrum locally. The PRPD map is input into a local pre-built deep learning model for feature extraction and classification calculation, resulting in a feature matching confidence index output by the model. It is then determined whether the feature matching confidence value is greater than or equal to a preset confidence threshold. If so, the partial discharge type label corresponding to the highest confidence value is directly output, including corona discharge type or surface discharge type. If not, the current PRPD map features are determined to be unclear, and an unknown weak feature label is assigned to the partial discharge data. Finally, the partial discharge type label or unknown weak feature label is output as the partial discharge feature analysis result.

7. The partial discharge intelligent monitoring and diagnosis system based on infrared and acoustic fusion of claim 6, wherein, The fusion diagnostic module acquires the partial discharge feature analysis results and multimodal target data with risk level labels, and combines the spatial coordinates of the sound source with the fused image to perform comprehensive anomaly assessment and evidence chain verification in sequence. In the comprehensive anomaly assessment process, the infrared temperature change rate, sound pressure level energy fluctuation value, and discharge pulse repetition rate in the partial discharge characteristic analysis results of the local area are extracted from the multimodal target data with risk level labels. The infrared temperature change rate, sound pressure level energy fluctuation value, and discharge pulse repetition rate are input into the preset comprehensive anomaly assessment model for weighted calculation to obtain the comprehensive anomaly assessment value of partial discharge. During the evidence chain verification process, the partial discharge type label and risk level label are used as joint query conditions. A table lookup and matching are performed in the preset discharge feature and risk benchmark mapping matrix to obtain the preset abnormal benchmark threshold that uniquely corresponds to the joint query condition. It is then determined whether the comprehensive abnormal assessment value of the partial discharge is greater than or equal to the preset abnormal benchmark threshold. If the comprehensive abnormal assessment value of partial discharge is greater than or equal to the preset abnormal benchmark threshold, it is determined that the multimodal data features under the current labeling system corroborate each other, the evidence chain is closed, and a comprehensive diagnostic result of multidimensional features including discharge type, risk level and spatial coordinates is directly output. If the comprehensive anomaly assessment value of partial discharge is less than the preset anomaly benchmark threshold, it is determined that there is a contradiction in the multimodal data characteristics and the evidence chain is broken. Environmental data is extracted from the time-synchronized multimodal target data. Based on the environmental data, dynamic environmental compensation and frequency band optimization calculations are performed on the effective acoustic data to obtain the optimized acoustic data after environmental compensation. The optimized acoustic data is fed back to the spectrum analysis module to trigger the spectrum analysis module to re-execute phase synchronization processing, PRPD spectrum reconstruction and deep learning model feature matching based on the optimized acoustic data, and to receive the partial discharge feature analysis results re-output by the spectrum analysis module; Extract the tag field attributes from the re-received partial discharge feature analysis results, and determine whether the tag field attributes belong to a preset set of known discharge type enumeration values, wherein the preset set of known discharge type enumeration values ​​includes corona discharge type and surface discharge type. If the attribute of the tag field belongs to the preset set of known discharge type enumeration values, the re-identification is determined to be successful. The newly received partial discharge feature analysis results are then fused and updated with the multimodal target data with risk level tags to output the final diagnostic conclusion. If the attribute of the tag field does not belong to the known discharge type enumeration value set, that is, the attribute of the tag field is an unknown weak feature tag, then the re-identification is determined to fail, and the corresponding video data, infrared temperature data, original acoustic features and environmental data are packaged to generate a manual intervention data package and output.

8. The partial discharge intelligent monitoring and diagnostic system based on the fusion of infrared and acoustic technologies according to claim 7, characterized in that, The diagnostic report generation module matches the risk alarm layout format with the discharge type, spatial coordinates, and risk level label categories in the comprehensive diagnostic results. It also associates and calls the fused screen, PRPD map, and local multimodal feature fragments extracted from time-synchronized multimodal target data. The module encapsulates the above data in a structured manner to generate a partial discharge diagnostic report and sends the partial discharge diagnostic report to the terminal for display.