Deep learning-based method and system for detecting acoustic vibration defects of a support insulator material

CN122814740APending Publication Date: 2026-09-25POWER RES INST OF STATE GRID SHAANXI ELECTRIC POWER CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610948190.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-29
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0004]本发明提供了一种基于深度学习的支柱绝缘子材料声振缺陷检测方法及系统,本发明解决了现有有监督检测方法对缺陷标注样本的强依赖性与现场缺陷样本稀缺性之间的冲突的技术问题,在完全不依赖缺陷标注样本的条件下完成模型训练,将现场小样本场景下的工程可行性从根本上确立,彻底摆脱了现有BP神经网络、SVM等有监督方法对批量缺陷样本采集与人工标注的依赖

Benefits of technology

[0042]本发明提供的技术方案中,本发明解决了现有有监督检测方法对缺陷标注样本的强依赖性与现场缺陷样本稀缺性之间的冲突的技术问题。本发明仅使用正常支柱绝缘子声振特征向量构建训练集,通过去噪重建训练范式驱动深度自动编码器学习正常声振信号的低维本质分布,在完全不依赖缺陷标注样本的条件下完成模型训练,将现场小样本场景下的工程可行性从根本上确立,彻底摆脱了现有BP神经网络、SVM等有监督方法对批量缺陷样本采集与人工标注的依赖。本发明在编码器中间层嵌入SE通道注意力模块,通过全局平均池化与门控子网络对各特征通道权重进行自适应学习,对频域功率谱中固有频率偏移等缺陷敏感通道实现自动强化、对时域均值等冗余通道实现自动抑制,使得正常与缺陷样本在重建误差上的分布间距拉大,相较于现有对所有通道均等处理的标准自动编码器方法,在小样本条件下的缺陷检测区分度明显提升。本发明基于正常样本重建误差的统计分布自动确定声振缺陷阈值,无需人工经验介入,消除了现有声振检测方法中依赖操作人员经验判读功率谱所引入的主观误差,实现了缺陷判决的全程数据驱动与客观量化。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122814740A_ABST
    Figure CN122814740A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of deep learning, and discloses a support insulator material acoustic vibration defect detection method and system based on deep learning, wherein the method comprises the following steps: performing time sequence alignment processing on acoustic signals and vibration signals of normal support insulators, constructing a first acoustic vibration feature vector set, applying self-adaptive mask noise disturbance, and obtaining a second acoustic vibration feature vector set; performing only normal sample-based denoising reconstruction training through a double-path attention deep automatic encoder, obtaining an acoustic vibration defect detection model, performing exponential weighted moving average calculation, and determining an acoustic vibration defect threshold; inputting a to-be-detected acoustic vibration feature vector of a to-be-detected support insulator into the acoustic vibration defect detection model for reconstruction, and outputting a defect detection result; and the model training is completed without relying on defect labeled samples, and full-process data driving and objective quantification of defect judgment are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of deep learning technology, and in particular to a method and system for detecting acoustic and vibration defects in post insulator materials based on deep learning. Background Technology

[0002] Post insulators are core load-bearing components of power grid transmission and transformation systems. If material defects are not detected in time, they can easily lead to sudden fracture accidents, causing large-scale power outages and equipment damage. Existing acoustic vibration detection technology, which identifies defects by applying excitation to the insulator and analyzing the frequency domain power spectrum characteristics of the vibration response signal, has the basic feasibility of live-line testing. However, the interpretation of the test results is highly dependent on the operator's experience. Differences in the technical background of different operators lead to significant uncertainty in the judgment results, resulting in a high rate of misdiagnosis.

[0003] To address the aforementioned issues, existing research attempts to introduce supervised machine learning methods such as BP neural networks and support vector machines (SVM) for automatic classification of acoustic and vibration signals. However, these methods all rely on training with paired normal and defective samples, requiring manual category labeling for each defect signal. However, insulator damage in substation operations is sudden and low-probability, making it extremely difficult to obtain batches of measured acoustic and vibration data for typical defects such as cracks and voids. This results in supervised models failing to obtain a sufficient number of defect samples for training in practical engineering. Furthermore, existing automated detection models lack an adaptive differentiation mechanism for the differences in defect sensitivity among components of the acoustic and vibration feature vectors, applying equal weights to all feature channels. Under small-sample training conditions, this easily leads to the model's attention being diverted to redundant feature dimensions, resulting in insufficient differentiation between normal and defective samples and a persistently high false negative rate. Summary of the Invention

[0004] This invention provides a method and system for detecting acoustic and vibration defects in post insulator materials based on deep learning. This invention solves the technical problem of the conflict between the strong dependence of existing supervised detection methods on defect-labeled samples and the scarcity of field defect samples. It completes model training without relying on defect-labeled samples, fundamentally establishing the engineering feasibility in field small sample scenarios, and completely getting rid of the dependence of existing supervised methods such as BP neural networks and SVM on batch defect sample collection and manual labeling.

[0005] In a first aspect, the present invention provides a method for detecting acoustic vibration defects in post insulator materials based on deep learning, the method comprising:

[0006] The acoustic and vibration signals of a normal post insulator are time-aligned to construct a first acoustic-vibration feature vector set. An adaptive mask noise perturbation is then applied based on the energy distribution of each sample acoustic-vibration feature vector in the first acoustic-vibration feature vector set to obtain a second acoustic-vibration feature vector set.

[0007] Based on the second acoustic vibration feature vector set, the dual-channel attention depth autoencoder that fuses the SE channel attention module and the bidirectional LSTM temporal attention module is trained on denoising and reconstruction based only on normal samples to obtain an acoustic vibration defect detection model. The reconstruction error of the first acoustic vibration feature vector set on the acoustic vibration defect detection model is calculated by exponential weighted moving average to determine the acoustic vibration defect threshold.

[0008] The acoustic and vibration feature vector of the post insulator to be tested is input into the acoustic and vibration defect detection model for reconstruction, and the reconstruction error is obtained. The obtained reconstruction error is compared with the acoustic and vibration defect threshold, and the defect detection result is output.

[0009] In conjunction with the first aspect, in a first implementation of the first aspect of the present invention, the step of performing time-series alignment processing on the acoustic signal and vibration signal of a normal post insulator to construct a first acoustic-vibration feature vector set, and applying adaptive masking noise perturbation based on the energy distribution of each sample acoustic-vibration feature vector in the first acoustic-vibration feature vector set to obtain a second acoustic-vibration feature vector set, includes:

[0010] Simultaneously, acoustic and vibration signals of normal support insulators are acquired, short-time Fourier transform is performed on the acoustic signals, wavelet packet transform is performed on the vibration signals, and the transformation results of the two are aligned on the time axis to obtain time-aligned acoustic and vibration transformation features.

[0011] The time-aligned acoustic transformation features and vibration transformation features are extracted using time-domain statistical features, frequency-domain power spectrum features, and Mel-domain cepstral coefficient features, respectively, to obtain time-domain statistical feature components, frequency-domain power spectrum feature components, and Mel-domain cepstral coefficient feature components; the time-domain statistical feature components, the frequency-domain power spectrum feature components, and the Mel-domain cepstral coefficient feature components are then vector-concatenated and standardized to obtain the first acoustic-vibration feature vector set;

[0012] An adaptive masking noise perturbation is applied to each sample acoustic vibration feature vector in the first acoustic vibration feature vector set based on the energy distribution at each time step to obtain the second acoustic vibration feature vector set.

[0013] In conjunction with the first aspect, in a second implementation of the first aspect of the present invention, the step of applying adaptive masking noise perturbation to each sample acoustic vibration feature vector in the first acoustic vibration feature vector set based on the energy distribution at each time step to obtain a second acoustic vibration feature vector set includes:

[0014] For each sample acoustic vibration feature vector in the first acoustic vibration feature vector set, the energy value of each time step is calculated. Time steps with energy values ​​higher than a preset high energy quantile threshold are defined as high energy regions, and time steps with energy values ​​lower than a preset low energy quantile threshold are defined as low energy regions, thus obtaining the time step division results of high energy regions and low energy regions.

[0015] Based on the time step division results, zero-mean Gaussian noise with a first preset standard deviation as a parameter is applied to each time step in the high-energy region, and zero-mean Gaussian noise with a second preset standard deviation as a parameter is applied to each time step in the low-energy region to obtain an adaptive mask noise vector. The adaptive mask noise vector is then added to the corresponding sample acoustic vibration feature vector step by step to obtain a noisy vibration feature vector, and all the noisy vibration feature vectors are used as the second acoustic vibration feature vector set; wherein, the first preset standard deviation is less than the second preset standard deviation.

[0016] In conjunction with the first aspect, in a third implementation of the first aspect of the present invention, the step of calculating the energy value of each time step for each sample acoustic vibration feature vector in the first acoustic vibration feature vector set, defining time steps with energy values ​​higher than a preset high energy quantile threshold as high energy regions, and defining time steps with energy values ​​lower than a preset low energy quantile threshold as low energy regions, to obtain the time step division results of high energy regions and low energy regions, includes:

[0017] For each time step of the sample acoustic vibration feature vector, the average value of the squared feature values ​​of each dimension of the time step is obtained to get the energy value of the time step. The energy values ​​of all time steps are arranged in order to obtain the energy value sequence of each time step.

[0018] The time steps whose energy values ​​are above a preset high energy quantile are defined as high energy regions, and the time steps whose energy values ​​are below a preset low energy quantile are defined as low energy regions, thus obtaining the time step division results of high energy regions and low energy regions.

[0019] In conjunction with the first aspect, in the fourth implementation of the first aspect of the present invention, the step of training a dual-channel attention depth autoencoder based solely on normal samples using the second acoustic feature vector set to perform denoising and reconstruction training on a dual-channel attention module fused with a bidirectional LSTM temporal attention module, thereby obtaining an acoustic defect detection model, and then calculating the acoustic defect threshold by performing an exponentially weighted moving average calculation on the reconstruction error of the first acoustic feature vector set on the acoustic defect detection model, includes:

[0020] The noisy vibration feature vectors in the second acoustic vibration feature vector set are input into a dual-channel attention deep autoencoder that fuses the SE channel attention module and the bidirectional LSTM temporal attention module for denoising and reconstruction training based only on normal samples, thus obtaining an acoustic vibration defect detection model.

[0021] The acoustic vibration feature vectors of each sample in the first acoustic vibration feature vector set are input into the acoustic vibration defect detection model for reconstruction to obtain the sample reconstruction output vector. The mean square reconstruction error between the sample acoustic vibration feature vector and the sample reconstruction output vector is calculated to obtain the normal sample reconstruction error set.

[0022] The sliding mean and sliding standard deviation of the reconstruction error are calculated by exponential weighted moving average based on the normal sample reconstruction error set. The sum of the sliding mean and the sliding standard deviation multiplied by a preset sensitivity coefficient is determined as the acoustic and vibration defect threshold.

[0023] In conjunction with the first aspect, in the fifth implementation of the first aspect of the present invention, the step of inputting the noisy vibration feature vectors from the second acoustic vibration feature vector set into a dual-channel attention deep autoencoder that fuses the SE channel attention module and the bidirectional LSTM temporal attention module for denoising and reconstruction training based solely on normal samples to obtain an acoustic vibration defect detection model includes:

[0024] The noisy vibration feature vector is input into the encoder of a dual-channel attention deep autoencoder that integrates the SE channel attention module and the bidirectional LSTM temporal attention module for feature compression. The SE channel attention module performs global average pooling on the encoder's intermediate layer feature vector to obtain the statistics of each channel. After two fully connected transformations through a gated subnetwork, the weights of each channel are output. At the same time, the bidirectional LSTM temporal attention module performs bidirectional sequence modeling on the encoder's intermediate layer feature vector to output the weights of each time step. The weights of each channel and the weights of each time step are multiplied by the intermediate layer feature vector channel by channel and time step by time to obtain the channel temporal weighted feature vector.

[0025] The channel temporal weighted feature vector is input into the decoder of the dual-channel attention deep autoencoder that fuses the SE channel attention module and the bidirectional LSTM temporal attention module to reconstruct the training reconstructed output vector. Then, backpropagation is performed on all parameters of the dual-channel attention deep autoencoder that fuses the SE channel attention module and the bidirectional LSTM temporal attention module to update the parameters for the current round.

[0026] Based on the current round update parameters, the dual-channel attention depth autoencoder that integrates the SE channel attention module and the bidirectional LSTM temporal attention module is iterated multiple times. The parameters corresponding to the lowest mean square reconstruction error of the validation set are taken as the optimal parameters to obtain the acoustic and vibration defect detection model.

[0027] In conjunction with the first aspect, in the sixth implementation of the first aspect of the present invention, the step of inputting the channel temporal weighted feature vector into the decoder of the dual-channel attention deep autoencoder that fuses the SE channel attention module and the bidirectional LSTM temporal attention module for reconstruction, obtaining the training reconstruction output vector, and performing backpropagation update on all parameters of the dual-channel attention deep autoencoder that fuses the SE channel attention module and the bidirectional LSTM temporal attention module to obtain the current round update parameters includes:

[0028] The channel temporal weighted feature vector is input into the decoder of the dual-channel attention deep autoencoder that fuses the SE channel attention module and the bidirectional LSTM temporal attention module to perform layer-by-layer dimensional augmentation and reconstruction, and the training reconstruction output vector is obtained.

[0029] The loss function value is calculated based on the reconstructed output vector from the training and the corresponding normal acoustic vibration feature vector.

[0030] Based on the ratio of the current training round to the total training rounds, the current learning rate of the Adam optimizer is dynamically updated using the cosine annealing formula to obtain the updated learning rate. Backpropagation is then performed on the loss function value using the updated learning rate to calculate the gradients of all parameters in the dual-channel attention deep autoencoder that integrates the SE channel attention module and the bidirectional LSTM temporal attention module, and the parameter updates are completed to obtain the updated parameters for the current round.

[0031] In conjunction with the first aspect, in the seventh implementation of the first aspect of the present invention, the step of inputting the acoustic vibration feature vectors of each sample in the first acoustic vibration feature vector set into the acoustic vibration defect detection model for reconstruction to obtain a sample reconstruction output vector, and calculating the mean square reconstruction error between the sample acoustic vibration feature vector and the sample reconstruction output vector to obtain a normal sample reconstruction error set, includes:

[0032] Each sample acoustic vibration feature vector in the first acoustic vibration feature vector set is sequentially input into the acoustic vibration defect detection model for forward propagation reconstruction to obtain the sample reconstruction output vector;

[0033] The normal sample reconstruction error is obtained by taking the square of the difference between each sample acoustic vibration feature vector and the sample reconstruction output vector in each dimension and then averaging the difference over all dimensions. All the normal sample reconstruction errors are then used as the normal sample reconstruction error set.

[0034] In conjunction with the first aspect, in the eighth implementation of the first aspect of the present invention, the step of inputting the acoustic vibration feature vector of the post insulator to be tested into the acoustic vibration defect detection model for reconstruction, obtaining the reconstruction error to be tested, comparing the obtained reconstruction error to be tested with the acoustic vibration defect threshold, and outputting the defect detection result includes:

[0035] The acoustic vibration feature vector of the post insulator to be tested is input into the acoustic vibration defect detection model for reconstruction, and the reconstruction output vector to be tested is obtained. The difference between the acoustic vibration feature vector to be tested and the reconstruction output vector to be tested is taken in each dimension, the square is taken, and the mean is calculated over all dimensions to obtain the reconstruction error to be tested.

[0036] Based on a preset attenuation factor, the acoustic and vibration defect threshold is updated online using an exponentially weighted moving average formula to obtain the current acoustic and vibration defect threshold.

[0037] The reconstruction error to be measured is compared with the current acoustic and vibration defect threshold. If the reconstruction error to be measured is greater than the current acoustic and vibration defect threshold, the defect detection result is output as a defect alarm. If the reconstruction error to be measured is less than or equal to the current acoustic and vibration defect threshold, the defect detection result is output as a normal state.

[0038] Secondly, the present invention provides a deep learning-based acoustic vibration defect detection system for post insulator materials, the deep learning-based acoustic vibration defect detection system for post insulator materials comprising:

[0039] The construction module is used to perform time-series alignment processing on the acoustic signal and vibration signal of a normal post insulator to construct a first acoustic-vibration feature vector set, and to apply adaptive mask noise perturbation based on the energy distribution of each sample acoustic-vibration feature vector in the first acoustic-vibration feature vector set to obtain a second acoustic-vibration feature vector set.

[0040] The training module is used to train the dual-channel attention depth autoencoder that fuses the SE channel attention module and the bidirectional LSTM temporal attention module based on the second acoustic vibration feature vector set, and to perform denoising and reconstruction training based only on normal samples to obtain an acoustic vibration defect detection model. The reconstruction error of the first acoustic vibration feature vector set on the acoustic vibration defect detection model is calculated by exponential weighted moving average to determine the acoustic vibration defect threshold.

[0041] The output module is used to input the acoustic vibration feature vector of the post insulator to be tested into the acoustic vibration defect detection model for reconstruction, obtain the reconstruction error to be tested, compare the obtained reconstruction error to be tested with the acoustic vibration defect threshold, and output the defect detection result.

[0042] The technical solution provided by this invention addresses the conflict between the strong dependence of existing supervised detection methods on defect-labeled samples and the scarcity of on-site defect samples. This invention uses only the acoustic and vibration feature vectors of normal post insulators to construct a training set. Through a denoising and reconstruction training paradigm, it drives a deep autoencoder to learn the low-dimensional essential distribution of normal acoustic and vibration signals, completing model training without relying on defect-labeled samples. This fundamentally establishes the engineering feasibility in small-sample scenarios, completely eliminating the dependence of existing supervised methods such as BP neural networks and SVMs on batch defect sample collection and manual annotation. This invention embeds an SE channel attention module in the middle layer of the encoder. Through global average pooling and a gated sub-network, it adaptively learns the weights of each feature channel, automatically strengthening defect-sensitive channels such as inherent frequency shifts in the frequency domain power spectrum and automatically suppressing redundant channels such as time-domain mean values. This widens the distribution gap in reconstruction errors between normal and defect samples, significantly improving defect detection discrimination under small-sample conditions compared to existing standard autoencoder methods that process all channels equally. This invention automatically determines the threshold of acoustic and vibration defects based on the statistical distribution of reconstruction errors of normal samples, without the need for human experience intervention. It eliminates the subjective errors introduced by the reliance on operator experience to interpret the power spectrum in existing acoustic and vibration detection methods, and realizes full data-driven and objective quantification of defect judgment.

[0043] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention are realized and obtained in accordance with the structures particularly pointed out in the description, claims and drawings.

[0044] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0045] Figure 1 This is a schematic diagram of an embodiment of the deep learning-based acoustic vibration defect detection method for post insulator materials according to the present invention;

[0046] Figure 2 This is a schematic diagram illustrating the construction of acoustic vibration features in an embodiment of the present invention;

[0047] Figure 3 This is a schematic diagram of denoising and reconstruction training based on normal samples in an embodiment of the present invention;

[0048] Figure 4 This is a schematic diagram of acoustic vibration defect detection in an embodiment of the present invention;

[0049] Figure 5This is a schematic diagram of an embodiment of the acoustic vibration defect detection system for post insulator materials based on deep learning in this invention. Detailed Implementation

[0050] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0051] The terms "comprising" and "having," and any variations thereof, used in the embodiments of this invention are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the steps or units listed, but may optionally include other steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices.

[0052] To facilitate understanding of this embodiment, a method for detecting acoustic and vibration defects in post insulator materials based on deep learning, as disclosed in this embodiment of the invention, will first be described in detail. For example... Figure 1 As shown, this method includes the following steps:

[0053] 101. The acoustic signal and vibration signal of the normal post insulator are time-aligned to construct the first acoustic-vibration feature vector set. An adaptive mask noise perturbation is applied based on the energy distribution of the acoustic-vibration feature vectors of each sample in the first acoustic-vibration feature vector set to obtain the second acoustic-vibration feature vector set.

[0054] 102. Based on the second acoustic and vibration feature vector set, the dual-channel attention deep autoencoder that fuses the SE channel attention module and the bidirectional LSTM temporal attention module is trained by denoising and reconstruction based only on normal samples to obtain the acoustic and vibration defect detection model. The reconstruction error of the first acoustic and vibration feature vector set on the acoustic and vibration defect detection model is calculated by exponential weighted moving average to determine the acoustic and vibration defect threshold.

[0055] 103. Input the acoustic vibration feature vector of the post insulator to be tested into the acoustic vibration defect detection model for reconstruction, obtain the reconstruction error, compare the obtained reconstruction error with the acoustic vibration defect threshold, and output the defect detection result.

[0056] This invention automatically determines the threshold of acoustic and vibration defects based on the statistical distribution of reconstruction errors of normal samples, without the need for human experience intervention. It eliminates the subjective errors introduced by the reliance on operator experience to interpret the power spectrum in existing acoustic and vibration detection methods, and realizes full data-driven and objective quantification of defect judgment.

[0057] In one specific embodiment, such as Figure 2 As shown, the process of executing step 101 can specifically include the following steps:

[0058] 1011. Simultaneously acquire acoustic and vibration signals of normal support insulators, perform short-time Fourier transform on the acoustic signals and wavelet packet transform on the vibration signals, and align the transformation results of the two on the time axis to obtain time-aligned acoustic and vibration transformation features.

[0059] 1012. Extract time-domain statistical features, frequency-domain power spectrum features, and Mel-domain cepstral coefficient features from the time-aligned acoustic transformation features and vibration transformation features, respectively, to obtain time-domain statistical feature components, frequency-domain power spectrum feature components, and Mel-domain cepstral coefficient feature components; perform vector concatenation and standardization on the time-domain statistical feature components, frequency-domain power spectrum feature components, and Mel-domain cepstral coefficient feature components to obtain the first acoustic-vibration feature vector set;

[0060] 1013. Based on the energy distribution of each time step, an adaptive mask noise perturbation is applied to the acoustic vibration feature vectors of each sample in the first acoustic vibration feature vector set to obtain the second acoustic vibration feature vector set.

[0061] Specifically, an active acoustic sensor is used to apply random excitation to a post insulator that has been manually confirmed to be defect-free. Simultaneously, acoustic and vibration signals are collected by the acoustic and vibration sensors, respectively. A signal conditioning unit performs filtering and amplification, and a data recording unit digitizes and stores the original signal sequence at a sampling rate of at least 20 kHz, obtaining the original acoustic and vibration signals. A short-time Fourier transform is performed on the original acoustic signal to convert the time-domain acoustic sequence into time-frequency domain acoustic transform features. A wavelet packet transform is performed on the original vibration signal to decompose the time-domain vibration sequence into multi-frequency band vibration transform features. Due to the order-of-magnitude difference in the propagation speeds of sound waves and structural waves in the post insulator, there is a timing misalignment between the two at the acquisition end. The short-time Fourier transform and wavelet packet transform results need to be aligned on the same time axis based on the sampling timestamps of the two signals to eliminate the timing deviation caused by the difference in propagation speed, resulting in time-aligned acoustic and vibration transform features.

[0062] Three types of characterization processing are performed in parallel on the acoustic transformation features and vibration transformation features aligned with the same time sequence. Among them, the time-domain statistical features are used to describe the overall shape and dispersion of the waveform, and can extract the mean, variance, kurtosis, number of zero crossings, and number of flips, so that the model can perceive the signal amplitude distribution, peak degree, and frequency of waveform changes. The frequency-domain power spectrum features are used to reflect the frequency response state of the structure after excitation. After performing a fast Fourier transform on the time-domain sequence, the resonant frequency, anti-resonant frequency, resonant peak amplitude, and power ratio in the target frequency band can be extracted from the power spectral density curve. The target frequency band can be determined based on the statistical results of the normal sample spectrum of the same type of post insulator under the same installation boundary, the same excitation method, and the same sensor layout conditions. For example, it can be taken as 3 kHz to 5 kHz. Factors such as internal material cracks, voids, or deterioration will cause changes in frequency position and energy distribution. The Mel-domain cepstral coefficient features are used to supplement the spectral envelope information, so that the subtle spectral differences in the acoustic and vibration signals that are not easily read directly from a single peak can also be stably encoded. For example, the first 13 cepstral coefficients can be taken as the characterization results.

[0063] Following a pre-defined field order, the time-domain statistical feature components, frequency-domain power spectrum feature components, and Mel-domain cepstral coefficient feature components are concatenated end-to-end to form a single acoustic vibration feature vector. After vector concatenation, each feature dimension is standardized to compress feature components with different physical meanings to a comparable numerical scale. This prevents frequency-based or amplitude-based features from dominating model training due to their large numerical range, and avoids count-based features from being ignored by the network due to their small magnitude. After standardization, all normal samples constitute the first acoustic vibration feature vector set. Using each sample vector in the first acoustic vibration feature vector set as a baseline, zero-mean Gaussian noise with different standard deviations is applied to high-energy and low-energy regions according to the energy distribution at each time step, constructing corresponding adaptive masked noisy versions. All noisy samples are then aggregated into the second acoustic vibration feature vector set.

[0064] In one specific embodiment, the process of performing step 1013 may specifically include the following steps:

[0065] For each sample acoustic vibration feature vector in the first acoustic vibration feature vector set, the energy value of each time step is calculated. Time steps with energy values ​​higher than the preset high energy quantile threshold are defined as high energy regions, and time steps with energy values ​​lower than the preset low energy quantile threshold are defined as low energy regions, thus obtaining the time step division results of high energy regions and low energy regions.

[0066] Based on the time step division results, zero-mean Gaussian noise with a first preset standard deviation as the parameter is applied to each time step in the high-energy region, and zero-mean Gaussian noise with a second preset standard deviation as the parameter is applied to each time step in the low-energy region to obtain an adaptive mask noise vector. The adaptive mask noise vector is then added to the corresponding sample acoustic vibration feature vector step by step to obtain the noisy vibration feature vector, and all noisy vibration feature vectors are used as the second acoustic vibration feature vector set; wherein, the first preset standard deviation is less than the second preset standard deviation.

[0067] Specifically, each sample acoustic vibration feature vector is sequentially fed into the noise generation stage. The energy value of each time step in the sample acoustic vibration feature vector is calculated by averaging the squares of the feature values ​​of each dimension of the time step. The energy values ​​of all time steps are arranged sequentially to form a sequence of energy values ​​for each time step. Each time step is divided into regions based on a preset high energy quantile threshold and a preset low energy quantile threshold. Time steps with energy values ​​higher than the preset high energy quantile threshold are assigned to the high energy region, and time steps with energy values ​​lower than the preset low energy quantile threshold are assigned to the low energy region. A first preset standard deviation is applied to the time steps in the high energy region. Zero-mean Gaussian noise with a second preset standard deviation is applied to the time steps of the low-energy region. The first preset standard deviation is smaller than the second, forcing the model to focus more on reconstructing the weak feature components of the low-energy region during training. The noise values ​​at each time step are concatenated into an adaptive mask noise vector, corresponding one-to-one with the original vector dimensions. This adaptive mask noise vector is then added to the corresponding sample acoustic vibration feature vectors in the same time-step order, ensuring that each component in the noise vector only affects the feature components in the original sample vector at the same position, resulting in a noisy vibration feature vector. After all samples have been superimposed time-step according to the same rules, all noisy vibration feature vectors are summarized in sample number order to form a second acoustic vibration feature vector set, while maintaining the one-to-one correspondence with the first acoustic vibration feature vector set.

[0068] In one specific embodiment, the process of performing the following steps for each sample acoustic vibration feature vector in the first acoustic vibration feature vector set, calculating the energy value at each time step, defining time steps with energy values ​​higher than a preset high energy quantile threshold as high energy regions, and defining time steps with energy values ​​lower than a preset low energy quantile threshold as low energy regions, to obtain the time step division results of high energy regions and low energy regions, can specifically include the following steps:

[0069] For each time step of the sample acoustic vibration feature vector, the average value of the squared feature values ​​of each dimension of the time step is obtained to get the energy value of the time step. The energy values ​​of all time steps are arranged in order to obtain the energy value sequence of each time step.

[0070] Time steps whose energy values ​​are above a preset high energy quantile are defined as high-energy regions, and time steps whose energy values ​​are below a preset low energy quantile are defined as low-energy regions, thus obtaining the time step division results of high-energy regions and low-energy regions.

[0071] Specifically, following a sample-by-sample processing approach, a separate set of zero-mean Gaussian noise vectors with identical dimensions to the original vector is generated for each sample's acoustic vibration feature vector. Using each feature dimension of the sample's acoustic vibration feature vector as a corresponding position, Gaussian sampling is performed sequentially on each dimension, ensuring that each feature component receives a noise sample value corresponding to its position. Since zero-mean Gaussian noise itself does not introduce a fixed-direction overall shift, it is more suitable for maintaining the approximate stability of the normal sample distribution center during training sample amplification, while simultaneously applying a slight random perturbation to each feature dimension, enabling the model to learn more robust normal feature representations during training. After calculating the Gaussian sample values ​​for each vector dimension, the vectors are concatenated end-to-end according to the original dimensional arrangement of the sample's acoustic vibration feature vector to form a zero-mean Gaussian noise vector. During the concatenation process, the order of the time-domain statistical feature components, frequency-domain power spectrum feature components, and Mel-domain cepstral coefficient feature components in the original vector remains unchanged.

[0072] In one specific embodiment, such as Figure 3 As shown, the process of executing step 102 can specifically include the following steps:

[0073] 1021. Input the noisy vibration feature vectors in the second acoustic vibration feature vector set into a dual-channel attention deep autoencoder that fuses the SE channel attention module and the bidirectional LSTM temporal attention module, and perform denoising and reconstruction training based only on normal samples to obtain the acoustic vibration defect detection model.

[0074] 1022. Input the acoustic vibration feature vectors of each sample in the first acoustic vibration feature vector set into the acoustic vibration defect detection model for reconstruction to obtain the sample reconstruction output vector. Calculate the mean square reconstruction error between the sample acoustic vibration feature vector and the sample reconstruction output vector to obtain the normal sample reconstruction error set.

[0075] 1023. Based on the normal sample reconstruction error set, the moving mean and moving standard deviation of the reconstruction error are calculated by exponential weighted moving average. The sum of the moving mean and the moving standard deviation multiplied by the preset sensitivity coefficient is determined as the acoustic and vibration defect threshold.

[0076] Specifically, the noisy vibration feature vectors in the second acoustic vibration feature vector set are fed into a dual-channel attention deep autoencoder that integrates the SE channel attention module and the bidirectional LSTM temporal attention module in batches. The noisy samples are used as network inputs, and the corresponding original normal acoustic vibration feature vectors are used as reconstruction targets. The denoising and reconstruction training is performed based solely on normal samples. During training, the encoder performs feature compression on the input vector. The SE channel attention module then performs global average pooling and a two-layer fully connected gated subnetwork transformation on each feature channel, outputting the channel weights. Simultaneously, the bidirectional LSTM temporal attention module performs bidirectional sequence modeling on the intermediate layer feature vector, outputting the weights for each time step. These two weights are then multiplied channel-by-channel and time-by-time on the intermediate layer feature vector to form a channel-time-weighted feature vector. The decoder then performs layer-by-layer reconstruction based on this channel-time-weighted feature vector and updates all trainable parameters through backpropagation based on the error between the reconstructed output vector and the corresponding normal acoustic vibration feature vector. During the parameter update phase, the optimizer learning rate is dynamically adjusted based on the ratio between the current training epoch and the total training epochs, maintaining a fast convergence speed in the early stages and gradually decreasing the step size in the later stages to stably approach the optimal parameter region. After multiple training epochs, the parameters corresponding to the optimal reconstruction error index on independent normal validation samples are used as the model parameters, resulting in the acoustic vibration defect detection model.

[0077] After the model parameters are determined, independent normal calibration samples that are not involved in parameter updates are input into the already trained acoustic and vibration defect detection model. This sequentially yields the reconstructed output vectors of the calibration samples. For each calibration sample's acoustic and vibration feature vector, the difference between each vector and the corresponding reconstructed output vector is calculated dimension by dimension, squared, and then the mean is calculated over all dimensions to form the mean squared reconstruction error for that calibration sample. The errors of all calibration samples are summed to obtain the normal calibration error set. Based on this normal calibration error set, the moving mean and moving standard deviation of the reconstruction error are calculated using an exponentially weighted moving average method. The sum of the moving mean and the moving standard deviation multiplied by a preset sensitivity coefficient is used as the acoustic and vibration defect threshold. This ensures that the threshold tracks the statistical characteristics of the normal sample reconstruction error distribution, rather than relying on static calculations of fixed means and standard deviations.

[0078] In one specific embodiment, the process of performing step 1021 may specifically include the following steps:

[0079] The noisy vibration feature vector is input into the encoder of the dual-channel attention deep autoencoder that integrates the SE channel attention module and the bidirectional LSTM temporal attention module for feature compression. The SE channel attention module performs global average pooling on the intermediate layer feature vector of the encoder to obtain the statistics of each channel. After two fully connected transformations through the gated sub-network, the weights of each channel are output. At the same time, the bidirectional LSTM temporal attention module performs bidirectional sequence modeling on the intermediate layer feature vector of the encoder to output the weights of each time step. The weights of each channel and the weights of each time step are multiplied by the intermediate layer feature vector channel by channel and time step by time to obtain the channel temporal weighted feature vector.

[0080] The channel temporal weighted feature vector is input into the decoder of the dual-channel attention deep autoencoder that integrates the SE channel attention module and the bidirectional LSTM temporal attention module to reconstruct the training reconstructed output vector. Then, backpropagation is performed on all parameters of the dual-channel attention deep autoencoder that integrates the SE channel attention module and the bidirectional LSTM temporal attention module to update the parameters for the current round.

[0081] Based on the updated parameters of the current round, the dual-channel attention depth autoencoder that integrates the SE channel attention module and the bidirectional LSTM temporal attention module is iterated multiple times. The parameters corresponding to the lowest mean square reconstruction error of the validation set are taken as the optimal parameters to obtain the acoustic and vibration defect detection model.

[0082] Specifically, the noisy vibration feature vectors in the second acoustic vibration feature vector set are input in batches into a dual-channel attention depth autoencoder that fuses the SE channel attention module and the bidirectional LSTM temporal attention module, and the encoder undertakes the task of compressing and mapping the normal acoustic vibration distribution. For each batch of input samples, the encoder progressively shrinks the feature dimension along the preset fully connected layers, compressing the original standardized acoustic and vibration features to a bottleneck representation after continuous nonlinear mapping. Simultaneously, an SE channel attention module and a bidirectional LSTM temporal attention module are embedded in the middle of the encoding stage. The SE channel attention module performs global average pooling on the intermediate layer feature vector to obtain the statistics for each channel, which are then transformed by two fully connected layers of a gated subnetwork to output normalized weights for each channel. This results in higher weights for feature channels that contribute significantly to reconstruction, while reducing the weights of redundant components. The bidirectional LSTM temporal attention module performs forward and backward sequence modeling on the same intermediate layer feature vector, capturing the forward and backward dependencies between time steps and outputting normalized weights for each time step. This results in higher weights for time steps that have a stronger indicative significance for the acoustic and vibration response of defects. The two weights are then multiplied channel-by-channel and time-by-time on the intermediate layer feature vector to form a channel-time-weighted feature vector that simultaneously focuses on the defect-sensitive frequency band and key time windows.

[0083] The channel temporal weighted feature vector is input into the decoder to perform reconstruction, so that the low-dimensional representation is gradually restored to the original feature space along the level corresponding to the encoder to obtain the training reconstruction output vector. An error constraint is established between the training reconstruction output vector and the corresponding normal acoustic vibration feature vector. Then, based on the error result, all parameters of the dual-channel attention deep autoencoder that fuses the SE channel attention module and the bidirectional LSTM temporal attention module are updated by backpropagation to form the update parameters for the current round.

[0084] The model undergoes multiple iterative training rounds based on updated parameters for each training round. After each round, a validation set is used to evaluate the model's reconstruction capability, with the mean squared reconstruction error of the validation set serving as a unified metric for model performance. This approach avoids the network forming only local memories of noisy samples in the training set; instead, it requires the network to maintain a low reconstruction bias even on validation samples that were not updated, thus ensuring the final model more closely approximates the stable distribution of normal acoustic and vibration signals. As training progresses, the parameters obtained in each round are recorded along with the corresponding mean squared reconstruction error of the validation set. After all rounds are completed, the set of parameters corresponding to the lowest mean squared reconstruction error in the validation set is selected as the optimal parameters, thus solidifying the acoustic and vibration defect detection model.

[0085] In one specific embodiment, the process of inputting the channel temporally weighted feature vector into the decoder of the dual-channel attention deep autoencoder that fuses the SE channel attention module and the bidirectional LSTM temporal attention module for reconstruction, obtaining the training reconstructed output vector, and performing backpropagation to update all parameters of the dual-channel attention deep autoencoder that fuses the SE channel attention module and the bidirectional LSTM temporal attention module to obtain the updated parameters for the current round can specifically include the following steps:

[0086] The channel temporal weighted feature vector is input into the decoder of the dual-channel attention deep autoencoder that fuses the SE channel attention module and the bidirectional LSTM temporal attention module to perform layer-by-layer dimensional augmentation and reconstruction, and the training reconstruction output vector is obtained.

[0087] The loss function value is calculated based on the reconstructed output vector and the corresponding normal acoustic vibration feature vector;

[0088] Based on the ratio of the current training round to the total training rounds, the current learning rate of the Adam optimizer is dynamically updated using the cosine annealing formula to obtain the updated learning rate. Backpropagation is then performed on the loss function value using the updated learning rate. The gradients of all parameters in the dual-channel attention deep autoencoder, which integrates the SE channel attention module and the bidirectional LSTM temporal attention module, are calculated and the parameters are updated to obtain the updated parameters for the current round.

[0089] Specifically, the channel-time weighted feature vector is input into the decoder, which is mirror-symmetric to the encoder. The decoder progressively expands the feature dimension layer by layer, gradually restoring the compressed low-dimensional representation to the original standardized acoustic vibration feature space. During the layer-by-layer expansion process, batch normalization and ReLU activation can be applied sequentially between each layer except the output layer, with a dropout rate of 0.2 added at the corresponding positions. This dropout is only enabled during the training phase to suppress parameter overfitting under small sample conditions. The last layer uses linear activation to ensure that the value range of the reconstructed output vector is consistent with the value range of the standardized normal acoustic vibration feature vector, avoiding additional reconstruction bias caused by artificial truncation. The training reconstructed output vector output by the decoder is a complete approximate reconstruction of the normal acoustic vibration distribution under the current network parameters. At the same time, all trainable parameters can be entered into the joint training process using Xavier uniform initialization, keeping the activation amplitudes of each layer relatively stable in the early stages and reducing the interference of gradient vanishing or exploding on the convergence process. The parameters of the gated subnetwork of the SE channel attention module, the sequence modeling parameters of the bidirectional LSTM temporal attention module, and the encoder-decoder parameters are jointly updated end-to-end under the same loss constraint, so that the channel weights converge in the direction of enhancing the feature channels that contribute more to reconstruction and the time step weights converge in the direction of enhancing the time steps that have stronger defect indices.

[0090] A sample-by-sample error constraint is established between the trained and reconstructed output vector and the corresponding normal acoustic vibration feature vector, and the mean square error is used as the loss function value for the current round. The calculation relationship is as follows:

[0091]

[0092] in, This represents the loss function value for the current round. This represents the complete set of trainable parameters for a deep autoencoder that incorporates the SE channel attention module. This represents the number of normal training samples. Indicates the first A noisy input vector, Indicates the relationship with the first The normal acoustic vibration feature vector corresponding to each noisy input vector. Represents encoding mapping, This indicates decoding and reconstructing the mapping. Indicates the first The sample in the first The reconstructed value in the dimension. According to this loss definition, the training process learns a stable low-dimensional manifold of normal acoustic and vibration signals, rather than a mechanical fit to noise details. At the same time, the SE gated sub-network parameters and the encoder / decoder parameters are jointly updated end-to-end under the same loss constraint, so that the channel weights converge in the direction of enhancing the feature components that contribute more to reconstruction and weakening the redundant feature components.

[0093] Based on the ratio between the current training epoch and the total number of training epochs, the current learning rate of the Adam optimizer is dynamically updated using cosine annealing. Then, backpropagation and parameter adjustments are performed using the updated learning rate. The learning rate update relationship is as follows:

[0094]

[0095] in, Indicates the first The updated learning rate for each training epoch. This represents the initial learning rate. This represents the lower bound of the learning rate. This represents the total number of training epochs. According to this scheduling relationship, a larger parameter update step size can be maintained in the early stages of training to accelerate the escape from the initial suboptimal region. In the later stages of training, the learning rate is gradually and smoothly decayed to a smaller learning rate, allowing the parameters to maintain more refined convergence near the loss trough. Then, backpropagation is performed on the loss function value with the updated learning rate to calculate the gradients of all parameters in the dual-channel attention deep autoencoder that integrates the SE channel attention module and the bidirectional LSTM temporal attention module. The first-order moment of the Adam optimizer is used to estimate the decay coefficient (0.9), the second-order moment is used to estimate the decay coefficient (0.999), and the numerical stability term is used. Complete this round of parameter updates and obtain the updated parameters for the current round.

[0096] In one specific embodiment, the process of performing step 1022 may specifically include the following steps:

[0097] Each sample acoustic vibration feature vector in the first acoustic vibration feature vector set is sequentially input into the acoustic vibration defect detection model for forward propagation reconstruction to obtain the sample reconstruction output vector;

[0098] The normal sample reconstruction error is obtained by subtracting the acoustic vibration feature vector of each sample from the sample reconstruction output vector in each dimension, taking the square, and then averaging over all dimensions. All normal sample reconstruction errors are then used as the normal sample reconstruction error set.

[0099] Specifically, each sample acoustic feature vector from the first acoustic feature vector set is sequentially input into the acoustic defect detection model according to its sample number. Each normal sample completes a full forward propagation along the predetermined encoding mapping, SE channel attention weighted mapping, bidirectional LSTM temporal attention weighted mapping, and decoding reconstruction mapping. During the inference phase, the Dropout random deactivation mechanism is disabled to ensure a unique reconstruction result for each input, resulting in a sample reconstruction output vector corresponding one-to-one with the input sample. For the same sample, a dimension-by-dimensional comparison between the input and reconstructed features is performed. Following the original feature dimension order, the value of each dimension of the sample acoustic feature vector is subtracted from the value of the corresponding dimension of the sample reconstruction output vector. The differences in each dimension are then squared to eliminate the influence of mutual cancellation of positive and negative deviations and further highlight feature components with larger reconstruction deviations. The average of the squared deviations across all dimensions yields the normal sample reconstruction error for the corresponding sample. When the acoustic defect detection model can effectively represent the distribution of normal acoustic features, the reconstruction results of most normal samples after forward propagation typically maintain high consistency with the input vector, and the corresponding reconstruction errors are concentrated in a low range. After all samples have undergone the same processing, the reconstruction errors of all normal samples are summarized in order to form a normal sample reconstruction error set.

[0100] In one specific embodiment, such as Figure 4 As shown, the process of executing step 103 can specifically include the following steps:

[0101] 1031. Input the acoustic vibration feature vector of the post insulator to be tested into the acoustic vibration defect detection model for reconstruction to obtain the reconstruction output vector. Subtract the acoustic vibration feature vector to be tested from the reconstruction output vector of the post insulator in each dimension, take the square, and calculate the mean over all dimensions to obtain the reconstruction error.

[0102] 1032. Based on the preset attenuation factor, the acoustic vibration defect threshold is updated online using the exponential weighted moving average formula to obtain the current acoustic vibration defect threshold.

[0103] 1033. Compare the reconstruction error to be measured with the current acoustic and vibration defect threshold. If the reconstruction error to be measured is greater than the current acoustic and vibration defect threshold, the output defect detection result is a defect alarm. If the reconstruction error to be measured is less than or equal to the current acoustic and vibration defect threshold, the output defect detection result is a normal state.

[0104] Specifically, the raw acoustic and vibration signals of the insulator under test are simultaneously acquired by acoustic and vibration sensors. Short-time Fourier transform and wavelet packet transform are performed respectively, and time alignment is completed. These signals are then converted into acoustic and vibration feature vectors according to the same processing standards used in the training phase, and then fed into the inference link. The test samples are extracted using the aforementioned time-domain statistical features, frequency-domain power spectrum features, and Mel-domain cepstral coefficient features to form test feature vectors. The standardized statistical parameters saved in the training phase are used to perform standardized processing on the test feature vectors, ensuring that the test input and training input are within the same data scale range. The dual-channel attention depth autoencoder, which integrates the SE channel attention module and the bidirectional LSTM temporal attention module in the inference phase, is switched to deterministic operating mode. This means that the Dropout random deactivation mechanism used in the training phase is disabled, ensuring that the same test input corresponds to only one reconstruction result under the same model parameters, avoiding unnecessary fluctuations in the inference results due to random discarding. The standardized acoustic vibration feature vector to be tested is input into the acoustic vibration defect detection model. The acoustic vibration feature vector to be tested is then processed sequentially through encoding compression, SE channel weight redistribution, bidirectional LSTM temporal weight redistribution, and decoding reconstruction. The output is the reconstructed output vector to be tested with the same dimension as the original input.

[0105] The acoustic vibration feature vector to be measured is compared with the reconstructed output vector to be measured one by one in the same dimension order. The difference and square processing are performed on the reconstruction deviation of each dimension, and the mean of the squared deviations of all dimensions is calculated, thereby forming the reconstruction error of a single sample.

[0106] After each reconstruction error to be measured is obtained, the acoustic and vibration defect threshold is updated online using an exponentially weighted moving average formula based on a preset attenuation factor: the acoustic and vibration defect threshold of the previous moment and the current reconstruction error to be measured are weighted and averaged according to the attenuation factor to obtain the current acoustic and vibration defect threshold; when factors such as changes in ambient temperature and load cause the normal sample reconstruction error distribution to drift, the current acoustic and vibration defect threshold can adaptively track the drift through error statistics within the sliding window, avoiding the problem of the false alarm rate increasing due to long-term operation of the static threshold.

[0107] The reconstruction error to be tested is compared with the current acoustic and vibration defect threshold: Since the acoustic and vibration defect detection model is trained only based on normal samples, when the material state of the post insulator to be tested is close to the distribution of normal samples, the overall deviation between the reconstruction output vector and the acoustic and vibration feature vector to be tested is small, the reconstruction error to be tested is less than or equal to the current acoustic and vibration defect threshold, and the output is normal; when the post insulator to be tested has cracks, voids, material deterioration, or other factors that cause abnormal changes in acoustic and vibration response, the deviation of the acoustic and vibration feature vector to be tested from the normal distribution increases, the consistency of model reconstruction decreases, the reconstruction error to be tested exceeds the current acoustic and vibration defect threshold, and a defect alarm is output.

[0108] The above describes the deep learning-based acoustic vibration defect detection method for post insulator materials in the embodiments of the present invention. The following describes the deep learning-based acoustic vibration defect detection system for post insulator materials in the embodiments of the present invention. Please refer to [link / reference]. Figure 5 One embodiment of the deep learning-based acoustic vibration defect detection system for post insulator materials in this invention includes:

[0109] The construction module 501 is used to perform time-series alignment processing on the acoustic signal and vibration signal of a normal post insulator to construct a first acoustic-vibration feature vector set, and to apply adaptive mask noise perturbation based on the energy distribution of each sample acoustic-vibration feature vector in the first acoustic-vibration feature vector set to obtain a second acoustic-vibration feature vector set.

[0110] Training module 502 is used to train the dual-channel attention depth autoencoder that fuses the SE channel attention module and the bidirectional LSTM temporal attention module based on normal samples only, according to the second acoustic vibration feature vector set, to obtain the acoustic vibration defect detection model, and to calculate the reconstruction error of the first acoustic vibration feature vector set on the acoustic vibration defect detection model by exponential weighted moving average to determine the acoustic vibration defect threshold.

[0111] The output module 503 is used to input the acoustic vibration feature vector of the post insulator to be tested into the acoustic vibration defect detection model for reconstruction, obtain the reconstruction error, compare the obtained reconstruction error with the acoustic vibration defect threshold, and output the defect detection result.

[0112] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0113] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0114] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for detecting acoustic and vibration defects in post insulator materials based on deep learning, characterized in that, include: The acoustic and vibration signals of a normal post insulator are time-aligned to construct a first acoustic-vibration feature vector set. An adaptive mask noise perturbation is then applied based on the energy distribution of each sample acoustic-vibration feature vector in the first acoustic-vibration feature vector set to obtain a second acoustic-vibration feature vector set. Based on the second acoustic vibration feature vector set, the dual-channel attention depth autoencoder that fuses the SE channel attention module and the bidirectional LSTM temporal attention module is trained on denoising and reconstruction based only on normal samples to obtain an acoustic vibration defect detection model. The reconstruction error of the first acoustic vibration feature vector set on the acoustic vibration defect detection model is calculated by exponential weighted moving average to determine the acoustic vibration defect threshold. The acoustic and vibration feature vector of the post insulator to be tested is input into the acoustic and vibration defect detection model for reconstruction, and the reconstruction error is obtained. The obtained reconstruction error is compared with the acoustic and vibration defect threshold, and the defect detection result is output.

2. The method for detecting acoustic and vibration defects in post insulator materials based on deep learning according to claim 1, characterized in that, The process involves time-aligning the acoustic and vibration signals of a normal post insulator to construct a first acoustic-vibration feature vector set. An adaptive masking noise perturbation is then applied based on the energy distribution of each sample acoustic-vibration feature vector in the first set to obtain a second acoustic-vibration feature vector set, including: Simultaneously, acoustic and vibration signals of normal support insulators are acquired, short-time Fourier transform is performed on the acoustic signals, wavelet packet transform is performed on the vibration signals, and the transformation results of the two are aligned on the time axis to obtain time-aligned acoustic and vibration transformation features. The time-aligned acoustic transformation features and vibration transformation features are extracted using time-domain statistical features, frequency-domain power spectrum features, and Mel-domain cepstral coefficient features, respectively, to obtain time-domain statistical feature components, frequency-domain power spectrum feature components, and Mel-domain cepstral coefficient feature components; the time-domain statistical feature components, the frequency-domain power spectrum feature components, and the Mel-domain cepstral coefficient feature components are then vector-concatenated and standardized to obtain the first acoustic-vibration feature vector set; An adaptive masking noise perturbation is applied to each sample acoustic vibration feature vector in the first acoustic vibration feature vector set based on the energy distribution at each time step to obtain the second acoustic vibration feature vector set.

3. The method for detecting acoustic and vibration defects in post insulator materials based on deep learning according to claim 2, characterized in that, The process involves applying an adaptive masking noise perturbation to each sample acoustic feature vector in the first acoustic feature vector set based on the energy distribution at each time step, resulting in a second acoustic feature vector set, including: For each sample acoustic vibration feature vector in the first acoustic vibration feature vector set, the energy value of each time step is calculated. Time steps with energy values ​​higher than a preset high energy quantile threshold are defined as high energy regions, and time steps with energy values ​​lower than a preset low energy quantile threshold are defined as low energy regions, thus obtaining the time step division results of high energy regions and low energy regions. Based on the time step division results, zero-mean Gaussian noise with a first preset standard deviation as a parameter is applied to each time step in the high-energy region, and zero-mean Gaussian noise with a second preset standard deviation as a parameter is applied to each time step in the low-energy region to obtain an adaptive mask noise vector. The adaptive mask noise vector is then added to the corresponding sample acoustic vibration feature vector step by step to obtain a noisy vibration feature vector, and all the noisy vibration feature vectors are used as the second acoustic vibration feature vector set; wherein, the first preset standard deviation is less than the second preset standard deviation.

4. The method for detecting acoustic and vibration defects in post insulator materials based on deep learning according to claim 3, characterized in that, For each sample acoustic vibration feature vector in the first acoustic vibration feature vector set, the energy value of each time step is calculated. Time steps with energy values ​​higher than a preset high energy quantile threshold are defined as high-energy regions, and time steps with energy values ​​lower than a preset low energy quantile threshold are defined as low-energy regions, resulting in the time step division results of high-energy and low-energy regions, including: For each time step of the sample acoustic vibration feature vector, the average value of the squared feature values ​​of each dimension of the time step is obtained to get the energy value of the time step. The energy values ​​of all time steps are arranged in order to obtain the energy value sequence of each time step. The time steps whose energy values ​​are above a preset high energy quantile are defined as high energy regions, and the time steps whose energy values ​​are below a preset low energy quantile are defined as low energy regions, thus obtaining the time step division results of high energy regions and low energy regions.

5. The method for detecting acoustic and vibration defects in post insulator materials based on deep learning according to claim 1, characterized in that, The step of training a dual-channel attention depth autoencoder based solely on normal samples using the second acoustic feature vector set to perform denoising and reconstruction training on the fused SE channel attention module and bidirectional LSTM temporal attention module, thereby obtaining an acoustic defect detection model, and then calculating the acoustic defect threshold by performing an exponentially weighted moving average calculation on the reconstruction error of the first acoustic feature vector set on the acoustic defect detection model, includes: The noisy vibration feature vectors in the second acoustic vibration feature vector set are input into a dual-channel attention deep autoencoder that fuses the SE channel attention module and the bidirectional LSTM temporal attention module for denoising and reconstruction training based only on normal samples, thus obtaining an acoustic vibration defect detection model. The acoustic vibration feature vectors of each sample in the first acoustic vibration feature vector set are input into the acoustic vibration defect detection model for reconstruction to obtain the sample reconstruction output vector. The mean square reconstruction error between the sample acoustic vibration feature vector and the sample reconstruction output vector is calculated to obtain the normal sample reconstruction error set. The sliding mean and sliding standard deviation of the reconstruction error are calculated by exponential weighted moving average based on the normal sample reconstruction error set. The sum of the sliding mean and the sliding standard deviation multiplied by a preset sensitivity coefficient is determined as the acoustic and vibration defect threshold.

6. The method for detecting acoustic and vibration defects in post insulator materials based on deep learning according to claim 5, characterized in that, The process involves inputting the noisy vibration feature vectors from the second acoustic vibration feature vector set into a dual-channel attention deep autoencoder that fuses the SE channel attention module and the bidirectional LSTM temporal attention module for denoising and reconstruction training based solely on normal samples, thereby obtaining an acoustic vibration defect detection model. This includes: The noisy vibration feature vector is input into the encoder of a dual-channel attention deep autoencoder that integrates the SE channel attention module and the bidirectional LSTM temporal attention module for feature compression. The SE channel attention module performs global average pooling on the encoder's intermediate layer feature vector to obtain the statistics of each channel. After two fully connected transformations through a gated subnetwork, the weights of each channel are output. At the same time, the bidirectional LSTM temporal attention module performs bidirectional sequence modeling on the encoder's intermediate layer feature vector to output the weights of each time step. The weights of each channel and the weights of each time step are multiplied by the intermediate layer feature vector channel by channel and time step by time to obtain the channel temporal weighted feature vector. The channel temporal weighted feature vector is input into the decoder of the dual-channel attention deep autoencoder that fuses the SE channel attention module and the bidirectional LSTM temporal attention module to reconstruct the training reconstructed output vector. Then, backpropagation is performed on all parameters of the dual-channel attention deep autoencoder that fuses the SE channel attention module and the bidirectional LSTM temporal attention module to update the parameters for the current round. Based on the current round update parameters, the dual-channel attention depth autoencoder that integrates the SE channel attention module and the bidirectional LSTM temporal attention module is iterated multiple times. The parameters corresponding to the lowest mean square reconstruction error of the validation set are taken as the optimal parameters to obtain the acoustic and vibration defect detection model.

7. The method for detecting acoustic and vibration defects in post insulator materials based on deep learning according to claim 6, characterized in that, The process involves inputting the channel temporal weighted feature vector into the decoder of a dual-channel attention deep autoencoder that fuses the SE channel attention module and the bidirectional LSTM temporal attention module for reconstruction, obtaining the trained reconstructed output vector, and then performing backpropagation to update all parameters of the dual-channel attention deep autoencoder to obtain the updated parameters for the current round, including: The channel temporal weighted feature vector is input into the decoder of the dual-channel attention deep autoencoder that fuses the SE channel attention module and the bidirectional LSTM temporal attention module to perform layer-by-layer dimensional augmentation and reconstruction, and the training reconstruction output vector is obtained. The loss function value is calculated based on the reconstructed output vector from the training and the corresponding normal acoustic vibration feature vector. Based on the ratio of the current training round to the total training rounds, the current learning rate of the Adam optimizer is dynamically updated using the cosine annealing formula to obtain the updated learning rate. Backpropagation is then performed on the loss function value using the updated learning rate to calculate the gradients of all parameters in the dual-channel attention deep autoencoder that integrates the SE channel attention module and the bidirectional LSTM temporal attention module, and the parameter updates are completed to obtain the updated parameters for the current round.

8. The method for detecting acoustic and vibration defects in post insulator materials based on deep learning according to claim 7, characterized in that, The process involves inputting the acoustic vibration feature vectors of each sample in the first acoustic vibration feature vector set into the acoustic vibration defect detection model for reconstruction, obtaining a sample reconstruction output vector, and calculating the mean square reconstruction error between the sample acoustic vibration feature vector and the sample reconstruction output vector to obtain a normal sample reconstruction error set, including: Each sample acoustic vibration feature vector in the first acoustic vibration feature vector set is sequentially input into the acoustic vibration defect detection model for forward propagation reconstruction to obtain the sample reconstruction output vector; The normal sample reconstruction error is obtained by taking the square of the difference between each sample acoustic vibration feature vector and the sample reconstruction output vector in each dimension and then averaging the difference over all dimensions. All the normal sample reconstruction errors are then used as the normal sample reconstruction error set.

9. The method for detecting acoustic and vibration defects in post insulator materials based on deep learning according to claim 1, characterized in that, The process involves inputting the acoustic vibration feature vector of the insulator under test into the acoustic vibration defect detection model for reconstruction, obtaining the reconstruction error, comparing the obtained reconstruction error with the acoustic vibration defect threshold, and outputting the defect detection result, including: The acoustic vibration feature vector of the post insulator to be tested is input into the acoustic vibration defect detection model for reconstruction, and the reconstruction output vector to be tested is obtained. The difference between the acoustic vibration feature vector to be tested and the reconstruction output vector to be tested is taken in each dimension, the square is taken, and the mean is calculated over all dimensions to obtain the reconstruction error to be tested. Based on a preset attenuation factor, the acoustic and vibration defect threshold is updated online using an exponentially weighted moving average formula to obtain the current acoustic and vibration defect threshold. The reconstruction error to be measured is compared with the current acoustic and vibration defect threshold. If the reconstruction error to be measured is greater than the current acoustic and vibration defect threshold, the defect detection result is output as a defect alarm. If the reconstruction error to be measured is less than or equal to the current acoustic and vibration defect threshold, the defect detection result is output as a normal state.

10. A deep learning-based acoustic vibration defect detection system for post insulator materials, characterized in that, The method for performing the deep learning-based acoustic vibration defect detection method for post insulator materials as described in any one of claims 1-9 includes: The construction module is used to perform time-series alignment processing on the acoustic signal and vibration signal of a normal post insulator to construct a first acoustic-vibration feature vector set, and to apply adaptive mask noise perturbation based on the energy distribution of each sample acoustic-vibration feature vector in the first acoustic-vibration feature vector set to obtain a second acoustic-vibration feature vector set. The training module is used to train the dual-channel attention depth autoencoder that fuses the SE channel attention module and the bidirectional LSTM temporal attention module based on the second acoustic vibration feature vector set, and to perform denoising and reconstruction training based only on normal samples to obtain an acoustic vibration defect detection model. The reconstruction error of the first acoustic vibration feature vector set on the acoustic vibration defect detection model is calculated by exponential weighted moving average to determine the acoustic vibration defect threshold. The output module is used to input the acoustic vibration feature vector of the post insulator to be tested into the acoustic vibration defect detection model for reconstruction, obtain the reconstruction error to be tested, compare the obtained reconstruction error to be tested with the acoustic vibration defect threshold, and output the defect detection result.