Cross-domain transformer abnormal sound detection method based on wavelet style enhanced prototype network

By combining wavelet style enhancement and prototype networks, a highly adaptable method for detecting abnormal sounds in transformers is constructed, which solves the problem of decreased detection accuracy between the source and target domains and achieves efficient detection and improved adaptability in the target domain.

CN120932674APending Publication Date: 2025-11-11SOUTH CHINA NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510959899.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-11
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing technologies for detecting abnormal sounds in transformers suffer from domain offset between the source and target domains, leading to a decrease in the model's detection accuracy in new environments. Furthermore, traditional methods are complex to train or difficult to control the quality of generated samples, making it difficult to meet the requirements for rapid deployment and real-time response.

Method used

By employing a wavelet style enhancement prototype network approach, style enhancement training samples with statistical characteristics similar to the target domain are constructed. A prototype guidance mechanism is introduced to fine-tune the pre-trained model, including a combination of wavelet decomposition, style transfer, and prototype network, to generate more representative training samples and improve the model's adaptability and detection performance in the target domain.

Benefits of technology

With limited samples in the target domain, the model's cross-domain adaptability and anomaly detection performance were improved, achieving higher detection accuracy and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120932674A_ABST
    Figure CN120932674A_ABST
Patent Text Reader

Abstract

The invention discloses a cross-domain transformer abnormal sound detection method based on a wavelet style enhanced prototype network, and belongs to the technical field of abnormal sound detection. The method comprises the following steps: acquiring source domain audio and a small number of target domain audio samples; performing wavelet decomposition on the source domain audio, and extracting low-frequency and high-frequency components of the source domain audio; performing style enhancement on the source domain low-frequency component by using the low-frequency statistical characteristics of the target domain sample, combining with the original high-frequency component, and reconstructing an enhanced sample through inverse wavelet transform; constructing a prototype network, taking a target domain sample as a support set, taking an enhanced sample as a query set, and carrying out prototype matching training in the feature space; prototype loss and category classification loss are jointly optimized, and the cross-domain adaptability of the model is improved; in the detection stage, a target domain test sample is input into the adjusted model, feature extraction and classification judgment are carried out, and abnormal sound recognition is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method for detecting abnormal sounds in cross-domain transformers, and more particularly to a method for detecting abnormal sounds in cross-domain transformers based on a wavelet style enhancement prototype network, belonging to the technical field of methods for detecting abnormal sounds in cross-domain transformers. Background Technology

[0002] With the continuous improvement of transformer capacity and performance, the safe and stable operation of transformers has become an important guarantee for power system maintenance. Transformer failures may lead to power outages or even serious safety accidents such as fires. Therefore, the demand for real-time monitoring and fault early warning technology for transformer operation status is increasing.

[0003] Sound signal-based anomaly detection technology has become a research hotspot in transformer condition monitoring due to its advantages such as non-contact operation, high sensitivity, and low cost. By collecting the sound generated during equipment operation and combining signal processing and machine learning algorithms, abnormal conditions can be effectively identified and faults can be prevented. Compared with traditional vibration and temperature monitoring, sound detection technology is not affected by the surface structure or temperature changes of the equipment, making it more suitable for online monitoring of enclosed equipment such as transformers.

[0004] However, the operating environment of transformers is complex and variable. Factors such as environmental noise, equipment model, and recording conditions cause significant differences in the collected sound signals under different scenarios. This difference in statistical characteristics between the source domain and the target domain is called domain shift, which seriously affects the detection performance of the model in new environments. Traditional models are mainly based on source domain data during training and cannot directly adapt to changes in the target domain, resulting in decreased detection accuracy and insufficient robustness.

[0005] In practical applications, due to the wide distribution of transformers, it is both time-consuming and laborious to re-collect a large amount of target domain data and retrain the model, which is difficult to meet the requirements of rapid deployment and real-time response. Therefore, how to effectively adjust and optimize the source domain model using limited target domain samples so that it can maintain good anomaly detection performance in the new environment is a technical problem that urgently needs to be solved.

[0006] In existing technologies for detecting abnormal sounds in transformers, some methods improve the model's generalization ability in different environments through style transfer techniques. One method employs a generative adversarial network (GAN), inputting labeled fault sound samples from the source domain and samples from the target domain into the generator network to generate auxiliary training samples with environmental features of the target domain, thus improving the model's adaptability to the target domain. Simultaneously, a feature selection module is used to extract a more generalizable sub-feature space, and the model is trained using a dual-branch classifier structure to enhance cross-domain classification performance. While this method possesses strong generalization ability, the adversarial training process is complex, the model training is unstable, and the quality of generated samples is difficult to control, affecting the anomaly detection effect.

[0007] Another approach for cross-domain infrared target detection proposes an image enhancement and domain alignment framework based on frequency domain style transfer. This method transforms the image into the frequency domain through Fourier transform, replaces the low-frequency amplitude of the source domain with the low-frequency information of the target domain to achieve style transfer, and completes domain adaptation through a multi-layer alignment module. This technique performs well in the image domain, but directly using the global frequency domain replacement method is difficult to adapt to the time-varying characteristics and local anomalous components of audio signals, which can easily lead to the loss or weakening of anomalous signal information and reduce detection performance. Summary of the Invention

[0008] The main objective of this invention is to provide a cross-domain transformer anomaly sound detection method based on a wavelet style-enhanced prototype network, suitable for practical applications where the target environment sample is limited and the source and target domains are offset. This method constructs style-enhanced training samples with statistical characteristics similar to the target domain and introduces a prototype guidance mechanism to fine-tune the pre-trained model, ultimately achieving accurate anomaly detection under cross-domain conditions. This invention includes the following steps:

[0009] S1: Acquire audio samples and perform time-frequency graph conversion;

[0010] S1.1 Obtain audio samples from the source and target domains;

[0011] Audio data from transformer equipment under various operating conditions was collected from historical operational data to construct a source domain sample set. These conditions included normal conditions and several typical abnormal conditions (such as loose windings, abnormal cooling, and partial discharge). Each type of sample must be accurately labeled. Simultaneously, a small number of target domain abnormal samples were collected. Although limited in number, these samples should cover as many common equipment fault types as possible to support subsequent style modeling. Both types of audio samples were formatted consistently, preferably using a 16kHz sampling rate and 16-bit mono WAV format for easier subsequent unified processing.

[0012] S1.2 performs preprocessing and time-frequency graph conversion on the audio samples;

[0013] After acquisition, normalization, silence clipping, and time alignment operations are performed on the source and target domain audio samples, respectively, to unify the sample duration to a fixed length (e.g., 10 seconds) to improve training stability and comparability. Next, the processed audio signal is converted into a two-dimensional time-spectrum using either a short-time Fourier transform (STFT) or a log-Mel spectrum method. The converted time-spectrum serves as the input basis for subsequent wavelet decomposition and style enhancement.

[0014] S2: Perform wavelet decomposition on the time-spectrum graph to extract low-frequency and high-frequency components;

[0015] S2.1 Select a suitable wavelet basis and perform two-dimensional wavelet decomposition;

[0016] In this step, to separate the global background from the local details in the time-spectrum graph, a two-dimensional discrete wavelet transform is performed on the time-spectrum graphs of the source domain and the target domain, respectively.

[0017] (Discrete Wavelet Transform, DWT). Wavelet decomposition preferably uses the Daubechies wavelet basis function, which has good compact support and frequency resolution, making it suitable for extracting non-stationary background and structural information from audio signals. Through DWT operations, the original spectrum X can be decomposed to obtain the low-frequency component X. low and high-frequency component X high :

[0018] X low ,X high =DWT(X);

[0019] Among them, X low It mainly includes global features such as background noise and environmental information in the audio, while X high It records relatively detailed structural information, such as the vibration and friction characteristics of mechanical components. Since changes in background noise are usually the main source of domain offset, this invention primarily targets X. low Make style adjustments to reduce the background differences between the source and target domains.

[0020] S2.2 Obtain the decomposition results to support subsequent style transfer;

[0021] After decomposition, the high-frequency components of the source domain samples are retained. Used for restoring structural information of subsequent samples; at the same time, low-frequency components As the target of style transfer processing. Low-frequency components of the target domain samples. Simultaneously extracted and used as a style reference. By modeling the differences in low-frequency components between the two domains, it is possible to achieve background style matching and enhancement while keeping the device structure unchanged.

[0022] S3: Style enhancement processing is performed on the low-frequency components of the source domain samples;

[0023] S3.1 Extract style statistics for the target domain;

[0024] To construct training samples that more closely resemble the characteristics of the target domain, this step first extracts the style statistics of the low-frequency components of the target domain samples. (Target domain low-frequency components) The "style" is determined by its mean and standard deviation Express.

[0025] The mean and standard deviation are calculated as follows, for a given time spectrum. The formulas for calculating its mean μ(X) and standard deviation σ(X) are as follows:

[0026]

[0027] Where T and F represent the time and frequency dimensions of the feature map, respectively, and X t,f The value of the time spectrum X is represented in the time t and frequency f dimensions, where ε is a very small constant used to ensure numerical stability.

[0028] S3.2 performs style transfer on the low-frequency components of the source domain;

[0029] After obtaining the style statistics of the target domain, an adaptive instance normalization method is used to enhance the style of the low-frequency components of the source domain samples. Specifically, the low-frequency components of the source domain are first... Perform mean and variance normalization operations to obtain the content statistics of the source domain low-frequency components. The content is composed of normalization terms. The distribution of low-frequency components in the source domain is then adjusted using mean normalization and standard deviation scaling to align its style features with the target domain. Given the low-frequency components of the source domain... and low-frequency components in the target domain The adaptive instance normalization transform is defined as follows:

[0030]

[0031] Through the above operations, while retaining the low-frequency structural information related to the device's operating status in the source domain signal, style statistics from the target domain are injected. This ensures that the adjusted low-frequency components maintain the stable characteristics of the source domain device operation and approximate the background noise of the target domain in terms of spectral distribution. This enables the enhanced samples to possess the statistical characteristics of the target environment, thereby improving their discrimination ability in the target domain.

[0032] S4: Perform inverse wavelet transform on the adjusted low-frequency components and the source domain high-frequency components to reconstruct the sample.

[0033] After completing the style enhancement, the enhanced low-frequency components will be adjusted to match the low-frequency components. Compared with the original high-frequency components The data is then combined and reconstructed into a complete two-dimensional time-frequency spectrum using the Inverse Discrete Wavelet Transform (IDWT). The calculation formula is as follows:

[0034]

[0035] The reconstruction operation uses the same wavelet basis functions as the decomposition step to ensure the reversibility and fidelity of the transformation process. The reconstruction result is denoted as the enhanced sample X.aug It maintains the same spectral structure as the source domain sample, but its background statistical characteristics are close to the target domain distribution, giving it a stronger cross-domain adaptability.

[0036] S5: Anomaly detection model based on source domain sample pre-training.

[0037] S5.1 Construct an anomaly detection model.

[0038] Before constructing the wavelet-style enhanced samples, this step first builds a preliminary classification model for fault type identification based on normal and abnormal audio samples collected in the source domain. This model aims to extract key feature representations from the audio and learn the differences in feature distributions corresponding to different fault states, thereby achieving discrimination among multiple fault states. The model structure can employ convolutional neural networks, temporal modeling networks, or combinations thereof, such as ResNet, MobileNetV2, and WaveNet. The model input is the two-dimensional time-spectrum obtained after the aforementioned processing, and the output is the corresponding fault category prediction result, supporting multi-class discrimination tasks.

[0039] S5.2 Use source domain samples for model training.

[0040] After the model structure is determined, supervised training is performed using labeled normal and abnormal audio samples from the source domain. During training, the fault category corresponding to the sample is used as the target label, and optimization is performed using the cross-entropy loss function or other classification loss functions to guide the model in learning the discrimination boundaries between different categories. This stage of training helps the model grasp the feature differences between normal and various abnormal states in the audio data. This training process is completed without the participation of target domain information, providing a foundational model for subsequent fine-tuning using target domain samples.

[0041] S6: Model fine-tuning and target domain feature alignment based on prototype network.

[0042] S6.1 Construct the support set and query set and organize the training task.

[0043] In this step, to achieve target domain adaptation, the target domain samples are used as the support set S, and the augmented samples are used as the query set Q, employing an N-way K-shot task construction approach. For example, in a 5-way 3-shot setting, each task contains 5 categories, and each category has 3 samples in the support set. Each meta-task T = {S, Q} consists of a support set S and a query set Q. N represents the number of categories in each task (e.g., the number of outlier categories selected in each task), and K represents the number of samples provided for each category in the support set. The support set S is defined as follows:

[0044]

[0045] in, For the target domain sample, This indicates the anomaly category label.

[0046] The corresponding query set consists of style-enhanced samples, represented as:

[0047]

[0048] in, To enhance the sample, y q This represents the anomaly category label, where M is the number of samples in the query set.

[0049] S6.2 Constructing Class Prototype Vectors Based on Support Set Samples

[0050] Input each class of support set samples into the feature extractor f θ Next, the prototype vector for each category is calculated. The prototype representation for category c is:

[0051]

[0052] in This represents the set of supporting samples for category c. For the feature extraction network, the temporal spectrogram is mapped to a one-dimensional embedding vector with dimension D. This is the prototype for class c. The prototype vector represents the center position of this class in the embedding space and is used for subsequent distance metrics and classification prediction.

[0053] S6.3 utilizes prototype metrics to classify query samples and fine-tune the model.

[0054] For each query sample First, feature extraction is performed, and then the Euclidean distance between the feature vector and the prototype of each category is calculated:

[0055]

[0056] During training, the model aims to minimize the distance between the query sample and the correct class prototype, while avoiding excessive proximity to prototypes of other classes. Therefore, the model optimizes the metric space by minimizing the negative log-likelihood loss.

[0057]

[0058] This loss function forces query samples to move closer to the correct class prototype while moving away from the incorrect class prototype, thus bringing the target domain samples and augmented samples closer in the feature space, thereby improving the model's discriminative ability in the target domain. The model is fine-tuned using negative log-likelihood as the optimization objective.

[0059]

[0060] The final model optimizes L metric This ensures that the query sample is close to the correct target domain prototype in the embedding space, thereby completing the target domain feature alignment.

[0061] S6.4 introduces a classification loss to improve feature discrimination capability.

[0062] To further improve the model's discriminative performance in the target domain, this step introduces an auxiliary classification loss function in addition to the prototype metric loss to enhance the model's feature discrimination ability. The classification loss constrains the query samples to form a more compact and clearly defined class distribution in the embedding space, avoiding the boundary ambiguity caused by relying solely on prototype distance. This classification loss can be trained with the query samples using a conventional cross-entropy loss function to guide the model in outputting correct class predictions.

[0063] Preferably, to further enhance the angular discriminativeness between embedded features, an angular spacing optimization strategy, such as ArcFace classification loss, can be introduced.

[0064] (AdditiveAngularMarginLoss). ArcFace effectively improves the clarity of inter-class boundaries by introducing a fixed margin in the angular space between features and weights, making similar samples more concentrated and dissimilar samples more spaced, thereby improving the model's classification stability and robustness in the target domain.

[0065] Finally, a weighted joint loss function is used to incorporate both the metric loss and the classification loss into the optimization objective, as expressed below:

[0066] L=λL metric +(1-λ)L class ;

[0067] Among them, L class λ represents the category classification loss on the query set, and λ is a hyperparameter used to balance the contributions of the two losses.

[0068] S7: Input the target domain test samples into the adjusted model for anomaly detection.

[0069] S7.1 performs preprocessing and feature extraction on the test samples in the target domain.

[0070] After the model has been fine-tuned, it enters the actual testing phase. This step first performs the same preprocessing operations as in the training phase on the test samples to be classified in the target domain, including audio normalization, silence removal, and short-time Fourier transform, to generate the corresponding two-dimensional time-spectrum. Subsequently, the time-spectrum is input into the fine-tuned fault classification model, and its deep feature representation is extracted through the model's forward inference process.

[0071] S7.2 Use the trained classification model to determine the fault category.

[0072] During the training phase, the model parameters were optimized to adapt to the feature distribution of the target domain through the synergistic effect of wavelet style enhancement and the prototype network. Specifically, wavelet style enhancement adjusts the low-frequency style of the source domain samples to better reflect the statistical characteristics of the target domain, effectively expanding the representative training samples under small sample conditions and improving the model's generalization ability to the target domain. Simultaneously, the prototype network constructs class prototypes using a small number of samples from the target domain and promotes the alignment of the feature space between the augmented data samples and the target domain samples through metric learning, thereby enhancing cross-domain adaptability.

[0073] Therefore, the target domain samples in the testing phase do not require additional adjustment and can be directly input into the trained model for forward inference. Based on the extracted features and the learned fault category discriminator, the model performs classification and outputs the corresponding fault category label.

[0074] Technical effects of the invention:

[0075] This invention proposes a cross-domain transformer anomalous sound detection method based on a combination of wavelet style enhancement and a prototype network, aiming to improve the detection performance of a source-domain trained model in a target-domain environment. This method targets the background noise and environmental information contained in the low-frequency components of the audio signal. It employs wavelet decomposition to perform style transformation on the low-frequency portion of the source-domain samples, making them closer to the noise characteristics of the target domain. This generates more representative training samples even with a limited number of target-domain samples, enhancing the model's generalization ability. Simultaneously, by combining a prototype network, a prototype-like model is constructed using a small number of target-domain samples, reducing the distance between the target-domain samples and the enhanced samples in the feature space. Based on this, the parameters of the source-domain trained model are fine-tuned, achieving effective adaptation of the model to the target domain and improving its adaptability and anomaly detection performance under small sample conditions. Attached Figure Description

[0076] Figure 1 This is an overall flowchart of the cross-domain transformer abnormal sound detection method based on wavelet style enhancement prototype network according to an embodiment of the present invention;

[0077] Figure 2 This is a schematic diagram of the wavelet style enhancement module in the method, used to illustrate the style enhancement process of the low-frequency components of the source domain samples;

[0078] Figure 3 This is a schematic diagram of the prototype network fine-tuning and metric classification module in the method, used to illustrate the feature alignment process between target domain samples and enhanced samples;

[0079] Figure 4 This is a schematic diagram of the target domain sample discrimination process in the model testing phase of the method of the present invention. Detailed Implementation

[0080] To enable those skilled in the art to understand the technical solution of the present invention more clearly, the present invention will be further described in detail below with reference to the embodiments and accompanying drawings, but the embodiments of the present invention are not limited thereto.

[0081] To more clearly illustrate the technical solution of this invention, the following detailed description, in conjunction with specific embodiments, provides a method for cross-domain abnormal sound detection based on a wavelet style enhancement prototype network. This embodiment is applied to a cross-domain transformer fault sound recognition task. The source domain contains multiple types of labeled normal and abnormal state audio samples, such as "normal," "loose winding," "abnormal cooling," and "partial discharge," with a sufficient number of samples in each category. The target domain, however, contains multiple types of labeled normal and abnormal state audio samples of the same category as the source domain, but each category contains only a few samples. The sample quantity is limited and there is background environmental bias, making it unsuitable for direct support of traditional supervised training. This invention improves the model's fault recognition accuracy and cross-domain generalization ability in the target domain through wavelet style enhancement and prototype guidance mechanisms. Figure 1 As shown, the method of the present invention includes seven stages: audio sample preprocessing, wavelet decomposition, low-frequency style enhancement, sample reconstruction, anomaly detection model pre-training, prototype network guided fine-tuning, and target domain sample detection. The overall process starts from the acquisition of source domain samples and gradually completes the construction of cross-domain detection capabilities.

[0082] S1. Acquire audio samples and perform time-frequency graph conversion;

[0083] S1.1 Source domain audio samples are collected from the transformer's operating history, covering multiple categories such as "normal," "winding loose," "abnormal cooling," and "partial discharge." Each sample is accurately labeled manually or by an automated diagnostic system. Simultaneously, several target domain audio samples (significantly fewer than the source domain) are collected, primarily representing operating sounds under different conditions in the target environment, uniformly formatted as 16kHz sampling rate, 16-bit mono WAV. S1.2 All collected audio samples are normalized, silence segments are trimmed, and their length is standardized (e.g., 10 seconds), then converted into a log-Mel spectrum.

[0084] X = log(M·||STFT(x)|| 2 );

[0085] in, It is a Log-Mel spectrum. This represents the Mel filter bank. Given a one-dimensional audio signal of length L, let STFT represent the short-time Fourier transform, which converts the audio signal into a spectrum of dimension T×B, where T represents the number of time frames, B represents the number of frequency bands, and M represents the filtering process of the STFT spectrum through the Mel filter bank M, where F is the number of Mel filters.

[0086] S2. Perform wavelet decomposition on the log-Mel spectrum to extract low-frequency and high-frequency components:

[0087] S2.1. Perform a two-dimensional discrete wavelet transform on the time-frequency spectra corresponding to each source and target domain sample, using Daubechies type wavelet basis functions, such as db4, to decompose the original spectra X and obtain the low-frequency components X. low and high-frequency component X high :

[0088] X low ,X high =DWT(X);

[0089] Among them, X low It mainly includes global features such as background noise and environmental information in the audio, while X high It records relatively detailed structural information, such as the vibration and friction characteristics of mechanical parts.

[0090] S2.2 Preserve the high-frequency components of the source domain samples. This is for use in subsequent reconstruction, and its low-frequency components are extracted. Used as a style transfer target. Low-frequency components of the target domain samples are extracted simultaneously. Used to adjust the style of low-frequency components in the source domain.

[0091] S3. Perform style enhancement processing on the low-frequency components of the source domain samples.

[0092] S3.1 Calculate the style statistics of the low-frequency components in the target domain, namely the mean and standard deviation, using the following formula:

[0093]

[0094] Where T and F represent the time and frequency dimensions of the feature map, respectively, and X t,f Let X represent the value of the log-Mel spectrum X in the time t and frequency f dimensions, where ε is a very small constant used to ensure numerical stability.

[0095] S3.2 performs style transfer on the low-frequency components of the source domain;

[0096] After obtaining the style statistics of the target domain, an adaptive instance normalization method is used to enhance the style of the low-frequency components of the source domain samples. First, the low-frequency components of the source domain are... Perform mean and variance normalization operations to obtain the content statistics of the low-frequency components in the source domain. Then perform the adaptive instance normalization transformation:

[0097]

[0098] Through the above operations, while retaining the low-frequency structural information related to the device's operating status in the source domain signal, style statistics from the target domain are injected. This ensures that the adjusted low-frequency components maintain the stable characteristics of the source domain device operation and approximate the background noise of the target domain in terms of spectral distribution. This enables the enhanced samples to possess the statistical characteristics of the target environment, thereby improving their discrimination ability in the target domain.

[0099] like Figure 2 As shown, the wavelet style enhancement module includes source domain content statistics extraction, target domain style statistics extraction, adaptive instance normalization style enhancement, and wavelet inverse transform reconstruction. The module's function is to preserve the low-frequency structure of the source domain samples while replacing their background style, making the enhanced samples more closely resemble the target domain in terms of background characteristics.

[0100] S4. Perform inverse wavelet transform on the adjusted low-frequency components and the source domain high-frequency components to reconstruct the sample.

[0101] After completing the style enhancement, the enhanced low-frequency components will be adjusted to match the low-frequency components. With the original high-frequency components By using inverse discrete wavelet transform, the complete log-Mel spectrum can be restored:

[0102]

[0103] The reconstruction operation uses the same wavelet basis functions as the decomposition step to ensure the reversibility and fidelity of the transformation process.

[0104] S5. Anomaly detection model based on source domain sample pre-training.

[0105] S5.1 Construct an anomaly detection model.

[0106] First, based on normal and abnormal audio samples collected from the source domain, a preliminary classification model for fault type identification is constructed. The model structure can adopt convolutional neural networks, temporal modeling networks, or combinations thereof, such as ResNet, MobileNetV2, and WaveNet. The model input is the log-Melogram obtained after the aforementioned processing, and the output is the corresponding fault category prediction result.

[0107] S5.2 Use source domain samples for model training.

[0108] The model is trained under supervision using labeled normal and abnormal audio samples from the source domain. During training, the fault category corresponding to the sample is used as the target label, and optimization is performed using the cross-entropy loss function or other classification loss functions.

[0109] S6: Fine-tuning the model and aligning it with target domain features based on the prototype network.

[0110] S6.1 Construct the support set and query set and organize the training task.

[0111] The target domain samples are used as the support set S, and the augmented samples are used as the query set Q. An N-way K-shot task construction approach is adopted. Each meta-task T = {S, Q} consists of a support set S and a query set Q. N represents the number of classes in each task (e.g., the number of anomaly classes selected in each task), and K represents the number of samples provided for each class in the support set. The support set S is defined as follows:

[0112]

[0113] in, For the target domain sample, This indicates the anomaly category label.

[0114] The corresponding query set consists of style-enhanced samples, represented as:

[0115] in, To enhance the sample, y q This represents the anomaly category label, where M is the number of samples in the query set.

[0116] S6.2 Constructing Class Prototype Vectors Based on Support Set Samples

[0117] Input each class of support set samples into the feature extractor f θ Next, the prototype vector for each category is calculated. The prototype representation for category c is as follows:

[0118]

[0119] in, This represents the set of supporting samples for category c. For the feature extraction network, the temporal spectrogram is mapped to a one-dimensional embedding vector with dimension D. This is the prototype for class c. The prototype vector represents the center position of this class in the embedding space and is used for subsequent distance metrics and classification prediction.

[0120] S6.3 utilizes prototype metrics to classify query samples and fine-tune the model.

[0121] For each query sample First, feature extraction is performed, and then the Euclidean distance between the feature vector and the prototype of each category is calculated:

[0122]

[0123] Optimize the metric space by minimizing the negative log-likelihood loss:

[0124]

[0125] The model was fine-tuned using negative log-likelihood as the optimization objective.

[0126]

[0127] The final model optimizes L metric This ensures that the query sample is close to the correct target domain prototype in the embedding space, thus completing the target domain feature alignment. The overall process is as follows: Figure 3 As shown.

[0128] S6.4 introduces a classification loss to improve feature discrimination capability.

[0129] To further improve the model's discriminative performance in the target domain, this step introduces an auxiliary classification loss function in addition to the prototype metric loss to enhance the model's feature discrimination ability. Preferably, to further enhance the angular discriminativeness between embedded features, an angular margin optimization strategy, such as ArcFace classification loss, can be introduced. Finally, a weighted joint loss function is used to incorporate both the metric loss and classification loss into the optimization objective, expressed as follows: L = λL metric +(1-λ)L class ;

[0130] Among them, L class The loss represents the category classification loss on the query set, and λ is a hyperparameter used to balance the contributions of the two losses, with a value range of [0,1]. In this embodiment, it is set to λ = 0.5 to achieve a compromise between the two.

[0131] S7: Input the target domain test samples into the adjusted model for anomaly detection.

[0132] S7.1 performs preprocessing and feature extraction on the test samples in the target domain.

[0133] After the model has been fine-tuned, it enters the actual testing phase. First, the test samples to be classified in the target domain undergo the same preprocessing operations as in the training phase to generate the corresponding log-Melograms. Then, these time-varying spectra are input into the fine-tuned anomaly detection model.

[0134] S7.2 Use the trained classification model to determine the fault category.

[0135] like Figure 4 As shown, during the testing phase, target domain samples are directly input into the trained model for forward inference. Based on the extracted features, the model performs classification using the learned fault category discriminator and outputs the corresponding fault category label.

[0136] To verify the performance advantages of the proposed method in cross-domain abnormal sound detection tasks, this paper designs a comparative experiment based on the DCASE2021 Task 2 dataset to evaluate the detection performance of the proposed method in the target domain. In the experiment, the source domain training model uses a dual-path channel attention WaveNet as the baseline model, and different target domain adaptation strategies are introduced to adjust the model parameters.

[0137] The comparison methods include fine-tuning, the optimization-based meta-learning method MAML (Model-Agnostic Meta-Learning), and a prototype network based on metric learning, and their performance is compared with the wavelet style enhancement prototype network proposed in this invention. All methods pre-train the model using source domain data and then fine-tune the model parameters using a very small number of target domain samples.

[0138] The performance evaluation metrics used were the area under the ROC curve (AUC) and partial AUC (pAUC). Tests were conducted on seven types of equipment: toy cars, toy trains, fans, gearboxes, pumps, slides, and valves. Cross-domain detection performance was statistically analyzed. The experimental results are shown in Tables 1 and 2.

[0139] Table 1. Performance comparison of AUC (%) of various methods on the DCASE 2021 Task 2 dataset.

[0140]

[0141] Table 2 compares the pAUC (%) performance of each method on the DCASE 2021 Task 2 dataset;

[0142]

[0143]

[0144] As can be seen from the data in the table, the proposed wavelet-enhanced prototype network method achieves the best performance across all device types, with particularly significant improvements in device categories with significant domain offsets, such as ToyTrain and Valve.

[0145] The above description is merely a further embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope disclosed in the present invention, based on the technical solution and concept of the present invention, shall fall within the scope of protection of the present invention.

Claims

1. A method for detecting transdomain transformer abnormal sounds based on wavelet style enhancement prototype networks, characterized in that: Includes the following steps: S1: Obtain source domain audio samples and target domain audio samples, and perform time-frequency graph conversion on the audio samples to obtain the corresponding time-frequency graph; S2: Perform discrete wavelet decomposition on the time-spectrum diagrams of the source domain sample and the target domain sample respectively to obtain the corresponding low-frequency components and high-frequency components; S3: Perform style enhancement processing on the low-frequency components of the source domain samples to make them statistically similar to the background style of the target domain, thereby obtaining style-enhanced low-frequency components. S4: Combine the style-enhanced low-frequency components with the original source domain high-frequency components, and reconstruct the time-spectrum using inverse wavelet transform to obtain an enhanced sample that closely resembles the characteristics of the target domain. S5: An anomaly detection model is obtained by pre-training a large number of source domain samples; S6: During the training phase, class prototypes are constructed based on the target domain samples and input into the anomaly detection model along with style enhancement samples. With the help of the prototype network, the distance between the target domain samples and the enhancement samples is narrowed in the feature space, thereby adjusting the parameters of the anomaly detection model to make it suitable for anomaly detection in the target domain. S7: During the testing phase, the target domain test samples are input into the adjusted model for feature extraction and classification.

2. The method for detecting cross-domain transformer abnormal sounds based on wavelet style enhancement prototype networks according to claim 1, characterized in that: This includes the original audio files collected in the target domain environment; The target domain environment refers to the actual application scenario in which the device under test is located; The time-spectrum image is a two-dimensional image obtained by performing short-time Fourier transform or Mel spectrum transformation on the audio sample.

3. The method for detecting abnormal sounds in cross-domain transformers based on wavelet style enhancement prototype networks according to claim 1, characterized in that: The wavelet decomposition in S2 uses discrete two-dimensional wavelet transform to decompose the time-spectrum graph and extract low-frequency and high-frequency components respectively. The low-frequency components retain the overall background and structural information, while the high-frequency components retain local transient features.

4. The method for detecting abnormal sounds in cross-domain transformers based on wavelet style enhancement prototype networks according to claim 3, characterized in that: The style enhancement process in S3 includes: S3.1: Extract the mean and standard deviation of the low-frequency components of the target domain samples to form the style statistics of the target domain; S3.2: Apply the target domain style statistics to the low-frequency components of the source domain samples, perform style transfer through adaptive instance normalization, and adjust the low-frequency distribution of the source domain samples.

5. The method for detecting cross-domain transformer abnormal sounds based on wavelet style enhancement prototype networks according to claim 4, characterized in that: In step S4, during the inverse wavelet transform, the low-frequency component after style enhancement is combined with the original high-frequency component, and the wavelet basis function corresponding to the wavelet decomposition is applied to reconstruct the complete time-frequency spectrum. The reconstructed time-frequency spectrum is used as the training sample after style enhancement to train and adjust the model parameters that have been pre-trained based on the source domain samples.

6. The method for detecting abnormal sounds in cross-domain transformers based on wavelet style enhancement prototype networks according to claim 5, characterized in that: It includes an audio preprocessing module, a style enhancement module, a prototype network module, and an anomaly detection module; The audio preprocessing module is used to perform time-frequency graph conversion and standardization on the audio samples. The style enhancement module is used to perform style transfer adjustment on the low-frequency components of source domain samples based on the style statistics of the target domain, and generate style-enhanced samples. The prototype network module constructs class prototypes based on target domain samples, enabling feature alignment with style enhancement samples and fine-tuning of model parameters; The anomaly detection module is used to extract features and classify the target domain samples to be tested, thereby completing anomaly detection.