A device anomaly monitoring method based on audio signal analysis

By employing adaptive feature enhancement and fusion, time-frequency decoupling feature extraction, and cross-modal feature fusion, this method addresses the shortcomings in robustness and generalization ability of existing equipment anomaly detection technologies, achieving efficient and accurate detection of equipment anomalies and making it suitable for condition monitoring of industrial equipment.

CN120766722BActive Publication Date: 2025-11-11TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL +1
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202511272147.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-08
Publication Date
2025-11-11
Estimated Expiration
2045-09-08

AI Technical Summary

Technical Problem

Existing audio analysis-based methods for detecting device anomalies have significant shortcomings in terms of robustness, generalization ability, and detection performance. In particular, they are difficult to reliably detect abnormal states of equipment in complex and ever-changing industrial audio scenarios. Traditional methods have limited ability to extract features from raw audio and log-Mel spectrograms, making it difficult to fully capture the temporal dynamics and abnormal features in sound signals.

Method used

We employ adaptive feature enhancement and fusion, time-frequency decoupling feature extraction, and cross-modal feature fusion. We generate multi-channel fused features through an adaptive gating module, perform deep feature extraction using a selective state-space model, and train the model by combining a self-supervised ID classification task and an angular interval loss function. We design a bi-branch Mamba network for anomaly scoring during spectral analysis.

Benefits of technology

It improves the robustness and generalization ability of equipment anomaly detection, enabling efficient and accurate anomaly detection under different equipment types and complex operating conditions, and provides a more reliable and economical anomaly detection solution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120766722B_ABST
    Figure CN120766722B_ABST
Patent Text Reader

Abstract

This invention discloses a device anomaly monitoring method based on audio signal analysis, comprising the following steps: S1. Generating a log-Mel spectrum through time-frequency transformation, generating a gate vector using an adaptive gating module based on median, root mean square, and variance, and then weighted and fused with the original spectrum to supplement audio features to form a multi-channel fused feature. S2. Dividing the multi-channel fused feature into blocks along the time and frequency axes, and inputting them into parallel time-domain and frequency-domain branches respectively, both using a selective state-space model. The time-domain branch captures long-term steady-state features and trend anomalies, while the frequency-domain branch analyzes cross-band energy distribution and harmonic structure. S3. Linearly aligning the features of the two branches and adding them element-wise to obtain the fused feature, training the model through a self-supervised ID classification task combined with an angular interval loss function, and using the negative logarithmic probability of the target ID category as the anomaly score during inference, thereby improving the robustness, generalization ability, and accuracy of detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to audio signal analysis and equipment anomaly monitoring technology, and in particular to a method for equipment anomaly monitoring based on audio signal analysis. Background Technology

[0002] Intelligent, automated, and refined operation and maintenance of equipment have become crucial for modern industrial production. However, equipment faces issues such as wear, aging, and malfunctions during long-term operation, which not only affect production efficiency but may also pose serious safety hazards. Therefore, how to detect abnormal equipment conditions in a timely and accurate manner and ensure reliable equipment operation has become a focus of attention in the industrial sector.

[0003] Traditional equipment anomaly detection methods primarily rely on physical sensors, such as vibration sensors, temperature sensors, and pressure sensors. These sensors collect the equipment's physical parameters to determine if it is in an abnormal state. However, these methods suffer from complex installation, high maintenance costs, and significant dependence on the equipment's structure and environment. In recent years, audio analysis-based anomaly detection methods have gradually become a hot technology for industrial equipment anomaly detection due to their non-contact, low-cost, and easy-to-deploy advantages.

[0004] The goal of anomaly detection based on audio analysis is to identify anomalies by learning the sound characteristics of devices during normal operation. Currently, research mainly focuses on anomaly detection methods based on self-supervised classification models and generative reconstruction models. Examples include: an improved AE-based machine anomalous sound detection method and apparatus (patent application number 202310933972.5); a self-supervised machine anomalous sound detection method based on dual-path WaveNet (patent application number 202411896447.1); an unsupervised machine anomalous sound detection method using lightweight networks (patent application number 202311316592.3); a domain-transfer-based self-supervised machine anomalous sound detection method (patent application number 202210863510.6); and a multi-class machine anomalous sound detection method based on a diffusion model (patent application number 202510161917.8).

[0005] However, the most attention-grabbing anomalous sound detection techniques, such as the multi-class machine anomalous sound detection methods based on diffusion models (patent applications 202310933972.5 and 202510161917.8), are improvements to AE (Automatic Anomalous Detection) methods that detect anomalies by learning the distribution of normal data. These models attempt to reconstruct or generate samples similar to normal data during training; when the model cannot reconstruct or generate similar samples, they are marked as anomalous. However, these methods are prone to overfitting; that is, the model performs well on training data but may fail to effectively detect anomalies when faced with new data or unseen anomalies. Another example is a self-supervised machine anomalous sound detection method based on dual-path WaveNet proposed in patent application 202411896447.1. This method models the frequency information of each time frame of the input sound signal and also models the time dimension. However, the receptive field of this model is limited to a local time window, making it difficult to effectively capture the characteristics of the machine during long-term steady-state operation, and it also cannot model trend-level anomalous patterns.

[0006] In summary, existing abnormal sound detection methods have significant shortcomings in robustness, generalization ability, and detection performance, especially in complex and variable industrial audio scenarios, where they struggle to reliably detect abnormal equipment states. Traditional methods have limited ability to extract features from raw audio and log-Mel spectrograms, making it difficult to fully capture the temporal dynamics and abnormal features in sound signals, thus limiting detection performance.

[0007] It should be noted that the information disclosed in the background section above is only for understanding the background of this application, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0008] The main objective of this invention is to overcome the deficiencies in the aforementioned background technology and provide a device anomaly monitoring method based on audio signal analysis.

[0009] To achieve the above objectives, the present invention adopts the following technical solution:

[0010] A device anomaly monitoring method based on audio signal analysis includes the following steps:

[0011] S1. Adaptive Feature Enhancement and Fusion: The original audio signal is transformed by time-frequency to generate a log-Mel spectrum. The median, root mean square and variance statistics of each time frame are calculated by an adaptive gating module based on third-order statistical indices. The gating vector is generated and fused with the original spectrum in a weighted manner. The audio features are then spliced ​​together to form a multi-channel fused feature.

[0012] S2. Time-frequency decoupling feature extraction: The multi-channel fused features are divided into time-domain block sequences and frequency-domain block sequences along the time axis and frequency axis, respectively, and input into parallel time-domain branches and frequency-domain branches; wherein the time-domain branch performs causal scanning on the time-domain block sequences through a selective state-space model to capture long-term steady-state features and trend anomalies, and the frequency-domain branch performs causal scanning on the frequency-domain block sequences through a selective state-space model to analyze cross-band energy distribution and harmonic structure;

[0013] S3. Cross-modal feature fusion and anomaly scoring: The feature vectors output by the time domain branch and the frequency domain branch are linearly aligned and then added element by element to obtain a fused feature representation; based on the fused feature representation, a self-supervised ID classification task combined with the angle interval loss function is used for model training, and the negative log probability of the target ID category is used as the anomaly score during the inference stage.

[0014] Furthermore, the adaptive gating module described in step S1 generates the gating vector through the following process:

[0015] The median, root mean square, and variance were calculated frame by frame on the transposed log-Mel plot.

[0016] The three statistics are summed and then subtracted from the preset offset, adjusted by the scaling factor, and activated by the Sigmoid function.

[0017] The weighted fusion is achieved by multiplying the gated vector element-wise with the original log-Mel spectrum to enhance the weight of time frames.

[0018] Furthermore, the multi-channel fusion feature described in step S1 is constructed in the following manner:

[0019] The gated enhanced log-Mel spectrum is used as the first channel;

[0020] The original log-Mel spectrum is retained as the second channel;

[0021] Supplement the temporal characteristics of the original audio as a third channel;

[0022] The three-channel features are spliced ​​along the channel dimension to form the final input features.

[0023] Furthermore, the processing flow of the time-domain branch described in step S2 specifically includes:

[0024] Flatten the temporal block sequence and add token-like and positional embeddings;

[0025] The process involves layer normalization, linear transformation, two-dimensional convolution, and activation function sequentially.

[0026] Selective state-space model parameters are dynamically generated based on input features;

[0027] Capture long-range dependencies through a state-space modeling layer;

[0028] The output is added to the original input through residual connection.

[0029] Furthermore, the frequency domain branching process described in step S2 specifically includes:

[0030] Flatten the frequency domain block sequence and add token-like and position embeddings;

[0031] The process involves layer normalization, linear transformation, two-dimensional convolution, and activation function sequentially.

[0032] Selective state-space model parameters are dynamically generated based on input features;

[0033] Frequency domain structural characteristics are analyzed through state-space modeling layers;

[0034] The output is added to the original input through residual connection.

[0035] Furthermore, in step S2, the selective state-space model achieves input-dependent state transitions through dynamic parameter calculation, specifically including:

[0036] Time step parameters are generated from intermediate features through a linear transformation;

[0037] The state transition matrix is ​​calculated using time step parameters and preset parameters;

[0038] The input projection matrix is ​​generated by linear transformation of intermediate features;

[0039] State-space modeling operations are performed using dynamically generated state transition matrices and input projection matrices.

[0040] Furthermore, the cross-modal feature fusion described in step S3 specifically includes:

[0041] Linear projection is performed on the output features of the time-domain branch and the output features of the frequency-domain branch respectively to achieve feature space alignment;

[0042] The two aligned feature vectors are added element-wise to generate the final fused feature.

[0043] Furthermore, the self-supervised training process described in step S3 specifically includes:

[0044] A classification task is constructed using the device machine ID as the category label;

[0045] The model is optimized using the angular interval loss function, which forces the features of different machine IDs to maximize the inter-class distance and minimize the intra-class distance in the angular space.

[0046] The angle interval loss function introduces an additive angle interval constraint feature distribution.

[0047] Furthermore, the method for calculating the abnormal score in step S3 includes:

[0048] Input the audio to be tested into the trained model;

[0049] Extract the predicted probability of its corresponding real machine ID category;

[0050] The negative natural logarithm of the predicted probability is calculated as the outlier score.

[0051] A computer program product includes a computer program that, when executed by a processor, implements the device anomaly monitoring method based on audio signal analysis.

[0052] The present invention has the following beneficial effects:

[0053] To address the problems of existing technologies, this invention proposes an audio analysis-based method for equipment anomaly detection. This method enhances feature extraction capabilities by introducing a temporal feature enhancement network, a multi-scale temporal feature extraction module, and a latent diffusion adversarial generative model to deeply mine multi-scale temporal information of audio signals. Simultaneously, it employs an innovative end-to-end deep learning framework for refined analysis of audio signals from industrial equipment operation. The overall technical solution consists of three core stages: first, adaptive feature enhancement and fusion based on multiple statistical indicators; second, deep feature extraction using a time-frequency decoupled dual-branch Mamba network; and finally, model training and anomaly score calculation through a self-supervised classification task and a unique loss function. This method not only improves the robustness and generalization ability of detection but also maintains efficient and accurate detection performance when facing different equipment types, complex operating conditions, and situations where no anomalies are observed, providing a more reliable and economical anomaly detection solution for industrial equipment.

[0054] Compared with existing technologies, this invention achieves significant improvements in both the mechanism and effectiveness of abnormal sound detection. Its fundamental advantage lies in the first-time integration of time-frequency decoupling modeling with the Selective State-Space Model (SSM) and its application to the field of abnormal sound detection. Through a designed spectral-temporal dual-branch Mamba (STMamba) backbone network, this method completely overcomes the limitations of traditional models confined to local receptive fields: the temporal Mamba (TMamba) branch in this architecture can perform long-range scanning along the time axis, effectively capturing the long-term acoustic characteristics and trend changes during steady-state machine operation; while the parallel spectral Mamba (SMamba) branch accumulates and analyzes cross-band energy distribution and harmonic structure along the frequency axis. This strategy of separately modeling and then fusing the two orthogonal dimensions of time and frequency allows the model to perform a more refined and comprehensive analysis of the machine's acoustic characteristics with unprecedented granularity, thus overcoming the inherent problem of existing methods neglecting dynamic time-frequency coupling due to independent processing of spectral frames.

[0055] Another significant benefit of this invention lies in its innovative feature enhancement mechanism, which fundamentally improves the model's sensitivity to and robustness in detecting anomalous patterns. The invention's unique TriStat-Gating (TSG) gating mechanism does not passively receive acoustic features but actively and adaptively enhances them: by constructing a gating signal based on three complementary statistical indices—median, root mean square, and variance—it intelligently amplifies time regions in the input spectrum that exhibit potential anomalies in stability, intensity, or distribution uniformity. This mechanism ensures that the enhanced features (ESTgrams) ultimately fed into the backbone network carry more discriminative information, making it easier for the model to capture subtle differences between normal and anomalous conditions. This intelligent front-end enhancement, combined with the powerful time-frequency decoupling modeling capabilities of the back-end, creates a synergistic effect, not only improving the model's detection accuracy when facing diverse equipment and complex operating conditions but, more importantly, enhancing its generalization ability and stability when facing unknown anomaly types, providing a more reliable and advanced solution for the condition monitoring of industrial equipment.

[0056] Other beneficial effects of the embodiments of the present invention will be further described below. Attached Figure Description

[0057] Figure 1 This is a flowchart illustrating the overall process of the device anomaly detection method based on audio analysis according to the present invention.

[0058] Figure 2 This is an algorithm architecture diagram of the device anomaly detection method based on audio analysis according to an embodiment of the present invention. Detailed Implementation

[0059] The embodiments of the present invention will be described in detail below. It should be emphasized that the following description is merely exemplary and not intended to limit the scope and application of the present invention.

[0060] Definitions of abbreviations and key terms:

[0061] Audio analytics: A technology that assesses the health of equipment by collecting and analyzing sound signals emitted during operation. It provides a non-contact, low-cost method for fault detection, enabling real-time monitoring of equipment anomalies through signal processing, feature extraction, and modeling.

[0062] Abnormal sound detection: A method for identifying unusual sounds produced by machines or equipment during operation. Its principle is to first analyze the sound characteristics of the equipment during normal operation; when these characteristics show abnormal changes, the system identifies and marks them as abnormal events. This method has wide applications in various fields such as industrial manufacturing, energy equipment, transportation, and robotics automation, aiming to improve the reliability and safety of equipment.

[0063] Selective State-Space Models (SSMs) are an advanced sequence modeling architecture that adaptively focuses on key information within a sequence by introducing an input-dependent state transition mechanism. SSMs are effective for modeling long program sequences and are well-suited for handling long-term steady-state characteristics and sudden anomalous events in machine audio signals.

[0064] TriStat-Gating (TSG) Gating Module: This invention proposes a feature enhancement module that enhances the input Mel spectrum by introducing an adaptive gating mechanism based on third-order statistical indices (median, root mean square, and variance). This module improves the model's sensitivity to unknown abnormal patterns.

[0065] Log-Mel spectrogram: A two-dimensional time-frequency representation of audio signals. It transforms the audio signal to the frequency domain using a short-time Fourier transform (STFT), then maps it to the Mel scale, and finally takes the logarithm. This representation method conforms to the auditory characteristics of the human ear and is a commonly used feature in the field of audio processing.

[0066] This invention aims to improve the robustness, generalization ability, and detection performance of equipment anomaly detection under different equipment types, complex operating conditions, and no-anomaly scenarios. It proposes an equipment anomaly detection method based on audio analysis. This method introduces a temporal feature enhancement network, a multi-scale temporal feature extraction module, and a latent diffusion adversarial generative model to deeply mine the multi-scale temporal information of audio signals, enhance feature extraction capabilities, and provide a more reliable and economical anomaly detection solution for industrial equipment.

[0067] See Figure 1 This invention provides a device anomaly monitoring method based on audio signal analysis, comprising the following steps:

[0068] Step S1, Adaptive Feature Enhancement and Fusion: The original audio signal is transformed by time-frequency to generate a log-Mel spectrum. The median, root mean square and variance statistics of each time frame are calculated by an adaptive gating module (TSG) based on third-order statistical indices. The gating vector is generated and fused with the original spectrum in a weighted manner. The audio features are then spliced ​​together to form a multi-channel fused feature.

[0069] To eliminate the dimensional differences between different statistical indicators, the three calculated statistical vectors can be standardized (e.g., Z-score standardization) to ensure a mean of 0 and a variance of 1. These three normalized statistics are then calculated frame-by-frame on the transposed spectrogram and concatenated to supplement audio features, forming a multi-channel fusion feature.

[0070] In some embodiments, the adaptive gating module generates a gating vector through the following process: calculating the median, root mean square, and variance frame by frame on the transposed log-Mel spectrogram; summing the three statistics and subtracting a preset offset, adjusting by a scaling factor and activating through the Sigmoid function; the weighted fusion is achieved by multiplying the gating vector element by element with the original log-Mel spectrogram to achieve time frame weight enhancement.

[0071] In some embodiments, the multi-channel fusion feature is constructed by: using a gated enhanced log-Mel spectrogram as the first channel; retaining the original log-Mel spectrogram as the second channel; supplementing the temporal features of the original audio as the third channel; and concatenating the three-channel features along the channel dimension to form the final input feature.

[0072] Step S2, Time-Frequency Decoupling Feature Extraction: The multi-channel fused features are divided into time-domain block sequences and frequency-domain block sequences along the time and frequency axes, and respectively input into parallel time-domain branches and frequency-domain branches (spectral-time dual-branch Mamba (STMamba) backbone network); wherein the time-domain branch performs causal scanning on the time-domain block sequences through a selective state-space model (SSM) to capture long-term steady-state features and trend anomalies, and the frequency-domain branch performs causal scanning on the frequency-domain block sequences through a selective state-space model (SSM) to analyze cross-band energy distribution and harmonic structure.

[0073] In some embodiments, the processing flow of the temporal branch specifically includes: flattening the temporal block sequence and adding class tokens and position embeddings; processing it sequentially through layer normalization, linear transformation, two-dimensional convolution and activation function; dynamically generating selective state space model parameters based on input features; capturing long-range dependencies through a state space modeling layer; and adding the output to the original input through residual connections.

[0074] In some embodiments, the processing flow of the frequency domain branch specifically includes: flattening the frequency domain block sequence and adding class tokens and position embeddings; processing it sequentially through layer normalization, linear transformation, two-dimensional convolution and activation function; dynamically generating selective state space model parameters based on input features; analyzing frequency domain structural features through state space modeling layers; and adding the output to the original input through residual connections.

[0075] In some embodiments, the selective state-space model realizes input-dependent state transitions through dynamic parameter calculation, specifically including: generating time step parameters from intermediate features through linear transformation; calculating a state transition matrix using the time step parameters and preset parameters; generating an input projection matrix from intermediate features through linear transformation; and performing state-space modeling operations using the dynamically generated state transition matrix and input projection matrix.

[0076] Step S3, Cross-modal feature fusion and anomaly scoring: The feature vectors output by the time domain branch and the frequency domain branch are linearly aligned and then added element by element to obtain the fused feature representation; Based on the fused feature representation, a self-supervised ID classification task combined with the angle interval loss function is used for model training, and the negative log probability of the target ID category is used as the anomaly score during the inference stage.

[0077] In some embodiments, the cross-modal feature fusion specifically includes: performing linear projection on the time-domain branch output features and the frequency-domain branch output features respectively to achieve feature space alignment; and performing element-wise addition on the aligned two feature vectors to generate the final fused features.

[0078] In some embodiments, the self-supervised training process specifically includes: constructing a classification task using device machine ID as a category label; optimizing the model using an angular interval loss function, forcing features of different machine IDs to maximize inter-class distance and minimize intra-class distance in the angular space; the angular interval loss function constrains feature distribution by introducing additive angular interval.

[0079] In some embodiments, the calculation of the anomaly score specifically includes: inputting the audio to be tested into the trained model; extracting the predicted probability of its corresponding real machine ID category; and calculating the negative natural logarithm of the predicted probability as the anomaly score.

[0080] The main technical advantage of this invention lies in its innovative time-frequency decoupling modeling and adaptive feature enhancement mechanism, which significantly improves the robustness, generalization ability, and detection accuracy of abnormal sound detection in industrial equipment. Specifically, it combines the Selective State-Space Model (SSM) with the time-frequency decoupling concept to design a spectral-temporal dual-branch Mamba network: the time-domain branch accurately captures the long-term acoustic characteristics and trend anomalies of steady-state equipment operation through long-range causal scanning, while the frequency-domain branch deeply analyzes cross-band energy distribution and harmonic structure to identify instantaneous abnormal events, thus completely breaking through the limitations of the local receptive field of traditional models and achieving full-granular analysis of machine acoustic features. Simultaneously, the TriStat-Gating gating mechanism actively enhances the input features, dynamically amplifying the temporal regions of stability, intensity, or distribution anomalies in the spectrum based on third-order statistical indices of median, root mean square, and variance, enabling the enhanced features (ESTgram) to carry stronger discriminative information. The front-end intelligent enhancement and the back-end time-frequency decoupling modeling create a synergistic effect, not only improving the detection accuracy of diverse equipment and complex operating conditions but also enhancing the generalization ability to unknown anomalies, providing a highly reliable and low-cost anomaly monitoring solution for industrial equipment.

[0081] The following further describes specific embodiments of the present invention and examples of its algorithm implementation.

[0082] This invention proposes an anomaly detection method called "Enhanced Spectral Time-Based Dual-Branch Mamba (ESTM)". This method utilizes an innovative, end-to-end deep learning framework to perform refined analysis of audio signals during industrial equipment operation, achieving efficient and accurate anomaly detection. The overall technical solution can be divided into three core stages: first, adaptive feature enhancement and fusion based on multiple statistical indicators; second, deep feature extraction using a time-frequency decoupled dual-branch Mamba network; and finally, model training and anomaly score calculation through a self-supervised classification task and a unique loss function.

[0083] like Figure 2The diagram illustrates the complete chain of the device anomaly detection method based on audio analysis according to an embodiment of the present invention: Starting from the input of the original audio signal, the original audio signal is processed by a time-spectral network to extract time-domain features (Tgram) and by a short-time Fourier transform to obtain spectral features (Sgram). Subsequently, the Sgram is enhanced by a TriStat-Gating (TSG) gating module to obtain enhanced spectral features (ESgram). Next, the Tgram, Sgram, and ESgram are concatenated and fused along the channel dimension to form a fused feature (ESTgram), where T represents the time dimension and S represents the spectral dimension. The fused feature (ESTgram) generation process in the first part forms the key input, and then the TMamba and SMamba branches in the STMamba network, which are described in detail in the second part, are used for feature extraction. Finally, anomaly detection is completed through a classification head.

[0084] I. Adaptive Feature Generation and Fusion Based on TriStat-Gating

[0085] The initial steps of this invention do not directly use traditional acoustic features, but instead generate an enhanced multi-channel fusion feature, ESTgram, that is more sensitive to anomalous patterns through an original process. This process begins with the normalization preprocessing of the original mono audio signal x, applying a short-time Fourier transform (STFT) with a window size of 1024 and a jump length of 512, and combining it with a 128-Mel filter bank to generate a basic two-dimensional logarithmic Mel spectrum. However, using only the spectrum cannot fully highlight all potential anomalous information. Therefore, the core of this invention is the TriStat-Gating (TSG) gating module, which adaptively enhances time frames containing key information within the spectrum. The TSG module operates on the principle that anomalous sounds from a machine often disrupt the stability, intensity consistency, or uniformity of its normal sound characteristics. Therefore, this module computes three complementary higher-order statistical metrics in parallel—median, root mean square (RMS), and variance—to quantify the acoustic characteristics of each time frame from different dimensions. These three statistics are displayed in the transposed spectrum. The calculation is performed frame by frame and combined into a gating signal, which is then passed through the Sigmoid activation function. Generate a gated vector between 0 and 1. This vector is then broadcast and compared with the original spectrum. This process can automatically identify and increase the weight of time frames whose statistical characteristics deviate. Its mathematical expression is as follows:

[0086]

[0087] in This is a scaling factor set to 2. It should be noted that the specific values ​​of key hyperparameters, such as the scaling factor α and offset μ in the TSG module, and the number of blocks in the STMamba network, can be systematically determined using standard optimization methods such as grid search or Bayesian optimization on the standard validation set to maximize the model's anomaly detection performance (e.g., AUC score on the validation set). Finally, to construct an information-complete input, the original log-Mel spectrum is... , and a complementary original audio feature The Tgram is concatenated along the channel dimension to form the final three-channel fused feature ESTgram. This fusion design based on orthogonal statistical indices ensures the richness and non-redundancy of the input features, providing a solid foundation for the subsequent network to capture subtle time-frequency differences introduced by anomalous patterns.

[0088] II. Time-Frequency Decoupled Two-Branch Mamba (STMamba) Backbone Network

[0089] After the first part of adaptive feature generation and fusion based on TriStat-Gating, which yields the three-channel fused feature ESTgram containing rich time-frequency information, the next step is feature extraction.

[0090] After obtaining the enhanced ESTgram features, this invention employs a novel spectral-temporal dual-branch Mamba (STMamba) backbone network for deep feature extraction. The core of this network architecture addresses the challenge of applying the advanced one-dimensional sequence model Mamba to two-dimensional temporal-spectral analysis. The solution is time-frequency decoupling modeling: instead of treating the temporal spectrum as a whole, it splits it into two orthogonal dimensions—time and frequency—and designs a dedicated Mamba processing branch for each dimension. Specifically, the input two-dimensional ESTgram is first patched, divided into 12 time-domain patches along the time axis and 16 frequency-domain patches along the frequency axis. The number of blocks mentioned above was determined primarily based on the following principles: the choice of 12 time-domain blocks aims to balance computational efficiency with the needs of capturing the complete operating cycle of the device, ensuring that each block contains sufficient temporal information while the total sequence length is suitable for long-range dependency modeling; the number of 16 frequency-domain blocks is determined based on the total number of Mel spectrum bands of 128, dividing the total frequency band into 16 sub-bands (each sub-band contains 8 Mel frequencies). This approach preserves key frequency band resolution while enabling the model to effectively learn the correlations between different frequency bands (such as fundamental frequency and harmonics). These blocks are flattened, linearly projected, and learnable class tokens are added. and position embedding This preserves the global and positional context of the sequences in the original spectrogram. The two sequences are then fed into a parallel, two-branch Mamba architecture:

[0091] TMamba (temporal branch) performs a left-to-right causal scan of the time-domain block sequence, leveraging the powerful long-range memory capability of the Selective State-Space Model (SSM) to accurately capture the steady-state operating patterns, periodic vibrations, and trend anomalies that evolve over time in machine sound; simultaneously,

[0092] SMamba (spectral branch) performs a top-down causal scan of frequency domain block sequences. Its design aims to accumulate and analyze cross-band energy distribution and harmonic structure, effectively identifying the spectrum caused by transient anomalous events such as impacts and friction. Two branches work in parallel, deeply mining audio features from a global time and frequency perspective. Finally, the feature vectors output from the two branches are linearly aligned and fused to generate a highly condensed and comprehensive final feature representation H.

[0093] The specific computational flow of this dual-branch Mamba backbone network is shown in Table 1. The algorithm embeds sequences from two inputs in the time and frequency domains (…). and The algorithm performs parallel but weight-shared processing. In each step of the core loop, the input sequence X first undergoes preliminary feature mapping through layer normalization (LayerNorm), linear transformation, and a 2D convolutional layer. Subsequently, the algorithm enters the most crucial dynamic parameter calculation stage of the Mamba model: generating selective parameters through a series of data-dependent operations. A, B, and C. Specifically, the parameters... B is obtained by analyzing the input features. The parameters are calculated after a linear transformation, enabling the State Space Model (SSM) to dynamically adjust its state transitions based on the input content, thus achieving "selective" focus on key information in the sequence. Using these dynamically calculated parameters, the input sequence is transformed through the core State Space Modeling (SSM) layer to capture long-range dependencies within the sequence with extremely high efficiency. Finally, through a residual connection, the output of the SSM module is linearly transformed and added to the original input—a widely used technique in deep learning designed to ensure smooth information flow and stability during training. After both branches (time-domain and frequency-domain branches) have independently completed the above processing steps, the algorithm enters the Cross-Podal Fusion stage. The output features obtained by each branch are then... and It will first go through a shared linear alignment layer (Linear) align The purpose of this approach is to project features from different modalities onto the same feature subspace to ensure semantic comparability. After feature alignment, the two feature vectors are finally fused through simple element-wise addition to obtain a more powerful and comprehensive final feature representation H that simultaneously contains long-term temporal domain information and global frequency domain information.

[0094] Table 1

[0095]

[0096] III. Self-supervised training and outlier score calculation based on ArcFace loss

[0097] This invention employs a self-supervised ID classification strategy for end-to-end training of the entire framework, cleverly transforming the output of this classification task into anomaly scores. During training, device metadata (i.e., machine type and unique machine ID) is used as readily available labels. The fusion features H extracted by the STMamba network are fed into a fully connected layer with a Softmax activation function to predict the machine ID of the input audio; this part of the network responsible for classification prediction and anomaly score generation can be called the classification head. Unlike the cross-entropy loss used in traditional methods, this invention uses the ArcFace loss function. The core advantage of ArcFace lies in its application of an additive angular interval to features in angular space, forcing the model to learn features with greater separation between different categories (i.e., different machine IDs), while simultaneously making the feature distribution within the same category more compact. This characteristic is crucial for anomaly detection because it makes the "normal" feature space more clearly and rigorously defined, making any samples deviating from this space easier to identify. The entire network is trained using the AdamW optimizer with a learning rate of 0.0001, for a total of 200 training epochs, and a batch size of 128. During the inference phase, when an audio sample is input into the model, its anomaly score is directly defined as the negative log probability of its corresponding ID category. Since the model has only learned the compact distribution of normal sounds, when it encounters anomalies whose acoustic characteristics do not match the model's training, the model cannot classify them into any known normal category with high confidence. This results in a lower prediction probability, leading to a higher negative log probability value, i.e., a higher anomaly score, thereby achieving accurate alerts for abnormal events.

[0098] In summary, this invention proposes a device anomaly monitoring method based on audio signal analysis. Compared with existing technologies, this invention achieves significant improvements in both the mechanism and effectiveness of device anomaly sound detection. Its fundamental advantage lies in the fact that this invention, for the first time, combines the time-frequency decoupling modeling concept with the Selective State-Space Model (Mamba) and introduces it into the field of anomaly sound detection. Through the designed spectral-temporal dual-branch Mamba (STMamba) backbone network, this method completely overcomes the limitations of traditional models that are restricted to local receptive fields. The temporal Mamba (TMamba) branch in this architecture can perform long-range scanning along the time axis, effectively capturing the long-term acoustic characteristics and trend changes of the machine during steady-state operation, while the parallel spectral Mamba (SMamba) branch accumulates and analyzes cross-band energy distribution and harmonic structure along the frequency axis. This strategy of separately modeling and then fusing the two orthogonal dimensions of time and frequency allows the model to perform a more refined and comprehensive analysis of the machine's acoustic characteristics with unprecedented granularity, thereby overcoming the inherent problem of existing methods neglecting the dynamic coupling of time and frequency due to independent processing of spectral frames.

[0099] Another significant benefit of this invention lies in its innovative feature enhancement mechanism, which fundamentally improves the model's sensitivity and robustness to anomaly patterns. The invention's unique TriStat-Gating (TSG) gating module does not passively receive acoustic features but actively and adaptively enhances them. This module intelligently amplifies time regions in the input spectrum that exhibit potential anomalies in stability, intensity, or distribution uniformity by constructing a gating signal based on three complementary statistical indices: median, root mean square, and variance. This mechanism ensures that the enhanced features (ESTgrams) ultimately fed into the backbone network carry more discriminative information, allowing the model to more easily capture subtle differences between normal and abnormal conditions. This intelligent front-end enhancement, combined with the powerful time-frequency decoupling modeling capabilities of the back-end, creates a synergistic effect. This not only improves the model's detection accuracy when facing diverse equipment and complex operating conditions but, more importantly, enhances its generalization ability and stability when dealing with unknown anomaly types, providing a more reliable and advanced technical solution for the condition monitoring of industrial equipment.

[0100] This invention also provides a storage medium for storing a computer program, which, when executed, performs at least the methods described above.

[0101] This invention also provides a control device, including a processor and a storage medium for storing a computer program; wherein the processor executes the computer program by performing at least the method described above.

[0102] This invention also provides a processor that executes a computer program, at least performing the methods described above.

[0103] The storage medium can be implemented by any type of non-volatile storage device, or a combination thereof. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc or CD-ROM; magnetic surface memory can be disk storage or magnetic tape storage. The storage media described in the embodiments of this invention are intended to include, but are not limited to, these and any other suitable types of memory.

[0104] In the several embodiments provided by this invention, it should be understood that the disclosed systems and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or units can be electrical, mechanical, or other forms.

[0105] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of this embodiment according to actual needs.

[0106] In addition, in the various embodiments of the present invention, each functional unit can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.

[0107] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0108] Alternatively, if the integrated units of this invention are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this invention, or the parts that contribute to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, RAM, magnetic disks, or optical disks.

[0109] The methods disclosed in the several method embodiments provided by this invention can be arbitrarily combined without conflict to obtain new method embodiments.

[0110] The features disclosed in the several product embodiments provided by this invention can be arbitrarily combined without conflict to obtain new product embodiments.

[0111] The features disclosed in the several method or device embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method or device embodiments.

[0112] The above description, in conjunction with specific preferred embodiments, provides a further detailed explanation of the present invention. It should not be construed that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art, various equivalent substitutions or obvious modifications can be made without departing from the concept of the present invention, and all such modifications, achieving the same performance or application, should be considered within the scope of protection of the present invention.

Claims

1. A method for monitoring device anomalies based on audio signal analysis, characterized in that, Includes the following steps: S1. Adaptive Feature Enhancement and Fusion: The original audio signal is transformed by time-frequency to generate a log-Mel spectrum. The median, root mean square and variance statistics of each time frame are calculated by an adaptive gating module based on third-order statistical indices. The gating vector is generated and fused with the original spectrum in a weighted manner. The audio features are then spliced ​​together to form a multi-channel fused feature. S2. Time-frequency decoupling feature extraction: The multi-channel fused features are divided into time-domain block sequences and frequency-domain block sequences along the time axis and frequency axis, and then input into parallel time-domain branches and frequency-domain branches respectively. The time-domain branch uses a selective state-space model to perform causal scanning on time-domain block sequences to capture long-term steady-state characteristics and trend anomalies, while the frequency-domain branch uses a selective state-space model to perform causal scanning on frequency-domain block sequences to analyze cross-band energy distribution and harmonic structure. S3. Cross-modal feature fusion and anomaly scoring: The feature vectors output from the time domain branch and the frequency domain branch are linearly aligned and then added element by element to obtain the fused feature representation; Based on the fused feature representation, a self-supervised ID classification task combined with the angular interval loss function is used for model training, and the negative log probability of the target ID category is used as the anomaly score during the inference stage.

2. The equipment anomaly monitoring method as described in claim 1, characterized in that, The adaptive gating module in step S1 generates the gating vector through the following process: The median, root mean square, and variance were calculated frame by frame on the transposed log-Mel plot. The three statistics are summed and then subtracted from the preset offset, adjusted by the scaling factor, and activated by the Sigmoid function. The weighted fusion is achieved by multiplying the gated vector element-wise with the original log-Mel spectrum to enhance the weight of time frames.

3. The equipment anomaly monitoring method as described in claim 1, characterized in that, The multi-channel fusion feature mentioned in step S1 is constructed in the following way: The gated enhanced log-Mel spectrum is used as the first channel; The original log-Mel spectrum is retained as the second channel; Supplement the temporal characteristics of the original audio as a third channel; The three-channel features are spliced ​​along the channel dimension to form the final input features.

4. The equipment anomaly monitoring method as described in claim 1, characterized in that, The processing flow of the time-domain branch in step S2 specifically includes: Flatten the temporal block sequence and add token-like and positional embeddings; The process involves layer normalization, linear transformation, two-dimensional convolution, and activation function sequentially. Selective state-space model parameters are dynamically generated based on input features; Capture long-range dependencies through a state-space modeling layer; The output is added to the original input through residual connection.

5. The equipment anomaly monitoring method as described in claim 1, characterized in that, The frequency domain branching process described in step S2 specifically includes: Flatten the frequency domain block sequence and add token-like and position embeddings; The process involves layer normalization, linear transformation, two-dimensional convolution, and activation function sequentially. Selective state-space model parameters are dynamically generated based on input features; Frequency domain structural characteristics are analyzed through state-space modeling layers; The output is added to the original input through residual connection.

6. The equipment anomaly monitoring method as described in claim 1, characterized in that, In step S2, the selective state-space model achieves input-dependent state transitions through dynamic parameter calculation, specifically including: Time step parameters are generated from intermediate features through a linear transformation; The state transition matrix is ​​calculated using time step parameters and preset parameters; The input projection matrix is ​​generated by linear transformation of intermediate features; State-space modeling operations are performed using dynamically generated state transition matrices and input projection matrices.

7. The equipment anomaly monitoring method as described in claim 1, characterized in that, The cross-modal feature fusion described in step S3 specifically includes: Linear projection is performed on the output features of the time-domain branch and the output features of the frequency-domain branch respectively to achieve feature space alignment; The two aligned feature vectors are added element-wise to generate the final fused feature.

8. The equipment anomaly monitoring method as described in claim 1, characterized in that, The self-supervised training process in step S3 specifically includes: A classification task is constructed using the device machine ID as the category label; The model is optimized using the angular interval loss function, which forces the features of different machine IDs to maximize the inter-class distance and minimize the intra-class distance in the angular space. The angle interval loss function introduces an additive angle interval constraint feature distribution.

9. The equipment anomaly monitoring method as described in claim 1, characterized in that, The abnormal score calculation method described in step S3 includes: Input the audio to be tested into the trained model; Extract the predicted probability of its corresponding real machine ID category; The negative natural logarithm of the predicted probability is calculated as the outlier score.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the device anomaly monitoring method based on audio signal analysis as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Method for detecting abnormal sound of self-supervised machine based on domain transfer

    CN115376554A

  • Machine abnormal sound detection method and device for improving AE by combining Transformer

    CN117253505A

  • Unsupervised machine abnormal sound detection method based on lightweight network

    CN117292714A

  • A Self-Supervised Machine Abnormal Sound Detection Method Based on Dual-Path WaveNet

    CN119360884B

  • A multi-class machine abnormal sound detection method based on diffusion model

    CN119626259B