Abnormal sound detection method for decoupling acoustic features by using gradient inversion, program, equipment and storage medium
By introducing a gradient inversion classifier into the abnormal sound detection method, decoupling the domain-related features, the problem that existing methods are difficult to distinguish the differences in different domain characteristics under domain offset conditions is solved, and the robustness of the model and abnormal sound detection performance are improved.
Patent Information
- Application Number
- CN202510263389.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-10
AI Technical Summary
The existing anomaly sound detection method based on self-supervised classification is difficult to distinguish the feature differences between different domains under domain offset conditions, resulting in insufficient abnormal discrimination ability of the model in the target domain.
A gradient inversion classifier is introduced to separate domain-independent features and domain-related features. The gradient of opposite values is taken during the backpropagation process through the gradient inversion layer to decouple domain-related features, thereby improving the performance of the model under domain offset conditions.
By separating domain-independent features and domain-related features, the feature difference between different domains is increased, and the robustness of the model and the performance of abnormal sound detection is improved, especially under domain offset conditions.
Smart Images

Figure CN120126501A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of abnormal sound detection, and particularly relates to a method, program, device and storage medium for detecting abnormal sounds by decoupling acoustic features using gradient reversal. Background Art
[0002] The purpose of anomalous sound detection (ASD) is to automatically identify whether there is an abnormal sound in a target (such as a machine or device), and then determine whether the target has an abnormal behavior or state. The diversity and occasional occurrence of abnormal sounds make it difficult to obtain, and often only normal sounds can be used to train the model.
[0003] With the application of deep learning in the field of audio processing, existing research has provided two methods for realizing anomalous sound detection: unsupervised and self-supervised. The existing unsupervised method is to train feature reconstruction through an autoencoder structure, minimize the reconstruction error to learn the features of normal sounds, and use the reconstruction error as a score to detect anomalies. This method for detecting abnormal sounds can improve the performance of abnormal sound detection, but its false detection rate is relatively high and it is greatly affected by the prior set threshold. The existing self-supervised method introduces the metadata attached to the audio data (such as machine type, machine operating power / speed, environmental noise, etc.) into the modeling process, and learns the feature representation through metadata attribute classification, so as to determine the state (normal / abnormal) of the audio data.
[0004] However, the working conditions of acoustic targets are often complex and changeable, and the audio recording conditions also vary occasionally. Therefore, the state distribution of the training data cannot cover all the working conditions and operating states of the machine. The changes in the working conditions, operating states, and environmental noises of acoustic targets introduce the problem of domain shift for anomalous sound detection. Domain shift refers to the difference in distribution between the data in the source domain and the target domain due to different observation conditions, target states, background noises, etc. When the model trained on the source domain data is tested with the target domain data, problems such as low accuracy and poor stability may occur. The acoustic characteristics of the source domain and the target domain change with the change of metadata attributes, resulting in the model trained with the sounds in the source domain may misidentify anomalies in the target domain. Specifically, there are differences in the source domain and the target domain in terms of operating speed, machine load, viscosity, heating temperature, environmental noise type, etc., and the feature distributions are different. At the same time, the source domain data accounts for the majority in the training data, while the target domain data only accounts for a small part in the training data. The difference in feature distribution and quantity between the two leads to the difficulty of the model in being applicable to the detection of abnormal sounds in the target domain. A specific domain refers to a certain fixed working condition, operating state, and environmental noise condition in the source domain or the target domain.
[0005] The introduction of domain shift problems greatly limits the performance of abnormal sound detection in practice. In the target domain, the sounds of machines are significantly different from those in the source domain due to different operating conditions. This difference causes the model to be unable to directly transfer the knowledge learned in the source domain to the target domain, thus affecting the effect of abnormal sound detection. Domain-related features are those data features that depend on a specific domain or environment, reflecting the specific attributes or background information of the domain, such as the operating mode of the machine, speed, or changes in environmental noise. Domain-agnostic features are those features that remain consistent across different domains, such as the periodic signals of normal operating sounds, the sharp spectral features of abnormal sounds, etc. A domain classifier is a model that discriminates whether the input features come from source domain data or target domain data. When an abnormal sound detection method based on self-supervised classification uses a domain classifier to improve detection performance, it often only focuses on domain-related features through domain classification learning and ignores the influence of domain-agnostic features. Domain-agnostic features (such as the periodic signals of normal operating sounds, the sharp spectral features of abnormal sounds, etc.) often do not change with the domain. The common domain-agnostic features across different domains make the features between different domains overlap and difficult to distinguish. This kind of neglect may weaken the model's ability to distinguish different domains, thereby reducing its ability to distinguish normal sounds from abnormal sounds, and ultimately reducing the performance of abnormal sound detection. Summary of the Invention
[0006] The purpose of the present invention is to solve the problem that existing abnormal sound detection methods based on self-supervised classification ignore the influence of domain-agnostic features, making it difficult to distinguish the feature differences between different domains and resulting in insufficient abnormal discrimination ability in the target domain under domain shift conditions. The present invention provides a different sound detection method, program, device, and storage medium that utilize gradient reversal to decouple acoustic features. By introducing a gradient reversal layer in the abnormal sound detection method based on self-supervised classification, the domain-agnostic features and domain-related features are separated, increasing the feature differences between different domains, and thus improving the performance of the industrial different sound detection system under domain shift conditions.
[0007] A different sound detection method that utilizes gradient reversal to decouple acoustic features includes the following steps:
[0008] Step 1: Obtain the original audio signals of multiple acoustic targets and construct a training set; convert the original audio signals in the training set into Log-Mel spectral features, and split and reorganize the metadata information in the original audio signals to establish a three-layer metadata hierarchical information structure including acoustic target types, domain shift grouping indexes, and attribute group information.
[0009] Step 2: Construct a feature extraction network, including a gradient reversal classifier and a backbone network; use the training set to train the feature extraction network. After the Log-Mel spectral features are input into the backbone network, domain-related feature decoupling is performed under the constraint of the gradient reversal classifier to obtain audio features; during training, the loss function includes the prediction loss of the domain shift grouping index and the attribute group information.
[0010] Step 3: Input each Log-Mel spectral sample corresponding to each attribute group information in the training set into the trained feature extraction network to obtain the corresponding audio features, and take the mean of all audio features as the feature center of this attribute group information.
[0011] Step 4: Obtain the audio signal to be detected, convert it into Log-Mel spectral features and then input it into the trained feature extraction network to obtain audio features; calculate the Mahalanobis distance between the audio features of the audio signal to be detected and the feature center of each attribute group information, and take the minimum value as the anomaly score; compare the anomaly score with the preset anomaly threshold. If the anomaly score exceeds the anomaly threshold, it is determined that the audio signal to be detected is an abnormal signal.
[0012] Further, in step 1, converting the original audio signal into Log-Mel spectral features is specifically as follows: the original audio signal x undergoes a short-time Fourier transform to be converted into spectrogram features; the spectrogram features undergo a filtering process through a Mel filter bank to obtain a Mel spectrogram; the Mel spectrogram is logarithmically scaled to obtain Log-Mel spectral features X;
[0013]
[0014] where, F represents the dimension of Mel filtering, and T represents the number of time frames of the spectral features; represents the Mel filter matrix; ||STFT(x)|| 2 represents taking the power spectrum of the spectral features, thereby ignoring the computational cost brought by the short-time Fourier transform result in the complex domain.
[0015] Further, the three-layer metadata hierarchical information structure in step 1 is specifically as follows:
[0016] The three-layer metadata hierarchical information structure is a three-layer tree structure, with the acoustic target type as the root node, the domain shift grouping index as the intermediate node in the second layer, and the attribute group information as the leaf node in the third layer; the domain shift grouping index refers to the reason for the domain shift under this index, and the attribute group information is used to describe the specific content of the domain shift condition.
[0017] Further, the backbone network in step 2 includes a feature extractor and two 2D convolutional layers; Log-Mel spectrogram features are input to the feature extractor Subsequently, the domain-related features z of the audio are obtained under the constraint of the gradient reversal classifier GRC(·). rev , z rev is input to the first 2D convolutional layer Conv2D sec , and the audio features z related to the domain shift type are obtained sec , z sec is input to the second 2D convolutional layer Conv2D att , and the audio features z related to the specific domain are obtained att .
[0018] Furthermore, in step 2, when training the feature extraction network, the loss function includes the prediction losses of the domain shift grouping index and the attribute group information. Specifically:
[0019] L total = αL rev + βL sec + γL att
[0020] where L total is the overall loss; α, β, and γ are weight coefficients; L rev is the Focal loss; L sec is the cross-entropy loss; L att is the attribute group Focal loss;
[0021]
[0022] where represents predicting the self-supervised label l rev of the attribute group via the gradient reversal classifier GRC(·) based on the decoupled domain-related features z att ; the self-supervised label l att of the attribute group is obtained according to the attribute group information in the three-layer metadata hierarchical information structure of the original audio signal; L Focal (·, ·) represents the Focal loss calculation function, represents the prediction result is the probability of the self-supervised label l att of the attribute group; η is the balance factor used to balance the weights of different classes; is the adjustment factor used to reduce the impact on high-confidence samples and improve robustness;
[0023]
[0024] where represents the linear classifier for predicting the self-supervised label l sec of the domain shift grouping index For the domain offset grouping index self-supervised label l sec The prediction result; the domain offset grouping index self-supervised label l sec Obtained according to the domain offset grouping index in the three-layer metadata hierarchical information structure of the original audio signal; CE(·, ·) represents calculating the cross-entropy loss of two parameters; Represents the prediction result For the domain offset grouping index self-supervised label l sec The probability;
[0025]
[0026] Among them, Represents the linear classifier for predicting the attribute group self-supervised label l att The linear classifier, For the attribute group self-supervised label l att The prediction result.
[0027] Furthermore, the gradient reversal classifier includes a gradient reversal layer. During the backpropagation process, the gradient reversal layer in the gradient reversal classifier takes the opposite value of the gradient learned by the model to reverse the gradient, and transmits the obtained reversed gradient back to the backbone network;
[0028] The gradient before reversal expects the backbone network to learn the feature differences between the source domain and the target domain, that is, to correctly predict the domain label of the input data as much as possible; while the reversed gradient guides the backbone network to optimize in the opposite direction, that is, to eliminate the feature differences between the source domain and the target domain as much as possible and decouple the domain-related features;
[0029] The backpropagation process is expressed as:
[0030]
[0031] Among them, GRL(·) represents the gradient reversal operation function; λ represents the gradient reversal intensity, which gradually increases during the training process; θ rev Represents the learnable parameters of the gradient reversal classifier.
[0032] Furthermore, in step 3, each Log-Mel spectrogram sample corresponding to each attribute group information m in the training set is input into the trained feature extraction network to obtain the corresponding audio feature f i = z att_i , and the mean value c of all audio features is taken m As the feature center of this kind of attribute group information m;
[0033]
[0034] Among them, N mis the number of Log-Mel spectrum samples with attribute group information m in the training set;
[0035] In step 4, the audio signal j to be detected is converted into Log-Mel spectrum features and then input into the trained feature extraction network to obtain audio features f j = z att_j ; Calculate the Mahalanobis distance between the audio features of the audio signal to be detected and the feature centers of each attribute group information, and take the minimum value as the anomaly score
[0036]
[0037] where M represents the total number of attribute group information in the training set; the covariance matrix ∑ is obtained from the features of all audios in the m-th attribute group under the same domain shift grouping index; ∑ -1 represents the inverse matrix of the covariance matrix ∑.
[0038] A computer device / system includes a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps of the above abnormal sound detection method using gradient reversal to decouple acoustic features.
[0039] A computer-readable storage medium stores a computer program / instructions. When the computer program / instructions are executed by a processor, the steps of the above abnormal sound detection method using gradient reversal to decouple acoustic features are implemented.
[0040] A computer program product includes a computer program / instructions. When the computer program / instructions are executed by a processor, the steps of the above abnormal sound detection method using gradient reversal to decouple acoustic features are implemented.
[0041] The beneficial effects of the present invention are as follows:
[0042] In the abnormal sound detection method based on self-supervised classification, the present invention introduces a gradient reversal classifier, uses the gradient reversal classifier to separate domain-independent features and domain-related features, increases the feature difference between different domains, and improves the model robustness. The present invention can alleviate the influence of domain-independent features, effectively distinguish the feature differences between different domains, and improve the performance of abnormal sound detection under domain shift conditions. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 is the overall technical roadmap of the present invention.
[0044] Figure 2 is a schematic diagram for explaining the hierarchical metadata information structure and domain-independent features and domain-related features in the present invention.
[0045] Among them, (a) is the structural diagram of the hierarchical metadata information structure in the present invention, and (b) is the conceptual schematic diagram of the audio feature latent space in the present invention, showing the relationship between domain-related features and domain-independent features.
[0046] Figure 3 It is the structural diagram of the backbone network in the present invention.
[0047] Figure 4 It is the structural diagram of the gradient reversal classifier GRC(·) in the present invention.
[0048] Figure 5 It is the performance comparison data table of the present invention and the existing abnormal sound detection methods under the domain shift condition with AUC, pAUC, AUC-s, AUC-t, and HAUC as evaluation indicators. Specific implementation manner
[0049] The present invention will be further described below with reference to the accompanying drawings.
[0050] An abnormal sound detection method using gradient reversal to decouple acoustic features includes the following steps:
[0051] Step 1: Obtain the original audio signals of multiple acoustic targets and construct a training set; convert the original audio signals in the training set into Log-Mel spectral features, and split and reorganize the metadata information in the original audio signals to establish a three-layer metadata hierarchical information structure including acoustic target types, domain shift grouping indexes, and attribute group information;
[0052] The original audio single-channel signal corresponding to the acoustic target is where 1 represents the channel dimension corresponding to the single channel, L represents the number of sampling points of the audio digital signal, reflecting the duration of the original audio, and the sampling frequency of the original acoustic signal is 16 kHz.
[0053] First, the original audio signal x undergoes a short-time Fourier transform to be converted into a spectrogram feature. The spectrogram contains all the frequency band ranges specified by the sampling frequency. The time window of the short-time Fourier transform in the calculation process is 1024 sampling points (i.e., 64 ms), and the overlap rate between adjacent time windows is 50%, that is, the offset step of the time window is 512 sampling points (32 ms).
[0054] Subsequently, the spectrogram feature undergoes a filtering process through a Mel filter bank to obtain a Mel spectrogram. To amplify the interval sensitive to human auditory perception in the spectral features, logarithmic scaling is performed on the Mel spectrogram to obtain Log-Mel spectral features. The overall calculation process of the spectral features can be summarized as follows:
[0055]
[0056] Among them, X represents the Log-Mel spectral features, F represents the dimension of the Mel filter, and T represents the number of time frames of the spectral features. represents the Mel filter matrix, whose dimension is 128-dimensional. ||STFT(x)|| 2 represents taking the power spectrum of the spectral features, thus ignoring the computational cost brought by the short-time Fourier transform result in the complex domain.
[0057] The metadata information is the text information carried by the original audio of the acoustic target, which includes the acoustic target type, the serial number of each group of domain shift conditions (domain shift grouping index), and the specific content of each group of domain shift conditions. The acoustic target type refers to the main sound-producing machine category in the audio, such as toy cars, fans, etc. The domain shift condition refers to the specific reason for causing the domain shift, such as the serial number, speed, microphone change, or background noise type of the toy car. The metadata representing the same domain shift condition is divided into the same attribute group. The attribute groups of the source domain and the target domain are different, and the attribute groups under the same source domain or target domain are also different. Different domain shift grouping indexes indicate different reasons for causing the domain shift.
[0058] Figure 2 In (a), it shows the structure diagram of the hierarchical metadata information structure. This structure is a three-layer tree structure, with the acoustic target type as the root node, the second-layer domain shift grouping index as the intermediate node, and the third-layer attribute group information as the leaf node. The domain shift grouping index abstractly refers to the reason for causing the domain shift under this index, and the attribute group information details the specific content of the domain shift condition. These two layers of information contain the acoustic feature changes caused by the domain shift.
[0059] Step 2: Construct a feature extraction network, including a gradient reversal classifier and a backbone network; use the training set to train the feature extraction network. After the Log-Mel spectral features are input into the backbone network, domain-related feature decoupling is performed under the constraint of the gradient reversal classifier to obtain audio features; during training, the loss function includes the prediction losses of the domain shift grouping index and the attribute group information;
[0060] (1) Hierarchical extraction of audio features
[0061] Input the Log-Mel spectral features into the backbone network of the model for processing. The backbone network of the model contains a feature extractor Two 2D convolutional layers. The Log-Mel spectral features are input into the feature extractor and then, under the constraint of the gradient reversal classifier GRC(·), the domain-related features z of the audio are obtained rev , and then z rev is input into a 2D convolutional layer Conv2D sec , and the audio features zs related to the domain shift type are obtained ec , and finally zsec Input a 2D convolutional layer Conv2D att to obtain audio features z related to a specific domain att . The above feature extraction process can be expressed by the following formula:
[0062]
[0063] z sec = Conv2D sec (z rev )
[0064] z att = Conv2D att (z sec )
[0065] Figure 2 In (b), the concept of the latent space of audio data is shown. The latent space contains not only domain-related features but also domain-unrelated features. Since existing domain classifiers often ignore the influence of domain-unrelated features, it is difficult to distinguish the feature differences between different domains and ultimately reduces the performance of abnormal sound detection. Under the constraint of the gradient reversal classifier GRC(·), the domain-unrelated features and domain-related features can be separated, increasing the feature differences between different domains, thereby improving the performance of abnormal sound detection under domain shift conditions. Figure 3 shows the structure of the backbone network in the model.
[0066] (2) Construction of the gradient reversal classifier
[0067] Based on the decoupled domain-related features z rev , the gradient reversal classifier predicts the self-supervised label l of the attribute group att . This prediction process can be expressed as follows:
[0068]
[0069] where GRC(·) represents the gradient reversal classifier used to predict the self-supervised label l of the attribute group through the domain-related features z rev , and the prediction result is denoted as att . Figure 4 shows the structure of the gradient reversal classifier GRC(·), which includes a gradient reversal layer and three modules composed of a linear unit layer, a batch normalization layer, and a ReLU activation function. The structures of the two linear classifiers are the same, both composed of a linear unit layer and a softmax activation function.
[0070] (3) Loss calculation of the gradient reversal classifier
[0071] Due to the imbalance of audio samples in different attribute groups, the categories of the self-supervised labels l of the attribute groups att are unbalanced. The present invention introduces the Focal loss L rev to constrain the process of learning domain-related audio features, so as to eliminate the influence of class imbalance on the model's classification of the self-supervised labels of the attribute groups. Its formula is as follows:
[0072]
[0073] where L Focal (·, ·) represents the Focal loss calculation function, represents the prediction result is the probability of the self-supervised label l of the attribute group, η is the balancing factor used to balance the weights of different classes, att and γ is the adjustment factor used to reduce the influence on high-confidence samples and improve robustness. When optimizing by backpropagation according to the loss, the gradient reversal classifier is different from the general backpropagation optimization process.
[0074] During the backpropagation process, the gradient reversal layer (GradientReversalLayer, GRL) in the gradient reversal classifier GRC(·) reverses the gradient learned by the model, and then transmits the obtained reversed gradient back to the backbone network. The gradient before reversal expects the backbone network to learn the feature differences between the source domain and the target domain, that is, to correctly predict the domain label of the input data as much as possible; while the reversed gradient guides the backbone network to optimize in the opposite direction, that is, to eliminate the feature differences between the source domain and the target domain as much as possible and decouple the domain-related features.
[0075] The gradient reversal layer changes the direction of the gradient during the backpropagation process, making the optimization goal of the backbone network opposite to the original classification optimization goal, thereby realizing the adversarial constraint on the domain-related features and improving the robustness of the feature representation between different domains. This mechanism ensures that the backbone network can weaken the differences between the source domain and the target domain while retaining the key information helpful for abnormal sound detection when extracting features. In this way, the gradient reversal classifier enables the model to not only focus on the universal features independent of the domain but also avoid affecting the detection performance on the target domain due to overfitting the domain-related features. Through the adversarial learning between the gradient reversal classifier GRC(·) and the backbone network, we constrain the backbone network to extract decoupled domain-related features from the source domain and the target domain. Its backpropagation process can be expressed by the following formula:
[0076]
[0077] where GRL(·) represents the gradient reversal operation function, λ represents the gradient reversal intensity and gradually increases during the training process, and θ rev Denote the learnable parameters of the gradient reversal classifier.
[0078] (4) Use the classifier to predict the self-supervised labels
[0079] While optimizing the gradient reversal classifier according to the loss, based on the decoupled domain-related feature z rev And the feature z related to the domain shift type extracted by two 2D convolutional layers sec And the feature z related to a specific domain att , two linear classifiers are used to predict the two-layer self-supervised labels. The above label prediction process can be expressed by the following formula:
[0080]
[0081] Where, Denote the linear classifier for predicting the self-supervised label l of the domain shift grouping index through the feature z sec , and the prediction result is denoted as sec Denote the linear classifier for predicting the self-supervised label l of the attribute group through the feature z , and the prediction result is denoted as att , and the prediction result is denoted as att , and the prediction result is denoted as
[0082] (5) Calculate the overall loss through the label prediction results
[0083] In addition to the aforementioned introduction of the Focal loss L rev To constrain the process of learning domain-related audio features, the domain shift grouping index cross-entropy loss L sec And the attribute group Focal loss L att Are also introduced to constrain the processes of learning features related to the domain shift type and features related to a specific domain. It should be noted that considering the imbalance of audio samples in different attribute groups, the present invention replaces the cross-entropy of the attribute group classification with the Focal loss to eliminate the influence of class imbalance. The domain shift grouping index cross-entropy loss and the attribute group Focal loss are calculated as follows:
[0084]
[0085] Where, CE(·, ·) denotes calculating the cross-entropy loss of two parameters, Denote the prediction result Is the probability that the prediction result sec Is the domain shift grouping index l Denote the prediction result Is the probability that the prediction result att Is the self-supervised label l of the attribute group.
[0086] Finally, combine the empirical weights α, β, and γ obtained during training with the Focal loss L rev , the cross-entropy loss L sec and the Focal loss L att as the overall loss L total , and this process can be expressed as the following formula:
[0087] L total = αL rev + βL sec + γL att
[0088] The above prediction label-loss calculation process is the constrained self-supervised classification method proposed by the present invention. In this method, the gradient reversal classifier constrains the domain-related audio features, the domain shift grouping index contains the type characteristics of the domain shift that constrains the feature learning related to the domain shift type, and the attribute group information contains the acoustic features of the domain shift that constrains the feature learning related to a specific domain. The backbone network trained under these three-layer constraints can effectively decouple the domain-related features from the complex audio features.
[0089] Step 3: Input each Log-Mel spectrogram sample corresponding to each attribute group information in the training set into the trained feature extraction network to obtain the corresponding audio features, and take the mean of all audio features as the feature center of this type of attribute group information;
[0090] Suppose there is an attribute group m containing N m training audio segments, and use f i = z att_i to represent the audio features containing attribute group information extracted from the i-th training audio. Then the calculation method of its feature center c m is as follows:
[0091]
[0092] Step 4: Obtain the audio signal j to be detected, convert the audio signal j to be detected into Log-Mel spectrogram features and then input them into the trained feature extraction network to obtain the audio feature f j = z att_j ; Calculate the Mahalanobis distance between the audio features of the audio signal to be detected and the feature centers of each attribute group information, and take the minimum value as the anomaly score
[0093]
[0094] Among them, M represents the total number of attribute group information in the training set; the covariance matrix ∑ is obtained from the features of all audios of the m-th attribute group under the same domain shift grouping index; ∑ -1 represents the inverse matrix of the covariance matrix ∑.
[0095] Because the attribute group feature center c m is the average of the domain-related features of the same-attribute audio, and the domain-related features of the audio are extracted under the constraints of the gradient reversal classifier that can decouple the domain-related features and the two layers of domain shift grouping index and attribute group information, which contain the domain shift type characteristics and acoustic feature constraints. Therefore, the attribute group feature center c m contains the domain shift-related acoustic features of the group of normal sounds. Under the domain shift condition, the attribute group feature center c m is very consistent with the distribution of the normal sound feature center.
[0096] For a certain type of specific acoustic target k, the anomaly score obtained by the above calculation for its corresponding original audio waveform is denoted as Then the processing of this acoustic target by the anomaly judgment mechanism can be expressed as follows:
[0097]
[0098] where represents the mathematical expression form of the anomaly judgment mechanism and is a binary function. θ is the anomaly threshold applicable to different acoustic target types selected by the anomaly judgment mechanism through the overall distribution of the training data. When the anomaly score exceeds θ, it can be considered that this acoustic target exceeds the distribution range of the normal audio features and is judged as abnormal; otherwise, it is considered that this acoustic target is within the distribution range of the normal audio features and is judged as normal.
[0099] Example 1: Anomaly sound detection performance under domain shift conditions:
[0100] The anomaly sound detection method using gradient reversal to decouple acoustic features proposed by the present invention separates domain-independent features and domain-related features, increases the feature difference between different domains, and improves the performance of the current abnormal sound detection method under domain shift conditions. Figure 5 Shows the performance comparison of the abnormal sound detection method proposed by the present invention and the existing abnormal sound detection methods under domain shift conditions with AUC, pAUC, AUC-s, AUC-t, and HAUC as evaluation indicators. Among them, AUC-s refers to the AUC value calculated on the source domain, and AUC-t refers to the AUC value calculated on the target domain.
[0101] From Figure 5 it can be seen that the abnormal sound detection method using gradient reversal to decouple acoustic features proposed by the present invention (the GRHD method described above) has an absolute advantage over other methods under domain shift conditions and ensures high accuracy for different acoustic targets.
[0102] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for detecting abnormal sound by using gradient inversion decoupling acoustic features, characterized in that: The following steps are involved: Step 1: Obtain the original audio signals of various acoustic targets and construct a training set; convert the original audio signals in the training set into Log-Mel spectrum features, split and reorganize the metadata information in the original audio signals, and establish a three-layer metadata hierarchical information structure containing acoustic target type, domain offset grouping index and attribute group information; Step 2: Construct a feature extraction network, including a gradient reversal classifier and a backbone network. Use the training set to train the feature extraction network. After the Log-Mel spectrum features are input into the backbone network, domain-related features are decoupled under the constraints of the gradient reversal classifier to obtain audio features. During training, the loss function includes the prediction loss of the domain offset grouping index and attribute group information; Step 3: Input the Log-Mel spectrum samples corresponding to each attribute group information in the training set into the trained feature extraction network to obtain the corresponding audio features, and take the mean of all audio features as the feature center of this attribute group information; Step 4: Obtain the audio signal to be detected, convert it into Log-Mel spectrum features, and then input it into the trained feature extraction network to obtain audio features; The Mahalanobis distance between the audio features of the audio signal to be detected and the feature center of each attribute group information is calculated, and the minimum value is taken as the anomaly score; the anomaly score is compared with the preset anomaly threshold. If the anomaly score exceeds the anomaly threshold, the audio signal to be detected is determined to be an anomaly signal.
2. The method for detecting abnormal sound using gradient inversion decoupling acoustic features according to claim 1, characterized in that: In step 1, the original audio signal is converted into a Log-Mel spectrum feature, specifically: the original audio signal x is converted into a spectrum feature through a short-time Fourier transform; the spectrum feature is filtered through a Mel filter group to obtain a Mel spectrum; the Mel spectrum is logarithmically scaled to obtain a Log-Mel spectrum feature X; in, F represents the dimension of Mel filtering, and T represents the number of time frames of spectral features; represents the Mel filter matrix; ||STFT(x)|| 2 It means taking the power spectrum of the spectral feature, thus ignoring the computational cost brought by the short-time Fourier transform result in the complex domain.
3. The method for detecting abnormal sound using gradient inversion decoupling acoustic features according to claim 1, characterized in that: The three-layer metadata hierarchical information structure in step 1 is specifically: The three-layer metadata hierarchical information structure is a three-layer tree structure, with the acoustic target type as the root node, the domain offset grouping index in the second layer as the intermediate node, and the attribute group information in the third layer as the leaf node; the domain offset grouping index refers to the cause of the domain offset under the index, and the attribute group information is used to describe the specific content of the domain offset condition.
4. The method for detecting abnormal sound using gradient inversion decoupling acoustic features according to claim 1, characterized in that: The backbone network in step 2 includes a feature extractor and two 2D convolutional layers; Log-Mel spectrum feature input feature extractor Then, the domain-dependent feature z of the audio is obtained under the constraint of the gradient reversal classifier GRC(·) rev , z rev Input to the first 2D convolutional layer Conv2D sec , get the audio feature z related to the domain shift type sec , z sec Input to the second 2D convolutional layer Conv2D att , get the audio features z related to the specific domain att .
5. The method for detecting abnormal sound using gradient inversion decoupling acoustic features according to claim 4, characterized in that: In step 2, the feature extraction network is trained, and the loss function includes the prediction loss of the domain offset grouping index and the attribute group information, specifically: L total =αL rev +βL sec +γL att Among them, L total is the overall loss; α, β, γ are weight coefficients; L rev is the focal loss; L sec is the cross entropy loss; L att Focal loss for attribute group; in, Represents the domain-related features z based on the decoupled rev , the attribute group self-supervised label l via the gradient reversal classifier GRC(·) att Make predictions; attribute group self-supervised label l att Acquire the attribute group information in the three-layer metadata hierarchical information structure of the original audio signal; L Focal (·,·) represents the Focal loss calculation function, Represents the prediction results is the attribute group self-supervisory label l att The probability of; η is the balance factor, which is used to balance the weights of different categories; It is an adjustment factor used to reduce the impact on high-confidence samples and improve robustness; in, represents the self-supervised label l for predicting domain shift group index sec The linear classifier of Index the self-supervised labels l for the domain shift grouping sec The prediction results of the domain offset group index self-supervised label l sec The domain offset grouping index is obtained according to the three-layer metadata hierarchical information structure of the original audio signal; CE(·,·) represents the calculation of the cross entropy loss of two parameters; Represents the prediction results Index the self-supervised labels l for the domain shift grouping sec probability; in, Represents the self-supervised label l used to predict the attribute group att Linear classifier, is the attribute group self-supervisory label l att prediction results.
6. The method for detecting abnormal sound using gradient inversion decoupling acoustic features according to claim 5, characterized in that: The gradient reversal classifier includes a gradient reversal layer. During the back propagation process, the gradient reversal layer in the gradient reversal classifier takes the opposite value of the gradient learned by the model to reverse the gradient, and transmits the obtained reversed gradient back to the backbone network; The gradient before inversion expects the backbone network to learn the feature differences between the source domain and the target domain, that is, to predict the domain label of the input data as correctly as possible; while the gradient after inversion guides the backbone network to optimize in the opposite direction, that is, to eliminate the feature differences between the source domain and the target domain as much as possible and decouple domain-related features; The back propagation process is expressed as: Among them, GRL(·) represents the gradient reversal operation function; λ represents the gradient reversal strength, which gradually increases during the training process; θ rev Represents the learnable parameters of the gradient reversal classifier.
7. The method for detecting abnormal sound using gradient inversion decoupling acoustic features according to claim 4, characterized in that: In step 3, each Log-Mel spectrum sample corresponding to each attribute group information m in the training set is input into the trained feature extraction network to obtain the corresponding audio feature f i =z att_i , take the mean c of all audio features m As the feature center of the attribute group information m; Among them, N m is the number of Log-Mel spectrum samples with attribute group information m in the training set; In step 4, the audio signal j to be detected is converted into Log-Mel spectrum features and then input into the trained feature extraction network to obtain the audio feature f j =z att_j ; Calculate the Mahalanobis distance between the audio feature of the audio signal to be detected and the feature center of each attribute group information, and take the minimum value as the anomaly score Where M represents the total number of attribute group information in the training set; the covariance matrix ∑ is obtained by the features of all audios of the mth attribute group under the same domain offset grouping index; ∑ -1 represents the inverse matrix of the covariance matrix ∑.
8. A computer device / equipment / system comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 7.
9. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer program product comprising a computer program / instructions, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Abnormal sound detection method for restraining self-supervised classification by utilizing metadata hierarchical information
CN117079668A