An Audio Classification Method and Device Based on Feature Decoupling and Contrastive Learning
Through the audio classification method of feature decoupling and contrast learning, an audio classification model ACM is constructed, and the Wav2Vec 2.0 model and a supervised comparison learning loss function with masking mechanism is used to solve the misjudgment and generalization of the speech event classification model, which improves the classification accuracy and robustness.
Patent Information
- Application Number
- CN202411131516.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-18
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-08-18
AI Technical Summary
Existing voice event classification models are susceptible to interference from other information in speech, resulting in misjudgment and reducing model generalization, especially in the problem of category imbalance, the accuracy of category recognition is low.
Using feature decoupling and contrast learning methods, the audio classification model ACM is constructed, including feature extraction module AFEM, reconstruction decoupling module RFDM and contrast classification module CLCM, combined with the variational fitting module VDFM, audio features are extracted using the Wav2Vec 2.0 model, and optimized through alternate training and a supervised comparison learning loss function with masking mechanism.
显著提升了模型的分类准确率和鲁棒性,减少了信息丢失,增强了模型的泛化能力,动态平衡了损失函数的权重分布以优化训练效果。
Smart Images

Figure CN119132331B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech event classification, and in particular, to an audio classification method and device based on feature decoupling and contrast learning. Background Art
[0002] Speech event classification is to process the sound signals in the collected audio data, analyze their acoustic features, identify the corresponding event categories, and at the same time convert them into symbolic descriptions corresponding to the events in the acoustic environment, so as to perceive and understand the information existing in the surrounding environment. It is an important branch in the field of speech processing research and is widely used in audio analysis tasks in different fields.
[0003] In the prior art, the speech event classification framework usually uses an end-to-end model combined with cross-entropy loss for training. However, the cross-entropy loss is easily interfered by other information in the speech, resulting in misjudgment of some event categories by the model, thus affecting the overall classification accuracy of the model. Moreover, the cross-entropy loss focuses on the classification boundary, often ignoring the global representation structure, making it difficult to obtain a robust representation and greatly reducing the generalization ability of the model. In addition, when dealing with the problem of class imbalance, it is easily affected by the categories with higher prediction frequencies, resulting in low recognition accuracy for the categories with fewer samples.
[0004] Therefore, in the related art, there is an urgent need for a method that can improve the accuracy and robustness of speech event classification. Summary of the Invention
[0005] The present invention provides an audio event classification method and device based on feature decoupling and contrast learning for the problems existing in the prior art.
[0006] To achieve the above invention purposes, the present invention provides the following technical solutions:
[0007] S1: Collect the audio files and their corresponding categories, and use the preprocessed audio waveform data as the input data of the model while generating corresponding classification labels to construct an audio classification data set;
[0008] S2: Construct an audio classification model ACM (Audio Classification Module), and optimize the parameters of the model ACM through an alternative training method;
[0009] The audio classification model ACM mainly includes a feature extraction module AFEM (Audio Feature Extraction Module), a reconstruction and disentanglement module RFDM (Reconstruction & Feature Disentanglement Module), and a contrastive classification module CLCM (Contrastive Learning & Classification Module):
[0010] The feature extraction module AFEM respectively obtains the target coarse-grained information and non-target coarse-grained information in the audio representation according to the model input data in step 1;
[0011] The reconstruction and disentanglement module RFDM combines reconstruction based on the upper bound of the mutual information between the target coarse-grained information and the non-target coarse-grained information to achieve leak-free disentanglement between the coarse-grained information, completely separating it in the semantic space, and completing the transformation of the refinement of the information disentanglement granularity;
[0012] The contrastive classification module CLCM completes high-precision classification based on supervised contrastive learning with a Mask mechanism combined with cross-entropy on the basis of learning the robust representation of the decoupled target fine-grained information;
[0013] S3. Construct an audio variational distribution fitting module VDFM (Variational Distribution Fitting Module), fit the variational probability distribution according to the target coarse-grained information and the non-target coarse-grained information, and obtain an accurate upper bound of the mutual information;
[0014] S4. Input the audio to be classified into the trained audio classification model ACM, and the corresponding category of the audio can be obtained.
[0015] As a preferred solution of the present invention, an audio classification method based on feature disentanglement and contrastive learning is characterized in that the following processes are included in step 1:
[0016] S11: Convert the audio signal into variable-length time-series waveform data at a sampling rate of 16 kHz, and determine the corresponding fixed duration according to the audio set to ensure that it can cover 90% of the length of the audio; in addition, filter the audio with a duration less than 10% of the fixed duration to ensure that the audio has sufficient semantic information;
[0017] S12: Compensate for the duration of the audio with insufficient duration by adding silence; when the audio with too long duration is used as the training set, fix the audio duration by random sampling, and when it is used as the test set, fix the audio duration by intercepting from the beginning.
[0018] As a preferred embodiment of the present invention, an audio classification method based on feature decoupling and contrast learning is characterized in that, in step 2, the feature extraction module AFEM includes a Wav2Vec 2.0 module, a target information extraction module TIEM (Target Information Extraction Module), and a non-target information extraction module NTIEM (Non-Target Information Extraction Module);
[0019] The audio is preprocessed to obtain audio waveform data a = {a1, a2, a3, …, a T}, and the audio representation I is extracted through the Wav2Vec 2.0 of the 960-hour pre-trained model of the Librispeech dataset t = Wav2vec2(a t ), and global average pooling is performed on it in the time series dimension to obtain the compressed audio representation I = GlobalAveragePooling(I t , dim = 1); On this basis, the target information coarse-grained representation t = TIEM(I) and the non-target information coarse-grained representation n = NTIEM(I) are preliminarily extracted through the target information extraction module TIEM and the non-target information extraction module NTIEM respectively.
[0020] As a preferred embodiment of the present invention, an audio classification method based on feature decoupling and contrast learning is characterized in that, in step 3, the variational fitting module VDFM includes a variational distribution average module AM (Average Module) and a variational distribution standard deviation module SDM (Standard Deviation Module), which are mainly composed of fully connected layers, ReLU, and Tanh activation functions;
[0021] This module assumes that the variational distribution q follows a Gaussian distribution, and the variational distribution is iteratively updated through the variational distribution average module AM and the variational distribution standard deviation module SDM to ensure that after using the variational distribution to replace the conditional distribution, the upper bound of the mutual information can still be unbiasedly estimated through the Club algorithm;
[0022] First, the target information coarse-grained representation respectively obtains the mean q of the variational distribution q through the average module AM and the standard deviation module SDM m = AM(t) and the standard deviation q d = SDM(t), where D is the feature dimension;
[0023] Through the variational distribution mean q m 、standard deviation q d and the non-target information coarse-grained representation The corresponding variational distribution is obtained as where θ is the network parameter;
[0024] Maximize its log-likelihood based on the variational distribution That is, minimize loss = -LH to achieve KL(p(t,n)||q θ (t,n)) ≤ KL(p(t)p(n)||q θ (t,n)), so as to ensure that when using the variational fitting module to output the variational distribution q to replace the conditional distribution p, the upper bound I of the mutual information between the coarse-grained representation t of the target information and the coarse-grained representation n of the non-target information calculated by the Club algorithm vCLUB (t,n) remains valid, where p(t,n) is the joint distribution between t and n, p(t) is the true distribution of t, and p(n) is the true distribution of n.
[0025] As a preferred solution of the present invention, an audio classification method based on feature disentanglement and contrast learning is characterized in that the reconstruction and disentanglement module RFDM in step 2 is composed of a disentanglement function FDF (Feature Disentanglement Function) and a reconstruction module RM (Reconstruction Module);
[0026] The disentanglement function FDF is based on the mean q of the variational distribution q m , standard deviation q d and the coarse-grained representation of non-target information , and by constructing and using this as the loss function to optimize the target information extraction module TIEM and the non-target information extraction module NTIEM, so that it can finely disentangle the fine-grained representation t of the target information from the compressed audio representation I fine-turn = TIEM(I) and the fine-grained representation n of non-target information fine-turn = NTIEM(I);
[0027] The reconstruction module RM is mainly composed of a fully connected layer, a BatchNorm layer, and a ReLU activation function. Its input is the concatenated information I concat = concat([t fine-turn ,n fine-turn , dim = -1), and the reconstructed compressed audio representation I rec = RM(I concat ) is obtained through RM, aiming to maximize the retention of audio information and minimize the information lost during the disentanglement process, so as to achieve leak-free disentanglement throughout the process.
[0028] As a preferred solution of the present invention, an audio classification method based on feature decoupling and contrastive learning is characterized in that the variational fitting module VDFM and the reconstruction decoupling module RFDM perform parameter learning in an alternating training manner, specifically as follows:
[0029] First, compress the audio representation I to preliminarily extract features through the target information extraction module TIEM and the non-target information extraction module NTIEM to obtain the target information coarse-grained representation t = TIEM(I) and the non-target information coarse-grained representation n = NTIEM(I);
[0030] Set the hyperparameter step of the loop training rounds. On the premise of freezing the parameters of the target information extraction module TIEM and the non-target information extraction module NTIEM, through the loss function Successively optimize the variational fitting module VDFM in a loop;
[0031] After the variational fitting module VDFM is optimized step times, freeze the variational fitting module VDFM as the upper bound extractor of mutual information, and unfreeze the target information extraction module TIEM and the non-target information extraction module NTIEM. According to the upper bound I of the mutual information vCLUB (t, n) combined with the concatenated information I concat = concat([t, n], dim=-1) input to the reconstruction module RM to obtain the reconstruction loss Jointly optimize the target information extraction module TIEM and the non-target information extraction module NTIEM to obtain the target information fine-grained representation t fine-turn and the non-target information fine-grained representation n fine-turn .
[0032] As a preferred solution of the present invention, an audio classification method based on feature decoupling and contrastive learning is characterized in that in step 2, the contrast classification module CLCM includes a contrast mapping module CMM (Contrastive Mapping Module) and a class classification module CCM (Class Classification Module);
[0033] The contrast mapping module CMM is composed of a fully connected layer and a ReLU activation function. On the basis of obtaining the target information fine-grained representation t fine-turn through feature decoupling, the feature mapping module CMM maps the target information fine-grained representation t fine-turn to the hidden space suitable for supervised contrastive learning training and obtains the corresponding hidden space tensor s = CMM(t fine-turn ); On this basis, perform a normalization operation on s to obtain the normalized tensor This ensures the stability of the loss function calculation, where n is the dimension of the latent space tensor s; after the above operations, the supervised contrastive learning loss function with a Mask mechanism is trained by normalizing the tensor z and its corresponding label, and the specific formula is as follows:
[0034]
[0035] In the formula, Loss sup is the supervised contrastive learning loss function, I is the Batchsize, P(i) is the set of positive samples, that is, samples with the same label, z is the latent space tensor, τ represents the temperature value, and α is the Mask mechanism. The Mask mechanism sets masks for unpaired similar data in the same Batch to ensure that, without data augmentation, supervised contrastive learning is used for clustering in the latent space.
[0036] The input of the category classification module CCM is t fine-turn , and based on this, the probability distribution of the predicted audio category is output and the cross-entropy loss function is constructed therefrom where C represents the number of classification categories, y i is the One-hot label, the probability predicted by the model;
[0037] As a preferred embodiment of the present invention, an audio classification method based on feature decoupling and contrastive learning is characterized in that the contrast classification module CLCM in the method balances the weights between the cross-entropy loss and the supervised contrastive learning loss function through the Uncertainty Loss algorithm;
[0038] First, a variance prediction module VPM (Variance Prediction Module) is constructed to respectively predict the variance values σ CE , σ SUPCON =VPM(), and a multi-task loss function is constructed by fusing the cross-entropy and the supervised contrastive learning loss function with the variance values While optimizing the classification model, the task weights are made adaptive.
[0039] An audio classification device based on feature decoupling and contrastive learning is characterized in that the device includes: an input device, an output device, a power supply, at least one processor, and a memory communicatively connected to the processor; the memory may simultaneously correspond to instructions specified by at least one processor, and the instructions are executed by the at least one executor so that the at least one processor can execute the method according to any one of claims 1 to 8.
[0040] The beneficial effects of the present invention are:
[0041] (1) The present invention utilizes the Wav2Vec 2.0 model to extract rich information contained in audio. While significantly reducing information loss compared to traditional acoustic features, it remarkably improves the classification accuracy of the model.
[0042] (2) The proposed Reconstruction and Feature Decoupling Module (RFDM) in the present invention realizes the decoupling of target information and non-target information in the feature space under the premise of only having a single label through the idea of reconstructing mutual information, reduces the interference of non-target information in the classification process, and effectively improves the classification and recognition accuracy of the model.
[0043] (3) The proposed Contrastive Learning Classification Module (CLCM) in the present invention, while performing classification, makes the audio features cluster in the latent space through a supervised contrastive learning loss function with a masking mechanism, making similar features approach each other and dissimilar features separate from each other, greatly enhancing the generalization of the model.
[0044] (4) The present invention dynamically balances the weight distribution between the cross-entropy loss function and the supervised contrastive learning loss function through the Uncertainty Loss algorithm, enabling the model to focus on the corresponding loss function at different training stages, maximizing the optimization effect of the loss function. Description of the Drawings
[0045] Figure 1 It is a schematic flow chart of an audio classification method based on feature decoupling and contrastive learning described in Embodiment 1 of the present invention.
[0046] Figure 2 It is a schematic diagram of the feature extraction module in an audio classification method based on feature decoupling and contrastive learning described in Embodiment 1 of the present invention.
[0047] Figure 3 It is a schematic diagram of the variational fitting module in an audio classification method based on feature decoupling and contrastive learning described in Embodiment 1 of the present invention.
[0048] Figure 4 It is a schematic diagram of the Reconstruction and Feature Decoupling Module and the Contrastive Learning Classification Module in an audio classification method based on feature decoupling and contrastive learning described in Embodiment 1 of the present invention.
[0049] Figure 5 It is a schematic diagram of an audio classification device based on feature decoupling and contrastive learning described in Embodiment 2 of the present invention. Detailed Embodiments
[0050] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0051] Example 1
[0052] As Figure 1 shown, an audio classification method based on feature decoupling and contrast learning includes the following steps:
[0053] S1: Collect audio files and their corresponding categories, and construct an audio classification dataset;
[0054] S2: Preprocess the audio in the dataset, convert it into audio waveform data as the input data of the model, and generate corresponding classification labels according to the category;
[0055] S3: Construct an audio classification model ACM, which includes a feature extraction module AFEM, a reconstruction decoupling module RFDM, and a contrast classification module CLCM, and input the preprocessed audio features into the audio classification model for forward propagation to obtain target coarse-grained information and non-target coarse-grained information respectively;
[0056] S4: Freeze the audio classification model ACM while constructing a variational fitting module VDFM, and train through the target coarse-grained information and non-target coarse-grained information to make the model input fit the variational upper bound of mutual information;
[0057] S5: Freeze the variational fitting module VDFM, train and optimize the weights of the audio classification model ACM by fusing the cross-entropy loss function, the supervised contrast learning loss function, the variational mutual information upper bound loss function, and the reconstruction function, and finally obtain the optimal weights;
[0058] S6: Load the optimal weights into the audio classification model and input the audio to be classified to obtain the corresponding label of the audio.
[0059] Furthermore, the preprocessing in step S2 specifically includes the following steps:
[0060] S21: Convert the audio signal into variable-length time-series waveform data at a sampling rate of 16 kHz, and determine the corresponding fixed duration according to the audio set to ensure that it can cover 90% of the audio length; in addition, filter out the audio with a duration less than 10% of the fixed duration to ensure that the audio has sufficient semantic information;
[0061] S22: Compensate for the duration of the audio with insufficient duration by adding silence; when the audio with too long duration is used as the training set, fix the audio duration by random sampling, and when it is used as the test set, fix the audio duration by intercepting from the beginning.
[0062] Furthermore, as Figure 2 shown, the feature extraction module AFEM includes a Wav2Vec 2.0 module, a target information extraction module TIEM, and a non-target information extraction module NTIEM;
[0063] The audio is preprocessed to obtain audio waveform data a = {a1, a2, a3, …, a T}, and the audio representation I is extracted by loading the Wav2Vec 2.0 of the 960-hour pre-trained model of the Librispeech dataset t = Wav2vec2(a t ), and global average pooling is performed on it in the time series dimension to obtain the compressed audio representation I = GlobalAveragePooling(I t , dim = 1); On this basis, the target information coarse-grained representation t = TIEM(I) and the non-target information coarse-grained representation n = NTIEM(I) are preliminarily extracted from the compressed audio representation through the target information extraction module TIEM and the non-target information extraction module NTIEM respectively.
[0064] Furthermore, as Figure 3 shown, the variational fitting module VDFM includes a variational distribution average module AM and a variational distribution standard deviation module SDM, which are mainly composed of fully connected layers, ReLU, and Tanh activation functions;
[0065] This module assumes that the variational distribution q satisfies a Gaussian distribution, and the variational distribution is iteratively updated through the variational distribution average module AM and the variational distribution standard deviation module SDM to ensure that after using the variational distribution to replace the conditional distribution, the upper bound of the mutual information can still be unbiasedly estimated by the Club algorithm;
[0066] First, the target information coarse-grained representation respectively obtains the mean q of the variational distribution q through the average module AM and the standard deviation module SDM m = AM(t) and the standard deviation q d = SDM(t), where D is the feature dimension;
[0067] Through the variational distribution mean q m 、standard deviation q d and the non-target information coarse-grained representation the corresponding variational distribution is obtained as where θ is the network parameter;
[0068] Maximize its log-likelihood on the basis of the variational distribution that is, minimize loss = -LH to achieve KL(p(t, n)||q θ (t, n)) ≤ KL(p(t)p(n)||q θ (t, n)), and further ensure that when using the variational distribution q output by the variational fitting module to replace the conditional distribution p, the upper bound of the mutual information I between the target information coarse-grained representation t and the non-target information coarse-grained representation n calculated by the Club algorithmvCLUB (t, n) remains valid, where p(t, n) is the joint distribution between t and n, p(t) is the true distribution of t, and p(n) is the true distribution of n.
[0069] Furthermore, as Figure 4 shown, the reconstruction decoupling module RFDM consists of a decoupling function FDF and a reconstruction module RM;
[0070] The decoupling function FDF is based on the mean q m of the variational distribution q, the standard deviation q d and the coarse-grained representation of non-target information. On this basis, by constructing and using this as the loss function to optimize the target information extraction module TIEM and the non-target information extraction module NTIEM, so that they can finely decouple the target information fine-grained representation t fine-turn = TIEM(I) and the non-target information fine-grained representation n fine-turn = NTIEM(I);
[0071] The reconstruction module RM is mainly composed of a fully connected layer, a BatchNorm layer, and a ReLU activation function. Its input is the concatenated information I concat = concat([t fine-turn , n fine-turn , dim = -1). The reconstructed compressed audio representation I rec = RM(I concat ) is obtained through RM, aiming to maximize the retention of audio information and minimize the information lost during the decoupling process, achieving a leak-free decoupling throughout the process.
[0072] Furthermore, as Figure 4 shown, the variational fitting module VDFM and the reconstruction decoupling module RFDM adopt an alternating training method for parameter learning, specifically as follows:
[0073] First, the compressed audio representation I is initially extracted by the target information extraction module TIEM and the non-target information extraction module NTIEM to obtain the target information coarse-grained representation t = TIEM(I) and the non-target information coarse-grained representation n = NTIEM(I);
[0074] Set the hyperparameter step of the loop training rounds. On the premise of freezing the parameters of the target information extraction module TIEM and the non-target information extraction module NTIEM, the variational fitting module VDFM is optimized iteratively through the loss function ;
[0075] After the variational fitting module VDFM is optimized for step times, freeze the variational fitting module VDFM as the upper bound extractor of mutual information, and unfreeze the target information extraction module TIEM and the non-target information extraction module NTIEM. According to the upper bound I of mutual information extracted by the variational fitting module VDFM vCLUB (t,n) Combine and splice the information I concat =concat([t,n],dim=-1) Input the reconstruction loss obtained by the reconstruction module RM Jointly optimize the target information extraction module TIEM and the non-target information extraction module NTIEM to obtain the fine-grained representation t of the target information fine-turn And the fine-grained representation n of the non-target information fine-turn .
[0076] Furthermore, as Figure 4 shown, the contrast classification module CLCM includes a contrast mapping module CMM and a class classification module CCM;
[0077] The contrast mapping module CMM is composed of a fully connected layer and a ReLU activation function. After feature decoupling, the fine-grained representation t of the target information is obtained fine-turn On the basis of, the feature mapping module CMM maps the fine-grained representation t of the target information fine-turn To the latent space suitable for supervised contrast learning training and obtain the corresponding latent space tensor s = CMM(t fine-turn ); On this basis, perform a normalization operation on s to obtain the normalized tensor To ensure the stability of the loss function calculation, where n is the dimension of the latent space tensor s; After the above operations, train through the normalized tensor z and its corresponding label to form a supervised contrast learning loss function with a Mask mechanism. The specific formula is as follows:
[0078]
[0079] In the formula, Loss sup Is the supervised contrast learning loss function, I is the Batchsize, P(i) is the set of positive samples, that is, samples with the same label, z is the latent space tensor, τ represents the temperature value, and α is the Mask mechanism. The Mask mechanism sets a mask for unpaired similar data in the same Batch to ensure that in the absence of data augmentation, supervised contrast learning is used for clustering in the latent space.
[0080] The input of the class classification module CCM is t fine-turn , And based on this, output the probability distribution of predicting the audio category And thus construct a cross-entropy loss function Where C represents the number of classification categories, y iIt is a One-hot label, The probability predicted by the model;
[0081] Furthermore, as Figure 4 shown, the contrast classification module CLCM balances the weights between the cross-entropy loss and the supervised contrast learning loss function through the Uncertainty Loss algorithm;
[0082] First, a variance prediction module VPM (Variance Prediction Module) is constructed to predict the variance values σ corresponding to the classification task and the clustering task respectively CE , σ SUPCON = VPM(), and a multi-task loss function is constructed by fusing the cross-entropy and the supervised contrast learning loss function with the variance value
[0083] While optimizing the classification model, the task weight is made adaptive.
[0084] Embodiment 2
[0085] As Figure 5 shown, an audio classification device based on feature decoupling and contrast learning, characterized in that it includes an input device, an output device, a power supply, at least one processor, and a memory communicatively connected to the processor; the memory may simultaneously correspond to instructions specified by at least one processor, and the instructions are executed by the at least one executor. The input-output interface includes a display, a keyboard, a mouse, and a USB interface for completing data interaction operations; the power supply may be an external power supply or a rechargeable battery to provide electrical energy for the electronic device.
[0086] Those skilled in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: various media such as a removable storage device, a read-only memory (ROM), a magnetic disk, or an optical disc that can store program codes.
[0087] When the above integrated units of the present invention are implemented in the form of software functional units and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of the present invention, in essence, or the parts that contribute to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes such as removable storage devices, ROMs, magnetic disks, or optical discs.
[0088] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. An audio classification method based on feature decoupling and contrastive learning, characterized in that, It includes the following steps: Step 1: Collect the audio files and their categories, and while using the preprocessed audio waveform data as the model input, generate corresponding classification labels to construct an audio classification dataset; Step 2: Construct an audio classification model ACM, and optimize the parameters of the model ACM through an alternating training method; The audio classification model ACM mainly includes a feature extraction module AFEM, a reconstruction decoupling module RFDM, and a contrast classification module CLCM; The feature extraction module AFEM respectively obtains the target coarse-grained information t = TIEM(I) and non-target coarse-grained information n = NTIEM(I) in the audio representation according to the model input data in Step 1; The reconstruction decoupling module RFDM consists of a decoupling function FDF and a reconstruction module RM; The decoupling function FDF is based on the mean q of the variational distribution q m , the standard deviation q d , and the non-target coarse-grained information . On this basis, by constructing and using this as the loss function to optimize the target information extraction module TIEM and the non-target information extraction module NTIEM, so that it can finely decouple the target information fine-grained representation t from the compressed audio representation I fine-turn = TIEM(I) and the non-target information fine-grained representation n fine-turn = NTIEM(I); The reconstruction module RM is mainly composed of a fully connected layer, a BatchNorm layer, and a ReLU activation function, and its input is the concatenated information I concat = concat([t fine-turn , n fine-turn , dim = -1), and the reconstructed compressed audio representation I is obtained through the reconstruction module RM rec = RM(I concat ), realizing the lossless decoupling between coarse-grained information, completely separating it in the semantic space, and completing the transformation of the refinement of the information decoupling granularity; The contrast classification module CLCM completes high-precision classification based on supervised contrast learning with a Mask mechanism combined with cross-entropy on the basis of learning the robust representation of the target fine-grained information after decoupling; Step 3: Construct an audio variational fitting module VDFM, fit the variational probability distribution according to the target coarse-grained information and non-target coarse-grained information, and obtain an accurate upper bound of mutual information; Step 4: Input the audio to be classified into the trained audio classification model ACM, and the corresponding category of the audio can be obtained.
2. The audio classification method based on feature decoupling and contrast learning according to claim 1, wherein The following process is included in Step 1: S11: Convert the audio signal into variable-length time-series waveform data at a sampling rate of 16 kHz, and determine the corresponding fixed duration according to the audio set to ensure that it can cover 90% of the audio length; in addition, filter the audio with a duration less than 10% of the fixed duration to ensure that the audio has sufficient semantic information; S12: Compensate for the duration of the audio with insufficient duration by adding silence; when the audio with too long duration is used as the training set, fix the audio duration by random sampling, and when it is used as the test set, fix the audio duration by intercepting from the beginning.
3. The audio classification method based on feature decoupling and contrast learning according to claim 1, wherein, In Step 2, the feature extraction module AFEM includes a Wav2Vec 2.0 module, a target information extraction module TIEM, and a non-target information extraction module NTIEM; The audio is preprocessed to obtain audio waveform data a = {a1, a2, a3, …, a T}, and the audio representation I t = Wav2vec2(a t ) is extracted by loading the Wav2Vec 2.0 pre-trained model of the Librispeech dataset for 960 hours, and global average pooling is performed on it in the time series dimension to obtain the compressed audio representation I = GlobalAveragePooling(I t , dim = 1); on this basis, the target information coarse-grained information t = TIEM(I) and the non-target information coarse-grained information n = NTIEM(I) are preliminarily extracted from the compressed audio representation through the target information extraction module TIEM and the non-target information extraction module NTIEM respectively.
4. An audio classification method based on feature decoupling and contrast learning according to claim 1, characterized in that In Step 3, the variational fitting module VDFM includes a variational distribution averaging module AM and a variational distribution standard deviation module SDM, which are mainly composed of fully connected layers, ReLU, and Tanh activation functions; This module assumes that the variational distribution q satisfies a Gaussian distribution, and iteratively updates the variational distribution through the variational distribution averaging module AM and the variational distribution standard deviation module SDM to ensure that after using the variational distribution to replace the conditional distribution, the upper bound of mutual information can still be unbiasedly estimated through the Club algorithm; First, the target information is coarse-grained information The mean q of the variational distribution q is obtained through the average module AM and the standard deviation module SDM respectively m = AM(t) and the standard deviation q d = SDM(t), where D is the feature dimension; Through the mean q of the variational distribution m and the standard deviation q d as well as the coarse-grained information of non-target information the corresponding variational distribution is obtained as where θ is the network parameter; Maximize its log-likelihood based on the variational distribution That is, minimize the loss = -LH to achieve KL(p(t,n)||q θ (t,n)) ≤ KL(p(t)p(n)||q θ (t,n)), and further ensure that when using the variational fitting module to output the variational distribution q to replace the conditional distribution p, the upper bound I vCLUB (t,n) of the mutual information between the coarse-grained representation t of the target information and the coarse-grained representation n of the non-target information calculated by the Club algorithm remains valid, where p(t,n) is the joint distribution between t and n, p(t) is the true distribution of t, and p(n) is the true distribution of n.
5. The audio classification method based on feature decoupling and contrast learning according to claim 4, wherein The variational fitting module VDFM and the reconstruction decoupling module RFDM adopt an alternating training method for parameter learning, specifically as follows: First, compress the audio representation I and preliminarily extract features through the target information extraction module TIEM and the non-target information extraction module NTIEM to obtain the target coarse-grained information t = TIEM(I) and non-target coarse-grained information n = NTIEM(I); Set the hyperparameter step of the cyclic training rounds. On the premise of freezing the parameters of the target information extraction module TIEM and the non-target information extraction module NTIEM, optimize the variational fitting module VDFM iteratively through the loss function successively in a cycle; After the variational fitting module VDFM is optimized for step times, freeze the variational fitting module VDFM as the upper bound extractor of mutual information, and unfreeze the target information extraction module TIEM and the non-target information extraction module NTIEM. According to the upper bound of mutual information I vCLUB (t,n) combined with the reconstructed loss obtained by inputting the concatenated information I concat =concat([t,n],dim=-1) into the reconstruction module RM Jointly optimize the target information extraction module TIEM and the non-target information extraction module NTIEM to obtain the fine-grained representation t fine-turn of the target information and the fine-grained representation n fine-turn of the non-target information.
6. The audio classification method based on feature decoupling and contrastive learning according to claim 1, wherein In Step 2, the contrast classification module CLCM includes a contrast mapping module CMM and an audio classification module CCM; The contrast mapping module CMM consists of a fully connected layer and a ReLU activation function, and obtains a fine-grained representation t of the target information through feature decoupling. fine-turn On the basis of this, the feature mapping module CMM maps the fine-grained representation t of the target information fine-turn to a latent space suitable for supervised contrastive learning training, and obtains the corresponding latent space tensor s = CMM(t fine-turn ); on this basis, a normalization operation is performed on s to obtain a normalized tensor to ensure the stability of the loss function calculation, where n is the dimension of the latent space tensor s; after the above operations, the normalized tensor z and its corresponding label are used to construct a supervised contrastive learning loss function with a Mask mechanism for training, and the specific formula is as follows: where Loss sup is the supervised contrastive learning loss function, I is the Batchsize, P(i) is the set of positive samples, that is, samples with the same label, z is the latent space tensor, τ represents the temperature value, and α is the Mask mechanism. The Mask mechanism sets masks for unpaired similar data in the same Batch to ensure that in the absence of data augmentation, supervised contrastive learning is used to cluster in the latent space; The input of the audio classification module CCM is t fine-turn and based on this, it outputs the probability distribution of predicting the audio category and constructs a cross-entropy loss function from this where C represents the number of classification categories, and y i is the One-hot label, the probability predicted by the model.
7. The audio classification method based on feature decoupling and contrastive learning according to claim 6, wherein The comparison and classification module CLCM balances the weights between the cross-entropy loss and the supervised contrastive learning loss function through the Uncertainty Loss algorithm; First, construct a variance prediction module VPM to predict the variance values σ corresponding to the classification task and the clustering task respectively CE , σ SUPCON = VPM(), and construct a multi-task loss function by fusing the variance value with the cross-entropy and supervised contrastive learning loss functions While optimizing the classification model, it realizes the self-adaptation of task weights.
8. An audio classification device based on feature decoupling and contrastive learning, characterized in that, The device includes: an input device, an output device, a power supply, at least one processor, and a memory communicatively connected to the processor; the memory can simultaneously correspond to instructions specified by at least one processor, and the instructions are executed by the at least one executor so that the at least one processor can implement the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Audio understanding and generating method based on large-scale audio representation language model
CN116741153A
Voice privacy protection method and system based on multi-task adversarial decoupling learning
CN117995198A