A supervised domain adaptive acoustic scene classification method based on multiple devices
By adopting supervised domain adaptive method and frequency band standardization in the sound scene classification technology, the problem of inconsistent data distribution among different devices is solved, and the high accuracy classification and model generalization capabilities are improved on multiple devices.
Patent Information
- Application Number
- CN202310369908.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-10
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2043-04-10
AI Technical Summary
The inconsistent data distribution of existing sound scene classification technologies between different recording devices leads to performance degradation, making it difficult to adapt to multiple devices, especially on low-cost mobile devices.
The supervised domain adaptive method based on multiple devices is adopted, and the linear and nonlinear distortion between devices is reduced through frequency band standardization and supervised domain adaptive acoustic scene classification model, the linear and nonlinear distortion between devices is extracted, and the domain invariant features are improved, and the generalization ability of the model is improved.
High accuracy classification on different devices is achieved, and the generalization ability of the model is significantly improved, making the sound scene classification method suitable for various devices and scenarios.
Smart Images

Figure CN116386599B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of sound scene classification, and specifically refers to a sound scene classification method based on supervised domain adaptation of multiple devices. Background Art
[0002] In the era of the Internet of Everything, acoustic scene classification can be applied in many fields, such as smart city construction, biodiversity monitoring, urban security monitoring, etc. The goal of the acoustic scene classification task is to classify the collected sound signals to be classified according to pre-defined acoustic scene categories, thereby providing target acoustic scene information for many fields. Researchers have conducted many studies on acoustic scene classification technology and collected acoustic scene samples for a variety of different sound pickup devices, trying to apply this technology to various sound pickup devices that have been deployed. However, due to the inconsistent data distribution obtained by different types of recording devices, the problem of device diversity has been raised for acoustic scene classification technology.
[0003] In actual scenarios, this data distribution inconsistency caused by device mismatch causes the trained sound scene classification model to show a significant performance degradation on other devices, making it impossible to apply it to people's lives. Recently, a large number of deep learning-based sound scene classification methods have been applied to classification tasks, and have attempted to solve the impact of device mismatch, which are mainly divided into audio sample data optimization and network model structure optimization. Audio sample data optimization mainly includes data enhancement, frequency band standardization, etc. Although such methods can increase the number of samples or correct some sample differences caused by different pickup devices, due to the complexity of sound scene audio sample data, sound scene audio sample data may contain many overlapping sounds or background noises, so such methods cannot completely make up for the differences brought by devices and have great limitations. The optimization of network model structure mainly includes designing large-scale network structure, deep residual network integrating high and low frequency path separation, network using two-stage classifier for classification, and network actively adjusting regularization coefficient through receptive field. Such methods can better extract key features through network model optimization to improve the generalization ability of the model, but such methods have more model parameters and higher complexity, which is not conducive to application on low-cost mobile devices. Summary of the invention
[0004] The object of the present invention is to provide a supervised domain adaptive sound scene classification method based on multiple devices with simple structure, good classification effect and wide adaptability.
[0005] The technical solution for achieving the above-mentioned purpose includes the following contents.
[0006] A method for acoustic scene classification based on supervised domain adaptation of multiple devices comprises the following steps:
[0007] S1: reading scene audio signals collected by various types of sound pickup devices, and preprocessing the scene audio signals to obtain preprocessed sample data;
[0008] S2: Performing Fourier transform on the sample data obtained in step S1, performing Mel filter processing on the sample data after Fourier transform processing, and then performing frequency band standardization correction, extracting three feature spectrum graphs, and fusing the three feature spectrum graphs to obtain three-dimensional acoustic features;
[0009] S3: inputting the three-dimensional acoustic features obtained in step S2 into a data enhancement module to obtain the three-dimensional acoustic features after data enhancement;
[0010] S4: Construct a supervised domain adaptive acoustic scene classification model;
[0011] Construct an acoustic scene classification model based on CNN model and domain adaptation method;
[0012] The sound scene classification model is composed of several feature-aligned convolutional blocks and fully connected layers;
[0013] The sound scene classification model divides the three-dimensional acoustic features into a source domain and a target domain according to the type of sound pickup device during the training phase. The source domain and the target domain are then respectively passed through the sound scene classification model, and the difference loss between the source domain and the target domain is calculated separately in each feature alignment convolution block to obtain a domain difference loss.
[0014] The feature alignment convolution block compares the output features of the source domain and the target domain, and calculates the domain difference loss;
[0015] The total loss of the acoustic scene classification model is a weighted sum of domain difference loss, source domain loss and target domain loss;
[0016] S5: inputting the three-dimensional acoustic features and their corresponding labels in step S3 into the sound scene classification model in step S4 for supervised training to obtain a trained supervised domain adaptive sound scene classification model;
[0017] S6: Input the scene audio signal to be classified into the supervised domain adaptive sound scene classification model in step S5 to obtain a classification result.
[0018] Furthermore, the preprocessing of the scene audio signal in step S1 includes pre-emphasis, frame division, and windowing.
[0019] Further, in step S2, the extracting of three feature spectrum graphs includes using a Mel filter group, a first-order difference filter group and a second-order difference filter group to extract and obtain a frequency-band-standardized logarithmic Mel feature spectrum graph, a first-order difference feature spectrum graph and a second-order difference feature spectrum graph.
[0020] Furthermore, in step S2, the frequency band normalization is first divided according to the type of device in the sample space of the training set; then the mean and standard deviation are calculated for each frequency band of the acoustic feature spectrum diagram of the device; and the frequency band normalization processing is performed according to the corresponding device type of the input acoustic feature.
[0021] Furthermore, in step S3, the Mixup module and the SpecAugment module are used to perform data enhancement processing on the three-dimensional acoustic features in step S2. The Mixup and SpecAugment data enhancement methods are used together to improve the generalization ability of the model from the perspective of increasing the amount of sample data.
[0022] Further, in step S4, the sound scene classification model is composed of three feature alignment convolution blocks and two fully connected layers.
[0023] Further, in step S4, the feature alignment convolution block includes two convolution layers, two batch normalization layers (Batch-Normalization, hereinafter referred to as BN layer), two activation function layers (Rectifield Linear Unit, hereinafter referred to as Relu layer), and one pooling layer.
[0024] The data distribution differences between different types of sound pickup devices are mainly divided into two parts: linear distortion and nonlinear distortion. Linear distortion can be corrected by frequency band standardization, and nonlinear distortion can be processed by unsupervised domain adaptation methods. Unsupervised domain adaptation methods require large-scale unlabeled data in the target domain to achieve good performance. However, collecting large-scale audio sample data requires a large investment. In the absence of a large amount of data, unsupervised domain adaptation methods cannot extract domain-invariant features well, and the classification effect cannot be guaranteed. In the sound scene classification method based on supervised domain adaptation of multiple devices of the present invention, frequency band standardization is based on the perspective that different types of sound pickup devices have different frequency responses, and the frequency spectrum characteristics of the samples are linearly corrected, thereby reducing the differences between the devices and improving the classification accuracy; the supervised domain adaptive sound scene classification model divides the source domain and the target domain by device type, and reduces the degree of difference between the two domains through domain difference loss during the training process, so that the model can correct the nonlinear distortion between the two domains and can better extract the domain invariant features of the acoustic features, thereby improving the generalization ability of the model; the combined use of frequency band standardization and the supervised domain adaptive sound scene classification model can respectively correct the linear distortion and the nonlinear distortion and thus reduce the differences between the devices, which not only improves the classification accuracy of the sound scene classification, but also generalizes the model to other invisible sound pickup devices, significantly improving the generalization ability of the model, so that the sound scene classification method in the technical solution of the present invention is applicable to various devices and various scenes. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 is a frequency band normalized mean curve diagram in the embodiment;
[0026] Figure 2 is a frequency band normalization standard deviation curve diagram in the embodiment;
[0027] Figure 3 Schematic diagram of a three-dimensional acoustic feature extraction method in an embodiment;
[0028] Figure 4 Schematic diagram of the structure of feature alignment convolutional blocks of a supervised domain adaptive sound scene classification model in an embodiment;
[0029] Figure 5 A schematic diagram of the network structure of a supervised domain adaptive sound scene classification model in an embodiment;
[0030] Figure 6 Schematic diagram of experimental results in the examples. DETAILED DESCRIPTION
[0031] The present invention is specifically described below in conjunction with embodiments, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar modules or modules with the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and should not be construed as limiting the present invention.
[0032] See also Figures 1 to 6 , a method for acoustic scene classification based on supervised domain adaptation of multiple devices, comprising the following steps,
[0033] S1: Reading scene audio signals collected by various types of sound pickup devices, and performing pre-processing such as pre-emphasis, framing, and windowing on the scene audio signals to obtain pre-processed sample information;
[0034] S2: Performing Fourier transform on the sample data obtained in step S1, performing Mel filter processing on the sample data after Fourier transform processing, and then performing frequency band standardization correction, extracting three feature spectrum graphs, and fusing the three feature spectrum graphs to obtain three-dimensional acoustic features;
[0035] The extraction of three characteristic spectrum graphs includes using a Mel filter group, a first-order difference filter group and a second-order difference filter group to extract and obtain a frequency-band-standardized logarithmic Mel characteristic spectrum graph, a first-order difference characteristic spectrum graph and a second-order difference characteristic spectrum graph.
[0036] Frequency band normalization is a method to standardize and correct the Mel frequency band characteristics of various types of pickup devices in the sample space, which can effectively solve the linear distortion problem on various types of pickup devices.
[0037] like Figure 1 and Figure 2 The figure shows the frequency band normalization mean and standard deviation curves in this embodiment. It can be clearly observed that the mean and standard deviation of different devices in the same frequency band show great differences, indicating that the frequency responses of different types of pickup devices are not consistent. The frequency band normalization method can effectively correct the linear distortion problem on different types of pickup devices.
[0038] like Figure 3 This is a schematic diagram of the three-dimensional acoustic feature extraction method in this embodiment. The sample information passes through the Mel filter group, the frequency band normalization module, the first-order difference filter group and the second-order difference filter group in sequence to obtain the logarithmic Mel feature spectrum diagram after frequency band normalization, the first-order difference feature spectrum diagram and the second-order difference feature spectrum diagram. Then the three feature spectrum diagrams are input into the feature fusion module, and the three feature spectrum diagrams are spliced in the channel dimension to finally obtain the spliced and fused three-dimensional acoustic features.
[0039] In step S2, frequency band standardization first divides the sample space of the training set according to the type of equipment, and then calculates the mean and standard deviation of each frequency band of the acoustic feature spectrum of the equipment. The calculation formula is:
[0040]
[0041] Where d is the device category of the training sample, N is the number of training samples, M is the number of time frames of the training sample, and k is the Mel frequency band of the training sample;
[0042] Finally, the frequency band is normalized according to the corresponding device type of the input acoustic characteristics, and the calculation formula is:
[0043]
[0044] where x dnmk is the input acoustic feature spectrum, It is the acoustic feature spectrum after frequency band standardization;
[0045] S3: Input the three-dimensional acoustic features obtained in step S2 into the data enhancement module to obtain acoustic features after data enhancement.
[0046] The data enhancement module of this embodiment is based on the combination of Mixup and SpecAugment methods, and obtains acoustic features after data enhancement by performing time distortion, time masking, frequency masking and hybrid data enhancement methods on the input acoustic features.
[0047] In step S3, the data enhancement method is composed of Mixup and SpecAugment, and its construction method is:
[0048] SpecAugment processing methods include:
[0049] Time warping: Overlay a time spectrogram of any length in the input acoustic feature spectrogram onto an arbitrary time spectrogram of the input acoustic feature spectrogram;
[0050] Time masking: mask the time spectrogram of any length in the input acoustic feature spectrum;
[0051] Frequency masking: mask the frequency spectrum of any length in the input acoustic feature spectrum;
[0052] Mixup's composition methods are:
[0053]
[0054]
[0055] Where i and j are positive integers, λ∈(0,1) and conform to the beta distribution, x i represents the i-th sample of the input acoustic feature, x j represents the jth sample of the input acoustic feature, represents the acoustic features obtained by hybrid enhancement, y i Indicates the label corresponding to the i-th sample of the input acoustic feature, y j Represents the label corresponding to the jth sample of the input acoustic feature, Indicates the label corresponding to the acoustic feature obtained by hybrid enhancement;
[0056] S4: Construct a supervised domain adaptive acoustic scene classification model;
[0057] Construct an acoustic scene classification model based on CNN model and domain adaptation method;
[0058] The acoustic scene classification model consists of three feature-aligned convolutional blocks and two fully connected layers;
[0059] During the training phase, the acoustic scene classification model divides the acoustic features into source domain and target domain according to the type of sound pickup device. Then the source domain and target domain are respectively passed through the acoustic scene classification model, and the difference loss between the source domain and the target domain is calculated separately in each feature alignment convolution block to obtain the domain difference loss.
[0060] The feature alignment convolution block compares the output features of the source domain and the target domain and calculates the domain difference loss;
[0061] The feature alignment convolution block includes: two convolutional layers, two BN layers, two RELU activation layers, and one pooling layer;
[0062] like Figure 4 The figure shows a schematic diagram of the feature alignment convolution block structure of the supervised domain adaptive sound scene classification model in the present embodiment. In the three feature alignment convolution blocks, the convolution kernel step size is set to 1, and the input feature is convolved with the convolution kernel in the convolution layer to obtain the extracted high-order abstract features; in the first feature alignment convolution block, the number of channels of the two convolution layers is 64, the convolution kernel size is 5×5, and the maximum pooling size is 4×4; in the second feature alignment convolution block, the number of channels of the two convolution layers is 128, the convolution kernel size is 3×3, and the average pooling size is 4×4; in the third feature alignment convolution block, the number of channels of the two convolution layers is 256, the convolution kernel size is 3×3, and the average pooling size is 2×2;
[0063] The maximum pooling layer and the average pooling layer in the feature alignment convolution block reduce the feature size by selecting the maximum value and the average value respectively, thereby reducing the parameters of the model.
[0064] The calculation formula of the RELU activation layer in the feature alignment convolution block is:
[0065] After passing through three feature alignment convolution blocks, the input features are first processed by Flatten to flatten the three-dimensional features into one-dimensional features, and then input into two fully connected layers for data merging. Based on the classification results, one-dimensional data of length 10 is obtained, and finally the final classification prediction result is obtained through the softmax layer.
[0066] The calculation formula of the Softmax layer is:
[0067] Among them, i, j are positive integers, z i Indicates the predicted value of the input feature prediction as the i-th category.
[0068] like Figure 5 The diagram shows a network structure diagram of a supervised domain adaptive sound scene classification model in the present embodiment. In the supervised domain adaptive sound scene classification model, in order to reduce the feature difference between the source domain and the target domain, the main device is divided into the source domain according to the device category in the training stage, and the other devices are divided into the target domain. In addition, each round of training uses the source domain samples and the target domain samples to measure the difference between the source domain and the target domain in the supervised domain adaptive sound scene classification model using the mean square error MSE loss function, and calculates the domain difference to obtain the domain difference loss. The domain difference loss plus the source domain and target domain classification loss can obtain the total loss, and the total loss is reversely updated to train the model.
[0069] The MSE loss function calculation formula is:
[0070] Among them, m is the number of elements of the input feature, is the value of the i-th element in the source domain, is the value of the i-th element of the target domain;
[0071] The supervised domain adaptive sound scene classification model has a total of three feature alignment convolution blocks, and their domain difference losses are Loss mse1 , Loss mse2 , Loss mse3 , so the total domain difference loss LOSS ddl =Loss mse1 +Loss mse2 +Loss mse3 ;
[0072] The classification loss of the source domain and the target domain adopts NLL loss, which is calculated as follows:
[0073]
[0074] The total loss of the supervised domain adaptive acoustic scene classification model is the weighted sum of domain difference loss, source domain loss and target domain loss;
[0075] The total loss function is Loss = Loss ddl +Loss t +Loss s
[0076] S5: extracting acoustic features of training samples according to S1-S3 methods, and inputting the acoustic features of the training samples and their corresponding labels into the supervised domain adaptive sound scene classification model in S4 for supervised training, thereby obtaining a trained supervised domain adaptive sound scene classification model;
[0077] S6: extracting acoustic features of the audio samples of the sound scene to be classified according to the S1 and S2 methods, and inputting them into the trained supervised domain adaptive sound scene classification model to obtain the classification result;
[0078] This embodiment is established in an experimental environment with Window 10 system, RTX3060 graphics card, R7-5800H CPU, and 16G memory; pytorch is used as the deep learning framework, and the DCASE2020Task1A multi-device sound scene classification dataset in the DCASE competition is used. The dataset contains a total of ten predefined sound scene categories. According to the official dataset division method, the number of training set samples is 13926, of which the number of main devices A is 10215, and the number of devices B, C, S1, S2, and S3 is approximately 750. The number of test set samples is 2968, of which the number of devices A, B, C, S1, S2, S3, S4, S5, and S6 is approximately 330, and devices S4, S5, and S6 do not appear in the training set.
[0079] like Figure 6 From the experimental results shown, it can be seen that when only the samples of device A in the training set are used for training, the accuracy rate can reach 76.7% on device A in the test set, but the accuracy rate on other untrained devices does not exceed 43%, indicating that the model trained on a single device cannot be generalized to other devices. When all visible devices in the official training set are used for training, the recognition on the main device A is reduced by 8%. However, since more types of devices are involved in the model training, the recognition performance of the model on other devices is greatly improved, and even the model is generalized to invisible devices. This shows that more types of devices participating in the model training can help the model better extract common features or domain-invariant features between different devices.
[0080] After adding the device matching method (frequency band normalization and supervised domain adaptation method) mentioned in this embodiment, the model shows different performances on visible devices and invisible devices. Frequency band normalization achieves an accuracy of 77% on the main device A, which is equivalent to the effect of training device A alone, and improves the accuracy of the remaining visible devices B-S3 by at least 3%, but is affected to a certain extent in the invisible devices S4-S6. This shows that frequency band normalization is mainly a linear correction of the frequency axis of the visible device, which can effectively improve the recognition accuracy of the visible device, and the average classification accuracy on the test set is increased by 2%.
[0081] At the same time, the supervised domain adaptation method improved the accuracy by 4% on device A and achieved certain improvements on other devices, especially on invisible devices S4-S6. This shows that during the training process, the supervised domain adaptation method actively aligns the features of the source domain and the target domain, enabling the model to better extract domain-invariant features, enhance the classification accuracy and generalization ability of the model, and improve the average classification accuracy by 4% on the test set.
[0082] This embodiment aims to solve the technical problem that the audio data collected by various types of sound pickup devices have device mismatch problems, resulting in linear and nonlinear distortions that affect the classification accuracy and generalization ability of the model, and proposes a sound processing method that integrates frequency band standardization and supervised domain adaptation methods. The present invention can effectively reduce the impact of device mismatch, improve the performance of the model on different sound pickup devices, and has broad application prospects.
[0083] The above is only a preferred embodiment of the present invention. This embodiment uses sample data of six visible sound pickup devices for processing. In other embodiments, more or fewer visible sound pickup devices can also be used. As long as the sound processing method and technical features of the present invention are met, they are within the scope of protection of the present invention. A person of ordinary skill in the art can understand and implement all or part of the processes of the above embodiment, and the equivalent changes made according to the requirements of the present invention are still within the scope of the invention.
Claims
1. A supervised domain adaptive acoustic scene classification method based on multiple devices, It is characterized in that The following steps are included: S1: reading scene audio signals collected by various types of sound pickup devices, and preprocessing the scene audio signals to obtain preprocessed sample data; S2: Performing Fourier transform on the sample data obtained in step S1, performing Mel filter processing on the sample data after Fourier transform processing, and then performing frequency band standardization correction, extracting three feature spectrum graphs, and fusing the three feature spectrum graphs to obtain three-dimensional acoustic features; S3: inputting the three-dimensional acoustic features obtained in step S2 into a data enhancement module to obtain the three-dimensional acoustic features after data enhancement; S4: Building a supervised domain-adaptive acoustic scene classification model, Construct an acoustic scene classification model based on CNN model and domain adaptation method; The sound scene classification model is composed of several feature-aligned convolutional blocks and fully connected layers; The sound scene classification model divides the three-dimensional acoustic features into a source domain and a target domain according to the type of sound pickup device during the training phase. The source domain and the target domain are then respectively passed through the sound scene classification model, and the difference loss between the source domain and the target domain is calculated separately in each feature alignment convolution block to obtain a domain difference loss. The feature alignment convolution block compares the output features of the source domain and the target domain, and calculates the domain difference loss; The total loss of the acoustic scene classification model is a weighted sum of domain difference loss, source domain loss and target domain loss; S5: inputting the three-dimensional acoustic features and their corresponding labels in step S3 into the sound scene classification model in step S4 for supervised training to obtain a trained supervised domain adaptive sound scene classification model; S6: Input the scene audio signal to be classified into the supervised domain adaptive sound scene classification model in step S5 to obtain a classification result.
2. The method for sound scene classification based on supervised domain adaptation of multiple devices according to claim 1, It is characterized in that The preprocessing of the scene audio signal in step S1 includes pre-emphasis, frame division, and windowing.
3. The method for sound scene classification based on supervised domain adaptation of multiple devices according to claim 1, It is characterized in that In step S2, the extracting of three characteristic spectrum graphs includes using a Mel filter group, a first-order difference filter group and a second-order difference filter group to extract and obtain a frequency-band-standardized logarithmic Mel characteristic spectrum graph, a first-order difference characteristic spectrum graph and a second-order difference characteristic spectrum graph.
4. The method for sound scene classification based on supervised domain adaptation of multiple devices according to claim 1, It is characterized in that In step S2, the frequency band normalization is first divided according to the type of device in the sample space of the training set; then the mean and standard deviation are calculated for each frequency band of the acoustic feature spectrum of the device; and the frequency band normalization processing is performed according to the corresponding device type of the input acoustic feature.
5. The method for sound scene classification based on supervised domain adaptation of multiple devices according to claim 1, It is characterized in that In step S3, the Mixup module and the SpecAugment module are used to perform data enhancement processing on the three-dimensional acoustic features described in step S2.
6. The method for sound scene classification based on supervised domain adaptation of multiple devices according to claim 1, It is characterized in that In step S4, the sound scene classification model is composed of three feature alignment convolution blocks and two fully connected layers.
7. The method for sound scene classification based on multiple devices with supervised domain adaptation according to claim 1, It is characterized in that In step S4, the feature alignment convolution block includes two convolution layers, two batch normalization layers (Batch-Normalization, referred to as BN layers), two activation function layers, and one pooling layer.
Citation Information
Patent Citations
Cross-library speech emotion recognition method based on one-dimensional convolution auto-encoder and adversarial domain self-adaption
CN114038480A
Model and method for multi-source domain adaptation by aligning partial features
US20220138495A1