Audio classification method and device, equipment, storage medium and program product
By extracting and stacking audio features layer by layer, combining the intermediate discrimination results, and outputting audio classification results, the problems of low accuracy and poor robustness of audio classification in the prior art are solved, and higher classification accuracy and model robustness are achieved.
Patent Information
- Application Number
- CN202411934836.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-06-03
AI Technical Summary
Existing audio classification methods rely on specific features, resulting in a decrease in classification accuracy when the model processes specific types of audio, and the feature aggregation method ignores feature information at different levels, affecting the accuracy and robustness of audio classification.
Audio features of different granularity in the audio data to be classified layer by layer, and these features are handed over to the discriminator of the corresponding layer for discrimination, thereby generating intermediate discrimination results. Then, features of different particle sizes are stacked to generate feature bases, and finally the audio classification results are output through the final discriminator combined with the intermediate discrimination results.
A more refined feature extraction is achieved, and the accuracy of audio classification results and the robustness of the model are improved through more detailed and richer audio features and multiple intermediate discriminant results containing classification probability.
Smart Images

Figure CN120086707A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of audio processing, and in particular, to an audio classification method, apparatus, device, storage medium, and program product. Background Art
[0002] Audio classification is an important field in audio signal processing, which involves classifying audio signals in a detailed and accurate manner according to their content or features. Audio classification can be applied to various scenarios, such as music classification, environmental sound recognition, and speech recognition.
[0003] Existing audio classification methods use a pre-trained audio encoder as a basis, fine-tune the model for a specific dataset to adapt to different classification tasks. In the backend stage of the model, the extracted audio features are integrated through feature aggregation technology to finally complete the classification work. However, in the process of implementing the technical solutions of the present application by the inventors of the present application, it is found that the above technologies have at least the following technical problems:
[0004] The pre-trained audio encoder depends on specific features, resulting in a decrease in the classification accuracy of the model when processing specific types of audio; most existing feature aggregation methods only rely on the features of the last layer of the model for decision-making, ignoring the feature information from different levels, affecting the accuracy and robustness of audio classification. Summary of the Invention
[0005] The present application provides an audio classification method, apparatus, device, storage medium, and program product to solve the problems of low accuracy and poor robustness of existing audio classification methods.
[0006] In a first aspect, to solve the above problems, the present application discloses an audio classification method, including the following steps:
[0007] Obtain audio data to be classified;
[0008] Extract audio features with different granularities in the audio data to be classified layer by layer, and hand over the audio features to discriminators corresponding to each layer for discrimination, respectively generating intermediate discrimination results, where the intermediate discrimination results include the classification probabilities of the audio features with corresponding granularities;
[0009] Stack the audio features with different granularities to generate a feature base, where the feature base represents a set of audio features with different granularities in the audio data to be classified;
[0010] Process the feature base based on a final discriminator, and combine the intermediate discrimination results to output an audio classification result, where the audio classification result includes a classification label for representing an audio event.
[0011] Further, it further includes the following steps:
[0012] Obtain a complete audio representation, where the complete audio representation characterizes a training dataset for model training;
[0013] Extract training features of different granularities in the complete audio representation layer by layer, and hand over the training features to discriminators of corresponding layers for discrimination, respectively generating intermediate training discrimination results, where the intermediate training discrimination results characterize the classification probabilities of the training features of corresponding granularities;
[0014] Stack the training features of different granularities to generate a training feature base, where the training feature base characterizes a set of training features of different granularities in the complete audio representation;
[0015] Based on a final discriminator, process the training feature base, and combine the intermediate training discrimination results to output a global discrimination result, where the global discrimination result is used to evaluate the classification accuracy of the complete audio representation;
[0016] Compare the global discrimination result with the true label of the complete audio representation to determine a training loss, where the training loss characterizes the fitting degree of the model under the current parameters;
[0017] Determine a parameter gradient according to the training loss, where the parameter gradient characterizes the direction in which the training loss decreases;
[0018] Update model parameters through the parameter gradient to generate an audio classification model, where the audio classification model is used to output the audio classification result.
[0019] Further, obtaining a complete audio representation includes the following steps:
[0020] Obtain original audio data;
[0021] A gating unit identifies and analyzes multiple original audio features in the original audio data, and assigns the original audio features to corresponding expert models, where the original audio features include original waveform features, filter bank features, and short-time Fourier transform features;
[0022] The expert models respectively extract key information in the original audio features, where the key information characterizes the frequency components, periodic changes, and waveform features of the original audio data;
[0023] Integrate the key information to generate the complete audio representation.
[0024] Further, the integrating the key information to generate the complete audio representation specifically includes:
[0025] Preprocess the key information to obtain preprocessed information;
[0026] According to a preset coupling strategy, couple the preprocessed information together to generate a feature vector;
[0027] Adjust the length of the feature vector to a target length to obtain a target vector, where the target length refers to the length of the complete audio representation;
[0028] Map the target vector to the complete audio representation.
[0029] Further, the following steps are further included:
[0030] According to the audio classification result, turn on the corresponding noise reduction mode, where the noise reduction mode refers to a mode of reducing the noise level.
[0031] Further, the following steps are further included:
[0032] According to the audio classification result, turn on the corresponding speech enhancement mode, where the speech enhancement mode refers to a mode of extracting and enhancing a pure original speech signal from a noisy audio.
[0033] In a second aspect, the present application also discloses an audio classification device, including:
[0034] An acquisition module, configured to acquire audio data to be classified;
[0035] An extraction module, configured to extract audio features of different granularities in the audio data to be classified layer by layer, and hand over the audio features to discriminators of corresponding layers for discrimination to generate intermediate discrimination results respectively, where the intermediate discrimination results include classification probabilities of the audio features of corresponding granularities;
[0036] A stacking module, configured to stack the audio features of different granularities to generate a feature base, where the feature base represents a set of audio features of different granularities in the audio data to be classified;
[0037] A discrimination module, configured to process the feature base based on a final discriminator, and combine the intermediate discrimination results to output an audio classification result, where the audio classification result includes a classification label for characterizing an audio event.
[0038] In a third aspect, the present application also discloses an electronic device, including:
[0039] A processor;
[0040] A memory, configured to store executable instructions of the processor;
[0041] Wherein, the processor is configured to execute the instructions to implement the audio classification method according to any one of the above.
[0042] In a fourth aspect, the present application also discloses a computer-readable storage medium,
[0043] When the instructions in the computer-readable storage medium are executed by a processor of a terminal, the terminal is enabled to execute the audio classification method according to any one of the claims.
[0044] In a fifth aspect, the present application also discloses a computer program product,
[0045] The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the audio classification method according to any one of the above.
[0046] Compared with the prior art, the present application has the following advantages:
[0047] By using the audio classification model to extract audio features with different granularities from the audio data to be classified layer by layer, and handing these features to the discriminators of the corresponding layers for discrimination, generating intermediate discrimination results containing classification probabilities, and stacking the audio features with different granularities to construct a feature base, a rich and comprehensive feature representation is obtained. Finally, the discriminator adopts a weighted combination strategy, gives different weights to each intermediate discrimination result according to the confidence and importance of each intermediate discrimination result, and combines the discrimination result of the feature base to output the audio classification result. It realizes more refined feature extraction, enables the audio classification model to output the audio classification result through more detailed and rich audio features and multiple intermediate discrimination results containing classification probabilities, improves the accuracy of the audio classification result and the robustness of the audio classification model, and effectively solves the technical problems of low accuracy and poor robustness of the existing audio classification methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 is a flowchart of a method for audio classification provided by an embodiment of the present application;
[0049] Figure 2 is a flowchart of a method for generating an audio classification model provided by an embodiment of the present application;
[0050] Figure 3 is a flowchart of a method for obtaining a complete audio representation provided by an embodiment of the present application;
[0051] Figure 4 is a flowchart of a method for generating a complete audio representation provided by an embodiment of the present application;
[0052] Figure 5 is a flowchart of a method for turning on the noise reduction mode provided by an embodiment of the present application;
[0053] Figure 6It is a flowchart of a method for enabling a voice enhancement mode provided by an embodiment of the present application;
[0054] Figure 7 It is a schematic structural diagram of an audio classification device provided by an embodiment of the present application;
[0055] Figure 8 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0056] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below with reference to the accompanying drawings and specific implementation manners.
[0057] To facilitate understanding of the technical solutions provided by the present application, a brief introduction to audio classification is given:
[0058] As an important part of the field of audio signal processing, audio classification focuses on dividing audio signals in detail and accurately according to their content or features. Through audio classification, complex audio data can be converted into classification information with clear meaning and practical value.
[0059] Existing audio classification methods widely use pre-trained audio encoders as the cornerstone. These encoders are trained with a large amount of data and can efficiently extract key features in audio signals. In current practice, these methods usually only limit to input a specific audio feature, such as Fbank feature, for the model to learn. To improve the adaptability of the model in different classification tasks, the pre-trained model is fine-tuned for a specific data set by adjusting the model parameters to optimize its performance on specific tasks.
[0060] In the backend processing stage of the model, the feature aggregation technology uses the features extracted by the last layer of the model to make classification decisions. The first-level discriminator at the backend of the model analyzes the aggregated feature vectors and learns the discrimination boundaries between different categories, so as to achieve accurate classification of audio signals.
[0061] In the process of implementing the technical solutions of the embodiments of the present application, the inventors of the present application found that the above technologies have at least the following technical problems:
[0062] Pre-trained models often only support the input of specific types of features such as Fbank. Although Fbank features can effectively reflect the spectral characteristics of audio in some scenarios, when processing audio signals rich in high-frequency information, the model may have a decrease in classification accuracy due to its inability to fully extract these key features.
[0063] In addition, most of the existing feature aggregation methods only rely on the features of the last layer of the model for decision-making. This approach ignores the potential associations and complementarities between different levels and dimensions in the feature vector. In fact, the feature aggregation based on the convolutional module can achieve fine-grained features from top to bottom, while the feature aggregation based on Transformer is good at extracting coarse-grained features from bottom to top. Therefore, only using the features of the last layer for aggregation cannot fully utilize this multi-level and multi-scale feature information, which limits the improvement of classification performance.
[0064] Due to the above problems, the existing audio classification methods perform poorly in some specific audio classifications and are prone to overfitting, which seriously affects the accuracy and robustness of audio classification.
[0065] Embodiment 1
[0066] Referring to Figure 1 , there is shown an audio classification method provided by an embodiment of the present application, including the following steps:
[0067] S10. Obtain the audio data to be classified.
[0068] In this embodiment, the audio data to be classified refers to the audio data that has not been classified, and the audio data to be classified includes at least one of voice data, music data, environmental sound effects, animal calls, and human activity sounds. Specifically, an audio classification model is used to classify the audio data to be classified to obtain an audio classification result representing an audio event. Exemplarily, the audio classification model includes a multi-layer network structure, and the audio classification model is a trained convolutional neural network or a Transformer (deep learning) model.
[0069] S20. Extract audio features of different granularities in the audio data to be classified layer by layer, and hand over the audio features to discriminators of corresponding layers for discrimination to generate intermediate discrimination results respectively, where the intermediate discrimination results include the classification probabilities of the audio features of the corresponding granularities.
[0070] In this embodiment, using the multi-layer network structure of the audio classification model, audio features from coarse to fine are extracted layer by layer. Each layer of the audio classification model focuses on extracting audio features of a specific granularity or abstraction level, reflecting different sound components and attributes in the audio data to be classified.
[0071] In the process of extracting audio features layer by layer, each layer will pass the extracted audio features to the discriminator of the corresponding layer. The discriminator is a trained classifier responsible for analyzing the audio features extracted in the current layer and generating intermediate discrimination results. The intermediate discrimination results include the classification probabilities of the audio features corresponding to the corresponding granularity. Using multiple discriminators to generate the classification probabilities of the audio features corresponding to the corresponding granularity for the corresponding layer helps to more comprehensively understand the structure and content of the audio data to be classified and improves the accuracy of the audio classification results.
[0072] S30. Stack the audio features of different granularities to generate a feature base, and the feature base represents a set of audio features of different granularities in the audio data to be classified.
[0073] In this embodiment, by stacking the audio features of different granularities together to form a feature base. Specifically, the feature base not only represents a set of audio features of different granularities in the audio data to be classified, but also provides a comprehensive and in-depth description of the audio data to be classified through the fusion and complementarity of the features. Exemplarily, the feature base is formed by stacking from coarser to finer granularity. Using the feature base to comprehensively reflect the audio data to be classified improves the accuracy and robustness of audio classification.
[0074] S40. Process the feature base based on the final discriminator, and combine the intermediate discrimination results to output the audio classification result, where the audio classification result includes a classification label for characterizing the audio event.
[0075] In this embodiment, the audio classification result is output by the final discriminator of the audio classification model. The final discriminator is set at the end of the audio classification model. By discriminating the feature base and combining the intermediate discrimination results, the audio classification result is output. Specifically, the classification labels for characterizing the audio events at least include inside a bus, inside a restaurant, on a noisy street, on a subway, in a hospital, in a school.
[0076] Exemplarily, the final discriminator adopts a weighted combination strategy. According to the confidence and importance of each intermediate discrimination result, different weights are given to each intermediate discrimination result, and combined with the discrimination result of the feature base, the audio classification result is output. Using the weighted combination strategy can improve the accuracy of the audio classification result.
[0077] Extract audio features of different granularities in the audio data to be classified layer by layer through an audio classification model, and hand these features to the discriminator of the corresponding layer for discrimination to generate intermediate discrimination results containing classification probabilities. Stack the audio features of different granularities to construct a feature base to obtain a rich and comprehensive feature representation. Finally, the discriminator adopts a weighted combination strategy. According to the confidence and importance of each intermediate discrimination result, different weights are given to each intermediate discrimination result, and combined with the discrimination result of the feature base, the audio classification result is output. It realizes more refined feature extraction, enables the audio classification model to output the audio classification result through more detailed and rich audio features and multiple intermediate discrimination results containing classification probabilities, improves the accuracy of the audio classification result and the robustness of the audio classification model, and effectively solves the technical problems of low accuracy and poor robustness of existing audio classification methods.
[0078] Refer to Figure 2 , and also includes the following steps:
[0079] S100. Obtain a complete audio representation, and the complete audio representation characterizes the training data set for model training.
[0080] In this embodiment, the complete audio representation comprehensively includes various audio features. Using various audio features for model training can ensure the classification accuracy of the trained audio classification model and avoid overfitting. Specifically, the model used for training includes a convolutional neural network or a Transformer (deep learning) model.
[0081] S200. Extract training features of different granularities in the complete audio representation layer by layer, hand the training features to the discriminator of the corresponding layer for discrimination, and generate intermediate training discrimination results respectively. The intermediate training discrimination results characterize the classification probabilities of the training features of the corresponding granularity.
[0082] In this embodiment, using the multi-layer network structure of the model, extract audio features from coarse to fine layer by layer. Each layer of the audio classification model focuses on extracting audio features of a specific granularity or abstraction level, reflecting different sound components and attributes in the audio data to be classified. The discriminator of each layer is responsible for discriminating its corresponding training features to generate intermediate training discrimination results, and the intermediate training discrimination results characterize the probability distribution of the training features of the corresponding granularity in the classification task. By extracting training features of different granularities in the complete audio representation layer by layer and handing the training features to the discriminator of the corresponding layer for discrimination, the classification probabilities of the training features of the corresponding granularity are obtained, which improves the accuracy and effectiveness of model training and enhances the generalization ability.
[0083] S300. Stack the training features of different granularities to generate a training feature base, and the training feature base characterizes the set of training features of different granularities in the complete audio representation.
[0084] In this embodiment, the training feature base not only includes various feature information of the complete audio representation, but also enhances the expression ability of the features in a stacked manner, ensuring the effectiveness of model training.
[0085] S400. Process the training feature base based on the final discriminator, and combine the intermediate training discrimination results to output a global discrimination result, which is used to evaluate the classification accuracy of the complete audio representation.
[0086] In this embodiment, the final discriminator is set at the end of the model. The final discriminator comprehensively analyzes the training feature base by combining the intermediate training discrimination results generated by each previous layer of discriminators, and outputs a global discrimination result. The global discrimination result is used to evaluate the classification accuracy of the model for the complete audio representation and is an important measure of the model training effect.
[0087] S500. Compare the global discrimination result with the true label of the complete audio representation to determine the training loss, which characterizes the fitting degree of the model under the current parameters.
[0088] In this embodiment, the true label represents the actual classification result of the complete audio representation. By comparing the global discrimination result with the true label of the complete audio representation, the training loss can be determined. The training loss characterizes the fitting degree of the model to the training dataset under the current parameter configuration. The smaller the training loss, the more accurate the global discrimination result.
[0089] S600. Determine the parameter gradient according to the training loss, where the parameter gradient characterizes the direction in which the training loss decreases.
[0090] In this embodiment, the parameter gradient is used to guide the optimization direction of the model. The parameter gradient reveals how each parameter in the model should be adjusted to reduce the training loss. Specifically, the parameter gradient points to the direction in which the value of the training loss function decreases fastest, providing a clear guidance for updating the model parameters.
[0091] S700. Update the model parameters through the parameter gradient to generate an audio classification model, which is used to output the audio classification result.
[0092] In this embodiment, the gradient descent optimization algorithm is used to update the model parameters. After each update, the model will update to be closer to the distribution of the true label, thereby improving the generalization ability. Finally, after multiple iterative trainings, an audio classification model is generated.
[0093] Refer to Figure 3 , S100. Obtain the complete audio representation, including the following steps:
[0094] S110. Obtain the original audio data.
[0095] In this embodiment, the original audio data is collected from at least actual scenarios such as streets, restaurants, buses, subways, schools, and hospitals. When obtaining the original audio data, it is necessary to preprocess the data format and quality. Exemplarily, the sampling rate and bitrate parameters of the preprocessed original audio data are kept consistent to avoid problems such as format incompatibility or quality degradation during subsequent processing.
[0096] S120. The gating unit identifies and analyzes multiple original audio features in the original audio data and assigns the original audio features to corresponding expert models. The original audio features include original waveform features, filter bank features, and short-time Fourier transform features.
[0097] In this embodiment, the gating unit is an intelligent processing unit that can identify and analyze multiple original audio features in the original audio data. The original audio features at least include original waveform features, filter bank (Fbank) features, and short-time Fourier transform (STFT) features. The original waveform features reflect the basic waveform information of the original audio data. The filter bank features refer to the feature information of different frequency bands extracted after processing the audio data through a series of filters. The short-time Fourier transform features refer to the spectral features extracted after performing local spectral analysis on the audio data by sliding a time window. The gating unit assigns the original audio features to corresponding expert models for processing according to the nature and importance of the original audio features, realizing the preliminary parsing and feature extraction of the original audio data and providing a rich information source for obtaining a complete audio representation.
[0098] Exemplarily, the gating unit and expert models in MOE (Mixture of Experts) are used to process the original audio data.
[0099] S130. The expert models respectively extract the key information in the original audio features. The key information characterizes the frequency components, periodic changes, and waveform features of the original audio data.
[0100] In this embodiment, the expert models are used to process specific types of original audio features. By extracting and analyzing the key information in the original audio features, information characterizing the frequency components, periodic changes, and waveform features of the original audio data is obtained. The frequency components reflect the energy distribution of different frequency components in the original audio data and are an important basis for audio classification and recognition. The periodic changes reflect the existence and changes of periodic components in the original audio data. The waveform features reflect the basic waveform shape and changes of the original audio data. The expert models further process the original audio features to extract this key information as the basis for generating a complete audio representation.
[0101] S140. Integrate the key information to generate a complete audio representation.
[0102] In this embodiment, after extracting the key information, by fusing the key information together, a complete audio representation is generated. The complete audio representation not only contains the basic information and features of the original audio data, but also reflects the internal structure and rules of the original audio data, ensuring the effectiveness of model training.
[0103] Use an expert model to extract the key information from specific types of original audio features respectively. By integrating the key information, a complete audio representation is generated, solving the problems of low audio classification accuracy and overfitting caused by relying on specific features for model training in the prior art. Using more comprehensive audio features for model training improves the effectiveness of model training, thus ensuring the accuracy of audio classification results.
[0104] Refer to Figure 4 , S140. Integrate the key information to generate a complete audio representation, which specifically includes:
[0105] S141. Preprocess the key information to obtain preprocessed information.
[0106] In this embodiment, by preprocessing the key information, the preprocessed information is made consistent in format, dimension, and data type.
[0107] S142. According to the preset coupling strategy, couple the preprocessed information together to generate a feature vector.
[0108] In this embodiment, the preset coupling strategy refers to feature concatenation and feature fusion. By coupling different preprocessed information together, a feature vector that can comprehensively reflect audio features is generated.
[0109] S143. Adjust the length of the feature vector to the target length to obtain a target vector, where the target length refers to the length of the complete audio representation.
[0110] In this embodiment, since the lengths of different feature vectors may be different, to ensure consistency, it is necessary to adjust the length of the feature vector to the target length. Vector truncation, zero padding, or interpolation methods are used to extend or shorten the length of the feature vector so that the length of the adjusted feature vector is the same as the target length, ensuring that the adjusted feature vector has the same dimension as the complete audio representation.
[0111] S144. Map the target vector to the complete audio representation.
[0112] In this embodiment, by mapping the target vector to a complete audio representation, the target vector is transformed into the form of a complete audio representation. The complete audio representation not only includes core audio feature information but also meets the format and dimensional requirements for model training.
[0113] By preprocessing the key information, the consistency of the preprocessed information in terms of format, dimension, and data type is ensured. Through a preset coupling strategy, the information to be processed is coupled together to generate a feature vector, which comprehensively reflects the audio features, solves the technical problem of poor training effect of a single audio feature model, and improves the effectiveness of model training. By adjusting the length of the feature vector and mapping the obtained target vector to a complete audio representation, the format and dimensional requirements for model training are ensured.
[0114] Refer to Figure 5 , and it further includes the following steps: S50A. According to the audio classification result, turn on the corresponding noise reduction mode, where the noise reduction mode refers to a mode of reducing the noise level.
[0115] Specifically, selecting and turning on the corresponding noise reduction strategy according to the audio classification result can ensure that users enjoy a clear and quiet auditory experience.
[0116] Exemplarily, when the audio classification result indicates being inside a bus, by starting a noise reduction algorithm designed specifically for buses, background noises such as vehicle driving and passenger conversations are weakened. When the classification result shows that the current environment is on the street, a noise mode for the urban outdoor environment is enabled to suppress complex sound sources such as traffic noise and construction noise. When the classification result shows being inside a restaurant, a gentle noise reduction technology is adopted to reduce the interference of noises such as noisy voices and tableware collisions to ensure the noise reduction effect.
[0117] Refer to Figure 6 , and it further includes the following steps: S50B. According to the audio classification result, turn on the corresponding voice enhancement mode, where the voice enhancement mode refers to a mode of extracting and enhancing the pure original voice signal from the noisy audio.
[0118] Exemplarily, when the audio classification result shows being inside a bus, start a voice enhancement algorithm designed specifically for buses to identify and weaken background noises such as vehicle driving and passenger conversations, thereby highlighting and enhancing the target voice signal; when the audio classification result shows being on the street, turn on an enhancement strategy more suitable for the noisy outdoor environment to ensure that the voice content can still be clearly captured and enhanced under the interference of traffic noise, construction noise, etc.; when the audio classification result shows being inside a restaurant, enable a more delicate and efficient voice enhancement technology to achieve precise suppression of specific noises such as noisy voices and tableware collisions, while maintaining the clarity and naturalness of the voice signal, providing users with a better voice communication experience.
[0119] In the embodiments of the present application, a gating unit is adopted to allocate different types of original audio features to corresponding expert models for processing. The expert models extract key information from specific types of original audio features, generate a complete audio representation by integrating the key information, use the complete audio representation for model training to generate an audio classification model, extract audio features with different granularities from the audio data to be classified layer by layer through the audio classification model, and hand these features to discriminators of corresponding layers for discrimination to generate intermediate discrimination results including classification probabilities, and stack the audio features with different granularities to construct a feature base, use a final discriminator to discriminate the feature base, and combine the intermediate discrimination results to output an audio classification result including a classification label representing an audio event, thus solving the problems of low audio classification accuracy and overfitting caused by relying on specific features for model training in the prior art. Using more comprehensive audio features for model training improves the effectiveness of model training, thereby ensuring the accuracy of the audio classification effect. At the same time, it solves the technical problems of low accuracy and poor robustness of existing audio classification methods, and realizes the technical effect of ensuring the accuracy of audio classification.
[0120] Embodiment 2
[0121] Referring to Figure 7 , the embodiments of the present application provide an audio classification device, including:
[0122] An acquisition module 510, configured to acquire audio data to be classified.
[0123] An extraction module 520, configured to extract audio features with different granularities from the audio data to be classified layer by layer, and hand the audio features to discriminators of corresponding layers for discrimination, respectively generating intermediate discrimination results, where the intermediate discrimination results include classification probabilities of the audio features with corresponding granularities.
[0124] A stacking module 530, configured to stack the audio features with different granularities to generate a feature base, where the feature base represents a set of audio features with different granularities in the audio data to be classified.
[0125] A discrimination module 540, configured to process the feature base based on a final discriminator, combine the intermediate discrimination results, and output an audio classification result, where the audio classification result includes a classification label used to represent an audio event.
[0126] In some embodiments, it further includes:
[0127] An audio representation acquisition module, configured to acquire a complete audio representation, where the complete audio representation is used for a training data set for model training.
[0128] A training feature module, which is used to extract training features of different granularities in the complete audio representation layer by layer, and deliver the training features to discriminators of corresponding layers for discrimination, respectively generating intermediate training discrimination results, where the intermediate training discrimination results represent the classification probabilities of the training features of corresponding granularities.
[0129] A feature stacking module, which is used to stack training features of different granularities to generate a training feature base, where the training feature base represents a set of training features of different granularities in the complete audio representation.
[0130] A result output module, which is used to process the training feature base based on a final discriminator, combine the intermediate training discrimination results, and output a global discrimination result, where the global discrimination result is used to evaluate the classification accuracy of the complete audio representation.
[0131] A comparison module, which is used to compare the global discrimination result with the true label of the complete audio representation to determine a training loss, where the training loss represents the fitting degree of the model under the current parameters.
[0132] A parameter gradient module, which is used to determine a parameter gradient according to the training loss, where the parameter gradient represents the direction in which the training loss decreases.
[0133] An update module, which is used to update model parameters through the parameter gradient to generate an audio classification model, where the audio classification model is used to output an audio classification result.
[0134] In some embodiments, the audio representation acquisition module includes:
[0135] An acquisition unit, which is used to acquire original audio data.
[0136] An identification unit, which is used to identify and analyze multiple original audio features in the original audio data by a gating unit, and assign the original audio features to corresponding expert models, where the original audio features include original waveform features, filter bank features, and short-time Fourier transform features.
[0137] An extraction unit, which is used for expert models to respectively extract key information in the original audio features, where the key information represents the frequency components, periodic changes, and waveform features of the original audio data.
[0138] An integration unit, which is used to integrate the key information to generate a complete audio representation.
[0139] In some embodiments, the integration unit is used to:
[0140] Preprocess the key information to obtain preprocessed information.
[0141] According to a preset coupling strategy, couple the preprocessed information together to generate a feature vector.
[0142] Adjust the length of the feature vector to a target length to obtain a target vector, where the target length refers to the length of a complete audio representation.
[0143] Map the target vector to a complete audio representation.
[0144] In some embodiments, it further includes:
[0145] A noise reduction module, configured to turn on a corresponding noise reduction mode according to the audio classification result, where the noise reduction mode refers to a mode of reducing the noise level.
[0146] In some embodiments, it further includes:
[0147] A voice enhancement module, configured to turn on a corresponding voice enhancement mode according to the audio classification result, where the voice enhancement mode refers to a mode of extracting and enhancing a pure original voice signal from a noisy audio.
[0148] For the apparatus embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For related parts, please refer to the partial description of the method embodiments.
[0149] Embodiment III
[0150] Referring to Figure 8 , the embodiments of the present application further provide an electronic device, including:
[0151] A processor.
[0152] A memory, configured to store executable instructions for the processor.
[0153] Wherein, the processor is configured to execute instructions to implement the audio classification method of any one of them.
[0154] In this embodiment, the computer device includes a processor, a memory, and a network interface connected through a system bus.
[0155] Wherein, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data samples. The network interface of the computer device is used to communicate with an external terminal through a network connection. The computer program, when executed by the processor, implements the audio classification method of any one of them.
[0156] Those skilled in the art can understand, Figure 8The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.
[0157] Embodiment 4
[0158] The embodiment of this application also provides a computer-readable storage medium. When the instructions in the computer-readable storage medium are executed by the processor of the terminal, the terminal can execute any of the audio classification methods.
[0159] The above computer-readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory, electrically erasable programmable read-only memory, erasable programmable read-only memory, programmable read-only memory, read-only memory, magnetic memory, flash memory, a magnetic disk or an optical disc. The readable storage medium can be any available medium accessible by a general-purpose or special-purpose computer.
[0160] Optionally, the readable storage medium is coupled to the processor so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can be located in an application specific integrated circuit (ASIC). Of course, the processor and the readable storage medium can also exist as discrete components in the device.
[0161] Embodiment 5
[0162] The embodiment of this application also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by the processor, it implements any of the audio classification methods.
[0163] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0164] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and combinations of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0165] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0166] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0167] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0168] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
[0169] Each embodiment in this specification is described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other.
[0170] The above has introduced in detail the audio classification method, apparatus, and device provided by this application. Specific examples are used in this article to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to this application.
Claims
1. An audio classification method, characterized in that: The steps include: Obtain audio data to be classified; Extracting audio features of different granularities from the audio data to be classified layer by layer, and handing the audio features over to the discriminator of the corresponding layer for discrimination, and generating intermediate discrimination results respectively, wherein the intermediate discrimination results include the classification probabilities of the audio features of the corresponding granularities; The audio features of different granularities are stacked to generate a feature base, where the feature base represents a set of audio features of different granularities in the audio data to be classified; The feature base is processed based on the final discriminator, combined with the intermediate discrimination result, and an audio classification result is output, wherein the audio classification result includes a classification label for characterizing an audio event.
2. The method according to claim 1, characterized in that The following steps are also included: Obtaining a complete audio representation, wherein the complete audio representation represents a training data set for model training; Extracting training features of different granularities in the complete audio representation layer by layer, handing the training features over to the discriminator of the corresponding layer for discrimination, and generating intermediate training discrimination results respectively, wherein the intermediate training discrimination results represent the classification probability of the training features of the corresponding granularity; Stacking the training features of different granularities to generate a training feature base, wherein the training feature base represents a set of training features of different granularities in the complete audio representation; Processing the training feature base based on the final discriminator, combining the intermediate training discrimination results, and outputting a global discrimination result, wherein the global discrimination result is used to evaluate the classification accuracy of the complete audio representation; Comparing the global discrimination result with the true label of the complete audio representation to determine a training loss, wherein the training loss represents a degree of fit of the model under current parameters; Determine a parameter gradient according to the training loss, the parameter gradient representing a direction in which the training loss decreases; The model parameters are updated by using the parameter gradients to generate an audio classification model, and the audio classification model is used to output the audio classification result.
3. The method according to claim 2, characterized in that Obtaining a complete audio representation includes the following steps: Get the original audio data; The gating unit identifies and analyzes a plurality of original audio features in the original audio data, and assigns the original audio features to corresponding expert models, wherein the original audio features include original waveform features, filter bank features, and short-time Fourier transform features; The expert model extracts key information from the original audio features respectively, wherein the key information represents the frequency components, periodic changes and waveform characteristics of the original audio data; The key information is integrated to generate the complete audio representation.
4. The method according to claim 3, characterized in that The integrating the key information to generate the complete audio representation specifically includes: Preprocessing the key information to obtain preprocessed information; According to a preset coupling strategy, the preprocessing information is coupled together to generate a feature vector; Adjusting the length of the feature vector to a target length to obtain a target vector, wherein the target length refers to the length of the complete audio representation; The target vector is mapped to a complete audio representation.
5. The method according to claim 1, characterized in that The following steps are also included: According to the audio classification result, a corresponding noise reduction mode is turned on, where the noise reduction mode refers to a mode for reducing the noise level.
6. The method according to claim 1, characterized in that The following steps are also included: According to the audio classification result, a corresponding speech enhancement mode is turned on, where the speech enhancement mode refers to a mode for extracting and enhancing a pure original speech signal from a noisy audio signal.
7. An audio classification device, characterized in that: include: An acquisition module, used to acquire audio data to be classified; An extraction module, used for extracting audio features of different granularities in the audio data to be classified layer by layer, and delivering the audio features to the discriminator of the corresponding layer for discrimination, and generating intermediate discrimination results respectively, wherein the intermediate discrimination results include the classification probability of the audio features of the corresponding granularity; A stacking module, used for stacking the audio features of different granularities to generate a feature base, wherein the feature base represents a set of audio features of different granularities in the audio data to be classified; The discriminant module is used to process the feature base based on the final discriminator, combine the intermediate discriminant results, and output an audio classification result, wherein the audio classification result includes a classification label for characterizing an audio event.
8. An electronic device, characterized in that: include: processor; A memory, configured to store instructions executable by the processor; The processor is configured to execute the instructions to implement the audio classification method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: When the instructions in the computer-readable storage medium are executed by a processor of a terminal, the terminal is enabled to execute the audio classification method according to any one of claims 1 to 6.
10. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the audio classification method according to any one of claims 1 to 6 is implemented.