Sound scene classification model generation method, sound scene classification method, device, storage medium and electronic equipment
By generating enhanced spectrograms and labels in the sound scene classification model, the problem of poor performance caused by insufficient training data is solved, the accuracy and generalization ability of sound scene classification are improved, and the acoustic properties of audio data are protected.
Patent Information
- Application Number
- CN202410848719.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-27
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2044-06-27
AI Technical Summary
Existing deep learning-based sound scene classification models perform poorly with limited training data, resulting in inaccurate classification results.
By randomly selecting source and target audio from the sound scene classification dataset, source Mel spectrograms and target Mel spectrograms are generated. Then, enhanced spectrograms and labels are generated using random mask maps, inverted random masks, source Mel spectrograms, and target Mel spectrograms. Based on the enhanced spectrograms and labels, a preset neural network is trained to generate a sound scene classification model.
It improves the accuracy of sound scene classification results, avoids the problem of poor performance due to insufficient training data, and protects the acoustic characteristics of the original audio data, thus achieving better generalization and classification performance.
Smart Images

Figure CN118658464B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio processing technology, specifically to a method for generating a sound scene classification model, a sound scene classification method, an apparatus, a storage medium, and an electronic device. Background Technology
[0002] People's daily activities involve a variety of sound events, and the combinations of these sound events constitute various sound scenes. Sound scene classification technology has a wide range of applications, such as audio monitoring, multimedia retrieval, autonomous driving assistance, and smart homes.
[0003] With the development of sound scene classification technology, intelligent algorithms for sound scene classification based on deep learning have been widely used. However, sound scene classification methods based on deep learning are highly dependent on training data. In practical applications, current sound scene classification models often suffer from poor performance due to insufficient training data, which limits the application of the trained models in real-world scenarios and results in inaccurate sound scene classification results. Summary of the Invention
[0004] This application provides a method for generating a sound scene classification model, a sound scene classification method, an apparatus, a storage medium, and an electronic device, which can improve the accuracy of sound scene classification results.
[0005] In a first aspect, embodiments of this application provide a method for generating an acoustic scene classification model, including:
[0006] Source and target audio are randomly selected from the sound scene classification dataset;
[0007] Generate a source Mel spectrogram and a target Mel spectrogram based on the source speech and the target speech;
[0008] A random mask image is generated based on the target Mel spectrogram, and an inverted random mask image is obtained from the random mask image.
[0009] An enhanced spectrogram and a label are generated based on the random mask image, the inverted random mask, the source Mel spectrogram, and the target Mel spectrogram;
[0010] The preset neural network is trained based on the enhanced spectrogram and the labels to generate a sound scene classification model.
[0011] In the sound scene classification model generation method provided in this application embodiment, the step of generating an enhanced spectrogram and labels based on a random mask image, the inverted random mask, the source Mel spectrogram, and the target Mel spectrogram includes:
[0012] An enhanced spectrogram is generated using the random mask image, the inverted random mask, the source Mel spectrogram, and the target Mel spectrogram;
[0013] A label is generated using the random mask image, the source Mel spectrogram, and the target Mel spectrogram.
[0014] In the sound scene classification model generation method provided in this application embodiment, the step of generating an enhanced spectrogram using the random mask image, the inverted random mask, the source Mel spectrogram, and the target Mel spectrogram includes:
[0015] The random mask image is multiplied by the target Mel spectrum image to generate a first intermediate result image.
[0016] The source Mel spectrum is multiplied by the inverted random mask to generate a second intermediate result image.
[0017] The first intermediate result image is added to the second intermediate result image to obtain the enhanced spectrum image.
[0018] In the sound scene classification model generation method provided in this application embodiment, the step of generating labels using the random mask image, the source Mel spectrogram, and the target Mel spectrogram includes:
[0019] Obtain the first proportion of the random mask image in the target Mel spectrogram;
[0020] Obtain the second proportion of the random mask image in the source Mel spectrogram;
[0021] The label is generated based on the first proportion and the second proportion.
[0022] In the sound scene classification model generation method provided in this application embodiment, the step of generating a source Mel spectrogram and a target Mel spectrogram based on the source speech and the target speech includes:
[0023] Perform Fast Fourier Transform on the source speech and the target speech respectively to generate a first spectrogram and a second spectrogram;
[0024] The first and second spectrograms are processed using a Mel filter bank to generate a source Mel spectrogram and a target Mel spectrogram.
[0025] Secondly, embodiments of this application provide a sound scene classification method, including:
[0026] Get the audio to be categorized;
[0027] The audio to be classified is input into the above-mentioned sound scene classification model to obtain the sound scene classification result of the audio to be classified.
[0028] Thirdly, embodiments of this application provide an apparatus for generating an acoustic scene classification model, comprising:
[0029] The audio selection unit is used to randomly select source and target audio from the sound scene classification dataset;
[0030] The first generation unit is configured to generate a source Mel spectrogram and a target Mel spectrogram based on the source speech and the target speech;
[0031] The second generation unit is used to generate a random mask map based on the target Mel spectrogram and obtain an inverted random mask map of the random mask map;
[0032] The third generation unit is used to generate an enhanced spectrogram and a label based on the random mask image, the inverted random mask, the source Mel spectrogram, and the target Mel spectrogram;
[0033] The fourth generation unit is used to train a preset neural network based on the enhanced spectrogram and the label to generate a sound scene classification model.
[0034] Fourthly, embodiments of this application provide a sound scene classification device, including:
[0035] The audio acquisition unit is used to acquire the audio to be classified.
[0036] The audio classification unit is used to input the audio to be classified into the above-mentioned sound scene classification model to obtain the sound scene classification result of the audio to be classified.
[0037] Fifthly, embodiments of this application provide a storage medium storing a plurality of instructions adapted for loading by a processor to execute the method described in any of the preceding claims.
[0038] Sixthly, embodiments of this application provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method described in any of the above-mentioned embodiments.
[0039] In summary, the sound scene classification model generation method provided in this application involves randomly selecting source and target audio from the sound scene classification dataset; generating source and target Mel spectrograms based on the source and target audio; generating a random mask image based on the target Mel spectrogram, and obtaining an inverted random mask image; generating an enhanced spectrogram and labels based on the random mask image, the inverted random mask, the source and target Mel spectrograms; and training a preset neural network based on the enhanced spectrogram and the labels to generate a sound scene classification model. This solution enhances the training data in the sound scene classification dataset using a random mask image, avoiding the problem of poor performance of the sound scene classification model due to insufficient training data, thereby improving the accuracy of the sound scene classification results. Attached Figure Description
[0040] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0041] Figure 1 This is a flowchart illustrating the sound scene classification model generation method provided in this application embodiment.
[0042] Figure 2 This is a flowchart illustrating the sound scene classification method provided in the embodiments of this application.
[0043] Figure 3 This is a schematic diagram of the structure of the sound scene classification model generation device provided in the embodiments of this application.
[0044] Figure 4 This is a schematic diagram of the sound scene classification device provided in the embodiments of this application.
[0045] Figure 5 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0046] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0047] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, components, features, and elements with the same names in different embodiments of this application may have the same meaning or different meanings, the specific meaning of which must be determined by its interpretation in that specific embodiment or further in conjunction with the context of that specific embodiment.
[0048] It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit this application.
[0049] In the following description, the use of suffixes such as "module," "part," or "unit" to denote elements is solely for the purpose of illustrative purposes and has no specific meaning in itself. Therefore, "module," "part," or "unit" may be used interchangeably.
[0050] In the description of this application, it should be noted that the terms "upper," "lower," "left," "right," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this application. In addition, terms such as "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0051] With the development of sound scene classification technology, intelligent algorithms for sound scene classification based on deep learning have been widely used. However, sound scene classification methods based on deep learning are highly dependent on training data. In practical applications, current sound scene classification models often suffer from poor performance due to insufficient training data, which limits the application of the trained models in real-world scenarios and results in inaccurate sound scene classification results.
[0052] Based on this, embodiments of this application provide a method for generating a sound scene classification model, a sound scene classification method, an apparatus, a storage medium, and an electronic device. Specifically, the apparatus can be integrated into an electronic device. The electronic device can be a server or a terminal, etc.; wherein, the terminal can include embedded devices such as mobile phones, wearable smart devices, and tablet computers, or laptops and personal computers (PCs); the server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.
[0053] The technical solutions shown in this application will be described in detail below through specific embodiments. It should be noted that the order of description of the following embodiments is not intended to limit the priority of the embodiments.
[0054] Please see Figure 1 , Figure 1 This is a flowchart illustrating the sound scene classification model generation method provided in this application embodiment. The specific flow of the sound scene classification model generation method is as follows:
[0055] 101. Randomly select source audio and target audio from the sound scene classification dataset.
[0056] The source audio and the target audio are two audio samples randomly selected from the sound scene classification dataset.
[0057] In the specific implementation process, it is necessary to first construct a sound scene classification dataset. This dataset contains audio recordings of sound scenes from multiple scenarios, with audio recordings from the same scenario grouped into the same category. For example, this sound scene classification dataset can include audio recordings of sound scenes such as city streets, parks, offices, and stadiums.
[0058] It should be noted that each category of audio scene has a corresponding original label, which is used to represent the category to which the audio scene belongs. For example, the original label of audio scene in category A is (1,0,0,0), the original label of audio scene in category B is (0,1,0,0), the original label of audio scene in category C is (0,0,1,0), and the original label of audio scene in category D is (0,0,0,1).
[0059] 102. Generate source Mel spectrogram and target Mel spectrogram based on source speech and target speech.
[0060] In some embodiments, the source speech and the target speech can be subjected to Fast Fourier Transform to generate a first spectrogram and a second spectrogram, respectively; then, the first spectrogram and the second spectrogram can be processed by a Mel filter bank to generate a source Mel spectrogram and a target Mel spectrogram.
[0061] In the field of speech processing, Mel spectrograms are an important tool that can transform speech signals into a form that is easier to analyze and process. When generating Mel spectrograms for source and target speech, the two speech signals first need to be preprocessed, including steps such as sampling, quantization, and framing.
[0062] Specifically, for the source speech, it can first be converted into a digital signal, and then segmented into a series of short time frames using a window function. Next, a Fast Fourier Transform (FFT) is performed on each short time frame to transform it from the time domain to the frequency domain. Subsequently, a Mel filter bank is applied to filter the FFT results to obtain the Mel spectrum. Finally, by taking the logarithm of the Mel spectrum and performing a Discrete Cosine Transform (DCT), the Mel cepstral coefficients (MFCCs) of the source speech are obtained, thus constructing the source Mel spectrum.
[0063] The same processing flow applies to the target speech. First, the target speech is converted into a digital signal and processed by frame segmentation. Then, an FFT transform is performed on each short frame to obtain its frequency domain representation. Next, filtering is performed using the same Mel filter bank as the source speech to obtain the Mel spectrum of the target speech. Finally, by performing logarithmic and DCT transforms on the Mel spectrum, we obtain the MFCC of the target speech, thus constructing the target Mel spectrum.
[0064] 103. Generate a random mask image based on the target Mel spectrum image, and obtain the inverted random mask image of the random mask image.
[0065] Specifically, a matrix of the same size as the source and target spectrograms can be created, and all elements in the matrix can be initialized to 0. Then, random operations can be performed on the mask matrix, such as random rotation, cropping, or translation of the elements. In the transformed matrix, some positions are randomly selected and the corresponding elements are assigned a value of 1. After the above random operations, a binary random mask matrix (i.e., a random mask image) is obtained. Positions with a value of 1 in the matrix represent selected areas, and positions with a value of 0 represent unselected areas.
[0066] Then, an element-wise inversion operation can be performed on the random mask to obtain a new mask matrix that is complementary to the random mask, which is the inverted random mask.
[0067] 104. Generate enhanced spectrograms and labels based on random mask images, inverted random masks, source Mel spectrograms, and target Mel spectrograms.
[0068] Specifically, an enhanced spectrogram can be generated using a random mask image, an inverted random mask, a source Mel spectrogram, and a target Mel spectrogram; and a label can be generated using the random mask image, the source Mel spectrogram, and the target Mel spectrogram.
[0069] In some embodiments, a dot product operation can be performed between the random mask image and the target Mel spectrogram to generate a first intermediate result image; a dot product operation can be performed between the source Mel spectrogram and the inverted random mask to generate a second intermediate result image; and then the first intermediate result image and the second intermediate result image can be added together to obtain an enhanced spectrogram.
[0070] In a specific embodiment, the random mask image is first multiplied by the target Mel spectrogram. The dot product operation is a basic mathematical operation that multiplies corresponding elements to obtain a first intermediate result image. In this process, the random mask image selectively emphasizes or suppresses specific frequency or time regions in the target Mel spectrogram. This helps the model generalize better and learn important features from the spectrogram. In this way, the model is forced to infer the content of the masked portion from the remaining spectral information, thereby improving its robustness and generalization ability.
[0071] In addition to simple random masking strategies, more complex masking methods can be employed, such as frequency- or time-based masking strategies. These strategies can be customized based on the specific properties of the spectrogram or task requirements, thereby better extracting key information.
[0072] Subsequently, the source Mel-ray spectrogram is multiplied by the inverted random mask to generate a second intermediate result image. The inverted random mask is a mask image obtained by transforming the original random mask (e.g., inverting or negating it). By multiplying it by the source Mel-ray spectrogram, a result image with a different feature distribution or information focus than the original spectrogram can be obtained. This operation helps to introduce certain characteristics of the source spectrogram, thereby enhancing the expressiveness of the target spectrogram.
[0073] After obtaining the first and second intermediate result images, they are added together to obtain the final enhanced spectrogram. This enhanced spectrogram not only retains the key information of the target spectrogram but also incorporates certain characteristics of the source spectrogram, thus achieving optimization and enhancement of the original spectrogram.
[0074] Among these methods, spectrogram enhancement has wide applications in practice. For example, in speech synthesis tasks, spectrogram enhancement can improve the naturalness and intelligibility of synthesized speech; in speech recognition tasks, spectrogram enhancement helps improve recognition accuracy and robustness; and in audio style transfer tasks, by adjusting the random mask map and inverting the random mask, the conversion and fusion between different audio styles can be achieved.
[0075] In summary, by performing a dot product operation on the random mask image, the target Mel spectrogram, and the source Mel spectrogram and then adding them together, an enhanced spectrogram can be obtained, thereby achieving the optimization and enhancement of the audio signal.
[0076] In some embodiments, a first proportion of the random mask image in the target Mel spectrogram can be obtained; a second proportion of the random mask image in the source Mel spectrogram can be obtained; and a label can be generated based on the first and second proportions.
[0077] For example, a sound scene classification dataset contains four categories: A, B, C, and D. The original label for sound scene audio in category A is (1,0,0,0), for category B it's (0,1,0,0), for category C it's (0,0,1,0), and for category D it's (0,0,0,1). A randomly selected source speech is categorized as A, and the target speech as B. The random mask image has a first percentage (0.8) in the target Mel-ray spectrogram and a second percentage (0.2) in the source Mel-ray spectrogram. The probability of the enhanced spectrogram belonging to category A is 0.2, and the probability of it belonging to category B is 0.8. In this case, the label is (0.2, 0.8, 0, 0).
[0078] 105. Train the preset neural network based on the enhanced spectrogram and labels to generate a sound scene classification model.
[0079] In the specific implementation process, a loss function can be constructed based on the enhanced spectrogram and labels; the preset neural network can be iteratively trained by minimizing the loss function; in each iteration, the parameters of the preset neural network can be adjusted according to the training results; when the loss function value reaches the preset threshold or the number of training iterations reaches the preset number, training is stopped, and the trained neural network is saved as a sound scene classification model.
[0080] In some embodiments, to verify the performance of the sound scene classification model, a separate test dataset can be used for evaluation. The test dataset also needs to be preprocessed and spectrogram generated before being input into the model for recognition. By comparing the labels output by the model with the true labels in the test dataset, performance metrics such as accuracy and recall can be calculated, thereby evaluating the model's performance.
[0081] It should be noted that the preset neural network is a lightweight deep separable convolutional network. This network contains only 5 layers of deep separable convolutions, reducing the model's parameters and computational complexity, lowering the risk of overfitting. The lightweight nature of the network allows for acceleration using ARM operators during deployment, enabling real-time processing on low-computing-power devices. In other words, this sound scene classification model has low complexity and low memory (RAM) and computing power requirements, allowing for real-time processing on low-computing-power devices such as Bluetooth headsets and speakers.
[0082] Furthermore, this lightweight depthwise separable convolutional network decomposes ordinary convolution into two steps: depthwise convolution and pointwise convolution. Depthwise convolution operates independently on each input channel, while pointwise convolution combines the outputs from different channels. By breaking down the convolution operation into two independent steps, the number of parameters and computational cost of depthwise separable convolution are significantly reduced. Simultaneously, this combination method better captures the spatial and channel correlations in the input data, thereby improving the model's expressive power. This results in better performance in computationally limited scenarios (such as mobile devices). The additional residual connections introduced between the input and output of the depthwise separable convolutional layer improve the stability of the model during training.
[0083] In summary, the sound scene classification model generation method provided in this application includes: randomly selecting source audio and target audio from the sound scene classification dataset; generating source Mel spectrograms and target Mel spectrograms based on the source and target audio; generating a random mask map based on the target Mel spectrogram and obtaining an inverted random mask map; generating an enhanced spectrogram and labels based on the random mask map, the inverted random mask, the source Mel spectrogram, and the target Mel spectrogram; and training a preset neural network based on the enhanced spectrogram and labels to generate a sound scene classification model. This solution, by performing random rotation, cropping, translation, and random assignment operations on the random mask region, can obtain effective enhanced data, increase the effective training data volume, and avoid the problem of poor performance of the sound scene classification model due to insufficient training data, thereby improving the accuracy of the sound scene classification results. It also effectively protects the time-domain and frequency-domain information of the non-masked region, preserves the acoustic characteristics of the original audio data, and uses the enhanced spectrogram data to train the model, resulting in a sound scene classification model with better generalization performance and higher classification performance.
[0084] To illustrate the practical application process of the sound scene classification model generation provided in this application embodiment, based on the sound scene classification model generated in the above embodiments, this application embodiment also provides a sound scene classification method. This sound scene classification method can be implemented using the aforementioned electronic device, such as... Figure 2 As shown, the specific sound scene classification method can be as follows:
[0085] 201. Obtain the audio to be categorized.
[0086] The audio to be classified is audio data randomly collected in a certain sound scene.
[0087] 202. Input the audio to be classified into the sound scene classification model to obtain the sound scene classification result of the audio to be classified.
[0088] For specific implementation methods described above, please refer to the embodiments of the sound scene classification model generation method described above, which will not be repeated here.
[0089] In summary, this sound scene classification method uses a sound scene classification model generated by a sound scene classification model generation method, which can improve the accuracy of sound scene classification results.
[0090] To facilitate better implementation of the sound scene classification model generation method provided in this application, this application also provides a sound scene classification model generation apparatus. The meanings of the terms used are the same as in the sound scene classification model generation method described above, and specific implementation details can be found in the descriptions within the method embodiments.
[0091] Please see Figure 3 , Figure 3 This is a schematic diagram of the structure of the sound scene classification model generation device provided in this application embodiment. The sound scene classification model generation device may include an audio selection unit 301, a first generation unit 302, a second generation unit 303, a third generation unit 304, and a fourth generation unit 305.
[0092] The audio selection unit 301 is used to randomly select source audio and target audio from the sound scene classification dataset;
[0093] The first generation unit 302 is used to generate a source Mel spectrogram and a target Mel spectrogram based on the source speech and the target speech.
[0094] The second generation unit 303 is used to generate a random mask map based on the target Mel spectrum map and obtain the inverted random mask map of the random mask map;
[0095] The third generation unit 304 is used to generate an enhanced spectrogram and a label based on a random mask image, an inverted random mask, a source Mel spectrogram, and a target Mel spectrogram.
[0096] The fourth generation unit 305 is used to train a preset neural network based on the enhanced spectrogram and labels to generate a sound scene classification model.
[0097] For specific implementation methods of each of the above units, please refer to the embodiments of the sound scene classification model generation method described above, which will not be repeated here.
[0098] In summary, the sound scene classification model generation device provided in this application can randomly select source audio and target audio from the sound scene classification dataset through the audio selection unit 301; generate source Mel spectrogram and target Mel spectrogram based on the source and target speech by the first generation unit 302; generate a random mask map based on the target Mel spectrogram and obtain an inverted random mask map; generate an enhanced spectrogram and labels based on the random mask map, the inverted random mask, the source Mel spectrogram, and the target Mel spectrogram by the third generation unit 304; and train a preset neural network based on the enhanced spectrogram and labels to generate a sound scene classification model. This solution, by performing random rotation, cropping, translation, and random assignment operations on the random mask region, can obtain effective enhanced data, increase the effective training data volume, and avoid the problem of poor performance of the sound scene classification model due to insufficient training data, thereby improving the accuracy of the sound scene classification results. It effectively protects the time and frequency domain information of the non-masked region, preserves the acoustic characteristics of the original audio data, and uses enhanced spectrogram data to train the model, resulting in a sound scene classification model with better generalization performance and higher classification performance.
[0099] To facilitate better implementation of the sound scene classification method provided in this application, this application also provides a sound scene classification device. The meanings of the terms used are the same as in the sound scene classification method described above, and specific implementation details can be found in the descriptions within the method embodiments.
[0100] Please see Figure 4 , Figure 4 This is a schematic diagram of the sound scene classification device provided in an embodiment of this application. The sound scene classification device may include an audio acquisition unit 401 and an audio classification unit 402.
[0101] The audio acquisition unit 401 is used to acquire the audio to be classified.
[0102] The audio classification unit 402 is used to input the audio to be classified into the sound scene classification model to obtain the sound scene classification result of the audio to be classified.
[0103] For specific implementation methods of each of the above units, please refer to the embodiments of the sound scene classification method described above, which will not be repeated here.
[0104] In summary, this sound scene classification method uses a sound scene classification model generated by a sound scene classification model generation method, which can improve the accuracy of sound scene classification results.
[0105] This application also provides an electronic device that may integrate the sound scene classification model generation device or sound scene classification device of this application, such as... Figure 5As shown, it illustrates a structural schematic diagram of the electronic device involved in the embodiments of this application, specifically:
[0106] The electronic device may include a radio frequency (RF) circuit 601, a memory 602 including one or more computer-readable storage media, an input unit 603, a display unit 604, a sensor 605, an audio circuit 606, a wireless Fidelity (WiFi) module 607, a processor 608 including one or more processing cores, and a power supply 609, etc. Those skilled in the art will understand that... Figure 5 The electronic device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0107] in:
[0108] RF circuit 601 can be used for receiving and transmitting signals during information transmission or calls. Specifically, it receives downlink information from the base station and hands it over to one or more processors 608 for processing; additionally, it transmits uplink data to the base station. Typically, RF circuit 601 includes, but is not limited to, an antenna, at least one amplifier, a tuner, one or more oscillators, a Subscriber Identity Module (SIM) card, a transceiver, a coupler, a Low Noise Amplifier (LNA), a duplexer, etc. Furthermore, RF circuit 601 can also communicate wirelessly with networks and other devices. Wireless communication can use any communication standard or protocol, including but not limited to Global System for Mobile Communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), etc.
[0109] The memory 602 can be used to store software programs and modules. The processor 608 executes various functional applications and information processing by running the software programs and modules stored in the memory 602. The memory 602 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device (such as audio data, telephone directory, etc.). In addition, the memory 602 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 602 may also include a memory controller to provide access to the memory 602 for the processor 608 and the input unit 603.
[0110] The input unit 603 can be used to receive input digital or character information, and to generate keyboard, mouse, joystick, optical, or trackball signal inputs related to user settings and function control. Specifically, in one embodiment, the input unit 603 may include a touch-sensitive surface and other input devices. The touch-sensitive surface, also known as a touch display or touchpad, can collect user touch operations on or near it (e.g., user operations using fingers, styluses, or any suitable object or accessory on or near the touch-sensitive surface), and drive corresponding connection devices according to a pre-set program. Optionally, the touch-sensitive surface may include a touch detection device and a touch controller. The touch detection device detects the user's touch orientation and the signal generated by the touch operation, transmitting the signal to the touch controller; the touch controller receives touch information from the touch detection device, converts it into touch point coordinates, sends it to the processor 608, and can receive and execute commands from the processor 608. Furthermore, various types of touch-sensitive surfaces, such as resistive, capacitive, infrared, and surface acoustic wave, can be used. In addition to the touch-sensitive surface, the input unit 603 may also include other input devices. Specifically, other input devices may include, but are not limited to, one or more of the following: physical keyboard, function keys (such as volume control buttons, power buttons, etc.), trackball, mouse, joystick, etc.
[0111] Display unit 604 can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces of electronic devices. These graphical user interfaces can be composed of graphics, text, icons, video, and any combination thereof. Display unit 604 may include a display panel, optionally configured as a liquid crystal display (LCD), organic light-emitting diode (OLED), or similar form. Furthermore, a touch-sensitive surface may cover the display panel. When the touch-sensitive surface detects a touch operation on or near it, it transmits the information to processor 608 to determine the type of touch event. Subsequently, processor 608 provides corresponding visual output on the display panel according to the type of touch event. Although in Figure 5 In this context, the touch-sensitive surface and the display panel are two separate components for implementing input and output functions. However, in some embodiments, the touch-sensitive surface and the display panel can be integrated to achieve both input and output functions.
[0112] The electronic device may also include at least one sensor 605, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor, wherein the ambient light sensor can adjust the brightness of the display panel according to the ambient light level, and the proximity sensor can turn off the display panel and / or backlight when the electronic device is moved to the ear. As a type of motion sensor, a gravity acceleration sensor can detect the magnitude of acceleration in various directions (generally three axes), and can detect the magnitude and direction of gravity when stationary. It can be used for applications that recognize the phone's posture (such as landscape / portrait switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc. Other sensors that may be configured in the electronic device, such as gyroscopes, barometers, hygrometers, thermometers, and infrared sensors, will not be described in detail here.
[0113] Audio circuitry 606, a speaker, and a microphone provide an audio interface between the user and the electronic device. Audio circuitry 606 converts received audio data into electrical signals, transmits them to the speaker, and the speaker converts them into sound signals for output. Conversely, the microphone converts collected sound signals into electrical signals, which are then received by audio circuitry 606, converted back into audio data, and processed by processor 608. The processed data is then transmitted via RF circuitry 601 to, for example, another electronic device, or output to memory 602 for further processing. Audio circuitry 606 may also include an earphone jack to facilitate communication between external headphones and the electronic device.
[0114] WiFi is a short-range wireless transmission technology. Electronic devices using the WiFi module 607 can help users send and receive emails, browse web pages, and access streaming media, providing users with wireless broadband internet access. Although Figure 5 The diagram shows a WiFi module 607, but it is understood that it is not a necessary component of an electronic device and can be omitted as needed without changing the nature of the invention.
[0115] The processor 608 is the control center of the electronic device. It connects various parts of the phone via various interfaces and lines, and performs various functions and processes data by running or executing software programs and / or modules stored in the memory 602, and by calling data stored in the memory 602, thereby providing overall monitoring of the phone. Optionally, the processor 608 may include one or more processing cores; preferably, the processor 608 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 608.
[0116] The electronic device also includes a power supply 609 (such as a battery) that supplies power to various components. Preferably, the power supply can be logically connected to the processor 608 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 609 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.
[0117] Although not shown, the electronic device may also include a camera, Bluetooth module, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 608 in the electronic device loads the executable files corresponding to the processes of one or more applications into the memory 602 according to the following instructions, and the processor 608 runs the applications stored in the memory 602 to realize various functions, such as:
[0118] Source and target audio are randomly selected from the sound scene classification dataset;
[0119] Generate source Mel spectrograms and target Mel spectrograms based on source and target speech;
[0120] Generate a random mask image based on the target Mel spectrogram, and obtain the inverted random mask image of the random mask image;
[0121] Enhanced spectrograms and labels are generated based on random mask images, inverted random masks, source Mel spectrograms, and target Mel spectrograms.
[0122] A sound scene classification model is generated by training a pre-defined neural network based on enhanced spectrograms and labels.
[0123] or,
[0124] Get the audio to be categorized;
[0125] The audio to be classified is input into the sound scene classification model to obtain the sound scene classification result of the audio to be classified.
[0126] In summary, the electronic device provided in this application can randomly select source and target audio from a sound scene classification dataset; generate source and target Mel spectrograms based on the source and target speech; generate a random mask map based on the target Mel spectrogram, and obtain an inverted random mask map; generate an enhanced spectrogram and labels based on the random mask map, the inverted random mask, the source Mel spectrogram, and the target Mel spectrogram; and train a preset neural network based on the enhanced spectrogram and labels to generate a sound scene classification model. This solution, by performing random rotation, cropping, translation, and random assignment operations on the random mask region, can obtain effective enhanced data, increase the effective training data volume, and avoid the problem of poor performance of the sound scene classification model due to insufficient training data, thereby improving the accuracy of the sound scene classification results. It also effectively protects the time and frequency domain information of the non-masked region, preserves the acoustic characteristics of the original audio data, and uses the enhanced spectrogram data to train the model, obtaining a sound scene classification model with better generalization performance and higher classification performance.
[0127] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the detailed description of the sound scene classification model generation method or sound scene classification method above, which will not be repeated here.
[0128] It should be noted that, regarding the sound scene classification model generation method or sound scene classification method in the embodiments of this application, those skilled in the art will understand that all or part of the process of implementing the sound scene classification model generation method or sound scene classification method in the embodiments of this application can be accomplished by a computer program controlling the relevant hardware. The computer program can be stored in a computer-readable storage medium, such as stored in the memory of a terminal, and executed by at least one processor in the terminal. During execution, it can include the process of the embodiments of the sound scene classification model generation method or sound scene classification method.
[0129] For the sound scene classification model generation device or sound scene classification device of the embodiments of this application, its functional modules can be integrated into a processing electronic device, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0130] To this end, embodiments of this application provide a storage medium storing multiple instructions that can be loaded by a processor to execute steps in any of the sound scene classification model generation methods or sound scene classification methods provided in embodiments of this application. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), random access memory (RAM), etc.
[0131] The above provides a detailed description of the sound scene classification model generation method, sound scene classification method, device, storage medium, and electronic device provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A method for generating a sound scene classification model, characterized in that, include: Source and target audio are randomly selected from the sound scene classification dataset; Generate a source Mel spectrogram and a target Mel spectrogram based on the source speech and the target speech; A random mask image is generated based on the target Mel spectrogram, and an inverted random mask image is obtained from the random mask image. An enhanced spectrogram and a label are generated based on the random mask image, the inverted random mask, the source Mel spectrogram, and the target Mel spectrogram; The preset neural network is trained based on the enhanced spectrogram and the labels to generate a sound scene classification model.
2. The sound scene classification model generation method as described in claim 1, characterized in that, The generation of enhanced spectrograms and labels based on the random mask image, the inverted random mask, the source Mel spectrogram, and the target Mel spectrogram includes: An enhanced spectrogram is generated using the random mask image, the inverted random mask, the source Mel spectrogram, and the target Mel spectrogram; A label is generated using the random mask image, the source Mel spectrogram, and the target Mel spectrogram.
3. The sound scene classification model generation method as described in claim 2, characterized in that, The step of generating an enhanced spectrogram using the random mask image, the inverted random mask, the source Mel spectrogram, and the target Mel spectrogram includes: The random mask image is multiplied by the target Mel spectrum image to generate a first intermediate result image. The source Mel spectrum is multiplied by the inverted random mask to generate a second intermediate result image. The first intermediate result image is added to the second intermediate result image to obtain the enhanced spectrum image.
4. The sound scene classification model generation method as described in claim 2, characterized in that, The step of generating a label using the random mask image, the source Mel spectrogram, and the target Mel spectrogram includes: Obtain the first proportion of the random mask image in the target Mel spectrogram; Obtain the second proportion of the random mask image in the source Mel spectrogram; The label is generated based on the first proportion and the second proportion.
5. The sound scene classification model generation method as described in claim 1, characterized in that, The step of generating source Mel spectrograms and target Mel spectrograms based on the source speech and the target speech includes: Perform Fast Fourier Transform on the source speech and the target speech respectively to generate a first spectrogram and a second spectrogram; The first and second spectrograms are processed using a Mel filter bank to generate a source Mel spectrogram and a target Mel spectrogram.
6. A sound scene classification method, characterized in that, include: Get the audio to be categorized; The audio to be classified is input into the sound scene classification model as described in any one of claims 1-5 to obtain the sound scene classification result of the audio to be classified.
7. A device for generating a sound scene classification model, characterized in that, include: The audio selection unit is used to randomly select source and target audio from the sound scene classification dataset; The first generation unit is configured to generate a source Mel spectrogram and a target Mel spectrogram based on the source speech and the target speech; The second generation unit is used to generate a random mask map based on the target Mel spectrogram and obtain an inverted random mask map of the random mask map; The third generation unit is used to generate an enhanced spectrogram and a label based on the random mask image, the inverted random mask, the source Mel spectrogram, and the target Mel spectrogram; The fourth generation unit is used to train a preset neural network based on the enhanced spectrogram and the label to generate a sound scene classification model.
8. A sound scene classification device, characterized in that, include: The audio acquisition unit is used to acquire the audio to be classified. An audio classification unit is used to input the audio to be classified into the sound scene classification model as described in any one of claims 1-5, and obtain the sound scene classification result of the audio to be classified.
9. A storage medium, characterized in that, The storage medium stores a plurality of instructions adapted for loading by a processor to execute the method according to any one of claims 1-5 or 6.
10. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method as described in any one of claims 1-5 or 6.
Citation Information
Patent Citations
Spectrum mask model training method and audio scene recognition method and system
CN111028861A
Data processing system and method of speech recognition model, and speech recognition method
CN115762489A