Audio scene classification method, device, terminal device and storage medium
By extracting frequency and time features from audio data, using deep separable convolutional neural networks and recurrent neural networks for audio scene classification, combined with the XGBoost model, the problem of low accuracy in audio scene classification in the existing technology is solved, and efficient audio scene recognition and processing in noisy environments is achieved.
Patent Information
- Application Number
- CN202111282304.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-01
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2041-11-01
AI Technical Summary
The existing audio scene classification methods are low in accuracy when facing audio data in similar scenes, especially in noisy environments, it is difficult to effectively distinguish different types of vehicle noise.
By collecting audio data in the target scene, extracting frequency and time feature information, dividing it into feature segments according to preset rules and performing fusion processing, the frequency and time domain features can be extracted using deep separation convolutional neural networks and recurrent neural networks, and scene classification is performed by combining the XGBoost model.
It improves the accuracy of audio scene classification, can better identify and process audio scenes in noisy environments, and improves the audio quality of headphones.
Smart Images

Figure CN114186094B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of computer technology, and particularly relates to an audio scene classification method, apparatus, terminal device, and storage medium. Background Art
[0002] With the progress of the times and the improvement of the public's quality, the uncivilized behavior of playing sounds externally in public places has gradually decreased. Therefore, the functional requirements for headphones in people's daily lives have also increased accordingly. For example, in a noisy environment, users also hope that the audio quality heard through headphones will not decrease.
[0003] Currently, in order to avoid as much as possible the mixing of external noisy environmental sounds (such as various noises like vehicle sounds, crowd sounds, and the sounds of merchants' players) in the audio, resulting in a decrease in the audio quality heard by users through headphones, it is usually necessary to process the external audio signals heard by users. The specific method is to effectively judge the environmental sounds through audio scene classification, and then perform corresponding processing on the environmental sounds according to different audio scene classification results. However, the existing audio scene classification method obtains the audio scene classification result by calculating the frequency characteristics of the audio. When facing audio data of similar scenes (such as different types of transportation tools), the accuracy of the audio scene classification result obtained by using the above classification method is low. Summary of the Invention
[0004] Embodiments of this application provide an audio scene classification method, apparatus, terminal device, and storage medium, which can effectively improve the accuracy of the audio scene classification result.
[0005] In a first aspect, embodiments of this application provide an audio scene classification method, including:
[0006] Collect audio data of a first preset duration in a target scene to obtain a first audio segment;
[0007] Extract first feature information and second feature information of the first audio segment, where the first feature information represents frequency feature information based on time, and the second feature information represents time feature information based on frequency;
[0008] Divide the first feature information into N first feature segments and divide the second feature information into N second feature segments respectively according to a preset rule, where N is a positive integer greater than 1;
[0009] Divide the N first feature segments and the N second feature segments into N feature groups according to the preset rule, where each feature group includes one first feature segment and one second feature segment;
[0010] Calculate the fusion feature of each of the N feature groups;
[0011] Determine a first scene classification result of the first audio clip according to the fusion features of each of the N feature groups.
[0012] In a second aspect, an embodiment of the present application provides an audio scene classification device, including:
[0013] A first acquisition module, configured to acquire audio data of a first preset duration in a target scene to obtain a first audio clip;
[0014] A first processing module, configured to extract first feature information and second feature information of the first audio clip, where the first feature information represents frequency feature information based on time, and the second feature information represents time feature information based on frequency;
[0015] A second processing module, configured to divide the first feature information into N first feature segments and divide the second feature information into N second feature segments respectively according to a preset rule, where N is a positive integer greater than 1;
[0016] A third processing module, configured to divide the N first feature segments and the N second feature segments into N feature groups according to the preset rule, where each feature group includes a first feature segment and a second feature segment;
[0017] A calculation module, configured to calculate the fusion features of each of the N feature groups;
[0018] A first classification processing module, configured to determine a first scene classification result of the first audio clip according to the fusion features of each of the N feature groups.
[0019] In a third aspect, an embodiment of the present application provides a terminal device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, where the processor implements the audio scene classification method according to any one of the above first aspects when executing the computer program.
[0020] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and the computer program implements the audio scene classification method according to any one of the above first aspects when being executed by a processor.
[0021] In a fifth aspect, an embodiment of the present application provides a computer program product, which causes a terminal device to execute the audio scene classification method according to any one of the above first aspects when the computer program product runs on the terminal device.
[0022] The beneficial effects of the embodiments of the first aspect of the present application compared with the prior art are as follows: First, audio data for a first preset duration in a target scenario is collected to obtain a first audio segment; then, first feature information and second feature information of the first audio segment are extracted, where the first feature information represents frequency feature information based on time, and the second feature information represents time feature information based on frequency; according to a preset rule, the first feature information is divided into N first feature segments and the second feature information is divided into N second feature segments respectively, where N is a positive integer greater than 1; according to the preset rule, the N first feature segments and the N second feature segments are divided into N feature groups, where each feature group includes one first feature segment and one second feature segment; then, the fusion features of the N feature groups are calculated respectively; since each feature group fuses the frequency domain features and time domain features of the first audio segment to obtain a set of feature expressions with both time-frequency characteristics, and it does not lose the information of the audio signal on the time axis, thereby improving the utilization degree of the time domain features. Finally, according to the fusion features of the N feature groups, the first scene classification result of the first audio segment is determined. In this way, the feature groups in the two dimensions of the frequency domain and the time domain are used together to determine the first scene classification result of the first audio frequency band, thereby improving the accuracy of the audio scene classification result.
[0023] It can be understood that the beneficial effects of the above-mentioned second aspect to the fifth aspect can refer to the relevant descriptions in the above-mentioned first aspect, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0025] Figure 1 It is a schematic flowchart of an audio scene classification method provided by an embodiment of the present application;
[0026] Figure 2 is Figure 1 a schematic diagram of the specific implementation process of S105 in
[0027] Figure 3 It is a schematic structural diagram of a spectrum feature information extraction model provided by an embodiment of the present application;
[0028] Figure 4 It is a schematic structural diagram of a timing feature information extraction model provided by an embodiment of the present application;
[0029] Figure 5Schematic flowchart of the audio scene classification method provided by another embodiment of this application;
[0030] Figure 6 Structural schematic diagram of the audio scene classification device provided by an embodiment of this application;
[0031] Figure 7 Structural schematic diagram of the terminal device provided by an embodiment of this application. Detailed implementation manners
[0032] In the following description, specific details such as specific system architectures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of this application. However, those skilled in the art should clearly understand that this application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of this application.
[0033] It should be understood that when used in the specification and appended claims of this application, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0034] As used in the specification and appended claims of this application, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" according to the context. Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]" according to the context.
[0035] In addition, in the description of the specification and appended claims of this application, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0036] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that specific features, structures, or characteristics described in connection with that embodiment are included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in some other embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized.
[0037] Figure 1 Fig. shows a schematic flowchart of the audio scene classification method provided by this application. Referring to Figure 1 , the details of this audio scene classification method are as follows:
[0038] S101, collect audio data of a first preset duration in the target scene to obtain a first audio segment.
[0039] In this step, the audio data of the first preset duration (such as 10 s) in the target scene can be obtained through an audio collection module (such as a microphone) set on the terminal device. Here, the target scene is the environmental scene where the user is located when the terminal device outputs audio to the user's ear through the headphone device.
[0040] For the specific content of the solution in this application, reference can be made to the description in Embodiment 1 below.
[0041] S102, extract first feature information and second feature information of the first audio segment, where the first feature information represents frequency feature information based on time, and the second feature information represents time feature information based on frequency.
[0042] Optionally, the implementation process of extracting the first feature information of the first audio segment may include:
[0043] First, perform a short-time Fourier transform on the first audio segment to obtain the time-frequency signal of the first audio segment. Then, perform filtering processing (such as Mel filtering) on the time-frequency signal of the first audio segment. Finally, perform a logarithmic calculation on the filtered time-frequency signal to obtain a spectrogram that changes with time of the first audio segment, that is, the first feature information of the first audio segment. This first feature information represents frequency feature information based on time.
[0044] It should be understood that the frequency feature information based on time is to statistically analyze the frequency features corresponding to different times. For example, statistically analyze the frequency values corresponding to different times, that is, the horizontal axis is time and the vertical axis is the frequency value.
[0045] Here, the second feature information represents time feature information based on frequency. Specifically, the second feature information can be represented as time variation information based on frequency amplitude.
[0046] It should be understood that the time feature information based on frequency is to statistically analyze the time features corresponding to different frequencies. For example, statistically analyze the time corresponding to different frequencies, that is, the horizontal axis is frequency and the vertical axis is time.
[0047] S103, divide the first feature information into N first feature segments and divide the second feature information into N second feature segments respectively according to a preset rule, where N is a positive integer greater than 1;
[0048] In this step, the preset rule can be a preset time interval or a preset frequency interval.
[0049] Continuing with the above example, if the first feature information and the second feature information are divided according to a preset time interval, the first feature segment is equivalent to the feature segment obtained by dividing the horizontal axis of the first feature information; the second feature segment is equivalent to the feature segment obtained by dividing the vertical axis of the second feature information.
[0050] If the first feature information and the second feature information are divided according to a preset frequency interval, the first feature segment is equivalent to the feature segment obtained by dividing the vertical axis of the first feature information; the second feature segment is equivalent to the feature segment obtained by dividing the horizontal axis of the second feature information.
[0051] S104, divide the N first feature segments and the N second feature segments into N feature groups according to the preset rule, where each feature group includes one first feature segment and one second feature segment.
[0052] Here, taking the preset rule as a preset time interval as an example, according to the above preset time interval, the first feature segment and the second feature segment within the same time interval are divided into one feature group, so that the N first feature segments and the N second feature segments are divided into N feature groups.
[0053] Taking the preset rule as a preset frequency interval as an example, according to the above preset frequency interval, the first feature segment and the second feature segment within the same frequency range are divided into one feature group, so that the N first feature segments and the N second feature segments are divided into N feature groups.
[0054] S105, calculate the respective fusion features of the N feature groups.
[0055] Here, calculate the fusion feature of each of the N feature groups respectively to obtain the respective fusion features of the N feature groups.
[0056] Specifically, for each of the N feature groups, the first feature segment and the second feature segment in the feature group are weighted and fused to calculate the fused feature of the feature group.
[0057] It should be noted that by effectively extracting more representative feature information, each feature group contains a first feature segment representing frequency features and a second feature segment representing time features. By fusing the two, a set of feature expressions with both time-frequency characteristics is obtained. It takes into account not only the frequency information of the audio signal but also the information on the time axis of the audio signal, thereby improving the utilization degree of time-domain features. Moreover, by increasing the parameter dimension (i.e., time-domain features), it helps to improve the accuracy of the scene classification result.
[0058] For the specific content of the solution in this application, refer to the description in Embodiment 2 below.
[0059] S106. Determine the first scene classification result of the first audio segment according to the fused features of the N feature groups respectively.
[0060] In this step, specifically, the fused features of each of the N feature groups are respectively input into the trained classification model to determine the first scene classification result of the first audio segment.
[0061] Here, the classification model can be a gradient boosting model, such as the XGBoost (Extreme Gradient Boosting) model. The classification model can be pre-trained to obtain the trained classification model. When applied, using the trained classification model to calculate the scene classification result of the audio segment can improve the calculation accuracy.
[0062] For the specific content of the solution in this application, refer to the description in Embodiment 3 below.
[0063] Embodiment 1
[0064] In a possible implementation manner, the implementation process of step S101 may include:
[0065] S1011. Collect the original audio of the first preset duration in the target scene.
[0066] Here, the original audio of the first preset duration in the target scene can be collected through the microphone on the terminal device. The volume of the original audio can be obtained through volume detection.
[0067] S1012. When the volume of the original audio is greater than the preset threshold, determine whether there is a target sound in the original audio.
[0068] Among them, if the volume of the original audio is less than the preset threshold (such as 50 dB), it is determined that the user is currently in a quiet environment. In order to better reduce power consumption and improve the standby time of the device, the subsequent operations of the method of this application do not need to be executed in this case.
[0069] If the volume of the original audio is greater than the preset threshold, it is determined that the user is currently in a relatively noisy environment. It is necessary to determine the scene classification result in this scenario, so as to provide an accurate basis for the subsequent effective judgment of the ambient sound, and then enable the device to perform corresponding processing on the ambient sound according to different audio scene classification results, so as to achieve the purpose that the audio quality heard by the user through the earphone will not be reduced finally.
[0070] In this step, the target sound specifically refers to the prominent human voice, that is, the user's own voice. Judging whether there is a target sound in the original audio is because during the use of the device, since the user is very close to the microphone, the human voice emitted by the user collected by the microphone is stronger, and may even cover up the representative sound features in the environment. Therefore, if there is a relatively strong human voice, it is necessary to separate the human voice and the ambient sound of the audio, filter the human voice, and retain the ambient sound.
[0071] It should be noted that parameters such as bone conduction and accelerometer can be used to judge whether the audio contains the target sound.
[0072] S1013, if there is a target sound in the original audio, filter out the target sound in the original audio to obtain the first audio segment.
[0073] This situation corresponds to the case where there is a target sound, that is, a prominent human voice, in the original audio with a volume greater than the preset threshold. Then, filter out the target sound in the original audio to obtain the first audio segment.
[0074] S1014, if there is no target sound in the original audio, determine the original audio as the first audio segment.
[0075] This situation corresponds to the case where there is no target sound in the original audio with a volume greater than the preset threshold. That is to say, the microphone does not collect the prominent human voice of the user. Then, directly determine the original audio as the first audio segment.
[0076] Embodiment 2
[0077] Refer to Figure 2 In a possible implementation manner, the implementation process of step S105 may include:
[0078] Step S1051, input the first feature segment in each feature group of the N feature groups into the spectral feature information extraction model to obtain the frequency domain features of the first feature segment in each feature group.
[0079] Optionally, the spectral feature information extraction model is trained by a depthwise separable convolutional neural network.
[0080] Specifically, the spectral feature information extraction model includes a set of consecutive depthwise separable convolutional layers, which can effectively reduce the number of parameters and the complexity of the neural network model. In practical engineering applications, this can improve the running speed and reduce power consumption.
[0081] In one example, referring to Figure 3 , the spectral feature information extraction model includes an input layer, a convolutional layer Conv1, a Dropout layer dropout1, a depthwise separable convolutional layer dsConv1, a depthwise separable convolutional layer dsConv2, a Dropout layer dropout2, a depthwise separable convolutional layer dsConv3, a depthwise separable convolutional layer dsConv4, a Dropout layer dropout3, a depthwise separable convolutional layer dsConv5, a depthwise separable convolutional layer dsConv6, and an average pooling layer pool1.
[0082] Next, referring to Figure 3 , in combination with the first feature segment in the feature group of the first audio segment of this application, the functions of each component in the spectral feature information extraction model are specifically described.
[0083] (1) Input layer: The first feature segment in each feature group of a single audio segment is used as the input of the model and input into this input layer.
[0084] Specifically, the first feature information extracted from the first audio segment is segmented into consecutive non-overlapping feature segments. The feature matrix dimension of a single feature segment is 128*128, and this feature segment is used as the input of the network model.
[0085] (2) Convolutional layer Conv1: The convolutional layer is a 2D convolutional layer. This convolutional layer contains 16 filters, the size of the filter is 2*2, and the activation function used is the relu function, which can introduce sparsity into the network and improve the training speed.
[0086] Specifically, the output of the convolutional layer Conv1 is a feature matrix with a dimension of 128*128*16, and each filter contains 128*128 = 16384 weight values.
[0087] (3) Dropout layer dropout1: To prevent overfitting of the model, a Dropout layer is added to randomly "deactivate" some neurons in the network, that is, assign zero weights. The dropout ratio is set to 0.2, and the size of the output matrix is still 128*128*16.
[0088] (4) Separable Convolution Layer dsConv1: This separable convolution layer is a 2D depthwise separable convolution layer, including a depthwise convolution operation and a pointwise convolution operation.
[0089] Among them, the depthwise convolution operation is called per-channel convolution. Each convolution kernel is responsible for one channel, and one channel can only be convolved by one convolution kernel. The number of feature maps after convolution is the same as the number of input channels. The depthwise convolution only operates independently on each channel of the input layer and cannot effectively utilize the spatial information between different channels. Therefore, a pointwise convolution operation is required to operate on the feature maps generated by the pointwise convolution.
[0090] Here, the pointwise convolution operation is called pointwise convolution, and the convolution kernel size is fixed at 1*1. The result of the convolution layer Conv1 will be input into the separable convolution layer dsConv1. In the separable convolution layer dsConv1, the number of convolution kernels is 16, the size of the depthwise convolution kernel is 3*3, and the size of the output matrix is 64*64*16.
[0091] (5) Separable Convolution Layer dsConv2: This separable convolution layer is a 2D depthwise separable convolution layer. The result of the separable convolution layer dsConv1 will be input into the separable convolution layer dsConv2. In the separable convolution layer dsConv2, the number of convolution kernels is updated to 32, and the rest of the settings follow the same logic as the separable convolution layer dsConv1. The size of the output matrix is 32*32*32.
[0092] (6) Dropout Layer dropout2: Following the same logic as the Dropout layer dropout1, the size of the output matrix remains 32*32*32.
[0093] (7) Separable Convolution Layer dsConv3: This separable convolution layer is a 2D depthwise separable convolution layer. The result of the Dropout layer dropout2 will be input into the separable convolution layer dsConv3. In the separable convolution layer dsConv3, the number of convolution kernels is updated to 64, and the rest of the settings follow the same logic as the separable convolution layer dsConv1. The size of the output matrix is 16*16*64.
[0094] (8) Separable Convolution Layer dsConv4: This separable convolution layer is a 2D depthwise separable convolution layer. The result of separable convolution layer dsConv3 will be input into separable convolution layer dsConv4. In separable convolution layer dsConv4, the number of convolutional kernels is updated to 128, and the rest of the settings follow the same logic as separable convolution layer dsConv1. The size of the output matrix is 8*8*128.
[0095] (9) Dropout Layer dropout3: Following the same logic as Dropout layer dropout1, the size of the output matrix remains 8*8*128.
[0096] (10) Separable Convolution Layer dsConv5: This separable convolution layer is a 2D depthwise separable convolution layer. The result of Dropout layer dropout3 will be input into separable convolution layer dsConv5. In separable convolution layer dsConv5, the number of convolutional kernels is updated to 256, and the rest of the settings follow the same logic as separable convolution layer dsConv1. The size of the output matrix is 4*4*256.
[0097] (11) Separable Convolution Layer dsConv6: This separable convolution layer is a 2D depthwise separable convolution layer. The result of separable convolution layer dsConv5 will be input into separable convolution layer dsConv6. In separable convolution layer dsConv6, the number of convolutional kernels is updated to 512, and the rest of the settings follow the same logic as separable convolution layer dsConv1. The size of the output matrix is 2*2*512.
[0098] (12) Average Pooling Layer pool1: To reduce the computational load and prevent overfitting, the average pooling method is used to process the feature map after convolution. The pooling size is 2*2, and the size of the output matrix is 1*1*512.
[0099] (13) Flatten Layer: The result of average pooling layer pool1 will be input into the Flatten layer (not shown in the figure), flattening the 1*1*512 matrix into a vector of length 512.
[0100] Here, this 512-dimensional vector will subsequently be fused with the second feature segment in each feature group of a single audio segment extracted by the temporal feature information extraction model.
[0101] It should be noted that the general process of training this model is as follows: First, initialize the convolutional layer and the Dropout layer respectively, initialize the bias to all zeros, and then input the first feature segment in each feature group of a single audio segment into the separable convolution layer, update the weights, and train the network model.
[0102] Among them, the cross-entropy objective function is adopted for model training, and the network parameters are updated through the backpropagation algorithm. When the validation set error is less than or equal to the expected value or the maximum training loop count is reached, the training ends.
[0103] In this application, by extracting the frequency domain features of the first feature segment of each feature group of the first audio segment through consecutive depthwise separable convolutional layers, the number of parameters can be significantly reduced and the operation speed can be improved. At the same time, in actual engineering applications, under the same power consumption requirements, a deeper network layer can be achieved compared to the ordinary convolutional structure, thereby improving the classification effect.
[0104] Step S1052: Input the second feature segment in each of the N feature groups into the temporal feature information extraction model to obtain the temporal domain features of the second feature segment in each feature group.
[0105] Optionally, the temporal feature information extraction model is trained by a recurrent neural network. Specifically, the recurrent neural network can be an LSTM (Long Short-Term Memory) neural network.
[0106] Specifically, the temporal feature information extraction model includes a group of LSTM layers with different hidden nodes, which can extract the temporal internal connection of the audio segment to obtain the temporal expression of the audio segment, that is, the temporal domain features of the second feature segment.
[0107] In an example, referring to Figure 4 , the temporal feature extraction model includes an input layer, an LSTM layer L1, an LSTM layer L2, an LSTM layer L3, and a Flatten layer (not shown in the figure).
[0108] Next, the functions of each component in the temporal feature extraction model will be specifically described in combination with the second feature segment in the feature group of the first audio segment of this application.
[0109] (1) Input layer: The second feature segment in each feature group of a single audio segment is used as the input of this model and input into this input layer.
[0110] Specifically, the first feature information extracted from the first audio segment is sliced into continuous and non-overlapping feature segments. The feature matrix dimension of a single feature segment is 128*128, and this feature segment is used as the input of the network model, that is, the dimension of the output matrix is 128*128.
[0111] (2) LSTM layer L1: The output result of the input layer will be input into the LSTM layer L1.
[0112] Specifically, referring to Figure 4, the LSTM structure consists of an input gate, a forget gate, and an output gate, which can retain the state of the previous moment and obtain the connection of features in the time domain. Among them, the LSTM layer L1 is defined with 32 hidden nodes, the activation function is the relu function, the return sequence is selected, and the dropout ratio is set to 0.2. Each output gate corresponds to an output, and the size of the output matrix is 128*32.
[0113] (3) LSTM layer L2: The output result of the LSTM layer L1 will be input into the LSTM layer L2. The number of hidden nodes in the LSTM layer L2 is updated to 8, and the rest of the settings follow the same logic as the LSTM layer L1. The size of the output matrix is 128*8.
[0114] (4) LSTM layer L3: The output result of the LSTM layer L2 will be input into the LSTM layer L3. The number of hidden nodes in the LSTM layer L3 is updated to 1, and the rest of the settings follow the same logic as the LSTM layer L1. The size of the output matrix is 128*1.
[0115] (5) Flatten layer: The output result of the LSTM layer L3 will be input into the Flatten layer, and the two-dimensional matrix of 128*1 will be flattened into a vector of length 128.
[0116] Here, this 128-dimensional vector will be fused with the first feature segment in each feature group of a single audio segment extracted by the spectral feature information extraction model later.
[0117] It should be noted that the general process of training this model is as follows: Initialize the bias to all zeros, input the data of each frequency point in different time tenses into the LSTM layer for training. There are 128 frequency points in this example, and the frequency range is from 20Hz to 8KHz. Among them, the cross-entropy objective function is used for training this model, and the network parameters are updated through the backpropagation algorithm. When the validation set error is less than or equal to the expected value or reaches the maximum number of training loops, the training ends.
[0118] Step S1053, perform a fusion process on the frequency domain features of the first feature segments and the time domain features of the second feature segments of the N feature groups respectively to obtain the fusion features of the N feature groups respectively.
[0119] Optionally, the implementation process of step S1053 may include:
[0120] Vector splice the frequency domain features of the first feature segments and the time domain features of the second feature segments in each of the N feature groups according to a preset weight to obtain the fusion features of the N feature groups respectively.
[0121] Specifically, first, multiply the frequency-domain features of the first feature segment in each of the N feature groups by a first preset weight, and multiply the time-domain features of the second feature segment in each of the N feature groups by a second preset weight. Then, perform vector concatenation on the frequency-domain features of the first feature segment and the time-domain features of the second feature segment in each processed feature group to obtain the fusion features of the respective N feature groups. The sum of the first preset weight and the second preset weight is 1.
[0122] Optionally, the implementation process of this step is as follows:
[0123] For each of the N feature groups, obtain the fusion feature of each feature group through F = concatenate(α * FP, (1 - α) * FT); where F represents the fusion feature of each feature group, concatenate() represents the vector concatenation operation, α represents the first preset weight, FP represents the first feature segment in each feature group, and FT represents the second feature segment in each feature group.
[0124] It should be noted that the value of α is selected based on experience. Optionally, α is 0.2.
[0125] Further, according to the calculated fusion features of each feature group, calculate the frequency of the fusion feature in each feature group; according to the frequency of the fusion feature in each feature group, obtain the fusion feature distribution of each feature group. For example, obtain the fusion feature distribution histogram of each feature group.
[0126] Embodiment III
[0127] In a possible implementation manner, the implementation process of step S106 may include:
[0128] S1061, input the respective fusion features of the N feature groups into the trained classification model to obtain N probability matrices.
[0129] When the classification model is an XGBoost model, optimize the objective function through a tree structure and use the leaf nodes to output the first scene classification result of the first audio segment. The XGBoost model is briefly described below. The predicted value of the XGBoost model for a certain sample is:
[0130]
[0131] where, f k is the base learner, and the final model is a combination of multiple base learners. The initial objective function of this model is:
[0132]
[0133] Among them, represents the predicted values of the first t - 1 ensemble learners for the samples, and f t (X i ) is the predicted value of the current learner for the samples, and Ω(f t ) is the regularization term of the t-th learner. Then, perform a second-order Taylor expansion on the objective function and specify the regularization as the following formula:
[0134]
[0135] Among them, T represents the number of leaf nodes. In this model, the number of leaf nodes is used as the L1 regularization term, and the weight values of the leaf nodes are used as the L2 regularization term. The weight of the leaf node is actually the predicted value. After simplifying the derivative, the objective function is finally determined as:
[0136]
[0137] Among them, g i is the first derivative of the l function with respect to , and h i is the second derivative of the l function with respect to , and I j ={i|q(X i ) = j}, and q(X i ) = j means that the sample X i is partitioned into the leaf node numbered j.
[0138] It should be noted that the maximum depth of the XGBoost model can be set in advance, the object objective can be set in advance, such as set to binary, and the logic can adopt binary logistic regression. The learning rate eta can be set in advance, such as set to 0.1.
[0139] Here, the samples described in the above model refer to the fusion features of the feature groups of audio sample segments in this application, specifically referring to the fusion feature distribution of the feature groups of audio sample segments. That is, the fusion feature distribution of the feature groups of audio sample segments is input into this model for training, and finally the trained model is obtained. The output of the trained model here is the probability matrix of the feature group.
[0140] S1062. Calculate the mean of the N probability matrices to obtain the final probability matrix.
[0141] In step S1062, calculating the mean of the N probability matrices specifically includes: summing and averaging the values at the corresponding positions of each probability matrix in the N probability matrices to obtain a final probability matrix.
[0142] S1063. Determine the class label to which the element with the largest value in the final probability matrix belongs as the first scene classification result of the first audio segment.
[0143] Here, the class label to which the element with the largest value in the final probability matrix belongs, that is, the class label with the highest probability in the target scene (i.e., the scene classification result), is the class label to which this element belongs.
[0144] In applications, the audio data in the target scene is collected in real time. Due to the continuity of real-time recording, there may be individual mutations in consecutive classification results, but such mutations can be ignored. Therefore, in order to ensure the robustness of the audio scene classification result, in a possible implementation manner, after step S106 of the method in the embodiments of the present application, it further includes:
[0145] S107. Continue to collect the audio data of the second preset duration in the target scene to obtain a second audio segment.
[0146] For the specific implementation manner corresponding to step S107, reference can be made to the description in the above step S101 part, which will not be elaborated here.
[0147] S108. Determine the second scene classification result of the second audio segment according to the fusion feature of the second audio segment;
[0148] For the specific implementation manner of obtaining the fusion feature of the second audio segment in step S108, reference can be made to the description in steps S102 - S105 part, and for the specific implementation manner corresponding to this step, reference can be made to the description in step S106 part, which will not be elaborated here.
[0149] S109. Determine the final scene classification result of the target scene according to the first scene classification result of the first audio segment and the second scene classification result of the second audio segment.
[0150] The implementation process of step S109 may specifically include:
[0151] Determine the classification result with the most occurrences among all scene classification results as the target final scene classification result, and all scene classification results include the first scene classification result and the second scene classification result.
[0152] Here, if the classification result with the most occurrences includes two different classification results, then by comparing the magnitudes of the largest values in their respective final probability matrices, the class label to which the element corresponding to the larger value belongs can be determined as the final scene classification result.
[0153] It should be noted that comparing this implementation manner with Figure 1As can be seen from the combination of the implementation manners shown, the final scene classification result of the target scene is determined based on the scene classification results of multiple collected audio segments. That is, the final scene classification result of the target scene is obtained through double determination. Here, the specific number of times of collecting audio segments can be set according to the actual situation and is not specifically limited herein.
[0154] It should be understood that the first determination in the double determination is manifested as the implementation manner as Figure 1 shown, and the second determination is manifested as this implementation manner.
[0155] For example, the classification result determination process can refer to Figure 5 , first taking a feature group (the feature group corresponds to an audio segment with a recording length of 1.28 s) as the input feature of the scene classification system, that is, each feature group outputs a group of probability matrices. The first determination is manifested as selecting eight consecutive feature groups (that is, corresponding to 8 audio segments with a length of 1.28 s, and these eight audio segments as a whole can be understood as the first audio segment described in the above embodiment, corresponding to Figure 5 where wav_index = 8), and taking the average value of the probability matrices output by them, and selecting the class label corresponding to the index with the largest probability as the output result of the current overall audio segment; the second determination is manifested as retaining the class with the most occurrences as the final classification result within a processing time of one minute (corresponding to Figure 5 where wav_time = 6) in
[0156] In practical applications, currently there is audio data M with a recording duration of one minute by a microphone. First, the audio is segmented into continuous and non-overlapping audio segments m1, m2,..., m6 with a duration of 10 s, and then the time-frequency features of a single audio segment m i are extracted and divided into 8 feature groups according to a preset time interval. Among them, the 8 feature groups are respectively equivalent to corresponding to continuous and non-overlapping audio segments m i1 , m i2 ,..., m i8 with a duration of 1.28 s. After that, the respective fusion features of the 8 feature groups are respectively input into the scene classification system (that is, the trained classification model), and 8 probability matrices are output. Specifically, it can be expressed by the following formula:
[0157] The current classification system is defined as The relationship between the system output and the input is:
[0158]
[0159] Then the output result of a single 10-s audio is:
[0160]
[0161] Among them, labels[·] represents the classification label matrix, and argmax(·) represents the index for obtaining the maximum value of the probability matrix.
[0162] Finally, the output results of the audio segments m1, m2, …, m6 are calculated respectively, and the result with the most frequent occurrences is calculated, which is the final scene classification result of the audio data M.
[0163] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0164] Corresponding to the method described in the above embodiments, Figure 6 The structural block diagram of the audio scene classification device provided by the embodiment of the present application is shown. For the convenience of description, only the parts related to the embodiment of the present application are shown.
[0165] Referring to Figure 6 , the device 200 may include: a first acquisition module 210, a first processing module 220, a second processing module 230, a third processing module 240, a calculation module 250, and a first classification processing module 260.
[0166] Among them, the first acquisition module 210 is used to acquire audio data of a first preset duration in a target scene to obtain a first audio segment; the first processing module 220 is used to extract the first feature information and the second feature information of the first audio segment, where the first feature information represents frequency feature information based on time, and the second feature information represents time feature information based on frequency; the second processing module 230 is used to divide the first feature information into N first feature segments and divide the second feature information into N second feature segments respectively according to a preset rule, where N is a positive integer greater than 1; the third processing module 240 is used to divide the N first feature segments and the N second feature segments into N feature groups according to the preset rule, where each feature group includes one first feature segment and one second feature segment; the calculation module 250 is used to calculate the fusion features of the N feature groups respectively; the first classification processing module 260 is used to determine the first scene classification result of the first audio segment according to the fusion features of the N feature groups respectively.
[0167] In a possible implementation manner, the acquisition module 210 may specifically be used for:
[0168] Collect the original audio for the first preset duration in the target scenario; when the volume of the original audio is greater than a preset threshold, determine whether there is a target sound in the original audio; if there is a target sound in the original audio, filter out the target sound in the original audio to obtain the first audio segment; if there is no target sound in the original audio, determine the original audio as the first audio segment.
[0169] In a possible implementation, the calculation module 250 may specifically include:
[0170] The first calculation unit is configured to input the first feature segment in each of the N feature groups into the spectral feature information extraction model to obtain the frequency domain features of the first feature segment in each feature group;
[0171] The second calculation unit is configured to input the second feature segment in each of the N feature groups into the temporal feature information extraction model to obtain the temporal domain features of the second feature segment in each feature group;
[0172] The third calculation unit is configured to perform a fusion process on the frequency domain features of the first feature segment and the temporal domain features of the second feature segment in each of the N feature groups to obtain the fusion features of each of the N feature groups.
[0173] Optionally, the spectral feature information extraction model is trained by a depthwise separable convolutional neural network; the temporal feature information extraction model is trained by a recurrent neural network.
[0174] In a possible implementation, the third calculation unit is specifically configured to:
[0175] Perform vector splicing on the frequency domain features of the first feature segment and the temporal domain features of the second feature segment in each of the N feature groups according to a preset weight to obtain the fusion features of each of the N feature groups.
[0176] In a possible implementation, the first classification processing module 260 may specifically be used to:
[0177] Input the fusion features of each of the N feature groups into the trained classification model respectively to obtain N probability matrices;
[0178] Calculate the mean of the N probability matrices to obtain the final probability matrix;
[0179] Determine the class label to which the element with the largest value in the final probability matrix belongs as the first scene classification result of the first audio segment.
[0180] In a possible implementation, the apparatus 200 further includes:
[0181] A second acquisition module, configured to continue acquiring audio data of a second preset duration in the target scenario to obtain a second audio segment;
[0182] A second classification processing module, configured to determine a second scenario classification result of the second audio segment according to the fusion feature of the second audio segment;
[0183] A third classification processing module, configured to determine a final scenario classification result of the target scenario according to the first scenario classification result of the first audio segment and the second scenario classification result of the second audio segment.
[0184] It should be noted that, for the information interaction, execution process, etc. between the above-mentioned device / units, since they are based on the same concept as the method embodiments of the present application, for their specific functions and the technical effects brought, reference can be specifically made to the method embodiment part, and details are not described herein again.
[0185] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above-mentioned functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working processes of the units and modules in the above-mentioned system can refer to the corresponding processes in the foregoing method embodiments, and details are not described herein again.
[0186] An embodiment of the present application further provides a terminal device. Refer to Figure 7 , the terminal device 300 may include: at least one processor 310, a memory 320, and a computer program stored in the memory 320 and executable on the at least one processor 310. When the processor 310 executes the computer program, the steps in any of the above method embodiments are implemented, such as Figure 1 the steps S101 to S106 in the illustrated embodiment. Or, when the processor 310 executes the computer program, the functions of each module / unit in the above device embodiments are implemented, such as Figure 6 the functions of the illustrated modules 210 to 260.
[0187] Exemplarily, a computer program can be divided into one or more modules / units. One or more modules / units are stored in the memory 320 and executed by the processor 310 to complete this application. The one or more modules / units can be a series of computer program segments capable of performing specific functions, and these program segments are used to describe the execution process of the computer program in the terminal device 300.
[0188] Those skilled in the art can understand that Figure 7 merely examples of the terminal device, which do not constitute a limitation on the terminal device, and may include more or fewer components than shown in the figure, or combine certain components, or different components, such as input / output devices, network access devices, buses, etc.
[0189] The processor 310 can be a central processing unit (CPU), or can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.
[0190] The memory 320 can be an internal storage unit of the terminal device or an external storage device of the terminal device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. The memory 420 is used to store the computer program and other programs and data required by the terminal device. The memory 420 can also be used to temporarily store data that has been output or will be output.
[0191] The bus can be an industry standard architecture (ISA) bus, a peripheral component interconnect (PCI) bus, or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience in representation, the bus in the drawings of this application is not limited to only one bus or one type of bus.
[0192] The audio scene classification method provided by the embodiments of this application can be applied to terminal devices such as computers, tablet computers, laptop computers, netbooks, personal digital assistants (PDAs), etc. The embodiments of this application do not impose any restrictions on the specific types of terminal devices.
[0193] In the above embodiments, the descriptions of the respective embodiments have their own focuses. For parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0194] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0195] In the embodiments provided by this application, it should be understood that the disclosed terminal devices, apparatuses, and methods can be implemented in other ways. For example, the terminal device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in an electrical, mechanical, or other form.
[0196] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0197] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0198] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-mentioned embodiment methods of this application, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by one or more processors, the steps of the above-mentioned various method embodiments can be implemented.
[0199] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-mentioned embodiment methods of this application, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by one or more processors, the steps of the above-mentioned various method embodiments can be implemented.
[0200] Similarly, as a computer program product, when the computer program product runs on a terminal device, it enables the terminal device to implement the steps in the above-mentioned various method embodiments when executed.
[0201] Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0202] The above-mentioned embodiments are only used to illustrate the technical solutions of this application, rather than to limit it; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of this application, and should all be included in the protection scope of this application.
Claims
1. An audio scene classification method, characterized in that, Including: Collecting audio data of a first preset duration in a target scenario through a microphone on a terminal device to obtain a first audio segment, where the target scenario is the environmental scenario where the terminal device outputs audio to the user's ear through a headphone device; Extracting first feature information and second feature information of the first audio segment, where the first feature information represents frequency feature information based on time, and the second feature information represents time feature information based on frequency; Dividing the first feature information into N first feature segments and dividing the second feature information into N second feature segments respectively according to a preset rule, where N is a positive integer greater than 1; Dividing the N first feature segments and the N second feature segments into N feature groups according to the preset rule, where each feature group includes a first feature segment and a second feature segment; the preset rule is a preset frequency interval, and according to the preset frequency interval, the first feature segment and the second feature segment belonging to the same frequency range are divided into a feature group; Calculating the respective fusion features of the N feature groups; Determining a first scene classification result of the first audio segment according to the respective fusion features of the N feature groups; Among them, the collecting audio data of a first preset duration in a target scenario to obtain a first audio segment includes: Collecting the original audio of the first preset duration in the target scenario; Judging whether there is a target sound in the original audio when the volume of the original audio is greater than a preset threshold; If there is a target sound in the original audio, filtering the target sound in the original audio to obtain the first audio segment; The calculating the respective fusion features of the N feature groups includes: Inputting the first feature segment in each feature group of the N feature groups into a spectrum feature information extraction model to obtain the frequency domain features of the first feature segment in each feature group, and the spectrum feature information extraction model is trained by a depthwise separable convolutional neural network; Inputting the second feature segment in each feature group of the N feature groups into a time series feature information extraction model to obtain the time domain features of the second feature segment in each feature group, and the time series feature information extraction model is trained by a recurrent neural network; Performing fusion processing on the frequency domain features of the first feature segments and the time domain features of the second feature segments of the N feature groups respectively to obtain the respective fusion features of the N feature groups.
2. The audio scene classification method according to claim 1, characterized in that, If there is no target sound in the original audio, determining the original audio as the first audio segment.
3. The audio scene classification method according to claim 1, wherein The performing fusion processing on the frequency domain features of the first feature segments and the time domain features of the second feature segments of the N feature groups respectively to obtain the respective fusion features of the N feature groups includes: Performing vector splicing on the frequency domain features of the first feature segments and the time domain features of the second feature segments in each feature group of the N feature groups according to a preset weight to obtain the respective fusion features of the N feature groups.
4. The audio scene classification method according to claim 1, wherein The determining a first scene classification result of the first audio segment according to the respective fusion features of the N feature groups includes: Input the fusion features of each of the N feature groups into the trained classification model respectively to obtain N probability matrices; Calculate the mean of the N probability matrices to obtain the final probability matrix; Determine the class label to which the element with the largest value in the final probability matrix belongs as the first scene classification result of the first audio segment.
5. The audio scene classification method according to claim 1, characterized in that After determining the first scene classification result of the first audio segment according to the fusion features of each of the N feature groups, the method further includes: Continuously collect audio data of a second preset duration in the target scene to obtain a second audio segment; Determine the second scene classification result of the second audio segment according to the fusion feature of the second audio segment; Determine the final scene classification result of the target scene according to the first scene classification result of the first audio segment and the second scene classification result of the second audio segment.
6. An audio scene classification device, characterized in that, Include: A first acquisition module, configured to collect audio data of a first preset duration in a target scene through a microphone on a terminal device to obtain a first audio segment, where the target scene is the environmental scene where the user is located when the terminal device outputs audio to the user's ear through a headphone device; A first processing module, configured to extract first feature information and second feature information of the first audio segment, where the first feature information represents frequency feature information based on time, and the second feature information represents time feature information based on frequency; A second processing module, configured to divide the first feature information into N first feature segments and divide the second feature information into N second feature segments respectively according to a preset rule, where N is a positive integer greater than 1; A third processing module, configured to divide the N first feature segments and the N second feature segments into N feature groups according to the preset rule, where each feature group includes a first feature segment and a second feature segment; the preset rule is a preset frequency interval, and according to the preset frequency interval, the first feature segments and the second feature segments belonging to the same frequency range are divided into one feature group; A calculation module, configured to calculate the fusion feature of each of the N feature groups; A first classification processing module, configured to determine the first scene classification result of the first audio segment according to the fusion features of each of the N feature groups; Wherein, the acquisition module is specifically configured to collect the original audio of the first preset duration in the target scene; in the case that the volume of the original audio is greater than a preset threshold, determine whether there is a target sound in the original audio; if there is a target sound in the original audio, filter out the target sound in the original audio to obtain the first audio segment; The calculation module is specifically configured to: Input the first feature segment in each feature group of the N feature groups into a spectral feature information extraction model to obtain the frequency domain feature of the first feature segment in each feature group, and the spectral feature information extraction model is trained by a depthwise separable convolutional neural network; Input the second feature segment in each of the N feature groups into the temporal feature information extraction model to obtain the temporal domain features of the second feature segment in each feature group, where the temporal feature information extraction model is trained by a recurrent neural network; Perform fusion processing on the frequency domain features of the first feature segment and the temporal domain features of the second feature segment in each of the N feature groups to obtain the fusion features of each of the N feature groups.
7. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the audio scene classification method according to any one of claims 1 to 5.
8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the audio scene classification method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Music genre identification method and device, equipment and storage medium
CN113450828A