Audio recognition model training method and audio recognition method
By performing frequency domain feature extraction and target weight set calculation on historical audio data, and training the audio recognition model with preset parameter sets, the problem of high construction cost and poor generalization of data sets in music scene detection by deep learning methods is solved, and the accuracy and equipment applicability of audio recognition are improved.
Patent Information
- Application Number
- CN202211102193.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-09
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-09-09
AI Technical Summary
In the detection of music scenes, existing deep learning methods have problems such as high labor cost for data set construction, unreasonable extraction of feature types of multiple sound types, poor generalization, and equipment differences affecting audio recognition results.
By acquiring multiple historical audio data, frequency domain feature extraction and target weight set calculation, combined with preset parameter sets for fusion processing, the target audio recognition model is trained, the recognition accuracy of different types of audio features is improved, and the impact of equipment differences is reduced.
Accurate recognition of different types of audio features is achieved, reducing the impact of device differences on audio recognition results, and improving the accuracy and user satisfaction of audio recognition.
Smart Images

Figure CN115497461B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of audio processing, and particularly relates to an audio recognition model training method and an audio recognition method. Background Art
[0002] The purpose of music scene detection is to detect the audio stream containing music segments, which is a type of sound event detection. Correctly detecting music segments can optimize the bandwidth and audio quality of real-time audio coding and transmission, and improve the user experience. The current mainstream music scene detection is divided into two categories: one is based on traditional machine learning algorithms, that is, adding a classifier to various audio features extracted from audio signals for detection and judgment.
[0003] Currently, audio is mainly recognized based on deep learning methods. However, there are still many problems with current deep learning methods, including: 1. The labor cost of dataset construction is too high, and the specific labels of each frame need to be accurately marked; 2. There are multiple types of sounds in music at the same time. However, since the parameters of the current deep learning method model are fixed during inference, this means that the same set of parameters is used for calculation and feature extraction for different types of sounds that exist simultaneously, which is obviously unreasonable because the characteristics of different types of sounds are different; 3. The generalization of current deep learning methods is poor. For the same audio, even after being picked up by different devices in the same environment, different results may be obtained due to differences in device background noise, frequency response curves, etc. Summary of the Invention
[0004] This application aims to at least solve one of the technical problems in the related art to some extent. For this reason, an object of this application is to propose an audio recognition model training method and an audio recognition method.
[0005] To solve the above technical problems, the embodiments of this application provide the following technical solutions:
[0006] An audio recognition model training method includes:
[0007] Obtain a plurality of historical audio data, and input the plurality of historical audio data into the audio recognition model to be trained; extract frequency domain features from each of the historical audio data to obtain a plurality of initial audio features;
[0008] Based on each of the initial audio features, calculate a target weight set that matches the initial audio feature;
[0009] Obtain a preset parameter set, and calculate based on the preset parameter set and each of the target weight sets respectively to obtain a plurality of fusion parameters; based on each of the fusion parameters, process the initial audio feature that matches each of the fusion parameters to obtain a plurality of target audio features;
[0010] Train the audio recognition model to be trained based on multiple target audio features to obtain a target audio recognition model.
[0011] Optionally, obtaining the historical audio data includes:
[0012] Obtain a training audio file, and crop the training audio file to obtain multiple audio segments;
[0013] Annotate each audio segment to obtain multiple historical audio data; wherein, each historical audio data includes an annotation; the annotation includes the audio type of each audio segment.
[0014] Optionally, calculating a set of target weights matching the initial audio feature based on each initial audio feature includes:
[0015] Perform pooling processing on each initial audio feature to obtain multiple dimensionality-reduced audio features;
[0016] Calculate a set of target weights matching the dimensionality-reduced audio feature based on each dimensionality-reduced audio feature.
[0017] Optionally, calculating a set of target weights matching the dimensionality-reduced audio feature based on each dimensionality-reduced audio feature includes:
[0018] Perform convolution processing on each dimensionality-reduced audio feature based on a preset number of first convolutional kernels, and output a target dimension;
[0019] Calculate a set of target weights matching the dimensionality-reduced audio feature based on the target dimension.
[0020] Optionally, obtaining a preset parameter set, and calculating based on the preset parameter set and each set of target weights respectively to obtain multiple fusion parameters includes:
[0021] Obtain the target weights matching each preset parameter; wherein, each preset parameter set includes M preset parameters; each set of target weights includes M target weights; M is a positive integer;
[0022] Calculate each preset parameter and the matching target weight to obtain multiple calculation results;
[0023] Fuse multiple calculation results to obtain a fusion parameter.
[0024] Optionally, processing the initial audio features matching each of the fusion parameters based on each of the fusion parameters to obtain a plurality of target audio features, including:
[0025] A preset processing rule;
[0026] Processing the initial audio features matching each of the fusion parameters based on the processing rule and each of the fusion parameters to obtain a plurality of the target audio features.
[0027] An embodiment of the present application further provides an audio recognition method, including the target audio recognition model as described above, and further including:
[0028] Obtaining a plurality of audio data to be processed;
[0029] Based on the plurality of audio data to be processed, obtaining a plurality of real-time audio features;
[0030] Based on the plurality of real-time audio features, obtaining a plurality of audio features to be input;
[0031] Inputting the plurality of audio features to be input into the target audio recognition model to obtain target audio features.
[0032] Optionally, the obtaining a plurality of the real-time audio features based on the plurality of audio data to be processed includes:
[0033] Resampling each of the real-time audio data to obtain a plurality of sampled audio features;
[0034] Framing each of the sampled audio features to obtain a plurality of frames of initial audio;
[0035] Performing transformation processing on each frame of the initial audio to obtain a frequency domain amplitude spectrum;
[0036] Performing feature extraction on each of the frequency domain amplitude spectra to obtain a plurality of the real-time audio features.
[0037] Optionally, the obtaining a plurality of audio features to be input based on the plurality of real-time audio features includes:
[0038] Performing conversion processing on each of the real-time audio features to obtain a plurality of audio feature information;
[0039] Dividing each of the audio feature information based on the frequency domain to obtain a plurality of initial frequency bands;
[0040] Based on the plurality of initial frequency bands, determining a plurality of frequency bands to be processed and reserved frequency bands;
[0041] Processing each of the frequency bands to be processed based on a preset value to obtain a plurality of feature enhanced frequency bands;
[0042] Based on multiple of the enhanced frequency bands and the reserved frequency bands of the features, the to-be-input audio features are obtained.
[0043] Optionally, processing each of the to-be-processed frequency bands based on a preset value to obtain multiple enhanced frequency bands of features includes:
[0044] Determining a preset weight for each of the to-be-processed frequency bands; wherein, the preset value includes multiple of the preset weights;
[0045] Fusing each of the to-be-processed frequency bands with the matching preset weight to obtain multiple of the enhanced frequency bands of features.
[0046] An embodiment of the present application further provides an audio recognition model training device, including:
[0047] An input unit, configured to obtain multiple historical audio data and input the multiple historical audio data into an audio recognition model to be trained;
[0048] A feature extraction unit, configured to perform frequency-domain feature extraction on each of the historical audio data to obtain multiple initial audio features;
[0049] A calculation unit, configured to calculate and obtain a set of target weights matching the initial audio feature based on each of the initial audio features;
[0050] A processing unit, configured to obtain a preset parameter set and perform calculations based on the preset parameter set and each of the sets of target weights respectively to obtain multiple fusion parameters; based on each of the fusion parameters, process the initial audio features matching each of the fusion parameters to obtain multiple target audio features;
[0051] A training unit, configured to train the audio recognition model to be trained based on multiple of the target audio features to obtain a target audio recognition model.
[0052] An embodiment of the present application further provides an audio recognition device, including:
[0053] An acquisition unit, configured to acquire multiple to-be-processed audio data;
[0054] A first obtaining unit, configured to obtain multiple real-time audio features based on multiple of the to-be-processed audio data;
[0055] A second obtaining unit, configured to obtain multiple to-be-input audio features based on multiple of the real-time audio features;
[0056] A third obtaining unit, configured to input multiple of the to-be-input audio features into a target audio recognition model to obtain target audio features.
[0057] An embodiment of the present application further provides an audio recognition component, including:
[0058] A weight calculation module, which obtains a plurality of historical audio data and inputs the plurality of historical audio data into a to-be-trained audio recognition model; performs frequency domain feature extraction on each of the historical audio data to obtain a plurality of initial audio features; and calculates a set of target weights that matches the initial audio feature based on each of the initial audio features.
[0059] A perception module, which is connected to the weight calculation module. The perception module is configured to obtain a preset parameter set, calculate a plurality of fusion parameters based on the preset parameter set and each of the sets of target weights respectively, and process the initial audio features that match each of the fusion parameters based on each of the fusion parameters to obtain a plurality of target audio features.
[0060] Train the to-be-trained audio recognition model based on the plurality of target audio features to obtain a target audio recognition model.
[0061] An embodiment of the present application further provides an electronic device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the above-mentioned method is implemented.
[0062] An embodiment of the present application further provides a computer-readable storage medium. The computer-readable storage medium includes a stored computer program. When the computer program runs, it controls the device where the computer-readable storage medium is located to execute the above-mentioned method.
[0063] The embodiment of the present application has the following technical effects:
[0064] In the above technical solution of the present application, 1) for each initial audio feature, a set of target weights that matches the initial audio feature is calculated, and based on the preset parameters and the set of target weights, the initial audio feature is processed to obtain a target audio feature, realizing that in the process of recognizing the initial audio feature, different types of initial audio features are recognized based on different types of sets of target weights, improving the accuracy of recognizing the initial audio feature and reducing the influence of device differences on the audio recognition result.
[0065] 2) The historical audio data used for training is easy to obtain and process, and the audio recognition component can be built into various devices, such as smartphones and computers, etc., to improve the recognition ability of various devices for various types of audio in the received audio data, reduce the differences in the recognition results of the same audio by different devices, and improve user satisfaction.
[0066] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present application. Description of the Drawings
[0067] Figure 1 is a schematic structural diagram of an audio recognition element provided by an embodiment of the present application;
[0068] Figure 2 is a schematic flowchart of a method for training an audio recognition model provided by an embodiment of the present application;
[0069] Figure 3 is a schematic flowchart of an audio recognition method provided by an embodiment of the present application;
[0070] Figure 4 is a schematic structural diagram of an apparatus for training an audio recognition model provided by an embodiment of the present application;
[0071] Figure 5 is a schematic structural diagram of an audio recognition apparatus provided by an embodiment of the present application. Detailed Description of the Embodiments
[0072] The embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the drawings are exemplary and are intended to explain the present application and should not be construed as limiting the present application.
[0073] For the convenience of those skilled in the art to understand the embodiments, some terms are explained:
[0074] (1) fc: Full Connection, fully connected.
[0075] (2) cnn: Convolutional Neural Network, convolutional neural network.
[0076] (3) rnn: Recurrent Neural Network, recurrent neural network.
[0077] (4) lstm: Long Short-Term Memory, long short-term memory network.
[0078] (5) gru: Gate Recurrent Unit, gated recurrent neural network.
[0079] (6) transformer: transformer.
[0080] (7) mel: the Mel Scale, the Mel scale.
[0081] (8) mfcc: Mel-scale Frequency Cepstral Coefficients, Mel-frequency cepstral coefficients.
[0082] (9) FFT: Fast Fourier transform, fast Fourier transform.
[0083] (10) ReLU: Linear rectification function, linear rectifier function.
[0084] (11) sigmoid: Sigmoid function, often used as an activation function for neural networks, mapping variables between 0 and 1.
[0085] (12) tanh: hyperbolic function.
[0086] As Figure 1 shown, an embodiment of the present application provides an audio recognition component 10, including:
[0087] A weight calculation module 12, which obtains a plurality of historical audio data and inputs the plurality of historical audio data into an audio recognition model to be trained; performs frequency-domain feature extraction on each of the historical audio data to obtain a plurality of initial audio features; and calculates, based on each of the initial audio features, a set of target weights that matches the initial audio feature.
[0088] A perception module 13, the perception module 13 is connected to the weight calculation module 12, and the perception module 13 is configured to obtain a preset parameter set and calculate, based on the preset parameter set, with each of the sets of target weights respectively, to obtain a plurality of fusion parameters; and process, based on each of the fusion parameters, the initial audio features that match each of the fusion parameters to obtain a plurality of target audio features.
[0089] An optional embodiment of the present application further includes a feature extraction module 11, configured to obtain a plurality of historical audio data and perform frequency-domain feature extraction on each of the historical audio data to obtain a plurality of initial audio features.
[0090] In an optional embodiment of the present application, the weight calculation module 12 includes a pooling layer, a feature layer, and a softmax layer; wherein, the pooling layer is configured to perform compression processing on the initial audio features to obtain dimension-reduced audio features.
[0091] The feature layer is configured to calculate a plurality of target weights based on the dimension-reduced audio features.
[0092] Among them, the feature layer can be implemented based on current neural network building blocks, such as: fc layer, cnn layer, rnn layer, lstm layer, gru layer, transformer layer, etc.;
[0093] The softmax layer is used to normalize multiple target weights so that the sum of the multiple target weights is 1, facilitating the subsequent algorithm calls and calculations.
[0094] The perception module 13 can be a two-dimensional convolutional layer, including a preset number of second convolutional kernels; and can be implemented based on current neural network building blocks, such as: fc layer, cnn layer, rnn layer, lstm layer, gru layer, transformer layer, etc.
[0095] Furthermore, the feature layer is provided with a preset number of first convolutional kernels, where the specific value of the preset number is determined based on the number of second convolutional kernels included in the perception module 13, and is used to obtain the target weight of each second convolutional kernel based on the feature layer.
[0096] Specifically, when the feature layer is a one-dimensional convolutional layer, the feature layer includes a preset number of first convolutional kernels; if the feature layer is an n-dimensional convolutional layer, the nth-dimensional convolutional layer of the feature layer includes a preset number of first convolutional kernels, where n is a positive integer.
[0097] In the embodiments of the present application, the historical audio data for training is easy to obtain and process, and the audio recognition component 10 can be built into various devices, such as smartphones and computers, etc., to improve the recognition ability of various devices for various types of audio in the received audio data, reduce the differences in the recognition results of the same audio by different devices, and improve user satisfaction.
[0098] As Figure 2 shown, the embodiments of the present application also provide an audio recognition model training method, which is applied to the audio recognition component 10 as Figure 1 shown, and includes:
[0099] Step S21: Obtain a plurality of historical audio data, and input the plurality of historical audio data into the audio recognition model to be trained;
[0100] In an optional embodiment of the present application, obtaining the historical audio data includes:
[0101] Obtain a training audio file, and crop the training audio file to obtain a plurality of audio segments;
[0102] Annotate each of the audio segments to obtain a plurality of historical audio data; where each historical audio data includes an annotation; the annotation includes the audio type of each audio segment.
[0103] In the embodiments of the present application, different operations can be performed according to different voice types in the future. Therefore, it is not necessary to label the audio data at each moment in each audio segment. One annotation for one audio segment is sufficient, and it is not necessary to classify each frame of audio. Compared with classifying each frame of audio, the annotation method of the present application is simple, time-saving and labor-saving.
[0104] Step S22: Extract frequency domain features from each of the historical audio data to obtain a plurality of initial audio features;
[0105] Step S23: Calculate a target weight set matching the initial audio feature based on each of the initial audio features;
[0106] In an alternative embodiment of the present application, calculating a target weight set matching the initial audio feature based on each of the initial audio features includes:
[0107] Perform pooling processing on each of the initial audio features to obtain a plurality of dimension-reduced audio features;
[0108] Calculate a target weight set matching the dimension-reduced audio feature based on each of the dimension-reduced audio features.
[0109] Further, calculating a target weight set matching the dimension-reduced audio feature based on each of the dimension-reduced audio features includes:
[0110] Perform convolution processing on each of the dimension-reduced audio features based on a preset number of first convolutional kernels, and output a target dimension;
[0111] Calculate a target weight set matching the dimension-reduced audio feature based on the target dimension.
[0112] In the embodiments of the present application, pooling processing is performed on each initial audio feature to compress the dimension of the initial audio feature for obtaining dimension-reduced audio features with appropriate dimensions; among them, in order to perform subsequent one-dimensional convolution operations normally, the maximum value on the last dimension of the to-be-input audio feature can be taken through a max pooling layer to obtain the dimension-reduced audio feature; for example: the dimension of the dimension-reduced audio feature is 120*1;
[0113] For example: the dimension of the obtained mel spectrum is 120*120;
[0114] Then, perform one-dimensional convolution operation on the dimension-reduced audio feature based on a preset number of first convolutional kernels of the one-dimensional convolutional layer;
[0115] Among them, the specific value of the preset quantity is determined based on the number of preset parameters in the preset parameter set, and is used to ensure obtaining the target weight of each preset parameter; the size dimension of the first convolutional kernel is 3*1;
[0116] Assume that the preset parameters included in the preset value are 16, that is, the perception module 13 includes 16 second convolutional kernels, and the size of each second convolutional kernel is 3*3;
[0117] Therefore, in order to obtain the target weight of each second convolutional kernel, the one-dimensional convolutional layer also needs to include 16 first convolutional kernels;
[0118] After performing convolutional processing on the dimension-reduced audio features based on 16 first convolutional kernels respectively, a target weight set (dimension 16*1) can be output.
[0119] Then, the dimension 16*1 is normalized to between 0 and 1 through softmax, and the sum of 16 values is 1.
[0120] Furthermore, the one-dimensional convolutional layer can be replaced with L one-dimensional convolutional layers, and only the last one-dimensional convolutional layer needs to include 16 first convolutional kernels; where L is a positive integer.
[0121] Step S24: Obtain a preset parameter set, and calculate based on the preset parameter set and each of the target weight sets respectively to obtain a plurality of fusion parameters; based on each of the fusion parameters, process the initial audio features matching each of the fusion parameters to obtain a plurality of target audio features;
[0122] In an optional embodiment of the present application, the obtaining a preset parameter set, and calculating based on the preset parameter set and each of the target weight sets respectively to obtain a plurality of fusion parameters includes:
[0123] Obtain the target weight matching each of the preset parameters; where each of the preset parameter sets includes M preset parameters; each of the target weight sets includes M target weights; M is a positive integer;
[0124] Calculate each of the preset parameters and the matching target weights to obtain a plurality of calculation results;
[0125] Fuse a plurality of the calculation results to obtain one of the fusion parameters.
[0126] Furthermore, the processing the initial audio features matching each of the fusion parameters based on each of the fusion parameters to obtain a plurality of target audio features includes:
[0127] Preset processing rules;
[0128] Process the initial audio features matching each of the fusion parameters based on the processing rules and each of the fusion parameters to obtain a plurality of the target audio features.
[0129] In an embodiment of the present application, each preset parameter P a is multiplied by a corresponding target weight Q a ; where P a is the a-th preset parameter, Q a is the a-th target weight, and 1 ≤ a ≤ 16; a is an integer;
[0130] Then the a-th calculation result S a can be calculated based on the following formula:
[0131] S a = P a * Q a ;
[0132] The fusion parameter T b corresponding to the b-th initial audio feature can be calculated based on the following formula, where b is a positive integer:
[0133] T b = S1 + S2... S 16 ;
[0134] Among them, the processing rule can be σ, and σ represents an activation function, specifically, it can be a ReLU function, a sigmoid function, a tanh function, etc.
[0135] Further, the target audio feature y can be calculated based on the following formula:
[0136]
[0137] In the formula, x b is the b-th initial audio feature; is an operation rule of the perception module 13, such as a convolution operation, etc.
[0138] In an embodiment of the present application, for each initial audio feature, a set of target weights matching the initial audio feature is calculated, and based on the preset parameters and the set of target weights, the initial audio feature is processed to obtain the target audio feature, realizing that in the process of recognizing the initial audio feature, different types of initial audio features are recognized based on different types of sets of target weights, improving the accuracy of recognizing the initial audio feature and reducing the influence of device differences on the audio recognition result.
[0139] Step S25: Train the audio recognition model to be trained based on the plurality of target audio features to obtain a target audio recognition model.
[0140] Specifically, repeat the above steps, and use the obtained multiple target audio features to perform multiple iterative trainings on the audio recognition model to be trained.
[0141] Among them, the loss function involved in the training process is the cross-entropy loss function. When training for 100 generations or the change value of the loss function is less than 10 -6 , the training stops, and the target audio recognition model is obtained.
[0142] As Figure 3 shown, an embodiment of the present application further provides an audio recognition method, including the target audio recognition model as described above, and further including:
[0143] Step S31: Obtain multiple audio data to be processed;
[0144] Step S32: Based on the multiple audio data to be processed, obtain multiple real-time audio features;
[0145] In an optional embodiment of the present application, the obtaining multiple real-time audio features based on the multiple audio data to be processed includes:
[0146] Resample each piece of real-time audio data to obtain multiple sampled audio features;
[0147] Perform frame splitting on each sampled audio feature to obtain multiple frames of initial audio;
[0148] Perform transformation processing on each frame of the initial audio to obtain a frequency-domain amplitude spectrum;
[0149] Extract features from each frequency-domain amplitude spectrum to obtain multiple real-time audio features.
[0150] Specifically, in the embodiment of the present application, audio features include spectrum, mel spectrum, mfcc, etc.
[0151] Taking the mel spectrum as an example, the embodiment of the present application is explained as follows. Specifically:
[0152] 1) Uniformly resample historical audio data with different sampling rates to 16000 Hz to obtain multiple sampled audio features;
[0153] 2) Use a Hanning window with a frame length of 512 points to perform frame splitting on each sampled audio feature to reduce spectral leakage. The overlap between frames is 50%. After performing Fourier transform on each frame of audio data of each sampled audio feature, a corresponding frequency-domain amplitude spectrum is obtained. Therefore, the spectrum dimension obtained based on a sampled audio feature with a duration of 2 s is approximately 120×257; among them, 120 in the spectrum dimension represents the time step, and 257 represents the number of frequency points.
[0154] 3) For each frequency-domain amplitude spectrum, perform feature extraction through mel-frequency transformation and mel filters in sequence to obtain a mel spectrum;
[0155] Among them, if a 120-dimensional mel filter bank is selected, a mel spectrum with a dimension of 120×120 can be obtained;
[0156] Specifically, the mel-frequency transformation of the frequency-domain amplitude spectrum can be implemented based on the following formula:
[0157] f mel = 2595·lg(1 + f / 700Hz)
[0158] In the formula, f mel represents the mel frequency; f represents the initial frequency of the frequency-domain amplitude spectrum;
[0159] For the feature extraction of f mel based on the mel filter, it can be implemented based on the following formula:
[0160]
[0161] In the formula, H mel (m) represents the m-th mel filter, and f mel (m) is the center frequency of the m-th mel filter; k represents the frequency point serial number. For example, for the 257-dimensional spectrum obtained above, k takes values from 1 to 257.
[0162] Furthermore, for the convenience of processing, take the logarithm of the obtained mel spectrum to get the log mel spectrum.
[0163] 4) Normalize the log mel spectrum obtained in step 3) to the interval [0, 1] according to the maximum and minimum values to obtain the initial audio features.
[0164] For the specific method steps of normalization, it can be implemented based on existing algorithms, and the embodiments of the present application do not make specific limitations.
[0165] In the embodiments of the present application, by processing each historical audio data based on components such as filters for analog acoustic signals, it is possible to adjust the target audio recognition parameters in various devices, reduce the influence of the differences in the device's own parameters or attributes on audio recognition, and thereby improve the applicability and accuracy of the target audio model.
[0166] Step S33: Based on the multiple real-time audio features, obtain multiple audio features to be input;
[0167] In an optional embodiment of the present application, the obtaining of multiple audio features to be input based on the multiple real-time audio features includes:
[0168] Perform transformation processing on each of the real-time audio features to obtain multiple pieces of audio feature information;
[0169] Divide each of the audio feature information based on the frequency domain to obtain multiple initial frequency bands;
[0170] Based on the multiple initial frequency bands, determine multiple frequency bands to be processed and reserved frequency bands;
[0171] Process each of the frequency bands to be processed based on a preset value to obtain multiple feature-enhanced frequency bands;
[0172] Based on the multiple feature-enhanced frequency bands and the reserved frequency bands, obtain the audio feature to be input.
[0173] An embodiment of the present application, wherein one real-time audio feature is time-domain audio data X(t), where t represents time;
[0174] Perform transformation on the time-domain audio data X(t) based on FFT to obtain audio feature information (frequency-domain audio data) X(t, f) corresponding to the real-time audio feature; where f is the frequency corresponding to the real-time audio at time t of the real-time audio feature.
[0175] Divide the audio feature information based on the frequency domain. For example, when the minimum frequency corresponding to the audio feature information is 0 Hz and the maximum frequency corresponding to the audio feature information is 200 Hz, then divide 0 - 200 Hz into multiple initial frequency bands, where the frequency interval between every two initial frequency bands can be the same or different; for example, 1) the audio feature information can be divided based on a frequency interval of 20 Hz in the frequency domain to obtain 11 frequency boundaries and 10 initial frequency bands with the same width;
[0176] 2) It is also possible to divide the real-time audio feature based on multiple frequency intervals such as 10 Hz, 15 Hz, 20 Hz, etc. at the same time to obtain 10 initial frequency bands with different widths;
[0177] Among them, in order to prevent the width of the obtained initial frequency band from being too small, a frequency interval threshold can be preset, for example: 5 Hz, that is, when dividing the real-time audio feature based on the frequency domain, the frequency interval is greater than 5 Hz.
[0178] Further, the process of processing each of the frequency bands to be processed based on a preset value to obtain multiple feature-enhanced frequency bands includes:
[0179] Determine the preset weight of each of the frequency bands to be processed; where the preset value includes multiple of the preset weights;
[0180] Fuse each of the to-be-processed frequency bands with the matching preset weights to obtain multiple feature-enhanced frequency bands.
[0181] In an optional embodiment of the present application, it is assumed that the following 10 initial frequency bands are obtained:
[0182] The first initial frequency band: 0 Hz to 15 Hz;
[0183] The second initial frequency band: 15 Hz to 30 Hz;
[0184] The third initial frequency band: 30 Hz to 50 Hz;
[0185] The fourth initial frequency band: 50 Hz to 75 Hz;
[0186] The fifth initial frequency band: 75 Hz to 100 Hz;
[0187] The sixth initial frequency band: 100 Hz to 120 Hz;
[0188] The seventh initial frequency band: 120 Hz to 140 Hz;
[0189] The eighth initial frequency band: 140 Hz to 160 Hz;
[0190] The ninth initial frequency band: 160 Hz to 180 Hz;
[0191] The tenth initial frequency band: 180 Hz to 200 Hz;
[0192] From these 10 initial frequency bands, randomly select 5 to-be-processed frequency bands, and the remaining 5 initial frequency bands are used as reserved frequency bands, and no processing is performed on the reserved frequency bands.
[0193] The preset values may include the following 5 preset weights W(i), 1 ≤ i ≤ 5, and i is an integer, and W(i) is the i-th preset weight:
[0194] 0.2, 0.6, 0.8, 2, and 0.9;
[0195] Among them, the specific value of each preset weight can be randomly preset, and the embodiments of the present application do not make specific limitations on this.
[0196] For the 5 to-be-processed frequency bands X(t, f j ), 1 ≤ j ≤ 5, and j is an integer, f j is the j-th to-be-processed frequency band;
[0197] Randomly select a preset weight;
[0198] Then, the j-th feature-enhanced frequency band X 增强 (t, f j) = X(t, f j ) * W(i).
[0199] Based on 5 enhanced frequency bands and 5 reserved frequency bands (initial frequency bands), an audio feature to be input is obtained.
[0200] Step S34: Input multiple of the audio features to be input into the target audio recognition model to obtain target audio features.
[0201] As Figure 4 shown, an embodiment of the present application further provides an audio recognition model training apparatus 40, including:
[0202] An input unit 41, configured to obtain multiple historical audio data and input the multiple historical audio data into an audio recognition model to be trained;
[0203] A feature extraction unit 42, configured to perform frequency domain feature extraction on each of the historical audio data to obtain multiple initial audio features;
[0204] A calculation unit 43, configured to calculate, based on each of the initial audio features, a set of target weights that matches the initial audio feature;
[0205] A processing unit 44, configured to obtain a preset parameter set and perform calculations based on the preset parameter set and each of the sets of target weights to obtain multiple fusion parameters; based on each of the fusion parameters, process the initial audio features that match each of the fusion parameters to obtain multiple target audio features;
[0206] A training unit 45, configured to train the audio recognition model to be trained based on the multiple target audio features to obtain a target audio recognition model.
[0207] Optionally, obtaining the historical audio data includes:
[0208] Obtain a training audio file and crop the training audio file to obtain multiple audio segments;
[0209] Annotate each of the audio segments to obtain multiple historical audio data; wherein, each of the historical audio data includes an annotation; the annotation includes the audio type of each of the audio segments.
[0210] Optionally, the calculating, based on each of the initial audio features, a set of target weights that matches the initial audio feature includes:
[0211] Perform pooling processing on each of the initial audio features to obtain multiple dimensionality-reduced audio features;
[0212] Based on each of the dimensionality-reduced audio features, a target weight set matching the dimensionality-reduced audio feature is calculated and obtained.
[0213] Optionally, the calculating and obtaining a target weight set matching the dimensionality-reduced audio feature based on each of the dimensionality-reduced audio features includes:
[0214] Performing convolution processing on each of the dimensionality-reduced audio features based on a preset number of first convolutional kernels, and outputting a target dimension;
[0215] Based on the target dimension, a target weight set matching the dimensionality-reduced audio feature is calculated and obtained.
[0216] Optionally, the obtaining a preset parameter set and calculating based on the preset parameter set with each of the target weight sets respectively to obtain a plurality of fusion parameters includes:
[0217] Obtaining the target weights matching each of the preset parameters; wherein each preset parameter set includes M preset parameters; each target weight set includes M target weights; M is a positive integer;
[0218] Calculating each of the preset parameters with the matching target weights to obtain a plurality of calculation results;
[0219] Fusing the plurality of calculation results to obtain one fusion parameter.
[0220] Optionally, the processing the initial audio features matching each of the fusion parameters based on each of the fusion parameters to obtain a plurality of target audio features includes:
[0221] Presetting a processing rule;
[0222] Processing the initial audio features matching each of the fusion parameters based on the processing rule and each of the fusion parameters to obtain a plurality of the target audio features.
[0223] As Figure 5 shown, an embodiment of the present application further provides an audio recognition device 50, including:
[0224] An acquisition unit 51, configured to acquire a plurality of audio data to be processed;
[0225] A first obtaining unit 52, configured to obtain a plurality of real-time audio features based on the plurality of audio data to be processed;
[0226] A second obtaining unit 53, configured to obtain a plurality of audio features to be input based on the plurality of real-time audio features;
[0227] A third obtaining unit 54, configured to input the multiple to-be-input audio features into a target audio recognition model to obtain target audio features.
[0228] Optionally, the obtaining multiple real-time audio features based on the multiple to-be-processed audio data includes:
[0229] Resampling each of the real-time audio data to obtain multiple sampled audio features;
[0230] Performing frame division on each of the sampled audio features to obtain multiple frames of initial audio;
[0231] Performing transformation processing on each frame of the initial audio to obtain a frequency-domain amplitude spectrum;
[0232] Performing feature extraction on each of the frequency-domain amplitude spectra to obtain the multiple real-time audio features.
[0233] Optionally, the obtaining multiple to-be-input audio features based on the multiple real-time audio features includes:
[0234] Performing conversion processing on each of the real-time audio features to obtain multiple audio feature information;
[0235] Dividing each of the audio feature information based on the frequency domain to obtain multiple initial frequency bands;
[0236] Determining multiple to-be-processed frequency bands and reserved frequency bands based on the multiple initial frequency bands;
[0237] Processing each of the to-be-processed frequency bands based on a preset value to obtain multiple feature-enhanced frequency bands;
[0238] Obtaining the to-be-input audio features based on the multiple feature-enhanced frequency bands and the reserved frequency bands.
[0239] Optionally, the processing each of the to-be-processed frequency bands based on a preset value to obtain multiple feature-enhanced frequency bands includes:
[0240] Determining a preset weight for each of the to-be-processed frequency bands; wherein the preset value includes the multiple preset weights;
[0241] Fusing each of the to-be-processed frequency bands with the matched preset weight to obtain the multiple feature-enhanced frequency bands.
[0242] An embodiment of the present application further provides an electronic device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the above-mentioned method is implemented.
[0243] An embodiment of the present application further provides a computer-readable storage medium, which includes a stored computer program. When the computer program runs, it controls the device where the computer-readable storage medium is located to execute the method described above.
[0244] In addition, the other configurations and functions of the device in the embodiments of the present application are known to those skilled in the art. To reduce redundancy, they will not be described in detail here.
[0245] It should be noted that the logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a defined sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection part with one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, then editing, interpreting, or processing it in other suitable ways if necessary, and then storing it in a computer memory.
[0246] It should be understood that each part of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application specific integrated circuits with appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0247] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0248] In the description of this application, it should be understood that the orientation or positional relationship indicated by terms such as "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", "axial", "radial", "circumferential", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing this application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to this application.
[0249] In addition, the terms "first" and "second" are only used for descriptive purposes and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of this application, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically and clearly defined.
[0250] In this application, unless otherwise clearly specified and limited, terms such as "install", "connect", "join", "fix", etc. should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or integrated; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two elements or the interaction relationship between two elements, unless otherwise clearly limited. For those of ordinary skill in the art, the specific meanings of the above terms in this application can be understood according to specific circumstances.
[0251] In this application, unless otherwise clearly specified or limited, the first feature being "on" or "under" the second feature may mean that the first and second features are in direct contact, or the first and second features are indirectly in contact through an intermediate medium. Moreover, the first feature being "above", "over" and "on top of" the second feature may mean that the first feature is directly above or obliquely above the second feature, or merely indicates that the first feature has a higher horizontal height than the second feature. The first feature being "under", "beneath" and "underneath" the second feature may mean that the first feature is directly below or obliquely below the second feature, or merely indicates that the first feature has a lower horizontal height than the second feature.
[0252] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present application.
Claims
1. A method for training an audio recognition model, characterized in that, Including: Obtain multiple historical audio data, and input the multiple historical audio data into the audio recognition model to be trained; Extract frequency domain features from each of the historical audio data to obtain multiple initial audio features; Based on each of the initial audio features, calculate and obtain a set of target weights that match the initial audio features; Obtain a preset parameter set, and calculate based on the preset parameter set and each of the target weight sets respectively to obtain multiple fusion parameters; Based on each of the fusion parameters, process the initial audio features that match each of the fusion parameters to obtain multiple target audio features, including: a preset processing rule, and the processing rule is an activation function; based on the processing rule and each of the fusion parameters, process the initial audio features that match each of the fusion parameters to obtain multiple target audio features; Train the audio recognition model to be trained based on the multiple target audio features to obtain a target audio recognition model; The step of obtaining a preset parameter set, and calculating based on the preset parameter set and each of the target weight sets respectively to obtain multiple fusion parameters, includes: Obtain the target weights that match each of the preset parameters, where each preset parameter set includes M preset parameters, each target weight set includes M target weights, and M is a positive integer; Calculate each of the preset parameters and the matching target weights to obtain multiple calculation results, including: multiply each of the preset parameters by the matching target weights to obtain multiple calculation results; Fuse the multiple calculation results to obtain one fusion parameter, including: add the multiple calculation results to obtain one fusion parameter.
2. The method according to claim 1, wherein Obtain the historical audio data, including: Obtain a training audio file, and crop the training audio file to obtain multiple audio segments; Annotate each of the audio segments to obtain multiple historical audio data; wherein, each historical audio data includes an annotation; the annotation includes the audio type of each audio segment.
3. The method according to claim 1, wherein The step of calculating and obtaining a set of target weights that match the initial audio features based on each of the initial audio features, includes: Perform pooling processing on each of the initial audio features to obtain multiple dimensionality-reduced audio features; Based on each of the dimensionality-reduced audio features, calculate and obtain a set of target weights that match the dimensionality-reduced audio features.
4. The method according to claim 3, characterized in that, The step of calculating and obtaining a set of target weights that match the dimensionality-reduced audio features based on each of the dimensionality-reduced audio features, includes: Perform convolution processing on each of the dimensionality-reduced audio features based on a preset number of first convolutional kernels, and output a target dimension; Based on the target dimension, calculate and obtain a set of target weights that match the dimensionality-reduced audio features.
5. An audio recognition method, characterized in that, Including: Obtain multiple audio data to be processed; Based on the multiple audio data to be processed, obtain multiple real-time audio features; Based on the multiple real-time audio features, obtain multiple audio features to be input; Input multiple of the to-be-input audio features into the target audio recognition model according to any one of claims 1-4 to obtain target audio features.
6. The method according to claim 5, wherein The obtaining of multiple real-time audio features based on multiple of the to-be-processed audio data includes: Resample each of the to-be-processed audio data to obtain multiple sampled audio features; Perform frame division on each of the sampled audio features to obtain multiple frames of initial audio; Perform transformation processing on each frame of the initial audio to obtain a frequency-domain amplitude spectrum; Extract features from each of the frequency-domain amplitude spectra to obtain multiple of the real-time audio features.
7. The method according to claim 5, wherein The obtaining of multiple to-be-input audio features based on multiple of the real-time audio features includes: Perform transformation processing on each of the real-time audio features to obtain multiple audio feature information; Divide each of the audio feature information based on the frequency domain to obtain multiple initial frequency bands; Based on multiple of the initial frequency bands, determine multiple to-be-processed frequency bands and reserved frequency bands; Process each of the to-be-processed frequency bands based on a preset value to obtain multiple feature-enhanced frequency bands; Based on multiple of the feature-enhanced frequency bands and the reserved frequency bands, obtain the to-be-input audio features.
8. The method according to claim 7, wherein The processing each of the to-be-processed frequency bands based on a preset value to obtain multiple feature-enhanced frequency bands includes: Determine a preset weight for each of the to-be-processed frequency bands; wherein, the preset value includes multiple of the preset weights; Fuse each of the to-be-processed frequency bands with the matching preset weight to obtain multiple of the feature-enhanced frequency bands.
9. An audio recognition model training device, characterized in that, Includes: An input unit, configured to obtain multiple historical audio data and input the multiple historical audio data into a to-be-trained audio recognition model; A feature extraction unit, configured to perform frequency-domain feature extraction on each of the historical audio data to obtain multiple initial audio features; A calculation unit, configured to calculate, based on each of the initial audio features, a set of target weights that matches the initial audio feature; A processing unit, configured to obtain a preset parameter set and calculate, based on the preset parameter set and each of the sets of target weights respectively, multiple fusion parameters; Based on each of the fusion parameters, process the initial audio features that match each of the fusion parameters to obtain multiple target audio features, including: a preset processing rule, the processing rule being an activation function; based on the processing rule and each of the fusion parameters, process the initial audio features that match each of the fusion parameters to obtain multiple of the target audio features; A training unit, configured to train the to-be-trained audio recognition model based on multiple of the target audio features to obtain a target audio recognition model; The obtaining of a preset parameter set and calculating, based on the preset parameter set and each of the sets of target weights respectively, multiple fusion parameters includes: Obtain the target weights that match each of the preset parameters, wherein each of the preset parameter sets includes M preset parameters, each of the sets of target weights includes M target weights, and M is a positive integer; Calculate each of the preset parameters with the matching target weights to obtain a plurality of calculation results, including: multiplying each of the preset parameters by the matching target weights to obtain a plurality of calculation results; Fuse the plurality of calculation results to obtain one fusion parameter, including: adding the plurality of calculation results to obtain one fusion parameter.
10. An audio recognition device, characterized in that, Including: An acquisition unit for acquiring a plurality of audio data to be processed; A first acquisition unit for obtaining a plurality of real-time audio features based on the plurality of audio data to be processed; A second acquisition unit for obtaining a plurality of audio features to be input based on the plurality of real-time audio features; A third acquisition unit for inputting the plurality of audio features to be input into the target audio recognition model according to any one of claims 1-4 to obtain target audio features.
11. An audio recognition component, characterized in that, Including: A weight calculation module that acquires a plurality of historical audio data and inputs the plurality of historical audio data into the audio recognition model to be trained; Extract frequency domain features from each of the historical audio data to obtain a plurality of initial audio features; Based on each of the initial audio features, calculate and obtain a set of target weights that match the initial audio features; A perception module, the perception module is connected to the weight calculation module, and the perception module is used to obtain a set of preset parameters and calculate each of the set of target weights based on the set of preset parameters to obtain a plurality of fusion parameters; Based on each of the fusion parameters, process the initial audio features that match each of the fusion parameters to obtain a plurality of target audio features, including: a preset processing rule, the processing rule is an activation function; based on the processing rule and each of the fusion parameters, process the initial audio features that match each of the fusion parameters to obtain a plurality of the target audio features; Train the audio recognition model to be trained based on the plurality of target audio features to obtain a target audio recognition model; The obtaining a set of preset parameters and calculating each of the set of target weights based on the set of preset parameters to obtain a plurality of fusion parameters includes: Obtain the target weights that match each of the preset parameters, where each set of preset parameters includes M preset parameters, each set of target weights includes M target weights, and M is a positive integer; Calculate each of the preset parameters with the matching target weights to obtain a plurality of calculation results, including: multiplying each of the preset parameters by the matching target weights to obtain a plurality of calculation results; Fuse the plurality of calculation results to obtain one fusion parameter, including: adding the plurality of calculation results to obtain one fusion parameter.
12. An electronic device, characterized in that, Including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the method according to any one of claims 1 to 4 or 5 to 8 is implemented.
13. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program, wherein when the computer program runs, it controls the device where the computer-readable storage medium is located to execute the method described in any one of claims 1 to 4 or claims 5 to 8.
Citation Information
Patent Citations
Ambient sound based scene recognition method and device and mobile terminal
CN103456301A
Speech recognition model training method and device, computer equipment and medium
CN114420108A