Audio recognition method, device, medium and chip system
By separating and fusing low-frequency and high-frequency audio components, the method enhances voice recognition accuracy by reducing interference from harmonics and improving MFCC feature generalization.
Patent Information
- Application Number
- CN202010759752.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-07-31
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2040-07-31
AI Technical Summary
In the prior art, in voiceprint recognition, the high-frequency harmonic part interferes with the Mel cepspectral coefficient (MFCC) feature extraction, affecting the recognition accuracy.
The audio is separated into low-frequency channel signals and high-frequency sound source signals through a linear predictor, and appropriate feature extraction algorithms are used respectively, combined with wavelet transform to extract time-frequency features, and fuse them into the final feature vector.
It improves the accuracy of audio recognition, enhances the generalization ability of MFCC feature parameters, and improves the voiceprint recognition effect.
Smart Images

Figure CN114067782B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of speech recognition, and in particular, to an audio recognition method, its device, medium, and chip system. Background Art
[0002] With the rapid development of the Internet and information technology, and the increasing improvement of people's living standards, the requirements for the quality of life and work are also getting higher and higher. As a medium in people's daily life and work, audio greatly affects daily life behaviors. Audio contains extremely rich information, such as environment, language or dialect, emotion, etc. Audio processing is to extract effective audio information in a complex speech environment. By analyzing the audio information extracted from the audio, it is possible to distinguish the type of noise in the environment corresponding to the audio, distinguish the voices of people or objects in the audio (voiceprint recognition), and so on.
[0003] Taking voiceprint recognition as an example, a voiceprint feature refers to the voice feature that can uniquely identify someone or something, and is the sound wave spectrum carrying voice information displayed by electroacoustic instruments. Voiceprint recognition technology is an application technology that realizes automatic identification of the attributes and categories of a pronunciation device based on the physiological and physical characteristics represented by the pronunciation device. Voiceprint recognition generally consists of three parts: audio preprocessing, voice feature parameter extraction, and voiceprint model training and decision-making. Among them, audio feature extraction, as one of the key parts of voiceprint recognition, aims to extract feature parameters that reflect the characteristics of the voice, and its selection will directly affect the overall effect of voiceprint recognition. The selection of voice feature parameters is preferably with the largest inter-class distance and the smallest intra-class distance. Commonly used voice feature parameters in the field of voiceprint recognition are Mel Frequency Cepstral Coefficents (MFCC). Summary of the Invention
[0004] Embodiments of this application provide an audio recognition method, its device, medium, and chip system to improve the accuracy of audio recognition.
[0005] The first aspect of this application provides an audio recognition method, including: obtaining an audio to be recognized; separating a first frequency band range part and a second frequency band range part from the audio to be recognized through a linear predictor, where the frequency of the frequency band included in the first frequency band range part is lower than the frequency of the frequency band included in the second frequency band range part; based on at least one of the first audio feature extracted from the first frequency band range part and the second audio feature extracted from the second frequency band range part, recognizing the audio to determine the type of the audio to be recognized.
[0006] In this method, the low-frequency part representing the channel characteristics in the audio, that is, the channel signal (the first frequency band range part), and the high-frequency harmonic part representing the sound source characteristics, that is, the sound source signal (the second frequency band range part), are separated, and audio features are extracted separately. It can avoid the interference of the high-frequency harmonic part on audio feature extraction algorithms that simulate the cochlear perception ability of the human ear, such as MFCC, so as to improve the accuracy of audio recognition.
[0007] In a possible implementation of the above first aspect, a second audio feature is extracted from the second frequency band range part through wavelet transform, where the second audio feature is a time-frequency feature obtained through wavelet transform.
[0008] In a possible implementation of the above first aspect, the first frequency band range part characterizes the characteristics of the channel of the sounding object that emits the audio to be recognized, and the second frequency band range part characterizes the characteristics of the sound source of the sounding object.
[0009] In a possible implementation of the above first aspect, separating the first frequency band range part and the second frequency band range part from the audio to be recognized through a linear predictor includes: separating the first frequency band range part from the audio to be recognized through a linear predictor, and using the remaining part of the audio to be recognized after separating the first frequency band range part as the second frequency band range part.
[0010] In a possible implementation of the above first aspect, it further includes: extracting a first audio feature from the first frequency band range part through an audio feature extraction algorithm that simulates the cochlear perception ability of the human ear.
[0011] In a possible implementation of the above first aspect, the audio feature extraction algorithm that simulates the cochlear perception ability of the human ear is the Mel Frequency Cepstral Coefficient (MFCC) extraction method, and the first audio feature is the Mel Frequency Cepstral Coefficient (MFCC).
[0012] In a possible implementation of the above first aspect, based on at least one of the first audio feature extracted from the first frequency band range part and the second audio feature extracted from the second frequency band range part, identifying the audio to determine the type of the audio to be recognized includes:
[0013] Matching the first audio feature or the second audio feature of the audio to be recognized with the first audio feature corresponding to the first audio type, and when the matching degree is greater than the first matching degree threshold, determining that the type of the audio to be recognized is the first audio type. That is, one of the first audio feature and the second audio feature is used for audio recognition. For example, by matching the first audio feature or the second audio feature of the audio to be recognized with the audio features of known audio types, it is determined whether the type of the audio to be recognized is the known audio type.
[0014] In a possible implementation of the above first aspect, identifying the audio based on at least one of the first audio feature extracted from the first frequency band range portion and the second audio feature extracted from the second frequency band range portion to determine the type of the audio to be identified includes:
[0015] Fusing the first audio feature and the second audio feature to obtain a fused audio feature, and matching the fused audio feature with the second audio feature corresponding to the second audio type, and when the matching degree is greater than the second matching degree threshold, determining that the type of the audio to be identified is the second audio type.
[0016] That is, in a fused manner, both the first audio feature and the second audio feature are used for audio identification. For example, when the first audio feature and the second audio feature are MFCC feature parameters and time-frequency feature parameters respectively, the two can be linearly fused to form a feature vector. Or they can be linearly fused after normalization processing, or they can be linearly fused after weighting to form a feature vector. Calculate the corresponding eigenvalue of the feature vector, and when the difference between the calculated eigenvalue and the eigenvalue corresponding to the second audio type is greater than the second matching degree threshold, the type of the audio to be identified is the second audio type.
[0017] In a possible implementation of the above first aspect, identifying the audio based on at least one of the first audio feature extracted from the first frequency band range portion and the second audio feature extracted from the second frequency band range portion to determine the type of the audio to be identified includes:
[0018] Inputting the first audio feature, the second audio feature, or the fused audio feature of the first audio feature and the second audio feature into a neural network model to obtain the type of the audio to be identified.
[0019] In a possible implementation of the above first aspect, the audio to be identified includes noise.
[0020] For example, a user wears noise-canceling headphones and takes the subway. The noise-canceling headphones collect the audio in the subway through a microphone. When the audio intensity exceeds the preset sound intensity threshold in the noise-canceling headphones, the noise-canceling headphones separate the channel signal and the sound source signal from the collected audio through a linear filter. Then, extract MFCC feature parameters from the channel signal and extract time-frequency feature parameters from the sound source signal. Finally, identify the audio according to the MFCC feature parameters and the time-frequency feature parameters, and perform noise cancellation through the noise-canceling headphones.
[0021] A second aspect of the present application provides an audio recognition device, including: an acquisition module for acquiring the audio to be recognized; a separation module for separating a first frequency band range part and a second frequency band range part from the audio to be recognized, where the frequency of the frequency band included in the first frequency band range part is lower than the frequency of the frequency band included in the second frequency band range part; and an identification module for identifying the audio based on at least one of a first audio feature extracted from the first frequency band range part and a second audio feature extracted from the second frequency band range part to determine the type of the audio to be recognized. The audio recognition device can implement any of the methods provided in the foregoing first aspect.
[0022] A third aspect of the present application provides a computer-readable medium, characterized in that instructions are stored on the computer-readable medium, and when the instructions are executed on a computer, the computer is caused to execute any of the methods provided in the foregoing first aspect.
[0023] A fourth aspect of the present application provides an electronic device, including: a processor, the processor being coupled to a memory, and the memory storing program instructions, and when the program instructions stored in the memory are executed by the processor, the electronic device is caused to execute any of the methods provided in the foregoing first aspect.
[0024] A fifth aspect of the present application provides a chip system, characterized in that the chip system includes a processor and a data interface, and the processor reads instructions stored on a memory through the data interface to execute any of the methods provided in the foregoing first aspect. Description of the Drawings
[0025] Figure 1 According to some embodiments of the present application, a scenario of noise recognition by the audio recognition method provided by the present application is shown;
[0026] Figure 2 According to some embodiments of the present application, shown is Figure 1 the hardware structure diagram of the noise-canceling earphone shown;
[0027] Figure 3 According to some embodiments of the present application, shown is a flowchart of training a noise scenario recognition model by a server and transplanting the trained noise scenario recognition model to a noise-canceling earphone to achieve intelligent noise cancellation;
[0028] Figure 4 According to some embodiments of the present application, shown is the process of extracting MFCC feature parameters from the channel signals separated in a subway scenario;
[0029] Figure 5(a) According to some embodiments of the present application, shows a time-domain waveform diagram of a sound source signal;
[0030] Figure 5(b) shows a time-domain waveform diagram of a pitch pulse signal extracted from the sound source signal shown in Figure 5(a) according to some embodiments of the present application;
[0031] Figure 5(c) shows a time-domain waveform diagram of the sound source signal shown in Figure 5(a) at different sub-band frequencies according to some embodiments of the present application;
[0032] Figure 6 According to some embodiments of the present application, a flowchart of an audio recognition method is shown;
[0033] Figure 7 According to some embodiments of the present application, a structural schematic diagram of an audio recognition device is shown;
[0034] Figure 8 According to some embodiments of the present application, a structural schematic diagram of an electronic device is shown;
[0035] Figure 9 According to some embodiments of the present application, a block diagram of a system-on-chip (SoC) is shown. Detailed implementation manners
[0036] Illustrative embodiments of the present application include, but are not limited to, an audio recognition method and its device, medium, and electronic device.
[0037] It can be understood that the terms "first", "second", etc. used in the present application may be used in this text to describe various elements, but unless otherwise specified, these elements are not limited by these terms. These terms are only used to distinguish the first element from another element.
[0038] Embodiments of the present application disclose an audio recognition method, its device, medium, and electronic device. The existing MFCC characterizes the vocal tract characteristics of the sound generating device, which contains rich low-frequency vocal tract signal features. However, the high-frequency sound source signal features reflecting the sound source characteristics of the sound generating device cannot be extracted, and the Mel cepstral coefficients are directly extracted from the original audio with mixed high and low frequencies, making it easily contaminated by high-frequency signals, affecting the generalization ability of the Mel cepstral coefficients, and further affecting the accuracy of voiceprint recognition. Some embodiments provided in the present application design a linear predictor to separate the low-frequency part (characterizing the vocal tract characteristics of the sound generating object that emits the audio) and the high-frequency harmonic part (characterizing the sound source characteristics of the sound generating object that emits the audio) in the audio, and then respectively adopt corresponding feature extraction algorithms for the separated low-frequency part and high-frequency harmonic part to extract features, obtaining low-frequency audio features corresponding to the low-frequency part of the audio (hereinafter simply referred to as vocal tract information) and high-frequency audio features corresponding to the high-frequency harmonic part (hereinafter simply referred to as sound source signal). For example, extracting the Mel-scale Frequency Cepstral Coefficients (MFCC) from the vocal tract signal in the audio can make the extracted MFCC feature parameters free from the interference of high-frequency harmonics, better describe the vocal tract characteristics of the sound generating object of the audio, and enhance the generalization ability of the MFCC feature parameters. And, for example, extracting time-frequency feature parameters from the sound source signal separated by the linear predictor in the audio through multi-scale wavelet transform can effectively characterize the sound source characteristics of the sound generating object of the audio. Finally, using the strong complementary characteristics of the vocal tract signal and the sound source signal in the low-frequency part and the high-frequency harmonic part, linearly fusing the above-mentioned MFCC feature parameters extracted from the low-frequency vocal tract signal and the time-frequency feature parameters extracted from the high-frequency sound source signal into a final feature vector, which can more accurately reflect the characteristics of the audio and help improve the effect of audio recognition (such as voiceprint recognition, noise recognition, etc.).
[0039] It can be understood that in the embodiments of the present application, the sound generating object can be the vocal organs of a living being (such as a human) or the sound generating devices of non-living beings (such as musical instruments, machinery, sound generating equipment), etc., various devices capable of generating audio.
[0040] The embodiments of the present application will be further described in detail below with reference to the accompanying drawings.
[0041] Figure 1 According to some embodiments of the present application, a scenario 10 for noise recognition by the audio recognition method provided by the present application is shown. Specifically, as Figure 1As shown in the figure, the scenario 10 includes an electronic device 100 and an electronic device 200. Among them, the electronic device 100 can identify the noise in the scenario where the user is located through the audio recognition method provided by the present application, so that the electronic device 100 determines the type of the scenario where the user is located according to the recognized noise type, and then adaptively adjusts the noise reduction mode according to the determined scenario type to meet the personalized noise reduction needs of the user in different scenarios (such as airports, railway stations, buses, subways, shopping malls, conference rooms, etc.), and improves the user experience.
[0042] It can be understood that the electronic device 100 and the electronic device 200 provided in some embodiments of the present application can be various electronic devices capable of performing audio recognition using the audio recognition method provided by the present application, including but not limited to noise-canceling headphones, servers, tablet computers, smart phones, laptop computers, desktop computers, wearable electronic devices, head-mounted displays, mobile email devices, portable game consoles, portable music players, etc., and a television set embedded or coupled with one or more processors, or other electronic devices capable of accessing the network. It can be understood that the electronic device 100 and the electronic device 200 can collect audio in different scenarios where the user is located through an audio collection device. The audio collection device can be a part of the electronic device 100 or the electronic device 200, or an independent device independent of the electronic device 100 and the electronic device 200, and can be data-connected to the electronic device 100 and the electronic device 200 to send the collected audio to the electronic device 100 and the electronic device 200.
[0043] For the convenience of description, in the following, the electronic device 100 is taken as the noise-canceling headphone 100 and the electronic device 200 is taken as the server 200 as an example to illustrate the technical solution of the present application.
[0044] Figure 2 According to some embodiments of the present application, a hardware structure diagram of a noise-canceling headphone 100 is shown. Specifically, as Figure 2 shown, the noise-canceling headphone 100 includes a data processing chip 110, an audio module 120, a power module 130, a noise reduction circuit 140, a neural-network processing unit (NPU) 150, a microphone 160, and a speaker 170. Among them,
[0045] The microphone 160 is used to collect the audio in the scenario where the user is located.
[0046] The audio module 120 is used to convert digital audio information into an analog audio signal for output by the speaker 170, and is also used to convert the analog audio signal collected by the microphone 160 into a digital audio signal. The audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 170 can be disposed in the processing chip 110, or some functional modules of the audio module 170 can be disposed in the processing chip 110.
[0047] The data processing chip 110 (such as a Digital Signal Processing (DSP) chip) is used to separate the mid-low frequency channel signal and the high-frequency sound source signal in the audio collected by the microphone 160, and can respectively extract MFCC feature parameters from the separated channel signal, extract time-frequency feature parameters from the sound source signal, and fuse the MFCC feature parameters and the time-frequency feature parameters (for example, linearly fuse the two to form a feature vector, or normalize the two, weight the two and then combine them, etc.) to obtain the fusion feature parameters corresponding to the audio. The neural network processor 150 is used to identify the type of noise scene where the user is located according to one of the extracted MFCC feature parameters and time-frequency feature parameters, or the fusion feature parameters obtained by fusing the two, and then adaptively match the corresponding noise reduction mode according to the identified noise scene type. In a possible implementation manner, the neural network processor 150 can be located outside the noise-canceling headset 100, for example, in an electronic device (such as a mobile phone) that works in cooperation with the noise-canceling headset 100.
[0048] The noise reduction circuit 140 is used to generate an electrical signal corresponding to the identified noise scene according to the noise reduction mode determined by the neural network processor 150 (for example, the electrical signal has a phase opposite to that of the noise in the identified scene and the same amplitude). The speaker 170 is used to convert the electrical signal generated by the noise reduction circuit 110 into a sound wave for output, so as to achieve the purpose of noise reduction. The power supply module 130 is used to supply power to the neural network processor 150, the data processing chip 110, the noise reduction circuit 140, and the audio module 120.
[0049] It can be understood that the hardware structure of the noise-canceling headset 100 provided by the embodiments of the present application does not constitute a specific limitation on the noise-canceling headset 100. In other embodiments of the present application, the noise-canceling headset 100 may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements.
[0050] Figure 3 According to some embodiments of the present application, there is shown Figure 1Specific noise reduction technology for the shown scenario. In this technical solution, a noise scenario recognition model is trained by the server 200, and then the trained noise scenario recognition model is transplanted onto the noise reduction headset 100, so that the noise reduction headset 100 can identify the noise type through this noise scenario recognition model and adaptively adjust the noise reduction mode according to the recognized noise scenario.
[0051] Specifically, Figure 3 The shown noise reduction technology mainly includes noise scenario recognition model training and model noise reduction. Among them, during the process of the server 200 training the noise scenario recognition model, a large amount of audio data collected in different scenarios can be respectively separated into the sound source signal and the channel signal corresponding to each audio data through a linear predictor. Then, the MFCC feature parameters of the channel signal and the time-frequency feature parameters of the sound source signal are respectively extracted, and then based on the extracted MFCC feature parameters, time-frequency feature parameters, and the corresponding scenario types, the neural network model is trained to train the noise scenario recognition model.
[0052] It should be noted that the audio recognition method provided in the embodiments of the present application can be applied to various neural network models, for example, Convolutional Neural Network (CNN), Deep Neural Networks (DNN), Recurrent Neural Networks (RNN), Binary Neural Network (BNN), etc. Among them, in specific implementation, the number of layers of the neural network model, the number of nodes in each layer, and the connection parameters of two connected nodes (i.e., the weights on the connection line between the two nodes) can all be preset according to actual needs.
[0053] Such as Figure 3 As shown, the training process of the above noise scenario recognition model includes:
[0054] S301: The server 200 obtains the audio data for model training.
[0055] It can be understood that the server 200 can obtain the audio data for model training in real time, or it can obtain the audio data that has been collected by the audio collection device.
[0056] S302: The server 200 selects the audio data for training by detecting the audio intensity.
[0057] Select the audio data whose audio intensity reaches the sound intensity threshold (for example, 10^-5 (W / m 2)) for training, or the server 200 monitors in real time the intensity of the audio collected by the audio acquisition device (such as a microphone). Only when the server 200 determines that the audio intensity of the collected audio data reaches the sound intensity threshold (such as 10^-5 (W / m 2 )) does the server 200 start subsequent processing. For example, when the server 200 determines that the audio intensity of the collected audio data is less than the sound intensity threshold, it can be considered that the current scene (such as an empty office after work) is very quiet and there is no noise. When the server 200 determines that the audio intensity of the collected audio data is greater than the sound intensity threshold, the server 200 separates the channel signal and the sound source signal in the collected audio, and then extracts the feature parameters of the separated channel signal and sound source signal respectively, obtaining the MFCC feature parameters corresponding to the channel signal and the time-frequency feature parameters corresponding to the sound source signal.
[0058] S303: The server 200 separates the channel signal and the sound source signal from the audio data through linear prediction.
[0059] When training the noise scene recognition model, the server 200 first needs to perform linear prediction on the audio data in multiple scenes used for training that meet the audio intensity condition to separate the low-frequency channel signal (such as the low-frequency part with a frequency below 200 Hz and the middle-frequency part with a frequency between 200 and 3000 Hz in the audio) and the high-frequency sound source signal (such as the part with a frequency above 3000 Hz in the audio).
[0060] For example, taking the extraction of the feature parameters of the audio in the subway scene by the server 200 as an example, in some embodiments, a P-order linear predictor is used to separate the channel signal and the sound source signal in the audio data collected in the subway scene. That is, the current sampling value of the audio x(n) in the subway scene can be predicted by the weighted sum of the sampling values of the audio in the subway scene at the past P historical moments (that is, the channel signal in the audio in the subway scene), then the channel signal in the audio in the subway scene can be expressed as:
[0061]
[0062] where a i is the linear prediction coefficient, and the order of the linear predictor is P-order.
[0063] The transfer function of the P-order linear predictor is:
[0064]
[0065] In order to find the linear prediction coefficient a i, define the error E between the audio x(n) in the subway scenario and its channel signal as follows:
[0066]
[0067] Let the partial derivative of the error E with respect to the linear prediction coefficient a i be equal to 0 to find the minimum value of the error E:
[0068]
[0069] Combining Equation 3 and Equation 4, we get:
[0070]
[0071] Equation 5 can be simplified as follows:
[0072]
[0073] Substitute Equation 6 into Equation 3 to get:
[0074]
[0075] Therefore, if φ(j,i) can be calculated, the linear prediction coefficient a i can be obtained from Equation 7. To find φ(j,i), the autocorrelation function of the audio x(n) in the subway scenario can be defined as follows:
[0076]
[0077] where L represents the length of the audio segment. Therefore, from Equation 6, we can get:
[0078] φ(j,i) = r(j - i) (Equation 9). Since r(j) is an even function, Equation 5 can be simplified to:
[0079]
[0080] The matrix representation form of Equation 10 is:
[0081]
[0082] Solving the above equation can obtain the value of the linear prediction coefficient a i . At this time is the channel signal. By finding the difference between the audio x(n) in the subway scenario and the channel signal , the source signal in the audio in the subway scenario can be obtained.
[0083] S304: The server 200 separately extracts features from the channel signal and the sound source signal separated from the audio data.
[0084] Figure 4 According to some embodiments of the present application, a process of extracting MFCC from the separated channel signal in the subway scenario is shown. Specifically, referring to Figure 4 , preprocessing such as pre-emphasis, framing, and windowing is performed on the channel signal separated from the audio in the subway scenario to enhance the signal-to-noise ratio of the channel signal and improve the processing accuracy. Then, the corresponding spectrum of each short-time analysis window is obtained through FFT (Fast Fourier Transformation) to obtain the spectrum of the channel signal distributed in different time windows on the time axis. The obtained spectrum is passed through a Mel filter bank to obtain the Mel spectrum, so as to convert the linear natural spectrum into the Mel spectrum reflecting the auditory characteristics of the human ear through the Mel spectrum. Then, cepstrum analysis is performed on the Mel spectrum, for example, taking the logarithm of the obtained Mel spectrum, and then the MFCC of the channel signal is obtained through DCT (Discrete Cosine Transform).
[0085] It can be understood that the existing MFCC extraction technologies are applicable to the technical solutions of the present application, not limited to Figure 4 the scheme shown, so no limitation is made here. In addition, the channel features of the above channel signal can also be extracted through other audio feature extraction algorithms that simulate the cochlear perception ability of the human ear, such as the Linear Prediction Cepstrum Coefficient (LPCC) extraction algorithm.
[0086] The following details the process of the server 200 using multi-scale wavelet transform to extract time-frequency features of the sound source signal.
[0087] In some embodiments, in order to eliminate the influence of sound intensity on the time-frequency feature extraction of the sound source signal, amplitude normalization is performed on the sound source signal to obtain the normalized sound source signal x e (n) is:
[0088]
[0089] Among them, xn in formula 12 represents the current sampling value of the audio in the subway scenario, represents the channel signal in the audio in the subway scenario.
[0090] For example, in the embodiment shown in Fig. 5(a), the sound source signal x e(n) consists of periodic pitch pulses, and a Hamming window with a window length of two pitch periods (for example, the pitch period here can be set to 10 ms) is used to extract each pitch pulse in the sound source signal x e (n). The pitch pulse signal x ew (n) as shown in Figure 5(b) is obtained.
[0091] Performing binary discrete wavelet transform on the pitch pulse signal x ew (n) can be calculated by the following formula:
[0092]
[0093] where N is the Hamming window length; ψ * (n) is the conjugate function of the Daubechies wavelet basis; a and b respectively represent the scale parameter and the time factor, reflecting the frequency information and time information of the pitch pulse x ew (n).
[0094] To extract the frequency characteristics, the pitch pulse x ew (n) is represented by K sub-bands W k with different frequency resolutions:
[0095]
[0096] Generally speaking, the frequency range of audio is 300 - 3400 Hz. Therefore, K = 4 can be set to obtain 4 sub-bands with different frequency ranges as shown in Figure 5(c): 2000 - 4000 Hz (W1), 1000 - 2000 Hz (W2), 500 - 1000 Hz (W3), 250 - 500 Hz (W4).
[0097] To retain the time information, each group of wavelet coefficients in formula (14) is divided into M subsets:
[0098]
[0099] where M = 4 is the number of subsets. Calculating the 2-norm of each wavelet coefficient w k in the subset, 4 sub-vectors can be obtained:
[0100]
[0101] where |||| represents the 2-norm. The time-frequency characteristic parameters of the sound source signal can be expressed as follows:
[0102] ω = [ω1, ω2, ω3, ω4] T (Formula 17)
[0103] It can be understood that, in addition to wavelet transform, other algorithms can also be used to extract the time-frequency features in the sound source signal, which is not limited here. For example, the method for extracting the fundamental period.
[0104] S305: Fuse the MFCC extracted from the vocal tract signal and the time-frequency features extracted from the sound source signal to obtain the feature vector of the audio data.
[0105] Specifically, in some embodiments, after the server 200 extracts the feature parameters of a large number of audio collected in different scenarios according to the above method, it can fuse the MFCC feature parameters of the vocal tract signal and the time-frequency feature parameters of the sound source signal in the audio collected in each scenario to obtain the fused feature parameters, and input the fused feature parameters into the neural network model for training. Since the fused feature parameters include both the feature parameters of the vocal tract signal that can reflect the mid-low frequency part in the audio and the feature parameters of the sound source signal that can reflect the high frequency part in the audio, the neural network model trained using the fused feature parameters can more accurately identify the scenario where the user is located, which is conducive to improving the scenario recognition effect.
[0106] For example, in some embodiments, the MFCC feature parameters and the time-frequency feature parameters can be linearly fused to form a feature vector, or they can be linearly fused after normalization, or they can be linearly fused after weighting. In other embodiments, they can also be non-linearly fused, such as multiplying them. In the specific implementation process, the fusion rule can be preset as needed, and this solution does not limit it.
[0107] In addition, it can be understood that in other embodiments, feature fusion may not be performed, but the MFCC features and the time-frequency features are directly input into the neural network model for training.
[0108] S306: Input the obtained fused feature vector into the neural network model for model training.
[0109] Specifically, the server 200 can input the feature vector obtained by fusing the MFCC feature parameters of the channel signal and the time-frequency feature parameters of the sound source signal in the audio collected in each scenario into the neural network model for training. For example, the feature vector obtained by linearly fusing the MFCC feature parameters of the channel signal and the time-frequency feature parameters of the sound source signal in the audio collected in the subway scenario is input into the neural network model. Then, the output of the model (i.e., the training result of training the model with the fused feature vector of the audio collected in the subway scenario) is compared with the data representing the subway scenario to calculate the error (i.e., the difference between the two), the partial derivative of the aforementioned error is calculated, and the weights are updated according to this partial derivative. Until finally the model outputs the data representing the subway scenario, it is considered that the model training is completed. It can be understood that the fused feature parameters in other scenarios can also be input to train the model. Thus, in the training of a large number of sample scenarios, by continuously adjusting the weights, when the output error reaches a very small value (for example, meeting the predetermined error threshold), it is considered that the neural network model converges, and a noise scenario recognition model is trained.
[0110] It can be understood that the trained noise scenario recognition model can only include the above-mentioned trained neural network model, or can also include one or more of the audio intensity detection function, linear prediction function, feature extraction function, and feature fusion function in steps S302 to 305 while having the above-mentioned trained neural network model.
[0111] For the noise-canceling headphones 100, the above-mentioned trained noise scenario recognition model can be transplanted into the noise-canceling headphones 100 for noise reduction processing during the use of the noise-canceling headphones 100. For example, after training the noise scenario recognition model on the server 200, an Android project can be established, the model can be read and parsed through the model reading interface in the aforementioned project, and then compiled to generate an APK (Android application package) file, which is installed in the noise-canceling headphones 100 to complete the transplantation of the noise scenario recognition model.
[0112] Then, the noise-canceling headphones can use the noise scenario recognition model transplanted onto the noise-canceling headphones 100 to perform scenario recognition, identify the corresponding noise scenario. After that, the noise-canceling headphones 100 set the noise reduction mode corresponding to the identified scenario to achieve personalized noise reduction functions.
[0113] It can be understood that the process of using the noise-canceling headphones 100 to perform noise reduction on audio in different scenarios is similar. Next, continue to refer to Figure 3 and combine with Figure 2 and 4, taking the noise reduction of audio in the subway scenario using the noise-canceling headphones 100 as an example, the process of using the noise-canceling headphones 100 for noise reduction will be introduced. In the following embodiments, the noise scenario recognition model only includes the above-mentioned trained neural network model. Specifically, the noise reduction process of the noise-canceling headphones 100 includes:
[0114] S307: Obtain the audio data to be recognized. For example, the microphone 160 of the noise-canceling headphones 100 collects analog audio data in the subway. After being subjected to analog-to-digital conversion by the audio module 120, the digital audio data to be recognized is obtained.
[0115] S308: The data processing chip 110 of the noise-canceling headphones 100 uses a P-order linear filter to separate the channel signal and the source signal from the collected audio data. The specific separation process is similar to the process of separating the channel signal and the source signal using the server 200 described above, and will not be elaborated here.
[0116] S309: The data processing chip 110 of the noise-canceling headphones 100 respectively extracts features from the separated channel signal and source signal, and performs feature fusion on the extracted features. The specific extraction process and fusion process are similar to the process of respectively extracting features from the channel signal and source signal by the server 200 described above, and will not be elaborated here.
[0117] S310: The noise-canceling headphones 100 input the fused feature vector into the noise scenario recognition model transplanted into the neural network processor 150 of the noise-canceling headphones 100 for scenario recognition. For example, for the above-mentioned audio data collected from the subway scenario, the finally recognized result indicates that the user is taking the subway, and the noise reduction scenario is the subway scenario.
[0118] S311: The noise reduction circuit 140 of the noise-canceling headphones 100 generates an electrical signal corresponding to the recognized scenario, then converts the electrical signal into audio data, combines it with the audio data normally output by the noise-canceling headphones 100 to generate noise-reduced audio data, and outputs the noise-reduced audio data.
[0119] It can be understood that different noise reduction modes can be adopted for different scenarios. For subway scenarios, bus scenarios, airport scenarios, etc., the noise-canceling headphones 100 can achieve noise reduction by generating an electrical signal with a phase opposite to and an amplitude equal to that of the noise in the scenario. For example, when a user takes the subway and uses electronic devices such as mobile phones, tablets, etc. for entertainment activities such as listening to music, playing games, watching movies, etc., in order to avoid being disturbed by noisy voices in the surrounding environment, the roar of vehicles driving, etc., the noise reduction circuit 110 in the noise-canceling headphones 100 can generate an electrical signal with a phase opposite to and an amplitude equal to that of the noise in the audio of the scenario, and then superimpose it on the audio data normally output by the noise-canceling headphones 100 to output the noise-reduced audio data. In some other scenarios, the degree of noise reduction should not be too strong. For example, when a user walks through an intersection, the noise-canceling headphones 100 do not adopt too strong noise reduction, because if the degree of noise reduction is too strong and the honking of vehicles driving on the road, the roar of the engine, etc. are completely removed, there is a situation where the user may cause a traffic safety accident due to not hearing external warning sounds. Therefore, the phase of the electrical signal generated by the noise reduction circuit 110 is opposite to that of the noise, but the amplitude is smaller than the amplitude of the noise.
[0120] The above embodiments disclose a solution for using the audio recognition technology of the present application for noise reduction. It can be understood that the audio recognition technology of the present application can also be used for voiceprint recognition. For example, it can be used for voice assistants of electronic devices, voice command recognition of vehicle infotainment systems, voiceprint recognition of users, etc. The embodiments of using the technical solution of the present application for voiceprint recognition of electronic devices are introduced below. Specifically, as Figure 6 shown, including:
[0121] S602: The electronic device collects the user's voice to obtain audio data.
[0122] S604: The electronic device uses a linear filter to separate the channel signal and the source signal from the collected audio data. The specific separation scheme is the same as that shown in Figure 3 and will not be elaborated here.
[0123] S606: The electronic device respectively extracts features from the separated channel signal and source signal, and performs feature fusion on the extracted features to obtain a feature vector. The specific extraction and fusion scheme is the same as that shown in Figure 3 and will not be elaborated here.
[0124] S608: The electronic device matches the obtained feature vector with the feature vector of the voice signal of the legitimate user that has been stored.
[0125] For example, a matching degree threshold A is configured for the feature vector of the voice of a legitimate user. The eigenvalue of the feature vector of the fused user's voice and the eigenvalue of the feature vector of the legitimate user's voice are calculated. When the difference between the two eigenvalues (the absolute value of the difference between the two) is greater than the matching degree threshold A, it is confirmed that the user who emits the voice is a legitimate user. The matching degree threshold A can be configured to 0.5. When the absolute value of the difference between the eigenvalue of the feature vector of the fused user's voice and the eigenvalue of the feature vector of the legitimate user's voice is greater than 0.5, it is confirmed that the user who emits the voice is a legitimate user.
[0126] If they match, go to 610; otherwise, remind the user that the voice cannot be recognized, ask the user to emit the voice again, and enter 602 again.
[0127] S610: The voiceprint recognition is passed, and the electronic device performs corresponding operations. For example, if the electronic device is an access control, the door is opened after the voiceprint recognition is passed. For another example, if the electronic device is a mobile phone, the mobile phone is unlocked after the voiceprint recognition is passed.
[0128] In addition, it can be understood that in other embodiments, in S604, the features of the extracted vocal tract signal and sound source signal may not be fused, but directly in S606, the features of the vocal tract signal and / or sound source signal in the voice of the stored legitimate user are respectively matched with the features of the vocal tract signal and sound source signal obtained in S604.
[0129] For example, matching degree thresholds B are respectively configured for the features of the vocal tract signal and / or sound source signal in the voice of a legitimate user. The eigenvalue of the feature vector of the vocal tract signal or sound source signal of the user's voice is calculated. When the difference between the eigenvalue of the vocal tract signal and the eigenvalue of the feature of the vocal tract signal in the voice of the legitimate user (the absolute value of the difference between the two) is greater than the matching degree threshold B, or when the difference between the eigenvalue of the sound source signal and the eigenvalue of the feature of the sound source signal in the voice of the legitimate user (the absolute value of the difference between the two) is greater than the matching degree threshold B, it is confirmed that the user who emits the voice is a legitimate user.
[0130] Figure 7 According to some embodiments of the present application, a structural schematic diagram of an audio recognition device 700 is provided. As Figure 7 shown, the audio recognition device 700 includes:
[0131] An acquisition module 702, configured to acquire the audio to be recognized;
[0132] A separation module 704, configured to separate the low-frequency part and the high-frequency harmonic part from the audio to be recognized;
[0133] The recognition module 706 recognizes the audio based on at least one of the low audio features extracted from the low-frequency part and the high audio features extracted from the high-frequency harmonic part to determine the type of the audio to be recognized.
[0134] It can be understood that Figure 7 the illustrated audio recognition device 700 corresponds to the audio recognition method provided in this application, and the technical details in the above specific description of the audio recognition method provided in this application still apply to Figure 7 the illustrated audio recognition device 700. For specific descriptions, please refer to the above, and details will not be repeated here.
[0135] According to an embodiment of the present application, Figure 8 a schematic structural diagram of an electronic device 800 is shown. The electronic device 800 can also execute the audio recognition method disclosed in the above embodiments of this application. In Figure 8 it, similar components have the same reference numerals. As Figure 8 shown, the electronic device 800 may include a processor 810, a power module 840, a memory 880, a mobile communication module 830, a wireless communication module 820, a sensor module 890, an audio module 850, a camera 870, an interface module 860, a button 801, a display screen 802, etc.
[0136] It can be understood that the structure schematically shown in the embodiments of the present invention does not constitute a specific limitation on the electronic device 800. In other embodiments of this application, the electronic device 800 may include more or fewer components than those shown, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0137] The processor 810 may include one or more processing units. For example, it may include processing modules or circuits such as a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), a Digital Signal Processing (DSP), a Micro-programmed Control Unit (MCU), an Artificial Intelligence (AI) processor, or a Field Programmable Gate Array (FPGA). Among them, different processing units may be independent devices or integrated in one or more processors. A storage unit may be provided in the processor 810 for storing instructions and data. In some embodiments, the storage unit in the processor 810 is a cache memory 880. The memory 880 mainly includes a program storage area 881 and a data storage area 882. Among them, the program storage area 881 can store an operating system and application programs required for at least one function (such as functions like voice playback and voice recognition). The data storage area 882 can store MFCC feature parameters and time-frequency feature parameters extracted from audio by using the method provided in this application. The neural network model provided in the embodiments of this application can be regarded as an application program in the program storage area 881 that can implement functions such as audio recognition.
[0138] The power supply module 840 may include a power supply, a power management component, etc. The power supply may be a battery. The power management component is used to manage the charging of the power supply and the power supply to other modules. In some embodiments, the power management component includes a charging management module and a power management module. The charging management module is used to receive a charging input from a charger; the power management module is used to connect the power supply, the charging management module, and the processor 810. The power management module receives the input of the power supply and / or the charging management module and supplies power to the processor 810, the display screen 802, the camera 870, the wireless communication module 820, etc.
[0139] The mobile communication module 830 may include, but is not limited to, an antenna, a power amplifier, a filter, a Low noise amplify (LNA), etc. The mobile communication module 830 can provide solutions for wireless communications such as 2G / 3G / 4G / 5G applied to the electronic device 800.
[0140] The wireless communication module 820 may include an antenna and realize the transceiver of electromagnetic waves via the antenna. The wireless communication module 820 may provide solutions for wireless communications applied to the electronic device 800, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR), etc. The electronic device 800 may communicate with the network and other devices through wireless communication technologies.
[0141] In some embodiments, the mobile communication module 830 and the wireless communication module 820 of the electronic device 800 may also be located in the same module.
[0142] The display screen 802 is used to display a human-machine interaction interface, images, videos, etc. The display screen 802 includes a display panel.
[0143] The sensor module 890 may include a proximity light sensor, a pressure sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a distance sensor, a fingerprint sensor, a temperature sensor, a touch sensor, an ambient light sensor, a bone conduction sensor, etc.
[0144] The audio module 850 is used to convert digital audio information into analog audio output, or convert analog audio input into digital audio. The audio module 850 may also be used for audio encoding and decoding. In some embodiments, the audio module 850 may be disposed in the processor 810, or some functional modules of the audio module 850 may be disposed in the processor 810. In some embodiments, the audio module 850 may include a speaker, a receiver, a microphone, and a headphone jack.
[0145] The camera 870 is used to capture still images or videos. An object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element converts the optical signal into an electrical signal, and then transmits the electrical signal to the Image Signal Processing (ISP) to convert it into a digital image signal. The electronic device 800 may implement the shooting function through the ISP, the camera 870, a video codec, a Graphic Processing Unit (GPU), the display screen 802, and an application processor, etc.
[0146] The interface module 860 includes an external memory interface, a universal serial bus (USB) interface, a subscriber identification module (SIM) card interface, etc. The external memory interface can be used to connect an external memory card, such as a Micro SD card, to implement the storage capacity expansion of the electronic device 800. The external memory card communicates with the processor 810 through the external memory interface to implement the data storage function. The universal serial bus interface is used for the electronic device 800 to communicate with other electronic devices. The subscriber identification module card interface is used to communicate with the SIM card installed in the electronic device 800, such as reading the phone number stored in the SIM card or writing the phone number into the SIM card.
[0147] In some embodiments, the electronic device 800 further includes keys 801, a motor, and an indicator, etc. Among them, the keys 801 may include volume keys, power on / off keys, etc. The motor is used to make the electronic device 800 generate a vibration effect, such as generating a vibration when the user's electronic device 800 is called to prompt the user to answer the incoming call of the electronic device 800. The indicator may include a laser indicator, a radio frequency indicator, an LED indicator, etc.
[0148] According to an embodiment of the present application, Figure 9 shows a block diagram of a System on Chip (SoC) 900. In Figure 9 wherein, similar components have the same reference numerals. Additionally, the dashed boxes are optional features of a more advanced SoC. In Figure 9 wherein, the SoC 900 includes: an interconnect unit 950, which is coupled to the application processor 910; a system agent unit 970; a bus controller unit 980; an integrated memory controller unit 940; one or a group of co-processors 920, which may include integrated graphics logic, an image processor, an audio processor, and a video processor; a static random access memory (SRAM) unit 930; a direct memory access (DMA) unit 960. In one embodiment, the co-processor 920 includes a dedicated processor, such as a network or communication processor, a compression engine, a GPU, a high-throughput MIC processor, or an embedded processor, etc.
[0149] Embodiments of the mechanisms disclosed in the present application can be implemented in hardware, software, firmware, or a combination of these implementation methods. Embodiments of the present application can be implemented as a computer program or program code executed on a programmable system, which includes at least one processor, a storage system (including volatile and non-volatile memories and / or storage elements), at least one input device, and at least one output device.
[0150] The program code can be applied to the input instructions to perform the various functions described in this application and generate output information. The output information can be applied to one or more output devices in a known manner. The processing system can include any system having a processor such as, for example, a Digital Signal Processing (DSP), a microcontroller, an Application-Specific Integrated Circuit (ASIC), or a microprocessor.
[0151] The program code can be implemented in a high-level procedural language or an object-oriented programming language to communicate with the processing system. When needed, the program code can also be implemented in assembly language or machine language. In fact, the mechanisms described in this application are not limited to the scope of any particular programming language. In any case, the language can be a compiled language or an interpreted language.
[0152] In some cases, the disclosed embodiments can be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments can also be implemented as instructions carried or stored on one or more transient or non-transient machine-readable (e.g., computer-readable) storage media, which can be read and executed by one or more processors. For example, the instructions can be distributed via a network or via other computer-readable media. Thus, the machine-readable media can include any mechanism for storing or transmitting information in a form readable by a machine (e.g., a computer), including but not limited to, floppy disks, optical disks, optical discs, CD-ROMs, magneto-optical disks, Read Only Memory (ROM), Random Access Memory (RAM), Erasable Programmable Read Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), magnetic or optical cards, flash memory, or tangible machine-readable memories for transmitting information (e.g., carrier waves, infrared signals, digital signals, etc.) in electrical, optical, acoustic, or other forms via the Internet. Thus, the machine-readable media includes any type of machine-readable media suitable for storing or transmitting electronic instructions or information in a form readable by a machine (e.g., a computer).
[0153] In the accompanying drawings, some structural or method features may be shown in a specific arrangement and / or order. However, it should be understood that such a specific arrangement and / or ordering may not be required. Instead, in some embodiments, these features may be arranged in a manner and / or order different from that shown in the illustrative drawings. Additionally, the inclusion of a structural or method feature in a particular figure does not imply that such a feature is required in all embodiments, and in some embodiments, these features may not be included or may be combined with other features.
[0154] It should be noted that each unit / module mentioned in the device embodiments of this application is a logical unit / module. Physically, a logical unit / module can be a physical unit / module, a part of a physical unit / module, or can be implemented as a combination of multiple physical units / modules. The physical implementation manner of these logical units / modules themselves is not the most important. The combination of the functions implemented by these logical units / modules is the key to solving the technical problems proposed in this application. In addition, to highlight the innovative part of this application, the above device embodiments of this application do not introduce units / modules that are not closely related to solving the technical problems proposed in this application. This does not mean that there are no other units / modules in the above device embodiments.
[0155] It should be noted that in the examples and descriptions of this patent, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non - exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one" does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.
[0156] Although this application has been illustrated and described by reference to certain preferred embodiments of the present application, those of ordinary skill in the art should understand that various changes can be made in form and detail without departing from the spirit and scope of the present application.
Claims
1. An audio recognition method, characterized in that, Including: Obtain the audio to be recognized; Separate a first frequency band range part and a second frequency band range part from the audio to be recognized through a linear predictor, wherein the frequencies of the frequency bands included in the first frequency band range part are lower than the frequencies of the frequency bands included in the second frequency band range part; the first frequency band range part characterizes the characteristics of the vocal tract of the sound-producing object that emits the audio to be recognized, and the second frequency band range part characterizes the characteristics of the sound source of the sound-producing object; Fuse the first audio feature extracted from the first frequency band range part and the second audio feature extracted from the second frequency band range part to obtain a fused feature parameter; Identify the audio based on the fused feature parameter to determine the type of the audio to be recognized.
2. The method according to claim 1, wherein Further including: Extract the second audio feature from the second frequency band range part through wavelet transform, wherein the second audio feature is a time-frequency feature obtained through the wavelet transform.
3. The method according to claim 2, characterized in that, The separating the first frequency band range part and the second frequency band range part from the audio to be recognized through a linear predictor includes: Separate the first frequency band range part from the audio to be recognized through a linear predictor, and use the remaining part of the audio to be recognized after separating the first frequency band range part as the second frequency band range part.
4. The method according to claim 3, characterized in that, Further including: Extract the first audio feature from the first frequency band range part through an audio feature extraction algorithm that simulates the perception ability of the human cochlea.
5. The method according to claim 4, wherein The audio feature extraction algorithm that simulates the perception ability of the human cochlea is the Mel Frequency Cepstral Coefficient (MFCC) extraction method, and the first audio feature is the Mel Frequency Cepstral Coefficient (MFCC).
6. The method according to any one of claims 1 to 5, characterized in that, The identifying the audio based on the fused feature parameter to determine the type of the audio to be recognized includes: Input the fused feature parameter into a neural network model to obtain the type of the audio to be recognized.
7. The method according to any one of claims 1 to 5, characterized in that The audio to be recognized includes noise.
8. An audio recognition device, characterized in that, Including: An acquisition module for acquiring the audio to be recognized; A separation module for separating a first frequency band range part and a second frequency band range part from the audio to be recognized, wherein the frequencies of the frequency bands included in the first frequency band range part are lower than the frequencies of the frequency bands included in the second frequency band range part; the first frequency band range part characterizes the characteristics of the vocal tract of the sound-producing object that emits the audio to be recognized, and the second frequency band range part characterizes the characteristics of the sound source of the sound-producing object; An identification module for fusing the first audio feature extracted from the first frequency band range part and the second audio feature extracted from the second frequency band range part to obtain a fused feature parameter, and identifying the audio based on the fused feature parameter to determine the type of the audio to be recognized.
9. A computer-readable medium, characterized in that, Instructions are stored on the computer-readable medium, and when executed on a computer, the computer executes the audio recognition method according to any one of claims 1 to 7.
10. An electronic device, characterized in that, Including: A processor, the processor is coupled to a memory, and the memory stores program instructions. When the program instructions stored in the memory are executed by the processor, the electronic device executes the audio recognition method according to any one of claims 1 to 7.
11. A chip system, characterized in that, The chip system includes a processor and a data interface, and the processor reads instructions stored in a memory through the data interface to execute the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Voice recognition method based on wavelet transformation
CN108198545A