Terminal device, terminal device plug-in, system on chip and related methods
By employing time-frequency masking filtering techniques optimized through frequency point processing and prior speech recognition models, the problem of poor speech enhancement in noisy environments using traditional dual-microphone arrays is solved, thereby improving the accuracy and wake-up rate of speech recognition.
Patent Information
- Application Number
- CN202011404544.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-03
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2040-12-03
AI Technical Summary
Traditional dual-microphone array speech enhancement algorithms are ineffective in real-world applications because the strong correlation between signal and noise means that delayed summation beamforming cannot effectively eliminate coherent noise, resulting in poor enhancement performance.
A frequency-part processing unit is used to perform frequency domain processing on the beamforming signal. By determining the phase difference between the signal and noise at the corrected frequency points, the frequency part dominated by noise is suppressed. Combined with the prior speech recognition model to optimize the time delay estimation, time-frequency masking filtering is realized to enhance the speech signal.
It improves the accuracy and wake-up rate of speech recognition, enhances the signal-to-noise ratio of speech signals, and improves the speech enhancement effect of terminal devices.
Smart Images

Figure CN114613381B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the field of electronics, and more specifically, to a terminal device, a terminal device plug-in, a system on chip and related methods. BACKGROUND
[0002] Automatic speech recognition (ASR) is a technology that recognizes human speech as text, and is widely used in fields such as robot dialogue, smart home, sound box, voice control application (APP), etc. For example, for a sound box, it is generally required that the user speak a specific word, and the sound box starts to work after recognizing the specific word. For example, for a smart refrigerator in a smart home, the smart refrigerator recognizes the instruction "open the refrigerator" or "close the refrigerator" spoken by the user, and performs the corresponding action. In recent years, terminal devices such as sound boxes and smart homes often use a microphone array to collect the user's voice and recognize the voice that wakes up the terminal device to work. A common microphone array is a dual-microphone array. In order to enhance the recognition effect of the voice, voice enhancement is needed. A typical voice enhancement algorithm is beamforming. The most widely used beamforming algorithm is delay-and-sum. That is, the voice signals collected by the two microphones are time-delayed and then the signals after the time-delay are superimposed. Since the voice enhancement effect of the beamforming of the dual-microphone array is limited, a post-filter is usually added after the beamforming.
[0003] The post-filter in the traditional sense is to assume that the signal and the noise are not related, and then to estimate the power spectrum of the signal and the noise, so as to realize Wiener filtering on the beamformed signal. This algorithm has a relatively high requirement for the assumption. In the actual use environment of the intelligent terminal device, the signal and the noise usually have a strong correlation, and the delay-and-sum beamforming does not eliminate the coherent noise, so the enhancement effect is poor. SUMMARY
[0004] Therefore, the present disclosure aims to improve the voice enhancement effect of the dual-microphone terminal device.
[0005] To achieve this purpose, according to an aspect of the present disclosure, the present disclosure provides a terminal device, comprising:
[0006] a first microphone;
[0007] a second microphone;
[0008] a beamformer configured to superimpose the time delay between a first signal and a second signal after the time delay is filled, the first signal being a signal received by the first microphone, and the second signal being a signal received by the second microphone, to become a beamformed signal.
[0009] a frequency bin part processing unit, configured to divide the first signal and the second signal into frequency bin parts in a frequency domain, and process each frequency bin part to enhance speech in the beamformed signal.
[0010] Optionally, the processing of each frequency bin part comprises:
[0011] determining a corrected phase difference of the first signal and the second signal at the frequency bin, determining whether the corrected phase difference satisfies a first predetermined condition, and suppressing the frequency bin part if the first predetermined condition is not satisfied.
[0012] Optionally, the first predetermined condition comprises that the corrected phase difference is less than a first threshold, and the suppressing comprises filtering out the frequency bin part.
[0013] Optionally, the first predetermined condition comprises that the corrected phase difference is less than a first threshold, and the suppressing comprises attenuating the frequency bin part by a predetermined ratio if the corrected phase difference is between the first threshold and a second threshold, and filtering out the frequency bin part if the corrected phase difference is greater than the second threshold, wherein the second threshold is greater than the first threshold.
[0014] Optionally, the determining the corrected phase difference of the first signal and the second signal at the frequency bin comprises:
[0015] determining a difference of phase angles of the first signal and the second signal;
[0016] determining a time delay of the first signal and the second signal;
[0017] subtracting a product of an angular frequency of the frequency bin and the time delay from the difference of the phase angles to obtain the corrected phase difference.
[0018] Optionally, the determining the difference of phase angles of the first signal and the second signal comprises:
[0019] determining a phase angle of the first signal and a phase angle of the second signal according to real and imaginary parts of the first signal and the second signal after being transformed into the frequency domain, respectively;
[0020] subtracting the phase angle of the second signal from the phase angle of the first signal to obtain the difference of the phase angles.
[0021] Optionally, the determining the time delay of the first signal and the second signal comprises:
[0022] obtaining a candidate time delay set;
[0023] For a candidate delay in the candidate delay set, a candidate modified phase difference is obtained by subtracting the product of the angular frequency of the frequency point and the candidate delay from the phase angle difference; if it is determined that the candidate modified phase difference of the frequency point does not satisfy a second predetermined condition, the frequency point part of the beamforming signal is suppressed, and the suppressed beamforming signal is input into a prior speech recognition model, and a probability of recognizing a specific word is output by the prior speech recognition model.
[0024] The candidate delay in the candidate delay set with the maximum output probability of the prior speech recognition model is taken as the determined delay.
[0025] Optionally, the terminal device further includes an identification unit configured to perform speech recognition on the signal output by the frequency point part processing unit.
[0026] Optionally, the terminal device further includes a processor configured to perform a corresponding action according to a speech recognition result.
[0027] Optionally, the terminal device includes a sound box, and the corresponding action includes turning on the sound box.
[0028] According to an aspect of the present disclosure, a terminal device is provided, including:
[0029] a reference microphone configured to receive a first signal;
[0030] a plurality of other microphones configured to respectively receive second signals;
[0031] a beamformer configured to pad each second signal to a time delay of the first signal, and superimpose each time-delay-padded second signal and the first signal to become a beamforming signal;
[0032] a frequency point part processing unit configured to divide the first signal and the second signals into frequency point parts in a frequency domain, and process each frequency point part to enhance speech in the beamforming signal.
[0033] Optionally, the processing of each frequency point part includes:
[0034] determining a modified phase difference of the first signal and each second signal at a frequency point; determining whether an average value of the determined modified phase differences satisfies a first predetermined condition; and if the first predetermined condition is not satisfied, suppressing the frequency point part.
[0035] According to an aspect of the present disclosure, a terminal device plug-in is provided for plugging into a terminal device having a first microphone, a second microphone, and a beamformer configured to superimpose first and second signals after time delay compensation therebetween to form a beamformed signal, the first signal being received by the first microphone and the second signal being received by the second microphone, the terminal device plug-in comprising:
[0036] a frequency bin processing unit configured to divide the first and second signals into frequency bins in a frequency domain, and to process each frequency bin to enhance speech in the beamformed signal.
[0037] According to an aspect of the present disclosure, a terminal device plug-in is provided for plugging into a terminal device having a reference microphone, a plurality of other microphones, and a beamformer configured to superimpose each of the second signals after time delay compensation therebetween with respect to the first signal to form a beamformed signal, the first signal being received by the reference microphone and the second signals being received by the plurality of other microphones, the terminal device plug-in comprising a frequency bin processing unit configured to divide the first and second signals into frequency bins in a frequency domain, and to process each frequency bin to enhance speech in the beamformed signal.
[0038] According to an aspect of the present disclosure, a system-on-chip is provided for connecting to inputs of first and second microphones of a terminal device and to an output of a beamformer of the terminal device, the beamformer configured to superimpose first and second signals after time delay compensation therebetween to form a beamformed signal, the first signal being received by the first microphone and the second signal being received by the second microphone, the system-on-chip comprising a frequency bin processing unit configured to divide the first and second signals into frequency bins in a frequency domain, and to process each frequency bin to enhance speech in the beamformed signal.
[0039] Optionally, the processing of each frequency bin includes determining a modified phase difference between the first and second signals at the frequency bin, determining whether the modified phase difference satisfies a first predetermined condition, and suppressing the frequency bin if the first predetermined condition is not satisfied.
[0040] Optionally, the system-on-chip further comprises a recognition unit configured to perform speech recognition on a signal output by the frequency bin processing unit.
[0041] According to an aspect of the present disclosure, there is provided an on-chip system connected with an input of a reference microphone and a plurality of other microphones of a terminal device and an output of a beamformer of the terminal device, the reference microphone receiving a first signal, the plurality of other microphones respectively receiving second signals, the beamformer complementing a time delay of each of the second signals with respect to the first signal, and superimposing each of the second signals after the time delay is complemented with the first signal to become a beamformed signal, the on-chip system comprising: a frequency bin processing unit configured to divide the first signal and the second signals into frequency bins in a frequency domain, and process each of the frequency bins to enhance speech in the beamformed signal.
[0042] Optionally, the processing of each of the frequency bins comprises: determining a corrected phase difference of the first signal and each of the second signals at the frequency bin; determining whether an average of the determined corrected phase differences satisfies a first predetermined condition; and suppressing the frequency bin if the first predetermined condition is not satisfied.
[0043] According to an aspect of the present disclosure, there is provided an audio processing method of a terminal device, wherein the terminal device has a first microphone and a second microphone, the method comprising:
[0044] superimposing a time delay complemented between a first signal and a second signal to become a beamformed signal, the first signal being a signal received by the first microphone, and the second signal being a signal received by the second microphone;
[0045] dividing the first signal and the second signal into frequency bins in a frequency domain;
[0046] processing each of the frequency bins to enhance speech in the beamformed signal.
[0047] Optionally, the processing of each of the frequency bins comprises:
[0048] determining a corrected phase difference of the first signal and each of the second signals at the frequency bin;
[0049] determining whether the corrected phase difference satisfies a first predetermined condition;
[0050] suppressing the frequency bin if the first predetermined condition is not satisfied.
[0051] According to an aspect of the present disclosure, there is provided an audio processing method of a terminal device, wherein the terminal device has a reference microphone and a plurality of other microphones, the reference microphone receiving a first signal, and the plurality of other microphones respectively receiving second signals, the method comprising:
[0052] delay the second signals relative to the first signal, and superimpose the delayed second signals and the first signal to form a beamforming signal;
[0053] divide the first signal and the second signals into frequency bin portions in a frequency domain;
[0054] process the frequency bin portions to enhance speech in the beamforming signal.
[0055] Optionally, the processing of the frequency bin portions includes determining a corrected phase difference of the first signal and the second signals at the frequency bin, determining whether an average of the determined corrected phase differences satisfies a first predetermined condition, and suppressing the frequency bin portion if the first predetermined condition is not satisfied.
[0056] Embodiments of the present disclosure adopt a time-frequency masking method, i.e., dividing a frequency domain of a signal received by a microphone into frequency bin portions, and processing the frequency bin portions respectively. If a sound source signal dominates at a frequency bin, the portion of the signal at the frequency bin is preserved. If noise dominates at the frequency bin, the portion of the signal at the frequency bin is suppressed. Such preservation or suppression of the frequency bin portions can more finely suppress the influence of noise, and is helpful to improve speech enhancement effect, compared with preservation or suppression of the entire signal. BRIEF DESCRIPTION OF DRAWINGS
[0057] The above and other objects, features and advantages of the present disclosure will become more apparent from the following description when taken in conjunction with the accompanying drawings, in which:
[0058] Figure 1A is an appearance view of a dual-microphone array terminal device according to an embodiment of the present disclosure;
[0059] Figure 1B is an appearance view of a multi-microphone array terminal device according to an embodiment of the present disclosure;
[0060] Figure 2A is a structural view of a dual-microphone array terminal device according to an embodiment of the present disclosure;
[0061] Figure 2B is a structural view of a dual-microphone array terminal device according to an embodiment of the present disclosure;
[0062] Figure 3 is a probability density function curve of a corrected phase difference at different signal-to-noise ratios according to an embodiment of the present disclosure;
[0063] Figure 4 is a table showing terminal device wake-up rates of multiple tests when different filtering strategies are adopted;
[0064] Figure 5is a flow chart of a dual microphone array terminal device audio processing method according to one embodiment of the present disclosure;
[0065] Figure 6 is a flow chart of a multi-microphone array terminal device audio processing method according to one embodiment of the present disclosure. DETAILED DESCRIPTION
[0066] The present disclosure is described in detail based on the embodiments below, but the present disclosure is not limited to only these embodiments. In the following detailed description of the present disclosure, some specific details are described in detail. The present disclosure can also be fully understood without the description of these details by those skilled in the art. In order to avoid confusion of the essence of the present disclosure, the well-known methods, processes, flows are not described in detail. In addition, the drawings are not necessarily drawn to scale.
[0067] The following terms are used herein.
[0068] Sound box: a box that plays sound. There is generally a speaker hole on the box, and sound is played from the speaker hole to the outside.
[0069] Microphone: an energy conversion device that converts a sound signal into an electrical signal.
[0070] Wake up: when a predetermined word or any word spoken by a user is recognized, the terminal device enters a working state or starts a certain function.
[0071] Wake-up rate: the probability that, when a user speaks a predetermined word or any word, the terminal device can correctly recognize the predetermined word or any word, and the terminal device enters a working state or starts a certain function. It can be determined by the number of times the predetermined word is correctly recognized by the terminal device divided by the total number of times the user speaks the predetermined word or any word in the test.
[0072] Wake-up word: the predetermined word spoken by the user, which makes the terminal device enter a working state or start a certain function. For example, the wake-up word is "hello, XX" or "please turn on the sound box", and when the sound box recognizes that the user speaks these wake-up words, it starts to work or starts a certain function.
[0073] Microphone array: an array composed of multiple microphones deployed on the terminal device, used to receive sound signals of a user (which may contain a wake-up word spoken by the user) respectively, for subsequent processing, so as to identify whether the user speaks a wake-up word or any word in the subsequent processing, and decide whether to wake up the terminal device.
[0074] Beamforming: A concept originated from adaptive antenna. Signal processing at the receiving end, by processing the signals received from multiple antenna elements, can synthesize the desired signal. From the perspective of antenna pattern, this is equivalent to forming a beam with a specified direction. Forming beams in multiple specified directions is called multi-directional beamforming.
[0075] Delay-and-sum: A beamforming method for two-microphone array, which delays the signals received from two microphones, and then superimposes the signals after delay.
[0076] Filtering: Due to the limited speech enhancement effect of beamforming, filtering is usually needed after beamforming, to further reduce the noise in the beamformed signal, and further enhance the speech.
[0077] Time-frequency masking: After converting the time-domain speech signal into frequency domain, the frequency-domain speech signal can be divided into frequency bin parts. Each frequency bin part can be dominated by the sound source signal or by noise. Time-frequency masking refers to suppressing the frequency bin parts dominated by noise, to improve the speech enhancement effect.
[0078] Terminal device plug-in: An attached device inserted into a general terminal device, to make the terminal device have a certain special function.
[0079] System on chip: Refers to a complete system integrated on a single chip, which groups all or part of necessary electronic circuits. A complete system generally includes central processing unit (CPU), memory, and peripheral circuits, etc. For speech processing, it can also include beamformer, filtering unit, etc.
[0080] Dual microphone array terminal device embodiment
[0081] Figure 1AThe appearance of the dual microphone array terminal device is illustrated by taking a sound box as an example. Those skilled in the art should understand that the terminal device can be a smart home terminal (such as a smart air conditioner that responds to a control voice spoken by a person to perform various adjustment actions of the air conditioner, a smart refrigerator that responds to a control voice spoken by a person to perform various adjustment actions of the refrigerator, a television remote control that responds to a voice command spoken by a person to adjust the channel and the sound size, a smart doorbell that automatically rings in response to a voice spoken by a person, and a smart lock that verifies the opening of the door in response to a voice spoken by a person), a vehicle terminal (such as a smart car machine), a conference terminal, etc. In addition, the present disclosure can be embodied as a system on a chip (chip) or a plug-in in addition to being embodied as a terminal device. The system on a chip (chip) or plug-in plays a role in enhancing voice, and after being inserted into the general sound box, smart home terminal, vehicle terminal, conference terminal, etc., the voice of the general sound box, smart home terminal, vehicle terminal, conference terminal, etc. is enhanced, and the accuracy of recognizing user voice is higher.
[0082] As shown in Figure 1A A first microphone 111 and a second microphone 112 are arranged on the outer surface of the terminal device 100 (such as a sound box) to form a dual microphone array. The first microphone 111 and the second microphone 112 are used to respectively receive a sound signal (which includes a signal of a sound source signal propagating to the first microphone 111 and the second microphone 112 and a noise signal) so as to be converted into respective electrical signals for subsequent processing of the electrical signals by the terminal device 100. In addition, the surface of the terminal device 100 can have an array of speaker holes (not shown). The sound played by the terminal device 100 after being activated and working is output through the array of speaker holes. It should be noted that the sound signal received by the first microphone 111 and the second microphone 112 is only used for the terminal device 100 to wake up (i.e., to start working), and the sound played by the terminal device 100 after starting to work can be a sound file transmitted by a built-in disk or Bluetooth, and does not necessarily come from the sound received by the first microphone 111 and the second microphone 112.
[0083] As shown in Figure 2A The terminal device 100 can further include a beamformer 120, a filtering unit 130, a recognition unit 140, and a processor 150.
[0084] Beamforming in the field of audio processing refers to a signal synthesis process of multiple microphones received signals, so as to achieve the effect of speech enhancement. One method of synthesis processing is to overlap after filling the time delay between signals. The beamformer 120 overlaps the first signal and the second signal after filling the time delay between them to become a beamformed signal. The first signal is the signal received by the first microphone 111, and the second signal is the signal received by the second microphone 112. The time for the sound source signal to propagate to the first microphone 111 is equal to the distance from the sound source to the first microphone 111 divided by the speed of sound. The time for the sound source signal to propagate to the second microphone 112 is equal to the distance from the sound source to the second microphone 112 divided by the speed of sound. Therefore, the time delay for the sound source signal to propagate to the first microphone 111 and the second microphone 112 is equal to the difference between the distance from the sound source to the first microphone 111 and the distance from the sound source to the second microphone 112 divided by the speed of sound. After determining the time delay, assuming that the second signal is later than the first signal, the first signal is delayed by the determined time delay to obtain a first delayed signal. Overlapping the first delayed signal with the second signal, the amplitude of the obtained signal is approximately doubled. In this way, the amplitude of the sound source signal is enhanced, and the effect of speech enhancement is achieved.
[0085] Due to the limited effect of speech enhancement by beamforming, it is usually necessary to filter after beamforming to further reduce noise in the beamformed signal and further enhance speech. In the traditional sense, filtering is performed by assuming that the signal and the noise are not related, and then performing power spectrum estimation of the signal and the noise, so as to realize Wiener filtering on the double microphone beamforming output signal. This algorithm has high requirements for the assumption and is suitable for environments where the signal and the noise are not related. In the actual use environment of intelligent terminal devices, the signal and the noise usually have strong correlation. By filtering through the above method, the speech enhancement effect is poor. In order to solve the above problem, the embodiment of the present disclosure improves the filtering scheme. On the basis of having obtained an accurate propagation delay, the relationship between the delay and the phase difference of each frequency point of the beamformed output signal is constructed, which is used as a judgment standard to realize binary time-frequency masking of the speech signal, thereby realizing speech enhancement.
[0086] After converting the time-domain speech signal into the frequency domain, the frequency-domain speech signal can be divided into signal parts on each frequency point, i.e. frequency point parts. In each frequency point part, the sound source signal may be dominant or the noise may be dominant. For the frequency point part in which the noise is dominant, suppression can be performed, such as direct filtering or attenuation by a certain proportion. In this way, the method of suppressing the frequency point part with more noise for each frequency point is more likely to increase the amplitude of the useful signal and suppress the amplitude of the noise, thereby enhancing the speech in the beamformed signal. The above process is time-domain masking, which is realized by the frequency point part processing unit 131.
[0087] The frequency bin part processing unit 131, when processing each frequency bin, determines the corrected phase difference between the first signal and the second signal at the frequency bin in the frequency domain. The phase difference generally refers to the difference between the phases of two signals in the frequency domain. However, in fact, the two signals are caused by the same sound source signal propagating to two different microphones, and the sound source signal reaches the two microphones with a time delay. Therefore, the time delay needs to be corrected to eliminate the influence of the phase difference caused by the phase subtraction in the frequency domain. The phase difference after the correction is called the corrected phase difference.
[0088] Suppose the time delay difference between the first microphone 111 and the second microphone 112 is . Only the signal of one of the microphones needs to be time-delayed (the second signal of the second microphone 112 is time-delayed below), and then the first signal and the second signal are summed and averaged, that is,
[0089]
[0090] where z(t) is the beamforming output signal of the double microphone, is the first signal, is the second signal.
[0091] The first signal collected by the first microphone 111 and the second signal collected by the second microphone 112 are Fourier transformed, that is (the attenuation coefficient is omitted)
[0092]
[0093] where, is the frequency domain signal of the first signal Fourier transformed to the frequency domain, is the result of the sound source signal transformed to the frequency domain. Since the attenuation of the sound source signal propagating to the first microphone 111 is ignored, that is =1, the sound source signal propagating to the first microphone 111 is still , t represents the time frame, w represents the angular frequency, w=2πf, f is the frequency of the frequency bin, is the frequency domain representation of the noise signal at the first microphone 111. is the frequency domain signal of the second signal Fourier transformed to the frequency domain. Due to the influence of the phase, the phase of the sound source signal propagating to the second microphone 112 is different from the phase of the sound source signal propagating to the first microphone 111 by a time delay difference of , so the signal of the sound source signal propagating to the second microphone 112 becomes , is the frequency domain representation of the noise signal at the second microphone 112.
[0094] Thus, the corrected phase difference of the first signal and the second signal at the frequency point can be determined As follows:
[0095]
[0096] is the phase angle of the above-mentioned frequency domain signal , which can be determined according to the real part and the imaginary part of . Let = a + jb, the phase angle of is equal to arctan (b / a). is the phase angle of the above-mentioned frequency domain signal , which can be determined according to the real part and the imaginary part of . Let = a + jb, the phase angle of is equal to arctan (b / a). is the difference of the phase angles of the first signal and the second signal without time delay correction. However, the signals reach the first microphone 111 and the second microphone 112 with time delay, so the phase difference subtraction is performed alone without considering the influence of the time delay. In formula 1-4, the difference of the phase angles is subtracted by the product of the angular frequency of the frequency point and the time delay , to obtain the corrected phase difference. w = 2pf, f is the frequency of the frequency point. The method for determining the time delay is described below.
[0097] In the absence of noise interference, should be , because in the absence of noise interference, the difference of the phase angles of the first signal and the second signal should be exactly caused by the time delay, so the difference of the phase angles after subtracting the influence of the time delay is exactly 0. In practice, due to the existence of noise, is not 0, but satisfies a certain rule. By using the numerical value of , it can be evaluated whether the sound source signal or the noise dominates at each frequency point.
[0098] In F. Mustiere, R. Nakagawa, K. Wojcicki, I. Merks, and T. Zhang, "Dual-microphone phase-difference-based SNR estimation with applications to speech enhancement," 2016 IEEE International Workshop on Acoustic Signal Enhancement (IWAENC), Xi'an, 2016, pp. 1-5, doi: 10.1109 / IWAENC.2016.7602935, the probability density function of the corrected phase difference in the dual-microphone scenario can be obtained as:
[0099]
[0100] wherein is a parameter , and SNR represents the signal-to-noise ratio. Figure 3 The curves of as a function of are shown in Figs. 1-5 for SNRs of 20, 10, 0, -10, and -20, respectively. In Fig. 1, the solid curve represents as a function of for an SNR of 20, the curve connected by x's represents as a function of for an SNR of 10, the curve connected by o's represents as a function of for an SNR of 0, the curve connected by *'s represents as a function of for an SNR of -10, and the curve connected by +'s represents as a function of for an SNR of -20. As the SNR increases, the probability density function becomes more concentrated, i.e., the probability at the center point 0 becomes greater and the slope becomes steeper. This means that the higher the frequency point signal-to-noise ratio, the more likely the corresponding is near 0. When SNR = 0, = 0.6538, which means that the phase difference correction value of each frequency point satisfying this signal-to-noise ratio has a probability of 0.6538 within [-1, 1]; fixed u (u is a positive number), as the signal-to-noise ratio increases, the probability of falling within this interval also increases; fixed the signal-to-noise ratio, as u increases, The probability of falling into this interval also increases.
[0101] In the whole frequency domain, the frequency points with high SNR will be concentrated around 0, while the frequency points with low SNR will be more likely to be far away from 0. Therefore, for a certain frequency point of the frequency domain signal, if the first predetermined condition is met when u is fixed, it means that the SNR of the frequency point is large, and the sound source signal dominates on the frequency point, so the part of the frequency point of the beamforming signal can be retained. The first predetermined condition can be that the absolute value of the modified phase difference is less than a first threshold, i.e. , where is the first threshold. Common values of u can be 1, etc. If the first predetermined condition is not met, it means that the SNR of the frequency point is small, and the noise dominates on the frequency point, so the part of the frequency point of the beamforming signal can be suppressed. The meaning of suppression is attenuation or complete filtering. In one embodiment, if the first predetermined condition is met, the part of the frequency point can be completely filtered out. In another embodiment, if the first predetermined condition is met, suppression can also be performed separately for different cases. If the modified phase difference exceeds the first threshold not too much, it can be considered that the noise of the part of the frequency point can be tolerated appropriately, i.e., partial attenuation, and only when it exceeds the first threshold much, complete filtering is performed. A second threshold r can be set, for example, taking the value of π. If the modified phase difference is between the first threshold u and the second threshold r, the part of the frequency point is attenuated by a predetermined ratio (e.g., 50%); if the modified phase difference is greater than the second threshold r, the part of the frequency point is completely filtered out.
[0102] The first predetermined condition can also be an asymmetric predetermined condition in addition to , for example , and are positive numbers.
[0103] When is adopted, is the frequency domain filtering coefficient, is the first threshold, and the embodiments of the present disclosure correspond to the following time-frequency domain filtering coefficient selection scheme:
[0104]
[0105] For each part of the frequency point of the frequency domain transformed signal of the beamforming signal, the corresponding is calculated according to the above formulas 1-6, i.e., the enhanced signal of the time-frequency masking post-filtering algorithm is obtained, i.e., the following formula:
[0106] That is
[0107]
[0108] The inverse Fourier transform is performed to transform back to the time domain, and a speech signal after speech enhancement of the embodiment of the present disclosure is obtained.
[0109] From the above, it can be seen that the selection of the first threshold u determines the effect of the embodiment of the present disclosure, which can be selected according to the actual situation. When the environmental signal-to-noise ratio is relatively large, the estimation of the direction of the sound source is relatively accurate, a smaller first threshold u can be selected, so that the speech enhancement effect is further improved. When the environmental signal-to-noise ratio is relatively small, a larger first threshold u can be selected to ensure that the frequency points of the signal will not be incorrectly classified due to incorrect direction estimation.
[0110] The above discussion is based on the time delay between the first signal and the second signal is known. However, in fact, the time delay is not easy to accurately estimate, therefore, the embodiment of the present disclosure proposes a time delay estimation scheme based on a prior speech model and a binary time-frequency masking filtering algorithm.
[0111] The determination process of the time delay between the above-mentioned first signal and the second signal is discussed below.
[0112] First, a candidate time delay set is obtained. The candidate time delay is a discrete point of possible value of the time delay between the first signal and the second signal. At first glance, the possible value of the candidate time delay is infinite, but in fact it is finite. The reason is that, as mentioned above, the time delay of the sound source signal propagating to the first microphone 111 and the second microphone 112 is equal to the difference between the distance from the sound source to the first microphone 111 and the distance from the sound source to the second microphone 112 divided by the speed of sound, and according to the principle that the difference between two sides of a triangle is less than the third side, the difference between the distance from the sound source to the first microphone 111 and the distance from the sound source to the second microphone 112 is not greater than the distance between the first microphone 111 and the second microphone 112. Therefore, the maximum value of the time delay is the distance between the first microphone 111 and the second microphone 112 divided by the speed of sound. Since the distance between the first microphone 111 and the second microphone 112 is generally small, and the number of sampling points corresponding to the time delay multiplied by the sampling frequency can only be an integer, the maximum time delay point number (the number of points used must be an integer, so if the product of the sampling frequency is not an integer, it should be rounded to the nearest integer to obtain the maximum time delay point number) can be obtained by multiplying the maximum time delay value of the distance between the first microphone 111 and the second microphone 112 divided by the speed of sound by the sampling frequency of the frequency point. Then, all positive integers not greater than the maximum time delay point number are possible time delay points to be taken. For example, if the maximum time delay point number is 10, then 1, 2, 3, …, 10 are possible time delay points to be taken. The positive integers not greater than the maximum time delay point number are converted back to the corresponding time delay, i.e. each is divided by the sampling frequency, and the time delay obtained is used as a candidate time delay in the candidate time delay set.
[0113] Then, for each candidate time delay in the candidate time delay set, the process of binary time-frequency masking filtering as described above is repeated. That is, for each candidate time delay in the candidate time delay set, the modified phase difference of the first signal and the second signal at the frequency point is subtracted from the product of the angular frequency of the frequency point and the candidate time delay to obtain a candidate modified phase difference. Then, it is determined whether the candidate modified phase difference satisfies a second predetermined condition. The second predetermined condition can be that the absolute value of the candidate modified phase difference is less than a third threshold value, where w is the third threshold value. The value of w is, for example, 1. If the second predetermined condition is satisfied, the part of the beamforming signal at the frequency point is preserved. If the second predetermined condition is not satisfied, the part of the beamforming signal at the frequency point is suppressed. As mentioned above, one kind of suppression can be complete filtering out, and the other kind of suppression can be setting a fourth threshold value v greater than w. If then partial attenuation (e.g. 50% attenuation) is performed; if If the ratio is greater than the third threshold value w, the signal is filtered out completely. In this way, the time-frequency mask filtering is completed. Then, the signal completed with the time-frequency mask filtering is input into a pre-trained prior speech recognition model, and a probability of recognizing a specific word, i.e., a wake-up rate, is output by the prior speech recognition model.
[0114] The third threshold value w selected above can be less than the first threshold value u selected. This is because the third threshold value w is only used to screen out a suitable time delay, and is not the final frequency point classification, and too high a third threshold value w is not conducive to the efficiency of screening. However, the first threshold value u is used for the final frequency point classification and filtering, and too small a first threshold value u is easy to cause the frequency points of the signal to be misclassified.
[0115] The prior speech recognition model is a predetermined trained speech recognition model, which can use a mixed Gaussian model, etc. The speech recognition model can be trained using a pure wake-up word sample set. The pure wake-up word is a voice signal received by a microphone, which only contains the voice signal of the wake-up word emitted by the sound source, and does not contain noise signals. The voice signal sample of the wake-up word spoken by the user without noise is input into the speech recognition model, and the text recognized by the speech recognition model is compared with the text corresponding to the voice signal sample. If they are consistent, it is considered to be a correct wake-up. If they are not consistent, it is considered to be an incorrect wake-up. After all the samples in the set are input into the speech recognition model, the ratio of the number of times of correct wake-up of the speech recognition model to the number of all samples is the wake-up rate. If the wake-up rate is higher than a predetermined wake-up rate threshold (for example, 95%), it is considered that the prior speech recognition model is successfully trained. Any signal with user voice input into the model can obtain a probability of whether the recognized text is a specific word, i.e., a wake-up rate. Since each candidate time delay in the candidate time delay set obtains a corresponding wake-up rate, finally, a candidate time delay with the maximum wake-up rate can be found as the determined time delay. The time delay found based on the prior speech model is more accurate, and can improve the speech enhancement effect of the embodiment of the present disclosure. The embodiment of the present disclosure proposes a sound source estimation scheme based on a prior speech model and an improved post-filtering algorithm, thereby ensuring the speech enhancement performance.
[0116] Figure 4Table 1 is a table showing the terminal device wake-up rates of multiple tests when different filtering strategies are adopted. When the signals entering the first microphone 111 and the second microphone 112 are pure sound source signals without any noise, and the terminal device does not perform any additional audio processing, in 5 tests, the wake-up rates of the recognition unit 140 are 57%, 75.8%, 83.6%, 90.8% respectively, and the average wake-up rate is 76.8%; when the signals entering the first microphone 111 and the second microphone 112 are sound source signals mixed with noise, and the terminal device does not perform any additional audio processing, in 5 tests, the wake-up rates of the recognition unit 140 are 9.6%, 10.6%, 15.4%, 33.4% respectively, and the average wake-up rate is 17.3%; when the signals entering the first microphone 111 and the second microphone 112 are sound source signals mixed with noise, and the terminal device performs a Delayed Summation (DS) beamforming algorithm on the signals received by the first and second microphones 111 and 112, and a traditional Generalized Cross-Correlation (GCC) algorithm to calculate the time delay, the wake-up rates of the recognition unit 140 are 14.0%, 11.4%, 15.8%, 30.4% respectively, and the average wake-up rate is 17.9%; when the signals entering the first microphone 111 and the second microphone 112 are sound source signals mixed with noise, and the terminal device performs a Delayed Summation (DS) beamforming algorithm on the signals received by the first and second microphones 111 and 112, and uses the above prior speech recognition model to calculate the time delay, the wake-up rates of the recognition unit 140 are 18.2%, 18.6%, 18.6%, 40.0% respectively, and the average wake-up rate is 23.9%; when the signals entering the first microphone 111 and the second microphone 112 are sound source signals mixed with noise, and the terminal device performs a Delayed Summation (DS) beamforming algorithm on the signals received by the first and second microphones 111 and 112, and uses a traditional Generalized Cross-Correlation (GCC) algorithm to calculate the time delay, and uses the time-frequency masking method of the present embodiment for post-filtering (PF), the wake-up rates of the recognition unit 140 are 13.6%, 13.4%, 16.4%, 30.6% respectively, and the average wake-up rate is 18.5%; when the signals entering the first microphone 111 and the second microphone 112 are sound source signals mixed with noise, and the terminal device performs a Delayed Summation (DS) beamforming algorithm on the signals received by the first and second microphones 111 and 112, and uses the prior speech recognition model of the present embodiment to calculate the time delay, and uses the time-frequency masking method of the present embodiment for post-filtering (PF), the wake-up rates of the recognition unit 140 are 46.8%, 31.6%, 76.2%, 89.6% respectively, and the average wake-up rate is 57.6%.
[0117] FromFigure 4 As can be seen from the table, the latency estimated by the prior speech recognition model improves the wake-up rate of the processed data for both the delay-sum (DS) algorithm and the delay-sum (DS) + post-filtering (PF) algorithm, which indicates that the latency estimation is more accurate than the generalized cross-correlation algorithm (GC). After adding the post-filtering (PF), the wake-up rate of the delay-sum (DS) + post-filtering (PF) + prior speech recognition model is significantly improved compared to the delay-sum (DS) + prior speech recognition model, which reflects the performance advantage of the time-frequency masking post-filtering (PF) algorithm.
[0118] The recognition unit 140 performs speech recognition on the signal output by the filtering unit 130. A speech recognition model can be used in the recognition unit 140. It is a machine learning model, which can be trained in advance by the following method: a set of sound signal samples is constructed in advance, each sound signal sample in the set is pre-known to correspond to a word label spoken by a user; each sound signal sample in the set of sound signal samples is input into the machine learning model, and the word spoken by the user in the sample is recognized by the machine learning model, if the recognized word is consistent with the word label, it is determined to be successful; if the ratio of the sound signal samples in the set that are determined to be successful by the machine learning model exceeds a predetermined ratio (e.g., 95%), the machine learning model is considered to be successfully trained; otherwise, the coefficients in the machine learning model are adjusted until the ratio of the sound signal samples in the set that are determined to be successful by the machine learning model exceeds the predetermined ratio. After the model is successfully trained, the signal output by the filtering unit 130 is input into the recognition unit 140 to obtain the recognized text.
[0119] When the recognition unit 140 recognizes the text, the user can be fed back the text to confirm whether the recognition result is accurate. The user can be fed back through voice playback, such as synthesized voice playing "Do you want to do XX operation?" through the speaker. If the terminal device has a screen, the text can also be displayed on the screen. If the user approves the recognition result, the processor 150 performs the corresponding action according to the recognition result. If the user does not approve the recognition result, the user speaks the voice of the action that the terminal device needs to perform again. The terminal device identifies again according to the foregoing process, and feeds back the confirmation to the user again until the user confirms that it is correct. This way of confirmation can improve the accuracy of the terminal device performing the action.
[0120] In addition, when the recognition unit 140 recognizes the text, the user can not be fed back the text to confirm every time. When the recognition unit 140 recognizes the text, a confidence of the text will be generated. When the confidence is higher than a predetermined confidence threshold, the user can not be fed back the text for confirmation, but the processor 150 directly performs the corresponding action. This way can balance the accuracy and efficiency of the terminal device performing the action.
[0121] Then, the processor 150 performs a corresponding action according to the text recognized by the recognition unit 140. When the terminal device is a sound box, the corresponding action can be to wake up the sound box and make the sound box start working (e.g., playing music). When the terminal device is a smart home terminal, the corresponding action is to complete a certain control of the smart home terminal, for example, for a smart air conditioner, the corresponding action is to turn on the air conditioner, adjust the temperature to a certain value, adjust the wind direction, etc. When the terminal device is a vehicle terminal, the corresponding action is a certain control of the vehicle terminal, for example, displaying navigation to the destination, etc. When the terminal device is a conference terminal, the corresponding action is the setting of conference parameters, functions, etc., such as increasing the microphone of a certain person, etc.
[0122] To sum up, in the scenario of the smart terminal device, due to the too close distance between the microphones, the noise correlation is very strong, and the effect of using the traditional Wiener post-filter is poor, therefore, the embodiment of the disclosure constructs a time delay based binary time-frequency masking coefficient, different processing is performed on the frequency points with different signal-to-noise ratios, and the wake-up rate of the smart terminal device for specific words is improved. In addition, in order to solve the problem of inaccurate time delay estimation when the signal is enhanced, the embodiment of the disclosure proposes a new time delay estimation algorithm based on the actual smart device scenario, and the time delay estimation accuracy is improved.
[0123] In addition, the embodiment of the disclosure also proposes a terminal device plug-in (not shown) which has a frequency point partial processing unit 131 required for the post-processing of the frequency point part by the embodiment of the disclosure, and can be inserted into a general terminal device 100 having a first microphone 111, a second microphone 112 and a beamformer 120 to help the general terminal device 100 improve the voice enhancement effect. Since the detailed structure and principle of the filter unit 130, the first microphone 111, the second microphone 112 and the beamformer 120 have been described above, they will not be described again. In addition, the recognition unit 140 can be included in the terminal device plug-in.
[0124] The above terminal device plug-in can be embodied in the form of a system on chip, i.e., in the form of a chip. The chip can be assembled in a general terminal device 100 to help the general terminal device 100 improve the voice enhancement effect. The above system on chip can be connected with the input of the first microphone 111 and the second microphone 112 of the terminal device 100 and the output of the beamformer 120 of the terminal device 100. The connection with the input of the first microphone 111 and the second microphone 112 of the terminal device 100 is to obtain the first signal and the second signal to obtain the corrected phase difference of the first signal and the second signal at each frequency point. The connection with the output of the beamformer 120 is to perform time-frequency masking processing on the signal output by the beamformer 120, i.e., to suppress the frequency point part when the corrected phase difference of the first signal and the second signal at the frequency point does not meet the first predetermined condition.
[0125] In addition, the time-frequency masking scheme of the present disclosure can also be used in a terminal device with a multi-microphone (more than two microphones) array. As shown in Figure 1B a six-microphone sound box, which has a reference microphone 113 and a plurality of (five in this embodiment) other microphones 114. The reference microphone 113 receives a first signal. The first signal includes a signal formed by the propagation of a sound source signal to the reference microphone 113 and a noise signal. The plurality of other microphones 114 respectively receive a second signal. The second signal includes a signal formed by the propagation of the sound source signal to the corresponding other microphone 114 and a noise signal. Figure 1B
[0126] As shown in Figure 2B , the terminal device 100 can further include a beamformer 120, a frequency bin processing unit 131, an identification unit 140, and a processor 150. Figure 2B The beamformer 120 of the embodiment is different from the beamformer 120 of the embodiment in that, since Figure 2A there are a plurality of other microphones 114, each of which receives a second signal, each second signal has a possible different time delay compared to the first signal, therefore, the beamformer 120 needs to fill in the time delay of each second signal compared to the first signal, and then superimpose each time-delayed second signal with the first signal to obtain a beamforming signal, while Figure 2B the beamformer 120 of the embodiment only needs to fill in the time delay of one second signal compared to the first signal and superimpose it. Figure 2A
[0127] Figure 2B The frequency bin processing unit 131 of the embodiment is used to divide the first signal and the second signal into frequency bins in the frequency domain, and process each frequency bin to enhance the speech in the beamforming signal. It is consistent with the frequency bin processing unit 131 of the embodiment except that Figure 2A the frequency bin processing unit 131 of the embodiment is different in that, since Figure 2A in the embodiment, there is only one second signal, so there is only one corrected phase difference with the first signal, but Figure 2B in the embodiment, there are a plurality of second signals, so there are a plurality of phase differences with the first signal, therefore, when determining whether the first predetermined condition is met, it is necessary to determine whether the average of each corrected phase difference meets the first predetermined condition. In addition to this, the other parts are consistent with the frequency bin processing unit 131 of the embodiment, which can be referred to the description of the frequency bin processing unit 131 of the embodiment. Figure 2A Figure 2A In addition, the identification unit 140 and the processor 150 of the embodiment are also respectively consistent with the identification unit 140 and the processor 150 of the embodiment, which can be referred to the description of the identification unit 140 and the processor 150 of the embodiment.
[0128] Figure 2B Figure 2A Figure 2A The relevant descriptions of the identification unit 140 and the processor 150.
[0129] Furthermore, this disclosure also proposes a terminal device plug-in (not shown) that includes a frequency point processing unit 131 required for post-processing of frequency points according to this disclosure. This plug-in can be inserted into a general-purpose terminal device 100 having a reference microphone 113, multiple other microphones 114, and a beamformer 120, thereby improving the voice enhancement effect of the general-purpose terminal device 100. Since the detailed structure and principle of the frequency point processing unit 131, the reference microphone 113, the multiple other microphones 114, and the beamformer 120 have been described above, they will not be repeated here. Additionally, the identification unit 140 may be included in this terminal device plug-in.
[0130] The aforementioned terminal device plug-in can be implemented as a system-on-a-chip (SoC), i.e., it exists in the form of a chip. The chip can be assembled into a general-purpose terminal device 100 to help improve the voice enhancement effect of the general-purpose terminal device 100. The aforementioned SoC can be connected to the inputs of the reference microphone 113 and multiple other microphones 114 of the terminal device 100, and the output of the beamformer 120 of the terminal device 100. The connection to the inputs of the reference microphone 113 and multiple other microphones 114 of the terminal device 100 is for acquiring a first signal and a second signal to obtain the corrected phase difference between the first signal and the second signal at each frequency point. The connection to the output of the beamformer 120 is for performing time-frequency masking processing on the signal output by the beamformer 120, that is, suppressing the frequency portion of the beamformed signal when the corrected phase difference between the first signal and the second signal at that frequency point does not meet a first predetermined condition.
[0131] like Figure 5 As shown, according to one embodiment of this disclosure, an audio processing method for a dual-microphone array terminal device is also provided, which is executed by a terminal device 100. The terminal device 100 has a first microphone 111 and a second microphone 112. The method includes:
[0132] Step 510: After the time delay between the first signal and the second signal is filled in, they are superimposed to form a beamforming signal. The first signal is the signal received by the first microphone 111, and the second signal is the signal received by the second microphone 112.
[0133] Step 520: Divide the first signal and the second signal into frequency point components in the frequency domain;
[0134] Step 530: Process each frequency segment to enhance the speech in the beam program signal.
[0135] The implementation details of this method are as described above. Figure 2AThe embodiments have been described and can be referred to. Figure 2A The embodiments are as described above, so they will not be elaborated upon.
[0136] like Figure 6 As shown, according to one embodiment of this disclosure, an audio processing method for a multi-microphone array terminal device is also provided, which is executed by a terminal device 100. The terminal device 100 has a reference microphone 113 and a plurality of other microphones 114. The reference microphone 113 receives a first signal, and the plurality of other microphones 114 respectively receive a second signal. The method includes:
[0137] Step 610: Pad the time delay of each second signal relative to the first signal, and superimpose the second signals after time delay padding with the first signal to form a beamforming signal;
[0138] Step 620: Divide the first signal and the second signal into frequency point components in the frequency domain;
[0139] Step 630: Process each frequency segment to enhance the speech in the beam program signal.
[0140] Optionally, step 630 includes: determining the corrected phase difference between the first signal and each of the second signals at a frequency point; determining whether the average value of the determined corrected phase differences meets a first predetermined condition; if the first predetermined condition is not met, suppressing the frequency point portion.
[0141] The implementation details of this method are as described above. Figure 2B The embodiments have been described and can be referred to. Figure 2B The embodiments are as described above, so they will not be elaborated upon.
[0142] Business value of the present disclosure
[0143] Based on a priori speech model and an improved post-frequency point processing algorithm, this disclosure proposes a microphone speech enhancement scheme for smart devices. Compared with traditional technologies, it greatly improves the speech enhancement effect, increasing the wake-up rate of smart terminal devices by 30%-50%, and has good market prospects.
[0144] It should be understood that the various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.
[0145] It is to be understood that the foregoing description is descriptive only, certain embodiments having been described in particularity. Other embodiments are within the scope of the claims. In some cases the acts or steps recited in the claims can be performed in a different order and still accomplish the desired results. Additionally, the processes depicted in the figures do not necessarily require the particular order shown, or sequential order to achieve the desired results. In certain implementations, multitasking and parallel processing can be advantageous or necessary.
[0146] It is to be understood that the use of singular herein or the use of "one" or "another" does not limit the scope of the claims to one item, unless specifically stated otherwise. Furthermore, a structure described or shown as including certain parts in a certain arrangement is not necessarily limited to only an arrangement of these parts, but can include subsystems with these parts arranged differently.
[0147] It is also to be understood that the terminology and phraseology employed herein are for descriptive purposes and should not be regarded as limiting. The use of such terms and phrases does not thereby erode the scope of the application, which is defined in the appended claims. It is also to be understood that the use of "including," "having," "having at least," "including at least," "comprising," or "comprising at least" does not exclude the presence of other elements or steps. Other modifications, variations, and alternatives are also possible. Accordingly, the claims as presented are intended to cover all such alternatives, modifications, and equivalents.
Claims
1. A terminal device, comprising: a first microphone; a second microphone; a beamformer configured to superimpose a first signal and a second signal after time delay compensation, the first signal being a signal received by the first microphone, the second signal being a signal received by the second microphone, the superimposed signal being a beamformed signal; a frequency bin processing unit configured to divide the first signal and the second signal into frequency bins in a frequency domain, and process each frequency bin to enhance speech in the beamformed signal; wherein the processing each frequency bin comprises: determining a modified phase difference of the first signal and the second signal at the frequency bin, and attenuating the frequency bin by a predetermined ratio if the modified phase difference is between a first threshold and a second threshold, and filtering out the frequency bin if the modified phase difference is greater than the second threshold, wherein the second threshold is greater than the first threshold.
2. The terminal device of claim 1, wherein, The determining the modified phase difference of the first signal and the second signal at the frequency bin comprises: determining a difference of phase angles of the first signal and the second signal; determining a time delay of the first signal and the second signal; subtracting a product of an angular frequency of the frequency bin and the time delay from the difference of the phase angles to obtain the modified phase difference.
3. The terminal device of claim 2, wherein, The determining the difference of phase angles of the first signal and the second signal comprises: determining a phase angle of the first signal and a phase angle of the second signal according to real and imaginary parts of the first signal and the second signal after being transformed into the frequency domain, respectively; subtracting the phase angle of the first signal from the phase angle of the second signal to obtain the difference of the phase angles.
4. The terminal device of claim 3, wherein, The determining the time delay of the first signal and the second signal comprises: obtaining a candidate time delay set; for a candidate time delay in the candidate time delay set, subtracting a product of an angular frequency of the frequency bin and the candidate time delay from the difference of the phase angles to obtain a candidate modified phase difference, and if it is determined that the candidate modified phase difference of the frequency bin does not satisfy a second predetermined condition, suppressing the frequency bin of the beamformed signal, and inputting the suppressed beamformed signal into a priori speech recognition model to output a probability of recognizing a wake-up word by the a priori speech recognition model; selecting a candidate time delay in the candidate time delay set with the maximum output probability of the a priori speech recognition model as the determined time delay.
5. The terminal device of claim 1, further comprising: a recognition unit configured to perform speech recognition on a signal output by the frequency bin processing unit.
6. The terminal device of claim 5, further comprising: a processor configured to perform a corresponding action according to a speech recognition result.
7. The terminal device of claim 6, wherein, The terminal device comprises a sound box, and the corresponding action comprises turning on the sound box. 8.A terminal device, comprising: a reference microphone configured to receive a first signal; a plurality of other microphones configured to respectively receive second signals; a beamformer configured to compensate time delays of the second signals with respect to the first signal, and superimpose the second signals after the time delay compensation with the first signal to obtain a beamformed signal; a frequency bin processing unit configured to divide the first signal and the second signals into frequency bins in a frequency domain, and process each frequency bin to enhance speech in the beamformed signal; wherein the processing each frequency bin comprises: determining a modified phase difference of the first signal and the second signal at a frequency bin, if the modified phase difference is between a first threshold value and a second threshold value, attenuating the frequency bin by a predetermined ratio; if the modified phase difference is greater than the second threshold value, filtering out the frequency bin, wherein the second threshold value is greater than the first threshold value. 9.A terminal device plug-in for plugging into a terminal device having a first microphone, a second microphone, and a beamformer configured to superimpose first and second signals after time delay compensation therebetween to form a beamformed signal, the first signal being received by the first microphone, the second signal being received by the second microphone, the terminal device plug-in comprising: a frequency bin processing unit configured to divide the first and second signals into frequency bins in a frequency domain, and process each frequency bin to enhance speech in the beamformed signal; wherein the processing each frequency bin comprises: determining a modified phase difference of the first signal and the second signal at a frequency bin, if the modified phase difference is between a first threshold value and a second threshold value, attenuating the frequency bin by a predetermined ratio; if the modified phase difference is greater than the second threshold value, filtering out the frequency bin, wherein the second threshold value is greater than the first threshold value. 10.A terminal device plug-in for plugging into a terminal device having a reference microphone, a plurality of other microphones, and a beamformer, the reference microphone receiving a first signal, the plurality of other microphones respectively receiving second signals, the beamformer compensating time delays of the second signals with respect to the first signal, and superimposing each time delay compensated second signal with the first signal to form a beamformed signal, the terminal device plug-in comprising: a frequency bin processing unit configured to divide the first and second signals into frequency bins in a frequency domain, and process each frequency bin to enhance speech in the beamformed signal; wherein the processing each frequency bin comprises: determining a modified phase difference of the first signal and the second signal at a frequency bin, if the modified phase difference is between a first threshold value and a second threshold value, attenuating the frequency bin by a predetermined ratio; if the modified phase difference is greater than the second threshold value, filtering out the frequency bin, wherein the second threshold value is greater than the first threshold value. 11.A system on chip connected to inputs of first and second microphones of a terminal device and to an output of a beamformer of the terminal device, the beamformer configured to superimpose first and second signals after time delay compensation therebetween to form a beamformed signal, the first signal being received by the first microphone, the second signal being received by the second microphone, the system on chip comprising: a frequency bin processing unit configured to divide the first and second signals into frequency bins in a frequency domain, and process each frequency bin to enhance speech in the beamformed signal; wherein the processing each frequency bin comprises: determining a modified phase difference of the first signal and the second signal at a frequency bin, if the modified phase difference is between a first threshold value and a second threshold value, attenuating the frequency bin by a predetermined ratio; if the modified phase difference is greater than the second threshold value, filtering out the frequency bin, wherein the second threshold value is greater than the first threshold value. determining a corrected phase difference of the first signal and the second signal at the frequency bin, if the corrected phase difference is between a first threshold and a second threshold, attenuating the frequency bin by a predetermined ratio, and if the corrected phase difference is greater than the second threshold, filtering out the frequency bin, wherein the second threshold is greater than the first threshold.
12. The system on chip of claim 11, further comprising: a recognition unit configured to perform speech recognition on the signal output by the frequency bin processing unit. 13.A system on chip, connected to inputs of a reference microphone and a plurality of other microphones of a terminal device and to an output of a beamformer of the terminal device, the reference microphone receiving a first signal, the plurality of other microphones respectively receiving second signals, the beamformer aligning a time delay of each second signal with respect to the first signal, and superimposing each time-delay-aligned second signal with the first signal to form a beamformed signal, the system on chip comprising: a frequency bin processing unit configured to divide the first signal and the second signals into frequency bins in a frequency domain, and process each frequency bin to enhance speech in the beamformed signal. wherein the processing of each frequency bin comprises: determining a corrected phase difference of the first signal and the second signal at the frequency bin, if the corrected phase difference is between a first threshold and a second threshold, attenuating the frequency bin by a predetermined ratio, and if the corrected phase difference is greater than the second threshold, filtering out the frequency bin, wherein the second threshold is greater than the first threshold.
14. A terminal device audio processing method, wherein, the terminal device having a first microphone and a second microphone, the method comprising: superimposing a time-delay-aligned combination of the first signal and the second signal to form a beamformed signal, the first signal being a signal received by the first microphone, and the second signal being a signal received by the second microphone; dividing the first signal and the second signals into frequency bins in a frequency domain; processing each frequency bin to enhance speech in the beamformed signal. wherein the processing of each frequency bin comprises: determining a corrected phase difference of the first signal and the second signal at the frequency bin, if the corrected phase difference is between a first threshold and a second threshold, attenuating the frequency bin by a predetermined ratio, and if the corrected phase difference is greater than the second threshold, filtering out the frequency bin, wherein the second threshold is greater than the first threshold.
15. A method of audio processing for a terminal device, wherein, the terminal device having a reference microphone and a plurality of other microphones, the reference microphone receiving a first signal, the plurality of other microphones respectively receiving second signals, the method comprising: aligning a time delay of each second signal with respect to the first signal, and superimposing each time-delay-aligned second signal with the first signal to form a beamformed signal; dividing the first signal and the second signals into frequency bins in a frequency domain; processing each frequency bin to enhance speech in the beamformed signal. wherein the processing of each frequency bin comprises: determining a corrected phase difference of the first signal and the second signal at the frequency bin, if the corrected phase difference is between a first threshold and a second threshold, attenuating the frequency bin by a predetermined ratio, and if the corrected phase difference is greater than the second threshold, filtering out the frequency bin, wherein the second threshold is greater than the first threshold. determining a modified phase difference of the first signal and the second signal at the frequency bin, if the modified phase difference is between a first threshold and a second threshold, attenuating the frequency bin portion by a predetermined ratio, if the modified phase difference is greater than the second threshold, filtering out the frequency bin portion, wherein the second threshold is greater than the first threshold.
16. A terminal device, comprising: a first microphone; a second microphone; a beamformer configured to superimpose time delays between a first signal and a second signal to form a beamformed signal, the first signal being received by the first microphone, the second signal being received by the second microphone; a frequency bin processing unit configured to divide the first signal and the second signal into frequency bins in a frequency domain, and process each frequency bin to enhance speech in the beamformed signal; wherein the processing each frequency bin comprises: determining a difference between phase angles of the first signal and the second signal, determining a time delay between the first signal and the second signal, subtracting a product of an angular frequency of the frequency bin and the time delay from the difference between the phase angles to obtain a modified phase difference of the first signal and the second signal at the frequency bin, determining whether the modified phase difference satisfies a first predetermined condition, and suppressing the frequency bin portion if the first predetermined condition is not satisfied.
17. The terminal device of claim 16, wherein, the first predetermined condition comprises that the modified phase difference is less than a first threshold, and the suppressing comprises filtering out the frequency bin portion.
18. The terminal device of claim 16, wherein, the first predetermined condition comprises that the modified phase difference is less than a first threshold, and the suppressing comprises attenuating the frequency bin portion by a predetermined ratio if the modified phase difference is between the first threshold and a second threshold, and filtering out the frequency bin portion if the modified phase difference is greater than the second threshold, wherein the second threshold is greater than the first threshold.
19. The terminal device of claim 16, wherein, the determining the difference between the phase angles of the first signal and the second signal comprises: determining a phase angle of the first signal and a phase angle of the second signal according to real and imaginary parts of the first signal and the second signal after being transformed into the frequency domain, respectively; and subtracting the phase angle of the first signal from the phase angle of the second signal to obtain the difference between the phase angles.
20. The terminal device of claim 16, wherein, the determining the time delay between the first signal and the second signal comprises: obtaining a candidate time delay set; for a candidate time delay in the candidate time delay set, subtracting a product of an angular frequency of the frequency bin and the candidate time delay from the difference between the phase angles to obtain a candidate modified phase difference, suppressing the frequency bin portion of the beamformed signal if the candidate modified phase difference of the frequency bin does not satisfy a second predetermined condition, and inputting the suppressed beamformed signal into a priori speech recognition model to output a probability of recognizing a wake-up word by the a priori speech recognition model; taking a candidate time delay in the candidate time delay set with the maximum output probability of the a priori speech recognition model as the determined time delay.
21. The terminal device of claim 16, further comprising: a recognition unit configured to perform speech recognition on a signal output by the frequency bin processing unit.
22. The terminal device of claim 21, further comprising: a processor configured to perform a corresponding action according to a speech recognition result.
23. The terminal device of claim 22, wherein, The terminal device comprises a sound box, and the corresponding action comprises turning on the sound box.
24. A terminal device comprising: a first microphone; a second microphone; a beamformer configured to superimpose first signals received by the first microphone and second signals received by the second microphone after time delays between the first signals and the second signals are filled in, to form a beamformed signal; a frequency bin processing unit configured to divide the first signals and the second signals into frequency bins in a frequency domain, and to process each frequency bin to enhance speech in the beamformed signal; wherein the processing of each frequency bin comprises: determining a difference between phase angles of the first signals and the second signals; determining a time delay between the first signals and the second signals; subtracting a product of an angular frequency of the frequency bin and the time delay from the difference between the phase angles, to obtain a corrected phase difference between the first signals and the second signals at the frequency bin; determining whether the corrected phase difference satisfies a first predetermined condition; and suppressing the frequency bin if the first predetermined condition is not satisfied.
25. A terminal device plug-in configured to be plugged into a terminal device having a first microphone, a second microphone, and a beamformer configured to superimpose first signals received by the first microphone and second signals received by the second microphone after time delays between the first signals and the second signals are filled in, to form a beamformed signal, the terminal device plug-in comprising: a frequency bin processing unit configured to divide the first signals and the second signals into frequency bins in a frequency domain, and to process each frequency bin to enhance speech in the beamformed signal; wherein the processing of each frequency bin comprises: determining a difference between phase angles of the first signals and the second signals; determining a time delay between the first signals and the second signals; subtracting a product of an angular frequency of the frequency bin and the time delay from the difference between the phase angles, to obtain a corrected phase difference between the first signals and the second signals at the frequency bin; determining whether the corrected phase difference satisfies a first predetermined condition; and suppressing the frequency bin if the first predetermined condition is not satisfied.
26. A terminal device plug-in configured to be plugged into a terminal device having a reference microphone configured to receive first signals, a plurality of other microphones configured to receive second signals, and a beamformer configured to fill in time delays between the second signals and the first signals, and to superimpose each of the second signals after the time delays are filled in, with the first signals, to form a beamformed signal, the terminal device plug-in comprising: a frequency bin processing unit configured to divide the first signals and the second signals into frequency bins in a frequency domain, and to process each frequency bin to enhance speech in the beamformed signal; wherein the processing of each frequency bin comprises: determining a difference between phase angles of the first signals and the second signals; determining a time delay between the first signals and the second signals; subtracting a product of an angular frequency of the frequency bin and the time delay from the difference between the phase angles, to obtain a corrected phase difference between the first signals and the second signals at the frequency bin; determining whether the corrected phase difference satisfies a first predetermined condition; and suppressing the frequency bin if the first predetermined condition is not satisfied. determining a difference between phase angles of the first signal and the second signal; determining a time delay of the first signal and the second signal; subtracting a product of an angular frequency of the frequency bin and the time delay from the difference between the phase angles, to obtain a corrected phase difference between the first signal and the second signal at the frequency bin; determining whether the corrected phase difference satisfies a first predetermined condition; and suppressing the frequency bin if the first predetermined condition is not satisfied. 27.A system-on-chip, connected to inputs of first and second microphones of a terminal device and to an output of a beamformer of the terminal device, the beamformer configured to superimpose second signals after time delays between the second signals and a first signal are filled in, the first signal being a signal received by the first microphone, the second signals being signals received by the second microphone, the system-on-chip comprising: a frequency bin processing unit configured to divide the first signal and the second signals into frequency bins in a frequency domain, and to process each frequency bin to enhance speech in a beamformed signal; wherein the processing each frequency bin comprises: determining a difference between phase angles of the first signal and the second signal; determining a time delay of the first signal and the second signal; subtracting a product of an angular frequency of the frequency bin and the time delay from the difference between the phase angles, to obtain a corrected phase difference between the first signal and the second signal at the frequency bin; determining whether the corrected phase difference satisfies a first predetermined condition; and suppressing the frequency bin if the first predetermined condition is not satisfied.
28. The system on chip of claim 27, further comprising: a recognition unit configured to perform speech recognition on a signal output by the frequency bin processing unit. 29.A system-on-chip, connected to inputs of a reference microphone and a plurality of other microphones of a terminal device and to an output of a beamformer of the terminal device, the reference microphone receiving a first signal, the plurality of other microphones respectively receiving second signals, the beamformer configured to superimpose each second signal with the first signal after a time delay between the second signal and the first signal is filled in, to obtain a beamformed signal, the system-on-chip comprising: a frequency bin processing unit configured to divide the first signal and the second signals into frequency bins in a frequency domain, and to process each frequency bin to enhance speech in the beamformed signal; wherein the processing each frequency bin comprises: determining a difference between phase angles of the first signal and the second signal; determining a time delay of the first signal and the second signal; subtracting a product of an angular frequency of the frequency bin and the time delay from the difference between the phase angles, to obtain a corrected phase difference between the first signal and the second signal at the frequency bin; determining whether the corrected phase difference satisfies a first predetermined condition; and suppressing the frequency bin if the first predetermined condition is not satisfied.
30. A method of audio processing for a terminal device, wherein, the terminal device having the first and second microphones, the method comprising: superimposing second signals after time delays between the second signals and a first signal are filled in, to obtain a beamformed signal, the first signal being a signal received by the first microphone, the second signals being signals received by the second microphone; dividing the first signal and the second signal into frequency bin parts in a frequency domain; processing each frequency bin part to enhance speech in the beamforming signal; wherein the processing each frequency bin part comprises: determining a difference of phase angles of the first signal and the second signal; determining a time delay of the first signal and the second signal; subtracting a product of an angular frequency of the frequency bin and the time delay from the difference of the phase angles to obtain a corrected phase difference of the first signal and the second signal at the frequency bin; determining whether the corrected phase difference satisfies a first predetermined condition; and suppressing the frequency bin part if the first predetermined condition is not satisfied.
31. A method of audio processing for a terminal device, wherein, The terminal device has a reference microphone receiving a first signal and a plurality of other microphones each receiving a second signal. The method comprises: puncturing each second signal with respect to a time delay of the first signal, and superimposing each punctured second signal with the first signal to obtain a beamforming signal; dividing the first signal and the second signal into frequency bin parts in a frequency domain; processing each frequency bin part to enhance speech in the beamforming signal; wherein the processing each frequency bin part comprises: determining a difference of phase angles of the first signal and the second signal; determining a time delay of the first signal and the second signal; subtracting a product of an angular frequency of the frequency bin and the time delay from the difference of the phase angles to obtain a corrected phase difference of the first signal and the second signal at the frequency bin; determining whether the corrected phase difference satisfies a first predetermined condition; and suppressing the frequency bin part if the first predetermined condition is not satisfied.
Citation Information
Patent Citations
Speech enhancement method applied to dual-microphone array
CN105788607A
Method and device for extracting voice signal of desired sound source
CN110610718A