Speaker, terminal device, speaker plug-in, system on chip and related methods
By using differential adaptive processing circuits to perform differential adaptive processing on the microphone array in small speakers, the problems of high computing power and poor signal-to-noise ratio of small speakers are solved, and the signal-to-noise ratio and accuracy of speech recognition are improved.
Patent Information
- Application Number
- CN202011211019.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-11-03
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2040-11-03
AI Technical Summary
Small speakers have high computing power and poor signal-to-noise ratio during speech recognition, so they cannot effectively recognize user voice, especially when the incident angle is uncertain.
Differential adaptive processing circuit is used to perform differential adaptive processing on the microphone array, assuming the incident angle is 0° and 180°, a differential adaptive signal is generated, and voice recognition is performed in combination with the multi-beam output signal.
Without adding too much computing power burden, the signal-to-noise ratio and accuracy of speech recognition are significantly improved, especially when the incident angle of the sound source is 0° or 180°.
Smart Images

Figure CN114531631B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of electronics, and more particularly, to a speaker, a terminal device, a speaker plug-in, and related methods. Background Art
[0002] Automatic speech recognition (ASR) is a technology that converts human speech into text. It is widely used in areas such as robot conversations, smart homes, speakers, and voice-controlled applications (apps). For example, in the speaker field, users are generally required to speak a specific word, and the speaker will start working after recognizing the specific word. In the future, there may be a demand for speakers to start working when the user speaks any word. In recent years, microphone arrays have been widely used to recognize specific words when small speakers start working. They collect sound from multiple angles of incidence and, after comprehensive processing, activate the speaker if the wake-up word is recognized.
[0003] Because these speakers are compact and portable, they can't deploy microphones over a large area to improve sound source acquisition. Common small speakers typically deploy two microphones on the speaker's surface. These microphones beamform the sound signals collected by these two microphones to produce multiple beam output signals. A recognition unit then uses these beam output signals to determine whether the user has spoken a specific word, thereby deciding whether the speaker should activate. Most algorithms used in these speakers require high computing power and suffer from poor signal-to-noise ratios. Summary of the Invention
[0004] In view of this, the present disclosure aims to improve the signal-to-noise ratio of user voice recognition of the speaker without increasing the computing power burden.
[0005] To achieve this object, according to one aspect of the present disclosure, the present disclosure provides a speaker, comprising:
[0006] First microphone;
[0007] Second microphone;
[0008] a differential adaptive processing circuit, configured to perform differential adaptive processing on the sound signals collected by the first microphone and the second microphone when the incident angles of the antenna waveforms constructed for the first microphone and the second microphone are 0° and 180°, to obtain a first differential adaptive signal and a second differential adaptive signal;
[0009] The recognition unit is configured to perform speech recognition based on a multi-path beam output signal, the first differential adaptive signal, and the second differential adaptive signal, wherein the multi-path beam output signal is obtained from sound signals collected by the first microphone and the second microphone.
[0010] Optionally, the speaker further includes: a processor, configured to execute corresponding actions according to the voice recognition result.
[0011] Optionally, the corresponding action includes starting a speaker.
[0012] Optionally, the corresponding action includes activating different functions of the speaker according to different voice recognition results.
[0013] Optionally, the speaker further includes: a beam former, configured to obtain the multi-path beam output signal by performing beamforming on the sound signals collected by the first microphone and the second microphone.
[0014] Optionally, the recognition unit obtains a first probability from the multi-path beam output signal, the first differential adaptive signal and the second differential adaptive signal respectively, and determines a speech recognition result based on the first probability, wherein the first probability is the probability of recognizing a specific word.
[0015] Optionally, the recognition unit determines that the specific word is recognized as long as the first probability obtained from any one of the multi-path beam output signal, the first differential adaptive signal and the second differential adaptive signal is greater than a predetermined probability threshold.
[0016] Optionally, the recognition unit obtains a second probability from the multi-path beam output signal, the first differential adaptive signal and the second differential adaptive signal respectively, and determines a speech recognition result based on the first probability, wherein the first probability is the probability of recognizing a human voice.
[0017] Optionally, the recognition unit determines that a human voice is recognized as long as the second probability obtained from any one of the multi-path beam output signal, the first differential adaptive signal and the second differential adaptive signal is greater than a predetermined probability threshold.
[0018] Optionally, the differential adaptive processing circuit includes:
[0019] a first delay device, configured to delay the sound signal collected by the second microphone by a maximum time delay difference, where the maximum time delay difference is obtained by dividing the distance between the first microphone and the second microphone by the speed of sound;
[0020] a second delay device, configured to delay the sound signal collected by the first microphone by a maximum time delay difference;
[0021] a first subtractor, configured to subtract the output signal of the first delay device from the sound signal collected by the first microphone to obtain a first intermediate signal;
[0022] a second subtractor, configured to subtract the sound signal collected by the second microphone from the output signal of the second delay device to obtain a second intermediate signal;
[0023] a gain multiplier, configured to multiply the second intermediate signal by an adaptive gain;
[0024] The third subtractor is configured to subtract the output result of the gain multiplier from the first intermediate signal to obtain a first differential adaptive signal.
[0025] Optionally, the differential adaptive processing circuit operates in cycles, and the adaptive gain is equal to the adaptive gain of the previous cycle of the current cycle, plus the product of the second intermediate signal of the previous cycle, the first differential adaptive signal and a predetermined constant.
[0026] Optionally, the differential adaptive processing circuit includes:
[0027] a first delay device, configured to delay the sound signal collected by the first microphone by a maximum time delay difference, where the maximum time delay difference is obtained by dividing the distance between the first microphone and the second microphone by the speed of sound;
[0028] a second delayer, configured to delay the sound signal collected by the second microphone by a maximum time delay difference;
[0029] a first subtractor, configured to subtract the output signal of the first delay device from the sound signal collected by the second microphone to obtain a first intermediate signal;
[0030] a second subtractor, configured to subtract the sound signal collected by the first microphone from the output signal of the second delay device to obtain a second intermediate signal;
[0031] a gain multiplier, configured to multiply the second intermediate signal by an adaptive gain;
[0032] The third subtractor is configured to subtract the output result of the gain multiplier from the first intermediate signal to obtain a second differential adaptive signal.
[0033] Optionally, the differential adaptive processing circuit operates in cycles, and the adaptive gain is equal to the adaptive gain of the previous cycle of the current cycle, plus the product of the second intermediate signal of the previous cycle, the second differential adaptive signal and a predetermined constant.
[0034] According to one aspect of the present disclosure, a terminal device is provided, including:
[0035] First microphone;
[0036] Second microphone;
[0037] a differential adaptive processing circuit, configured to perform differential adaptive processing on the sound signals collected by the first microphone and the second microphone when the incident angles of the antenna waveforms constructed for the first microphone and the second microphone are 0° and 180°, to obtain a first differential adaptive signal and a second differential adaptive signal;
[0038] The recognition unit is configured to perform speech recognition based on a multi-path beam output signal, the first differential adaptive signal, and the second differential adaptive signal, wherein the multi-path beam output signal is obtained from sound signals collected by the first microphone and the second microphone.
[0039] Optionally, the recognition unit obtains a first probability from the multi-path beam output signal, the first differential adaptive signal and the second differential adaptive signal respectively, and determines a speech recognition result based on the first probability, wherein the first probability is the probability of recognizing a specific word.
[0040] Optionally, the recognition unit determines that the specific word is recognized as long as the first probability obtained from any one of the multi-path beam output signal, the first differential adaptive signal and the second differential adaptive signal is greater than a predetermined probability threshold.
[0041] Optionally, the recognition unit obtains a second probability from the multi-path beam output signal, the first differential adaptive signal and the second differential adaptive signal respectively, and determines a speech recognition result based on the first probability, wherein the first probability is the probability of recognizing a human voice.
[0042] Optionally, the recognition unit determines that a human voice is recognized as long as the second probability obtained from any one of the multi-path beam output signal, the first differential adaptive signal and the second differential adaptive signal is greater than a predetermined probability threshold.
[0043] According to one aspect of the present disclosure, there is provided a speaker, comprising:
[0044] Multiple microphones;
[0045] a differential adaptive processing circuit configured to perform differential adaptive processing on the sound signals collected by any two microphones of the plurality of microphones when the incident angles of the antenna waveforms constructed for the two microphones are 0° and 180°, thereby obtaining a differential adaptive signal pair;
[0046] The recognition unit is configured to perform speech recognition based on a multi-path beam output signal and the differential adaptive signal pair of any two microphones, wherein the multi-path beam output signal is obtained from sound signals collected by the multiple microphones.
[0047] Optionally, the speaker further includes: a beam former, configured to obtain the multi-path beam output signals by performing beamforming on the sound signals collected by the multiple microphones.
[0048] Optionally, the recognition unit obtains a first probability from the multi-path beam output signal and the differential adaptive signal pair of any two microphones, and determines a speech recognition result based on the first probability, wherein the first probability is a probability of recognizing a specific word.
[0049] Optionally, the recognition unit determines that the specific word is recognized as long as the first probability obtained from the multi-path beam output signal and any one signal of the differential adaptive signal pair of any two microphones is greater than a predetermined probability threshold.
[0050] Optionally, the recognition unit obtains a second probability from the multi-path beam output signal and the differential adaptive signal pair of any two microphones, and determines a speech recognition result based on the second probability, wherein the second probability is a probability of recognizing a human voice.
[0051] Optionally, the recognition unit determines that a human voice is recognized as long as the second probability obtained from the multi-path beam output signal and any one signal of the differential adaptive signal pair of any two microphones is greater than a predetermined probability threshold.
[0052] According to one aspect of the present disclosure, a speaker plug-in is provided for plugging into a speaker having a first microphone, a second microphone, and an identification unit, comprising:
[0053] a differential adaptive processing circuit, configured to perform differential adaptive processing on the sound signals collected by the first microphone and the second microphone when the incident angles of the antenna waveforms constructed for the first microphone and the second microphone are 0° and 180°, to obtain a first differential adaptive signal and a second differential adaptive signal;
[0054] The output interface is used to output the first differential adaptive signal and the second differential adaptive signal to the recognition unit, so that the recognition unit performs speech recognition based on the multi-path beam output signal, the first differential adaptive signal and the second differential adaptive signal.
[0055] According to one aspect of the present disclosure, a system-on-chip is provided that is connected to the outputs of a first microphone and a second microphone of a speaker. The system-on-chip includes: a differential adaptive processing circuit for performing differential adaptive processing on sound signals collected by the first microphone and the second microphone when the incident angles of the antenna waveforms constructed for the first microphone and the second microphone are 0° and 180°, thereby obtaining first and second differential adaptive signals for use in a speech recognition process.
[0056] Optionally, the system on chip also includes: a recognition unit for performing speech recognition based on a multi-path beam output signal, the first differential adaptive signal, and the second differential adaptive signal, wherein the multi-path beam output signal is obtained from the sound signal collected by the first microphone and the second microphone.
[0057] According to one aspect of the present disclosure, a speaker audio processing method is provided, wherein the speaker has a first microphone and a second microphone, and the speaker audio processing method includes:
[0058] When the incident angles of the antenna waveforms constructed for the first microphone and the second microphone are 0° and 180°, performing differential adaptive processing on the sound signals collected by the first microphone and the second microphone to obtain a first differential adaptive signal and a second differential adaptive signal;
[0059] Speech recognition is performed based on a multi-path beam output signal, the first differential adaptive signal, and the second differential adaptive signal, wherein the multi-path beam output signal is obtained from sound signals collected by the first microphone and the second microphone.
[0060] According to one aspect of the present disclosure, a terminal device audio processing method is provided, wherein the terminal device has a first microphone and a second microphone, and the terminal device audio processing method includes:
[0061] When the incident angles of the antenna waveforms constructed for the first microphone and the second microphone are 0° and 180°, performing differential adaptive processing on the sound signals collected by the first microphone and the second microphone to obtain a first differential adaptive signal and a second differential adaptive signal;
[0062] Speech recognition is performed based on a multi-path beam output signal, the first differential adaptive signal, and the second differential adaptive signal, wherein the multi-path beam output signal is obtained from sound signals collected by the first microphone and the second microphone.
[0063] According to one aspect of the present disclosure, a method for processing audio from a speaker is provided, wherein the speaker has multiple microphones, and the method includes:
[0064] For a combination of any two microphones from the plurality of microphones, when the incident angles of the antenna waveforms constructed for the two microphones are 0° and 180°, performing differential adaptive processing on the sound signals collected by the two microphones to obtain a differential adaptive signal pair;
[0065] Speech recognition is performed based on a multi-path beam output signal and the differential adaptive signal pair of any two microphones, wherein the multi-path beam output signal is obtained from the sound signals collected by the multiple microphones.
[0066] The disclosed embodiments cleverly introduce differential adaptive processing into microphone audio processing. It's generally believed that differential adaptive processing only ensures optimal sound source recognition when the antenna waveform of a dual-microphone array has an incident angle of 0° and 180°, and is less applicable at other incident angles. Small speakers can encounter a variety of incident angles, leading to a misconception in the field that fixed-direction noise reduction using differential adaptive processing is unsuitable for small speakers. The inventors of the present invention have overcome this misconception by discovering that, in addition to the normal multi-beam output signals generated by the speaker, differential adaptive processing can be performed additionally when the antenna waveform of the dual-microphone array has an incident angle of 0° and 180°. This generates first and second differential adaptive signals, which are then provided to the recognition unit along with the multi-beam output signals. If the actual sound source's position precisely causes the incident angle on the dual-microphone array to be 0° or 180°, one of the two signals is selected. These two signals, generated using differential adaptive processing, have a very high signal-to-noise ratio, significantly improving speech recognition. If this is not the case, at most recognition is performed within the multi-beam output signal, which at least does not degrade the recognition performance. The disclosed embodiment only adds two-way differential adaptive processing, which does not significantly increase the computing power requirements. Therefore, without significantly increasing the computing power burden, the signal-to-noise ratio and performance of speech recognition are greatly improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] The above and other objects, features and advantages of the present disclosure will become more apparent through description of the embodiments of the present disclosure with reference to the following drawings, in which:
[0068] Figure 1A is an appearance diagram of a dual-microphone array speaker according to an embodiment of the present disclosure;
[0069] Figure 1B is an appearance diagram of a multi-microphone array speaker according to an embodiment of the present disclosure;
[0070] Figure 1C is a schematic external view of a terminal device with a dual microphone array according to an embodiment of the present disclosure;
[0071] Figure 2A is a structural diagram of a dual-microphone array speaker according to an embodiment of the present disclosure;
[0072] Figure 2B is a structural diagram of a multi-microphone array speaker according to an embodiment of the present disclosure;
[0073] Figure 2C is a structural diagram of a terminal device with a dual microphone array according to an embodiment of the present disclosure;
[0074] Figure 3A is a circuit diagram of a portion of a differential adaptive processing circuit for obtaining a first differential adaptive signal according to an embodiment of the present disclosure;
[0075] Figure 3B is a circuit diagram of a portion of a differential adaptive processing circuit for obtaining a second differential adaptive signal according to an embodiment of the present disclosure;
[0076] Figure 4 This is a comparison table of the additional overhead in terms of memory and main frequency caused by one embodiment of the present disclosure and the prior art;
[0077] Figure 5A -D shows examples of four different arrangement scenarios of sound sources and interference sources to which the embodiments of the present disclosure are applicable;
[0078] Figure 6 Shown Figure 5A -Comparison of effect data under four scenarios of D;
[0079] Figure 7 1 shows the arrangement of the speaker insert 128 in the speaker 100 according to one embodiment of the present disclosure;
[0080] Figure 8 A flow chart of a dual-microphone array speaker audio processing method according to an embodiment of the present disclosure is shown;
[0081] Figure 9 A flow chart of an audio processing method for a dual-microphone array terminal device according to an embodiment of the present disclosure is shown;
[0082] Figure 10 A flow chart of a multi-microphone array speaker audio processing method according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0083] The present disclosure is described below based on examples, but the present disclosure is not limited to these examples. Certain specific details are described in detail in the detailed description of the present disclosure below. Those skilled in the art will appreciate that the present disclosure is fully understood without these details. To avoid obscuring the essence of the present disclosure, well-known methods, processes, and procedures have not been described in detail. The accompanying drawings are not necessarily drawn to scale.
[0084] The following terms are used in this document.
[0085] Speaker: A box that plays sound. It usually has speaker holes, through which the sound is played to the outside.
[0086] Microphone: An energy conversion device that converts sound signals into electrical signals.
[0087] Wake-up: When the user recognizes that they have spoken a predetermined word or any word, the speaker enters the working state or starts a certain function.
[0088] Wake-up word: A predefined word spoken by the user to activate the speaker or start a function. For example, the wake-up word could be "hello, XX" or "Please turn on the speaker." When the speaker recognizes the user speaking these words, it starts operating or starts a function.
[0089] Microphone array: An array of multiple microphones deployed on a speaker, used to separately receive the user's voice signals (which may contain the wake-up word spoken by the user) for subsequent processing. This allows the system to identify whether the user has spoken the wake-up word or any word, and decide whether to wake up the speaker.
[0090] Beamforming: A concept originating from adaptive antennas. Signal processing at the receiving end can form the desired ideal signal by weighted synthesis of the various signals received by multiple antenna elements. From the perspective of the antenna pattern, this is equivalent to forming a beam in a specified direction. Forming beams in multiple specified directions forms beamforming in multiple directions. Microphone arrays use the concept of type antennas, that is, by weighted synthesis of the various sound signals received by each microphone in the microphone array to form the desired ideal signal. From the perspective of the pattern, this is equivalent to forming a beam in the specified sound incident direction. For example, for a dual-microphone array, beams in the sound incident direction are formed at 30°, 90°, and 150° relative to the line connecting the two microphones.
[0091] Speaker plug-in: An attachment device that is inserted into a general speaker to give the speaker some special functions.
[0092] Dual-microphone array loudspeaker embodiment
[0093] Figure 1AThe figure shows the appearance of a dual-microphone array speaker. A first microphone 111 and a second microphone 112 are provided on the outer surface of the speaker 100 to form a dual-microphone array. The first microphone 111 and the second microphone 112 are used to receive sound signals respectively, and then convert them into respective electrical signals so that the speaker can perform subsequent processing on the respective electrical signals. In addition, the surface of the speaker may have a speaker hole array (not shown). The sound played after the speaker 100 is activated is output through the speaker hole array. Note that the embodiment of the present disclosure is only used for waking up the speaker 100. The sound played after the speaker 100 is activated may come from a built-in disk or a sound file transmitted by Bluetooth, and is not the sound received by the first microphone 111 or the second microphone 112.
[0094] like Figure 2A As shown, the speaker 100 may also include a beamformer 130, a differential adaptive processing circuit 120, an identification unit 140, and a processor 150. The beamformer 130 beamforms the sound signals collected by the first microphone 111 and the second microphone 112 to form a multi-path beam output signal 131. The sound signals collected by the first microphone 111 and the second microphone 112 include the sound signals emitted by the sound source 10 and reaching the first microphone 111 and the second microphone 112, as well as noise signals. The beamformer 130 weights and synthesizes the sound signals received by the first microphone 111 and the second microphone 112 to form the desired ideal signal. From a directional pattern perspective, this is equivalent to forming a beam in the specified sound incident direction.
[0095] The recognition unit 140 performs speech recognition based on the multi-beamform output signal, the first differential adaptive signal, and the second differential adaptive signal. Speech recognition here may refer to recognizing a specific word (used to identify a situation where a user speaks a specific word to activate the speaker) or any word spoken by the user (used to identify a situation where any word spoken by the user can activate the speaker).
[0096] In the case of specific word recognition, the recognition unit 140 obtains a first probability from the multi-path beam output signal, the first differential adaptive signal, and the second differential adaptive signal. The first probability is the probability of recognizing the specific word. The recognition unit 140 then determines the speech recognition result based on the first probability.
[0097] The portion of recognition unit 140 that obtains the first probability can employ a speech recognition model. This is a machine learning model that can be pre-trained in the following manner: a set of sound signal samples is constructed in advance, where each sound signal sample is known to correspond to a predetermined word label spoken by the user; each sound signal sample in the set is input into the machine learning model, and the model determines the corresponding predetermined word. If the determined predetermined word matches the predetermined word label, the training is considered successful; if the success rate of the machine learning model in the sound signal sample set exceeds a predetermined rate (e.g., 95%), the machine learning model is considered successfully trained; otherwise, the coefficients in the machine learning model are adjusted until the success rate of the machine learning model in the sound signal sample set exceeds the predetermined rate. In this way, by inputting any sound signal into the trained machine learning model, the probability of recognizing a specific word can be obtained, i.e., the first probability.
[0098] In one embodiment, as long as the first probability obtained from any one of the multi-path beam output signal, the first differential adaptive signal, and the second differential adaptive signal is greater than a predetermined probability threshold, the recognition unit 140 determines that the specific word is recognized.
[0099] In the case of non-specific word recognition, the recognition unit 140 obtains a second probability from the multi-path beam output signal, the first differential adaptive signal, and the second differential adaptive signal. The second probability is the probability of recognizing a human voice. The recognition unit 140 then determines a speech recognition result based on the second probability.
[0100] The portion of recognition unit 140 that derives the second probability can employ a voice recognition model. This machine-learning model can be pre-trained in the following manner: a sound sample set is constructed, some of which contain human speech, while others do not. Each sound sample in the sound sample set is input into the voice recognition model, which determines whether it contains human speech and compares it with previously known human speech detection results. If a match is found, the model is considered successful. If the success rate determined by the voice recognition model in the sound sample set exceeds a predetermined rate (e.g., 95%), the voice recognition model is considered successfully trained. Otherwise, the coefficients in the voice recognition model are adjusted until the success rate determined by the voice recognition model in the sound sample set exceeds the predetermined rate. In this way, any sound input into the trained model can yield the probability of recognizing a human voice, i.e., the second probability.
[0101] In one embodiment, the recognition unit determines that a human voice is recognized as long as the second probability obtained from any one of the multi-path beamforming output signal, the first differential adaptive signal, and the second differential adaptive signal is greater than a predetermined probability threshold.
[0102] The advantage of this embodiment is that it is very friendly to some scenarios that require energy saving. For example, if someone is identified in the hall, the speaker will play refreshing background music for them, and if no one is there, it will not work. In such a scenario, applying the above embodiment is very beneficial.
[0103] The differential adaptive processing circuit 120 is used to perform differential adaptive processing on the sound signals collected by the first microphone 111 and the second microphone 112, assuming that the sound source 10 is located on both sides of the extension line of the connection line between the first microphone 111 and the second microphone 112 (end-fire and opposite end-fire directions), to obtain a first differential adaptive signal 121 and a second differential adaptive signal 122, which are provided together with the beam output signals 131 formed by the beam former 130 to the recognition unit 140 to determine whether to wake up the speaker.
[0104] Differential adaptive processing is an algorithm that uses signal differentiation to eliminate noise signals in a specific direction. This ensures that the sound signal at certain incident angles is not attenuated, while the noise signal is nearly eliminated. However, at opposite incident angles, the noise signal is attenuated to zero. It is commonly used in hearing aids. By forcing the sound incident angle to the microphone array to a fixed direction, the sound in that direction is heard most clearly. That is, at this incident angle, the sound source signal is not attenuated, while the noise signal is nearly eliminated. Differential adaptive processing only ensures that the sound source signal is not attenuated at incident angles of 0° or 180° in the antenna waveform of a dual-microphone array, while attenuating the noise signal at opposite incident angles to zero. Therefore, it is extremely suitable for sound sources with incident angles of 0° or 180° in a dual-microphone array, significantly improving the received signal-to-noise ratio at these angles. However, in typical small speaker applications, the incident angle of the sound source 10 is uncertain. Therefore, there is a misconception in the field that differential adaptive processing for noise reduction in a fixed direction is not suitable for small speakers. The inventors of the present invention have overcome this prejudice by discovering that two additional signals can be added to the beamformed output signals 131 formed by the beamformer 130 through differential adaptive processing. Assuming the sound source is located at an angle of incidence of 0° or 180° in the antenna waveform of the dual microphones, differential adaptive processing is performed on the sound signals collected by the dual microphones to generate a first differential adaptive signal 121 and a second differential adaptive signal 122. These signals, along with the multiple beamformed output signals 131 obtained by the beamformer 130, serve as the basis for speech recognition by the recognition unit 140. For the recognition unit 140, as long as the probability of recognizing a specific word or any word from one signal exceeds a predetermined threshold, the speaker is activated. Having two more signals to choose from increases the probability of the speaker being activated. If the actual sound source 10 happens to have an angle of incidence of 0° or 180° on the dual microphones, the probability of one of the two signals is likely to exceed the predetermined probability threshold. Because these two paths are adaptively derived through differential processing, the sound source signal is barely attenuated, while the noise signal is nearly attenuated to zero, resulting in a very high signal-to-noise ratio (SNR), significantly improving speech recognition. If the antenna waveform's incident angles are not 0° or 180°, the situation remains the same as in the prior art, at least without compromising the SNR.
[0105] In one embodiment, Figure 3A As shown, the differential adaptive processing circuit 120 includes a first delay 210 , a second delay 202 , a first subtractor 203 , a second subtractor 204 , a gain multiplier 205 , and a third subtractor 206 .
[0106] Assume that f(t) and b(t) are the sound signals collected by the first microphone 111 and the second microphone 112 respectively. Assume that the maximum delay difference between the two microphones is , which can be obtained by dividing the distance between the first microphone 111 and the second microphone 112 by the speed of sound. A portion of the sound signal collected by each microphone is the result of the sound source signal propagating to the microphone, and the other portion is a noise signal.
[0107] Assume that the sound signal f(t) collected by the first microphone 111 consists of two parts:
[0108] Formula 1
[0109] Where s(t) represents the signal when the sound source signal propagates to the first microphone 111, and v(t) represents the noise signal at the first microphone 111. The sound signal b(t) collected by the second microphone 112 can be expressed as:
[0110] Formula 2
[0111] Since the maximum time delay difference between the two microphones is Since it is assumed that the incident angle of the sound source 10 relative to the antenna waveform of the first microphone 111 and the second microphone 112 is 0°, the signal when the sound source signal propagates to the second microphone 112 is , and the signal when the noise signal propagates to the second microphone 112 is , where i is the time delay of the noise signal propagating to the second microphone 112.
[0112] The first delay device 201 delays the sound signal collected by the second microphone 112 Maximum delay difference ,get The second delay 202 delays the sound signal f(t) collected by the first microphone 111 by a maximum delay difference of ,get .
[0113] The first subtractor 203 subtracts the output signal of the first delay 201 from the sound signal f(t) collected by the first microphone 111. , get the first intermediate signal , as follows:
[0114] Formula 3
[0115] The second subtractor 204 converts the output signal of the second delay 202 into Subtract the sound signal b(t) collected by the second microphone 112 to obtain the second intermediate signal , as follows:
[0116] Formula 4
[0117] The gain multiplier 205 converts the second intermediate signal Multiply by the adaptive gain , the output is , as follows:
[0118] Formula 5
[0119] The third subtractor 206 converts the first intermediate signal Subtract the output of the gain multiplier from , and obtain the first differential adaptive signal 121, namely , as follows:
[0120] Formula 6
[0121] In the above process, the differential adaptive processing circuit 120 operates in clock cycles. t represents the current clock cycle, t-1 represents the cycle before the current clock cycle, and t+1 represents the cycle after the current clock cycle. The adaptive gain W(t) is equal to the adaptive gain of the cycle before the current cycle. , plus the second intermediate signal of the previous cycle , the first differential adaptive signal With a predetermined constant The product of is as follows:
[0122] Formula 7
[0123] In other words, the adaptive gain W(t+1) of the next cycle after the current cycle is equal to the adaptive gain of the previous cycle , plus the second intermediate signal of the previous cycle , the first differential adaptive signal With a predetermined constant The product of is as follows: Formula 8
[0124] That is, the adaptive gain used in each cycle It is not fixed, but depends on the adaptive gain in the previous cycle. This is called adaptive because it adjusts the values of certain parameters. Through this adaptive differentiation, as long as the incident angle of the sound source is fixed at 0°, the sound source signal in that direction will not be attenuated, and the noise signal will almost disappear, thus achieving the effect of improving the signal-to-noise ratio.
[0125] The differential adaptive processing circuit 120 also has the effect of improving the signal-to-noise ratio when the incident angle of the sound source is 180°. Figure 3B As shown, the differential adaptive processing circuit 120 includes components similar to Figure 3A The difference is that f(t) and b(t) are swapped, with b(t) representing the sound signal collected by the first microphone 111 and f(t) representing the sound signal collected by the second microphone 112.
[0126] At this time, the sound signal b(t) collected by the first microphone 111 consists of two parts:
[0127] Formula 9
[0128] Where s(t) represents the signal when the sound source signal propagates to the first microphone 111, and v(t) represents the noise signal at the first microphone 111. The sound signal f(t) collected by the second microphone 112 can be expressed as:
[0129] Formula 10
[0130] Since the maximum time delay difference between the two microphones is Since it is assumed that the incident angle of the sound source 10 relative to the first microphone 111 and the second microphone 112 is 180° instead of 0°, the signal when the sound source signal propagates to the second microphone 112 is , and the signal when the noise signal propagates to the second microphone 112 is , where i is the time delay of the noise signal propagating to the second microphone 112.
[0131] The first delay device 201 delays the sound signal collected by the second microphone 112 Maximum delay difference ,get The second delay 202 delays the sound signal b(t) collected by the first microphone 111 by a maximum delay difference of ,get .
[0132] The first subtractor 203 subtracts the output signal of the first delay 201 from the sound signal b(t) collected by the first microphone 111. , get the first intermediate signal , as follows:
[0133] Formula 11
[0134] The second subtractor 204 converts the output signal of the second delay 202 into Subtract the sound signal f(t) collected by the second microphone 112 to obtain the second intermediate signal , as follows:
[0135] Formula 12
[0136] The gain multiplier 205 converts the second intermediate signal Multiply by the adaptive gain , the output is , as follows:
[0137] Formula 13
[0138] The third subtractor 206 converts the first intermediate signal Subtract the output of the gain multiplier from , and obtain the second differential adaptive signal 122, namely , as follows:
[0139] Formula 14
[0140] In the above process, the differential adaptive processing circuit 120 operates in clock cycles. t represents the current clock cycle, t-1 represents the cycle before the current clock cycle, and t+1 represents the cycle after the current clock cycle. The adaptive gain W(t) is equal to the adaptive gain of the cycle before the current cycle. , plus the second intermediate signal of the previous cycle , the second differential adaptive signal With a predetermined constant The product of is as follows:
[0141] Formula 15
[0142] In other words, the adaptive gain W(t+1) of the next cycle after the current cycle is equal to the adaptive gain of the previous cycle , plus the second intermediate signal of the previous cycle , the second differential adaptive signal With a predetermined constant The product of is as follows: Formula 16
[0143] The recognition unit 140 performs speech recognition based on the multi-beamform output signal 131 , the first differential adaptive signal 121 , and the second differential adaptive signal 122 .
[0144] In one embodiment, the processor 150 performs a corresponding action based on the voice recognition result. That is, when it is recognized that the voice spoken by the user is a specific word or the user is recognized to have made a voice, the corresponding action is performed. Otherwise, the corresponding action is not performed. In one embodiment, the corresponding action includes starting the speaker. That is, when it is recognized that the voice spoken by the user is a specific word or the user is recognized to have made a voice, the speaker is activated and starts playing the sound to be played, such as music. In another embodiment, the corresponding action includes activating different functions of the speaker based on different voice recognition results. For example, when the user says "play music", music is played for the user; when the user says "louder", the volume of the playback is turned up; when the user says "shut down", the speaker is turned off.
[0145] In the latter case, the specific word may include multiple specific words, each corresponding to a different function of the speaker, such as "louder" corresponding to increasing the volume, and "play music" corresponding to playing music for the user. The recognition unit 140 outputs the probability of recognizing each specific word from the multi-beam output signal 131, the first differential adaptive signal 121, and the second differential adaptive signal 122. As long as the probability of recognizing a specific word from any one of the multi-beam output signal 131, the first differential adaptive signal 121, and the second differential adaptive signal 122 exceeds a probability threshold corresponding to the specific word, the speaker's function corresponding to the specific word is activated.
[0146] The disclosed embodiment utilizes the low computing power and high performance characteristics of the differential adaptive algorithm, and uses it to provide two more auxiliary outputs for the currently mature speaker algorithm, thereby improving the voice recognition effect of the speaker in special scenarios (end-fire or end-fire reverse direction).
[0147] Figure 4 The figure shows the main frequency and memory overhead of existing commercial algorithms, as well as the additional main frequency and memory overhead incurred by the disclosed embodiments. Clearly, the additional main frequency and memory overhead incurred by the disclosed embodiments is minimal compared to the four existing mature commercial algorithms. Therefore, the disclosed embodiments can assist these mature commercial algorithms.
[0148] Figure 5A -D shows examples of four different arrangement scenarios of sound sources and interference sources to which the embodiments of the present disclosure are applicable.
[0149] exist Figure 5A In the figure, the sound signal of the sound source 10 is incident along the 0° incident angle of the dual-microphone antenna waveform. There is only one interference source 11, which is incident at a 180° incident angle along the dual-microphone antenna waveform. Figure 6As shown, the interference source 11 can be conversation / news / music / cafe noise. In the case of four interference sources, specific words are identified from the sound signal received by microphone 1 respectively. When sound source 10 says the specific word 200 times, the number of times the correct specific word is identified from the sound signal received by microphone 1 is 13 / 4 / 19 / 14 respectively. That is, when the interference source 11 is a conversation, 13 out of 200 times are correctly identified; when the interference source 11 is news, 4 out of 200 times are correctly identified; when the interference source 11 is music, 19 out of 200 times are correctly identified; and when the interference source 11 is cafe noise, 13 out of 200 times are correctly identified. When sound source 10 says the specific word 200 times, the number of times the correct specific word is identified from the sound signal received by microphone 2 is 16 / 5 / 18 / 13 respectively. When the sound source 10 said a specific word 200 times, the number of times the wake-up device using commercial algorithm 1 correctly recognized the specific word was 137 / 116 / 39 / 154, the number of times the wake-up device using commercial algorithm 2 correctly recognized the specific word was 158 / 141 / 170 / 194, the number of times the wake-up device using commercial algorithm 3 correctly recognized the specific word was 0 / 155 / 0 / 1, the number of times the wake-up device using commercial algorithm 4 correctly recognized the specific word was 200 / 191 / 200 / 200, and the number of times the wake-up device using the embodiment of the present disclosure correctly recognized the specific word was 200 / 199 / 200 / 200. As can be seen from this, the recognition success rate of the embodiment of the present disclosure is the highest.
[0150] exist Figure 5B In the figure, the sound signal of the sound source 10 is incident along the 0° incident angle of the dual-microphone antenna waveform. There are two interference sources 11, and their interference noises are incident along the 120° and 180° incident angles of the dual-microphone antenna waveform respectively. Figure 6 As shown, the two interference sources 11 are conversation and cafe noise. When sound source 10 says a specific word 200 times, the correct specific word is recognized 1 time from the sound signal received by microphone 1, 0 times from the sound signal received by microphone 2, 132 times from the recognition unit of commercial algorithm 1, 158 times from the recognition unit of commercial algorithm 2, 0 times from the recognition unit of commercial algorithm 3, 65 times from the recognition unit of commercial algorithm 4, and 175 times from the recognition unit of the embodiment of the present disclosure. As can be seen from this, the recognition success rate of the embodiment of the present disclosure is the highest.
[0151] exist Figure 5CIn the figure, the sound signal of the sound source 10 is incident along the 0° incident angle of the dual-microphone antenna waveform. There are three interference sources 11, which are incident along the 120°, 180° and -150° incident angles of the dual-microphone antenna waveform. Figure 6 As shown, the three interference sources 11 are conversation, music, and cafe noise. When sound source 10 says a specific word 200 times, the number of times the correct specific word is recognized from the sound signal received by microphone 1 is 0, the number of times the correct specific word is recognized from the sound signal received by microphone 2 is 0, the number of times the correct specific word is recognized from the recognition unit of commercial algorithm 1 is 16, the number of times the correct specific word is recognized from the recognition unit of commercial algorithm 2 is 132, the number of times the correct specific word is recognized from the recognition unit of commercial algorithm 3 is 0, the number of times the correct specific word is recognized from the recognition unit of commercial algorithm 4 is 123, and the number of times the correct specific word is recognized from the recognition unit of the embodiment of the present disclosure is 161. As can be seen from this, the recognition success rate of the embodiment of the present disclosure is the highest.
[0152] exist Figure 5D In the figure, the sound signal of the sound source 10 is incident along the 0° incident angle of the dual-microphone antenna waveform. There are four interference sources 11, which are incident along the 120°, 180°, -150° and -120° incident angles of the dual-microphone antenna waveform. Figure 6 As shown, the four interference sources 11 are conversation, music, news, and cafe noise. When sound source 10 says a specific word 200 times, the number of times the correct specific word is recognized from the sound signal received by microphone 1 is 0, the number of times the correct specific word is recognized from the sound signal received by microphone 2 is 0, the number of times the correct specific word is recognized from the recognition unit of commercial algorithm 1 is 14, the number of times the correct specific word is recognized from the recognition unit of commercial algorithm 2 is 102, the number of times the correct specific word is recognized from the recognition unit of commercial algorithm 3 is 0, the number of times the correct specific word is recognized from the recognition unit of commercial algorithm 4 is 28, and the number of times the correct specific word is recognized from the recognition unit of the embodiment of the present disclosure is 123. As can be seen from this, the recognition success rate of the embodiment of the present disclosure is the highest.
[0153] from Figure 6 It can be seen that the embodiment of the present disclosure has extremely high noise reduction capabilities when the incident angle of the sound source relative to the antenna waveform of the dual-microphone array is 0° and 180°, and the interference source is located in the opposite direction plane, and is less affected by the amount of interference. The recognition success rate of various existing commercial solutions will drop rapidly as the amount of noise increases, while the embodiment of the present disclosure does not drop significantly. Although the recognition success rate of commercial algorithm 2 is also relatively good, Figure 4Compared to the main frequency and memory requirements of commercial algorithm 2, the disclosed embodiment surpasses its recognition performance with far lower memory and main frequency requirements. Therefore, without significantly increasing the computing power burden, the disclosed embodiment significantly improves the signal-to-noise ratio of speaker voice recognition, which is a good supplement to existing mature commercial algorithms.
[0154] In addition to being applicable to dual-microphone speakers, the embodiments of the present disclosure can also be applied to multi-microphone speakers, such as Figure 1B It should be understood that the above number of microphones is only an example, and the embodiments of the present disclosure can also be applied to speakers with other numbers of microphones.
[0155] Figure 2B This is the structural diagram of the multi-microphone speaker. Figure 2A The same parts will not be repeated. Figure 2B There are multiple microphones 110, and the functions of the corresponding beamformer 130, differential adaptive processing circuit 120 and identification unit 140 are changed accordingly. The beamformer 130 does not beamform the sound signals collected by two microphones to form a multi-path beam output signal, but beamforms the sound signals collected by multiple microphones 110 to form a multi-path beam output signal, but the beamforming method is the same. The differential adaptive processing circuit 120 performs the same beamforming on any combination of two microphones 110 in accordance with the Figure 2A In a similar manner to the embodiment, differential adaptive processing is performed to obtain a pair of differential adaptive signals. Assuming that the number of microphones is M, we can actually obtain The identification unit 140 is based on the multi-path beam output signal, the differential adaptive signal pair for any two microphones of the plurality of microphones (a total of Yes), perform speech recognition.
[0156] As for the structure of the differential adaptive processing circuit 120, it is also the same as that in the dual-microphone speaker, that is, Figure 3A -B, which is just for the combination of any two microphones among the plurality of microphones, one of the microphones is used as the first microphone 111 and the other microphone is used as the second microphone 112. A combination of microphones to be processed This is just the first time, so I won’t go into details.
[0157] As for the recognition unit 140, it can be the same as the recognition unit 140 in the above embodiment, except that the differential adaptive signal pair of any two microphones is Yes, whereas in the above embodiment there is only one pair.
[0158] The above embodiment extends the application of differential adaptive processing in a dual-microphone speaker to a multi-microphone speaker, thereby greatly improving the signal-to-noise ratio of the multi-microphone speaker without increasing the computing power burden of the multi-microphone speaker.
[0159] In addition to being applied to speakers, the embodiments of the present disclosure can also be applied to other terminal devices 101 that need to collect the sound of the sound source and perform subsequent processing, such as intelligent conversation robots and video conferencing systems. Like speakers, intelligent conversation robots may require users to say specific words or any words to wake them up in order to work properly. In a video conferencing system, participants or hosts may also need to say specific words such as "meeting" to start the video conferencing system. Since they all need to collect the sound signal of a person (sound source) to determine whether a specific word has been said, they may both need a first microphone 111 and a second microphone 112 to collect sound signals (including the signal attenuated when the sound signal of the sound source propagates to the first microphone 111 and the second microphone 112). Figure 2C As shown, it also needs to include a beamformer 130, a differential adaptive processing circuit 120 and a recognition unit 140. The beamformer 130 performs beamforming on the sound signals collected by the first microphone 111 and the second microphone 112 to form a multi-path beam output signal 131. The differential adaptive processing circuit 120 performs differential adaptive processing on the sound signals collected by the first microphone 111 and the second microphone 112 when the incident angles of the antenna waveform diagrams constructed for the first microphone 111 and the second microphone 112 are 0° and 180°, respectively, to obtain a first differential adaptive signal 121 and a second differential adaptive signal 122. The recognition unit 140 performs speech recognition based on the multi-path beam output signal 131, the first differential adaptive signal 121 and the second differential adaptive signal 122. The first microphone 111, the second microphone 112, the beamformer 130, the differential adaptive processing circuit 120 and the recognition unit 140 are connected to the multi-path beam output signal 131. Figure 2A The embodiments are generally the same, and their structures and working processes can refer to the above Figure 2A The embodiments are not described in detail.
[0160] The terminal device 101 may also include a processor 150 for performing corresponding actions based on the speech recognition results. When the terminal device 101 is an intelligent dialogue robot, the corresponding action may include responding to the recognized voice. In this case, the speech recognition result of the recognition unit 140 may not be used to start the intelligent dialogue robot, but to respond to user questions after startup. When the terminal device 101 is a video conferencing system, the corresponding action may include triggering switches for different operations based on different speech recognition results. For example, it recognizes that the user says "connect to branch X venue" to connect to branch X venue, and in response to the user saying "I can't hear the speaker clearly", the volume of the speaker in the venue where the speaker is located is turned up.
[0161] This embodiment extends the scope of use of differential adaptive processing to terminal devices, not just speakers, thus broadening the application of the embodiments of the present disclosure.
[0162] In addition, if Figure 7 As shown, the embodiment of the present disclosure also proposes a speaker plug-in 128, which is inserted into the universal speaker 100, so that the universal speaker 100 has the function of improving the signal-to-noise ratio of the speaker without increasing the computing power burden proposed by the embodiment of the present disclosure. The universal speaker 100 includes a first microphone 111, a second microphone 112, a beamformer 130, and an identification unit 140. The functions of the first microphone 111, the second microphone 112, the beamformer 130, and the identification unit 140 are described in detail in the following examples. Figure 2A The speaker plug-in 128 of the embodiment of the present disclosure includes a differential adaptive processing circuit 120, the structure and function of which refer to Figure 2A As described in the embodiment. The speaker plug-in 128 further includes an output interface 129, which is equivalent to an interface with the recognition unit 140. After the speaker plug-in 128 is inserted into the universal speaker 100, it is used to output the first differential adaptive signal 121 and the second differential adaptive signal 122 to the recognition unit 140, so that the recognition unit 140 performs speech recognition based on the multi-path beam output signal 131, the first differential adaptive signal 121 and the second differential adaptive signal 122.
[0163] This embodiment is applicable to the case where a plug-in unit having the differential adaptive processing circuit 120 is produced separately, rather than the case where an entire speaker box having the differential adaptive processing circuit 120 is produced.
[0164] In addition, embodiments of the present disclosure further provide a system-on-chip (SoC), similar to the aforementioned speaker plug-in 128, which is specifically configured to perform differential adaptive processing on the sound signals collected by the first microphone 111 and the second microphone 112. When the SoC is installed in an IoT device or a standard terminal device, the IoT device or standard terminal device can utilize not only the multi-path beamforming output signal 131 obtained from the sound signals collected by the first microphone 111 and the second microphone 112 during speech recognition, but also the first differential adaptive signal 121 and the second differential adaptive signal 122 obtained by the SoC, thereby improving the signal-to-noise ratio of speech recognition in the IoT device or standard terminal device.
[0165] The SoC itself can incorporate the aforementioned recognition unit 140 for performing speech recognition based on the multi-path beamforming output signal 131, the first differential adaptive signal 121, and the second differential adaptive signal 122. This recognition unit 140 can also be external to the SoC. When the SoC includes the recognition unit 140, it becomes a speech recognition chip, specifically responsible for speech recognition in IoT devices or general terminal devices.
[0166] like Figure 8 As shown, according to one embodiment of the present disclosure, an audio processing method for a dual-microphone speaker 100 is also provided, wherein the speaker 100 has a first microphone 111 and a second microphone 112, and the audio processing method includes:
[0167] Step 310: When the incident angles of the antenna waveforms constructed for the first microphone 111 and the second microphone 112 are 0° and 180°, perform differential adaptive processing on the sound signals collected by the first microphone 111 and the second microphone 112 to obtain a first differential adaptive signal 121 and a second differential adaptive signal 122.
[0168] Step 320 : Perform speech recognition based on the multi-beam output signal 131 , the first differential adaptive signal 121 , and the second differential adaptive signal 122 , wherein the multi-beam output signal is obtained from the sound signals collected by the first microphone 111 and the second microphone 112 .
[0169] The implementation details of this method are as described above. Figure 2A The embodiments have been described in Figure 2A The embodiments of the present invention are not described in detail.
[0170] like Figure 9 As shown, according to one embodiment of the present disclosure, an audio processing method of a terminal device 101 is further provided, wherein the terminal device 101 has a first microphone 111 and a second microphone 112, and the audio processing method of the terminal device 101 includes:
[0171] Step 410: When the incident angles of the antenna waveforms constructed for the first microphone 111 and the second microphone 112 are 0° and 180°, perform differential adaptive processing on the sound signals collected by the first microphone 111 and the second microphone 112 to obtain a first differential adaptive signal 121 and a second differential adaptive signal 122.
[0172] Step 420 : Perform speech recognition based on the multi-beam output signal 131 , the first differential adaptive signal 121 , and the second differential adaptive signal 122 , wherein the multi-beam output signal 131 is obtained from the sound signals collected by the first microphone 111 and the second microphone 112 .
[0173] The implementation details of this method are as described above. Figure 2C The embodiments have been described in Figure 2C The embodiments of the present invention are not described in detail.
[0174] like Figure 10 As shown, according to one embodiment of the present disclosure, there is also provided an audio processing method for a multi-microphone speaker 100, wherein the speaker 100 has a plurality of microphones 110, and the speaker audio processing method includes:
[0175] Step 510: For a combination of any two microphones 110 from the plurality of microphones 110, when the incident angles of the antenna waveforms constructed for the two microphones 110 are 0° and 180°, perform differential adaptive processing on the sound signals collected by the two microphones 110 to obtain differential adaptive signal pairs 121 and 122.
[0176] Step 520: Perform speech recognition based on the multi-path beam output signal 131 and the differential adaptive signal pairs 121 and 122 of any two microphones, wherein the multi-path beam output signal 131 is obtained from the sound signals collected by the multiple microphones.
[0177] The implementation details of this method are as described above. Figure 2B The embodiments have been described in Figure 2B The embodiments of the present invention are not described in detail.
[0178] The commercial value of this disclosure
[0179] Compared with the four existing mature commercial algorithms, the additional main frequency and memory size overhead caused by the embodiment of the present disclosure is only slightly increased on the basis of them, but the recognition success rate and signal-to-noise ratio are greatly improved. It is estimated that the sales of speakers or terminal devices will increase by 20%, and it has good market prospects.
[0180] It should be understood that the various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments.
[0181] It should be understood that the foregoing description of this specification is based on specific embodiments. Other embodiments are within the scope of the claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0182] It should be understood that an element described herein in the singular or shown in the drawings as only one does not limit the number of the element to one. In addition, modules or elements described or shown herein as separate may be combined into a single module or element, and modules or elements described or shown herein as single may be split into multiple modules or elements.
[0183] It should also be understood that the terms and expressions used herein are for descriptive purposes only, and the one or more embodiments of this specification should not be limited to these terms and expressions. The use of these terms and expressions does not mean to exclude any equivalent features of the illustrations and descriptions (or portions thereof), and it should be recognized that various modifications that may exist should also be included in the scope of the claims. Other modifications, variations, and substitutions may also exist. Accordingly, the claims should be deemed to cover all such equivalents.
Claims
1. A speaker, comprising: First microphone; Second microphone; a differential adaptive processing circuit, configured to perform differential adaptive processing on the sound signals collected by the first microphone and the second microphone when the incident angles of the antenna waveforms constructed for the first microphone and the second microphone are 0° and 180°, to obtain a first differential adaptive signal and a second differential adaptive signal; A recognition unit is used to perform speech recognition based on a multi-beam output signal, the first differential adaptive signal, and a second differential adaptive signal, including: obtaining a first probability from the multi-beam output signal, the first differential adaptive signal, and the second differential adaptive signal, respectively, and determining a speech recognition result based on the first probability, wherein the multi-beam output signal is obtained from a sound signal collected by the first microphone and the second microphone, and the first probability is a probability of recognizing a specific word.
2. The speaker according to claim 1, further comprising: The processor is used to perform corresponding actions according to the speech recognition results.
3. The speaker according to claim 2, wherein: The corresponding action includes starting the speaker to work.
4. The speaker according to claim 2, wherein The corresponding actions include activating different functions of the speaker according to different voice recognition results.
5. The speaker according to claim 1, further comprising: The beamformer is configured to obtain the multi-path beam output signal by performing beamforming on the sound signals collected by the first microphone and the second microphone.
6. The speaker according to claim 5, wherein: As long as a first probability obtained from any one of the multi-path beam output signal, the first differential adaptive signal, and the second differential adaptive signal is greater than a predetermined probability threshold, the recognition unit determines that the specific word is recognized.
7. The speaker according to claim 1, wherein The recognition unit further obtains a second probability from the multi-path beam output signal, the first differential adaptive signal and the second differential adaptive signal respectively, and determines a speech recognition result based on the second probability, wherein the second probability is the probability of recognizing a human voice.
8. The speaker according to claim 7, wherein: As long as the second probability obtained from any one of the multi-path beam output signal, the first differential adaptive signal, and the second differential adaptive signal is greater than a predetermined probability threshold, the recognition unit determines that a human voice is recognized.
9. The speaker according to claim 1, wherein The differential adaptive processing circuit includes: a first delay device, configured to delay the sound signal collected by the second microphone by a maximum time delay difference, where the maximum time delay difference is obtained by dividing the distance between the first microphone and the second microphone by the speed of sound; a second delay device, configured to delay the sound signal collected by the first microphone by a maximum time delay difference; a first subtractor, configured to subtract the output signal of the first delay device from the sound signal collected by the first microphone to obtain a first intermediate signal; a second subtractor, configured to subtract the sound signal collected by the second microphone from the output signal of the second delay device to obtain a second intermediate signal; a gain multiplier, configured to multiply the second intermediate signal by an adaptive gain; The third subtractor is configured to subtract the output result of the gain multiplier from the first intermediate signal to obtain a first differential adaptive signal.
10. The speaker according to claim 9, wherein: The differential adaptive processing circuit operates in cycles, and the adaptive gain is equal to the adaptive gain of the previous cycle of the current cycle plus the product of the second intermediate signal of the previous cycle, the first differential adaptive signal and a predetermined constant.
11. The speaker according to claim 1, wherein The differential adaptive processing circuit includes: a first delay device, configured to delay the sound signal collected by the first microphone by a maximum time delay difference, where the maximum time delay difference is obtained by dividing the distance between the first microphone and the second microphone by the speed of sound; a second delayer, configured to delay the sound signal collected by the second microphone by a maximum time delay difference; a first subtractor, configured to subtract the output signal of the first delay device from the sound signal collected by the second microphone to obtain a first intermediate signal; a second subtractor, configured to subtract the sound signal collected by the first microphone from the output signal of the second delay device to obtain a second intermediate signal; a gain multiplier, configured to multiply the second intermediate signal by an adaptive gain; The third subtractor is configured to subtract the output result of the gain multiplier from the first intermediate signal to obtain a second differential adaptive signal.
12. The speaker according to claim 11, wherein The differential adaptive processing circuit operates in cycles, and the adaptive gain is equal to the adaptive gain of the previous cycle of the current cycle, plus the product of the second intermediate signal of the previous cycle, the second differential adaptive signal and a predetermined constant.
13. A terminal device comprising: First microphone; Second microphone; a differential adaptive processing circuit, configured to perform differential adaptive processing on the sound signals collected by the first microphone and the second microphone when the incident angles of the antenna waveforms constructed for the first microphone and the second microphone are 0° and 180°, to obtain a first differential adaptive signal and a second differential adaptive signal; A recognition unit is used to perform speech recognition based on a multi-beam output signal, the first differential adaptive signal, and a second differential adaptive signal, including: obtaining a first probability from the multi-beam output signal, the first differential adaptive signal, and the second differential adaptive signal, respectively, and determining a speech recognition result based on the first probability, wherein the multi-beam output signal is obtained from a sound signal collected by the first microphone and the second microphone, and the first probability is a probability of recognizing a specific word.
14. The terminal device according to claim 13, further comprising: The processor is used to perform corresponding actions according to the speech recognition results.
15. The terminal device according to claim 14, wherein: The corresponding action includes responding to the recognized voice.
16. The terminal device according to claim 14, wherein: The corresponding actions include triggering switches for different operations according to different voice recognition results.
17. The terminal device according to claim 13, wherein: As long as a first probability obtained from any one of the multi-path beam output signal, the first differential adaptive signal, and the second differential adaptive signal is greater than a predetermined probability threshold, the recognition unit determines that the specific word is recognized.
18. The terminal device according to claim 13, wherein: The recognition unit obtains a second probability from the multi-path beam output signal, the first differential adaptive signal and the second differential adaptive signal respectively, and determines a speech recognition result based on the second probability, wherein the second probability is a probability of recognizing a human voice.
19. The terminal device according to claim 18, wherein: As long as the second probability obtained from any one of the multi-path beam output signal, the first differential adaptive signal, and the second differential adaptive signal is greater than a predetermined probability threshold, the recognition unit determines that a human voice is recognized.
20. A speaker, comprising: Multiple microphones; a differential adaptive processing circuit configured to perform differential adaptive processing on sound signals collected by any two microphones of the plurality of microphones when the incident angles of the antenna waveforms constructed for the two microphones are 0° and 180°, to obtain a first differential adaptive signal and a second differential adaptive signal; A recognition unit is used to perform speech recognition based on a multi-path beam output signal, the first differential adaptive signal and the second differential adaptive signal of any two microphones, comprising: obtaining a first probability from the multi-path beam output signal, the first differential adaptive signal and the second differential adaptive signal respectively, and determining a speech recognition result based on the first probability, wherein the multi-path beam output signal is obtained from the sound signal collected by the multiple microphones, and the first probability is the probability of recognizing a specific word.
21. The speaker according to claim 20, further comprising: The processor is used to perform corresponding actions according to the speech recognition results.
22. The speaker according to claim 21, wherein The corresponding action includes starting the speaker to work.
23. The speaker according to claim 21, wherein The corresponding actions include activating different functions of the speaker according to different voice recognition results.
24. The speaker of claim 20, further comprising: The beamformer is configured to obtain the multi-path beam output signals by performing beamforming on the sound signals collected by the multiple microphones.
25. The speaker of claim 20, wherein The recognition unit determines that the specific word is recognized as long as the first probability obtained from any one of the multi-path beam output signal, the first differential adaptive signal and the second differential adaptive signal of any two microphones is greater than a predetermined probability threshold.
26. The speaker of claim 20, wherein The recognition unit obtains a second probability from the multi-path beam output signal, the first differential adaptive signal and the second differential adaptive signal of any two microphones respectively, and determines a speech recognition result based on the second probability, wherein the second probability is the probability of recognizing a human voice.
27. The speaker of claim 26, wherein The recognition unit determines that a human voice is recognized as long as the second probability obtained from any one of the multi-path beam output signal, the first differential adaptive signal and the second differential adaptive signal of any two microphones is greater than a predetermined probability threshold.
28. A speaker plug-in for plugging into a speaker having a first microphone, a second microphone, and an identification unit, comprising: a differential adaptive processing circuit, configured to perform differential adaptive processing on the sound signals collected by the first microphone and the second microphone when the incident angles of the antenna waveforms constructed for the first microphone and the second microphone are 0° and 180°, to obtain a first differential adaptive signal and a second differential adaptive signal; an output interface, configured to output the first differential adaptive signal and the second differential adaptive signal to the recognition unit, so that the recognition unit performs speech recognition based on the multi-path beam output signal, the first differential adaptive signal and the second differential adaptive signal; The speech recognition based on the multi-path beam output signal, the first differential adaptive signal and the second differential adaptive signal includes: obtaining a first probability from the multi-path beam output signal, the first differential adaptive signal and the second differential adaptive signal respectively, and determining a speech recognition result based on the first probability, wherein the first probability is the probability of recognizing a specific word.
29. A system on a chip connected to outputs of a first microphone and a second microphone of a speaker, the system on a chip comprising: a differential adaptive processing circuit, configured to perform differential adaptive processing on the sound signals collected by the first microphone and the second microphone when the incident angles of the antenna waveforms constructed for the first microphone and the second microphone are 0° and 180°, to obtain a first differential adaptive signal and a second differential adaptive signal for use in a speech recognition process; The speech recognition process includes: obtaining first probabilities from the multi-path beam output signal, the first differential adaptive signal and the second differential adaptive signal respectively, and determining a speech recognition result based on the first probabilities, where the first probability is the probability of recognizing a specific word.
30. The system on chip according to claim 29, further comprising: The recognition unit is configured to perform speech recognition based on a multi-path beam output signal, the first differential adaptive signal, and the second differential adaptive signal, wherein the multi-path beam output signal is obtained from sound signals collected by the first microphone and the second microphone.
31. A speaker audio processing method, wherein: The speaker has a first microphone and a second microphone, and the speaker audio processing method includes: When the incident angles of the antenna waveforms constructed for the first microphone and the second microphone are 0° and 180°, performing differential adaptive processing on the sound signals collected by the first microphone and the second microphone to obtain a first differential adaptive signal and a second differential adaptive signal; Speech recognition is performed based on the multi-beam output signal, the first differential adaptive signal and the second differential adaptive signal, including: obtaining a first probability from the multi-beam output signal, the first differential adaptive signal and the second differential adaptive signal respectively, and determining a speech recognition result based on the first probability, wherein the multi-beam output signal is obtained from the sound signal collected by the first microphone and the second microphone, and the first probability is the probability of recognizing a specific word.
32. A terminal device audio processing method, wherein: The terminal device has a first microphone and a second microphone, and the terminal device audio processing method includes: When the incident angles of the antenna waveforms constructed for the first microphone and the second microphone are 0° and 180°, performing differential adaptive processing on the sound signals collected by the first microphone and the second microphone to obtain a first differential adaptive signal and a second differential adaptive signal; Speech recognition is performed based on the multi-beam output signal, the first differential adaptive signal and the second differential adaptive signal, including: obtaining a first probability from the multi-beam output signal, the first differential adaptive signal and the second differential adaptive signal respectively, and determining a speech recognition result based on the first probability, wherein the multi-beam output signal is obtained from the sound signal collected by the first microphone and the second microphone, and the first probability is the probability of recognizing a specific word.
33. A speaker audio processing method, wherein: The speaker has a plurality of microphones, and the speaker audio processing method includes: For a combination of any two microphones among the plurality of microphones, when the incident angles of the antenna waveforms constructed for the two microphones are 0° and 180°, performing differential adaptive processing on the sound signals collected by the two microphones to obtain a first differential adaptive signal and a second differential adaptive signal; Speech recognition is performed based on the multi-beam output signal, the first differential adaptive signal and the second differential adaptive signal of any two microphones, including: obtaining a first probability from the multi-beam output signal, the first differential adaptive signal and the second differential adaptive signal respectively, and determining a speech recognition result based on the first probability, wherein the multi-beam output signal is obtained from the sound signal collected by the multiple microphones, and the first probability is the probability of recognizing a specific word.
Citation Information
Patent Citations
Audio signal processing method, audio signal processing device, audio signal processing system, equipment and storage medium
CN110556103A
Endfire linear array microphone
US20190387311A1