Microphone with digital output determined at different power consumption levels

By using a threshold detector circuit and a combination of analog-to-digital converters with different power levels in the acoustic activation device to dynamically adjust power consumption, the problem of high power consumption in acoustic wake-up detection is solved, and efficient wake-up detection with low power consumption is achieved.

CN114175153BActive Publication Date: 2025-10-21QUALCOMM INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202080035410.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-03-14
Filing Date
2020-03-16
Publication Date
2025-10-21
Estimated Expiration
2040-03-16

AI Technical Summary

Technical Problem

Existing acoustic activation devices require a large amount of data capture during wake-up detection, resulting in high power consumption. In addition, existing technologies find it difficult to effectively reduce power consumption during infrequent acoustic signal intervals.

Method used

A threshold detector circuit is used, combined with a low-power successive approximation register analog-to-digital converter and a high-power Sigma-Delta analog-to-digital converter. By switching, the system dynamically adjusts between low power consumption and high signal-to-noise ratio to achieve acoustic wake-up detection.

Benefits of technology

Before the wake-up word is detected, power consumption is significantly reduced while ensuring the capture quality of necessary data, achieving low-power acoustic wake-up detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114175153B_ABST
    Figure CN114175153B_ABST
Patent Text Reader

Abstract

An acoustic device is described including an acoustic sensor element configured to sense acoustic energy and produce an output signal and a threshold detector circuit including a switch having an input coupled to an output of the acoustic sensor element to receive the output signal, a control port to receive a control signal, and first and second output ports, a first channel including an analog-to-digital converter operating at a first power level, a second analog-to-digital converter operating at a second, higher power level relative to the first power level, and a threshold level detector receiving output from the first analog-to-digital converter to produce the control signal having a first state if the first digitized output signal meets threshold criteria, the control signal causing the switch to feed the output signal from the acoustic sensor element to the second analog-to-digital converter.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Claim priority

[0002] This application claims priority under 35 U.S. Patent Code § 119(e) to U.S. Provisional Application No. 62 / 818,216, filed on March 14, 2019, the entire contents of which are incorporated herein by reference. Background Art

[0003] The present disclosure relates generally to acoustic sensing, and more particularly to the use of sensors such as microphones in voice-activated devices such as smart speakers and other types of acoustically activated devices.

[0004] With the growth of the Internet of Things (IoT) and the increasing use of acoustically activated devices, one of the challenges facing these devices is reducing power consumption. Typically, acoustically activated devices sense acoustic signals (sound, vibration, etc.) that may occur at infrequent intervals. One approach to addressing power consumption in acoustically activated devices is acoustic wake-up detection.

[0005] In the case of acoustic wake-up detection, an acoustic detector circuit is included in the acoustic activation device and remains in an active state that consumes power, while the wake-up circuit and / or the rest of the acoustic activation device is in an off or dormant state. When the acoustic detector circuit detects an event, the acoustic detector circuit generates a signal that causes power to be switched to the wake-up detection circuit and / or the acoustic activation device. The acoustic detector circuit can also be an algorithm executed by a processor. Summary of the Invention

[0006] Some methods of acoustic wakeup detection may require a large amount of data (e.g., approximately 500 milliseconds of data using current technology) before detecting the wakeup word utterance. If a threshold-based wakeup system or a voice detection-based wakeup system is used to turn on the analog-to-digital converter (ADC), digital signal processor (DSP), or other components of the acoustically activated device, the system may not be able to provide the necessary amount of data (e.g., 500 milliseconds of data) when the wakeup word causes the system to wake up.

[0007] The need for this data hinders the use of many power saving techniques because an ADC and audio buffer are required to capture this data. However, this data may not need to be of high quality relative to the rest of the utterance. By continuously buffering data to provide the necessary amount of data before the wake word utterance (e.g., 500 milliseconds of data or other time units called for by the specific application), significant power savings can be achieved by switching from a low-power, low-quality ADC to a higher-quality, higher-power ADC using threshold-based wake-up or voice detection-based wake-up.

[0008] According to one aspect, a threshold detector circuit is configured to receive a signal from an acoustic sensor element and generate an output signal to wake up an acoustic control device, the threshold detector circuit comprising: a switch having an input coupled to the output of the acoustic sensor element to receive the output signal from the acoustic sensor element, a control port for receiving a control signal, and a first output port and a second output port; a first analog-to-digital converter having an input coupled to the first output port of the switch and having an output for converting the output signal from the acoustic sensor element into a first digitized output signal, the first analog-to-digital converter operating at a first power level; a second analog-to-digital converter having an input coupled to the second output port of the switch and having an output for converting the output signal from the acoustic sensor element into a second digitized output signal, the second analog-to-digital converter operating at a second power level higher than the first power level; and a threshold level detector receiving the output from the first analog-to-digital converter to generate the control signal having a first state if the first digitized output signal meets a threshold criterion, the control signal causing the switch to feed the output signal from the acoustic sensor element to the second analog-to-digital converter.

[0009] Some embodiments may include one of the following features or a combination of two or more of the following features.

[0010] A conversion circuit is coupled between the output of the first analog-to-digital converter and the output of the second analog-to-digital converter to format the first digitized output signal into an audio signal format; and a buffer is coupled to the output of the first analog-to-digital converter and the output of the second analog-to-digital converter and is configured to store the first digitized output signal or the second digitized output signal in accordance with the control signal. The threshold level detector receives the output from the second analog-to-digital converter. When the second digitized output signal falls below the threshold level, the threshold level detector generates the control signal having a second state, which causes the switch to feed the output signal from the acoustic sensor element to the first analog-to-digital converter. The threshold detector circuit is configured to provide the output signal from the first analog-to-digital converter or the second analog-to-digital converter to the acoustic control device. The acoustic control device is a sensor device. The acoustic sensor element is a piezoelectric MEMS microphone. The buffer stores data in time units. The first analog-to-digital converter is a successive approximation register type analog-to-digital converter, and the second analog-to-digital converter is a sigma-delta type analog-to-digital converter. The microphone is a MEMS microphone and the threshold detector is a voice activity detector configured to detect when an input audio signal has an amplitude greater than a threshold amplitude. The microphone is a MEMS piezoelectric microphone and the threshold detector is a voice activity detector configured to detect when an input audio signal has an amplitude greater than a threshold amplitude.

[0011] According to another aspect, a threshold detector circuit is configured to receive an input signal from an acoustic sensor element and generate an output signal to wake up an acoustic control device, and includes: a switch having an input coupled to the output of the acoustic sensor element to receive the output signal from the acoustic sensor element, a control port for receiving a control signal, and a first output port and a second output port; a first channel including a per-band energy level detector circuit, the per-band energy level detector circuit dividing the output signal from the acoustic sensor element into frequency bands and buffering the per-band energy level; a second channel including an analog-to-digital converter, the analog-to-digital converter having an input coupled to the second output port of the switch and having an output for converting the output signal from the acoustic sensor element into a second digitized output signal, and the analog-to-digital converter operates at a second power level higher than the first power level; a threshold level detector receiving the output from the first channel to generate the control signal having a first state when the first digitized output signal meets the threshold standard, the control signal causing the switch to feed the output signal from the acoustic sensor element to the second analog-to-digital converter.

[0012] Some embodiments may include one of the following features or a combination of two or more of the following features.

[0013] The energy level per frequency band is calculated temporally in units of frames. The threshold detector circuit includes one or more buffer circuits. The first channel provides a precursor for calculating Mel-frequency cepstral coefficients. The threshold detector circuit also includes a wake-up sound signal detection circuit. The threshold detector circuit also includes a filter bank having a plurality of frequency bands whose sizes are determined using Mel-frequency levels.

[0014] According to another aspect, an acoustic device comprises: an acoustic sensor element configured to sense acoustic energy and generate an output signal; and a threshold detector circuit configured to receive an input signal from the acoustic device and generate an output signal to wake up an acoustic control device, the threshold detector circuit comprising: a switch having an input coupled to an output of the acoustic device to receive the output signal from the acoustic device, a control port for receiving a control signal, and a first output port and a second output port; a first analog-to-digital converter having an input coupled to the first output port of the switch and having an output for converting the output signal from the acoustic sensor element into a first digitized output signal. output, and the first analog-to-digital converter operates at a first power level; a second analog-to-digital converter having an input coupled to the second output port of the switch and having an output for converting the output signal from the acoustic sensor element into a second digitized output signal, and the second analog-to-digital converter operates at a second power level higher than the first power level; and a threshold level detector for receiving the output from the first analog-to-digital converter to generate the control signal having a first state if the first digitized output signal meets a threshold criterion, the control signal causing the switch to feed the output signal from the acoustic sensor element to the second analog-to-digital converter.

[0015] Some embodiments may include one of the following features or a combination of two or more of the following features.

[0016] a conversion circuit coupled between the output of the first analog-to-digital converter and the output of the second analog-to-digital converter to format the first digitized output signal into an audio signal format; and a buffer coupled to the output of the first analog-to-digital converter and the output of the second analog-to-digital converter and configured to store the first digitized output signal or the second digitized output signal according to the control signal.

[0017] The threshold level detector receives the output from the second analog-to-digital converter. When the second digitized output signal falls below the threshold level, the threshold level detector generates the control signal having a second state, which causes the switch to feed the output signal from the acoustic sensor element to the first analog-to-digital converter. The threshold detector circuit is configured to provide the output signal from the first analog-to-digital converter or the second analog-to-digital converter to an acoustically actuated device. The acoustically actuated device is a sensor device. The acoustically actuated device is a piezoelectric MEMS microphone. The buffer stores data per time unit. The first analog-to-digital converter is a successive approximation register (SACR)-type SAR-type ADC, and the second SAR-type Sigma-Delta (Sigma-Delta)-type ADC. The microphone is a MEMS microphone, and the threshold detector is a detector for determining whether a signal contains information of interest. The detector may be a threshold detector, a voice activity detector, an acoustic energy detector, or the like. The acoustic sensor element is a MEMS piezoelectric microphone, and the threshold level detector is implemented as a voice activity detector, which is configured to detect when the input audio signal has an amplitude greater than a threshold amplitude and the information of interest, and the microphone is packaged together with the voice activity detector in a hybrid circuit structure.

[0018] According to another aspect, an acoustic device includes: an acoustic sensor element configured to sense acoustic energy and generate an output signal; and a threshold detector circuit configured to receive an input signal from the acoustic device and generate an output signal to wake up an acoustic control device, and the threshold detector circuit includes: a switch having an input coupled to the output of the acoustic sensor element to receive the output signal, a control port for receiving a control signal, and a first output port and a second output port; a first channel including a per-band energy level detector circuit, the per-band energy level detector circuit dividing the output signal into frequency bands and buffering the per-band energy level; a second channel including an analog-to-digital converter having an input coupled to the second output port of the switch and having an output for converting the output signal from the acoustic sensor element into a second digitized output signal, and the analog-to-digital converter operates at a second power level higher than the first power level; a threshold level detector receiving the output from the first channel to generate the control signal having a first state when the first digitized output signal meets a threshold criterion, the control signal causing the switch to feed the output signal from the acoustic sensor element to the second analog-to-digital converter.

[0019] Some embodiments may include one of the following features or a combination of two or more of the following features.

[0020] The per-band energy level is calculated in units of frames over time. The threshold detector circuit includes one or more buffer circuits. The first channel provides a precursor for calculating Mel-frequency cepstral coefficients. The threshold detector circuit also includes a wake-up sound signal detection circuit. The threshold detector circuit also includes a filter bank having a plurality of frequency bands whose sizes are determined using Mel-frequency levels. The acoustic sensor element is a MEMS piezoelectric microphone, and the threshold level detector is a voice activity detector, which is configured to detect when an input audio signal has an amplitude greater than a threshold amplitude, and the microphone is packaged together with the voice activity detector in a hybrid circuit structure.

[0021] The details of one or more embodiments of the present disclosure are set forth in the accompanying drawings and the description below. Other features, objects, and advantages of the present disclosure will become apparent from the description and drawings, and from the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 is a block diagram of an exemplary networking system.

[0023] Figure 2 is a block diagram of an exemplary smart speaker.

[0024] Figure 3-7 is a block diagram of an exemplary detector circuit.

[0025] Figure 7A is a schematic diagram of a single microphone and the equivalent circuit.

[0026] Figure 8 is a block diagram of an exemplary processing circuit. DETAILED DESCRIPTION

[0027] Piezoelectric devices have the inherent ability to be actuated by stimulation even in the absence of a bias voltage, due to the so-called "piezoelectric effect," which causes piezoelectric materials to separate charges and provide a voltage potential difference between a pair of electrodes sandwiching the piezoelectric material. This physical property enables piezoelectric devices to provide ultra-low-power detection of a wide range of stimulation signals.

[0028] Microelectromechanical systems (MEMS) can include both piezoelectric and capacitive devices. Microphones manufactured as capacitive devices require a charge pump to provide the polarization voltage, whereas piezoelectric devices do not. The charge generated by the piezoelectric effect is generated by mechanical stress in the material caused by a stimulus. Therefore, ultra-low-power circuits can be used to deliver the generated charge through simple gain circuits.

[0029] Now refer to Figure 1, shows an exemplary distributed network architecture 10 for interconnecting IoT devices 20 having embedded processors and being acoustically activated. The distributed network architecture 10 embodies principles associated with the so-called "Internet of Things" (IoT), a term referring to the interconnection of uniquely identifiable devices 20, which may be sensors, detectors, appliances, process controllers, smart speakers, etc. Figure 1 In the context of , the device 20 is a voice detection based system that wakes up when acoustic energy is detected. These devices include threshold based wake up circuits or voice detection based wake up circuits.

[0030] The distributed network architecture 10 includes a gateway 16 located in a central, convenient location, such as within a separate building or structure. These gateways 16 communicate with the servers 14, whether they are stand-alone, dedicated servers or cloud-based servers running cloud applications using web programming techniques. Typically, the servers 14 also communicate with a database 17. The servers are networked together using established network technologies, such as the Internet Protocol, or a dedicated network that may not use the Internet or may use a portion of the Internet. The details of the distributed network 10 and the communications with these devices 20 are well known.

[0031] Now refer to Figure 2 , an exemplary IoT device 20a is shown. The IoT device 20a is a so-called smart speaker (hereinafter referred to as smart speaker 20a), and includes a microphone 22, an acoustic threshold detector circuit 24, and a circuit for receiving a signal (S) from the acoustic threshold detector circuit 24. OUT ) wake-up circuit 26. The smart speaker 20a also includes smart speaker electronic circuitry 28, which is part of the entire smart speaker 20 and includes various circuits not explicitly shown, such as: circuitry to wake up the smart speaker 20a in response to a name, computing circuitry for voice interaction, music playback, setting alarms, streaming podcasts, and playing audiobooks, in addition to providing weather, traffic, and other real-time information from the Internet. In some implementations, the smart speaker 20a can control other devices, thereby acting as a home automation hub. The smart speaker 20a has circuitry to connect to the Internet (wired and / or wirelessly) and has short-range communications, such as Bluetooth, to connect to other similarly enabled devices.

[0032] Now refer to Figure 3 , showing in detail the microphone 22 and the detector circuit 24. In one embodiment, the microphone 22 is a piezoelectric-based microphone. More specifically, the microphone 22 is a MEMS (micro-electromechanical system) piezoelectric microphone fabricated on a die. The piezoelectric-based MEMS microphone 22 is Figure 2, an equivalent circuit is represented by a capacitor in series with a voltage source, shunted by a resistor. The voltage source represents the equivalent voltage generated by the piezoelectric element in response to acoustic energy. The capacitor and resistor represent the equivalent capacitance and equivalent resistance of piezoelectric-based MEMS microphone 24. In some embodiments, piezoelectric-based MEMS microphone 24 is coupled to detector circuit 24, while in other embodiments, piezoelectric-based MEMS microphone 24 and detection circuit 24 are hybrid-integrated.

[0033] The threshold detection circuit 24 includes a switch 32 (at Figure 3 The input of switch 32 is coupled to the output of piezoelectric-based MEMS microphone 24. Switch 32 also has a first output coupled to first channel 34 and a second output coupled to second channel 36. Switch 32 also has a control port that is fed with a control signal to control whether switch 32 couples the input to the first output or the second output of switch 32.

[0034] The first channel 34 includes a first analog front end 34a, a successive approximation register (SAR) analog-to-digital converter (ADC) 34b, and a digital voltage level detector 34c. The first analog front end 34a has an output coupled to the input of the SAR ADC 34b. The output of the SAR ADC 34b is coupled to the input of the digital voltage level detector 34c. Each of the first analog front end 34a, the SAR ADC 34b, and the digital voltage level detector 34c is an ultra-low power device. The SAR ADC 34b is an ADC that converts a continuous input analog signal into a digital representation, using a binary search across all quantization levels to converge on a digital output during each conversion. This method introduces quantization error and quantization noise. However, the SAR ADC generally consumes much less power than other more accurate ADCs, such as Sigma-Delta ADCs.

[0035] In some embodiments, the digital voltage level detector 34c is an amplitude detector. That is, the digital voltage level detector 34c can measure when the amplitude of the digital data from the SAR ADC 34b reaches or exceeds a threshold value within one or more time increments. In other embodiments, the threshold detector is a voice activity detector, which is configured to detect when the input audio signal has a frequency that includes a frequency band threshold frequency corresponding to speech (e.g., 20HZ to 20000Hz). See, for example, U.S. patent application 62 / 818,140, ​​entitled “A Piezoelectric MEMS Device with an AdaptiveThreshold for Detection of an Acoustic Stimulus,” filed on March 14, 2019, the entire contents of which are incorporated herein by reference. That is, the digital voltage level detector 34c can be a threshold detector, a voice activity detector (VAD), or one of many other types of detectors for determining whether a signal of interest is present. For example, a VAD algorithm can determine the ratio of signal energy to zero crossing (signal offset between positive and negative levels) within a time interval. High energy levels with few zero crossings indicate that the signal is more likely speech, while low energy levels and / or high energy levels with many zero crossings indicate that the signal is more likely noise.For those systems that perform functions other than detecting speech as a signal of interest, other types of detection schemes may be used.

[0036] The second channel 36 includes a second analog front end 36a and a Sigma-Delta ADC (SD ADC) 36b. SD ADC 36b includes a Sigma-Delta modulator 37a and a digital filter, also commonly referred to as a decimation circuit 37b. The output of decimation circuit 37b (e.g., the output of ADC SD ADC 36b) is coupled to the input of digital voltage level detector 34c. The components in the second channel 36 (particularly SD ADC 36b and possibly the second analog front end 36a) will typically consume higher levels of power than the components in the first channel 34. The use of a conventional SD ADC 36b in the second channel 36 allows the input analog signal received from the analog front end 36a to undergo delta modulation, in which changes in the signal (e.g., deltas) are encoded rather than the absolute value of the signal, resulting in a pulse stream that passes through a 1-bit DAC and is added (sigma) to the input signal prior to delta modulation.

[0037] SD ADC 36b has significantly reduced quantization error, such as quantization error noise, which is common with simpler and lower-power types of ADCs, such as SAR ADC 34b. Thus, channel 34, despite being at a lower power level, will have higher quantization error and, therefore, higher quantization noise than channel 36.

[0038] Both channel 34 and channel 36 have signal outputs that are fed to conversion circuitry 40, which converts the digital signal received from channel 34 or channel 36 into a common digital audio format, depending on the state of a control signal from VAD 34c, as applied to switch 32. Conversion circuitry 40 has an output that is fed to buffer 42. Buffer 42 stores a time unit amount of digitized acoustic signal values ​​(S) captured by microphone 22. OUT In a quiet environment, the output signal from microphone 22 is coupled to channel 34 (a low-power channel relative to channel 36). The output signal is processed by channel 34, and the digitized, converted output signal from channel 34 is stored or buffered for a time unit amount of data, such as 500 milliseconds.

[0039] The digitized output signal from the SAR ADC 34 b is fed to the digital voltage level detector 34 c, and when the digital voltage level detector 34 c determines that there is speech or high ambient sound, the digital voltage level detector 34 c changes the state of the control signal to cause the switch 32 to switch to the channel 36 and the SD ADC 36 b, thereby providing better quality audio compared to the channel 34 and the SAR ADC 34 b.

[0040] On the other hand, the digitized output signal from the SD ADC 36 b is also fed to the digital voltage level detector 34 c, and when the digital voltage level detector 34 c determines that there is no longer speech or high ambient sound, the digital voltage level detector 34 c again changes the state of the control signal to switch the switch 32 to the channel 34 and the SAR ADC 34 b, thereby providing audio with lower quality but lower power consumption compared to the channel 36 and the SD ADC 36 b.

[0041] References to low power and relatively high power do not require or imply that a high-power SD ADC 36b should be used. Rather, it should be understood that, for a given set of requirements for a particular application, all components will utilize the lowest possible power consumption, taking into account performance and cost criteria. However, it is apparent that, given the properties of a typical SAR ADC 34b and due to the operating principles and complexity of a typical SD ADC 36b, a typical SD ADC 36b will generally consume more power for a given resolution than a typical SAR ADC 34b. Therefore, all components may be low-power components.

[0042] Now refer to Figure 4 , an alternative embodiment of a detection circuit is shown. A microphone 22, such as a piezoelectric based microphone, and an alternative detector circuit 44 are shown in detail.

[0043] The threshold detection circuit 24 includes an attenuation switch 41 for attenuating the output signal from the microphone 22 (eg, by a fixed decibel amount) and a switch 32 (eg, as shown in FIG. 1 ). Figure 3 However, the switch 32 is inserted between the attenuation switch 41 and the alternative first channel 34' and the second channel 36. In addition, the switch 32 and Figure 3 The control port operates similarly to that described in , where it is fed with a control signal from the digital voltage level detector 34c.

[0044] The first channel 34' includes a first analog front end 34a (e.g., Figure 3 As shown), threshold circuit 44a, SAR ADC34b (for example, as Figure 3 ) and a digital voltage level detector 34c (e.g., as Figure 3 ). The first channel 34' includes a threshold circuit 44a that can be used to "gate" the SAR ADC 34b to operate when the output signal from the front end 34a exceeds a threshold.

[0045] The second channel 36 includes a second analog front end 36a and an SD ADC 36b, which includes a Sigma-Delta modulator 37a and a digital filter also commonly referred to as a decimation circuit 37b. Figure 3 As shown, the output of the decimation circuit 37b (eg, the output of the ADC SD ADC 36b) is coupled to the input of the digital voltage level detector 34c. Figure 3 As shown, the components in the second channel 36 (particularly the SD ADC 36 b and possibly the second analog front end 36 a ) will typically consume higher levels of power than the components in the first channel 34 .

[0046] Both channel 34′ and channel 36 have signal outputs that are fed to a conversion circuit 40, which converts the digital signal received from channel 34′ or channel 36 into a common digital audio format depending on the state of the control signal from VAD 34 c applied to switch 32. As described above, the conversion circuit 40 has an output that is fed to a buffer 42, which buffers a time unit of data (S), such as a 500 millisecond amount of data. OUT ).

[0047] The digitized output signal from the SAR ADC 34b is fed to a digital voltage level detector 34c. Figure 3 As discussed, when the digital voltage level detector 34c determines that there is speech or high ambient sound, the digital voltage level detector 34c changes the state of the control signal to switch the switch 32 to the channel 36 and the SD ADC 36b, thereby providing better quality audio than the channel 34' and the SAR ADC 34b. Figure 3 As discussed, when the digital voltage level detector 34 c determines that speech or high ambient sound is no longer present, the digital voltage level detector 34 c again changes the state of the control signal to cause the switch 32 to switch to the channel 34′ and the SAR ADC 34 b, thereby providing audio with lower power consumption, albeit lower audio quality, compared to the channel 36 and the SD ADC 36 b.

[0048] Now refer to Figure 5 , shows another alternative embodiment of the detection circuit. In this embodiment, there is a pair of microphones 22a, 22b arranged in a differential configuration 23, wherein the differential configuration 23 has a reference line coupled to a reference potential and output lines each coupled to a switch arrangement 32' as a double pole double throw configuration.

[0049] The switch arrangement 32' has a pair of inputs that receive output signals from a pair of microphones 22a, 22b. The switch arrangement 32' also has two pairs of outputs coupled to an alternative first analog front end 34a' and an alternative second analog front end 36a', each of the two pairs of outputs having a differential input. The switch arrangement 32' determines whether the signal from the switch arrangement 32' is fed to the alternative first analog front end 34a' or to the alternative second analog front end 36a'. The SD ADC 36b' may have differential inputs, and the SDADC 36b' may include a digital filter 48 inserted between the Sigma-Delta modulator 37a and the decimation circuit 37b to attenuate the output from the SD ADC 36b', the digital filter 48 being used for outputs that are above a bandwidth of interest depending on the application of the circuit. As described above, the detection circuit includes a conversion circuit 40 having an output that is fed to a buffer 42, which buffers an amount of data (S) in a time unit, such as an amount of data of 500 milliseconds. OUT ).

[0050] Now refer to Figure 6 , shows another alternative embodiment of the detection circuit. In this embodiment, as Figure 5 As in the example, there is a differential structure 23 ( Figure 5 ) arrangement of a pair of microphones 22a, 22b coupled to an alternative first analog front end 34a' and an alternative second analog front end 36a'.

[0051] Figure 6 A third channel 49 is included that can accommodate an analog wake-up sound circuit. An example is of the type described in the following patent applications: U.S. patent application 16 / 081,015, filed on August 29, 2018, entitled "A Piezoelectric Mems Device for Producing a Signal Indicative of Detection of an Acoustic Stimulus," and co-pending U.S. patent application 62 / 818,140, ​​filed on March 14, 2019, entitled "A Piezoelectric MEMS Device with an Adaptive Threshold for Detection of an Acoustic Stimulus," both of which are incorporated herein by reference in their entireties, and as described therein, each of the two patent applications provides an output signal (D OUT ).

[0052] Now refer to Figure 7 , shows another alternative embodiment of the detection circuit. In this embodiment, there is a signal SOUT Channel 36 (see Figure 4 ) and provide signal S OUTa to S OUTn Another alternative channel 34". Channel 34" includes an alternative first analog front end 34a' (e.g., Figure 5 ), filter bank 52, per-band energy level detector circuit 54, and wake-up sound signal detection circuit 56. Instead of using a SAR ADC (e.g., as Figure 5 and 6 ) to digitize the output signals from the microphones 22a, 22b, the channel 34" includes a filter bank 52 that divides the output signals into frequency bands, and a per-band energy level detector circuit 54 calculates the per-band energy level in units of frames or time windows (e.g., every 20 milliseconds). These values ​​are fed to a per-band analog-to-digital converter, and the output from the analog-to-digital converter is stored in a per-band buffer to buffer the per-band energy level signal. The per-band energy level signal is calculated in units of frames in time, such as every 20 milliseconds.

[0053] By storing (buffering) in this format, channel 34" provides a precursor for calculating Mel Frequency Cepstral Coefficients (MFCCs) to compress the audio signal. In a typical digital system, MFCCs will be calculated every 20 millisecond interval to essentially provide an average of the squared voltage over that interval. Figure 7 In a system with a squaring operation, the system may first calculate the average over the time interval (to use the instantaneous information of the wake-up algorithm), or use the order of operations typical of digital systems. The wake-up sound signal detection circuit operates on a frequency band using, for example, the detection scheme described in the aforementioned provisional application.

[0054] Converting to MFCCs can also compress audio signals. This conversion is done by first framing the signal into short frames (e.g., 25 milliseconds), applying a discrete Fourier transform (DFT) to the framed signal to transform the signal into separate frequency bands corresponding to so-called Mel frequency levels, and calculating the natural logarithm (log) of the signal energy in each Mel frequency band. This conversion also involves calculating the discrete cosine transform (DCT) of the new signal (energy levels in a range of frequency bands), and in some instances removing the higher coefficients and retaining the remaining coefficients as MFCCs.

[0055] Therefore, the Mel frequency level can be used to determine Figure 7 The filter bank size in , and in this case, Figure 7A set of operations that is functionally equivalent to the previous steps of the MFCC conversion will be described. If the output is converted to logarithmic level, it may be more efficient to store in the buffer (fewer bits can be used to store a better signal representation). Because MFCCs are commonly used in speech recognition systems, the stored MFCCs can be transmitted and used by the rest of the system instead of the original audio signal.

[0056] Now refer to Figure 7A , showing a single microphone with a differential output set as Figure 7 A replacement for the microphone.

[0057] As described above, the switch 32 (or 32 ′) receives a control signal from the digital voltage level detector 34 c , which switches the state of the control signal according to the outputs from the respective channels 34 , 36 .

[0058] Alternatively, the switch 32 (or 32') may be switched from a processing device (e.g. Figure 8 The processing device 80 shown receives a control signal. The processing device begins with the control signal in a first state, which causes the output from the microphone 22 to be fed to the first channel 34 having the SAR ADC 34b. The processing device analyzes the SAR ADC signal at the output and determines when the output signal reaches or exceeds a signal level having an amplitude of interest. If a signal having an amplitude of interest is present, the processing device changes the state of the control signal to a second state, which causes the output from the microphone 22 to be fed to the second channel 36 having the SD ADC 36b.

[0059] After a period of time has elapsed in which the processing means has not detected any signal of interest, the process again changes the state of the control signal back to the first state so that the output from the microphone 22 is fed to the first channel 34 having the SAR ADC 34b.

[0060] As another alternative, the switch 32 (or 32') receives data from the processing device and Figure 6 The control signal is determined by the wake-up sound circuit.

[0061] As another alternative, Figure 7 In the circuit, instead of buffering the actual audio signal, the signal is filtered into frequency bands, and these frequency bands can be used to wake up the sound signal detection circuit, and these frequency bands can also be used to generate data to be buffered.

[0062] Now refer to Figure 8, shows an example of an embedded processing device 80 that can be used to process the digitized output from the buffer 42. The processing device 80 includes a processor / controller 82, which can be an embedded processor, a central processing unit, or can be manufactured as an ASIC (application-specific integrated circuit), etc. The processing device 80 also includes a memory 84 (memory), a storage device 86, and an I / O (input / output) circuit 88, all of which are connected to the processor / controller 82 via a bus 89. The I / O circuit 88 receives the digitized output signal from the buffer 42, processes the signal, and generates, for example, the smart speaker 20a ( Figure 2 ) and other circuits in the IoT device 20 suitable for the wake-up signal.

[0063] In some implementations, the processing device 80 performs the function of the threshold detector 34b to detect when the acoustic input of, for example, a microphone, equals or exceeds a threshold level, for example, by detecting when the digitized output from the buffer equals or exceeds an amplitude level or is within a frequency band. Because the detection is performed by the processing device 80, rather than being included in the acoustic device as in, for example, a hybrid integrated microphone / detector, the processing device 80 needs to remain powered on to detect the audio stimulus.

[0064] A number of embodiments of the technology have been described. However, it will be appreciated that various modifications can be made without departing from the spirit and scope of the present disclosure. Accordingly, other embodiments are within the scope of the following claims.

Claims

1. A threshold detector circuit configured to receive a signal from an acoustic sensor element and generate an output signal to wake up an acoustic control device, the threshold detector circuit comprising: a switch having an input configured to be coupled to an output of the acoustic sensor element to receive an output signal from the acoustic sensor element, a control port for receiving a control signal, and a first output port and a second output port; a first analog front end having an input and an output, wherein the input of the first analog front end is coupled to the first output port of the switch; a second analog front end having an input and an output, wherein the input of the second analog front end is coupled to the second output port of the switch; a first analog-to-digital converter having an input coupled to the output of the first analog front end and having an output, the first analog-to-digital converter operating at a first power level; a second analog-to-digital converter having an input coupled to the output of the second analog front end and having an output, the second analog-to-digital converter operating at a second, higher power level relative to the first power level; and a digital voltage level detector having an output, a first input, and a second input, wherein the first input of the digital voltage level detector is coupled to the output of the first analog-to-digital converter, wherein the second input of the digital voltage level detector is coupled to the output of the second analog-to-digital converter, wherein the output of the digital voltage level detector is coupled to the control port of the switch, and wherein the digital voltage level detector is configured to generate the control signal at the control port for selecting the second output port of the switch based on a determination that a first digitized output signal from the first analog-to-digital converter satisfies a threshold criterion.

2. The threshold detector circuit of claim 1 , further comprising: a conversion circuit coupled to an output of the first analog-to-digital converter and an output of the second analog-to-digital converter to format the first digitized output signal into an audio signal format; as well as A buffer is coupled to an output of the first analog-to-digital converter and an output of the second analog-to-digital converter and is configured to store the first digitized output signal or the second digitized output signal according to the control signal.

3. The threshold detector circuit according to claim 1 , wherein: The digital voltage level detector is configured to receive an output from the second analog-to-digital converter.

4. The threshold detector circuit according to claim 1, wherein The digital voltage level detector is configured to generate the control signal having a second state if the second digitized output signal falls below the threshold level, the control signal causing the switch to feed the output signal from the acoustic sensor element to the first analog-to-digital converter.

5. The threshold detector circuit according to claim 1, wherein The acoustic sensor element is a piezoelectric-based MEMS microphone.

6. The threshold detector circuit according to claim 2, wherein: The buffer stores a time unit amount of data.

7. The threshold detector circuit according to claim 1, wherein The first analog-to-digital converter is a successive approximation register type analog-to-digital converter, and the second analog-to-digital converter is a Sigma-Delta type analog-to-digital converter.

8. An acoustic device comprising: an acoustic sensor element configured to sense acoustic energy and generate an output signal; as well as A threshold detector circuit comprising: a switch having an input configured to be coupled to an output of the acoustic sensor element to receive the output signal, a control port for receiving a control signal, and a first output port and a second output port; a first analog front end having an input and an output, wherein the input of the first analog front end is coupled to the first output port of the switch; a second analog front end having an input and an output, wherein the input of the second analog front end is coupled to the second output port of the switch; a first analog-to-digital converter having an input coupled to the output of the first analog front end and having an output, the first analog-to-digital converter operating at a first power level; a second analog-to-digital converter having an input coupled to the output of the second analog front end and having an output, the second analog-to-digital converter operating at a second, higher power level relative to the first power level; and a digital voltage level detector having an output, a first input, and a second input, wherein the first input of the digital voltage level detector is coupled to the output of the first analog-to-digital converter, wherein the second input of the digital voltage level detector is coupled to the output of the second analog-to-digital converter, wherein the output of the digital voltage level detector is coupled to the control port of the switch, and wherein the digital voltage level detector is configured to generate the control signal at the control port for selecting the second output port of the switch based on a determination that a first digitized output signal from the first analog-to-digital converter satisfies a threshold criterion.

9. The acoustic device according to claim 8, further comprising: a conversion circuit coupled to an output of the first analog-to-digital converter and an output of the second analog-to-digital converter to format the first digitized output signal into an audio signal format; as well as A buffer is coupled to an output of the first analog-to-digital converter and an output of the second analog-to-digital converter and is configured to store the first digitized output signal or the second digitized output signal according to the control signal.

10. The acoustic device according to claim 8, wherein The digital voltage level detector is configured to receive an output from the second analog-to-digital converter.

11. The acoustic device according to claim 8, wherein The digital voltage level detector is configured to generate the control signal having a second state if the second digitized output signal falls below the threshold level, the control signal causing the switch to feed the output signal from the acoustic sensor element to the first analog-to-digital converter.

12. The acoustic device according to claim 8, wherein The acoustic sensor element is a piezoelectric-based MEMS microphone.

13. The acoustic device according to claim 9, wherein The buffer stores a time unit amount of data.

14. The acoustic device according to claim 8, wherein The first analog-to-digital converter is a successive approximation register type analog-to-digital converter, and the second analog-to-digital converter is a Sigma-Delta type analog-to-digital converter.

15. The acoustic device according to claim 8, wherein The acoustic sensor element is a MEMS piezoelectric microphone, and the digital voltage level detector is implemented as a voice activity detector, which is configured to detect when an input audio signal has an amplitude greater than a threshold amplitude, and wherein the microphone is packaged together with the voice activity detector in a hybrid circuit structure.

Citation Information

Patent Citations

  • Piezoelectric mems device for producing a signal indicative of detection of an acoustic stimulus

    US10715922B2

  • Extraction and analysis of audio feature data

    US20130110521A1

  • Analog-to-digital converter (ADC) dynamic range enhancement for voice-activated systems

    US20160314805A1